Files
log2ml/3-3-security-and-performance-log-analysis-sw/Elasticsearch-Pandas-vs-Polars-May-15-2024.ipynb
T
2024-06-25 16:48:23 +02:00

237 KiB

Elasticsearch and tabular integration

Elasticsearch is a NoSQL database, which indexes JSON records. In the following the Winlog Beat index gets queried, which holds Windows EventLog data. The Elasticsearch SQL endpoint is used to define a query, and the resulting data is retrieved as a JSON stream. The data gets read into in-memory dataframe objects which allow data-manipulation tasks.

In-memory processing can be difficult if the datasets grow large. Therefore a comparison is made between two polular in-memory dataframe libraries:

1.) Pandas 2.) Polars

The memory footprint is assessed, because runtime memory is the limiting factor for the implementations.

Versions

In [1]:
!pip show pandas | grep -E 'Name:|Version:'
Name: pandas
Version: 2.1.4
In [2]:
!pip show polars | grep -E 'Name:|Version:'
Name: polars
Version: 0.20.26

Elasticsearch API

The Elasticsearch API uses HTTP and is available on port 9200.

The index "winlogbeat-" contains data from the period. It's a periodically rotating index.

Here the Elasticsearch DSL is used, and an event timeline is being retrieved, in time-descending order.

The resulting JSON data is piped to the jq utility, which is prettier on a command-line. Only the first JSON record is analyzed.

The output shows the index and the timestamp.

In [21]:
%%bash
curl -s -X GET "http://192.168.20.106:9200/winlogbeat-*/_search" -H 'Content-Type: application/json' -d '{
  "size": 1,
  "sort": [
    {
      "@timestamp": {
        "order": "desc"
      }
    }
  ]
}' | jq '.hits.hits[0] | {index: ._index, timestamp: ._source["@timestamp"]}'
{
  "index": "winlogbeat-7.10.0-2024.05.15-000008",
  "timestamp": "2024-05-15T15:57:22.877Z"
}

The following Bash command shows a SQL query.

The Limit 1 is a common SQL statement. The output is further limited with the head command. Only the first fields of the first record are shown.

By default the order of records doesn't represent a timeline, but the order of records in the index.

In [27]:
%%bash
curl -s -X POST "http://192.168.20.106:9200/_sql/translate" -H 'Content-Type: application/json' -d '{
  "query": "SELECT * FROM \"winlogbeat-7.10.0-2024.05.15-*\" LIMIT 1"
}' | jq | head -n 3
{
  "size": 1,
  "_source": {

Elasticsearch tabular-integration and Pandas

Pandas is the de-facto standard for data-manipulation of small to medium datasets in Data Science. It offers robust functions for in-memory data transactions and tabular feature integration.

In the following the expansion of JSON data is used to allow a simple feature selection for further processing. The data is returned from Elasticsearch, from an SQL query.

The data is provided via a Scrolling API, which delivers a portion of the data each time. This simplifies batch processing of large datasets.

In [64]:
import requests
import pandas as pd
import json

# Function to recursively normalize nested columns in a DataFrame
def recursively_normalize(data):
    df = pd.json_normalize(data)
    while True:
        nested_cols = [col for col in df.columns if isinstance(df[col].iloc[0], (dict, list))]
        if not nested_cols:
            break
        for col in nested_cols:
            if isinstance(df[col].iloc[0], dict):
                normalized = pd.json_normalize(df[col])
                df = df.drop(columns=[col]).join(normalized)
            elif isinstance(df[col].iloc[0], list):
                df = df.explode(col)
                normalized = pd.json_normalize(df[col])
                df = df.drop(columns=[col]).join(normalized)
    return df

# Function to fetch the next batch using the cursor
def fetch_next_batch(cursor):
    response = requests.post(
        f"{base_url}/_sql?format=json",
        headers={"Content-Type": "application/json"},
        json={"cursor": cursor}
    ).json()
    return response

# Elasticsearch base URL
base_url = "http://192.168.20.106:9200"
# Index name
index = "winlogbeat-*"

# SQL query for initial search
sql_query = """
SELECT "@timestamp", host.hostname, host.ip, log.level, winlog.event_id, winlog.task, message FROM "winlogbeat-7.10.0-2024.05.15-*"
LIMIT 5000
"""

# Initial search request to start scrolling
initial_response = requests.post(
    f"{base_url}/_sql?format=json",
    headers={"Content-Type": "application/json"},
    json={
        "query": sql_query,
        "field_multi_value_leniency": True
    }
).json()

# Extract the cursor for scrolling
cursor = initial_response.get('cursor')
rows = initial_response.get('rows')
columns = [col['name'] for col in initial_response['columns']]

# Initialize CSV file (assumes the first batch is not empty)
if rows:
    df = pd.DataFrame(rows, columns=columns)
    df = recursively_normalize(df.to_dict(orient='records'))
    df.to_csv("lab_logs_normal_activity.csv", mode='w', index=False, header=True)

# Track total documents retrieved
total_documents_retrieved = len(rows)
print(f"Retrieved {total_documents_retrieved} documents.")

# Loop to fetch subsequent batches of documents until no more documents are left
while cursor:
    # Fetch next batch of documents using cursor
    response = fetch_next_batch(cursor)
    
    # Update cursor for the next batch
    cursor = response.get('cursor')
    rows = response.get('rows')
    
    # If no rows, break out of the loop
    if not rows:
        break
    
    # Normalize data and append to CSV
    df = pd.DataFrame(rows, columns=columns)
    df = recursively_normalize(df.to_dict(orient='records'))
    
    # Append to CSV file without headers
    df.to_csv("lab_logs_normal_activity.csv", mode='a', index=False, header=False)
    
    # Convert DataFrame to JSON, line by line
    json_lines = df.to_json(orient='records', lines=True).splitlines()
    # Append each line to an existing JSON file
    with open("lab_logs_normal_activity.json", 'a') as file:
        for line in json_lines:
            file.write(line + '\n')  # Append each line and add a newline
        
    # Update total documents retrieved
    total_documents_retrieved += len(rows)
    
    print(f"Retrieved {total_documents_retrieved} documents.")

print("Files have been written.")
Retrieved 1000 documents.
Retrieved 2000 documents.
Retrieved 3000 documents.
Retrieved 4000 documents.
Retrieved 5000 documents.
Files have been written.

Alternative approach with polars

Polars is a newer tabular-integration library, which challenges Pandas. It's supposed to me more memory efficient, because it's backend is written in Rust.

In [ ]:
%pip install polars
In [63]:
import requests
import polars as pl
import json

# Function to recursively unnest nested columns in a DataFrame
def recursively_unnest(df):
    nested = True
    while nested:
        nested = False
        for col in df.columns:
            if df[col].dtype == pl.List:
                df = df.explode(col)
                nested = True
            elif df[col].dtype == pl.Struct:
                df = df.unnest(col)
                nested = True
    return df

# Function to fetch the next batch using the cursor
def fetch_next_batch(cursor):
    response = requests.post(
        f"{base_url}/_sql?format=json",
        headers={"Content-Type": "application/json"},
        json={"cursor": cursor}
    ).json()
    return response

# Elasticsearch base URL
base_url = "http://192.168.20.106:9200"
# Index name
index = "winlogbeat-*"

# SQL query for initial search
sql_query = """
SELECT "@timestamp", host.hostname, host.ip, log.level, winlog.event_id, winlog.task, message FROM "winlogbeat-7.10.0-2024.05.15-*"
LIMIT 5000
"""

# Initial search request to start scrolling
initial_response = requests.post(
    f"{base_url}/_sql?format=json",
    headers={"Content-Type": "application/json"},
    json={
        "query": sql_query,
        "field_multi_value_leniency": True
    }
).json()

# Extract the cursor for scrolling
cursor = initial_response.get('cursor')
rows = initial_response.get('rows')
columns = [col['name'] for col in initial_response['columns']]

# Initialize CSV file (assumes the first batch is not empty)
if rows:
    df = pl.DataFrame(rows, schema=columns)
    df = recursively_unnest(df)
    df.write_csv("lab_logs_normal_activity.csv", include_header=True)

# Track total documents retrieved
total_documents_retrieved = len(rows)
print(f"Retrieved {total_documents_retrieved} documents.")

# Loop to fetch subsequent batches of documents until no more documents are left
while cursor:
    # Fetch next batch of documents using cursor
    response = fetch_next_batch(cursor)
    
    # Update cursor for the next batch
    cursor = response.get('cursor')
    rows = response.get('rows')
    
    # If no rows, break out of the loop
    if not rows:
        break
    
    # Normalize data and append to CSV
    df = pl.DataFrame(rows, schema=columns)
    df = recursively_unnest(df)
    
    # Manually write the CSV to avoid headers
    with open("lab_logs_normal_activity.csv", 'a') as f:
        df.write_csv(f, include_header=False)
    
    # Convert DataFrame to JSON, line by line
    json_lines = [json.dumps(record) for record in df.to_dicts()]
    # Append each line to an existing JSON file
    with open("lab_logs_normal_activity.json", 'a') as file:
        for line in json_lines:
            file.write(line + '\n')  # Append each line and add a newline
        
    # Update total documents retrieved
    total_documents_retrieved += len(rows)
    
    print(f"Retrieved {total_documents_retrieved} documents.")

print("Files have been written.")
Retrieved 1000 documents.
Retrieved 2000 documents.
Retrieved 3000 documents.
Retrieved 4000 documents.
Retrieved 5000 documents.
Files have been written.

Memory footprint and profile comparison

A JSON schema is provided in both cases to improve the comparison.

In [27]:
%reset
Once deleted, variables cannot be recovered. Proceed (y/[n])?  y
In [ ]:
!pip install git+https://github.com/H4dr1en/jupyterflame.git
In [ ]:
# directly on the shell within the conda env: conda install -y perl
In [28]:
%load_ext jupyterflame
The jupyterflame extension is already loaded. To reload it, use:
  %reload_ext jupyterflame
In [29]:
import pandas as pd

# Read a small chunk of the JSON file
file_path = "lab_logs_normal_activity.json"
pd_df = pd.read_json(file_path, lines=True, nrows=10)

print(pd_df.dtypes)
@timestamp         object
host.hostname      object
host.ip            object
log.level          object
winlog.event_id     int64
winlog.task        object
message            object
dtype: object
In [46]:
import polars as pl

# Define the mapping from Pandas dtype to Polars dtype
dtype_mapping = {
    "object": pl.Utf8,
    "int64": pl.Int64,
    "float64": pl.Float64,
    # Add more mappings if needed
}

pandas_dtype_mapping = {
    "object": "str",
    "int64": "int64",
    "float64": "float64",
    # Add more mappings if needed
}


# Generate the schema for Polars from Pandas dtype
polars_schema = {col: dtype_mapping[str(dtype)] for col, dtype in pd_df.dtypes.items()}
print("Polars Schema:", polars_schema)

pandas_schema = {col: pandas_dtype_mapping[str(dtype)] for col, dtype in pd_df.dtypes.items()}
print("Pandas Schema:", pandas_schema)
Polars Schema: {'@timestamp': String, 'host.hostname': String, 'host.ip': String, 'log.level': String, 'winlog.event_id': Int64, 'winlog.task': String, 'message': String}
Pandas Schema: {'@timestamp': 'str', 'host.hostname': 'str', 'host.ip': 'str', 'log.level': 'str', 'winlog.event_id': 'int64', 'winlog.task': 'str', 'message': 'str'}
In [53]:
def test_polars():
    # Read the JSON file using the defined schema
    lazy_df = pl.scan_ndjson(file_path)

    # Collect the LazyFrame to a DataFrame
    pl_df = lazy_df.collect()

    # Convert columns to the correct data types according to the schema
    pl_df = pl_df.with_columns([pl.col(col).cast(dtype) for col, dtype in polars_schema.items()])

    # Print the DataFrame and its memory usage
    print(pl_df)

    num_rows_polars = pl_df.shape[0]

    print(f"Polars DataFarme number of rows:  {num_rows_polars}")
    print(f"Polars DataFrame memory usage: {pl_df.estimated_size() / (1024 ** 2):.2f} MB")
In [48]:
%%flame -q --inverted
test_polars()
Out [48]:
shape: (8_000, 7)
┌──────────────┬─────────────┬─────────────┬─────────────┬─────────────┬─────────────┬─────────────┐
│ @timestamp   ┆ host.hostna ┆ host.ip     ┆ log.level   ┆ winlog.even ┆ winlog.task ┆ message     │
│ ---          ┆ me          ┆ ---         ┆ ---         ┆ t_id        ┆ ---         ┆ ---         │
│ str          ┆ ---         ┆ str         ┆ str         ┆ ---         ┆ str         ┆ str         │
│              ┆ str         ┆             ┆             ┆ i64         ┆             ┆             │
╞══════════════╪═════════════╪═════════════╪═════════════╪═════════════╪═════════════╪═════════════╡
│ 2024-05-15T1 ┆ win10       ┆ fe80::24b4: ┆ information ┆ 13          ┆ Registry    ┆ Registry    │
│ 5:57:18.471Z ┆             ┆ 3691:44a6:3 ┆             ┆             ┆ value set   ┆ value set:  │
│              ┆             ┆ 8a1         ┆             ┆             ┆ (rule:      ┆ RuleName: … │
│              ┆             ┆             ┆             ┆             ┆ Regi…       ┆             │
│ 2024-05-15T1 ┆ win10       ┆ fe80::24b4: ┆ information ┆ 13          ┆ Registry    ┆ Registry    │
│ 5:57:18.471Z ┆             ┆ 3691:44a6:3 ┆             ┆             ┆ value set   ┆ value set:  │
│              ┆             ┆ 8a1         ┆             ┆             ┆ (rule:      ┆ RuleName: … │
│              ┆             ┆             ┆             ┆             ┆ Regi…       ┆             │
│ 2024-05-15T1 ┆ win10       ┆ fe80::24b4: ┆ information ┆ 13          ┆ Registry    ┆ Registry    │
│ 5:57:18.471Z ┆             ┆ 3691:44a6:3 ┆             ┆             ┆ value set   ┆ value set:  │
│              ┆             ┆ 8a1         ┆             ┆             ┆ (rule:      ┆ RuleName: … │
│              ┆             ┆             ┆             ┆             ┆ Regi…       ┆             │
│ 2024-05-15T1 ┆ win10       ┆ fe80::24b4: ┆ information ┆ 13          ┆ Registry    ┆ Registry    │
│ 5:57:18.471Z ┆             ┆ 3691:44a6:3 ┆             ┆             ┆ value set   ┆ value set:  │
│              ┆             ┆ 8a1         ┆             ┆             ┆ (rule:      ┆ RuleName: … │
│              ┆             ┆             ┆             ┆             ┆ Regi…       ┆             │
│ 2024-05-15T1 ┆ win10       ┆ fe80::24b4: ┆ information ┆ 13          ┆ Registry    ┆ Registry    │
│ 5:57:18.471Z ┆             ┆ 3691:44a6:3 ┆             ┆             ┆ value set   ┆ value set:  │
│              ┆             ┆ 8a1         ┆             ┆             ┆ (rule:      ┆ RuleName: … │
│              ┆             ┆             ┆             ┆             ┆ Regi…       ┆             │
│ …            ┆ …           ┆ …           ┆ …           ┆ …           ┆ …           ┆ …           │
│ 2024-05-15T1 ┆ win10       ┆ fe80::24b4: ┆ information ┆ 4663        ┆ Removable   ┆ An attempt  │
│ 6:10:07.128Z ┆             ┆ 3691:44a6:3 ┆             ┆             ┆ Storage     ┆ was made to │
│              ┆             ┆ 8a1         ┆             ┆             ┆             ┆ access …    │
│ 2024-05-15T1 ┆ win10       ┆ fe80::24b4: ┆ information ┆ 4663        ┆ Removable   ┆ An attempt  │
│ 6:10:07.136Z ┆             ┆ 3691:44a6:3 ┆             ┆             ┆ Storage     ┆ was made to │
│              ┆             ┆ 8a1         ┆             ┆             ┆             ┆ access …    │
│ 2024-05-15T1 ┆ win10       ┆ fe80::24b4: ┆ information ┆ 4663        ┆ Removable   ┆ An attempt  │
│ 6:10:07.136Z ┆             ┆ 3691:44a6:3 ┆             ┆             ┆ Storage     ┆ was made to │
│              ┆             ┆ 8a1         ┆             ┆             ┆             ┆ access …    │
│ 2024-05-15T1 ┆ win10       ┆ fe80::24b4: ┆ information ┆ 4663        ┆ Removable   ┆ An attempt  │
│ 6:10:07.149Z ┆             ┆ 3691:44a6:3 ┆             ┆             ┆ Storage     ┆ was made to │
│              ┆             ┆ 8a1         ┆             ┆             ┆             ┆ access …    │
│ 2024-05-15T1 ┆ win10       ┆ fe80::24b4: ┆ information ┆ 4663        ┆ Removable   ┆ An attempt  │
│ 6:10:07.149Z ┆             ┆ 3691:44a6:3 ┆             ┆             ┆ Storage     ┆ was made to │
│              ┆             ┆ 8a1         ┆             ┆             ┆             ┆ access …    │
└──────────────┴─────────────┴─────────────┴─────────────┴─────────────┴─────────────┴─────────────┘
Pandas DataFarme number of rows:  8000
Polars DataFrame memory usage: 4.76 MB
 
Icicle Graph Reset Zoom Search /home/marius/anaconda3/lib/python3.11/site-packages/polars/lazyframe/frame.py:1683:collect /.. /tmp/ipykernel_67611/2832609216.py:1:test_polars /ho.. ~.. ~:0:<method 'collect' of 'builtins.PyLazyFrame' objects> ~:.. /ho.. /.. /home/marius/an.. ~:0:<bu.. <string>:1:<module> ~:0:<built-in method builtins.exec> /home/mar..
In [49]:
%%prun
test_polars()
shape: (8_000, 7)
┌──────────────┬─────────────┬─────────────┬─────────────┬─────────────┬─────────────┬─────────────┐
│ @timestamp   ┆ host.hostna ┆ host.ip     ┆ log.level   ┆ winlog.even ┆ winlog.task ┆ message     │
│ ---          ┆ me          ┆ ---         ┆ ---         ┆ t_id        ┆ ---         ┆ ---         │
│ str          ┆ ---         ┆ str         ┆ str         ┆ ---         ┆ str         ┆ str         │
│              ┆ str         ┆             ┆             ┆ i64         ┆             ┆             │
╞══════════════╪═════════════╪═════════════╪═════════════╪═════════════╪═════════════╪═════════════╡
│ 2024-05-15T1 ┆ win10       ┆ fe80::24b4: ┆ information ┆ 13          ┆ Registry    ┆ Registry    │
│ 5:57:18.471Z ┆             ┆ 3691:44a6:3 ┆             ┆             ┆ value set   ┆ value set:  │
│              ┆             ┆ 8a1         ┆             ┆             ┆ (rule:      ┆ RuleName: … │
│              ┆             ┆             ┆             ┆             ┆ Regi…       ┆             │
│ 2024-05-15T1 ┆ win10       ┆ fe80::24b4: ┆ information ┆ 13          ┆ Registry    ┆ Registry    │
│ 5:57:18.471Z ┆             ┆ 3691:44a6:3 ┆             ┆             ┆ value set   ┆ value set:  │
│              ┆             ┆ 8a1         ┆             ┆             ┆ (rule:      ┆ RuleName: … │
│              ┆             ┆             ┆             ┆             ┆ Regi…       ┆             │
│ 2024-05-15T1 ┆ win10       ┆ fe80::24b4: ┆ information ┆ 13          ┆ Registry    ┆ Registry    │
│ 5:57:18.471Z ┆             ┆ 3691:44a6:3 ┆             ┆             ┆ value set   ┆ value set:  │
│              ┆             ┆ 8a1         ┆             ┆             ┆ (rule:      ┆ RuleName: … │
│              ┆             ┆             ┆             ┆             ┆ Regi…       ┆             │
│ 2024-05-15T1 ┆ win10       ┆ fe80::24b4: ┆ information ┆ 13          ┆ Registry    ┆ Registry    │
│ 5:57:18.471Z ┆             ┆ 3691:44a6:3 ┆             ┆             ┆ value set   ┆ value set:  │
│              ┆             ┆ 8a1         ┆             ┆             ┆ (rule:      ┆ RuleName: … │
│              ┆             ┆             ┆             ┆             ┆ Regi…       ┆             │
│ 2024-05-15T1 ┆ win10       ┆ fe80::24b4: ┆ information ┆ 13          ┆ Registry    ┆ Registry    │
│ 5:57:18.471Z ┆             ┆ 3691:44a6:3 ┆             ┆             ┆ value set   ┆ value set:  │
│              ┆             ┆ 8a1         ┆             ┆             ┆ (rule:      ┆ RuleName: … │
│              ┆             ┆             ┆             ┆             ┆ Regi…       ┆             │
│ …            ┆ …           ┆ …           ┆ …           ┆ …           ┆ …           ┆ …           │
│ 2024-05-15T1 ┆ win10       ┆ fe80::24b4: ┆ information ┆ 4663        ┆ Removable   ┆ An attempt  │
│ 6:10:07.128Z ┆             ┆ 3691:44a6:3 ┆             ┆             ┆ Storage     ┆ was made to │
│              ┆             ┆ 8a1         ┆             ┆             ┆             ┆ access …    │
│ 2024-05-15T1 ┆ win10       ┆ fe80::24b4: ┆ information ┆ 4663        ┆ Removable   ┆ An attempt  │
│ 6:10:07.136Z ┆             ┆ 3691:44a6:3 ┆             ┆             ┆ Storage     ┆ was made to │
│              ┆             ┆ 8a1         ┆             ┆             ┆             ┆ access …    │
│ 2024-05-15T1 ┆ win10       ┆ fe80::24b4: ┆ information ┆ 4663        ┆ Removable   ┆ An attempt  │
│ 6:10:07.136Z ┆             ┆ 3691:44a6:3 ┆             ┆             ┆ Storage     ┆ was made to │
│              ┆             ┆ 8a1         ┆             ┆             ┆             ┆ access …    │
│ 2024-05-15T1 ┆ win10       ┆ fe80::24b4: ┆ information ┆ 4663        ┆ Removable   ┆ An attempt  │
│ 6:10:07.149Z ┆             ┆ 3691:44a6:3 ┆             ┆             ┆ Storage     ┆ was made to │
│              ┆             ┆ 8a1         ┆             ┆             ┆             ┆ access …    │
│ 2024-05-15T1 ┆ win10       ┆ fe80::24b4: ┆ information ┆ 4663        ┆ Removable   ┆ An attempt  │
│ 6:10:07.149Z ┆             ┆ 3691:44a6:3 ┆             ┆             ┆ Storage     ┆ was made to │
│              ┆             ┆ 8a1         ┆             ┆             ┆             ┆ access …    │
└──────────────┴─────────────┴─────────────┴─────────────┴─────────────┴─────────────┴─────────────┘
Pandas DataFarme number of rows:  8000
Polars DataFrame memory usage: 4.76 MB
 
         256 function calls (253 primitive calls) in 0.020 seconds

   Ordered by: internal time

   ncalls  tottime  percall  cumtime  percall filename:lineno(function)
        2    0.014    0.007    0.014    0.007 {method 'collect' of 'builtins.PyLazyFrame' objects}
        1    0.003    0.003    0.003    0.003 {built-in method new_from_ndjson}
        2    0.001    0.001    0.001    0.001 {built-in method posix.stat}
        1    0.000    0.000    0.000    0.000 {method 'as_str' of 'builtins.PyDataFrame' objects}
        1    0.000    0.000    0.020    0.020 <string>:1(<module>)
        1    0.000    0.000    0.020    0.020 2832609216.py:1(test_polars)
        1    0.000    0.000    0.000    0.000 socket.py:543(send)
        2    0.000    0.000    0.000    0.000 wrap.py:12(wrap_df)
        6    0.000    0.000    0.000    0.000 iostream.py:610(write)
        1    0.000    0.000    0.020    0.020 {built-in method builtins.exec}
        2    0.000    0.000    0.014    0.007 frame.py:1683(collect)
        1    0.000    0.000    0.004    0.004 ndjson.py:86(scan_ndjson)
        1    0.000    0.000    0.000    0.000 2832609216.py:9(<listcomp>)
      2/1    0.000    0.000    0.004    0.004 deprecation.py:130(wrapper)
        2    0.000    0.000    0.000    0.000 wrap.py:16(wrap_ldf)
        7    0.000    0.000    0.000    0.000 expr.py:1917(cast)
        1    0.000    0.000    0.001    0.001 various.py:182(normalize_filepath)
    41/39    0.000    0.000    0.000    0.000 {built-in method builtins.isinstance}
        7    0.000    0.000    0.000    0.000 col.py:20(_create_col)
        3    0.000    0.000    0.001    0.000 {built-in method builtins.print}
        1    0.000    0.000    0.000    0.000 frame.py:4006(with_columns)
        2    0.000    0.000    0.000    0.000 {method 'optimization_toggle' of 'builtins.PyLazyFrame' objects}
        1    0.000    0.000    0.000    0.000 frame.py:7998(lazy)
        3    0.000    0.000    0.000    0.000 frame.py:316(_from_pyldf)
        1    0.000    0.000    0.001    0.001 frame.py:8164(with_columns)
        7    0.000    0.000    0.000    0.000 {method 'cast' of 'builtins.PyExpr' objects}
        1    0.000    0.000    0.000    0.000 iostream.py:243(schedule)
        7    0.000    0.000    0.000    0.000 convert.py:388(py_type_to_dtype)
       19    0.000    0.000    0.000    0.000 {built-in method __new__ of type object at 0x860f60}
        6    0.000    0.000    0.000    0.000 iostream.py:505(_is_master_process)
        1    0.000    0.000    0.000    0.000 {method 'with_columns' of 'builtins.PyLazyFrame' objects}
        1    0.000    0.000    0.000    0.000 <frozen os>:674(__getitem__)
        7    0.000    0.000    0.000    0.000 {col}
        7    0.000    0.000    0.000    0.000 col.py:145(__new__)
        1    0.000    0.000    0.000    0.000 <frozen genericpath>:39(isdir)
        7    0.000    0.000    0.000    0.000 wrap.py:24(wrap_expr)
        1    0.000    0.000    0.000    0.000 <frozen posixpath>:229(expanduser)
        6    0.000    0.000    0.000    0.000 {built-in method posix.getpid}
       14    0.000    0.000    0.000    0.000 expr.py:131(_from_pyexpr)
        1    0.000    0.000    0.000    0.000 {method 'lazy' of 'builtins.PyDataFrame' objects}
        1    0.000    0.000    0.000    0.000 typing.py:1579(__subclasscheck__)
        2    0.000    0.000    0.000    0.000 frame.py:439(_from_pydf)
        1    0.000    0.000    0.000    0.000 parse_expr_input.py:56(<listcomp>)
        7    0.000    0.000    0.000    0.000 convert.py:146(is_polars_dtype)
        1    0.000    0.000    0.000    0.000 threading.py:1185(is_alive)
        1    0.000    0.000    0.001    0.001 <frozen genericpath>:16(exists)
        1    0.000    0.000    0.000    0.000 {method 'estimated_size' of 'builtins.PyDataFrame' objects}
        6    0.000    0.000    0.000    0.000 iostream.py:532(_schedule_flush)
        1    0.000    0.000    0.000    0.000 frame.py:980(__str__)
        1    0.000    0.000    0.000    0.000 parse_expr_input.py:59(_parse_inputs_as_iterable)
        1    0.000    0.000    0.000    0.000 <frozen _collections_abc>:771(get)
        1    0.000    0.000    0.000    0.000 threading.py:1118(_wait_for_tstate_lock)
        7    0.000    0.000    0.000    0.000 parse_expr_input.py:85(parse_as_expression)
        1    0.000    0.000    0.000    0.000 frame.py:3600(estimated_size)
        1    0.000    0.000    0.000    0.000 parse_expr_input.py:50(_parse_positional_inputs)
        1    0.000    0.000    0.000    0.000 <frozen os>:756(encode)
        1    0.000    0.000    0.000    0.000 frame.py:591(shape)
        1    0.000    0.000    0.000    0.000 {built-in method builtins.issubclass}
        1    0.000    0.000    0.000    0.000 iostream.py:127(_event_pipe)
        1    0.000    0.000    0.000    0.000 parse_expr_input.py:72(_is_iterable)
        1    0.000    0.000    0.000    0.000 parse_expr_input.py:20(parse_as_list_of_expressions)
        1    0.000    0.000    0.000    0.000 typing.py:1304(__instancecheck__)
        1    0.000    0.000    0.000    0.000 <frozen abc>:121(__subclasscheck__)
        1    0.000    0.000    0.000    0.000 {built-in method _abc._abc_subclasscheck}
        1    0.000    0.000    0.000    0.000 {method 'shape' of 'builtins.PyDataFrame' objects}
        1    0.000    0.000    0.000    0.000 {method 'acquire' of '_thread.lock' objects}
        7    0.000    0.000    0.000    0.000 {built-in method builtins.len}
        6    0.000    0.000    0.000    0.000 {method 'write' of '_io.StringIO' objects}
        6    0.000    0.000    0.000    0.000 {method '__exit__' of '_thread.RLock' objects}
        1    0.000    0.000    0.000    0.000 {built-in method _stat.S_ISDIR}
        1    0.000    0.000    0.000    0.000 various.py:210(scale_bytes)
        1    0.000    0.000    0.000    0.000 {method 'disable' of '_lsprof.Profiler' objects}
        2    0.000    0.000    0.000    0.000 {method 'get' of 'dict' objects}
        2    0.000    0.000    0.000    0.000 deprecation.py:143(_rename_keyword_argument)
        1    0.000    0.000    0.000    0.000 {method 'encode' of 'str' objects}
        1    0.000    0.000    0.000    0.000 {method 'items' of 'dict' objects}
        1    0.000    0.000    0.000    0.000 {method 'startswith' of 'str' objects}
        1    0.000    0.000    0.000    0.000 _utils.py:58(parse_row_index_args)
        1    0.000    0.000    0.000    0.000 threading.py:568(is_set)
        1    0.000    0.000    0.000    0.000 {method 'append' of 'collections.deque' objects}
        1    0.000    0.000    0.000    0.000 {built-in method posix.fspath}
In [50]:
def test_pandas():
    # Load the JSON file into a Pandas DataFrame
    pd_df = pd.read_json(file_path, lines=True, dtype=pandas_schema)
    pd_memory_usage = pd_df.memory_usage(deep=True).sum()

    # Get the number of rows in the Pandas DataFrame
    num_rows_pandas = pd_df.shape[0]

    print(pd_df)

    print(f"Pandas DataFarme number of rows:  {num_rows_pandas}")
    print(f"Pandas DataFrame memory usage: {pd_memory_usage / (1024 ** 2):.2f} MB")    
In [51]:
%%flame -q --inverted
test_pandas()
Out [51]:
                    @timestamp host.hostname                    host.ip  \
0     2024-05-15T15:57:18.471Z         win10  fe80::24b4:3691:44a6:38a1   
1     2024-05-15T15:57:18.471Z         win10  fe80::24b4:3691:44a6:38a1   
2     2024-05-15T15:57:18.471Z         win10  fe80::24b4:3691:44a6:38a1   
3     2024-05-15T15:57:18.471Z         win10  fe80::24b4:3691:44a6:38a1   
4     2024-05-15T15:57:18.471Z         win10  fe80::24b4:3691:44a6:38a1   
...                        ...           ...                        ...   
7995  2024-05-15T16:10:07.128Z         win10  fe80::24b4:3691:44a6:38a1   
7996  2024-05-15T16:10:07.136Z         win10  fe80::24b4:3691:44a6:38a1   
7997  2024-05-15T16:10:07.136Z         win10  fe80::24b4:3691:44a6:38a1   
7998  2024-05-15T16:10:07.149Z         win10  fe80::24b4:3691:44a6:38a1   
7999  2024-05-15T16:10:07.149Z         win10  fe80::24b4:3691:44a6:38a1   

        log.level  winlog.event_id                               winlog.task  \
0     information               13  Registry value set (rule: RegistryEvent)   
1     information               13  Registry value set (rule: RegistryEvent)   
2     information               13  Registry value set (rule: RegistryEvent)   
3     information               13  Registry value set (rule: RegistryEvent)   
4     information               13  Registry value set (rule: RegistryEvent)   
...           ...              ...                                       ...   
7995  information             4663                         Removable Storage   
7996  information             4663                         Removable Storage   
7997  information             4663                         Removable Storage   
7998  information             4663                         Removable Storage   
7999  information             4663                         Removable Storage   

                                                message  
0     Registry value set:\nRuleName: InvDB-Ver\nEven...  
1     Registry value set:\nRuleName: InvDB-Path\nEve...  
2     Registry value set:\nRuleName: InvDB-Pub\nEven...  
3     Registry value set:\nRuleName: InvDB-CompileTi...  
4     Registry value set:\nRuleName: InvDB-Ver\nEven...  
...                                                 ...  
7995  An attempt was made to access an object.\n\nSu...  
7996  An attempt was made to access an object.\n\nSu...  
7997  An attempt was made to access an object.\n\nSu...  
7998  An attempt was made to access an object.\n\nSu...  
7999  An attempt was made to access an object.\n\nSu...  

[8000 rows x 7 columns]
Pandas DataFarme number of rows:  8000
Pandas DataFrame memory usage: 7.56 MB
 
Icicle Graph Reset Zoom Search /tmp/ipykernel_67611/1231667944.py:1:test_pandas /.. /home/marius/anaconda3/lib/python3.11/site-packages/pandas/io/json/_json.py:500:read_json ~:0:<built-in method pandas._libs.js.. /home/m.. /home/marius.. /home/marius/anaco.. ~:.. /home/marius/anaconda3/lib/python3.11/site-packages/pandas/io/json/_json.py:1022:_get_o.. ~:0:<.. /home/marius/anaco.. /home/marius/anaconda3/lib/python3.11/site-packages/pandas/io/json/_json.py:1172:parse <f.. /home/.. /home/ma.. ~:0:<.. /hom.. /hom.. /home.. /home/m.. ~:0:<built-in method builtins.exec> /.. /home/marius.. ~:.. /.. /ho.. ~:.. /home/mariu.. /home/.. /home/marius/ana.. /home/marius/anac.. /h.. /home/marius.. /.. /home/marius/an.. ~:0:<method 'rea.. ~:.. /home/mariu.. /home/.. /home/marius/an.. /.. /home/mariu.. /.. /home/marius/anaconda3/lib/python3.11/site-packages/pandas/io/.. /home/marius/anaco.. /ho.. /home.. /h.. /.. /home/m.. /home/ma.. /home/.. ~:0:<buil.. ~:0:<pandas._l.. /.. /home/m.. /home/ma.. /home/marius/anaconda3/lib/python3.11/site-packages/pandas/io/json/_json.py:980:read /h.. /home/marius/ana.. /home/.. /home/marius/anacon.. <string>:1:<module>
In [52]:
%%prun
test_pandas()
                    @timestamp host.hostname                    host.ip  \
0     2024-05-15T15:57:18.471Z         win10  fe80::24b4:3691:44a6:38a1   
1     2024-05-15T15:57:18.471Z         win10  fe80::24b4:3691:44a6:38a1   
2     2024-05-15T15:57:18.471Z         win10  fe80::24b4:3691:44a6:38a1   
3     2024-05-15T15:57:18.471Z         win10  fe80::24b4:3691:44a6:38a1   
4     2024-05-15T15:57:18.471Z         win10  fe80::24b4:3691:44a6:38a1   
...                        ...           ...                        ...   
7995  2024-05-15T16:10:07.128Z         win10  fe80::24b4:3691:44a6:38a1   
7996  2024-05-15T16:10:07.136Z         win10  fe80::24b4:3691:44a6:38a1   
7997  2024-05-15T16:10:07.136Z         win10  fe80::24b4:3691:44a6:38a1   
7998  2024-05-15T16:10:07.149Z         win10  fe80::24b4:3691:44a6:38a1   
7999  2024-05-15T16:10:07.149Z         win10  fe80::24b4:3691:44a6:38a1   

        log.level  winlog.event_id                               winlog.task  \
0     information               13  Registry value set (rule: RegistryEvent)   
1     information               13  Registry value set (rule: RegistryEvent)   
2     information               13  Registry value set (rule: RegistryEvent)   
3     information               13  Registry value set (rule: RegistryEvent)   
4     information               13  Registry value set (rule: RegistryEvent)   
...           ...              ...                                       ...   
7995  information             4663                         Removable Storage   
7996  information             4663                         Removable Storage   
7997  information             4663                         Removable Storage   
7998  information             4663                         Removable Storage   
7999  information             4663                         Removable Storage   

                                                message  
0     Registry value set:\nRuleName: InvDB-Ver\nEven...  
1     Registry value set:\nRuleName: InvDB-Path\nEve...  
2     Registry value set:\nRuleName: InvDB-Pub\nEven...  
3     Registry value set:\nRuleName: InvDB-CompileTi...  
4     Registry value set:\nRuleName: InvDB-Ver\nEven...  
...                                                 ...  
7995  An attempt was made to access an object.\n\nSu...  
7996  An attempt was made to access an object.\n\nSu...  
7997  An attempt was made to access an object.\n\nSu...  
7998  An attempt was made to access an object.\n\nSu...  
7999  An attempt was made to access an object.\n\nSu...  

[8000 rows x 7 columns]
Pandas DataFarme number of rows:  8000
Pandas DataFrame memory usage: 7.56 MB
 
         46681 function calls (46472 primitive calls) in 0.113 seconds

   Ordered by: internal time

   ncalls  tottime  percall  cumtime  percall filename:lineno(function)
        1    0.029    0.029    0.029    0.029 {built-in method pandas._libs.json.ujson_loads}
        6    0.010    0.002    0.010    0.002 {pandas._libs.lib.memory_usage_of_objects}
        1    0.009    0.009    0.012    0.012 {method 'read' of '_io.TextIOWrapper' objects}
     8001    0.005    0.000    0.005    0.000 construction.py:915(<genexpr>)
        1    0.005    0.005    0.005    0.005 {pandas._libs.lib.dicts_to_array}
       94    0.004    0.000    0.004    0.000 {method 'split' of 'str' objects}
        1    0.003    0.003    0.009    0.009 {pandas._libs.lib.fast_unique_multiple_list_gen}
        1    0.003    0.003    0.010    0.010 _json.py:960(_combine_lines)
        1    0.003    0.003    0.003    0.003 {built-in method _codecs.utf_8_decode}
     64/8    0.003    0.000    0.003    0.000 {method 'join' of 'str' objects}
        1    0.003    0.003    0.113    0.113 <string>:1(<module>)
        1    0.002    0.002    0.049    0.049 _json.py:1360(_parse)
        6    0.002    0.000    0.002    0.000 {built-in method pandas._libs.lib.ensure_string_array}
        4    0.002    0.000    0.002    0.001 managers.py:2194(_stack_arrays)
     8002    0.002    0.000    0.003    0.000 _json.py:965(<genexpr>)
        1    0.002    0.002    0.002    0.002 construction.py:922(<listcomp>)
        1    0.001    0.001    0.004    0.004 _json.py:965(<listcomp>)
        1    0.001    0.001    0.091    0.091 _json.py:500(read_json)
     8001    0.001    0.000    0.001    0.000 {method 'strip' of 'str' objects}
        1    0.001    0.001    0.001    0.001 {built-in method io.open}
        9    0.001    0.000    0.001    0.000 {method 'astype' of 'numpy.ndarray' objects}
        2    0.001    0.001    0.002    0.001 managers.py:2224(_merge_blocks)
    53/51    0.001    0.000    0.001    0.000 {built-in method numpy.core._multiarray_umath.implement_array_function}
        1    0.001    0.001    0.010    0.010 _json.py:1422(_try_convert_types)
2422/2396    0.001    0.000    0.001    0.000 {built-in method builtins.isinstance}
     8001    0.001    0.000    0.001    0.000 {method 'keys' of 'dict' objects}
       32    0.001    0.000    0.001    0.000 generic.py:6147(__finalize__)
        1    0.001    0.001    0.001    0.001 {built-in method posix.stat}
        2    0.000    0.000    0.002    0.001 _json.py:1282(_try_convert_to_date)
        2    0.000    0.000    0.005    0.003 construction.py:96(arrays_to_mgr)
       98    0.000    0.000    0.000    0.000 generic.py:6206(__setattr__)
       24    0.000    0.000    0.001    0.000 base.py:510(find)
1415/1290    0.000    0.000    0.000    0.000 {built-in method builtins.len}
       29    0.000    0.000    0.001    0.000 {pandas._libs.lib.maybe_convert_objects}
        1    0.000    0.000    0.016    0.016 construction.py:793(to_arrays)
       21    0.000    0.000    0.000    0.000 managers.py:991(iget)
      661    0.000    0.000    0.000    0.000 format.py:428(len)
        1    0.000    0.000    0.000    0.000 socket.py:543(send)
       21    0.000    0.000    0.001    0.000 common.py:1587(pandas_dtype)
       47    0.000    0.000    0.000    0.000 {built-in method numpy.empty}
       93    0.000    0.000    0.000    0.000 config.py:127(_get_single_key)
        6    0.000    0.000    0.001    0.000 format.py:1332(_format_strings)
      198    0.000    0.000    0.000    0.000 base.py:236(construct_from_string)
       41    0.000    0.000    0.000    0.000 generic.py:274(__init__)
        5    0.000    0.000    0.001    0.000 base.py:478(__new__)
       19    0.000    0.000    0.001    0.000 construction.py:519(sanitize_array)
        7    0.000    0.000    0.001    0.000 series.py:371(__init__)
       18    0.000    0.000    0.000    0.000 {method 'reduce' of 'numpy.ufunc' objects}
       67    0.000    0.000    0.000    0.000 printing.py:162(pprint_thing)
       60    0.000    0.000    0.001    0.000 format.py:1355(_format)
        4    0.000    0.000    0.000    0.000 {pandas._libs.tslib.array_with_unit_to_datetime}
       63    0.000    0.000    0.000    0.000 base.py:5350(__getitem__)
       89    0.000    0.000    0.001    0.000 config.py:145(_get_option)
       91    0.000    0.000    0.000    0.000 config.py:633(_get_root)
       67    0.000    0.000    0.000    0.000 printing.py:193(as_escaped_string)
      182    0.000    0.000    0.000    0.000 config.py:647(_get_deprecated_option)
      212    0.000    0.000    0.000    0.000 generic.py:42(_instancecheck)
        6    0.000    0.000    0.000    0.000 {pandas._libs.lib.map_infer}
      212    0.000    0.000    0.000    0.000 generic.py:37(_check)
       21    0.000    0.000    0.001    0.000 frame.py:4402(_get_item_cache)
        1    0.000    0.000    0.004    0.004 format.py:843(_get_strcols_without_index)
       21    0.000    0.000    0.001    0.000 frame.py:3776(_ixs)
       35    0.000    0.000    0.000    0.000 numeric.py:290(full)
       18    0.000    0.000    0.001    0.000 format.py:1909(_make_fixed_width)
        2    0.000    0.000    0.002    0.001 managers.py:2137(_form_blocks)
        9    0.000    0.000    0.002    0.000 format.py:1217(format_array)
       10    0.000    0.000    0.003    0.000 astype.py:56(_astype_nansafe)
       61    0.000    0.000    0.000    0.000 {built-in method builtins.max}
      180    0.000    0.000    0.000    0.000 format.py:1932(just)
       14    0.000    0.000    0.000    0.000 base.py:5300(__contains__)
       15    0.000    0.000    0.000    0.000 {built-in method numpy.array}
       67    0.000    0.000    0.000    0.000 inference.py:373(is_sequence)
        9    0.000    0.000    0.001    0.000 indexing.py:1006(_getitem_lowerdim)
        2    0.000    0.000    0.021    0.011 frame.py:665(__init__)
      198    0.000    0.000    0.000    0.000 format.py:1923(<genexpr>)
       21    0.000    0.000    0.001    0.000 frame.py:4384(_box_col_values)
       73    0.000    0.000    0.000    0.000 missing.py:184(_isna)
       24    0.000    0.000    0.000    0.000 warnings.py:466(__enter__)
       24    0.000    0.000    0.001    0.000 frame.py:1392(items)
       88    0.000    0.000    0.001    0.000 config.py:271(__call__)
        7    0.000    0.000    0.003    0.000 format.py:890(format_col)
      301    0.000    0.000    0.000    0.000 {built-in method builtins.getattr}
       24    0.000    0.000    0.000    0.000 warnings.py:181(_add_filter)
        2    0.000    0.000    0.000    0.000 {pandas._libs.lib.array_equivalent_object}
       17    0.000    0.000    0.000    0.000 cast.py:1147(maybe_infer_to_datetimelike)
        9    0.000    0.000    0.007    0.001 _json.py:1204(_try_convert_data)
        7    0.000    0.000    0.004    0.001 managers.py:308(apply)
        5    0.000    0.000    0.000    0.000 printing.py:28(adjoin)
       18    0.000    0.000    0.000    0.000 format.py:1938(<listcomp>)
       56    0.000    0.000    0.000    0.000 {pandas._libs.lib.is_list_like}
       48    0.000    0.000    0.000    0.000 {built-in method numpy.asarray}
       62    0.000    0.000    0.000    0.000 {built-in method builtins.all}
        9    0.000    0.000    0.001    0.000 indexing.py:1139(__getitem__)
        2    0.000    0.000    0.010    0.005 _json.py:1396(_process_converter)
       89    0.000    0.000    0.000    0.000 config.py:686(_warn_if_deprecated)
       35    0.000    0.000    0.000    0.000 <__array_function__ internals>:177(copyto)
       15    0.000    0.000    0.000    0.000 blocks.py:2388(new_block)
       15    0.000    0.000    0.000    0.000 inference.py:273(is_dict_like)
        6    0.000    0.000    0.000    0.000 base.py:836(__iter__)
        7    0.000    0.000    0.010    0.001 base.py:1135(_memory_usage)
       24    0.000    0.000    0.000    0.000 {method 'remove' of 'list' objects}
       20    0.000    0.000    0.000    0.000 printing.py:65(<listcomp>)
        7    0.000    0.000    0.000    0.000 blocks.py:247(make_block)
       60    0.000    0.000    0.000    0.000 {built-in method _abc._abc_instancecheck}
        8    0.000    0.000    0.000    0.000 managers.py:1825(from_array)
      201    0.000    0.000    0.000    0.000 {method 'replace' of 'str' objects}
       28    0.000    0.000    0.000    0.000 printing.py:69(<listcomp>)
        2    0.000    0.000    0.001    0.000 concat.py:618(get_result)
       41    0.000    0.000    0.000    0.000 flags.py:53(__init__)
        1    0.000    0.000    0.000    0.000 construction.py:928(_finalize_columns_and_data)
        1    0.000    0.000    0.110    0.110 1231667944.py:1(test_pandas)
        9    0.000    0.000    0.001    0.000 indexing.py:1651(_getitem_tuple)
       40    0.000    0.000    0.000    0.000 {built-in method builtins.any}
       31    0.000    0.000    0.000    0.000 generic.py:562(_get_axis)
       93    0.000    0.000    0.000    0.000 config.py:674(_translate_key)
       16    0.000    0.000    0.000    0.000 format.py:903(_get_formatter)
       14    0.000    0.000    0.000    0.000 numpy_.py:98(__init__)
       93    0.000    0.000    0.000    0.000 config.py:615(_select_options)
       32    0.000    0.000    0.000    0.000 managers.py:1960(internal_values)
        3    0.000    0.000    0.000    0.000 {method '_rebuild_blknos_and_blklocs' of 'pandas._libs.internals.BlockManager' objects}
       10    0.000    0.000    0.003    0.000 astype.py:158(astype_array)
       32    0.000    0.000    0.000    0.000 generic.py:335(_from_mgr)
       16    0.000    0.000    0.000    0.000 blocks.py:2317(maybe_coerce_values)
       60    0.000    0.000    0.000    0.000 __init__.py:33(using_copy_on_write)
       14    0.000    0.000    0.000    0.000 {method 'get_loc' of 'pandas._libs.index.IndexEngine' objects}
        6    0.000    0.000    0.000    0.000 {built-in method pandas._libs.missing.isnaobj}
        7    0.000    0.000    0.005    0.001 generic.py:6368(astype)
       72    0.000    0.000    0.000    0.000 base.py:909(__len__)
        2    0.000    0.000    0.000    0.000 construction.py:596(_homogenize)
       48    0.000    0.000    0.000    0.000 printing.py:60(justify)
        1    0.000    0.000    0.000    0.000 {method 'close' of '_io.TextIOWrapper' objects}
        7    0.000    0.000    0.004    0.001 blocks.py:588(astype)
        1    0.000    0.000    0.000    0.000 concat.py:94(concatenate_managers)
        9    0.000    0.000    0.001    0.000 indexing.py:1681(_getitem_axis)
       21    0.000    0.000    0.000    0.000 frame.py:654(_constructor_sliced_from_mgr)
       48    0.000    0.000    0.000    0.000 format.py:431(justify)
       73    0.000    0.000    0.000    0.000 missing.py:101(isna)
      140    0.000    0.000    0.000    0.000 {built-in method builtins.hasattr}
        4    0.000    0.000    0.001    0.000 datetimes.py:721(to_datetime)
        6    0.000    0.000    0.000    0.000 iostream.py:610(write)
       19    0.000    0.000    0.001    0.000 base.py:7521(ensure_index)
       20    0.000    0.000    0.000    0.000 blocks.py:2346(get_block_type)
       63    0.000    0.000    0.000    0.000 common.py:149(cast_scalar_indexer)
      142    0.000    0.000    0.000    0.000 {pandas._libs.lib.is_scalar}
       14    0.000    0.000    0.000    0.000 managers.py:1949(dtype)
        2    0.000    0.000    0.000    0.000 {method 'get_slice' of 'pandas._libs.internals.BlockManager' objects}
      136    0.000    0.000    0.000    0.000 {built-in method builtins.issubclass}
        1    0.000    0.000    0.015    0.015 construction.py:891(_list_of_dict_to_arrays)
       24    0.000    0.000    0.000    0.000 warnings.py:487(__exit__)
        9    0.000    0.000    0.000    0.000 indexing.py:931(_validate_tuple_indexer)
       21    0.000    0.000    0.000    0.000 series.py:1372(_set_as_cached)
        7    0.000    0.000    0.000    0.000 _dtype.py:344(_name_get)
        7    0.000    0.000    0.000    0.000 missing.py:261(_isna_array)
       18    0.000    0.000    0.000    0.000 dtypes.py:1266(construct_from_string)
        7    0.000    0.000    0.000    0.000 warnings.py:130(filterwarnings)
       14    0.000    0.000    0.000    0.000 base.py:3763(get_loc)
        6    0.000    0.000    0.000    0.000 missing.py:380(notna)
       30    0.000    0.000    0.000    0.000 format.py:1617(<lambda>)
       27    0.000    0.000    0.000    0.000 construction.py:485(ensure_wrapped_if_datetimelike)
        3    0.000    0.000    0.000    0.000 base.py:1418(_format_with_header)
       23    0.000    0.000    0.000    0.000 {method 'match' of 're.Pattern' objects}
       18    0.000    0.000    0.000    0.000 dtypes.py:814(construct_from_string)
        9    0.000    0.000    0.000    0.000 indexing.py:1614(_is_scalar_access)
       18    0.000    0.000    0.000    0.000 dtypes.py:332(construct_from_string)
       66    0.000    0.000    0.000    0.000 {built-in method pandas._libs.missing.checknull}
       10    0.000    0.000    0.000    0.000 format.py:479(get_adjustment)
        7    0.000    0.000    0.000    0.000 base.py:649(_simple_new)
      217    0.000    0.000    0.000    0.000 {method 'ljust' of 'str' objects}
        1    0.000    0.000    0.075    0.075 _json.py:980(read)
       16    0.000    0.000    0.000    0.000 common.py:137(is_object_dtype)
       32    0.000    0.000    0.000    0.000 flags.py:89(allows_duplicate_labels)
       24    0.000    0.000    0.000    0.000 warnings.py:440(__init__)
        1    0.000    0.000    0.113    0.113 {built-in method builtins.exec}
       18    0.000    0.000    0.000    0.000 dtypes.py:1021(construct_from_string)
        9    0.000    0.000    0.002    0.000 format.py:1328(get_result)
        1    0.000    0.000    0.000    0.000 string.py:119(_join_multiline)
        3    0.000    0.000    0.000    0.000 _strptime.py:309(_strptime)
       30    0.000    0.000    0.000    0.000 construction.py:420(extract_array)
       21    0.000    0.000    0.000    0.000 blocks.py:1007(iget)
        1    0.000    0.000    0.012    0.012 frame.py:3471(memory_usage)
       76    0.000    0.000    0.000    0.000 {pandas._libs.lib.is_np_dtype}
       77    0.000    0.000    0.000    0.000 format.py:884(<genexpr>)
       10    0.000    0.000    0.000    0.000 generic.py:760(_set_axis)
       18    0.000    0.000    0.000    0.000 indexing.py:1536(_validate_key)
       17    0.000    0.000    0.000    0.000 warnings.py:165(simplefilter)
        4    0.000    0.000    0.000    0.000 _asarray.py:31(require)
       25    0.000    0.000    0.000    0.000 common.py:1425(_is_dtype_type)
        6    0.000    0.000    0.000    0.000 concat.py:322(_get_block_for_concat_plan)
       60    0.000    0.000    0.000    0.000 <frozen abc>:117(__instancecheck__)
        7    0.000    0.000    0.004    0.001 astype.py:192(astype_array_safe)
       25    0.000    0.000    0.000    0.000 common.py:96(is_bool_indexer)
        1    0.000    0.000    0.001    0.001 common.py:652(get_handle)
        7    0.000    0.000    0.000    0.000 construction.py:1028(convert)
      157    0.000    0.000    0.000    0.000 {method 'append' of 'list' objects}
        2    0.000    0.000    0.000    0.000 concat.py:403(__init__)
        9    0.000    0.000    0.000    0.000 indexing.py:2678(check_dict_or_set_indexers)
        2    0.000    0.000    0.000    0.000 base.py:1683(_validate_names)
        3    0.000    0.000    0.000    0.000 cast.py:119(maybe_convert_platform)
        3    0.000    0.000    0.000    0.000 {built-in method numpy.arange}
       14    0.000    0.000    0.000    0.000 managers.py:1964(array_values)
      103    0.000    0.000    0.000    0.000 {pandas._libs.lib.is_integer}
        4    0.000    0.000    0.001    0.000 base.py:1038(astype)
        7    0.000    0.000    0.000    0.000 string.py:129(<listcomp>)
        1    0.000    0.000    0.062    0.062 _json.py:1022(_get_object_parser)
        1    0.000    0.000    0.000    0.000 concat.py:296(_get_combined_plan)
       47    0.000    0.000    0.000    0.000 range.py:963(__len__)
        2    0.000    0.000    0.000    0.000 parse.py:374(urlparse)
        6    0.000    0.000    0.000    0.000 concat.py:389(is_na)
       39    0.000    0.000    0.000    0.000 inference.py:334(is_hashable)
        6    0.000    0.000    0.000    0.000 fromnumeric.py:69(_wrapreduction)
        7    0.000    0.000    0.004    0.001 managers.py:405(astype)
        2    0.000    0.000    0.000    0.000 format.py:956(_get_formatted_index)
        3    0.000    0.000    0.000    0.000 cast.py:1544(construct_1d_object_array_from_listlike)
        8    0.000    0.000    0.000    0.000 range.py:198(_simple_new)
        9    0.000    0.000    0.000    0.000 common.py:1066(is_numeric_dtype)
        5    0.000    0.000    0.000    0.000 format.py:434(adjoin)
        4    0.000    0.000    0.001    0.000 datetimes.py:216(_maybe_cache)
       49    0.000    0.000    0.000    0.000 {built-in method __new__ of type object at 0x860f60}
       36    0.000    0.000    0.000    0.000 managers.py:1799(__init__)
        1    0.000    0.000    0.000    0.000 format.py:915(_get_formatted_column_labels)
        2    0.000    0.000    0.000    0.000 generic.py:4296(_slice)
        9    0.000    0.000    0.000    0.000 series.py:653(name)
       18    0.000    0.000    0.000    0.000 string_.py:135(construct_from_string)
        8    0.000    0.000    0.000    0.000 series.py:581(_constructor_from_mgr)
        6    0.000    0.000    0.000    0.000 __init__.py:272(_compile)
       61    0.000    0.000    0.000    0.000 printing.py:57(<genexpr>)
        4    0.000    0.000    0.000    0.000 datetimes.py:526(_to_datetime_with_unit)
        2    0.000    0.000    0.004    0.002 managers.py:2068(create_block_manager_from_column_arrays)
       27    0.000    0.000    0.000    0.000 indexing.py:1144(<genexpr>)
        2    0.000    0.000    0.000    0.000 {pandas._libs.lib.is_all_arraylike}
       10    0.000    0.000    0.000    0.000 base.py:73(_validate_set_axis)
       32    0.000    0.000    0.000    0.000 series.py:750(_values)
      154    0.000    0.000    0.000    0.000 {method 'rjust' of 'str' objects}
       18    0.000    0.000    0.000    0.000 dtypes.py:2180(construct_from_string)
       14    0.000    0.000    0.000    0.000 indexing.py:1629(_validate_integer)
        6    0.000    0.000    0.000    0.000 {method 'reshape' of 'numpy.ndarray' objects}
        1    0.000    0.000    0.011    0.011 frame.py:3561(<listcomp>)
        5    0.000    0.000    0.000    0.000 base.py:574(_ensure_array)
        1    0.000    0.000    0.000    0.000 _parser.py:666(_parse)
        1    0.000    0.000    0.007    0.007 frame.py:1229(to_string)
       72    0.000    0.000    0.000    0.000 {built-in method _warnings._filters_mutated}
       71    0.000    0.000    0.000    0.000 {method 'get' of 'dict' objects}
       21    0.000    0.000    0.000    0.000 frame.py:651(_sliced_from_mgr)
        1    0.000    0.000    0.000    0.000 {built-in method _operator.gt}
        1    0.000    0.000    0.000    0.000 string.py:89(_insert_dot_separator_vertical)
        4    0.000    0.000    0.000    0.000 _asarray.py:112(<setcomp>)
        1    0.000    0.000    0.003    0.003 <frozen codecs>:319(decode)
        3    0.000    0.000    0.000    0.000 format.py:1619(<listcomp>)
       40    0.000    0.000    0.000    0.000 inference.py:300(<genexpr>)
        7    0.000    0.000    0.000    0.000 common.py:1268(is_extension_array_dtype)
        5    0.000    0.000    0.000    0.000 cast.py:1483(construct_1d_arraylike_from_scalar)
       14    0.000    0.000    0.000    0.000 blocks.py:2241(array_values)
        3    0.000    0.000    0.000    0.000 _parser.py:77(get_token)
       13    0.000    0.000    0.000    0.000 common.py:296(maybe_iterable_to_list)
        7    0.000    0.000    0.000    0.000 managers.py:1812(from_blocks)
       18    0.000    0.000    0.000    0.000 dtypes.py:1789(construct_from_string)
        1    0.000    0.000    0.000    0.000 array_ops.py:290(comparison_op)
       38    0.000    0.000    0.000    0.000 generic.py:548(_get_axis_number)
        5    0.000    0.000    0.000    0.000 base.py:69(shape)
       48    0.000    0.000    0.000    0.000 {method 'startswith' of 'str' objects}
       14    0.000    0.000    0.000    0.000 dtypes.py:1407(__init__)
        9    0.000    0.000    0.000    0.000 numerictypes.py:356(issubdtype)
       12    0.000    0.000    0.000    0.000 base.py:7616(maybe_extract_name)
        1    0.000    0.000    0.000    0.000 expressions.py:95(_evaluate_numexpr)
        1    0.000    0.000    0.000    0.000 iostream.py:243(schedule)
        1    0.000    0.000    0.000    0.000 range.py:902(_concat)
       14    0.000    0.000    0.000    0.000 construction.py:695(_sanitize_ndim)
        7    0.000    0.000    0.000    0.000 _json.py:1442(is_ok)
        1    0.000    0.000    0.000    0.000 numeric.py:2407(array_equal)
        7    0.000    0.000    0.000    0.000 series.py:703(name)
       10    0.000    0.000    0.000    0.000 common.py:1322(is_ea_or_datetimelike_dtype)
        2    0.000    0.000    0.000    0.000 config.py:153(_set_option)
        1    0.000    0.000    0.014    0.014 _json.py:816(__init__)
        4    0.000    0.000    0.000    0.000 blocks.py:297(slice_block_columns)
        1    0.000    0.000    0.000    0.000 blocks.py:2375(new_block_2d)
        8    0.000    0.000    0.001    0.000 <__array_function__ internals>:177(concatenate)
        1    0.000    0.000    0.007    0.007 frame.py:1123(__repr__)
       10    0.000    0.000    0.000    0.000 common.py:1562(validate_all_hashable)
        1    0.000    0.000    0.002    0.002 managers.py:2207(_consolidate)
        9    0.000    0.000    0.000    0.000 indexing.py:948(_is_nested_tuple_indexer)
        1    0.000    0.000    0.001    0.001 format.py:564(__init__)
       10    0.000    0.000    0.000    0.000 managers.py:225(set_axis)
        2    0.000    0.000    0.000    0.000 common.py:1155(_is_binary_mode)
       42    0.000    0.000    0.000    0.000 base.py:5127(_values)
        3    0.000    0.000    0.000    0.000 range.py:234(_data)
       18    0.000    0.000    0.000    0.000 common.py:367(apply_if_callable)
        1    0.000    0.000    0.000    0.000 common.py:228(asarray_tuplesafe)
        2    0.000    0.000    0.000    0.000 indexing.py:978(_getitem_tuple_same_dim)
        1    0.000    0.000    0.000    0.000 {pandas._libs.internals.get_concat_blkno_indexers}
        4    0.000    0.000    0.000    0.000 datetimes.py:369(_convert_listlike_datetimes)
       35    0.000    0.000    0.000    0.000 {method 'format' of 'str' objects}
        9    0.000    0.000    0.000    0.000 blocks.py:2467(extend_blocks)
       17    0.000    0.000    0.000    0.000 __init__.py:43(using_pyarrow_string_dtype)
        1    0.000    0.000    0.012    0.012 _json.py:896(_preprocess_data)
        1    0.000    0.000    0.000    0.000 cast.py:1569(maybe_cast_to_integer_array)
        1    0.000    0.000    0.000    0.000 cast.py:774(infer_dtype_from_scalar)
        2    0.000    0.000    0.000    0.000 base.py:5519(equals)
        3    0.000    0.000    0.007    0.002 {built-in method builtins.print}
        6    0.000    0.000    0.000    0.000 missing.py:305(_isna_string_dtype)
        5    0.000    0.000    0.000    0.000 api.py:379(default_index)
        1    0.000    0.000    0.000    0.000 _parser.py:62(__init__)
        1    0.000    0.000    0.004    0.004 construction.py:423(dict_to_mgr)
        1    0.000    0.000    0.002    0.002 _json.py:1185(_convert_axes)
        6    0.000    0.000    0.000    0.000 fromnumeric.py:2432(all)
        1    0.000    0.000    0.000    0.000 common.py:289(_get_filepath_or_buffer)
        2    0.000    0.000    0.001    0.001 concat.py:157(concat)
       53    0.000    0.000    0.000    0.000 {built-in method builtins.hash}
        1    0.000    0.000    0.000    0.000 base.py:2293(is_unique)
        2    0.000    0.000    0.000    0.000 base.py:5422(append)
        2    0.000    0.000    0.000    0.000 _json.py:1049(close)
       36    0.000    0.000    0.000    0.000 {method 'insert' of 'list' objects}
        1    0.000    0.000    0.000    0.000 common.py:538(infer_compression)
        7    0.000    0.000    0.010    0.001 series.py:5223(memory_usage)
        3    0.000    0.000    0.000    0.000 concat.py:572(_is_uniform_join_units)
        6    0.000    0.000    0.000    0.000 managers.py:2212(<lambda>)
        7    0.000    0.000    0.000    0.000 generic.py:6189(__getattr__)
       18    0.000    0.000    0.000    0.000 indexing.py:2651(is_label_like)
       13    0.000    0.000    0.000    0.000 blocks.py:187(is_extension)
        1    0.000    0.000    0.000    0.000 base.py:1427(<listcomp>)
        5    0.000    0.000    0.000    0.000 common.py:173(_expand_user)
        1    0.000    0.000    0.002    0.002 _json.py:912(_get_data_from_filepath)
        9    0.000    0.000    0.000    0.000 common.py:131(<lambda>)
        2    0.000    0.000    0.000    0.000 common.py:121(close)
        1    0.000    0.000    0.000    0.000 construction.py:487(<listcomp>)
        7    0.000    0.000    0.000    0.000 {method 'add_index_reference' of 'pandas._libs.internals.BlockValuesRefs' objects}
        9    0.000    0.000    0.000    0.000 format.py:1300(__init__)
        2    0.000    0.000    0.000    0.000 concat.py:492(_clean_keys_and_objs)
        3    0.000    0.000    0.000    0.000 blocks.py:198(_consolidate_key)
        1    0.000    0.000    0.000    0.000 string.py:128(<listcomp>)
        1    0.000    0.000    0.005    0.005 format.py:1077(to_string)
        1    0.000    0.000    0.000    0.000 range.py:489(copy)
       18    0.000    0.000    0.000    0.000 numerictypes.py:282(issubclass_)
       24    0.000    0.000    0.000    0.000 managers.py:169(blknos)
       14    0.000    0.000    0.000    0.000 construction.py:734(_sanitize_str_dtypes)
        7    0.000    0.000    0.000    0.000 frame.py:1539(__len__)
        9    0.000    0.000    0.000    0.000 concat.py:597(<genexpr>)
       66    0.000    0.000    0.000    0.000 generic.py:393(flags)
        9    0.000    0.000    0.000    0.000 common.py:514(is_string_or_object_np_dtype)
       63    0.000    0.000    0.000    0.000 {pandas._libs.lib.is_float}
       34    0.000    0.000    0.000    0.000 {method 'endswith' of 'str' objects}
        2    0.000    0.000    0.000    0.000 concat.py:478(_get_ndims)
       68    0.000    0.000    0.000    0.000 {built-in method builtins.iter}
        1    0.000    0.000    0.062    0.062 _json.py:1172(parse)
        1    0.000    0.000    0.000    0.000 base.py:842(_engine)
       20    0.000    0.000    0.000    0.000 common.py:1581(<genexpr>)
        1    0.000    0.000    0.000    0.000 {pandas._libs.missing.is_float_nan}
        6    0.000    0.000    0.000    0.000 <__array_function__ internals>:177(all)
        1    0.000    0.000    0.000    0.000 base.py:7092(_cmp_method)
        1    0.000    0.000    0.001    0.001 format.py:825(_truncate_vertically)
       18    0.000    0.000    0.000    0.000 indexing.py:966(_validate_key_length)
        9    0.000    0.000    0.000    0.000 <frozen importlib._bootstrap>:1207(_handle_fromlist)
       27    0.000    0.000    0.000    0.000 indexing.py:955(<genexpr>)
        2    0.000    0.000    0.000    0.000 concat.py:52(concat_compat)
       27    0.000    0.000    0.000    0.000 indexing.py:2685(<genexpr>)
       27    0.000    0.000    0.000    0.000 indexing.py:1143(<genexpr>)
        5    0.000    0.000    0.000    0.000 <frozen posixpath>:229(expanduser)
        6    0.000    0.000    0.000    0.000 generic.py:487(_validate_dtype)
        1    0.000    0.000    0.000    0.000 managers.py:1740(<listcomp>)
        1    0.000    0.000    0.000    0.000 series.py:3159(_append)
       25    0.000    0.000    0.000    0.000 managers.py:185(blklocs)
       34    0.000    0.000    0.000    0.000 flags.py:57(allows_duplicate_labels)
        1    0.000    0.000    0.000    0.000 {method 'take' of 'numpy.ndarray' objects}
       14    0.000    0.000    0.000    0.000 managers.py:2124(_grouping_func)
       34    0.000    0.000    0.000    0.000 generic.py:358(attrs)
       14    0.000    0.000    0.000    0.000 series.py:626(dtype)
        6    0.000    0.000    0.000    0.000 {built-in method posix.getpid}
        1    0.000    0.000    0.000    0.000 format.py:487(get_dataframe_repr_params)
       42    0.000    0.000    0.000    0.000 {method 'lower' of 'str' objects}
       11    0.000    0.000    0.000    0.000 common.py:306(is_null_slice)
        1    0.000    0.000    0.000    0.000 console.py:9(get_console_size)
        1    0.000    0.000    0.000    0.000 {method 'argsort' of 'numpy.ndarray' objects}
       29    0.000    0.000    0.000    0.000 managers.py:1902(_block)
       13    0.000    0.000    0.000    0.000 frame.py:949(axes)
        1    0.000    0.000    0.016    0.016 construction.py:506(nested_data_to_arrays)
       11    0.000    0.000    0.000    0.000 indexing.py:150(iloc)
       10    0.000    0.000    0.000    0.000 format.py:425(__init__)
       15    0.000    0.000    0.000    0.000 base.py:831(_reset_identity)
        4    0.000    0.000    0.000    0.000 common.py:233(stringify_path)
        5    0.000    0.000    0.000    0.000 base.py:592(_dtype_to_subclass)
       27    0.000    0.000    0.000    0.000 indexing.py:2694(<genexpr>)
        2    0.000    0.000    0.000    0.000 base.py:7592(trim_front)
        7    0.000    0.000    0.000    0.000 _dtype.py:330(_name_includes_bit_suffix)
        2    0.000    0.000    0.000    0.000 construction.py:765(_try_cast)
        1    0.000    0.000    0.000    0.000 format.py:1163(save_to_buffer)
        2    0.000    0.000    0.000    0.000 managers.py:1734(_consolidate_check)
        7    0.000    0.000    0.005    0.001 _json.py:1429(<lambda>)
       14    0.000    0.000    0.000    0.000 series.py:791(array)
        2    0.000    0.000    0.002    0.001 managers.py:1744(_consolidate_inplace)
        9    0.000    0.000    0.000    0.000 concat.py:587(<genexpr>)
        2    0.000    0.000    0.000    0.000 frozen.py:73(__getitem__)
        1    0.000    0.000    0.000    0.000 nanops.py:76(_f)
        1    0.000    0.000    0.000    0.000 array_ops.py:191(_na_arithmetic_op)
       35    0.000    0.000    0.000    0.000 multiarray.py:1079(copyto)
        2    0.000    0.000    0.000    0.000 concat.py:713(_get_concat_axis)
        9    0.000    0.000    0.000    0.000 common.py:556(require_length_match)
       16    0.000    0.000    0.000    0.000 common.py:123(<lambda>)
        1    0.000    0.000    0.000    0.000 string.py:189(_binify)
        1    0.000    0.000    0.000    0.000 format.py:683(_initialize_justify)
       21    0.000    0.000    0.000    0.000 common.py:1255(is_1d_only_ea_dtype)
        6    0.000    0.000    0.000    0.000 iostream.py:505(_is_master_process)
        2    0.000    0.000    0.000    0.000 indexing.py:1718(_get_slice_axis)
        1    0.000    0.000    0.001    0.001 format.py:789(truncate)
        1    0.000    0.000    0.005    0.005 string.py:40(_get_string_representation)
       24    0.000    0.000    0.000    0.000 blocks.py:583(dtype)
       18    0.000    0.000    0.000    0.000 indexing.py:1627(<genexpr>)
       12    0.000    0.000    0.000    0.000 format.py:633(is_truncated_horizontally)
        1    0.000    0.000    0.000    0.000 nanops.py:604(nansum)
        7    0.000    0.000    0.000    0.000 _json.py:1462(<lambda>)
        8    0.000    0.000    0.000    0.000 common.py:1366(_is_dtype)
        1    0.000    0.000    0.005    0.005 string.py:28(to_string)
        3    0.000    0.000    0.000    0.000 frame.py:641(_constructor_from_mgr)
        3    0.000    0.000    0.000    0.000 locale.py:396(normalize)
       14    0.000    0.000    0.000    0.000 {built-in method builtins.setattr}
       28    0.000    0.000    0.000    0.000 {pandas._libs.lib.is_iterator}
        1    0.000    0.000    0.000    0.000 format.py:1023(__init__)
        1    0.000    0.000    0.000    0.000 generic.py:12031(_min_count_stat_function)
       27    0.000    0.000    0.000    0.000 indexing.py:915(<genexpr>)
       14    0.000    0.000    0.000    0.000 construction.py:754(_maybe_repeat)
        2    0.000    0.000    0.000    0.000 {method 'all' of 'numpy.ndarray' objects}
        1    0.000    0.000    0.000    0.000 format.py:949(<listcomp>)
       17    0.000    0.000    0.000    0.000 blocks.py:1003(shape)
       15    0.000    0.000    0.000    0.000 generic.py:659(ndim)
       16    0.000    0.000    0.000    0.000 common.py:121(classes)
        2    0.000    0.000    0.000    0.000 _methods.py:61(_all)
       14    0.000    0.000    0.000    0.000 utils.py:62(is_list_like_indexer)
        2    0.000    0.000    0.000    0.000 concat.py:695(new_axes)
        9    0.000    0.000    0.000    0.000 indexing.py:909(_expand_ellipsis)
        1    0.000    0.000    0.000    0.000 _parser.py:199(split)
        5    0.000    0.000    0.000    0.000 printing.py:48(<listcomp>)
        6    0.000    0.000    0.000    0.000 fromnumeric.py:70(<dictcomp>)
        1    0.000    0.000    0.000    0.000 fromnumeric.py:51(_wrapfunc)
        1    0.000    0.000    0.000    0.000 range.py:341(nbytes)
        2    0.000    0.000    0.000    0.000 common.py:557(condition)
        1    0.000    0.000    0.000    0.000 managers.py:918(_verify_integrity)
        1    0.000    0.000    0.000    0.000 console.py:79(in_ipython_frontend)
        8    0.000    0.000    0.000    0.000 {method 'max' of 'numpy.ndarray' objects}
        2    0.000    0.000    0.000    0.000 missing.py:466(array_equivalent)
        1    0.000    0.000    0.000    0.000 generic.py:6337(dtypes)
        2    0.000    0.000    0.000    0.000 concat.py:698(<listcomp>)
        6    0.000    0.000    0.000    0.000 enum.py:193(__get__)
        1    0.000    0.000    0.000    0.000 string.py:67(_insert_dot_separators)
        1    0.000    0.000    0.000    0.000 base.py:674(_with_infer)
       15    0.000    0.000    0.000    0.000 base.py:71(<genexpr>)
        2    0.000    0.000    0.000    0.000 common.py:277(is_fsspec_url)
        3    0.000    0.000    0.000    0.000 format.py:1613(_format_strings)
        3    0.000    0.000    0.000    0.000 base.py:1396(format)
        2    0.000    0.000    0.000    0.000 common.py:145(is_url)
        2    0.000    0.000    0.000    0.000 missing.py:564(_array_equivalent_object)
        2    0.000    0.000    0.000    0.000 concat.py:565(<listcomp>)
        2    0.000    0.000    0.000    0.000 <string>:1(<lambda>)
        3    0.000    0.000    0.000    0.000 frame.py:4399(_clear_item_cache)
        1    0.000    0.000    0.001    0.001 _json.py:1432(_try_convert_dates)
        1    0.000    0.000    0.000    0.000 construction.py:1068(<listcomp>)
       18    0.000    0.000    0.000    0.000 {method 'search' of 're.Pattern' objects}
       14    0.000    0.000    0.000    0.000 format.py:877(<genexpr>)
        4    0.000    0.000    0.000    0.000 missing.py:642(na_value_for_dtype)
        1    0.000    0.000    0.000    0.000 managers.py:278(get_dtypes)
        2    0.000    0.000    0.000    0.000 range.py:996(_getitem_slice)
        5    0.000    0.000    0.000    0.000 format.py:2024(_has_names)
        4    0.000    0.000    0.000    0.000 base.py:1751(_get_names)
        1    0.000    0.000    0.000    0.000 common.py:62(new_method)
       26    0.000    0.000    0.000    0.000 base.py:7606(<genexpr>)
        1    0.000    0.000    0.005    0.005 string.py:34(_get_strcols)
        1    0.000    0.000    0.000    0.000 series.py:6094(_reduce)
       44    0.000    0.000    0.000    0.000 typing.py:2256(cast)
        3    0.000    0.000    0.000    0.000 blocks.py:265(make_block_same_class)
        3    0.000    0.000    0.000    0.000 locale.py:593(getlocale)
        2    0.000    0.000    0.000    0.000 parse.py:119(_coerce_args)
        6    0.000    0.000    0.000    0.000 {built-in method builtins.sum}
       25    0.000    0.000    0.000    0.000 {built-in method builtins.callable}
        1    0.000    0.000    0.005    0.005 format.py:611(get_strcols)
        2    0.000    0.000    0.000    0.000 format.py:974(<listcomp>)
        2    0.000    0.000    0.000    0.000 concat.py:543(_get_sample_object)
        7    0.000    0.000    0.000    0.000 series.py:784(_references)
       19    0.000    0.000    0.000    0.000 base.py:1657(name)
        6    0.000    0.000    0.000    0.000 blocks.py:203(_can_hold_na)
        6    0.000    0.000    0.000    0.000 __init__.py:225(compile)
        4    0.000    0.000    0.000    0.000 base.py:448(size)
        1    0.000    0.000    0.000    0.000 construction.py:1006(convert_object_array)
        1    0.000    0.000    0.000    0.000 api.py:106(_get_distinct_objs)
        1    0.000    0.000    0.000    0.000 inference.py:404(is_dataclass)
       10    0.000    0.000    0.000    0.000 _parser.py:203(isword)
        7    0.000    0.000    0.000    0.000 inspect.py:292(isclass)
        1    0.000    0.000    0.000    0.000 base.py:1795(set_names)
        1    0.000    0.000    0.000    0.000 managers.py:279(<listcomp>)
        1    0.000    0.000    0.000    0.000 threading.py:1185(is_alive)
        3    0.000    0.000    0.000    0.000 {built-in method _locale.setlocale}
        6    0.000    0.000    0.000    0.000 generic.py:6182(<genexpr>)
        3    0.000    0.000    0.000    0.000 format.py:645(has_index_names)
        1    0.000    0.000    0.001    0.001 shape_base.py:223(vstack)
        1    0.000    0.000    0.000    0.000 construction.py:481(<listcomp>)
        3    0.000    0.000    0.000    0.000 inference.py:105(is_file_like)
        2    0.000    0.000    0.000    0.000 concat.py:773(_concat_indexes)
        6    0.000    0.000    0.000    0.000 base.py:6625(_validate_indexer)
        3    0.000    0.000    0.000    0.000 frame.py:966(shape)
        3    0.000    0.000    0.000    0.000 format.py:653(show_row_idx_names)
        2    0.000    0.000    0.000    0.000 managers.py:1726(is_consolidated)
        1    0.000    0.000    0.000    0.000 arraylike.py:54(__gt__)
        1    0.000    0.000    0.000    0.000 _json.py:1126(__init__)
        1    0.000    0.000    0.000    0.000 config.py:469(__init__)
        1    0.000    0.000    0.000    0.000 nanops.py:389(new_func)
        1    0.000    0.000    0.000    0.000 contextlib.py:104(__init__)
        7    0.000    0.000    0.000    0.000 _dtype.py:24(_kind_name)
        2    0.000    0.000    0.000    0.000 base.py:773(_view)
        6    0.000    0.000    0.000    0.000 iostream.py:532(_schedule_flush)
        3    0.000    0.000    0.000    0.000 _strptime.py:26(_getlang)
        1    0.000    0.000    0.000    0.000 construction.py:950(_validate_or_indexify_columns)
        1    0.000    0.000    0.000    0.000 base.py:1243(copy)
        1    0.000    0.000    0.000    0.000 base.py:5458(_concat)
        9    0.000    0.000    0.000    0.000 common.py:126(_classes_and_not_datetimelike)
        4    0.000    0.000    0.000    0.000 nanops.py:79(<genexpr>)
        1    0.000    0.000    0.000    0.000 shape_base.py:81(atleast_2d)
        1    0.000    0.000    0.000    0.000 _parser.py:221(__init__)
        1    0.000    0.000    0.000    0.000 {built-in method builtins.sorted}
        4    0.000    0.000    0.000    0.000 {pandas._libs.lib.maybe_indices_to_slice}
        1    0.000    0.000    0.001    0.001 <frozen genericpath>:16(exists)
        1    0.000    0.000    0.000    0.000 api.py:120(_get_combined_index)
        4    0.000    0.000    0.000    0.000 base.py:675(empty)
        1    0.000    0.000    0.000    0.000 config.py:477(__enter__)
        2    0.000    0.000    0.000    0.000 _json.py:1105(__exit__)
        4    0.000    0.000    0.000    0.000 generic.py:568(_get_block_manager_axis)
        2    0.000    0.000    0.000    0.000 base.py:4190(_validate_positional_slice)
        1    0.000    0.000    0.000    0.000 range.py:352(memory_usage)
        1    0.000    0.000    0.000    0.000 function.py:411(validate_func)
        8    0.000    0.000    0.000    0.000 common.py:1390(_get_dtype)
       15    0.000    0.000    0.000    0.000 blocks.py:239(mgr_locs)
        9    0.000    0.000    0.000    0.000 contextlib.py:428(__init__)
        1    0.000    0.000    0.000    0.000 string.py:126(<listcomp>)
        8    0.000    0.000    0.000    0.000 _methods.py:39(_amax)
        1    0.000    0.000    0.000    0.000 range.py:1030(_cmp_method)
        1    0.000    0.000    0.000    0.000 _parser.py:395(__init__)
        9    0.000    0.000    0.000    0.000 series.py:577(_constructor)
        2    0.000    0.000    0.000    0.000 config.py:215(get_default_val)
        1    0.000    0.000    0.000    0.000 api.py:72(get_objs_combined_axis)
       13    0.000    0.000    0.000    0.000 {method 'pop' of 'dict' objects}
        1    0.000    0.000    0.000    0.000 concat.py:202(_maybe_reindex_columns_na_proxy)
        3    0.000    0.000    0.000    0.000 locale.py:479(_parse_localename)
        2    0.000    0.000    0.000    0.000 base.py:5453(<setcomp>)
        2    0.000    0.000    0.000    0.000 construction.py:196(mgr_to_mgr)
        9    0.000    0.000    0.000    0.000 contextlib.py:434(__exit__)
        4    0.000    0.000    0.000    0.000 construction.py:687(_sanitize_non_ordered)
        3    0.000    0.000    0.000    0.000 _parser.py:189(__next__)
        9    0.000    0.000    0.000    0.000 concat.py:584(<genexpr>)
        1    0.000    0.000    0.000    0.000 construction.py:532(treat_as_nested)
        2    0.000    0.000    0.000    0.000 format.py:657(show_col_idx_names)
        3    0.000    0.000    0.000    0.000 base.py:791(is_)
        1    0.000    0.000    0.000    0.000 format.py:751(_adjust_max_rows)
        1    0.000    0.000    0.000    0.000 generic.py:12070(sum)
        3    0.000    0.000    0.000    0.000 _strptime.py:565(_strptime_datetime)
        4    0.000    0.000    0.000    0.000 {built-in method sys.getsizeof}
        2    0.000    0.000    0.000    0.000 format.py:629(is_truncated)
        1    0.000    0.000    0.000    0.000 base.py:1754(_set_names)
        1    0.000    0.000    0.000    0.000 <__array_function__ internals>:177(array_equal)
        4    0.000    0.000    0.000    0.000 format.py:637(is_truncated_vertically)
        4    0.000    0.000    0.000    0.000 range.py:347(<genexpr>)
        4    0.000    0.000    0.000    0.000 managers.py:920(<genexpr>)
        2    0.000    0.000    0.000    0.000 common.py:521(is_string_dtype)
        1    0.000    0.000    0.001    0.001 common.py:1141(file_exists)
        2    0.000    0.000    0.000    0.000 base.py:7607(<listcomp>)
        2    0.000    0.000    0.000    0.000 generic.py:4314(_set_is_copy)
        1    0.000    0.000    0.000    0.000 base.py:5153(_get_engine_target)
        5    0.000    0.000    0.000    0.000 managers.py:896(__init__)
        1    0.000    0.000    0.000    0.000 concat.py:703(_get_comb_axis)
       11    0.000    0.000    0.000    0.000 {method 'items' of 'dict' objects}
        1    0.000    0.000    0.000    0.000 fromnumeric.py:1038(argsort)
        4    0.000    0.000    0.000    0.000 blocks.py:1016(_slice)
        1    0.000    0.000    0.000    0.000 contextlib.py:132(__enter__)
        2    0.000    0.000    0.000    0.000 format.py:649(has_column_names)
        1    0.000    0.000    0.000    0.000 expressions.py:226(evaluate)
        1    0.000    0.000    0.000    0.000 common.py:977(is_numeric_v_string_like)
        1    0.000    0.000    0.000    0.000 config.py:483(__exit__)
        3    0.000    0.000    0.000    0.000 generic.py:2073(<genexpr>)
        2    0.000    0.000    0.000    0.000 base.py:782(_rename)
        1    0.000    0.000    0.000    0.000 threading.py:1118(_wait_for_tstate_lock)
        3    0.000    0.000    0.000    0.000 nanops.py:72(check)
        1    0.000    0.000    0.000    0.000 missing.py:131(dispatch_fill_zeros)
        6    0.000    0.000    0.000    0.000 enum.py:1249(value)
        4    0.000    0.000    0.000    0.000 datetimes.py:156(should_cache)
        1    0.000    0.000    0.000    0.000 expressions.py:67(_evaluate_standard)
        2    0.000    0.000    0.000    0.000 format.py:1179(get_buffer)
        3    0.000    0.000    0.000    0.000 blocks.py:192(_can_consolidate)
        1    0.000    0.000    0.000    0.000 iostream.py:127(_event_pipe)
        1    0.000    0.000    0.000    0.000 range.py:484(_view)
        4    0.000    0.000    0.000    0.000 generic.py:6177(<genexpr>)
        1    0.000    0.000    0.000    0.000 <frozen codecs>:309(__init__)
        1    0.000    0.000    0.000    0.000 {method 'sum' of 'numpy.ndarray' objects}
        1    0.000    0.000    0.000    0.000 managers.py:2233(<listcomp>)
       14    0.000    0.000    0.000    0.000 base.py:6612(_maybe_cast_indexer)
        9    0.000    0.000    0.000    0.000 range.py:377(dtype)
        1    0.000    0.000    0.000    0.000 common.py:85(consensus_name_attr)
        1    0.000    0.000    0.000    0.000 _validators.py:450(check_dtype_backend)
        1    0.000    0.000    0.000    0.000 base.py:459(_engine_type)
        1    0.000    0.000    0.000    0.000 nanops.py:455(newfunc)
        1    0.000    0.000    0.001    0.001 <__array_function__ internals>:177(vstack)
        6    0.000    0.000    0.000    0.000 common.py:1107(<lambda>)
        1    0.000    0.000    0.000    0.000 {built-in method _codecs.lookup}
        1    0.000    0.000    0.000    0.000 console.py:54(in_interactive_session)
        1    0.000    0.000    0.000    0.000 contextlib.py:287(helper)
        1    0.000    0.000    0.000    0.000 common.py:80(ensure_str)
        1    0.000    0.000    0.000    0.000 _parser.py:322(weekday)
        1    0.000    0.000    0.000    0.000 {method 'any' of 'numpy.ndarray' objects}
        2    0.000    0.000    0.000    0.000 concat.py:689(_get_result_dim)
        9    0.000    0.000    0.000    0.000 {pandas._libs.lib.item_from_zerodim}
        7    0.000    0.000    0.000    0.000 managers.py:335(<dictcomp>)
        4    0.000    0.000    0.000    0.000 config.py:663(_get_registered_option)
        6    0.000    0.000    0.000    0.000 concat.py:351(__init__)
        4    0.000    0.000    0.000    0.000 {method 'upper' of 'str' objects}
       12    0.000    0.000    0.000    0.000 base.py:363(ndim)
        1    0.000    0.000    0.000    0.000 string.py:22(__init__)
       11    0.000    0.000    0.000    0.000 {method 'read' of '_io.StringIO' objects}
        7    0.000    0.000    0.000    0.000 {method 'write' of '_io.StringIO' objects}
        3    0.000    0.000    0.000    0.000 concat.py:167(<listcomp>)
        1    0.000    0.000    0.000    0.000 <string>:2(__init__)
        1    0.000    0.000    0.000    0.000 series.py:6195(sum)
        1    0.000    0.000    0.000    0.000 <__array_function__ internals>:177(argsort)
        2    0.000    0.000    0.000    0.000 {built-in method builtins.next}
        1    0.000    0.000    0.000    0.000 nanops.py:253(_get_values)
        1    0.000    0.000    0.000    0.000 base.py:1900(rename)
        3    0.000    0.000    0.000    0.000 range.py:281(start)
        2    0.000    0.000    0.000    0.000 base.py:346(shape)
        1    0.000    0.000    0.000    0.000 nanops.py:324(_get_dtype_max)
        1    0.000    0.000    0.000    0.000 {method 'fill' of 'numpy.ndarray' objects}
        1    0.000    0.000    0.000    0.000 concat.py:303(<listcomp>)
        2    0.000    0.000    0.000    0.000 generic.py:638(_info_axis)
        1    0.000    0.000    0.000    0.000 contextlib.py:141(__exit__)
        1    0.000    0.000    0.000    0.000 format.py:641(dimensions_info)
        1    0.000    0.000    0.000    0.000 generic.py:2015(empty)
        1    0.000    0.000    0.000    0.000 base.py:5462(<listcomp>)
        1    0.000    0.000    0.000    0.000 concat.py:747(<listcomp>)
        1    0.000    0.000    0.000    0.000 dataclasses.py:1256(is_dataclass)
        4    0.000    0.000    0.000    0.000 {pandas._libs.algos.ensure_object}
        1    0.000    0.000    0.000    0.000 _parser.py:208(isnum)
        1    0.000    0.000    0.000    0.000 common.py:1107(_maybe_memory_map)
        1    0.000    0.000    0.000    0.000 api.py:102(<listcomp>)
        6    0.000    0.000    0.000    0.000 fromnumeric.py:2427(_all_dispatcher)
        8    0.000    0.000    0.000    0.000 multiarray.py:152(concatenate)
       10    0.000    0.000    0.000    0.000 {method 'isalpha' of 'str' objects}
        2    0.000    0.000    0.000    0.000 construction.py:916(<genexpr>)
        1    0.000    0.000    0.000    0.000 frame.py:1114(_info_repr)
        1    0.000    0.000    0.000    0.000 {method 'acquire' of '_thread.lock' objects}
        2    0.000    0.000    0.000    0.000 config.py:897(is_nonnegative_int)
        1    0.000    0.000    0.000    0.000 format.py:732(_calc_max_rows_fitted)
        6    0.000    0.000    0.000    0.000 common.py:175(<genexpr>)
        2    0.000    0.000    0.000    0.000 common.py:171(not_none)
        1    0.000    0.000    0.000    0.000 shape_base.py:218(_vhstack_dispatcher)
        1    0.000    0.000    0.000    0.000 generic.py:1948(__iter__)
        1    0.000    0.000    0.000    0.000 <__array_function__ internals>:177(atleast_2d)
        1    0.000    0.000    0.000    0.000 range.py:946(<listcomp>)
        6    0.000    0.000    0.000    0.000 {method '__exit__' of '_thread.RLock' objects}
        3    0.000    0.000    0.000    0.000 {method 'clear' of 'dict' objects}
        1    0.000    0.000    0.000    0.000 _parser.py:342(ampm)
        2    0.000    0.000    0.000    0.000 series.py:3169(<genexpr>)
        1    0.000    0.000    0.000    0.000 expressions.py:76(_can_use_numexpr)
        1    0.000    0.000    0.000    0.000 common.py:1025(needs_i8_conversion)
        1    0.000    0.000    0.000    0.000 format.py:721(_calc_max_cols_fitted)
        9    0.000    0.000    0.000    0.000 contextlib.py:431(__enter__)
        3    0.000    0.000    0.000    0.000 range.py:316(step)
        1    0.000    0.000    0.000    0.000 _methods.py:55(_any)
        6    0.000    0.000    0.000    0.000 {pandas._libs.lib.is_int_or_none}
        7    0.000    0.000    0.000    0.000 _json.py:1401(<lambda>)
        1    0.000    0.000    0.000    0.000 config.py:478(<listcomp>)
        1    0.000    0.000    0.000    0.000 concat.py:631(<listcomp>)
        1    0.000    0.000    0.000    0.000 _parser.py:329(month)
        3    0.000    0.000    0.000    0.000 range.py:911(<genexpr>)
        1    0.000    0.000    0.000    0.000 format.py:623(should_show_dimensions)
        2    0.000    0.000    0.000    0.000 base.py:539(<genexpr>)
        7    0.000    0.000    0.000    0.000 series.py:1381(_clear_item_cache)
        1    0.000    0.000    0.000    0.000 {method 'disable' of '_lsprof.Profiler' objects}
        2    0.000    0.000    0.000    0.000 format.py:765(_is_in_terminal)
        1    0.000    0.000    0.000    0.000 _parser.py:319(jump)
        5    0.000    0.000    0.000    0.000 {method 'add' of 'set' objects}
        3    0.000    0.000    0.000    0.000 frame.py:637(_constructor)
        2    0.000    0.000    0.000    0.000 concat.py:73(<listcomp>)
        3    0.000    0.000    0.000    0.000 {method '__exit__' of '_thread.lock' objects}
        1    0.000    0.000    0.000    0.000 concat.py:720(<listcomp>)
        1    0.000    0.000    0.000    0.000 _parser.py:213(isspace)
        4    0.000    0.000    0.000    0.000 {pandas._libs.lib.is_bool}
        1    0.000    0.000    0.000    0.000 format.py:689(_initialize_columns)
        1    0.000    0.000    0.000    0.000 {built-in method pandas._libs.lib.is_interval}
        5    0.000    0.000    0.000    0.000 {built-in method posix.fspath}
        2    0.000    0.000    0.000    0.000 {pandas._libs.lib.dtypes_all_equal}
        1    0.000    0.000    0.000    0.000 common.py:503(get_compression_method)
        1    0.000    0.000    0.000    0.000 {built-in method builtins.min}
        1    0.000    0.000    0.000    0.000 _methods.py:47(_sum)
        1    0.000    0.000    0.000    0.000 dispatch.py:17(should_extension_dispatch)
        1    0.000    0.000    0.000    0.000 {method 'getvalue' of '_io.StringIO' objects}
        1    0.000    0.000    0.000    0.000 format.py:665(_initialize_sparsify)
        1    0.000    0.000    0.000    0.000 range.py:922(<listcomp>)
        1    0.000    0.000    0.000    0.000 threading.py:568(is_set)
        1    0.000    0.000    0.000    0.000 managers.py:536(nblocks)
        1    0.000    0.000    0.000    0.000 shape_base.py:207(_arrays_for_stack_dispatcher)
        1    0.000    0.000    0.000    0.000 inference.py:306(is_named_tuple)
        1    0.000    0.000    0.000    0.000 string.py:63(_need_to_wrap_around)
        1    0.000    0.000    0.000    0.000 _validators.py:226(validate_bool_kwarg)
        2    0.000    0.000    0.000    0.000 base.py:974(dtype)
        3    0.000    0.000    0.000    0.000 range.py:299(stop)
        2    0.000    0.000    0.000    0.000 managers.py:235(items)
        1    0.000    0.000    0.000    0.000 {built-in method _codecs.lookup_error}
        2    0.000    0.000    0.000    0.000 indexing.py:2665(need_slice)
        3    0.000    0.000    0.000    0.000 {built-in method builtins.id}
        2    0.000    0.000    0.000    0.000 format.py:959(<dictcomp>)
        1    0.000    0.000    0.000    0.000 managers.py:2242(<listcomp>)
        2    0.000    0.000    0.000    0.000 concat.py:766(_maybe_check_integrity)
        1    0.000    0.000    0.000    0.000 <frozen codecs>:260(__init__)
        1    0.000    0.000    0.000    0.000 nanops.py:209(_maybe_get_mask)
        1    0.000    0.000    0.000    0.000 {method 'pop' of 'list' objects}
        1    0.000    0.000    0.000    0.000 base.py:2756(_is_multi)
        1    0.000    0.000    0.000    0.000 function.py:64(__call__)
        1    0.000    0.000    0.000    0.000 format.py:697(_initialize_colspace)
        1    0.000    0.000    0.000    0.000 {method 'append' of 'collections.deque' objects}
        1    0.000    0.000    0.000    0.000 {method 'isdigit' of 'str' objects}
        2    0.000    0.000    0.000    0.000 {built-in method numpy.asanyarray}
        1    0.000    0.000    0.000    0.000 {method 'isspace' of 'str' objects}
        1    0.000    0.000    0.000    0.000 nanops.py:1491(_maybe_null_out)
        1    0.000    0.000    0.000    0.000 fromnumeric.py:1034(_argsort_dispatcher)
        1    0.000    0.000    0.000    0.000 _parser.py:1056(_could_be_tzname)
        1    0.000    0.000    0.000    0.000 numeric.py:2403(_array_equal_dispatcher)
        1    0.000    0.000    0.000    0.000 {method 'values' of 'dict' objects}
        2    0.000    0.000    0.000    0.000 {function FrozenList.__getitem__ at 0x7f0665c34860}
        1    0.000    0.000    0.000    0.000 _parser.py:186(__iter__)
        2    0.000    0.000    0.000    0.000 parse.py:108(_noop)
        1    0.000    0.000    0.000    0.000 {method 'reverse' of 'list' objects}
        2    0.000    0.000    0.000    0.000 _json.py:1102(__enter__)
        2    0.000    0.000    0.000    0.000 base.py:1954(nlevels)
        1    0.000    0.000    0.000    0.000 interactiveshell.py:637(get_ipython)
        1    0.000    0.000    0.000    0.000 range.py:228(_constructor)
        1    0.000    0.000    0.000    0.000 format.py:670(_initialize_formatters)
        1    0.000    0.000    0.000    0.000 {pandas._libs.lib.is_period}
        1    0.000    0.000    0.000    0.000 shape_base.py:77(_atleast_2d_dispatcher)

Results

Polars and Pandas both processed the same data (8000 rows, categorical data represented as strings).

Versions

  • Pandas: 2.1.4
  • Polars: 0.20.26

Memory usage comparison

File on disk: 6,0 MB (du -sh), 8000 rows, 7 columns.

  • Polars: 4,76 MB
  • Pandas: 7,56 MB

-> Polars was more memory efficient: ~ 1,6 times less memory

Profile comparison

  • Polars: 256 function calls (253 primitive calls) in 0.020 seconds
  • Pandas: 46681 function calls (46472 primitive calls) in 0.113 seconds

-> Polars was ~ 5,6 times faster and needed ~ 180x less function and primitive calls.

Conclusion

Polars should be used whenever possible.

In [ ]: