Why KLR.gg set a 2-second processing budget per map
On our KLR.gg project, processing a match map must complete under 2 seconds. When a server running MatchZy finishes a match, our Python worker pulls raw CSV artifacts and demo logs over FTP, parses round events, and calculates HLTV Rating 2.0, Round Swing, and Glicko-2 skill distributions across 30 rounds and roughly 15,000 tick-level combat events. In benchmarks on 14 March 2024, pandas 2.2 with NumPy 1.26 backends parsed and scored maps in 1.84 s total end-to-end execution. Polars 1.0 took 2.45 s on the exact same stateful workload. Match telemetry arrives in sudden bursts when tournament rounds finish simultaneously across dozens of servers. Profiling showed that 72% of worker CPU time goes to stateful rating calculations rather than string ingestion. CSV parsing speed mattered less than downstream memory layout compatibility for our rating engine.
Delayed log processing caused stale UI states on our Next.js front end and pushed automated payout settlements past our 5-second queue SLA.
Arrow chunked memory killed our Numba loops
Apache Arrow memory layouts inside Polars group data into zero-copy, chunked columnar buffers ideal for parallel filtering and aggregations. CS2 match scoring works differently: round N updates depend on round N-1 player ratings and weapon economy states, requiring sequential state mutations. Passing Arrow chunked arrays to C-extensions or scalar Python loops forces Polars to handle chunk boundary checks or allocate new memory via to_numpy(). Standard pandas C-contiguous NumPy float64 arrays store contiguous blocks in memory, allowing direct memory pointer transfers into Numba-compiled functions.
Arrow multi-chunk layout means elements in a single column often reside in non-contiguous memory addresses. When calculating Glicko-2 rating deviations across 10 players over 30 rounds, memory access patterns dominate execution latency. Passing a Polars column to a Python loop requires either calling to_numpy() (which allocates new memory) or using Polars expression wrappers that introduce C-layer call overhead on every iteration.
380 ms penalty: Polars UDFs vs contiguous NumPy memory
Calculating Glicko-2 and Round Swing variance requires updating player variance parameters step-by-step across sequential match events. Polars handles native vectorized expressions efficiently, but custom User-Defined Functions (UDFs) fall back to Python execution or require custom Rust plugins. Moving intermediate data out of Rust into Python objects to update iterative ratings introduced a 340 ms per-map penalty in Polars 0.20 and 380 ms in Polars 1.0. Remaining in pandas with float64 NumPy arrays let us feed raw pointers directly into JIT-compiled @numba.njit loops without heap allocations.
On 22 March 2024, we attempted writing a custom Rust plugin for Polars to calculate Round Swing, but maintaining Rust bindings for rapid scoring rule iterations cost us three days of dev time without beating flat NumPy pointer iteration.
We tested three patterns: Polars native expressions with map_batches(), pure pandas vectorized methods, and pandas extract arrays passed to Numba. The Numba approach operating over flat NumPy buffers executed in 14 milliseconds per map, compared to 380 milliseconds for Polars map_batches(). Glicko-2 updating requires step-by-step matrix mutation based on opponent ratings, so standard relational joins and column mappings in Polars could not express the algorithm without falling back to iterative Python execution.
pandas 2.2 vs Polars 1.0 on CS2 match log ingestion
Framework choice depends on where CPU time is spent. Polars parses raw CSV files 3.6x faster than pandas due to multithreaded Rust readers (85 ms vs 310 ms on a 30-round log). However, initial CSV file parsing accounts for less than 15% of total map processing time in our pipeline. Pandas combined with C-contiguous NumPy arrays executes non-vectorizable rating loops in 14 ms, beating Polars map_batches() by 366 ms.
| Metric / Operation | pandas 2.2 (NumPy backend) | Polars 1.0 (Arrow backend) | Engine impact |
|---|---|---|---|
| MatchZy CSV ingestion | 310 ms | 85 ms | Polars Rust parser is 3.6x faster on raw I/O |
| Round grouping & filters | 120 ms | 40 ms | Polars expressions run multithreaded natively |
| Glicko-2 & Rating 2.0 computation | 14 ms (via Numba ptr) | 380 ms (via map_batches) | pandas contiguous block avoids Arrow conversion penalty |
| Memory buffer conversion overhead | 0 ms (zero-copy pointer) | 165 ms (to_numpy array allocation) | Polars requires array materialization for custom C-loops |
| Total pipeline latency per map | 1.84 s | 2.45 s | pandas/NumPy wins by 610 ms on end-to-end match compute |
14 milliseconds in Numba: the exact code we deployed
To achieve sub-2-second map ingestion, we separated structural grouping from mathematical score generation. Pandas handles file parsing and event filtering, while a JIT-compiled Numba kernel computes player impact, Round Swing, and Glicko-2 state updates directly on 1D NumPy arrays. Passing memory views of underlying pandas Series into compiled code eliminates Python object creation inside inner execution loops.
import numba as nb
import numpy as np
@nb.njit(fastmath=True)
def compute_round_swing(kills, deaths, damage, player_ids, out_rating):
n_events = kills.shape[0]
for i in range(n_events):
p_id = player_ids[i]
impact = (kills[i] * 2.13) - (deaths[i] * 0.98) + (damage[i] * 0.008)
out_rating[p_id] += impact * 0.15
# Zero-copy extraction from pandas block memory
kills_buf = df['kills'].to_numpy(dtype=np.float64, copy=False)
deaths_buf = df['deaths'].to_numpy(dtype=np.float64, copy=False)
damage_buf = df['damage'].to_numpy(dtype=np.float64, copy=False)
p_ids_buf = df['player_id'].to_numpy(dtype=np.int64, copy=False)
out_buf = np.zeros(10, dtype=np.float64)
compute_round_swing(kills_buf, deaths_buf, damage_buf, p_ids_buf, out_buf)Using copy=False inside to_numpy() guarantees that Numba receives direct access to the memory pointers of pandas DataFrame blocks. Polars stores columns as Arrow chunked arrays, so calling to_numpy() in Polars forces a memory copy whenever a column spans multiple Arrow chunks.
What broke when we picked pandas over Polars
Choosing pandas 2.2 over Polars meant sacrificing multi-core CPU scaling during raw CSV ingestion and taking on manual memory alignment duties. Polars automatically parallelizes filtering and CSV reading across CPU cores via its Rust engine, whereas pandas runs single-threaded per process. We resolved backfill ingestion bottlenecks by spawning process pools with Python's concurrent.futures, pinning each match map task to an isolated worker process.
We also hit higher peak memory consumption during complex DataFrame merges. Pandas creates memory copies during full DataFrame merges unless index alignment is strictly managed. For our pipeline, worker nodes configured with 4 GB RAM per core handle peak loads without swapping. Pipeline tasks that run pure vectorized aggregation over multi-gigabyte logs without custom C math extensions remain better suited for Polars.