Parallel parsing¶
When your input is many independent pieces that each parse on their own — records,
log entries, top-level definitions —
walk_parallel
parses them across a thread pool. The native parse releases the GIL, so the parses
overlap across cores; one listener is produced per chunk, in input order.
from my_listener import MyGrammarEventListener
chunks = MyGrammarEventListener.split_on_token(
text, MyGrammarEventListener.RECORD, where="before"
)
for listener in MyGrammarEventListener.walk_parallel(
chunks, start_rule="record"
):
... # one fully-walked listener per chunk, in order
Input: strings or positioned chunks¶
walk_parallel accepts an iterable of str or
Chunk:
- a bare
stris treated as contiguous with the previous item, and its source position is computed for you; - a
Chunkpins an explicitoffset/line/column(and optionalsourcename), so callbacks reportspan/line_colagainst the whole source rather than the chunk. AChunkalso re-anchors the bare strings that follow it.
The chunking helpers yield positioned Chunks, so they drop straight
in.
Bounded, lazy, in order¶
The result is a lazy iterator: chunks are pulled and parsed on demand with at
most max_workers parses in flight (also the ordering window), so neither the whole
input nor all results need to be held in memory — consume it incrementally, or
list(...) it if you want them all. Combined with a streaming chunker
(stream_on_token /
stream_on_pattern), the whole
pipeline — read, chunk, parse — stays bounded in the input size.
max_workers defaults to os.cpu_count(); pass 1 to run inline without a pool.
For an end-to-end pipeline over a file too large to hold in memory — streaming
chunker into walk_parallel, plus preamble handling, terminator preservation, and
recovering off-channel metadata — see the
streaming-records recipe.
How much speedup?¶
The per-event Python dispatch still holds the GIL, so the parallel speedup scales with how parse-heavy the work is relative to per-callback Python work. See the Parallel parsing section of Performance & limitations for the scaling details and measured numbers.