Chunking¶
When your input consists of many independent pieces, such as records, log lines,
or top-level definitions, you can split it into Chunks
and parse them in parallel with
walk_parallel. The generated <Grammar>EventListener
provides chunking classmethods that produce those chunks, each carrying its
exact source position so callbacks still report positions against the whole source.
The token-based and rule-based chunkers use the facade's lexer and parser and
its token-type constants, so you call them on the class without passing a lexer or
parser.
There are four kinds of chunker. The first three trade correctness against speed; the fourth, streaming, bounds memory use. Choose by your constraints rather than by speed alone, because the splitting cost is usually small compared with the per-chunk parse that follows.
Token-based — split at lexer tokens¶
A single lexer pass in C++, with no parser, finds the boundaries, so a delimiter inside a string or comment never causes a split. The splitter asks the lexer for the boundary tokens only, so little data crosses into Python.
# each chunk begins with a delimiter token (one type, or several):
chunks = MyGrammarEventListener.split_on_token(
text, MyGrammarEventListener.RECORD, where="before"
)
# or one chunk per open..close region (optionally balanced):
chunks = MyGrammarEventListener.split_between_tokens(
text, (MyGrammarEventListener.BEGIN, MyGrammarEventListener.END), nested=True
)
split_on_token: split at each delimiter token.where="before"orwhere="after"puts the delimiter at the start or end of each chunk.split_between_tokens: one chunk per region between an opening and a closing token. With several bracket kinds, each opener matches its own closer, andnested=Truematches balanced pairs.lex: the underlying token pass without a parser, exposed directly in case you want the token stream.
Regex-based — split on a pattern¶
These chunkers use no lexer. They find delimiters about 7× faster (about 4× faster end to end), but they are not token-aware, so a match inside a string still causes a split. Prefer them when the delimiter text cannot appear anywhere else in the input.
split_on_pattern: the regex counterpart ofsplit_on_token; the pattern matches the delimiter.chunk_by_pattern: the pattern matches a whole record, so each match is a chunk.
Rule-based — split on grammar structure¶
chunk_by_rule parses the
input entirely in C++ and yields each occurrence of a grammar rule as a chunk, so it
splits on actual grammar structure rather than on a token or pattern heuristic. It
costs a full parse, but only the chunk spans cross into Python. It is worthwhile
when the per-chunk walk_parallel callback work dominates, or when no delimiter
clearly marks a record.
See Performance & limitations for measurements of the token, regex, and rule chunkers and the trade-off between speed and correctness.
Streaming — bounded memory¶
For input too large to hold in memory, the streaming chunkers read incrementally
and yield chunks without keeping the whole source. Combined with walk_parallel,
the whole pipeline runs in bounded memory.
stream_on_token: the streaming form ofsplit_on_token. The native layer opens the file and lexes it over a sliding window (UTF-8).stream_on_pattern: the streaming form ofsplit_on_pattern. The regex runs in Python, so it reads any text source (a path, an open file, or an iterable ofstr) in any encoding.stream_by_rule: the streaming form ofchunk_by_rule, for input that is a top-level sequence of records, each an occurrence of a grammarrule(or one of several). It parses one record at a time, feeding the lexer into the parser with bounded memory. Records must be directly adjacent, with only whitespace and comments that the lexer skips between them. When there are several candidate rules, the next token decides which one to parse, so the rules should start with different tokens (for exampleclassanddef). If an on-channel token does not begin any candidate rule (an unsupported header or separator), it raises a clear error naming the token and candidates rather than silently truncating. Unlikechunk_by_rule, it does not search a full parse for the rule at any depth. For arbitrarily nested or comma-separated records, usechunk_by_ruleorstream_on_token.
Source positions and names¶
Every chunk carries its start (offset, line, column) against the whole source, so
callbacks still report positions against the original after splitting. Each chunker
also accepts sourcename=, a file name for diagnostics that is recorded on every
chunk. The streaming chunkers default it to the file path. During the parse it is
available through
FacadeListener.sourcename(),
so a callback can report a position as sourcename:line:column.
By default each chunk is trimmed of surrounding whitespace and whitespace-only
regions are dropped. Pass trim=False to a delimiter chunker to keep every region
exactly as it appears. This preserves leading or trailing whitespace that the
grammar needs, such as a record's mandatory newline; only empty, zero-length
regions are dropped.
Leading and trailing regions (the preamble)¶
With where="before", any content before the first delimiter becomes its own
leading chunk; this is often a file's header or preamble. With where="after", any
content after the last delimiter becomes a trailing chunk. These chunks are not
marked in any way (each is simply the first or last chunk), so skip or handle them
explicitly. The
streaming-records recipe walks through preamble handling,
trim=False terminators, absolute positions, and recovering off-channel metadata end
to end.