First parser
Sigma gives you a set of small parsers and the functions to combine them. There is no separate grammar file and no code generation step. A parser is an ordinary value, and you build a bigger one by passing smaller ones to a function.
This page builds one up from a single literal to a parser for a small settings format, and explains what comes back at each step.
name = sigma
retries = 3
verbose = true
tags = [core, docs, bench]Running a parser
The smallest useful parser matches a literal. string builds one, and run feeds it some input.
const Parser = string('hello')run(Parser).with('hello'){
isOk: true,
start: 0,
end: 5,
pos: 5,
value: 'hello',
errors: []
}isOk tells you which shape you're holding, and it narrows the type. Inside an if (result.isOk) branch TypeScript knows value exists. start and end are offsets into the input describing the region the parser consumed. pos is where the cursor ended up. errors collects failures the parser recovered from, which the error recovery guide covers.
Nothing requires a parser to reach the end of the input. Given more to work with, it still stops after what it matched.
run(Parser).with('hello there'){
isOk: true,
start: 0,
end: 5,
pos: 5,
value: 'hello',
errors: []
}So a parser can report success on a file it only understood the first line of.
When a parser fails
run(Parser).with('help'){
isOk: false,
start: 0,
end: 4,
pos: 0,
expected: 'hello',
label: null,
errors: []
}expected is a short description of what would have matched, written so it reads well after the word "expected" in a message to a user. label stays null until something commits to a parse, which the error recovery guide covers.
The three offsets do different jobs on a failure, and the difference matters once you start rendering diagnostics. start and end are the region the failed parser highlights. The parser that failed decides how wide that region is. string highlights as many characters as its literal is long, clamped to the input, so here the region spans all four characters of help. Most other parsers highlight a single point. pos is where the failure gets reported, which is wherever the parser that gave up was standing. In a larger grammar that is usually somewhere in the middle of the input. The cursor is separate from all three. A parser that fails leaves the cursor where it found it, and that is what lets the next alternative try the same input.
Putting parsers in a row
sequence applies parsers one after another and collects the results into a tuple, typed per position.
const Setting = sequence(letters(), string('='), integer())run(Setting).with('retries=3'){
isOk: true,
start: 0,
end: 9,
pos: 9,
value: [ 'retries', '=', 3 ],
errors: []
}integer gives back a number rather than the digits it matched, so the tuple's type is [string, string, number].
The '=' in the middle carries no information. Rather than destructure it away later, use a selector combinator to drop it up front. outer runs three parsers and keeps the outer two.
const Setting = sequence(letters(), string('='), integer())
const Setting = outer(letters(), string('='), integer()) run(Setting).with('retries=3'){
isOk: true,
start: 0,
end: 9,
pos: 9,
value: [ 'retries', 3 ],
errors: []
}inner keeps the middle one and discards the delimiters around it, which is what you want for brackets. first and last do the same job for pairs.
Shaping the result
Tuples can get unreadable quickly. map applies a function to whatever a parser produced, and it hands that function a Span as a second argument, so a node can record where it came from.
const Setting = map(
outer(letters(), string('='), integer()),
([name, value], span) => ({ name, value, span }),
)run(Setting).with('retries=3'){
isOk: true,
start: 0,
end: 9,
pos: 9,
value: { name: 'retries', value: 3, span: { start: 0, end: 9 } },
errors: []
}Choosing between alternatives
A setting's value isn't always a number. choice takes several parsers and returns the result of the first one that matches, with the result type being the union of theirs.
const Flag = map(choice(string('true'), string('false')), (text) => ({
kind: 'flag' as const,
value: text === 'true',
}))
const Num = map(integer(), (value) => ({ kind: 'number' as const, value }))
const Text = map(letters(), (value) => ({ kind: 'text' as const, value }))
const Value = choice(Flag, Num, Text)choice is ordered. It takes the first alternative that succeeds and never reconsiders, even if a later one would have matched more input. Put Text first and true stops being a flag, because letters matches the whole word.
run(choice(Text, Flag, Num)).with('true'){
isOk: true,
start: 0,
end: 4,
pos: 4,
value: { kind: 'text', value: 'true' },
errors: []
}run(choice(Flag, Num, Text)).with('true'){
isOk: true,
start: 0,
end: 4,
pos: 4,
value: { kind: 'flag', value: true },
errors: []
}Neither run failed, so nothing points at the mistake. When two alternatives can match the same input, put the specific one first.
Setting can now take any of the three.
const Setting = map(
outer(letters(), string('='), integer()),
outer(letters(), string('='), Value),
([name, value], span) => ({ name, value, span }),
)Whitespace
So far the grammar only accepts retries=3 with no spaces. Sigma has no separate lexer, so you have to handle whitespace yourself. The usual approach is to pick a convention and apply it in one place, e.g. every token consumes the whitespace that follows it.
const ws = optional(whitespace())
function token<T>(parser: Parser<T>): Parser<T> {
return first(parser, ws)
}whitespace fails on an empty run, so optional wraps it to make the trailing space optional. first keeps the left result and discards the whitespace. With token in hand, the punctuation is defined once and the rules below never mention spacing again.
const Word = token(letters())
const Equals = token(string('='))
const Comma = token(string(','))
const Open = token(string('['))
const Close = string(']')const Setting = map(
outer(letters(), string('='), Value),
outer(Word, Equals, Value),
([name, value], span) => ({ name, value, span }),
)run(Setting).with('retries = 3'){
isOk: true,
start: 0,
end: 11,
pos: 11,
value: {
name: 'retries',
value: { kind: 'number', value: 3 },
span: { start: 0, end: 11 }
},
errors: []
}Close is not a token, so a setting's span ends at its last character instead of swallowing the newline after it. The whitespace between settings is handled further down, by wrapping the whole setting in token.
Lists
A bracketed list needs a value repeated with separators between them, which is sepBy. It returns an array, discards the separators, and succeeds on an empty run.
const List = map(inner(Open, sepBy(Word, Comma), Close), (value) => ({
kind: 'list' as const,
value,
}))
const Value = choice(Flag, Num, Text)
const Value = choice(List, Flag, Num, Text)
const Setting = map(
outer(Word, Equals, Value),
([name, value], span) => ({ name, value, span }),
) Parsers are values, so Setting holds the Value it was built from and has to be rebuilt to see the new one. The same applies to every redefinition further down the page.
run(Setting).with('tags = [core, docs, bench]'){
isOk: true,
start: 0,
end: 26,
pos: 26,
value: {
name: 'tags',
value: { kind: 'list', value: [ 'core', 'docs', 'bench' ] },
span: { start: 0, end: 26 }
},
errors: []
}run(Setting).with('tags = []'){
isOk: true,
start: 0,
end: 9,
pos: 9,
value: {
name: 'tags',
value: { kind: 'list', value: [] },
span: { start: 0, end: 9 }
},
errors: []
}Use sepBy1 instead when an empty list should be a syntax error rather than an empty array.
More than one setting
many applies a parser until it stops matching and collects the results. Zero matches gives an empty array rather than a failure.
run(many(token(Setting))).with('name = sigma\nretries = 3\n'){
isOk: true,
start: 0,
end: 25,
pos: 25,
value: [
{
name: 'name',
value: { kind: 'text', value: 'sigma' },
span: { start: 0, end: 12 }
},
{
name: 'retries',
value: { kind: 'number', value: 3 },
span: { start: 13, end: 24 }
}
],
errors: []
}The token around Setting is what consumes the newlines between them. Each setting's own span still stops at its last character.
The whole input
many stops at the first thing it can't parse, so a broken file can look fine.
run(inner(ws, many(token(Setting)), ws)).with('\nname = sigma\n???\n'){
isOk: true,
start: 0,
end: 14,
pos: 14,
value: [
{
name: 'name',
value: { kind: 'text', value: 'sigma' },
span: { start: 1, end: 13 }
}
],
errors: []
}The run above produced one setting and dropped the ??? without a word. eof fixes this. It matches only at the end of the input, so putting it last forces the parser to account for every character.
const Settings = inner(ws, many(token(Setting)), ws)
const Settings = inner(ws, many(token(Setting)), eof()) run(Settings).with('\nname = sigma\n???\n'){
isOk: false,
start: 14,
end: 14,
pos: 14,
expected: 'end of input',
label: null,
errors: []
}Now it's a failure, but a poor one. Position 14 is the start of ???, and "expected end of input" tells the user nothing about what they got wrong. The error recovery guide shows how to turn it into a real diagnostic and get the rest of the file parsed anyway.
Naming what you expected
Parse a setting with nothing after the =:
run(Setting).with('retries = '){
isOk: false,
start: 10,
end: 10,
pos: 10,
expected: '[',
label: null,
errors: []
}Every alternative in Value failed at position 10, and when alternatives tie like that choice reports the first one, so the message names a bracket for a value that was never going to be a list. error replaces the expectation with one you chose.
const Value = choice(List, Flag, Num, Text)
const Value = error(choice(List, Flag, Num, Text), 'value')
const Setting = map(
outer(Word, Equals, Value),
([name, value], span) => ({ name, value, span }),
) run(Setting).with('retries = '){
isOk: false,
start: 10,
end: 10,
pos: 10,
expected: 'value',
label: null,
errors: []
}Do this for anything a user will see. error leaves committed failures alone, so it relabels ordinary expectations without concealing real syntax errors.
When you'd rather throw
run returns a result for a failed parse instead of throwing, which suits code that wants to inspect the failure. When a failure is not something the caller can do anything about, tryRun returns the Success directly and throws a ParserError otherwise.
tryRun(Setting).with('retries = ')ParserError {
name: 'ParserError',
message: 'value',
span: { start: 10, end: 10 },
pos: 10,
label: null,
errors: []
}The message is the expectation, and the span, position and label are carried on the error. tryRun also throws when a parse succeeded but recovered from failures along the way, so a result you get back from it is always clean.
Where to go next
The settings format is flat on purpose: no value contains another setting, so every rule could be defined before it was used. Recursive grammars guide covers formats that nest, where a rule refers to itself and can't be written as a plain const.
Expressions guide covers arithmetic, which needs more than recursion because the obvious grammar for it loops forever.
And error recovery guide covers what to do when stopping at the first mistake isn't good enough.