v0.19: Add std/markup/djot #1509

Open
jay wants to merge 17 commits from jay/gerbil:v0.19-djot into v0.19-staging
Member

Adds a native Djot parser and renderer under :std/text/markup/djot.

  • Typed syntax tree with optional source locations and :std/io input/output.
  • SXML and HTML rendering, including tables, footnotes, attributes, math, and raw content. Output is not sanitized.
  • Typed parser access and interface-based rendering.
  • Documentation, regression tests, and a native Gerbil benchmark. Also fixes SXML printing of literal symbol attribute names.

Benchmarks

Complete parsing into syntax trees, source locations disabled, no rendering. Gerbil compiled with gxc -O -exe, measured on an AMD Ryzen 9 7940HS at commit be13dea6.

Median milliseconds per parse across five timed batches. File reading and process startup are excluded; garbage collection during parsing is included.

Input Gerbil
README 0.381
Syntax reference 1.156
Tutorial 0.140
Large repeated README (~1 MB) 31.555
Plain text 3.029
Malformed markup 3.858

Warning

LLM-assisted with gpt-6-astra and gerbil-mcp.

Adds a native Djot parser and renderer under `:std/text/markup/djot`. - Typed syntax tree with optional source locations and `:std/io` input/output. - SXML and HTML rendering, including tables, footnotes, attributes, math, and raw content. Output is not sanitized. - Typed parser access and interface-based rendering. - Documentation, regression tests, and a native Gerbil benchmark. Also fixes SXML printing of literal symbol attribute names. ### Benchmarks Complete parsing into syntax trees, source locations disabled, no rendering. Gerbil compiled with `gxc -O -exe`, measured on an AMD Ryzen 9 7940HS at commit `be13dea6`. Median milliseconds per parse across five timed batches. File reading and process startup are excluded; garbage collection during parsing is included. | Input | Gerbil | |---|---:| | README | 0.381 | | Syntax reference | 1.156 | | Tutorial | 0.140 | | Large repeated README (~1 MB) | 31.555 | | Plain text | 3.029 | | Malformed markup | 3.858 | > [!WARNING] > LLM-assisted with `gpt-6-astra` and `gerbil-mcp`.
Emit symbol names rather than their Scheme write representation, which
can add bars around numeric or colon-ending attribute names. This
matches string attribute keys and existing element-name emission.
Keep all 302 pinned functional fixtures byte-for-byte unchanged. Adapt
only the three checkbox expectations to the existing HTML printer's
void-element spelling, without general normalization or skipped cases.
vyzo left a comment

just a readability comment: can you ask astra to make the code sparser so that it is easier to read?

just a readability comment: can you ask astra to make the code sparser so that it is easier to read?
Use compact slice maps and demand-grown token storage to reduce
allocation and repeated scanning. Preserve eager AST construction,
source coordinates, and resource checks.

Source-range lookup trades constant-time indexing for binary search
over slices; a located multiline recovery workload regresses 5.7%.
Author
Member

@vyzo wrote in #1509 (comment):

just a readability comment: can you ask astra to make the code sparser so that it is easier to read?

sure thing, done. still need to do some benchmarking and write up a description. you are welcome to look over the code though, it's basically done

@vyzo wrote in https://git.cons.io/mighty-gerbils/gerbil/pulls/1509#issuecomment-2960: > just a readability comment: can you ask astra to make the code sparser so that it is easier to read? sure thing, done. still need to do some benchmarking and write up a description. you are welcome to look over the code though, it's basically done
jay changed title from WIP: v0.19: Add std/markup/djot to v0.19: Add std/markup/djot 2026-09-27 20:31:19 +00:00
Author
Member

actually, now working on some more optimizations, i found this https://github.com/dcampbell24/djot-implementations/
and I see we are quite a bit behind the C parser in performance. unacceptable!!

Input Gerbil parse JS parse Gerbil → HTML C → HTML
README 0.378 ms 0.318 ms 2.558 ms 0.059 ms
Pandoc manual 11.909 ms 14.010 ms 85.021 ms 1.561 ms
Tartan 127.874 ms 107.001 ms 421.235 ms 11.613 ms
actually, now working on some more optimizations, i found this https://github.com/dcampbell24/djot-implementations/ and I see we are quite a bit behind the C parser in performance. unacceptable!! | Input | Gerbil parse | JS parse | Gerbil → HTML | C → HTML | |---|---:|---:|---:|---:| | README | 0.378 ms | 0.318 ms | 2.558 ms | 0.059 ms | | Pandoc manual | 11.909 ms | 14.010 ms | 85.021 ms | 1.561 ms | | Tartan | 127.874 ms | 107.001 ms | 421.235 ms | 11.613 ms |
Owner

ok, let me review and i can probably give some optimization advice. the gap to C is pretty big, but let's not throw safety out the window to close it. Also it provides a good benchmark we can use for the type-gerbil optimizer work.

ok, let me review and i can probably give some optimization advice. the gap to C is pretty big, but let's not throw safety out the window to close it. Also it provides a good benchmark we can use for the type-gerbil optimizer work.
Owner

so is the benchmark checked in? one obvious potential pitfall is using a port for the read; you should use directly a Reader, you can get one with call-with-file-reader from std/io.

so is the benchmark checked in? one obvious potential pitfall is using a port for the read; you should use directly a Reader, you can get one with `call-with-file-reader` from std/io.
Owner

we probably need an optimized writer for sxml, using a BufferedWriter instead of a port. maybe a bit out of scope for this pr, but definitely relevant to benchmarking.

we probably need an optimized writer for sxml, using a BufferedWriter instead of a port. maybe a bit out of scope for this pr, but definitely relevant to benchmarking.
Owner

S-expressions are not a very efficient data representation for HTML, either, if what you're looking for is performance.

S-expressions are not a very efficient data representation for HTML, either, if what you're looking for is performance.
vyzo left a comment

first pass review, mainly looking at how we can improve the code and its performance:

  • use dotted notation for all annotated types, and use using liberally. this results in compiler checkable and eliminable contracts and unchecked accessors.
  • avoid big linear scans on types with cond; use an interface and dispatch directly.
  • big linear case checks could benefit from using hash tables, but this is something to measure, it might not.
first pass review, mainly looking at how we can improve the code and its performance: - use dotted notation for all annotated types, and use using liberally. this results in compiler checkable and eliminable contracts and unchecked accessors. - avoid big linear scans on types with cond; use an interface and dispatch directly. - big linear case checks could benefit from using hash tables, but this is something to measure, it might not.
@ -0,0 +96,4 @@
(footnotes : :list))
transparent: #t)
(defstruct (djot-paragraph djot-container)
Owner

you should mark the leaves of the AST with final: #t. results in faster predicate and accessors/mutators.

you should mark the leaves of the AST with `final: #t`. results in faster predicate and accessors/mutators.
Owner

also note that transparent: #t is unnecessary, it is the default.

also note that transparent: #t is unnecessary, it is the default.
@ -0,0 +109,4 @@
(if (or (= p end) (memq (attribute-parser-state parser) '(done fail))) p
(let ((c (string-ref input p))
(state (attribute-parser-state parser)))
(case state
Owner

this could benefit from a hash table.

this could benefit from a hash table.
@ -0,0 +154,4 @@
(extra : :list := [])) => :list
(render-element tag node extra (children) newlines))
(cond
Owner

this is a big linear scan, i think it could benefit from an interface and dispatching the prerequisite methods. probably something to measure.

this is a big linear scan, i think it could benefit from an interface and dispatching the prerequisite methods. probably something to measure.
@ -0,0 +620,4 @@
(def (inline-container (kind : :symbol)
(source :? djot-source-range)
(children : :list)) => djot-container
(case kind
Owner

this could probably benefit from a hash table.

this could probably benefit from a hash table.
@ -0,0 +7,4 @@
(export string->djot)
;;; Public Gambit Unicode character sets have no Gerbil wrapper in v0.19.
(extern namespace: #f char-set:punctuation char-set-contains?)
Owner

this should be added to the runtime module in the prelude.

this should be added to the runtime module in the prelude.
@ -0,0 +114,4 @@
candidate)))))))
(def (visit (node : djot-node)) => :void
(let (identifier (node-identifier node))
Owner

you have a type annotation, use dotted notation!

you have a type annotation, use dotted notation!
Owner

i agree with fare; SSXML is ancient and not the fastest of cats. We could make a properly typed xml module, to replace sxml and html could be implemented using that. probably not in this pr though, it is already large, but it could be follow up work where we can measure improvements.

now, for immediate performance improvements see my comments, and also use a Reader instead of a port directly for reading the file to benchmark, this will make an immediate (but maybe small) difference.

i agree with fare; SSXML is ancient and not the fastest of cats. We could make a properly typed xml module, to replace sxml and html could be implemented using that. probably not in this pr though, it is already large, but it could be follow up work where we can measure improvements. now, for immediate performance improvements see my comments, and also use a Reader instead of a port directly for reading the file to benchmark, this will make an immediate (but maybe small) difference.
jay changed title from v0.19: Add std/markup/djot to WIP: v0.19: Add std/markup/djot 2026-09-28 18:59:58 +00:00
jay changed title from WIP: v0.19: Add std/markup/djot to v0.19: Add std/markup/djot 2026-09-28 20:14:35 +00:00
Author
Member

@vyzo addressed your feedback, i kept the SXML intermediary for now
there is definitely a lot of meat left on the bones in terms of performance gains, though

@vyzo addressed your feedback, i kept the SXML intermediary for now there is definitely a lot of meat left on the bones in terms of performance gains, though
Owner

yeah, but at least my quick feedback already produced gains; i looked in the benchmark.md and there seems to be quite an improvement -- clearly ahead of js now.

I think maybe we could improve performance by stopping parsing strings (and prereading a giant string) and just reading bytes/utf8 characters straight from the reader. Also it is likely that location tracking is expensive, we could have a flag to turn that off. It is useful for reporting errors, but it does cost, potentially quite a bit. Can you check with astra how hard it would be to make this transition to parsing from a buffered reader? Probably not so hard. This will also substantially improve memory usage as well, those big strings are heavy and memory hungry.

next we should figure out what to do with sxml; i think std/markup/xml and std/markup/html packages with properly structured xml and html are the way to go and now we have a nice benchmark we can use to measure things. But as I said this is for another pr, you want to take it?

yeah, but at least my quick feedback already produced gains; i looked in the benchmark.md and there seems to be quite an improvement -- clearly ahead of js now. I think maybe we could improve performance by stopping parsing strings (and prereading a giant string) and just reading bytes/utf8 characters straight from the reader. Also it is likely that location tracking is expensive, we could have a flag to turn that off. It is useful for reporting errors, but it does cost, potentially quite a bit. Can you check with astra how hard it would be to make this transition to parsing from a buffered reader? Probably not so hard. This will also substantially improve memory usage as well, those big strings are heavy and memory hungry. next we should figure out what to do with sxml; i think std/markup/xml and std/markup/html packages with properly structured xml and html are the way to go and now we have a nice benchmark we can use to measure things. But as I said this is for another pr, you want to take it?
Author
Member

@vyzo yeah I can take it, and you were thinking to remove sxml yes? like fully replace with std/markup/xml?

@vyzo yeah I can take it, and you were thinking to remove sxml yes? like fully replace with std/markup/xml?
Owner

yea, sxml is ancient and not exactly a joy to work with.

yea, sxml is ancient and not exactly a joy to work with.
Owner

@jay what do you think of trying the parser working directly with a (buffered) reader instead of parsing strings? We need to measure its performance, and we have benchmarks already.

@jay what do you think of trying the parser working directly with a (buffered) reader instead of parsing strings? We need to measure its performance, and we have benchmarks already.
This pull request can be merged automatically.
You are not authorized to merge this pull request.
View command line instructions

Checkout

From your project repository, check out a new branch and test the changes.
git fetch -u v0.19-djot:jay-v0.19-djot
git switch jay-v0.19-djot

Merge

Merge the changes and update on Forgejo.

Warning: The "Autodetect manual merge" setting is not enabled for this repository, you will have to mark this pull request as manually merged afterwards.

git switch v0.19-staging
git merge --no-ff jay-v0.19-djot
git switch jay-v0.19-djot
git rebase v0.19-staging
git switch v0.19-staging
git merge --ff-only jay-v0.19-djot
git switch jay-v0.19-djot
git rebase v0.19-staging
git switch v0.19-staging
git merge --no-ff jay-v0.19-djot
git switch v0.19-staging
git merge --squash jay-v0.19-djot
git switch v0.19-staging
git merge --ff-only jay-v0.19-djot
git switch v0.19-staging
git merge jay-v0.19-djot
git push origin v0.19-staging
Sign in to join this conversation.
No description provided.