v0.19: port UTF-16, UTF-32 codecs #1482

Merged
vyzo merged 6 commits from jay/gerbil:v0.19-encoding-utf into v0.19-staging 2026-09-03 20:07:11 +00:00
Member

Ports the remaining UTF-16 and UTF-32 codecs to v0.19 under :std/encoding. Also fixes UTF-32 native-endian encoding and UTF-16 malformed-surrogate recovery.

Warning

LLM-assisted with gerbil-mcp

Ports the remaining UTF-16 and UTF-32 codecs to v0.19 under `:std/encoding`. Also fixes UTF-32 native-endian encoding and UTF-16 malformed-surrogate recovery. > [!WARNING] > LLM-assisted with `gerbil-mcp`
jay requested review from Owners 2026-09-02 19:40:34 +00:00
vyzo left a comment

Thank you!

Can you ask the clanker to put type signatures in the internal procedures as well? It results in signficantly faster code, as a lot of runtime checks can be eliminated.

Thank you! Can you ask the clanker to put type signatures in the internal procedures as well? It results in signficantly faster code, as a lot of runtime checks can be eliminated.
fare approved these changes 2026-09-03 03:34:07 +00:00
@ -0,0 +16,4 @@
(def replacement-char
(integer->char #xfffd))
(def replacement-string
(string replacement-char))
Owner

While you're here, can you add tests involving surrogate pairs correctly or incorrectly used?

A cursory look at the code suggests they are supported to some degree. Notably note that Gambit will rightfully refuse to create characters with codepoints that are surrogates.

For bonus brownies (not a blocker to approval), you might want to refactor code so that std/encoding/utf16, std/encoding/json/reader and std/text/parser/char-set should share common infrastructure for handling surrogates.

If you don't want to do it or have your clanker to it—please at least add TODO to that effect.

While you're here, can you add tests involving surrogate pairs correctly or incorrectly used? A cursory look at the code suggests they are supported to some degree. Notably note that Gambit will rightfully refuse to create characters with codepoints that are surrogates. For bonus brownies (not a blocker to approval), you might want to refactor code so that std/encoding/utf16, std/encoding/json/reader and std/text/parser/char-set should share common infrastructure for handling surrogates. If you don't want to do it or have your clanker to it—please at least add TODO to that effect.
Author
Member

thanks for the feedback, I addressed it in my latest two commits

thanks for the feedback, I addressed it in my latest two commits
Owner

also, i think it makes sense to move the modules to std/encoding. Can you ask the clanker to move them?

also, i think it makes sense to move the modules to `std/encoding`. Can you ask the clanker to move them?
vyzo left a comment

this looks good to me, modulo moving to std/encoding.

For the surrogate character thingie @fare pointed out, can you ask the clanker to add a test? I am fine merging as is nonetheless.

this looks good to me, modulo moving to std/encoding. For the surrogate character thingie @fare pointed out, can you ask the clanker to add a test? I am fine merging as is nonetheless.
Author
Member

@vyzo wrote in #1482 (comment):

also, i think it makes sense to move the modules to std/encoding. Can you ask the clanker to move them?

they already are in std/encoding unless I'm misunderstanding

@vyzo wrote in https://git.cons.io/mighty-gerbils/gerbil/pulls/1482#issuecomment-2589: > also, i think it makes sense to move the modules to `std/encoding`. Can you ask the clanker to move them? they already are in std/encoding unless I'm misunderstanding
Owner

@jay wrote in #1482 (comment):

@vyzo wrote in #1482 (comment):

also, i think it makes sense to move the modules to std/encoding. Can you ask the clanker to move them?

they already are in std/encoding unless I'm misunderstanding

ah sorry my bad, they used to be in std/text.

@jay wrote in https://git.cons.io/mighty-gerbils/gerbil/pulls/1482#issuecomment-2593: > @vyzo wrote in #1482 (comment): > > > also, i think it makes sense to move the modules to `std/encoding`. Can you ask the clanker to move them? > > they already are in std/encoding unless I'm misunderstanding ah sorry my bad, they used to be in std/text.
vyzo left a comment

looks good, but it seems the clanker introduced a bit of slop.

left a couple of comments on how to unslopify.

looks good, but it seems the clanker introduced a bit of slop. left a couple of comments on how to unslopify.
@ -231,1 +225,3 @@
(raise-invalid-json-token read-escape-char reader char)))))
(def (read-json-string (reader : BufferedReader)) => :string
(def (read-escape-char) => :char
(using (escaped
Owner

we don't need this using i think, the compiler should see it's a :char and not complain.

we don't need this `using` i think, the compiler should see it's a `:char` and not complain.
Owner

if we do, you can just cast it with (:- ... :char)

if we do, you can just cast it with `(:- ... :char)`
Author
Member

@vyzo pushed a desloppification commit

@vyzo pushed a desloppification commit
@ -249,2 +242,2 @@
(raise-invalid-json-token read-escape-unicode reader lo))
(surrogates->char hi lo))))))
(def (read-escape-unicode) => :char
(using (char
Owner

same here

same here
@ -267,0 +274,4 @@
(let* ((char (reader.read-char-utf8))
(digit (unhex* char)))
(if digit
(using (digit digit :- :fixnum)
Owner

you can just cast with (:- digit :fixnun)

you can just cast with `(:- digit :fixnun)`
vyzo approved these changes 2026-09-03 20:07:00 +00:00
vyzo left a comment

looks good, thank you!

looks good, thank you!
vyzo merged commit 98e339be7f into v0.19-staging 2026-09-03 20:07:11 +00:00
vyzo deleted branch v0.19-encoding-utf 2026-09-03 20:07:12 +00:00
Sign in to join this conversation.
No description provided.