v0.19: port the remaining modules of std/os #1485

Merged
fare merged 11 commits from v0.19-std-os into v0.19-staging 2026-09-07 22:03:20 +00:00
Owner

ports std/os modules flock, pipe and inotify.

done with sol with assistance from gerbilosaurus rex.

ports std/os modules flock, pipe and inotify. done with sol with assistance from gerbilosaurus rex.
vyzo requested review from fare 2026-09-07 11:07:24 +00:00
Owner

While it LGTM, GPT 5.6 High had this to say:

  1. Blocker: O_TRUNC happens before the lock is acquired. The new open-file-writer/lock defaults to O_CREAT|O_TRUNC, then eventually reaches open/lock, which first calls open(... flags ...) and only afterwards calls flock-fd/block. O_TRUNC truncates an existing writable regular file as part of open() itself. Thus, suppose A holds the exclusive flock while preserving some valuable contents. B calls open-file-writer/lock and waits for the lock—or even times out without ever obtaining it. B has already truncated the file underneath A. This rather defeats the point of the locked writer. The current timeout test only contends using open-file-reader/lock, so it cannot catch this. I would put the review comment on the __open call in src/std/os/flock.ss:131. I think the right fix is for open/lock to strip O_TRUNC, acquire the lock, then ftruncate(fd, 0) if truncation was requested; fixing it at that level also handles callers of open-file/lock that explicitly pass O_TRUNC. A regression test should put contents in a file, hold an exclusive lock, try a default locked writer with a short timeout, then verify that the original contents survived.

  2. Blocker-ish: flock can operate on a stale, reused FD after its OSDevice is closed. flock and flock/block immediately extract dev.fd without doing the usual do-check-device-open. But device-close marks dev.dir closed and closes the raw port without changing dev.fd; the rest of the device API has explicit closed-device checks. Therefore (flock closed-dev ...) doesn't reliably mean EBADF: if that descriptor number has since been reused, it can lock an entirely unrelated file. More seriously, flock/block freezes the integer FD and then repeatedly sleeps and retries. During one of those cooperative sleeps another green thread can close the device and open something else that receives the same FD number; the next retry can then lock the wrong object. This does not require SMP. I would comment around src/std/os/flock.ss:50-59. flock should check do-check-device-open; flock/block needs the check on each polling iteration, not merely once before entering it. The fd-only helper is still useful for open/lock, since there the descriptor hasn't yet escaped.

  3. SMP/fork-exec: pipe has the classic FD_CLOEXEC race. It does pipe() first and then two separate F_SETFD FD_CLOEXEC operations. ([cons.io git][5]) If another OS thread executes fork()+exec() in between, a pipe FD leaks into the child. This is exactly the race creation-time CLOEXEC flags are intended to prevent; the Linux documentation explicitly warns that post-creation F_SETFD is insufficient in multithreaded programs. On Linux, pipe2(..., O_CLOEXEC) sets it atomically on both ends. Given the SMP work, I'd mention this now. I wouldn't necessarily hold up this PR over portable handling: use pipe2(O_CLOEXEC) on Linux and retain the pipe()+fcntl fallback elsewhere, or leave a conspicuous TODO tied to fork/SMP semantics. I would put this comment on __pipe-syscall around src/std/os/pipe.ss:51.

While it LGTM, GPT 5.6 High had this to say: 1. **Blocker: `O_TRUNC` happens before the lock is acquired.** The new `open-file-writer/lock` defaults to `O_CREAT|O_TRUNC`, then eventually reaches `open/lock`, which first calls `open(... flags ...)` and only afterwards calls `flock-fd/block`. `O_TRUNC` truncates an existing writable regular file as part of `open()` itself. Thus, suppose A holds the exclusive flock while preserving some valuable contents. B calls `open-file-writer/lock` and waits for the lock—or even **times out without ever obtaining it**. B has already truncated the file underneath A. This rather defeats the point of the locked writer. The current timeout test only contends using `open-file-reader/lock`, so it cannot catch this. I would put the review comment on the `__open` call in `src/std/os/flock.ss:131`. I think the right fix is for `open/lock` to strip `O_TRUNC`, acquire the lock, then `ftruncate(fd, 0)` if truncation was requested; fixing it at that level also handles callers of `open-file/lock` that explicitly pass `O_TRUNC`. A regression test should put contents in a file, hold an exclusive lock, try a default locked writer with a short timeout, then verify that the original contents survived. 2. **Blocker-ish: `flock` can operate on a stale, reused FD after its `OSDevice` is closed.** `flock` and `flock/block` immediately extract `dev.fd` without doing the usual `do-check-device-open`. But `device-close` marks `dev.dir` closed and closes the raw port without changing `dev.fd`; the rest of the device API has explicit closed-device checks. Therefore `(flock closed-dev ...)` doesn't reliably mean EBADF: if that descriptor number has since been reused, it can lock an entirely unrelated file. More seriously, `flock/block` freezes the integer FD and then repeatedly sleeps and retries. During one of those cooperative sleeps another green thread can close the device and open something else that receives the same FD number; the next retry can then lock the wrong object. This does **not require SMP**. I would comment around `src/std/os/flock.ss:50-59`. `flock` should check `do-check-device-open`; `flock/block` needs the check on **each polling iteration**, not merely once before entering it. The fd-only helper is still useful for `open/lock`, since there the descriptor hasn't yet escaped. 3. **SMP/fork-exec: `pipe` has the classic `FD_CLOEXEC` race.** It does `pipe()` first and then two separate `F_SETFD FD_CLOEXEC` operations. ([cons.io git][5]) If another OS thread executes `fork()+exec()` in between, a pipe FD leaks into the child. This is exactly the race creation-time CLOEXEC flags are intended to prevent; the Linux documentation explicitly warns that post-creation `F_SETFD` is insufficient in multithreaded programs. On Linux, `pipe2(..., O_CLOEXEC)` sets it atomically on both ends. Given the SMP work, I'd mention this now. I wouldn't necessarily hold up this PR over portable handling: use `pipe2(O_CLOEXEC)` on Linux and retain the `pipe()+fcntl` fallback elsewhere, or leave a conspicuous TODO tied to fork/SMP semantics. I would put this comment on `__pipe-syscall` around `src/std/os/pipe.ss:51`.
Owner

I'm impressed by the attention to detail and knowledge of the Unix API of GPT.

Asking for lock-on-open, GPT tells me it exists on FreeBSD and Darwin, with O_EXLOCK / O_SHLOCK, but not on Linux or other BSDs.

I'm impressed by the attention to detail and knowledge of the Unix API of GPT. Asking for lock-on-open, GPT tells me it exists on FreeBSD and Darwin, with `O_EXLOCK` / `O_SHLOCK`, but not on Linux or other BSDs.
Author
Owner

review issue addressed, please rereview.

review issue addressed, please rereview.
fare force-pushed v0.19-std-os from 4fc7cfb8cb to 16452acdc3 2026-09-07 22:01:31 +00:00 Compare
fare merged commit d7ec60e8ca into v0.19-staging 2026-09-07 22:03:20 +00:00
fare deleted branch v0.19-std-os 2026-09-07 22:03:37 +00:00
Owner

Followup: Leave a TODO for O_{EX|SH}LOCK on FreeBSD/Darwin ? Have a clanker do it?

Followup: Leave a TODO for O_{EX|SH}LOCK on FreeBSD/Darwin ? Have a clanker do it?
Sign in to join this conversation.
No description provided.