One entry per notable day in the Zipp repository, reconstructed from commit subjects, tag annotations and the handoff notes, each linked to the commits it describes. Search by word or filter by topic; every entry also has its own page and appears in the RSS feed.
45 entries29 May 2026 to 25 September 2026RSS feed
Native dict, set and instance storage, direct calls and inlined attribute access take Python to about 1.75x CPython 3.13 with the JIT; a hello-world starts in 20 ms; the compiler reads Zipp's own syntax tree and emits the same bytecode every run.
Direct calls, fused Python-only instructions, native helpers, a JIT tier for hot loops, per-VM caches and small integers in the value word take the Python benchmark suite from 22.3x CPython 3.13 to about 2.1x.
Python without torch becomes a smaller WebAssembly variant with torch as a verified package, the Python parser becomes hand-written and then Zipp's own, and a prepared GPU step falls to 0.27 ms against PyTorch's 0.78.
Graph protocol versions 3 and 4 add masks, dropout, slicing and gathers bit for bit on five backends, and zipp py runs the same graphs on Vulkan, Direct3D 12 or Metal through wgpu.
Silent errors in eager torch are fixed, the layers, optimizers, dtypes, linalg, fft and distributions ordinary training code uses are added, and eager torch's heavy loops run natively: about 4.8x faster, same bytes.
Before building further, a pass over the model host, gpu-lab sessions, the WebAssembly host boundary, the dev server, the Python frontend and the JIT fixes the defects found, each with a regression test.
Qwen3 splits into stages that reproduce the whole model bit for bit, matmul_fixed accumulates in integers so every backend gives the same bits, the kernels build reproducibly, and v0.0.20 ships.
An experimental local-model plugin overlay arrives, and the GPU protocol learns to keep a checkpoint's Q4_K and Q6_K blocks and decode them inside the matmul on every backend.
v0.0.19 was held because the README promised more of torch.compile than it did; four tracks made it true - protocol-v2 capture, resident GPU sessions, native kernels, a second pass over the Python hot paths - and the release shipped with the two longest gate jobs run locally.
The pinned Test262 corpus passes 95,671 of 95,680 executions unmodified and 95,680 of 95,680 with five documented corrections; the DateTimeFormat shard passes 488 of 488; v0.0.18 is published.
zipp.org gains crawlable pages for the engine, Python, WebAssembly, GPU compute, Test262, benchmarks and architecture, all rendered from one JSON file, plus this journal with an RSS feed.
The experimental Python 3 frontend lands, compiling to the same register bytecode as JavaScript, together with WASM build variants, a folder-based playground, GPU graphs from Python and a Torch subset with GPU training.
Two releases close 38 tickets from the 11 September audits: VM rooting, coercion, TypedArray semantics, parser, GC and metadata, while preserving the exact 95,939 of 95,942 result on the previous corpus.
Embedding-API and host-boundary audit findings are fixed, browser Worker smoke tests and a reference host adapter land, releases are gated on the exact tagged commit, and a fresh benchmark capture shows the cost of strict call order.
Method calls now evaluate in specification order by default; the fused lowering that could be observed by a getter or proxy trap is used only where argument shapes provably cannot see it.
v0.0.14 adds the accel bridge; paired A/B measurements close five WASM interpreter pathologies, including a join with 300K retained strings from 8,423 ms to 138 ms.
v0.0.11 and v0.0.12 land register classes for tokenizer loops and inline stores into young dense arrays; the clean PGO capture at 8229b3fc becomes the public benchmark, with 21 of 30 rows faster than Node.
Six releases in one day, each raising a ceiling a real embedder had hit: a 32 MiB ArrayBuffer, a renewable instruction budget, byte-scheduled collections, a 512 MB heap.
The host boundary stops walking the whole heap per re-entry; an O(heap) walk is removed from the interpreter-only WASM build; JSON.stringify ×4000 drops from 919 ms to 8.5 ms; wasm-opt is dropped after measurement.
The first native and WebAssembly release; B249 takes bytecode-vm from 3.68× to 0.93× Node; module cycle roots restore 95,939 of 95,942; the performance ledger is archived at 881 KB.
The generational nursery's first stage is refuted by its own measurements before the third stage lands with a prover; the first landing page is committed the same day.
Benchmark artifacts begin recording engine commit, dirty flag, competitor hashes and per-repetition schedules, after a harness was found able to name a commit it had never measured.
Two phantom executions removed, then 22 of 30 remaining failures traced to the repository or runner rather than the engine; 99.995% on 95,846 executions.
A new front end cuts parse-negative failures from 607 files to 80; honouring YAML list-form flags drops the denominator by 181 executions and the score is republished.
The WebAssembly embedding lands as a persistent VM that compiles once and re-enters later; fat LTO with one codegen unit buys about 2%; regex literals validated at compile time add 712 executions.
After six silent weeks, the Test262 runner is found to have been scoring a single execution mode; the honest pass rate is 96.97%. The same day, every crate but zipp-vm is deleted and the docs are rewritten against measurements.
The whole-function JIT tier is enabled by default (parse 4.1× → 3.3× Node) and PERF_ROADMAP.md is written as a checkable path to V8 parity. Then the log goes quiet for six weeks.
Strings gain real lone-surrogate support, BigInt gains a two-tier i128 plus num-bigint implementation, and bench/real arrives as a ten-workload five-engine suite.
Module mode lands with top-level await and a Test262 module runner; modules become runtime structures on the explicit frame stack like everything else.
A forked regress engine for regular expressions, eleven TypedArray constructors, Proxy traps, Temporal type by type, the whole Intl namespace, a mark-sweep GC, and a 17,400-line vm.rs split into a folder.
The register VM gets an x86-64 JIT for hot integer functions, OSR loop regions, call-free inline caches, scalar replacement, rope strings and a microtask loop, and beats V8 on loop.js at 26.5 ms against 28.0.
The ES5 object model arrives, a bytecode compiler and VM are built, two engines are tried, and by evening the clean-sheet explicit-frame register VM exists. So does the first retraction of a false performance claim.
The initial commit is a tree-walking interpreter for a small typed language; by the end of the day it has a Cranelift JIT, an LLVM tier, a mark-sweep GC, a gas-metered WebAssembly profile, a zero-knowledge back end and a TypeScript front end.
No entries match that search. Try a different word, or clear the filters.