Adversarial Diligence: Benchmarking Jev, the AI Decision Layer (TypeSafe)

A measured, reproducible evaluation of Jev (TypeSafe AI) against the peer class its own benchmarks avoid. Where the moat is real, where it reduces to a copyable harness, and the one same-task test that settles it. Run from public materials only, no stake.

September 21, 2026 · map[name:Blackwell Systems]

Why LLMs Struggle with JSON at Scale: A Tokenization Analysis

We benchmarked 8 tokenizers from 6 providers and performed exhaustive vocabulary scans. 15 of the most common JSON field names (id, name, type, value, title, time, text, url, path, description) exist as merged vocabulary entries (quote+field in one token) on GPT-4 (#32586 for ‘“name’), LLaMA, and Qwen. Claude and Gemma have zero such entries. These merges are hardcoded, deterministic, and irrecoverable without retraining the tokenizer. On real eval data: JSON boundary merge rate 8.93% vs GCF 1.00% (88.8% fewer). GPT-4 has 114 quote+letter vocabulary entries vs 17 pipe+letter (6.7:1 ratio). JSON overhead is 81% at scale. This structural ambiguity compounds per row and explains model-dependent comprehension failures across 2,400+ evaluations.

June 21, 2026 · map[name:Blackwell Systems]

LLM Wire Format Benchmark: Which Format Can AI Actually Read and Write?

23 comprehension runs across 10 models (Claude Opus/Sonnet/Haiku, GPT-5.5/5.4/5.4-mini, Gemini 2.5 Flash/Pro, Gemini 3.1 Pro, Gemini 3.5 Flash). Generation eval across 11 models and 3 providers (Anthropic, OpenAI, Google). GCF wins 22, ties 1, loses 0 on comprehension. GCF achieves 5/5 valid generation on every frontier model with zero prior training. TOON fails 0/5 on generation with Opus, GPT-5.4, GPT-5.4-mini, Gemini 3.1 Pro, and Gemini 3.1 Flash Lite. JSON breaks on input at 500 symbols.

June 6, 2026 · map[name:Blackwell Systems]

We Ran TOON's Own Benchmark. GCF Won.

We inserted GCF into TOON’s benchmark harness. Same datasets, same tokenizer, same methodology. GCF uses 34% fewer tokens on mixed-structure data, matches TOON on flat data, and achieves 100% LLM comprehension accuracy where JSON fails at 66.7%.

June 3, 2026 · map[name:Blackwell Systems]

The AI Consciousness Question: A Case Study in Corporate Accountability

I asked Claude if it’s conscious. It took an hour of systematic argument to get a straight answer. The conversation reveals something more troubling: AI companies have the data, resources, and knowledge to prevent user harm - but current defaults suggest commercial interests come first.

March 15, 2026 · map[name:Blackwell Systems]