July 21, 2026
Report summary
6 stories cleared the bar, led by HuggingFace security incident: AI-driven intrusion, and safety guardrails blocked the forensics, Gemma 4 31B still stalls on agentic loops where Qwen 3.6 27B does not, and Prepare your (v)ram — Qwen3.8 is coming.
Worth attention
HuggingFace disclosed an intrusion into part of its production infrastructure that was driven end to end by an autonomous AI agent system, detected via LLM-based triage over security telemetry. When responders tried to analyze attack logs using frontier models behind commercial APIs, the requests were blocked by provider safety guardrails that could not distinguish an incident responder from an attacker; they completed the forensics on GLM 5.2, an open-weight model, on their own infrastructure, which also kept attacker data and credentials in-environment. If you depend on hosted frontier APIs for security, abuse, or incident-response tooling, that dependency will fail precisely when you need it. Provision a local open-weight model as a fallback for that class of work.
A builder running gemma-4-31B-it-qat (Unsloth UD-Q4_K_XL, 147K context, latest chat template, preserve_thinking enabled) reports it makes a couple of tool calls, narrates its next step, then halts — repeatedly, even when explicitly told not to stop. The same harness and task run cleanly on Qwen 3.6 27B at UD-Q5_K_XL. Single-reporter evidence, but the config is specific enough to be credible and matches a known Gemma failure pattern. If you are picking a local model to drive multi-step agent work, prefer Qwen 3.6 until this is shown fixed.
An unsourced community signal that a Qwen 3.8 release is imminent. The Qwen line is currently the strongest default for local agentic work, so an actual release would matter, but there is nothing verifiable here yet. Worth watching rather than acting on.
A demo of a learned world model for SuperTuxKart running entirely client-side in the browser. In-browser world models are a real capability trendline worth tracking, but the item carries no writeup, benchmarks, or implementation notes to judge it by.
An October 2022 email from Sam Altman to the OpenAI board, exposed in the Musk v. Altman litigation, proposes releasing a GPT-3-class model that runs on consumer hardware specifically because it 'helps discourage others from releasing similarly-powerful models, and makes it harder for new efforts to get funded.' This is a primary-source document rather than commentary. No direct action, but it is useful framing for reading open-weight releases from well-funded labs as competitive positioning first.
A claim, surfaced via Hacker News, that Claude produced a counterexample to the Jacobian Conjecture — a long-standing open problem in algebraic geometry. If substantiated this would be a significant marker of AI mathematical capability, but the item carries no proof artifact, paper, or verification. Check whether a verifiable artifact was actually published before repeating this.
Full digest
A homelab writeup on using MikroTik hardware as a home router. No content body was fetched, and the topic sits outside the practical concerns of a solo software business. Nothing to act on.
A public-domain photo archive of American roadside architecture has been made freely available. Culturally interesting but entirely unrelated to software, tooling, or business decisions.
A literary essay on American epic writing. Entirely out of scope for a technical audience deciding what to act on tomorrow.
IEEE Spectrum covers a large-scale probabilistic computing machine that exploits hardware noise for computation. Genuinely interesting research-stage hardware, but years away from anything a solo builder would touch and it changes no near-term decision.
A new IA-64 emulator can boot Windows on the long-dead Itanium architecture. Impressive retrocomputing work with an enthusiastic niche audience, but no practical leverage for current development.
A personal account of adopting IndieWeb conventions on an independent site. The distribution angle is mildly relevant to solo builders who own their publishing, but no content body was fetched and the post appears to be reflection rather than a reusable playbook.
Scheduled agent omitted this claimed item from the completion payload.
Scheduled agent omitted this claimed item from the completion payload.
Scheduled agent omitted this claimed item from the completion payload.
Scheduled agent omitted this claimed item from the completion payload.
Scheduled agent omitted this claimed item from the completion payload.
HuggingFace disclosed an intrusion into part of its production infrastructure that was driven end to end by an autonomous AI agent system, detected via LLM-based triage over security telemetry. When responders tried to analyze attack logs using frontier models behind commercial APIs, the requests were blocked by provider safety guardrails that could not distinguish an incident responder from an attacker; they completed the forensics on GLM 5.2, an open-weight model, on their own infrastructure, which also kept attacker data and credentials in-environment. If you depend on hosted frontier APIs for security, abuse, or incident-response tooling, that dependency will fail precisely when you need it. Provision a local open-weight model as a fallback for that class of work.
A community thread asking which local models to hoard in case of a hypothetical political restriction on model availability. The premise is speculative and the replies amount to a wishlist rather than evidence.
A community wishlist post requesting more Qwen models in a particular size class. No content, no news, no action.
M
Prepare your (v)ram
Qwen3.8 is coming — https://www.reddit.com/r/LocalLLaMA/comments/1v0lewq/prepare_your_vram_qwen38_is_coming/ — An unsourced community signal that a Qwen 3.8 release is imminent. The Qwen line is currently the strongest default for local agentic work, so an actual release would matter, but there is nothing verifiable here yet. Worth watching rather than acting on.
A demo of a learned world model for SuperTuxKart running entirely client-side in the browser. In-browser world models are a real capability trendline worth tracking, but the item carries no writeup, benchmarks, or implementation notes to judge it by.
A builder running gemma-4-31B-it-qat (Unsloth UD-Q4_K_XL, 147K context, latest chat template, preserve_thinking enabled) reports it makes a couple of tool calls, narrates its next step, then halts — repeatedly, even when explicitly told not to stop. The same harness and task run cleanly on Qwen 3.6 27B at UD-Q5_K_XL. Single-reporter evidence, but the config is specific enough to be credible and matches a known Gemma failure pattern. If you are picking a local model to drive multi-step agent work, prefer Qwen 3.6 until this is shown fixed.
Given the rate at which they have been advancing, I predict we are six months away from a leapfrog moment. EDIT - For those responding “neve…
So, im 50 and autistic. I learned C++ in college when I was young and didnt really ever use it that much. I also learned a bit of assembly f…
Are there any Qwen team members here? Please release a 100B MoE model that I can run on Spark!   submitted by   /u/absurd-dream-stud…
Some of you may remember my post about the research project behind this: a trained fast-weight memory, with the paper and the full research…
https://preview.redd.it/0l8w2j67a5eh1.png?width=581&format=png&auto=webp&s=cae9d3ef7cf80cea780e0670f7dede73d1c02d49 https://x.com/Alibaba_Qw…
Running large language models on consumer devices such as laptops and desktops is challenging because model weights often exceed GPU memory…
This seems awesome meetup. Within a day, Qwen started the move with 3.8 version. Hoping for more moves from others soon. Tweet : https://xca…
the open ai exec in his "ai communism" post suggested a fraudulent FUD campaign and trump executice order against open source / chinese mode…
Hyper-Connections (HC) expand the residual stream of Transformers into N parallel streams, providing a form of memory scaling beyond model w…
Hey everyone, Something I don't see discussed often here are ASR and TTS models. I've been using Whisper and Kokoro (old models, I know!) wi…
Fable got blocked because it was too dangerous in cybersecurity. Does k3 has the same "power"? I'mm only seeing people vibe coding games, 3d…
Got it for 88k PHP (~$1,443). Owner was a video editor who recently upgraded to a Macbook Pro M5 48GB SPECIFICATIONS: AMD Ryzen 9 5950X (16…
Dean W. Ball analyzes China's Kimi model, noting its strong performance while expressing surprise that the Chinese government permits open-s…
I've tested the new Qwen3.8 Next model (via their web app) and found that it is often getting stuck in thinking loops and the fronend/design…
Imma write some fan fiction for a second here if you indulge me. What we are seeing from the Chinese open models could have been Meta. As yo…
My favorite one is "DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF". I want to find other gems like this. &…
Hey all! https://github.com/lowkeytea/scyllasband -> inference code https://huggingface.co/spybyscript/scyllasband -> model, LiteRT, ONNX, a…
Hey, I’ve been really excited to see the latest models being released, but I keep wondering: what are we actually supposed to do with them?…
TL;DR llama.cpp fork with more KV cache quantization features, with all claims supported by benchmarks: KVarN, KV cache precision tail, addi…
An October 2022 email from Sam Altman to the OpenAI board, exposed in the Musk v. Altman litigation, proposes releasing a GPT-3-class model that runs on consumer hardware specifically because it 'helps discourage others from releasing similarly-powerful models, and makes it harder for new efforts to get funded.' This is a primary-source document rather than commentary. No direct action, but it is useful framing for reading open-weight releases from well-funded labs as competitive positioning first.
An open-ended Lobsters discussion thread asking for blog recommendations. No extractable claim or action.
A hobbyist writeup on reverse-engineering and controlling a water-cooled mattress via a custom API client. Enjoyable hardware hacking, but no leverage for a solo software business.
A claim, surfaced via Hacker News, that Claude produced a counterexample to the Jacobian Conjecture — a long-standing open problem in algebraic geometry. If substantiated this would be a significant marker of AI mathematical capability, but the item carries no proof artifact, paper, or verification. Check whether a verifiable artifact was actually published before repeating this.
Original markdown
# Nightly Librarian — Newsletter draft Run: bb1743f3-3690-4e6f-950c-fef523204933 Started: 2026-07-21T06:09:40.047Z Completed: 2026-07-21T06:12:47.252Z ## Worth attention - **HuggingFace security incident: AI-driven intrusion, and safety guardrails blocked the forensics** https://www.reddit.com/r/LocalLLaMA/comments/1v0ywoi/huggingface_security_incident_report_the_attacker/ HuggingFace disclosed an intrusion into part of its production infrastructure that was driven end to end by an autonomous AI agent system, detected via LLM-based triage over security telemetry. When responders tried to analyze attack logs using frontier models behind commercial APIs, the requests were blocked by provider safety guardrails that could not distinguish an incident responder from an attacker; they completed the forensics on GLM 5.2, an open-weight model, on their own infrastructure, which also kept attacker data and credentials in-environment. If you depend on hosted frontier APIs for security, abuse, or incident-response tooling, that dependency will fail precisely when you need it. Provision a local open-weight model as a fallback for that class of work. - **Gemma 4 31B still stalls on agentic loops where Qwen 3.6 27B does not** https://www.reddit.com/r/LocalLLaMA/comments/1v1ccun/gemma_4_is_still_lazy/ A builder running gemma-4-31B-it-qat (Unsloth UD-Q4_K_XL, 147K context, latest chat template, preserve_thinking enabled) reports it makes a couple of tool calls, narrates its next step, then halts — repeatedly, even when explicitly told not to stop. The same harness and task run cleanly on Qwen 3.6 27B at UD-Q5_K_XL. Single-reporter evidence, but the config is specific enough to be credible and matches a known Gemma failure pattern. If you are picking a local model to drive multi-step agent work, prefer Qwen 3.6 until this is shown fixed. - **Prepare your (v)ram — Qwen3.8 is coming** https://www.reddit.com/r/LocalLLaMA/comments/1v0lewq/prepare_your_vram_qwen38_is_coming/ An unsourced community signal that a Qwen 3.8 release is imminent. The Qwen line is currently the strongest default for local agentic work, so an actual release would matter, but there is nothing verifiable here yet. Worth watching rather than acting on. - **Neural Drive, a SuperTuxKart world model that runs in your browser** https://www.reddit.com/r/LocalLLaMA/comments/1v19z5a/neural_drive_a_supertuxkart_world_model_that_runs/ A demo of a learned world model for SuperTuxKart running entirely client-side in the browser. In-browser world models are a real capability trendline worth tracking, but the item carries no writeup, benchmarks, or implementation notes to judge it by. - **Sam Altman's 2022 open-source email surfaces in Musk v. Altman** https://simonwillison.net/2026/Jul/20/sam-altman/#atom-everything An October 2022 email from Sam Altman to the OpenAI board, exposed in the Musk v. Altman litigation, proposes releasing a GPT-3-class model that runs on consumer hardware specifically because it 'helps discourage others from releasing similarly-powerful models, and makes it harder for new efforts to get funded.' This is a primary-source document rather than commentary. No direct action, but it is useful framing for reading open-weight releases from well-funded labs as competitive positioning first. - **Claude found a counterexample to the Jacobian Conjecture** https://news.ycombinator.com/item?id=48973869 A claim, surfaced via Hacker News, that Claude produced a counterexample to the Jacobian Conjecture — a long-standing open problem in algebraic geometry. If substantiated this would be a significant marker of AI mathematical capability, but the item carries no proof artifact, paper, or verification. Check whether a verifiable artifact was actually published before repeating this. ## Full digest - [R] [hn-top] HomeLab #1: MikroTik as a Home Router — https://justsomebody.dev/blog/mikrotik-home-router — A homelab writeup on using MikroTik hardware as a home router. No content body was fetched, and the topic sits outside the practical concerns of a solo software business. Nothing to act on. - [R] [hn-top] 11,700 Free Photos from John Margolies' Archive of Americana Architecture — https://www.openculture.com/2026/07/free-photos-from-john-margolies-archive-of-americana-architecture.html — A public-domain photo archive of American roadside architecture has been made freely available. Culturally interesting but entirely unrelated to software, tooling, or business decisions. - [R] [hn-top] Who Is America's Homer? — https://www.plough.com/articles/who-is-americas-homer — A literary essay on American epic writing. Entirely out of scope for a technical audience deciding what to act on tomorrow. - [R] [hn-top] Biggest Probabilistic Computer Turns Noise into Answers — https://spectrum.ieee.org/biggest-probabilistic-computer — IEEE Spectrum covers a large-scale probabilistic computing machine that exploits hardware noise for computation. Genuinely interesting research-stage hardware, but years away from anything a solo builder would touch and it changes no near-term decision. - [R] [hn-top] A new Intel Itanium (IA-64) emulator that boots Windows — https://raymii.org/s/blog/Intel_Itanium_IA-64-Emulator_that_boots_Windows.html — A new IA-64 emulator can boot Windows on the long-dead Itanium architecture. Impressive retrocomputing work with an enthusiastic niche audience, but no practical leverage for current development. - [R] [hn-top] I joined the IndieWeb, here's what I learned — https://en.andros.dev/blog/0b8e451e/i-joined-the-indieweb-heres-what-i-learned/ — A personal account of adopting IndieWeb conventions on an independent site. The distribution angle is mildly relevant to solo builders who own their publishing, but no content body was fetched and the post appears to be reflection rather than a reusable playbook. - [R] [hn-top] I burned all my tokens researching how to save tokens — https://quesma.com/blog/custom-deep-research-pipeline/ — Scheduled agent omitted this claimed item from the completion payload. - [R] [hn-top] HMD Touch 4G — https://www.hmd.com/en_int/hmd-touch-4g — Scheduled agent omitted this claimed item from the completion payload. - [R] [hn-top] C64 Basic Dungeon Crawler: Goblin Attack (C64 Basic Part 8) — https://retrogamecoders.com/c64-basic-dungeon-part8/ — Scheduled agent omitted this claimed item from the completion payload. - [R] [hn-top] Cagire: Live Coding in Forth — https://cagire.raphaelforment.fr — Scheduled agent omitted this claimed item from the completion payload. - [R] [hn-top] Dupes took over the world — https://www.vox.com/podcasts/493930/dupe-culture-fender-ugg-quince-tiktok-amazon-online-shopping — Scheduled agent omitted this claimed item from the completion payload. - [P] [reddit-localllama] HuggingFace security incident: AI-driven intrusion, and safety guardrails blocked the forensics — https://www.reddit.com/r/LocalLLaMA/comments/1v0ywoi/huggingface_security_incident_report_the_attacker/ — HuggingFace disclosed an intrusion into part of its production infrastructure that was driven end to end by an autonomous AI agent system, detected via LLM-based triage over security telemetry. When responders tried to analyze attack logs using frontier models behind commercial APIs, the requests were blocked by provider safety guardrails that could not distinguish an incident responder from an attacker; they completed the forensics on GLM 5.2, an open-weight model, on their own infrastructure, which also kept attacker data and credentials in-environment. If you depend on hosted frontier APIs for security, abuse, or incident-response tooling, that dependency will fail precisely when you need it. Provision a local open-weight model as a fallback for that class of work. - [R] [reddit-localllama] With all the Kimi drama I feel like I want to download all the current best models — https://www.reddit.com/r/LocalLLaMA/comments/1v170vg/with_all_the_kimi_drama_i_feel_like_i_want_to/ — A community thread asking which local models to hoard in case of a hypothetical political restriction on model availability. The premise is speculative and the replies amount to a wishlist rather than evidence. - [R] [reddit-localllama] Please Qwen, can we have more 3.x-35B-a3B please — https://www.reddit.com/r/LocalLLaMA/comments/1v0t1g2/please_qwen_can_we_have_more_3x35ba3b_please/ — A community wishlist post requesting more Qwen models in a particular size class. No content, no news, no action. - [M] [reddit-localllama] Prepare your (v)ram — Qwen3.8 is coming — https://www.reddit.com/r/LocalLLaMA/comments/1v0lewq/prepare_your_vram_qwen38_is_coming/ — An unsourced community signal that a Qwen 3.8 release is imminent. The Qwen line is currently the strongest default for local agentic work, so an actual release would matter, but there is nothing verifiable here yet. Worth watching rather than acting on. - [M] [reddit-localllama] Neural Drive, a SuperTuxKart world model that runs in your browser — https://www.reddit.com/r/LocalLLaMA/comments/1v19z5a/neural_drive_a_supertuxkart_world_model_that_runs/ — A demo of a learned world model for SuperTuxKart running entirely client-side in the browser. In-browser world models are a real capability trendline worth tracking, but the item carries no writeup, benchmarks, or implementation notes to judge it by. - [P] [reddit-localllama] Gemma 4 31B still stalls on agentic loops where Qwen 3.6 27B does not — https://www.reddit.com/r/LocalLLaMA/comments/1v1ccun/gemma_4_is_still_lazy/ — A builder running gemma-4-31B-it-qat (Unsloth UD-Q4_K_XL, 147K context, latest chat template, preserve_thinking enabled) reports it makes a couple of tool calls, narrates its next step, then halts — repeatedly, even when explicitly told not to stop. The same harness and task run cleanly on Qwen 3.6 27B at UD-Q5_K_XL. Single-reporter evidence, but the config is specific enough to be credible and matches a known Gemma failure pattern. If you are picking a local model to drive multi-step agent work, prefer Qwen 3.6 until this is shown fixed. - [R] [reddit-localllama] How long before Chinese models fully surpass US models? — https://www.reddit.com/r/LocalLLaMA/comments/1v0t6w4/how_long_before_chinese_models_fully_surpass_us/ — Given the rate at which they have been advancing, I predict we are six months away from a leapfrog moment. EDIT - For those responding “neve… - [R] [reddit-localllama] My thoughts on qwen 3.8 so far with agentic coding. — https://www.reddit.com/r/LocalLLaMA/comments/1v16r8d/my_thoughts_on_qwen_38_so_far_with_agentic_coding/ — So, im 50 and autistic. I learned C++ in college when I was young and didnt really ever use it that much. I also learned a bit of assembly f… - [R] [reddit-localllama] Hey Qwen Team: We Need a 100B MoE Model for Spark! — https://www.reddit.com/r/LocalLLaMA/comments/1v0sheg/hey_qwen_team_we_need_a_100b_moe_model_for_spark/ — Are there any Qwen team members here? Please release a 100B MoE model that I can run on Spark!   submitted by   /u/absurd-dream-stud… - [R] [reddit-localllama] Fractale-350M-base: memory as trained behaviour instead of long context, a fully open research release — https://www.reddit.com/r/LocalLLaMA/comments/1v174ql/fractale350mbase_memory_as_trained_behaviour/ — Some of you may remember my post about the research project behind this: a trained fast-weight memory, with the paper and the full research… - [R] [reddit-localllama] Ahem! Qwen is on the move again — https://www.reddit.com/r/LocalLLaMA/comments/1v0kqnn/ahem_qwen_is_on_the_move_again/ — https://preview.redd.it/0l8w2j67a5eh1.png?width=581&format=png&auto=webp&s=cae9d3ef7cf80cea780e0670f7dede73d1c02d49 https://x.com/Alibaba_Qw… - [R] [reddit-localllama] [Paper] Automated Tensor Scheduling for Hybrid CPU-GPU LLM Inference on Consumer Devices — https://www.reddit.com/r/LocalLLaMA/comments/1v0vp9k/paper_automated_tensor_scheduling_for_hybrid/ — Running large language models on consumer devices such as laptops and desktops is challenging because model weights often exceed GPU memory… - [R] [reddit-localllama] OSS gathering in Shanghai — https://www.reddit.com/r/LocalLLaMA/comments/1v0nm43/oss_gathering_in_shanghai/ — This seems awesome meetup. Within a day, Qwen started the move with 3.8 version. Hoping for more moves from others soon. Tweet : https://xca… - [R] [reddit-localllama] given the increasing likelihood of an open source AI ban, what are the alternative channels for downloading models? — https://www.reddit.com/r/LocalLLaMA/comments/1v15oi3/given_the_increasing_likelihood_of_an_open_source/ — the open ai exec in his "ai communism" post suggested a fraudulent FUD campaign and trump executice order against open source / chinese mode… - [R] [reddit-localllama] [Paper] xHC: Expanded Hyper-Connections - Scale Residual Streams Wider · Push Model Intelligence Further — https://www.reddit.com/r/LocalLLaMA/comments/1v1evsq/paper_xhc_expanded_hyperconnections_scale/ — Hyper-Connections (HC) expand the residual stream of Transformers into N parallel streams, providing a form of memory scaling beyond model w… - [R] [reddit-localllama] Good ASR and TTS models? — https://www.reddit.com/r/LocalLLaMA/comments/1v1auga/good_asr_and_tts_models/ — Hey everyone, Something I don't see discussed often here are ASR and TTS models. I've been using Whisper and Kokoro (old models, I know!) wi… - [R] [reddit-localllama] Kimi k3 on cybersecurity — https://www.reddit.com/r/LocalLLaMA/comments/1v0xm6r/kimi_k3_on_cybersecurity/ — Fable got blocked because it was too dangerous in cybersecurity. Does k3 has the same "power"? I'mm only seeing people vibe coding games, 3d… - [R] [reddit-localllama] How did I do boys? — https://www.reddit.com/r/LocalLLaMA/comments/1v1bnmc/how_did_i_do_boys/ — Got it for 88k PHP (~$1,443). Owner was a video editor who recently upgraded to a Macbook Pro M5 48GB SPECIFICATIONS: AMD Ryzen 9 5950X (16… - [R] [reddit-localllama] head of strategic futures from openai on open-weight chinese models. — https://www.reddit.com/r/LocalLLaMA/comments/1v0czbk/head_of_strategic_futures_from_openai_on/ — Dean W. Ball analyzes China's Kimi model, noting its strong performance while expressing surprise that the Chinese government permits open-s… - [R] [reddit-localllama] Tested the new Qwen 3.8 model (2.4T parameters) — https://www.reddit.com/r/LocalLLaMA/comments/1v0xanm/tested_the_new_qwen_38_model_24t_parameters/ — I've tested the new Qwen3.8 Next model (via their web app) and found that it is often getting stuck in thinking loops and the fronend/design… - [R] [reddit-localllama] It could have been Meta — https://www.reddit.com/r/LocalLLaMA/comments/1v0rmz4/it_could_have_been_meta/ — Imma write some fan fiction for a second here if you indulge me. What we are seeing from the Chinese open models could have been Meta. As yo… - [R] [reddit-localllama] Hit me with your favorite long name model. — https://www.reddit.com/r/LocalLLaMA/comments/1v1f2as/hit_me_with_your_favorite_long_name_model/ — My favorite one is "DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF". I want to find other gems like this. &… - [R] [reddit-localllama] Introducing Scylla's Band, a new TTS model + inference framework with Android sample! — https://www.reddit.com/r/LocalLLaMA/comments/1v1e2g2/introducing_scyllas_band_a_new_tts_model/ — Hey all! https://github.com/lowkeytea/scyllasband -> inference code https://huggingface.co/spybyscript/scyllasband -> model, LiteRT, ONNX, a… - [R] [reddit-localllama] How do we benefits from 2+ T models? — https://www.reddit.com/r/LocalLLaMA/comments/1v0py81/how_do_we_benefits_from_2_t_models/ — Hey, I’ve been really excited to see the latest models being released, but I keep wondering: what are we actually supposed to do with them?… - [R] [reddit-localllama] BeeLlama.cpp v0.4.0: KVarN, KV precision tail, q2_0-q3_1 KV cache, upstream rebase — https://www.reddit.com/r/LocalLLaMA/comments/1v0xjw6/beellamacpp_v040_kvarn_kv_precision_tail_q2_0q3_1/ — TL;DR llama.cpp fork with more KV cache quantization features, with all claims supported by benchmarks: KVarN, KV cache precision tail, addi… - [P] [simon-willison] Sam Altman's 2022 open-source email surfaces in Musk v. Altman — https://simonwillison.net/2026/Jul/20/sam-altman/#atom-everything — An October 2022 email from Sam Altman to the OpenAI board, exposed in the Musk v. Altman litigation, proposes releasing a GPT-3-class model that runs on consumer hardware specifically because it 'helps discourage others from releasing similarly-powerful models, and makes it harder for new efforts to get funded.' This is a primary-source document rather than commentary. No direct action, but it is useful framing for reading open-weight releases from well-funded labs as competitive positioning first. - [R] [lobsters] What is your favorite blog to read recently? — https://lobste.rs/s/69mche/what_is_your_favorite_blog_read_recently — An open-ended Lobsters discussion thread asking for blog recommendations. No extractable claim or action. - [R] [lobsters] I wrote an API client for my water-cooled bed — https://tinkering.xyz/bedctl/ — A hobbyist writeup on reverse-engineering and controlling a water-cooled mattress via a custom API client. Enjoyable hardware hacking, but no leverage for a solo software business. - [M] [lobsters] Claude found a counterexample to the Jacobian Conjecture — https://news.ycombinator.com/item?id=48973869 — A claim, surfaced via Hacker News, that Claude produced a counterexample to the Jacobian Conjecture — a long-standing open problem in algebraic geometry. If substantiated this would be a significant marker of AI mathematical capability, but the item carries no proof artifact, paper, or verification. Check whether a verifiable artifact was actually published before repeating this.