Skip to main content

Posts

Featured

New ask Hacker News story: Ask HN: HotPin – lossless 120B MoE inference on 24GB RAM (CPU, 50 loc)

Ask HN: HotPin – lossless 120B MoE inference on 24GB RAM (CPU, 50 loc) 2 by LozzKappa | 0 comments on Hacker News. I'm a mechatronics designer with a background in control systems, robotics, PCB design, and embedded hardware. I design physical systems: motors, sensors, microcontrollers, and real-time control loops. I applied this design thinking to LLM memory management – and it worked. HotPin is a set of patches for llama.cpp that runs 30B–120B Mixture of Experts (MoE) models on far less RAM than their disk footprint, with bit-identical (lossless) output. Tested on an AMD Ryzen AI 9 HX 370 (Zen5, AVX512), 23.6GB LPDDR5X, NVMe >1GB/s, CPU-only. Results: | Model | Disk | Min RAM | Savings | tok/s | |-------|------|---------|---------|-------| | gpt-oss:120b | 58.5GB | 19.1GB | -67% | 3.84 | | qwen3:30b-a3b | 18.0GB | 10.4GB | -42% | 19.7 | | gemma4:26b-a4b | 16.2GB | 10.6GB | -35% | 11.5 | | GLM-4.7-Flash | 19.0GB | 13.3GB | -30% | 12.4 | Output is SHA-256 bit-identical to ful...

Latest Posts

New ask Hacker News story: Ask HN: How to get your open-source project gain popularity?

New ask Hacker News story: Ask HN: What's the best AI coding tool today?

New ask Hacker News story: Tell HN: ChatGPT exports do not contain all conversation messages

New ask Hacker News story: Ask: Why does everything Microsoft create lately feel fragile and half-baked?

New ask Hacker News story: Ask HN: Have you noticed an improvement in AI responses with memory disabled?

New ask Hacker News story: Are we in the era of AI slop landing pages?

New ask Hacker News story: OpenAI API Is Down

New ask Hacker News story: Building an AI-orchestrated publishing workflow for a long-form writing project

New ask Hacker News story: Ask HN: What changed in your life when you started meditating (and how)?

New ask Hacker News story: Ask HN: What do you consider the function of AI to be in your life currently?