98
2
Inkling: A New Open-Weight 975B Moe with a Few Surprises (sebastianraschka.com)
5
Claude Code's Real Secret Sauce Isn't the Model (sebastianraschka.com)
2
The State of LLMs 2025: Progress, Problems, and Predictions (sebastianraschka.com)
1
A Researcher's Field Guide to Non-Standard LLM Architectures (sebastianraschka.com)
3
Explanation of Gated DeltaNet (Qwen3-Next and Kimi Linear) (github.com/rasbt)
3
The Core Components of Modern LLMs and the Models Beyond Transformers [video] (youtube.com)
3
Popular Attention Alternatives: GQA, MLA, SWA (sebastianraschka.com)
4
Multi-Head Latent Attention (sebastianraschka.com)
4
Thinking Machines Lab Co-Founder Departs for Meta (wsj.com)
3
OpenAI's internal Slack messages could cost it billions in copyright suit (sherwood.news)
3
LLM Evaluation from Scratch: Multiple Choice, Verifiers, Leaderboards, LLM Judge (sebastianraschka.com)
107
Gemma 3 270M re-implemented in pure PyTorch for local tinkering (github.com/rasbt)
112
GPT-OSS vs. Qwen3 and a detailed look how things evolved since GPT-2 (sebastianraschka.com)
2
LLM Research Papers: The 2024 List (sebastianraschka.com)
1