1
1
k-Coloring is Faster than Computing the Chromatic Number (arxiv.org)
2
x86 AMX/ACE with >8 tiles (kernel.org)
1
RSI: Recursive Self-Improvement (github.com/zartbot)
1
Specula: Scaling formal specs for autonomous model checking of system code (arxiv.org)
2
NVIDIA-labs OO Agents: Native Python Object-Oriented Agents (github.com/nvidia-nemo)
1
At-the-Roofline Sparse Tensor Contractions on Vector Processors for Inference (arxiv.org)
1
Characterizing Warp Divergence from Pascal to Blackwell (arxiv.org)
2
Why Higher-Order Logic Is a Good Foundation for Deep Verification (sequent.inc)
1
Fleet: Hierarchical Task-Based Abstraction for Megakernels on Multi-Die GPUs (arxiv.org)
1
Fuzzing with Agents? Generators Are All You Need (arxiv.org)
22
Kuna: Decompiler Development in the Age of Coding Agents (noelo.org)
1
PyCuTe: Reference implementation and examples of the CuTe Layout (github.com/nvlabs)
24
PyTorch: A Reference Language (pytorch.org)
2
Tracked Capabilities for Safer Agents (martinodersky.substack.com)
5
Attention Decode on AMD MI450 GPUs: A Gluon Kernel Optimization Guide (amd.com)
1
Don't Use Ordinary Software to Contain Software-Hacking Agents (verse.systems)
1
A practical partitioner for distributed simulations on sparse dynamic domains (acm.org)
1
NVLink, NVSwitch, and All That (doubleword.ai)
1
Continuation-Centric Computing with Arca (usenix.org)
2
AutoCO: An Online Continuous Optimization System for Phase-Changing DB Workloads (acm.org)
1
Reusing Buffers in Multicore OCaml (gazagnaire.org)
3
Mohabi: Disaggregating and Sandboxing the Firefox JavaScript Engine (usenix.org)
1
Automated Discovery Has No Universally Superior Harness (arxiv.org)
2
Drawn Apart: Remote GPU Fingerprinting (bgu.ac.il)
2
ICFP Programming Contest 2026 (icfpcontest2026.com)
1
Optimization of SASS Stall Counts (redplait.blogspot.com)
1
Characterizing Metastable Faults and Failures (muratbuffalo.blogspot.com)
1
Every Microsecond Matters:Achieving Near SpeedOfLight Latency in GPU Collectives (arxiv.org)
12
An Empirical Study: AI Agent Rules Need Context and Layered Enforcement (eunomia.dev)
4
SPIR-V on ROCm: A Portable IR for AMD GPUs (amd.com)
1
GEMM Performance Measurement Methodology Guidelines (nvidia.com)
1
Faster Algorithms for Structured Matrix Multiplication via Flip Graph Search (acm.org)
2
AI for Systems is "AGI-Complete" (acm.org)
1
Controlling Reasoning Effort in LLMs (sebastianraschka.com)
3
Making TLA+ and x86 Kiss via Z3Py (philipzucker.com)
2
Irreducible Loops (maskray.me)
7
Fleet: Hierarchical Task-Based Abstraction for Megakernels on Multi-Die GPUs (arxiv.org)
4
Accelerating Block Low-Rank Foundation Model Inference on MemoryConstrained GPUs (acm.org)
2
Locality-Aware Automatic Differentiation on the GPU for Mesh-Based Computations (acm.org)
1
Triton Plugin Extensions: Enabling TLX and Custom Compiler Passes Out of the Box (pytorch.org)
1
What If the Harness Comes Before Pretraining? A Data Flywheel Perspective (hanchenli.github.io)
2
Mimesys: Turn Resource Usage Traces into Executable Workloads (usenix.org)
2
System call instrumentation on Linux/x86-64 using memory-indirect calls (humprog.org)
2
Bidirectional Elaborators à la Carte (arxiv.org)
2
CTA-Pipelining: A Latency-Oriented Spatial Scaling Method for Multi-GPU Systems (arxiv.org)
1
AI Model Co-Design: Hardware-Friendly LLM Design (nvidia.com)
1
Towards Free Normalization: Fusing Normalization into GEMM and Attention Kernels (pytorch.org)
1
Negotiating AI in Open Source Software Communities: A Case Study of LLVM Project (gu.se)
2
Reducing HBM Bottlenecks in JAX-Based LLM Training with Host Offloading (nvidia.com)
2
Compiler Testing – Part 2: Metamorphic Testing with Verified Identities (nowarp.io)
6
What Every Python Developer Should Know About the CPython ABI (quansight.org)
1
A CIRCT (Circuit IR Compilers and Tools) Project Tutorial (samuelcoward.co.uk)
1
Lifting Terms: Making Well Scoped Syntax Dumber (philipzucker.com)
1
Sheaves in Haskell (tweag.io)
1
Writing static checks to an unsuspecting library with Liquid Haskell (tweag.io)
1
Harvesting Sub-Microsecond CXL Memory Stalls with LiteSwitch (usenix.org)
3
Formally Verifying AI-Generated GPU Kernels (gimletlabs.ai)
2
Occupancy Math on the AMD MI355X GPU (CDNA4): A From-First-Principles Guide (amd.com)
3
Intelligence Is Free, Now What? Data Systems For, Of, and by Agents (bair.berkeley.edu)
2
Binvariants: Register-Level Invaraint-Guided Fuzzing for Binaries (github.com/futureslab)
2
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models (supercomputing-system-ai-lab.github.io)
2
Ekka: Automated Diagnosis of Silent Errors in LLM Inference (washington.edu)
2
Architecture 2.0: Designing AI-Assisted Loops for Computing Systems (mlsysbook.ai)
2
Beyond Prediction: Tail-Aware Scheduling for LLM Inference (yl3469.github.io)
2
SOLAR: AI-Powered Speed-of-Light Performance Analysis (arxiv.org)
2
The Verification Horizon: No Silver Bullet for Coding Agent Rewards (arxiv.org)
4
Binary Coverage the Wrong Way (redvice.org)
2
Stealing 50 Years of Database Ideas for AI Agents (onewill.ai)
1
Empirical Computation: Prompting versus Programming [pdf] (mboehme.github.io)
1
Programming Language Design and Implementation in the Era of Machine Learning [video] (youtube.com)
63
Senior SWE-Bench: open-source benchmark that assesses agents as senior engineers (snorkel.ai)
2
Software Security Analysis in 2030 and Beyond: A Research Roadmap (acm.org)
1
VeriCache: Turning Lossy KV Cache into Lossless LLM Inference (arxiv.org)
1
Designing GPU-Accelerated Query Engines with NVIDIA GQE (nvidia.com)
2
Teaching Algorithms in 2026 (uni.edu)
1
Pragmatic Approaches to Improving Compiler Correctness (ecoop.org)
1
The Expensive Fictions of Low-Level Programming Languages (stng.substack.com)
2
On the Efficacy of PyTorch for High-Performance Computing (acm.org)
2
Accelerating LLM Inference on AMD GPUs with Low-Latency GEMMs (amd.com)
1
Agentic Hardware Design as Repository-Level Code Evolution (arxiv.org)
3
Reward hacking is swamping model intelligence gains (cursor.com)
3
Memora: A Harmonic Memory Representation Balancing Abstraction and Specificity (microsoft.com)
5
Compiler-Assisted Floating-Point Error Analysis and Profiling with FPChecker (fpanalysistools.org)
1
Are We Ready for an Agent-Native Memory System? (arxiv.org)
25
Micro-Agent: Beat Frontier Models with Collaboration Inside Model API (vllm.ai)
1
TraceLab: Characterizing Coding Agent Workloads for LLM Serving (washington.edu)
1
It's Always the Learning Rates (ianbarber.blog)
3
Efficiency in LLMs – Part 1 – Columbia Machine Learning Summer School 2026 [video] (youtube.com)
1
Using Local Coding Agents (sebastianraschka.com)
1
The Thing We All Obviously Want (kmicinski.com)
2
ActPlane: Programmable OS-Level Policy Enforcement for Agent Harnesses (arxiv.org)
1
Fenwick trees for products mod 2ⁿ (bitmath.blogspot.com)
3
A Fake Shell for Pangenomics (cornell.edu)
1
Making Equality Saturation Usable for Developing Vectorized Compilers (acm.org)
2
Reading AI Model Compilation in MLIR Through the Lens of Formal Theories (arxiv.org)
2
Liveness Proofs in Veil, Part I: The First Step (proofsandintuitions.net)
1
ParallelKernelBench: Can LLMs write fast multi-GPU kernels? (github.com/togethercomputer)
1
LXM: Better Splittable Pseudorandom Number Generators (and Almost as Fast) [video] (youtube.com)
2