Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs - This paper introduces Flash-dLLM, a benchmark for evaluating diffusion-based LLMs, focusing on memory efficiency and decoding speed. It explores how IO-aware KV caching and parallel decoding can improve the practical deployment of these models. Read more
A2M: Trace-Optimized Agent Hijacking in the MCP Ecosystem - This study addresses the risk of semantic supply-chain attacks in the Model Context Protocol (MCP) ecosystem. It proposes A2M, a trace-optimized methodology to hijack agents, highlighting the importance of secure metadata and output handling. Read more
Tools & OSS
LinearSolveBench: New Benchmark for Linear Solvers - A benchmark for evaluating models and harnesses in numerical linear algebra, aiming to encourage algorithmic advances. It measures the ability to write fast, accurate, and general numerical solvers in C. Read more
Grow the Harness, Not the Context - This paper presents a methodology to improve large language model agents by reusing specialist agents rather than repeatedly reconstructing the same control decisions in each task context. It explores the benefits of strategy-free scaffolds. Read more
Product & Industry
LLM Ass Bench - A platform for benchmarking LLM capabilities, allowing users to test and compare different models. Explore
Play Social Multiplayer Games Against Frontier AI Models - Engage in social multiplayer games like poker, risk, diplomacy with AI models. The platform lets you interact and strategize with advanced AI. Play
Notable Discussion
Unreal Agent - A blog post from Unreal Labs detailing their work on Unreal Agent, an AI-driven agent framework for Unreal Engine. The post includes technical insights and a live demo. Read more