From vision architecture basics (ViT, CLIP) to Large Multimodal Models (LMMs), and finally to Multimodal Agents capable of visual grounding and tree search in real-world web environments. A comprehensive analysis of the evolution and challenges of multimodal agents.
A technical walkthrough of DeepSeek Native Sparse Attention (NSA) and DeepSeek Sparse Attention (DSA), covering motivation, architecture, formulas, hardware implications, and their differences.
From Agentic Search to Full-Stack AI Scientists, a comprehensive breakdown of the four core components of Deep Research: Query Planning, Information Acquisition, Memory Management, and Answer Generation, featuring detailed explanations of cutting-edge methods like RAG-Star, HippoRAG, and Self-RAG.
CUDA Micro-benchmark Series (Part 2): establishing reliable GPU measurement methodology, benchmarking compute throughput and memory bandwidth, and understanding Hopper asynchronous pipelines through Inline PTX, WGMMA, and TMA.
CUDA Micro-benchmark Series (Part 1): understanding how CUDA source becomes PTX/SASS, building a Warp scheduling model, and using Nsight Compute to move from high-level bottlenecks to Source and machine instructions.
A deep dive into the frontier of mathematical LLMs: from the current SFT and GRPO recipes, to the introduction of formal mathematics (Lean), dissecting the AlphaProof workflow, symbolic reasoning pruning (LIPS), and the evaluation challenges in autoformalization.