造得出世界,走不进去:Karpathy 的 10 美元实验暴露了 LLM 的结构性盲区
Karpathy 用 Opus 5 将指环王文字渲染为 3D 场景,10 美元产出 5500 行代码。真正的故事不是 AI 多强,而是它造得出世界却看不见自己造的世界。
AdTech Algorithm Engineer. System Builder. AI Explorer.
Interested in
I build the prediction and pricing stack for a programmatic ad exchange. CTR models that convert CPM bids into CPC prices, serving millions of auction requests daily. The work is half ML (DNN with FM interactions, isotonic calibration, PCOC monitoring) and half systems (Go on 300+ K8s pods, multi-region TF Serving, Kafka pipelines).
Recently I've been designing training pipelines from the ground up: Parquet ingestion, feature canonicalization, cross-feature generation, incremental retraining with cold-start handling. I also lead an AI Agent project that automates ad traffic allocation, turning hours of daily manual ops into LLM-driven analysis and one-click config.
Side projects: an LLM API gateway aggregating 40+ providers with real-token probing for quality monitoring, and a lightweight agent framework. I'm drawn to the overlap between recommendation systems and large language models, specifically how transformer architectures can improve conversion prediction at scale.
Agent Harness Observability — detect errors, context rot, and regressions in AI agent systems.
A native macOS voice-to-text app — press Fn, speak, and polished text lands at your cursor in any app.
A production-ready multi-agent platform with sandboxed execution, budget control, and observability.
A Claude Code skill that generates daily AI/tech intelligence reports from Hacker News and HuggingFace Papers.
A Claude Code skill that generates importable Excalidraw architecture diagrams from source code.
Karpathy 用 Opus 5 将指环王文字渲染为 3D 场景,10 美元产出 5500 行代码。真正的故事不是 AI 多强,而是它造得出世界却看不见自己造的世界。
NVIDIA NOOA 用 253 行代码在 SWE-bench 跑出 82.2%,token 消耗减半。同时一份烧了 60B token 的 AGENTS.md 8 条规则走红。两者共同指向:约束比模型能力更重要。
Tw93 用三个月写了六篇'你不知道的'系列文章,从 Claude Code 到具身智能构建了完整的 AI 工程师认知框架,成为中文社区 AI 面试准备的非官方教材。