二十分之一的算力,凭什么只落后一年
DeepSeek 用美国 1/20 的算力做到只落后一到两年。这不是省钱,是约束迫使他们走了一条完全不同的架构路线:MoE 稀疏激活、FP8 全流程训练、自研稀疏注意力、昇腾 Day 0 适配。
AdTech Algorithm Engineer. System Builder. AI Explorer.
Interested in
I build the prediction and pricing stack for a programmatic ad exchange. CTR models that convert CPM bids into CPC prices, serving millions of auction requests daily. The work is half ML (DNN with FM interactions, isotonic calibration, PCOC monitoring) and half systems (Go on 300+ K8s pods, multi-region TF Serving, Kafka pipelines).
Recently I've been designing training pipelines from the ground up: Parquet ingestion, feature canonicalization, cross-feature generation, incremental retraining with cold-start handling. I also lead an AI Agent project that automates ad traffic allocation, turning hours of daily manual ops into LLM-driven analysis and one-click config.
Side projects: an LLM API gateway aggregating 40+ providers with real-token probing for quality monitoring, and a lightweight agent framework. I'm drawn to the overlap between recommendation systems and large language models, specifically how transformer architectures can improve conversion prediction at scale.
Agent Harness Observability — detect errors, context rot, and regressions in AI agent systems.
A native macOS voice-to-text app — press Fn, speak, and polished text lands at your cursor in any app.
A production-ready multi-agent platform with sandboxed execution, budget control, and observability.
A Claude Code skill that generates daily AI/tech intelligence reports from Hacker News and HuggingFace Papers.
A Claude Code skill that generates importable Excalidraw architecture diagrams from source code.
DeepSeek 用美国 1/20 的算力做到只落后一到两年。这不是省钱,是约束迫使他们走了一条完全不同的架构路线:MoE 稀疏激活、FP8 全流程训练、自研稀疏注意力、昇腾 Day 0 适配。
Harness Engineering 有一个结构性盲区:RL 训练缺少代码可维护性的奖励信号。这让所有试图关灯运行的 Software Factory 在 3-6 个月后被迫重新开灯。