An ongoing, hands-on log of testing Claude / GPT / Gemini / DeepSeek / MiniMax / GLM — both the base models and each vendor's own chatbot/agent products — to find the right brain for Nova (my agent on OpenClaw), and to learn which brain fits which scene. Not a leaderboard; a usage diary that keeps getting updated.
An ongoing log of testing frontier models — Claude, GPT, Gemini, DeepSeek, MiniMax, and Zhipu GLM — not just as raw models but through each vendor's own chatbot and agent products. The goal started practical: find the right brain to wire into Nova, my agent running on OpenClaw, and learn which model fits which scene. I compare context handling, multimodality, cost, hard-task quality, response speed, and specific generation skills (image / video / voice / text). It's a living document, not a finished scorecard — and it's honest about its limits: I never benchmarked coding on purpose (I leave coding to Gemini and Claude), and ~80% of my use is Chinese, so I didn't deliberately compare Chinese vs. English.
The Chinese version below covers the motivation, the evaluation dimensions and their boundaries, my (admittedly informal) methodology, a real usage history for each model, and the findings I keep revising.