zhaoweiguo 的知识库
Search
搜索
暗色模式
亮色模式
目录
标签: benchmark
此标签下有39条笔记。
2026年9月22日
DuMateBench(智能体真实交付能力基准)
ai-ml
benchmark
ecosystem
noteworthy
2026年9月22日
Multi-SWE-bench(多语言 issue 修复基准)
ai-ml
benchmark
ecosystem
noteworthy
2026年9月22日
DuMateBench 榜单(20 组 agent × 模型快照)
ai-ml
benchmark
ecosystem
2026年9月22日
Multi-SWE-bench 榜单(论文基线 + 社区提交快照)
ai-ml
benchmark
ecosystem
2026年9月22日
DuMateBench 打榜指南
ai-ml
benchmark
ecosystem
2026年9月22日
Multi-SWE-bench 打榜指南
ai-ml
benchmark
ecosystem
2026年9月22日
RepoLaunch — 仓库构建/测试环境自动化框架
ai-ml
benchmark
ecosystem
devtool
2026年9月22日
SWE-agent-for-eval 参考 Harness 解析
ai-ml
benchmark
ecosystem
devtool
2026年9月22日
microsoft/SWE-bench-Live
ai-ml
benchmark
ecosystem
devtool
2026年9月19日
Coding Agent 榜单总览
ai-ml
benchmark
ecosystem
recommended
2026年9月19日
Coding Agent 打榜选型(自研 harness + 对外曝光)
ai-ml
benchmark
ecosystem
2026年9月19日
打榜指南目录
ai-ml
benchmark
ecosystem
2026年9月19日
SWE-bench-Live 打榜指南
ai-ml
benchmark
ecosystem
2026年9月19日
Terminal-Bench 打榜指南
ai-ml
benchmark
ecosystem
2026年9月18日
harbor-framework/harbor
ai-ml
library
cli
devtool
benchmark
recommended
2026年8月13日
LLM 评测集总览
ai-ml
benchmark
ecosystem
2026年8月13日
Apex Shortlist(MathArena Apex)
ai-ml
benchmark
noteworthy
2026年8月13日
Codeforces(LLM 竞赛编程评测口径)
ai-ml
benchmark
recommended
2026年8月13日
DeepSWE(Datacurve)
ai-ml
benchmark
recommended
2026年8月13日
DSBench-FullStack(DeepSeek 内部全栈测试集)
ai-ml
benchmark
noteworthy
2026年8月13日
Humanity's Last Exam(HLE)
ai-ml
benchmark
recommended
2026年8月13日
SuperCLUE(中文大模型测评基准)
ai-ml
benchmark
noteworthy
2026年8月13日
SWE-bench Verified
ai-ml
benchmark
recommended
2026年8月13日
Terminal-Bench(4.0)
ai-ml
benchmark
recommended
2026年8月13日
Toolathlon-Verified(Tool Decathlon)
ai-ml
benchmark
recommended
2026年8月13日
Vibe Code Bench(Vals AI)
ai-ml
benchmark
noteworthy
2026年8月13日
LLM 榜单总览
ai-ml
benchmark
ecosystem
2026年8月13日
Apex Shortlist 分数(快照 2026-08-13)
ai-ml
benchmark
2026年8月13日
Codeforces rating 进展榜(快照 2026-08-13)
ai-ml
benchmark
2026年8月13日
DeepSWE 榜单(快照 2026-08-13)
ai-ml
benchmark
2026年8月13日
DSBench-FullStack 分数(快照 2026-08-13)
ai-ml
benchmark
2026年8月13日
Humanity's Last Exam 榜单(快照 2026-08-13)
ai-ml
benchmark
2026年8月13日
SuperCLUE 榜单(快照 2026-08-13)
ai-ml
benchmark
2026年8月13日
SWE-bench Verified 榜单(快照 2026-08-13)
ai-ml
benchmark
2026年8月13日
Terminal-Bench 榜单(4.0 与 2.1 快照)
ai-ml
benchmark
2026年8月13日
Toolathlon 榜单(快照 2026-08-13)
ai-ml
benchmark
2026年8月13日
Vibe Code Bench 榜单(快照 2026-08-12)
ai-ml
benchmark
2026年6月11日
Agents' Last Exam (ALE): A Benchmark for Economically Valuable Long-Horizon Agent Tasks
ai-ml
benchmark
multi-agent
2026年5月18日
WildClawBench — 真实长时域 Agent 评测基准
ai-ml
benchmark
multi-agent
evaluation
cli