强智能体进行自主机器学习工程需要多少支撑框架?

Recent autonomous machine learning engineering (MLE) agents have made significant progress on public leaderboards. Often motivated by progress stagnation over long-horizon cycles and limited Large Language Model (LLM) primitives, modern MLE agents are deployed on top of increasingly elaborate machinery:multi-agent orchestrators, dedicated retrieval subagents, and more. While such harnesses expand, the use of more primitive but improved coding agents - where LLMs have direct access to the execution environment through read, write, and bash primitives - has received little attention in the field. In this paper we find that, under an equal time budget and the same frontier LLM backbone, open-source state-of-the-art harnesses provide no advantages over a single session of a minimal-harness coding agent baseline, pointing to the backbone as the primary driver for performance. Via a series of large-scale systematic ablation studies, we argue that the machinery layers become redundant in the coding agent setting. We conclude that the effort spent elaborating hand-crafted harnesses around strong models yields poor returns for current MLE benchmarks.
评论
    公告

    AI千集是一个专注于数字员工的智能平台
    在这里您可以获得本平台自训练的
    数字员工
    和小伙伴一起玩转AI,做自己的AI数字员工
    来AI千集,赋能智慧快人一步
    扫一扫,快速获取解决方案与报价
    立即咨询

    AISet
    连接写作与商家获客的桥梁
    让写作融入商家运营
    登陆小程序
    AI数字人随身守护
    智慧管理更高效
    获客能力悄然升级

    AISet

    积分排行