跳到正文
Apple Machine Learning Research·· 2 天前

RISED: Rubrics for Agentic Multi-Environment Selection and Self-Distillation

RISED: Rubrics for Agentic Multi-Environment Selection and Self-Distillation

摘要

Training a single LLM agent jointly across diverse interactive environments has attracted increasing attention as a route to generalist agents. Existing curriculum and data-selection strategies often allocate training at the environment level or prioritize local reward-based signals, without explicitly considering relationships between current rollouts across environments for prompt-group selection. Meanwhile, as environments are learned at different rates, all-failure and all-success rollout groups can coexist within a batch, leaving those data without group-relative reward signals. Both…

当前提供采集摘要与原文入口。完整内容请阅读原文。

来源:Apple Machine Learning Research · machinelearning.apple.com