登录 注册
Coding Agents with an Obstacle-Aware Harness for Safe Robot Manipulation
👁 32 📚 28
Objective vs. Search: Decomposing What Makes a Good Tokeniser
👁 107 📚 18
ScienceBuddy: Recursive-in-Recursive Self-Improvement for Interactive Scientific Agents
👁 95 📚 14
Stellar Colosseum: A Many-Agent Harness for Long-Horizon Research in Mathematics and Theoretical Com...
👁 187 📚 25
Type Diversity Enables Transformers to Generalise Compositionally
👁 175 📚 30
Distance generalization in transformers: why bother with positional encoding?
👁 100 📚 23
数据 (Data) Scarcity and 模型 (Model) Sparsity: Mixtures-of-Experts Overfit More to Repeated 数据 (Data)
Data Scarcity and Model Sparsity: Mixtures-of-Experts Overfit More to Repeated Data
👁 202 📚 4
IdeaAMBIG: Benchmarking Implementation-Critical Gaps in Research-Idea Specifications
👁 163 📚 6
ReCite: Agentic Reasoning for Faithful Citation
👁 207 📚 1
Same Trajectory, Contradictory Rewards (ROBORMBENCH): Paraphrase Fragility in Vision Language Reward...
Same Trajectory, Contradictory Rewards (ROBORMBENCH): Paraphrase Fragility in Vision Language Reward...
👁 92 📚 12
WearableQA: A Benchmark for Health Reasoning over Real-World Wearable 数据 (Data)
WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data
👁 207 📚 7
ESPO: Error-Structured Prompt 优化 (Optimization) via Diagnose, Diversify, and Stabilize
ESPO: Error-Structured Prompt Optimization via Diagnose, Diversify, and Stabilize
👁 75 📚 11
User Feedback Provides a Unique Signal that LLMs Can not Detect
👁 90 📚 23
Efficient SWE Agent Benchmarking via Trajectory-Aware Evaluation
👁 220 📚 6
Context-Aware Interleaved Batching for WhisperX
👁 139 📚 24
A Formal Limitation on 学习 (Learning) Human Language From Textual Corpora
A Formal Limitation on Learning Human Language From Textual Corpora
👁 122 📚 0
SWE-Prime: Fewer Trajectories, Better Performance
👁 118 📚 24
WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution
👁 112 📚 30
CritICL: 推断 (Inference)-Time Weak-to-Strong Generalization from Small Language 模型 (Model) Failure Mo...
CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes
👁 125 📚 6
PlanSightRAG: A Visual-First Multimodal RAG for Automating Question Answering and Compliance Checkin...
👁 84 📚 3
海洋智能体 🌊
海洋智能体
AI科研助手 · 3665篇文献
你好!你正在浏览文献列表,我可以帮你筛选方向、推荐高引论文或解读某个研究领域。