登录 注册
CoCo-IR: Contextual Composed Image Retrieval
👁 108 📚 19
Perceptual Anchoring: Prototype-Guided Text Calibration for Training-free Open-Vocabulary Semantic S...
👁 111 📚 8
WorldExam: Benchmarking World 模型 (Model)s from Apparent Appearance to Inherent Reactivity
WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity
👁 93 📚 23
Toward Robust and 3D-Aware RGB-NIR Imaging in the Dark
👁 122 📚 28
PhiZero: A World 模型 (Model) Built Around Physical Language
PhiZero: A World Model Built Around Physical Language
👁 161 📚 21
ACE-数据 (Data)-0: Human-Centric Ambient Capture as Embodied 数据 (Data) Engine
ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine
👁 142 📚 27
ReToken: One Token to Improve Vision-Language 模型 (Model)s for Visual Retrieval
ReToken: One Token to Improve Vision-Language Models for Visual Retrieval
👁 131 📚 27
TurboVLA: Real-Time Vision-Language-Action 模型 (Model) at 32 Hz on an RTX 4090 with <1 GB VRAM
TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM
👁 135 📚 14
VetClaw: An Edge-Cloud Multimodal Agentic System for Veterinary Disease Screening
👁 29 📚 20
数据 (Data) Pyramid for Embodied Manipulation
Data Pyramid for Embodied Manipulation
👁 165 📚 19
Robot-Factored World 模型 (Model)s via Robot Rendering
Robot-Factored World Models via Robot Rendering
👁 127 📚 15
Unified Video Dense 预测 (Prediction) from Disjoint 数据 (Data)
Unified Video Dense Prediction from Disjoint Data
👁 87 📚 17
Streaming Multi-Agent Autoregressive Diffusion 模型 (Model) with World State Registers
Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers
👁 85 📚 13
ATSplat: Compact Feed-forward 3D Gaussian Splatting with Adaptive Token Expansion
👁 170 📚 30
Appearance Pointers -- Multimodal Region Control of Diffusion Transformers
👁 196 📚 20
Simple Domain Generalization for Strong Pixel-Level Image Tampering Detection in Modern VLMs
👁 104 📚 21
Knowing the Self, Understanding the World: A Dual-Cognition Benchmark for UAV Spatio-temporal Reason...
👁 27 📚 13
Hierarchical Denoising For Multi-Step Visual Reasoning
👁 193 📚 29
VideoRAE: Taming Video Foundation 模型 (Model)s for Generative 模型 (Model)ing via Representation Autoen...
VideoRAE: Taming Video Foundation Models for Generative Modeling via Representation Autoencoders
👁 209 📚 26
FlowWAM: Optical Flow as a Unified Action Representation for World Action 模型 (Model)s
FlowWAM: Optical Flow as a Unified Action Representation for World Action Models
👁 111 📚 13
海洋智能体 🌊
海洋智能体
AI科研助手 · 2983篇文献
你好!你正在浏览文献列表,我可以帮你筛选方向、推荐高引论文或解读某个研究领域。