登录 注册
找到 620 个结果

Self-Evolving World Models for LLM Agent Planning

World models offer a principled way to equip long-horizon LLM agents with foresight: predictions of action consequences before execution. However, unreliable foresight can be ignored, misused, or even...

👤 Xuan Zhang | Wenxuan Zhang | See-Kiong N... 📰 arXiv 📅 2026 👁 43 📚 19

Optimal Adversarial Testing: Extracting Honest Test Results from Dishonest Test Takers

In applications, it is often required to test objects or people to determine their qualities in terms of certain metrics. However, besides being naturally noisy, the test results can be corrupted by a...

👤 Owen Cox | April Xu | Weiyu Xu 📰 arXiv 📅 2026 👁 43 📚 19

Goodness-of-Fit Tests and Calibration Machine-Learning Algorithms for Logistic Regression with Sparse Data

Assessing the goodness-of-fit of a logistic regression model is a critical prerequisite before the model is used for inference. However, goodness-of-fit (GOF) tests such as the chi-square and deviance...

👤 Ebrahim Khaled Ebrahim 📰 arXiv 📅 2026 👁 39 📚 19

LongMemEval-V2: Evaluating Long-Term Agent Memory Toward Experienced Colleagues

Long-term memory is crucial for agents in specialized web environments, where success depends on recalling interface affordances, state dynamics, workflows, and recurring failure modes. However, exist...

👤 Di Wu | Zixiang Ji | Asmi Kawatkar | Bry... 📰 arXiv 📅 2026 👁 37 📚 19

Leveraging Phytolith Research using Artificial Intelligence

Phytolith analysis is a crucial tool for reconstructing past vegetation and human activities, but traditional methods are severely limited by labour-intensive, time-consuming manual microscopy. To add...

👤 Andrés G. Mejía Ramón|Kate Dudgeon|Nina ... 📰 Artificial Intelligence 📅 2026 👁 378 📚 18

Modeling diesel output particulate matter as the Ornstein-Uhlenbeck process

Diesel engine particulate matter (PM) is one of the most challenging emission constituents to predict. As engines become cleaner and emissions levels drop, manufacturers need reliable methods to quant...

👤 Maxwell Bolt|Alex Alberts|Akash S. Desai... 📰 arXiv 📅 2026 👁 367 📚 18

Beyond Passive Aggregation: Active Auditing and Topology-Aware Defense in Decentralized Federated Learning

Decentralized Federated Learning (DFL) remains highly vulnerable to adaptive backdoor attacks designed to bypass traditional passive defense metrics. To address this limitation, we shift the defensive...

👤 Sheng Pan, Niansheng Tang 📰 arXiv 📅 2026 👁 236 📚 18

Effects of motion cueing on longitudinal acceleration perception in a driving simulator

The driveability of a new heavy-truck driveline is traditionally assessed using physical prototypes. Enabling early evaluation of the driving experience in a human-in-the-loop driving simulator using ...

👤 Erik Gustaf Lilljebjörn | Sogol Kharrazi... 📰 arXiv 📅 2026 👁 219 📚 18

1-Lipschitz Neural Networks on Hadamard Manifolds

Controlling the Lipschitz constant of a neural network is a standard way to promote robustness and stability. Most existing constraining strategies are designed for Euclidean spaces. In this work, we ...

👤 Davide Murari | Marta Ghirardelli | Ben ... 📰 arXiv 📅 2026 👁 191 📚 18

Estimation and Hypothesis Testing of Fixed Effects Models-Based Uncertainty for Factor Designs

To analyze the uncertain data frequently encountered in practice, this paper proposes novel fixed-effects models that incorporate an uncertain measure to investigate variables of interest and nuisance...

👤 Fan Zhang, Zhiming Li 📰 arXiv 📅 2026 👁 157 📚 18

Bellman-Ford in Almost-Linear Time

We consider the single-source shortest paths problem on a directed graph with real-valued (possibly negative) edge weights and solve this problem in $m^{1+o(1)}$ time.

👤 Isaac M. Hair | George Z. Li | Jason Li ... 📰 arXiv 📅 2026 👁 142 📚 18

In-Place Test-Time Training

The static ``train then deploy" paradigm fundamentally limits Large Language Models (LLMs) from dynamically adapting their weights in response to continuous streams of new information inherent in real...

👤 Guhao Feng | Shengjie Luo | Kai Hua | Ge... 📰 arXiv 📅 2026 👁 135 📚 18

Clinical Reasoning Graphs: Structured Evaluation of LLM Diagnostic Reasoning Reveals Competence Without Consistency

Modern large language models (LLMs) reach 60-70% diagnostic accuracy on complex clinical case benchmarks, but accuracy alone cannot distinguish stable clinically-grounded reasoning from pattern matchi...

👤 Nisarg A. Patel 📰 arXiv 📅 2026 👁 130 📚 18

AI Economist Agent: An Agentic Framework for Model-Grounded Economic Analysis with RAG, Knowledge Graphs, and Large Language Models

We propose a model-grounded RAG-based AI economist with an agentic framework for economic scenario analysis using large language models (LLMs) and knowledge graphs. While LLMs can generate fluent econ...

👤 Masahiro Kato 📰 arXiv 📅 2026 👁 127 📚 18

Kernelized Stein Discrepancy for Goodness-of-Fit Tests and Stein Sampling in R

Stein's method constructs computable discrepancies between a target distribution and a candidate distribution without requiring the target distribution's normalizing constant. These discrepancies supp...

👤 Junhao Gao | Ery Arias-Castro 📰 arXiv 📅 2026 👁 119 📚 18

Shrinkage Regularization for (Non)Linear Serial Dependence Test

This paper introduces a regularized test of the null hypothesis of the absence of linear and nonlinear serial dependence for high-dimensional non-Gaussian time series. Our approach extends the portman...

👤 Francesco Giancaterini|Alain Hecq|Joann ... 📰 arXiv 📅 2026 👁 115 📚 18

critband: A Python Package for Critical Bandwidth Analysis of Multimodal Distributions

Multimodal density estimation is a fundamental problem in scientific computing, but Python has lacked a cohesive implementation of critical bandwidth analysis and related mode-counting tools. We prese...

👤 Ruiyu Zhang | Qihao Wang 📰 arXiv 📅 2026 👁 103 📚 18

A Quasi-Regression Method for the Mediation Analysis of Zero-Inflated Single-Cell Data

Recent advances in single-cell technologies have advanced our understanding of gene regulation and cellular heterogeneity at single-cell resolution. Single-cell data contain both gene expression level...

👤 Seungjun Ahn | Donald Porchia | Panos Ro... 📰 arXiv 📅 2026 👁 83 📚 18

An Augmented Rating System for Test cricket: adapting Glicko's model

ICC's current ranking system does not adequately account for key contextual factors such as home advantage, toss impact and scheduling imbalances; leading to inconsistencies in team evaluation in Test...

👤 Rhitankar Bandyopadhyay|Diganta Mukherje... 📰 arXiv 📅 2026 👁 65 📚 18

Optimal Stratified Allocation for Rare-Event Onset Forecasting in Dependent Sequences

Let a finite population of n labelled examples carry a class-weighted loss, with pi*n in a rare positive class weighted by N0/N1. We study estimation of total risk from a subsample K << n under design...

👤 Jaskaran Singh 📰 arXiv 📅 2026 👁 63 📚 18
海洋智能体 🌊
海洋智能体
AI科研助手 · 3725篇文献
你在高级搜索页面,告诉我你想找什么方向的文献,我来帮你定位。