登录 注册
找到 620 个结果

Zero-Flow Two-Sample Tests

We propose a new approach to two-sample testing for deciding whether two sets of samples are drawn from the same distribution. The test is built on a statistical discrepancy based on the zero-flow cri...

👤 Yakun Wang | Leyang Wang | Song Liu | Ta... 📰 arXiv 📅 2026 👁 28 📚 13

Trojan horse hunt in deep forecasting models: Insights from the European Space Agency competition

Forecasting plays a crucial role in modern safety-critical applications, such as space operations. However, the increasing use of deep forecasting models introduces a new security risk of trojan horse...

👤 Krzysztof Kotowski|Ramez Shendy|Jakub Na... 📰 arXiv 📅 2026 👁 532 📚 12

Behavioral Fingerprints for LLM Endpoint Stability and Identity

The consistency of AI-native applications depends on the behavioral consistency of the model endpoints that power them. Traditional reliability metrics such as uptime, latency and throughput do not ca...

👤 Jonah Leshin, Manish Shah, Ian Timmis, D... 📰 arXiv 📅 2026 👁 487 📚 12

CATEKAPPA: An R Shiny Application for Design and Analysis of Consistency Tests Based on the Kappa Statistic for Categorical Responses

The kappa statistic is the most widely used measure of inter-rater agreement for categorical data. Despite its popularity, applied researchers often encounter two major hurdles: (i) determining the sa...

👤 Zheng Gai | Li Xincheng | Jiang Wangying... 📰 arXiv 📅 2026 👁 169 📚 12

Towards Computational Provenance: Carrying Causal-State Evidence in Generated Text

A language model's output does not by itself provide verifiable evidence about the internal computation that produced it. We study computational provenance: whether generated text can carry detectable...

👤 Benjamin Belay 📰 arXiv 📅 2026 👁 158 📚 12

SceneCritic: A Symbolic Evaluator for 3D Indoor Scene Synthesis

Large Language Models (LLMs) and Vision-Language Models (VLMs) increasingly generate indoor scenes through intermediate structures such as layouts and scene graphs, yet evaluation still relies on LLM ...

👤 Kathakoli Sengupta | Kai Ao | Paola Casc... 📰 arXiv 📅 2026 👁 140 📚 12

A Design-Based Approach to Testing and Inference in (Quasi-)Experiments with Spillovers

Economic policies rarely affect only their direct targets. To study these spillovers, researchers summarize who else was treated with a simple exposure measure, such as the share of treated neighbors ...

👤 Yechan Park 📰 arXiv 📅 2026 👁 140 📚 12

Lead-Lag Relationships in Financial Markets: A Comparison of Multiple Clustering Algorithms

Lead-lag relationships are widely used in financial time series, and many clustering algorithms based on them have been developed. The traditional DTW-KMedoids algorithm performs well both on the synt...

👤 Ruichen Deng | Yichi Zhang 📰 arXiv 📅 2026 👁 115 📚 12

Score-Based Ideal Observer Approximation via Denoising Score Matching for Signal-Known-Exactly Detection Tasks

The Bayesian Ideal Observer (IO) establishes the theoretical upper bound on task performance for binary detection tasks. However, analytical computation of the IO test statistic is generally intractab...

👤 Weimin Zhou 📰 arXiv 📅 2026 👁 68 📚 12

Functional CLT for general sample covariance matrices

This paper studies the central limit theorems (CLTs) for linear spectral statistics (LSSs) of general sample covariance matrices, when the test functions belong to $C^3$, the class of functions with c...

👤 Jian Cui|Zhijun Liu|Jiang Hu|Zhidong Bai 📰 arXiv 📅 2026 👁 341 📚 11

HiPath: Hierarchical Vision-Language Alignment for Structured Pathology Report Prediction

Pathology reports are structured, multi-granular documents encoding diagnostic conclusions, histological grades, and ancillary test results across one or more anatomical sites; yet existing pathology ...

👤 Ruicheng Yuan|Zhenxuan Zhang|Anbang Wang... 📰 arXiv 📅 2026 👁 284 📚 11

Ideological Bias in LLMs' Economic Causal Reasoning

Do large language models (LLMs) exhibit systematic ideological bias when reasoning about economic causal effects? As LLMs are increasingly used in policy analysis and economic reporting, where directi...

👤 Donggyu Lee | Hyeok Yun | Jungwon Kim | ... 📰 arXiv 📅 2026 👁 226 📚 11

An Efficient Likelihood Ratio Test for Online Changepoint Detection in the Presence of Autocorrelation

Changepoint detection methods have seen considerable development in recent years, with online algorithms capable of identifying structural changes in streaming data in near real time. However, the maj...

👤 Yuntang Fan | Paul Fearnhead | Idris A. ... 📰 arXiv 📅 2026 👁 210 📚 11

SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?

Modern software systems accumulate technical debt over decades of development, which makes migration expensive and largely manual. As coding agents become increasingly capable at bug fixing, can they ...

👤 Deyao Hong | Yizhe Chi | Wenyi Li | Xiao... 📰 arXiv 📅 2026 👁 175 📚 11

From Manipulation to Mistrust: Explaining Diverse Micro-Video Misinformation for Robust Debunking in the Wild

The rise of micro-videos has reshaped how misinformation spreads, amplifying its speed, reach, and impact on public trust. Existing benchmarks typically focus on a single deception type, overlooking t...

👤 Zhi Zeng | Yifei Yang | Jiaying Wu | Xul... 📰 arXiv 📅 2026 👁 168 📚 11

ClinHallu: A Benchmark for Diagnosing Stage-Wise Hallucinations in Medical MLLM Reasoning

Building trustworthy medical multimodal large language models (MLLMs) is critical for reliable clinical decision support. Existing medical hallucination benchmarks mainly focus on data collection, but...

👤 Sicheng Yang | Hangjie Yuan | Wenjun Zha... 📰 arXiv 📅 2026 👁 160 📚 11

Path-Explosive Behaviour in Economic Time Series: A Realization-Centred Exploratory Framework

We propose a descriptive, realization-centred framework for detecting and characterising explosive and co-explosive behaviour in economic time series, which we term path-explosive behaviour. Departing...

👤 José Francisco Perles-Ribes 📰 arXiv 📅 2026 👁 148 📚 11

EEVEE: Towards Test-time Prompt Learning in the Real World for Self-Improving Agents

In this paper, we propose EEVEE, the first multi-dataset test-time prompt learning framework for LLM agents, enabling test-time prompt learning under real-world task streams. Existing methods are larg...

👤 Weixian Xu | Shilong Liu | Mengdi Wang 📰 arXiv 📅 2026 👁 114 📚 11

Evaluating HWE and Association in Genome Wide Association Studies: A Unified Procedure

In genome wide association studies (GWASs) based on a case-control design, single nucleotide polymorphisms (SNPs) are typically evaluated for an association test and a Hardy-Weinberg equilibrium (HWE)...

👤 Stefan Böhringer | Hajo Holzmann 📰 arXiv 📅 2026 👁 80 📚 11

Derivative-Informed Operator Learning for Finance: On-the-Fly Greeks, Surfaces, Hedging, and Control

Financial decision systems require fast surrogate models for pricing, calibration, hedging, XVA, stress testing, and portfolio optimization. Standard neural surrogates reproduce prices or risk quantit...

👤 Miquel Noguer I Alonso 📰 arXiv 📅 2026 👁 77 📚 11
海洋智能体 🌊
海洋智能体
AI科研助手 · 3725篇文献
你在高级搜索页面,告诉我你想找什么方向的文献,我来帮你定位。