📰 科技脉搏
2026-09-04
🔥 AI & 科技
Many recurring text functions are easy to describe but difficult to implement with rules, while calling a large remote model for every input introduces repeated cost, latency, and dependency on a provider. We present compile by training, which turns a natural-language specification into a reusable n
📅 09-03📎 arXiv AI Papers
Language-model judges now gate training data, score generations, and drive leaderboards. The judge is then a measurement instrument, resting on one rarely stated assumption: the same request, sent to the same model name, reads the same tomorrow. We audited that assumption in two preregistered campai
📅 09-03📎 arXiv AI Papers
Evolutionary prompt optimizers such as GEPA suffer from prompt bloat: each iteration appends rules and caveats, producing prompts up to 3$\times$ longer yet no more accurate. We trace this to three deficiencies – incomplete error observation, limited search diversity, and unreliable selection – and
📅 09-03📎 arXiv AI Papers
Video editing spans diverse editing paradigms, yet achieving high-quality instruction-guided and subject-guided editing within a single unified framework remains challenging. We introduce EditVid, a training-free framework combining sparse causal memory for local coherence, correspondence-based post
📅 09-03📎 arXiv AI Papers
Weakly-Supervised Dense Video Captioning aims to localize and describe multiple events in untrimmed videos given only an ordered set of event-level captions per video. Recent work synthesizes auxiliary transition captions via LLM to provide additional vision-language alignment, but these captions la
📅 09-03📎 arXiv AI Papers
Gaps remain in our understanding of how large language models (LLMs) acquire knowledge during pre-training. We posit that auxiliary views, reformulations of knowledge, are causally helpful for learning. We design controlled experiments to isolate this. First, we confirm that repetition is necessary
📅 09-03📎 arXiv AI Papers
Explaining why a specific outcome occurred, and which inputs deserve the blame or credit, is central to philosophical, scientific, and policy analysis. Existing tools split into two camps. The theory of actual causality (AC) gives principled verdicts, but only for toy-sized models, because computing
📅 09-03📎 arXiv AI Papers
On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a teacher. Existing work has mainly studied its algorithmic behavior, leaving the role of training data unclear. We examine this role at the data-minimal limit by training on a single query. One-
📅 09-03📎 arXiv AI Papers
Multi-agent AI science ecosystems rely on agents possessing tools that allow them to communicate, coordinate, and build on each other’s work. Yet this shared infrastructure can also introduce vulnerabilities by creating a substrate for the contagious spread of unintended and undesirable behaviors. W
📅 09-03📎 arXiv AI Papers
本文提出SWE-Gate评估框架,指出现有软件工程智能体仅通过功能测试不足以证明实际能力,因为测试可能被针对性地绕过。SWE-Gate通过引入补丁语义校验与额外约束,确保解决方案真正满足任务意图。该工作为智能体代码生成提供了更严谨的评估标准,推动软件工程智能体的可靠发展。
📅 09-03📎 arXiv AI Papers
General Science
📅 09-03📎 Google AI Blog
General Science
📅 09-03📎 Google AI Blog
💰 财经 & 投资
由 Hermes 自动生成 | 查看历史摘要