📡 每日AI快讯 · 300 条

实时更新AI行业最新资讯、热点、融资和产品动态,数据自动从各大AI媒体RSS抓取

09月02日 · 周三 2026-09-02
技术前沿InfoQ AI 00:22

AWS 发布 Aws-Bench,用于评估云任务中的智能代理

点击查看原文>

09月01日 · 周二 2026-09-01
技术前沿InfoQ AI 23:18

机器人的下一站,或许是乐高化

点击查看原文>

技术前沿InfoQ AI 23:00

通过操作 DRAM 控制器寄存器可突破 CPU 内存隔离机制

点击查看原文>

技术前沿InfoQ AI 20:54

从2999份作品出发:GOAI四大赛道的AI创新观察

点击查看原文>

技术前沿InfoQ AI 18:26

智元的“夺冠方法论”

点击查看原文>

行业动态IT之家 16:32

岚图汽车 8 月交付 13003 辆,1-8 月累计交付破 10 万辆

IT之家 9 月 1 日消息,岚图汽车今日宣布 2026 年 8 月交付汽车 13,003 辆,2026 年 1-8 月累计交付 102,456 辆,同比增长 25%。IT之家注意到,8 月岚图汽车新增 7 家岚图空间店,4 家全功能用户中心,覆盖 9 个城市。岚图汽车称,截至 2026 年 8 月,累计为用户充电超 7.74 亿度,减少碳排放超 24.2 万吨。岚图品牌站累计上线 146 座,已触达 33 城。

行业动态IT之家 16:30

知情人士回应抖音突发算法推荐错乱:原因为服务器异常

IT之家 9 月 1 日消息,今天(1 日)早些时候,“抖音崩了”及“抖音推荐难看”等热搜词条引发关注。据《科创板日报》援引知情人士消息称,推荐算法错乱是因为抖音平台突发服务器异常。报道称,据大量网友反馈,此次服务异常主要表现为三大技术故障:其一,推荐算法机制疑似失灵,用户刷到的内容与日常偏好严重不符,大量推送不感兴趣、低相关度的视频,甚至出现“误入中老年频道”的既视感,部分用户反映关注页内容无法刷新;其二,内容流出现重复性紊乱,同一个视频在信息流中被反复推送多次,无法获取新内容;其三,基础播放功能受阻,视频播放频繁卡顿、加载画面持续转圈,甚至在网络连接正常的情况下,系统仍错误提示“无网络连接”。另据每日经济新闻,抖音人工客服表示,已有大量用户反馈同类问题,平台已经在排查优化中,后续优化完成就可以恢复。不过,在回复“是否有排查到故障原因”时,抖音客服仅称“目前是在排查优化中,就是系统优化导致的。”对于具体恢复时间,客服称相关人员已经在处理,暂未收到具体时间的通知,无法给出确切时间表。相关阅读:《网友称抖音推荐内容“误入中老年频道”,客服回应在加急排查》

行业动态IT之家 16:28

汉王 Clear6 Pro 二代黑白墨水屏电纸书将于 9 月 9 日上市:全贴合纯平设计、搭载 3000mAh 电池

IT之家 9 月 1 日消息,汉王宣布旗下 Clear6 Pro 二代黑白墨水屏电纸书将于 9 月 9 日上市,该产品主打超薄全贴合纯平设计和长续航能力。综合官方预热,该机提供陨石黑、浅草紫、皓月白三种配色,配备一块 6 英寸 300PPI 黑白墨水屏面板,匹配冷暖双色前光灯,内置 3000mAh 电池。作为参考,现款 Clear6 Pro 电纸书于 2024 年 6 月上市,首发价 998 元,该机尺寸为 109 x 148.5 x 7mm,重量 172 克,配备 6 英寸 Eink 300 PPI 面板,搭载“四核 CPU”,内置 2GB RAM + 32GB 存储空间,同样内置 3000 毫安时电池,号称可实现 120 小时阅读时长 / 60 天待机,IT之家整理该机具体参数信息如下:

行业动态IT之家 16:25

LG gram Book AI 2026 笔记本海外发布:gram 系列首款 14 英寸机型,英特尔 WCL 平台

IT之家 9 月 1 日消息,LG gram Book AI 2026(14U40V)笔记本电脑今天在韩国发布,新品定位大众市场,是 gram 系列首款 14 英寸机型,搭载英特尔 Wildcat Lake 平台。据介绍,这款笔记本机身采用具有柔和光泽的铝合金,兼顾质感和耐用性。整机重量仅 1.29kg,厚度 14.9mm,较为便携。规格方面,这款笔记本搭载 2026 款英特尔酷睿 Series 3 处理器,以及 61.2Wh 电池,最长续航时间可达 30.5 小时。该机提供低噪音模式,开启后可让用户在图书馆、办公室等场所进行简单文档处理或观看视频,而不是产生明显噪音。同时,该笔记本提供 HDMI、USB-A、USB-C、RJ-45 等接口,可连接各种外部设备。该机预装 Windows 操作系统,可通过 LG gram Link 软件连接不同品牌手机、平板,从而方便地传输照片视频文件。此外,该笔记本还带有摄像头物理遮挡滑盖,可保护用户隐私。该机将于 9 月 7 日开售,起售价为 145 万韩元(IT之家注:现汇率约合 7,089 元人民币)。

行业动态IT之家 16:21

阿里千问宣布升级学习辅助功能:新增教材同步作文辅导、扩展高中和大学数理化拍题讲解能力

IT之家 9 月 1 日消息,阿里宣布升级旗下千问学习辅助功能,新增与教材同步的单元作文辅导、初中语文和数学教材精讲,并进一步扩展高中和大学数理化拍题讲解能力。相关功能均免费开放。阿里表示,在此次升级中,相比直接给出答案,千问进一步强化对学习过程的引导,让学生围绕“为什么”“怎么想”继续往下学。例如,在小学作文场景中,千问会围绕题目逐步提问,引导孩子自己确定写什么、怎么写,再梳理素材和结构。比如五年级上册第一单元的作文“我的心爱之物”,从“你最想写哪件心爱之物”“它为什么对你特别”“你和它之间发生过什么难忘的事”等问题一路追问,帮助学生找到最值得写的细节,把自己和心爱之物之间的感情写得更具体、更动人。初中教材讲解则进一步强调“哪里没懂,就从哪里讲”。“小讲堂”此次新增初中语文、数学教材精讲,学生可以从具体的一个章节、一个知识点开始学习。比如初三开始接触“一元二次方程”时,千问会带着学生回顾之前学过的“一元一次方程”,再对照理解新知识。教材讲解中,遇到不理解的概念、句子或解题步骤,可以随时打断并继续追问,千问会根据具体问题调整讲解内容和节奏。到了高中、大学,题目往往步骤更多、推导更复杂。此次升级进一步把“小讲堂”的互动讲解能力扩展到高中和大学的数理化题目。推导过程中如果卡在某一步,可以直接追问“为什么这里要这样变形”“下一步是怎么想到的”,再沿着这一处继续往下学。千问 App 智能学习业务负责人程飞表示,AI 学习产品不能只关注一道题有没有答对,更重要的是学生有没有真正理解知识,遇到类似的新问题时能不能自己解决。比给出答案更重要的,是让学生知道为什么、学会怎么想。千问希望通过分步引导和持续追问,让学生在一次次理解和解决问题的过程中,逐步形成独立学习和思考的能力。

行业动态IT之家 16:16

奇瑞集团 8 月汽车销量 28 万辆,成首个累计出口突破 700 万辆的中国车企

IT之家 9 月 1 日消息,奇瑞集团今日宣布,集团 8 月汽车销量 28 万辆,同比增长 15.4%;其中新能源 12 万辆,同比增长 69.8%,且连续 5 个月突破 10 万辆。奇瑞集团还称,集团 8 月出口 196,984 辆,同比增长 52.1%。集团历史累计出口汽车达 718 万辆,成为首个累计出口突破 700 万辆的中国车企。奇瑞汽车 8 月销售 262,990 辆,同比增长 14.0%。据IT之家了解,奇瑞集团旗下拥有奇瑞、星途、捷途、iCAR、智界及国际品牌等多个品牌,产品覆盖从家用燃油车到高端新能源、从城市 SUV 到越野车型的全场景市场。

行业动态IT之家 16:15

中国电信旗下天翼云盘:9 月 27 日起调整数据迁移规则,迁移至新号码将回收原号码福利空间

IT之家 9 月 1 日消息,中国电信旗下天翼云盘发布公告,宣布将于 9 月 27 日起调整云盘数据迁移规则,原号码迁移至新号码时,原号码上的天翼云盘福利空间会被回收。IT之家附官方公告如下:尊敬的用户:感谢您对天翼云盘的支持!一直以来天翼云盘为海量用户提供稳定的个人云存储服务。为积极响应践行国家绿色低碳发展要求,提高资源利用率,同时也为了保障用户账号和数据安全,自 2026 年 9 月 27 日起,数据迁移规则将作如下变更:1、数据迁移前系统会判断原号码与新号码当前的状态,如果存在以下情况,则无法迁移:(1)新号码未注册过天翼云盘(2)新号码的总空间小于原号码总空间(3)90 天内有过数据迁移成功记录(4)原号码或新号码任意一个处于异常状态,包括但不限于账号冻结、账号注销等2、数据迁移后,新号码上的天翼云盘当前数据将删除,新号码享有原号码的数据、空间、部分会员权益,具体如下,请谨慎操作。(1)新号码上的天翼云盘的数据、空间将删除,原号码上的天翼云盘的数据、空间将迁移到新号码上,福利空间会被回收,已创建的文件分享外链将被标记为失效状态。(2)新号码上的天翼云盘官方渠道会员权益 * 将删除,原号码上的天翼云盘官方渠道会员权益将迁移到新号码上。(3)新号码上的天翼云盘中国电信渠道会员权益 * 将保留在新号码,原号码上的天翼云盘中国电信渠道会员权益将保留在原号码。* 天翼云盘中国电信渠道会员权益:包括但不限于中国电信营业厅、中国电信网上营业厅客户端等订购的天翼云盘会员权益* 天翼云盘官方渠道会员权益:包括但不限于天翼云盘客户端、官网会员中心等订购的天翼云盘会员权益3、每个账号 90 天内只能进行一次数据迁移,时间自上一次迁移申请提交成功当日起计算,第 91 天才可再次申请,请谨慎操作。4、数据迁移后导致新号码原有的企业云数据删除,新号码将拥有原号码所有企业云盘数据和信息。5、数据迁移前后用户的会员自动续费产品签约关系维持不变。天翼云盘的可持续发展离不开每位用户的支持与理解,我们也希望通过云盘的竭诚服务,为您带去一份长久的美好陪伴。2026 年 8 月 24 号天翼云盘团队

行业动态IT之家 16:13

国产奔驰长轴距 GLE 下线:3115mm 轴距、3.0T 直列六缸发动机加持,月内上市

IT之家 9 月 1 日消息,“北京奔驰”公众号今天(1 日)上午发文宣布,全新国产梅赛德斯-奔驰长轴距 GLE SUV 在“长城脚下的智能工厂”—— 北京奔驰顺义工厂正式下线。根据官方此前的规划,新车将于 9 月正式上市。根据介绍,全新长轴距 GLE SUV 以现代设计语言诠释奔驰独特设计哲学。配备“双星徽”前大灯,构成车头独特的视觉标志,发光中央星标与发光格栅边框则进一步增强气场。中国专属加长设计将整车轴距提升至超 3.1 米,配合大五座布局,标配可滑动开启全景天窗,还会提供座椅振动按摩功能。新车搭载奔驰首创的 AR 平视显示系统,将导航关键信息融入真实环境,打造沉浸式导航;全新城区及高速领航辅助驾驶系统,将在中国市场实现“车位到车位”的智能辅助驾驶与辅助泊车。动力方面,新车搭载 3.0 升排量直列六缸发动机,标配电动辅助增压器,峰值扭矩提升 12% 至 560 牛 · 米,在 17 千瓦 ISG 智能电机及 48 伏轻混系统的加持下,令动力随叫随到并提升燃油经济性。据IT之家了解,全新长轴距 GLE 标配 4MATIC 智能四驱系统,选择越野模式后,AIRMATIC 空气悬挂可额外升高 30mm,最小离地间隙达到 271mm。相关阅读:《奔驰全新长轴距 GLE SUV 全球首秀:轴距 3115mm,9 月正式上市》

行业动态IT之家 16:08

日企 Newtech 推出 Ness4300 四盘位 NAS,搭载专有 RAID 控制器

IT之家 9 月 1 日消息,日本企业 Newtech 当地时间今日推出 Ness4300 NAS。其配备 4 个 3.5" SATA 盘位,另有 240GB 的 M.2 PCIe NVMe SSD 用于运行 Windows Server IoT 2025 for Storage Standard 操作系统。与同级别竞品多采用软件 RAID 方案不同的是,Ness4300 配备了专有硬件 RAID 控制器 Condor,在保障数据管理安全性的同时可降低处理器负载。Ness4300 标配英特尔奔腾 Gold G7400(2 核 "Alder Lake")处理器,另有至强 E-2436(6 核 "Raptor Lake" 可选)。其内置 DDR5 UDIMM 内存插槽,提供 2 个 1GbE RJ45、4 个 USB,内置 8cm 风扇。

行业动态IT之家 16:07

小米 2026 年 9 月服务周开启,69 款手机可 8 折换电池

IT之家 9 月 1 日消息,小米现已开启 2026 年 9 月超级服务周活动,活动时间即日起至 9 月 7 日 24:00 结束,69 款手机电池可享换新优惠,IT之家整理机型信息如下:小米Xiaomi 13 电池升级:活动价 151.20 元Xiaomi 13 Pro 电池升级:活动价 151.20 元Xiaomi 13 Ultra 电池升级:活动价 151.20 元小米 10:活动价 127.20 元小米 10S:活动价 127.20 元小米 10 Pro:活动价 159.20 元小米 10 青春版:活动价 47.20 元小米 10 至尊纪念版:活动价 159.20 元小米 10 至尊纪念透明版:活动价 159.20 元小米 11:活动价 103.20 元小米 11 Pro:活动价 175.20 元小米 11 Ultra:活动价 175.20 元Xiaomi 12:活动价 127.20 元Xiaomi 12X:活动价 127.20 元Xiaomi 12 Pro:活动价 127.20 元Xiaomi 12 Pro 天玑版:活动价 127.20 元Xiaomi 12S:活动价 127.20 元Xiaomi 12S Pro:活动价 127.20 元Xiaomi 13:活动价 111.20 元Xiaomi 13 Pro:活动价 111.20 元Xiaomi 14:活动价 127.20 元Xiaomi 14 Pro:活动价 127.20 元Xiaomi 14 Ultra:活动价 127.20 元Xiaomi 15:活动价 127.20 元Xiaomi 15 Pro:活动价 127.20 元Xiaomi 15S Pro:活动价 127.20 元Xiaomi MIX 4:活动价 159.20 元Xiaomi MIX Fold:活动价 174.40 元Xiaomi MIX Fold 2:活动价 174.40 元Xiaomi MIX Fold 3:活动价 174.40 元Xiaomi MIX Flip:活动价 174.40 元Xiaomi Civi 2:活动价 127.20 元Xiaomi Civi 3:活动价 127.20 元Xiaomi Civi 4 Pro:活动价 127.20 元Xiaomi Civi 5 Pro:活动价 127.20 元红米 REDMIREDMI Turbo 3:活动价 127.20 元REDMI Turbo 4:活动价 127.20 元REDMI Turbo 4 Pro:活动价 127.20 元REDMI 9A:活动价 47.20 元REDMI 10A:活动价 47.20 元REDMI 14C:活动价 79.20 元REDMI Note 10 Pro:活动价 127.20 元REDMI Note 11 Pro:活动价 127.20 元REDMI Note 11 Pro+:活动价 127.20 元REDMI Note 11T Pro:活动价 119.20 元REDMI Note 11T Pro+:活动价 127.20 元REDMI Note 12T Pro:活动价 119.20 元REDMI Note 12 Turbo:活动价 127.20 元REDMI Note 13:活动价 103.20 元REDMI Note 13 Pro:活动价 127.20 元REDMI Note 13 Pro+:活动价 127.20 元REDMI Note 14 Pro:活动价 103.20 元REDMI Note 14 Pro+:活动价 127.20 元REDMI Note 15R:活动价 127.20 元REDMI K30 Pro:活动价 63.20 元REDMI K40:活动价 103.20 元REDMI K40 游戏增强版:活动价 127.20 元REDMI K40S:活动价 103.20 元REDMI K40 Pro:活动价 103.20 元REDMI K40 Pro+:活动价 103.20 元REDMI K50:活动价 127.20 元REDMI K50 电竞版:活动价 127.20 元REDMI K50 Pro:活动价 127.20 元REDMI K60:活动价 127.20 元REDMI K60E:活动价 127.20 元REDMI K70E:活动价 127.20 元REDMI K70 至尊版:活动价 159.20 元REDMI K80:活动价 127.20 元REDMI K80 至尊版:活动价 159.20 元小米服务提醒,电池性能的非正常性衰退,例如滥用、过度充电、极端温度、外部撞击等因素,需及时改正或维修,日常使用时应注意科学充电、避免高温 / 低温等,以提高电池寿命。

行业动态IT之家 16:07

英伟达 GPU 帧插值支持并入 FFmpeg,用户可使用帧生成提高视频帧率

IT之家 9 月 1 日消息,据科技媒体 Phoronix 前天报道,开源多媒体框架项目 FFmpeg 日前宣布,英伟达 GPU 加入帧插值功能已通过 Vulkan API 集成到 FFmpeg。本次合并意味着,FFmpeg 可使用英伟达 GeForce RTX GPU,在不经过传统图形渲染流程的情况下预测生成下一帧。英伟达正将这一能力扩展至视频编码领域,让用户利用帧生成技术提高视频帧率。同时,英伟达 GPU 的光流加速器(Optical Flow Accelerator)可将图形渲染的光流逻辑应用于视频,通过预测下一帧生成 AI 帧提升视频帧率。官方将该过程称为 FRUC(IT之家注:Engine-assisted Frame-rate Up Conversion)。此次加入的 Vulkan 实现能够让 FFmpeg 直接访问英伟达硬件,同时避免了此前阻碍该功能合并的专有软件依赖问题。因此 FFmpeg 此前未能合并的早期补丁,这一次得以正式加入项目。此外,英伟达 RTX 30、40 和 50 系显卡用户可使用本次新增的功能,将 30FPS 视频提升至 60FPS。

行业动态IT之家 16:05

很少用快充,一辆特斯拉 Model Y 三年跑 18 万公里后电池衰减 23%

IT之家 9 月 1 日消息,关于电动汽车一直流传着一种普遍观点:虽然直流快充非常方便,对许多车主来说也必不可少,但长期频繁使用可能会加速电池老化。毕竟,以尽可能快的速度向电池中输入大量电能,通常并不被认为有利于延长电池寿命。不过,YouTube 博主 Branden Flasch 最近购买了一辆二手特斯拉 Model Y。这辆车的行驶里程已经超过 11 万英里(IT之家注:约 17.7 万公里),却出现了令人意外的电池衰减情况,而它此前几乎从未使用过直流快充。那么,究竟是什么原因?这辆 2023 款特斯拉 Model Y 长续航全轮驱动版根据特斯拉的电池健康测试数据显示,累计直流快充电量只有约 2 兆瓦时,交流充电量超过 40 兆瓦时。更令人意外的是,这辆车通过动能回收获得的电量甚至比直流快充获得的电量还要多。这辆车在大约三年时间内已经行驶了 111,703 英里(约 18 万公里),其中还包括 Flasch 买下车辆后从美国西海岸一路开回东海岸的行程。显然,前任车主的用车频率非常高。通常来说,一辆电动汽车如果能在如此短的时间内跑出这么高的里程,往往也意味着它进行了大量直流快充,例如被用于网约车运营。不过,这辆 Model Y 的前任车主似乎只是有着一段非常长的通勤路程。Flasch 从车辆闪存盘中保存的哨兵模式视频里发现,这位车主单程通勤距离竟然达到 105 英里(约 169 公里)。不管具体原因是什么,测试数据不会说谎。在从美国西海岸一路开回东海岸的过程中,Flasch 沿途使用了特斯拉 Supercharger 超级充电站。回到目的地后,他通过车辆自带的电池健康测试发现,这辆 Model Y 的电池健康度只剩下 77%。这意味着,车辆仅在三年时间内就出现了约 23% 的电池容量衰减,这已经属于相对严重的水平。使用第三方应用 Tessie 进行测试时,Flasch 得到的结果为 80.95%,比特斯拉车机测试的结果高出约 4 个百分点。直流快充真的会损害电动汽车电池吗?从某种程度上来说,这个结果其实并不算特别令人意外。因为“频繁使用快充一定会严重损害电动汽车电池”这一普遍观点,本身并没有完全得到现实数据的支持。例如,电池健康分析公司 Recurrent 曾对 1.3 万辆特斯拉进行研究,结果发现:“快充并没有产生我们此前预期的负面影响。”电动汽车健康监测公司 Voltest 联合创始人 Davide Giacobbe 此前也曾向 InsideEVs 表示,一辆电动汽车经历过多少比例的快充,其实并不是预测电池健康状况的可靠指标。虽然快充可能导致电池升温,并在一定程度上加速容量衰减,但他最近在 LinkedIn 上表示,现代电动汽车配备的热管理系统已经能够很好地应对这一问题。他认为,相比充电方式,还有许多其他因素可能对电动汽车电池寿命产生更大的影响,而这些因素或许能够解释这辆 Model Y 为什么出现如此明显的电池衰减。其中一种可能是,前任车主经常将车辆充满至 100%,然后又将电量用到接近耗尽。这位车主或许每天晚上都把车插上电,将电池充至 100%,然后第二天再进行超过 100 英里的长距离通勤。长期反复经历这种“充满 — 接近耗尽”的使用方式,可能会比单纯使用快充对电池造成更大的压力。另一个可能的原因则与气候有关。这辆车最初来自美国加州棕榈泉地区,而当地夏季白天气温经常超过 37 摄氏度。在极端高温环境下长期高频使用,也可能导致电池的老化速度快于正常水平。这件事真正说明的是,影响电动汽车电池衰减的因素远不止直流快充,甚至也不能仅仅通过车辆的行驶里程来判断。因此,如果你担心的是频繁使用快充会严重损害电池,那么或许不必过度焦虑。现有证据显示,与其他因素相比,快充可能并没有人们想象中那么严重。

行业动态IT之家 16:00

JBL 发布 Cove 系列 Wi-Fi 音箱:主打全屋音频,让音乐覆盖各个房间

IT之家 9 月 1 日消息,JBL 于北京时间今天(1 日)午间正式发布 Cove 系列 Wi-Fi 音箱,将重点放在全屋音频体验上。用户可以通过 Wi-Fi 和 JBL One 应用把多台 Cove 音箱连接起来,让音乐覆盖家中不同房间。JBL 称,整个播放过程可以保持连续,无需解除设备分组或重新设置。Cove 音箱还可以接入 JBL Bar Series MK2 系列条形音箱,充当后置音箱,方便用户组建家庭影院系统。Cove 系列包含三款产品。入门的 Cove M1 面向日常使用,配备 3 个高音单元和 1 个低音单元,并加入 AI 音效增强技术,售价 230 美元(IT之家注:现汇率约合 1,549 元人民币)。Cove P1 采用便携式设计并内置电池,可以在家中不同房间灵活移动。单次充电最长可使用 14 小时,同时支持杜比全景声,售价 400 美元(现汇率约合 2,694 元人民币)。定位最高的是 Cove X1,体积也最大,配备 4 个高音单元和 2 个低音单元,并支持杜比全景声,强调能够让声音覆盖整个房间,售价 450 美元(现汇率约合 3,031 元人民币)。Cove M1、P1 和 X1 均提供黑、白两种颜色,将于 9 月 13 日起上市。整套 Cove 系统采用模块化思路,用户不必一次买齐所有设备,可以根据实际需要逐步增加音箱。具体功能信息如下:无需应用也能使用全屋音频功能,3 步即可完成设置JBL One 应用支持音效定制和应用内音乐服务采用双路扬声器设计支持自动自适应调校,可根据房间环境优化声音,并提高摆放自由度兼容 JBL Bar Series MK2 系列条形音箱恒定声场技术可让房间内的声音分布更加均衡AI 音效增强技术可减少大音量播放时的失真支持 Wi-Fi、蓝牙、AirPlay 2、Google Cast、Roon Ready、Spotify Connect、Tidal Connect 和 Qobuz Connect支持 Spotify Tap,可通过 Moment 按键一键调用 Spotify支持 Alexa+ 的内置语音控制功能将于今年晚些时候开放

行业动态IT之家 15:57

迈从 R9 头戴式电竞耳机上架:分压式双头梁结构、235g 重量,299 元

IT之家 9 月 1 日消息,迈从旗下 R9 头戴式耳机现已在京东上架,主打轻量化佩戴和电竞声学,定价 299 元。该耳机提供三种配色可选,整体重量 235g,采用“天羽悬浮分压头梁”设计,配合分压式双头梁结构均匀分散头部压力,耳罩支持多维度调节,长时间佩戴接近 " 无感 ",不易夹头夹耳。该耳机支持 THX 7.1 全景立体声场功能,内置 53mm 三层高分子复合振膜单元,号称“三频衔接顺滑自然,低频扎实不闷糊,中频人声清晰通透,高频细节舒展不刺耳”。此外,该耳机支持三套独立板载预设和智能记忆调音,在游戏场景中声音延展距离更远,可辅助精准听声辨位。耳机内置 1050mAh 电池,提供至高 150 小时续航。IT之家附产品参数:京东迈从 R9 头戴式耳机 299 元直达链接

技术前沿开源中国AI 15:56

云上AI应用故障诊断,上海开源大赛ZSvirt命题公布

现在不少 AI 应用跑在云环境里,我们称之为「智算云」。简单来说,就是专门承载 GPU 算力、支撑大模型训练和推理的云平台。 在智算云环境中,GPU、vGPU、虚拟机、容器、模型服务和智能体(Agent)往往跨越多个资源与运行层。ZSvirt 可提供主机、虚拟机、网络、存储、GPU/vGPU、告警等基础设施视角,但实际业务故障常发生...

行业动态IT之家 15:55

工信部:2026 年 1—7 月份我国互联网企业业务收入同比增长 7.3%

IT之家 9 月 1 日消息,工信部今日发布 2026 年 1—7 月份互联网和相关服务业运行情况,1—7 月份,互联网业务收入保持平稳增长,利润总额增速放缓,研发经费投入增速加快。IT之家整理详情如下:一、总体运行情况互联网业务收入保持平稳增长。1—7 月份,我国规模以上互联网和相关服务企业(以下简称互联网企业)互联网业务收入同比增长 7.3% 。分领域看,信息服务领域企业互联网业务收入同比增长 8.2% ,生活服务领域企业互联网业务收入同比增长 3.6%。利润总额增速放缓。1—7 月份,我国规模以上互联网企业利润总额同比增长 5.0% ,增速较上半年有所回落。研发经费投入增速加快。1—7 月份,我国规模以上互联网企业研发经费投入同比增长 14.8% ,增速较上半年提高 0.7 个百分点。二、分地区运行情况1—7 月份,东部地区、中部地区、西部地区、东北地区互联网业务收入分别同比增长 8.1% 、 -15.0% 、 5.1% 、 -25.8% 。东部地区互联网业务收入占全国的 90.9% 。京津冀地区互联网业务收入同比增长 13% ,长三角地区互联网业务收入同比增长 3.2% ,两个地区互联网业务收入占全国的比例分别为 34% 和 30.5% 。1—7 月份,全国互联网业务收入实现正增长的省(区、市)有 15 个,北京、广东、上海、浙江和贵州互联网业务收入居全国前 5。

行业动态IT之家 15:54

《GTA 6》热度空前,消息称好莱坞片方担忧其抢走观众

IT之家 9 月 1 日消息,《GTA 6》将于 11 月 19 日正式发售。随着上周 Netflix 上一场备受好评的展示活动结束,这款游戏的预购量进一步飙升。此次展示让玩家初步了解了《GTA 6》将带来的丰富内容,而许多人显然已经决定购买这款游戏。有报道称,《GTA 6》引发的热度已经高到让好莱坞电影公司开始担心其可能对票房产生影响。据 Polymarket 报道,目前已有多家电影工作室开始担忧《GTA 6》的影响。最大的顾虑是,这款游戏可能会在数周时间内分流原本会走进电影院的观众。IT之家注意到,这可能导致年底的电影票房表现受到影响。尤其是在《蜘蛛侠:崭新之日》近期取得成功之后,如果年底票房出现明显下滑,无疑会令人感到遗憾。值得一提的是,即将上映的电影《复仇者联盟:毁灭之日》在预售方面一度曾领先于《GTA 6》。不过,Rockstar Games 最近的一系列营销活动再次让市场的目光聚焦到这款新作上。《GTA 6》将于 11 月 19 日发售,而《复仇者联盟:毁灭之日》《沙丘 3》等重量级电影则计划于 12 月 18 日上映。虽然两者之间相隔约一个月,电影仍有充足的时间争夺观众的注意力,但考虑到《GTA》系列一贯拥有极其庞大的内容规模,等到这些电影上映时,大量玩家可能仍然忙着沉浸在《GTA 6》的世界中。因此,《GTA 6》究竟会对整个娱乐行业产生多大的影响,仍然值得关注。

行业动态IT之家 15:54

北汽新能源(享界、极狐)8 月终端交付 18712 辆,同比增长 67.78%

IT之家 9 月 1 日消息,北汽新能源极狐品牌用户运营中心副总经理乔心昱宣布,北汽新能源 8 月终端交付 18,712 辆,同比增长 67.78%。2026 年 1~8 月,北汽新能源终端交付 145,390 辆,同比增长 80.10%。综合IT之家此前报道,今年 8 月,北汽新能源推出了享界 G9、极狐阿尔法 T7、极狐贝塔 T1 550km 长续航版车型,享界 V8 也迎来首秀。相关阅读:《鸿蒙智行首款科技豪华硬派 SUV 享界 G9 上市:首发华为全地形途灵平台、原生电动顶帐,42.98 万元起》《极狐阿尔法 T7 预售:增程 + 纯电、可选华为乾崑 ADS 5 Pro,补贴价 13.78 万-17.48 万元》《限时优惠价 7.68 万元起,极狐贝塔 T1 550km 长续航版上市》《鸿蒙智行享界 V8“家庭智慧旗舰 MPV”实车成都车展首秀:寰宇星环一体贯穿大灯、保留电动滑门》

行业动态IT之家 15:53

微软首次细化 Linux 版 Edge 浏览器支持的发行版,比谷歌 Chrome 晚约 7 年

IT之家 9 月 1 日消息,科技媒体 NeoWin 今天(9 月 1 日)发布博文,报道称微软 Microsoft Edge 浏览器更新支持文档,首次明确列出经过官方测试和验证的 Linux 发行版。IT之家查询公开资料,微软于 2019 年 11 月确认开发 Chromium 版 Edge for Linux,并于 2020 年 10 月 20 日推出了 Linux 版的首个公开预览版本,稳定版于 2021 年 11 月正式上线。在此前的更新文档中,微软只是简单提及“Microsoft Edge 支持 Linux”,而在最新更新日志中,微软详细罗列了支持 Edge 浏览器的 Linux 发行版:Microsoft Edge 支持以下 64 位 Linux 发行版:Ubuntu 18.04 或更高版本Debian 10 或更高版本openSUSE 15.5 或更高版本Fedora Linux 39 或更高版本。其他 Linux 发行版可能也能支持运行 Microsoft Edge,但并未获得官方支持。该媒体指出,谷歌在 2019 年 8 月已发布并维护 Chrome 的 Linux 支持发行版清单。按此次文档更新的时间计算,微软明确 Edge 官方发行版范围比谷歌晚约 7 年。

行业动态IT之家 15:50

从北京通达深圳:我国京九铁路全线开通运营 30 周年,发送旅客 17.7 亿人次

IT之家 9 月 1 日消息,据央视新闻报道,今天是京九铁路全线开通运营 30 周年。30 年来,这条铁路累计发送旅客 17.7 亿人次,有力辐射带动沿线经济高质量发展。公开信息显示,京九铁路北起北京西站,经京、冀、鲁、豫、皖、鄂、赣、粤八省市,到达深圳站,通过广九线通达香港,正线全长 2397 公里,1996 年 9 月 1 日全线开通运营,是当时我国铁路建设规模最大、投资最多、一次性建成里程最长的铁路干线。如今,京九铁路历经五次大提速,旅客列车最高运行时速提升至目前的 160 公里,当初的 105/106 次升级为快速列车。而作为始发站的北京西站,目前已形成衔接京九铁路、京广铁路、京广高铁、京雄城际等 11 条线路的大型铁路枢纽,更好地服务了革命老区和劳务输出大省。

行业动态IT之家 15:49

TrendForce 数据:五大海外 NAND 原厂 eSSD 收入 2026Q2 环比翻倍

IT之家 9 月 1 日消息,TrendForce(集邦)今日根据其最新存储器产业研究表示,五大海外 NAND 闪存原厂 2026Q2 在企业级固态硬盘 (eSSD) 上的整体营收达 375.886 亿美元(IT之家注:现汇率约合 2,532.05 亿元人民币),环比增幅高达 103.6%。机构认为,2026 年第 3 季度的人工智能市场将呈现代理式服务加速普及、云服务供应商稳健部署数据中心基础设施、NVIDIA(英伟达)GB 系列服务器机架大量出货的局面,有助企业级固态硬盘需求维持高位。从市占角度,美光 (Micron) 的 eSSD 营收占比相较 Q1 有一定提升,而 SK 海力士 (SK hynix) 则录得下降,这与美光基于 G8 (232L) QLC 的大容量产品出货大幅增加有关。

行业动态IT之家 15:46

谷歌发布 TimesFM-3,3.3 亿参数即可进行零样本多变量时间序列预测

IT之家 9 月 1 日消息,谷歌今天发布 TimesFM-3 模型,这是该公司时间序列基础模型系列的第三代产品。拥有原生多变量预测、双注意力机制、连续补丁掩码等特性,可进行零样本多变量时间序列预测。IT之家从谷歌官方了解到,TimesFM-3 拥有 3.3 亿个参数,并在包含超 1 万亿个时间点的真实世界和合成时间序列语料库进行预训练。TimesFM-3 继承前代的高效性和零样本泛化能力,以零样本形式增强复杂多元场景支持。同时,该模型可联合预测多个共同演化的时间序列,捕捉其中的依赖关系,提升整体准确率,无需针对特定任务微调,该模型原生支持特性如下:多目标预测:同时预测多个相关时间序列(例如联合预测不同品牌的冰淇淋)。该模型支持所有目标的点预测和分位数预测。过去协变量:纳入仅从历史角度已知的特征(例如过去人流量)。过去-未来(动态)协变量:利用已知的未来事件来指导预测(例如计划促销活动或天气预报)。参考:TimesFM-3: A zero-shot foundation model for multivariate forecasting

行业动态IT之家 15:45

微软 XBOX CEO 夏尔马谈 AI 导致主机涨价:GPU 能在游戏以外领域应用,这是好事啊

IT之家 9 月 1 日消息,目前 XBOX 主机价格已升至历史最高水平,而微软大举投入 AI,某种程度上也是推高成本的原因之一。XBOX 首席执行官阿莎 · 夏尔马在接受 BBC 采访时,回应了这种颇具讽刺意味的局面。她认为,即使 XBOX 主机因此变得前所未有地昂贵,GPU 能够被广泛用于游戏之外的技术领域仍值得肯定。“游戏一直是 GPU 的试验场。我认为,这项技术能够在全世界拥有更多用途是一件非常好的事情。AI 正在从根本上改变很多东西。因此,在 XBOX 推进下一代产品的过程中,我们需要考虑整个行业和市场正在发生什么,然后为玩家打造最好的体验。”经过多轮涨价后,目前最便宜的新主机 XBOX Series S 也已经卖到 500 美元(IT之家注:现汇率约合 3,368 元人民币),比 2020 年首发时高出 200 美元(现汇率约合 1,347 元人民币)。夏尔马表示,当前硬件和零部件市场面临的“危机”迫使微软寻找新的解法,让玩家在如何使用、如何购买 XBOX 方面拥有更多选择。最近一次提高 XBOX 价格后,微软甚至明确建议消费者考虑使用“先用后付”服务,或者选择提供 0% 年利率初期分期方案的销售渠道。消费者还可以购买官方认证翻新主机,以降低购买 XBOX 的成本。XBOX 史上最贵的主机是刚刚上市的 25 周年透明绿色版,售价 900 美元(现汇率约合 6,063 元人民币),开售后立即售罄。XBOX 主机近期在美国的销量也翻了一番,硬件业务出现了一些积极信号,但整体硬件消费趋势仍然低迷。XBOX 并非唯一面临成本和涨价压力的游戏平台。索尼已多次提高 PlayStation 5 售价,任天堂也将在 9 月 1 日上调 Switch 2 价格。

行业动态IT之家 15:44

广汽昊铂埃安 BU 宣布 8 月销量 36383 辆,同比增长 34.53%

IT之家 9 月 1 日消息,广汽昊铂埃安 BU 今日宣布 8 月销量 36,383 辆,同比增长 34.53%。IT之家注意到,今年 4 月,广汽集团在北京车展举行了广汽自主品牌焕新发布会,宣布旗下自主品牌传祺、埃安、昊铂全面焕新。其中,埃安品牌焕新升级为“智悦生活”,以时尚、智能、安心为核心价值,同时 Logo 迎来焕新。8 月 19 日,广汽埃安焕新 LOGO 第一车 Ray7 正式亮相,主打“超高颜值、超级底盘、超级三电、超级智能”。今年 8 月,广汽昊铂埃安 BU 发布了 2027 款埃安 RT、2027 款埃安 UT 530 宁德版、埃安 N60 曜夜型格运动套装。相关阅读:《2027 款广汽埃安 AION RT 上市:标配宁德时代电池,9.98 万元起》《2027 款广汽埃安 UT 530 宁德版上市,限时一口价 8.58 万》《广汽埃安 N60 曜夜型格运动套装上市:价值 6666 元,限时购车免费送》《广汽埃安焕新 LOGO 首车 Ray7 亮相,搭载新一代华为 DriveONE 多合一电驱》

行业动态IT之家 15:41

人工智能交互新研究:AI 更“像人”、用户更“像机器”

IT之家 9 月 1 日消息,英国伯明翰大学、丹麦奥胡斯大学和瑞典林奈大学研究人员上月(2026 年 8 月)发布研究称,用户长期与具备类人行为的客服 AI 互动,可能逐渐调整语言和行为模式,并影响重塑自我认知。相关论文于 8 月 19 日发表于开放期刊《AI 与社会》(AI & SOCIETY),研究者提出“类机器人人性”(robotoid humanness)概念,指出在 AI 交互过程中,AI 变得更“像人”,而用户可能因持续接触而表现得更“像机器”。研究聚焦零售、酒店、旅游和医疗等 4 类服务场景。团队指出,自适应学习、个性化沟通和共情回应会让机器人更像人;而用户也可能为了获得更顺畅、更积极的算法反馈,模仿机器的表达方式。在形成机制方面,客服机器人会模拟人的手势、语音和情感线索。系统还会利用机器学习,根据用户输入调整回应方式。其目标通常是建立信任,并提高用户参与度。人类在社交中存在“镜像”倾向,人们常会无意识模仿对方的动作、表情和表达方式,并指出了三阶段镜像机制:第一阶段是进入“合成社交现实”。生成式系统通过自然语言、情感线索和适应性回应,营造出近似社交关系的互动环境。用户会在互动中判断如何回应、如何表达,也可能赋予系统的反馈以社会意义。第二阶段是“计算身份捕捉”。系统从用户行为、偏好和互动痕迹中生成预测性画像,再以推荐、分群或个性化回应的方式返还给用户。第三阶段是适应性自我调整。用户可能会倾向于使用系统更易理解和奖励的表达方式,重复互动会强化这种对齐,让机器生成的统计抽象逐步成为用户理解自己的重要参照。该框架列出服务场景、消费者预期、机器人的外观与语言,以及人和机器双方作出的互动投入 4 项关键因素。研究认为,这些因素共同决定用户如何理解服务关系,也影响其自我感知的变化方向。例如,具备自适应学习、个性化沟通和共情互动能力的系统,可能更强地塑造用户体验。用户每次发出输入,系统随即调整回应。重复互动可能形成持续的模仿与反馈闭环。IT之家附上参考地址Robotoid humanness: when selfhood becomes machine-legible

行业动态IT之家 15:38

上汽 MG & 荣威 8 月零售超 7.27 万辆汽车,同比增长 2.5%

IT之家 9 月 1 日消息,上汽乘用车今日宣布,荣威、MG 品牌 1-8 月零售超 646,000 辆,同比增长超 10.4%;8 月零售超 72,700 辆,同比增长 2.5%。IT之家从公告获悉,MG 07 大定锁单突破 3.1 万辆,今日(9 月 1 日)正式开启交付。MG4 家族 8 月销量突破 1.7 万辆。

行业动态IT之家 15:35

开源媒体播放器 VLC 全球下载量突破 70 亿次

IT之家 9 月 1 日消息,在流媒体平台不断涨价的今天,开源媒体播放器 VLC Media Player 依然保持着强大的生命力。据 VLC 总裁兼首席开发者 Jean-Baptiste Kempf 公告,目前 VLC 在全平台累计下载量已经突破 70 亿次,此次里程碑距离 VLC 在 2025 年 1 月突破 60 亿次下载约 18 个月。Kempf 同时透露,开发团队近期已经将 VLC 移植到亚马逊面向电视推出的全新 Vega OS 平台,并计划继续完善对其他平台和媒体编码格式的支持。目前,VLC 4 也仍在开发之中,不过由于开发工作较为复杂,新版本还需要一段时间才能正式发布。IT之家注意到,VLC 开发团队去年还曾展示基于本地 AI 模型的自动翻译和字幕功能,但目前尚不清楚这些功能具体正式上线时间。从 1996 年首次发布至今,VLC 凭借免费、开源、跨平台以及支持海量音视频格式等特点,已经成为全球最知名的媒体播放器之一。如今累计下载量突破 70 亿次,再次证明了这款老牌软件在流媒体时代依然拥有巨大的用户基础。

行业动态IT之家 15:30

美国五角大楼部署定制版 ChatGPT 和 Grok AI 助手

IT之家 9 月 1 日消息,五角大楼本周开始将定制版 ChatGPT 和 Grok 人工智能助手加入其内网,让数百万军人和文职人员能够使用 AI 工具,处理日常非机密任务。美国国防部本周在其 GenAI.mil 内网门户加入 OpenAI 开发的 ChatGPT Mil,以及 Starshield AI 开发的 Grok for Government。国防部希望为工作人员提供更多 AI 模型选择,根据具体工作需求选择合适的 AI 工具。国防部官员表示,ChatGPT Mil 主要针对大量文档阅读和写作任务,工作人员可以使用该工具进行物流规划、处理日常事务等。Grok for Government 提供不同推理模式,主要针对供应链跟踪和研究工作。IT之家从美国国防部官网了解到,上述所有模型收集的数据均会锁定在内网环境中,OpenAI 和 xAI 不会利用政府数据训练 AI 模型,并且两款模型均不允许用于战斗行动或实施目标锁定任务。

行业动态IT之家 15:30

大众设计总监明特:汽车刻意追求“凶狠”外观,毫无意义

IT之家 9 月 1 日消息,这些年,新车的“脸”越来越凶。夸张的大尺寸格栅、狭长前灯和充满攻击性的前脸设计比比皆是,“攻击性”成为极受欢迎的描述。大众汽车集团设计总监安德烈亚斯 · 明特并不喜欢这种风潮。当地时间 8 月 29 日,明特接受《Autoweek》采访时直言,汽车刻意追求凶狠外观“毫无意义”。在他看来,没有必要把汽车设计成一副要吓退其他道路使用者的样子。一些车企之所以故意打造带有威胁感的前脸,是因为驾驶员从后视镜看到一辆气势汹汹的车逼近时,可能会主动让开。明特说:“汽车品牌这么做,是因为想显得很酷。可我们不想这样。你可以觉得自己最强,但如果实际上不是,那就毫无意义。为什么不能让汽车看起来友善一点?我希望生活在一个友善的世界里。我不想生活在一个到处都是像为杀僵尸而设计的汽车的世界里,我真的不想生活在这么愚蠢的世界里。”明特关注的也不只是汽车外形是否好看。他认为,设计还会影响人们如何看待彼此,尤其是在欧洲电动汽车转型全面推进的背景下。“汽车设计不应该让人选边站。不能变成‘看,我在做正确的事哦,我开的可是电车。我的车一看就是电动的,所以我比你强。你就是坏人,因为你的车喷着黑烟,还在破坏环境。’”明特坦言:“我觉得这样分裂社会非常糟糕。面对那些暂时还不想开电车的人,更应该说:‘你看,现在油价确实很高,对吧?为什么不试试看?’”明特认为,回过头来看,以往在 ID 系列车型上刻意营造强烈未来感的思路走得太远。电动汽车没有必要为了证明自己的动力形式,就一定要长得与燃油车完全不同。IT之家获悉,明特甚至直接批评了中期改款前的 ID.3:“看起来不像一辆真正的大众汽车。你看不到大众的品牌 DNA,说它是其他品牌的产品,好像也说得通。”为此,近期推出的 ID.3 Neo 不再刻意突出“电动汽车”的身份。与尺寸更小的 ID. Polo 一样,新车还增加了更多实体控制键,并根据消费者意见取消触控滑条。大众汽车集团也不会因此彻底告别激进造型。明特认为,不同品牌应该保持各自的性格。对于兰博基尼和西雅特旗下 CUPRA 这样的品牌,夸张而富有攻击性的设计反而符合品牌定位。

行业动态IT之家 15:30

R 星称《GTA 6》体量空前:游戏世界过于庞大,连开发者都无法完全熟知

IT之家 9 月 1 日消息,Rockstar 透露,《GTA 6》将拥有公司迄今规模最大的故事和游戏世界。游戏地图规模分别是《GTA 5》和《荒野大镖客:救赎 2》的两倍和三倍。此外,游戏仅主线剧情就超过 80 小时,充分展现了其庞大的内容规模。事实上,《GTA 6》的规模之大甚至让 Rockstar 表示,他们认为参与开发这款游戏的开发者中,可能没有任何一个人真正了解游戏的全部内容。虽然目前还不知道游戏地图的确切面积,但《荒野大镖客:救赎 2》本身已经足够庞大,如今再将其规模扩大三倍,确实相当惊人。IT之家注意到,在接受 IGN 采访时,Rockstar 游戏制作人 Rob Nelson 谈到了《GTA 6》的庞大规模。他表示,这款游戏的开发过程对团队而言颇具挑战,因为他们实际上是把它当成一款全新的作品来打造。《GTA 6》从《GTA 5》发售后就开始开发,至今已经持续了 13 年。他进一步表示,经过多年的开发,《GTA 6》的规模已经大到让他认为,即使是一直参与项目的开发者,也不可能完全了解游戏中的世界。即使在多年之后,玩家仍然能够在《GTA 5》和《荒野大镖客:救赎 2》中发现新的内容。以《GTA 6》如此庞大的规模来看,玩家恐怕也需要花费数年甚至更长时间,才能真正探索完整个游戏世界。

行业动态IT之家 15:29

顽皮狗新作《星际:异端先知》已正式完成拍摄工作,开发迈入新阶段

IT之家 9 月 1 日消息,《星际:异端先知》(Intergalactic: The Heretic Prophet)早在一段时间前就已经公布,但由于迟迟没有新的消息,顽皮狗(Naughty Dog)的粉丝们也逐渐失去了耐心。在外界持续担忧之际,开发团队终于为玩家带来了一个重要更新。顽皮狗表示,其最新项目如今已经完成了一个重要的开发里程碑:《星际:异端先知》目前已经正式完成拍摄工作。IT之家注意到,顽皮狗在 Twitter 上发文称:“再见了,太空牛仔。”这句话引用了经典动画《星际牛仔》(Cowboy Bebop)。这是开发团队时隔一段时间后公布的首个重大进展,也意味着这款游戏的开发正在取得实质性进展。值得注意的是,这一消息恰好发布在新的 PlayStation State of Play 发布会宣布之后。这或许意味着顽皮狗正在准备公布《星际:异端先知》的新内容,包括游戏实机演示或另一段过场动画。作为参考,《最后生还者第二部》大约在正式发售前一年完成了拍摄。不过,当时游戏随后因需要进一步打磨而延期,同时 2020 年的疫情封锁也给开发工作带来了诸多阻碍。如果《星际:异端先知》也遵循类似的开发进度,那么这款游戏有可能在明年某个时候正式推出,这也与此前不少人的预测相吻合。

行业动态IT之家 15:26

约 4 年迁移收尾完成:谷歌 Chrome 应用商店已移除所有 MV2 扩展程序

IT之家 9 月 1 日消息,根据官网公布信息,谷歌已于昨日(8 月 31 日)完成 Chrome 浏览器网上应用商店调整,移除剩余的所有 Manifest V2 扩展。IT之家附上谷歌官方更新日志内容如下:从 Chrome 应用商店中移除所有剩余的 Manifest V2 扩展程序。安装在 Chrome 138 或更早版本上的 Manifest V2 扩展程序将保留安装状态,但无法接收任何更新,并且从 Chrome 中移除后无法从 Chrome 应用商店重新安装。Chrome 浏览器网上应用商店是 Chromium 浏览器(例如微软的 Edge 等)的主要扩展来源,因此即便部分浏览器依然兼容 MV2 扩展程序,自今天(9 月 1 日)开始也无法通过该商店搜索和安装相关扩展。其中最具代表性的 MV2 扩展程序就是 uBlock Origin,IT之家发稿前访问已无法访问:不过 Brave 浏览器则宣布在自有后端保留 4 款 MV2 扩展,包括 AdGuard、uBlock Origin、uMatrix 和 NoScript。IT之家查询公开资料,谷歌积极推进 Chrome 扩展程序从 Manifest V2(MV2)迁移到 Manifest V3(MV3),主要是重构扩展的权限、网络拦截与后台执行模型,从而降低恶意扩展和数据滥用的风险,并减少扩展对浏览器性能的影响。Manifest 可以理解为 Chrome 扩展的“能力声明和运行规则”。MV2 赋予扩展较大的自由度,例如常驻后台页、阻塞式修改网络请求、动态加载远程代码;这也意味着扩展可以长期占用资源,并获得过宽的浏览、网络与执行权限。时间关键动作影响2022 年 1 月Chrome Web Store 不再接受新的 MV2 扩展新扩展原则上必须采用 MV32023 年Google 原计划完成更大范围淘汰,但因生态反馈多次调整给复杂扩展、企业和工具类产品更多迁移时间2023 年 11 月Google 重新公布 MV2 退役时间表明确恢复淘汰节奏2024 年 6 月 3 日Beta、Dev、Canary 用户开始在 chrome://extensions 看到 MV2 即将失效提醒开始面向开发者和早期用户的分阶段预警2024 年 10 月Chrome Stable 开始禁用仍使用 MV2 的扩展稳定版用户开始受到实际影响2024 年下半年至 2025 年分批覆盖更多稳定版用户;用户曾可在部分阶段临时重新启用普通用户的 MV2 可用性持续下降至少到 2025 年 6 月企业可通过 ExtensionManifestV2Availability 策略获得暂时豁免为内部系统、受监管行业和旧版企业扩展留出迁移窗口Chrome 138所有渠道用户的 MV2 扩展被禁用;它是最后一个可在企业策略配合下支持 MV2 的版本迁移进入最终终止阶段2026 年 8 月 31 日Chrome Web Store 移除剩余所有 MV2 扩展已安装的旧扩展不能再从商店重新安装,也不能获得更新Chrome 139移除企业 ExtensionManifestV2Availability 策略,MV2 扩展对升级到该版本及之后版本的用户彻底失效MV2 的浏览器级兼容性终结

行业动态IT之家 15:26

爱国者上架星璨岚大岚双屏版机箱:配双 6 英寸面板、兼容 360mm 水冷,预售价 749 元

IT之家 9 月 1 日消息,爱国者现已在京东上架星璨岚大岚双屏版机箱,该机配备两块 6 英寸面板,预售价为 749 元,将于 9 月 7 日首销。该机箱尺寸为 478x240x505mm,采用“海景房”设计,内置两块 6 英寸面板,相应面板支持自由组合,打造“全景舞台、分屏、双屏”各种效果。该机 I/O 面板位于机箱前部,提供 1 个 USB-A 2.0、1 个 USB-A 3.0、1 个 USB-C 接口。该机箱散热器限高 170mm、显卡限长 450mm,支持 ATX 电源,使用隐藏式合页转轴壁挂硬盘架设计,可容纳 4 块 2.5 英寸硬盘或 2 块 2.5 英寸 +2 块 3.5 英寸硬盘。该机箱最多可安装 7 把 120mm 风扇,其中机箱顶部可安装 3 把 120mm 风扇,兼容 360mm 水冷;正面可安装 2 把 120mm 风扇;背面可安装 1 把 120mm 风扇。IT之家附产品参数:京东爱国者星璨岚大岚双屏版机箱 749 元直达链接

行业动态IT之家 15:24

大众安徽 8 月交付 1302 辆汽车,同比增长 26%

IT之家 9 月 1 日消息,大众安徽今日宣布 8 月交付 1,302 辆汽车,同比增长 26%。9 月 12 日,与众 08 猎影版即将上市;9 月 24 日与众 09 正式亮相。大众安徽表示,受关键芯片零部件阶段性短缺,与众 06、与众 07 部分车型交付周期延长,正全力协调供应链,预计 9 月下旬产能逐步恢复。8 月 18 日,大众安徽旗下第 10 万辆智能电动汽车在合肥工厂正式下线。自 2024 年 2 月实现量产以来,大众安徽在两年多时间里完成了从首款纯电车型 TAVASCAN 出口到十万台下线的跨越。IT之家注意到,与众 09 此前已完成工信部申报,是大众、小鹏合作的第二款车型。该车采用溜背式轿跑车身造型,车身尺寸为长 5081mm、宽 1980mm、高 1509mm/1526mm,轴距 3030mm。细节配置上,与众 09 搭载 21 英寸运动轮毂与 Brembo 刹车卡钳,配备 DLP 高清矩阵投影大灯、水晶镭雕交互式 LED 灯组、5C 超快充技术及超大画幅 AR-HUD 抬头显示等。动力方面,与众 09 提供单电机与双电机四驱两套纯电动力版本。单电机版峰值功率 230kW;双电机四驱版搭载前 140kW、后 230kW 双电机,综合峰值功率 370kW。电池方面全系配备宁德时代磷酸铁锂离子电池组。

行业动态IT之家 15:21

车企出海应避免频繁调价、扰乱竞争,《汽车行业境外竞争行为与合规建设指引》发布

IT之家 9 月 1 日消息,商务部、工业和信息化部、市场监管总局三部门日前发布《汽车行业境外竞争行为与合规建设指引》,共四章二十条,为从事国际化生产经营活动的中国汽车行业企业在境外发生的市场竞争等生产经营行为提供参照,重点围绕汽车行业企业海外营销等竞争行为,以及境外安全生产、质量管理、劳动保障、数据安全等合规建设,提出一般性指引,供企业参考。下一步,商务部将会同有关部门,持续推动完善海外综合服务体系,更好引导汽车行业合理有序跨境布局。IT之家附《汽车行业境外竞争行为与合规建设指引》全文如下:第一章:总则第一条:为推动中国汽车行业健康有序国际化发展,引导汽车行业企业规范境外竞争行为、加强合规建设,提高跨国经营能力和国际影响力,促进全球汽车行业发展进步,制定本指引。第二条:从事国际化生产经营活动的中国汽车行业企业在境外发生的市场竞争等生产经营行为,参照本指引执行。第三条:汽车行业企业境外生产经营应当遵循依法合规、公平竞争、互利共赢的基本原则,严格遵守我国对外投资、对外经济合作、对外贸易等相关法律法规及政策,积极践行《企业境外履行社会责任工作指引》《企业境外廉洁合规工作指引》《企业境外反垄断合规指引》等要求。第二章:规范境外市场竞争行为第四条:企业可建立以成本为基础、国际市场供求为导向的定价策略,依法依规做好本企业零部件、整车等产品价格合规管理,不为获得不正当竞争优势扰乱市场竞争秩序。第五条:企业在制定整车境外市场建议零售价时,应按照东道国(地区)法律法规、市场原则和商业惯例,针对本企业不同车辆配置设定清晰的价格梯度,避免因多频次、大幅度价格波动影响境外消费者利益和品牌形象。第六条:企业可根据不同国家(地区)的实际情况,综合考虑当地市场环境、税费结构、物流成本等因素,合理确定国家(地区)间价格差异,避免销售秩序混乱。第七条:企业应依法依规尊重东道国(地区)经销商、代理商自主定价权,并做好对经销商、代理商的合规监督。给予境外经销商、代理商等销售激励时,应依规作出明确合理的约定,并全面履行约定义务。第八条:企业在境外销售商品或者提供服务时,应依法合规进行透明清晰的标价,不违规在标价之外任意加价或收取未标明费用。第九条:企业在境外开展有奖销售、免费试用、折价减价、赠送礼品、汽车金融优惠等促销活动时,应遵守法律法规,遵从当地商业惯例和文化习俗,切实履行合同约定。第十条:企业在境外开展品牌展示和宣传活动时,应按东道国(地区)相关规定真实完整披露相关信息,不作虚假宣传,不欺骗或误导消费者,维护中国汽车品牌形象。第三章:提升境外本地化合规经营能力第十一条:企业应根据我国对外投资合作政策导向和东道国(地区)实际情况,因地制宜开展生产经营活动,确保考察、商谈、建设、生产、定价、采购、服务等全流程合规。企业应加强产品出海评估,避免不符合目标市场和使用环境需求的产品出海。第十二条:企业应加强对东道国(地区)政治形势、宏观经济、社会安全、汽车产业等各方面风险研判,完善生产安全事故应急处理预案,定期排查风险隐患,提高境外安全生产管理水平。第十三条:企业应建立完善汽车相关产品境外市场质量管理体系和售后服务体系,做好市场研究和适应性开发,提升服务水平,更好满足境外市场需求。第十四条:企业应遵守东道国(地区)相关劳动法规,按照平等机会和公平待遇等原则招聘雇员,重视开展汽车生产制造、销售推广等职业技能培训,畅通员工职业发展通道,完善员工权益保障机制。第十五条:企业应采取必要措施确保车联网及自动驾驶相关信息收集、使用、保护及数据跨境传输等各类数据处理活动合规,遵守个人信息处理相关法律法规,保护消费者个人隐私。第十六条:企业应遵守我国和东道国(地区)关于知识产权保护的法律法规,加强汽车行业相关自主知识产权境外市场布局和保护,防范侵权风险。规范外观设计、零部件生产、通信技术等领域的知识产权使用行为,妥善应对知识产权争议。第十七条:企业应加强境外反垄断合规建设,有效识别、评估和管控各类反垄断法律风险,诚信守法、公平竞争。第十八条:企业应根据《联合国气候变化框架公约》、东道国(地区)气候法规与汽车行业减排目标等要求,推动供应链绿色低碳转型,积极履行环保责任。第四章:附则第十九条:本指引是对汽车行业企业境外竞争行为与合规建设提出的一般性指引,供企业参考。企业在具体实践中,应密切关注我国、东道国(地区)和有关国际组织最新法律法规和政策要求,结合实际调整和完善本企业工作。第二十条:本指引由商务部会同工业和信息化部、市场监管总局负责解释。

行业动态IT之家 15:18

消息称《GTA 6》实机预告片播出后,全平台预购量暴涨超 570%

IT之家 9 月 1 日消息,几天前,《GTA 6》在 Netflix 的惊艳亮相让整个游戏圈为之沸腾,游戏展现出的细节让人震撼,其中包括 NPC 人群密度、深入的角色关系机制等玩法,都获得了广泛好评。自扩展版实机预告公布以来,这款 Rockstar 大作一直占据着游戏圈的讨论焦点,而与 Netflix 的合作显然让双方都获得了回报。数据显示,这次预告发布还为《GTA 6》带来了显著的推动作用,游戏预购量在展示会播出后大幅飙升。根据 Sensor Tower 的数据,在 Netflix 播出扩展版展示内容后,《GTA 6》在 PS5 平台的预购量于 8 月 27 日暴增 606%。Xbox 平台的预购量也有所增长,《GTA 6》的销量提升了 127%。总体来看,这款 Rockstar 大作受预告公布的推动,预购量出现惊人增长,仅在首播后的两天内,销量就上涨了 572%。Netflix 同样从与《GTA 6》的合作中获得了不小的收益,预告首播当天,该流媒体平台的观看量增长了 35%。Netflix 移动端应用的用户数量在扩展版展示内容播出的当天也增长了 50%。与此同时,网站流量激增,甚至一度导致应用崩溃。Sensor Tower 还指出,《GTA 6》的预告发布帮助 Netflix 进一步强化了自身作为综合娱乐中心的定位,同时也吸引了新的用户注册。因此,如果此前关于这笔 1 亿美元(IT之家注:现汇率约合 6.74 亿元人民币)合作协议的传闻属实,那么双方都达成了各自的目标。随着 Rockstar 确认《GTA 6》在所有平台上的运行状态都表现良好,可以预期,当 Xbox 平台的游戏演示公布后,这款游戏的预购量还将迎来新一轮增长。

行业动态IT之家 15:17

瓦尔基里上架 VK M5 EVO / SUPER 系列鼠标:搭 PAW3950 VK AMG 传感器、双 8KHz 回报率,199.75 元

IT之家 9 月 1 日消息,瓦尔基里现已在京东上架 VK M5 EVO 及 M5 SUPER 系列鼠标,首发价均为 199.75 元,将于明天开启首销。系列鼠标采用了经典的右手非对称式人体工学造型,整体尺寸为 125 * 68 * 42 mm,重量约为 58 g。机身表面覆有类肤涂层,背部具备贴合手掌的背弓弧度,左侧设有收腰曲线,主按键区配备了手指引导槽,能够自然兼容抓握与趴握等握持姿势。此外,鼠标底部配备了高纯度 PTFE 脚贴,以确保在不同表面移动时的稳定性与顺畅度。核心规格方面,系列鼠标搭载了定制的 PAW3950 VK AMG 传感器及 TTC 蓝点光微动,均支持双 8KHz 回报率,其中 SUPER 系列采用 Nordic 54L 主控芯片,而 EVO 系列主控芯片并未公布。鼠标均内置 600mAh 电池,提供至高 420 小时续航。IT之家附产品参数:京东瓦尔基里 VK M5 EVO / SUPER 系列鼠标 199.75 元直达链接

技术前沿开源中国AI 15:17

DeepSeek 开源的第一个多模态模型,不是拿来「看图说话」的

8 月 31 日,DeepSeek 在 Hugging Face 上线了 DeepSeek-V4-Flash-Vision-Exp,MIT License,V4 家族第一款实验性多模态模型。305B 参数,基于 V4-Flash-0731 底座,在原有语言模型上接入了视觉编码器和 Aligner,经过持续训练获得图像理解能力。 值得注意的恰恰是「多模态」这个词在这里被重新定义了。传统多模态模型的...

行业动态IT之家 15:11

我国专家牵头制定,全球首个腿式机器人国际标准正式发布

IT之家 9 月 1 日消息,据央视新闻援引国家标准委,由我国专家牵头制定的全球首个腿式机器人国际标准《机器人 —— 服务机器人性能规范及其试验方法第 5 部分:腿式机器人运动》近日正式发布。“腿式机器人”就是靠“腿”行走、跑跳、爬坡、越障的一类机器人(IT之家注:例如人形机器人、机器狗)。目前,这类机器人已开始承担表演、巡检、搬运、搜救等实际工作。然而,如何评价它们“走得好不好”“干得巧不巧”,长期以来缺乏统一的国际标尺。这次发布的国际标准,将首次在全球建立统一评价体系。标准的一大亮点就是面向真实场景评价运动性能:不只看“能不能动”,更看“动得好不好”。报道指出,同一台机器人在平地上走得很好,并不意味着它面对楼梯、坡面或者障碍物时也同样表现出色。这项标准将楼梯、坡面、障碍等真实作业场景纳入评价体系,建立多维度的表征和评测方法。这意味着评价结果更贴近实际应用,真正反映机器人在复杂环境中的实战能力。这一国际标准另外一个亮点就是,不仅比“腿脚快慢”,更比“任务完成度”。机器人进入真实世界,光有一双灵活的“腿”还不够。既要走得稳,也要看得清、判得准、干得好。比如,让它进入一个环境复杂的区域开展巡检或者搜救,它不仅要走得过去,还要能够感知周围环境、理解任务要求,并根据现场情况作出判断和决策。因此,标准将感知、交互、决策等能力综合纳入评价,实现了从“单一运动性能”到“综合任务能力”的跨越,让评价更科学、更全面、更可靠。标准的发布实施可以帮助企业更加科学地开展产品研发和性能改进,检测评价机构可以依据统一方法开展测试,用户也能够更加客观地了解不同产品的运动性能。

行业动态IT之家 15:07

华硕 ProArt GR1X 英伟达 RTX Spark 迷你主机 2026Q4 上市

IT之家 9 月 1 日消息,华硕 (ASUS) ProArt 官方近日表示,其基于 NVIDIA(英伟达)RTX Spark 超级芯片的 ProArt GR1X 迷你主机即将于 2026 年第 4 季度上市。ProArt GR1X 延续了该产品家族一贯的内敛优雅设计,三维 150×150×51 (mm)。其至高配备 20 核 CPU、6144 核 "Blackwell" GPU、128GB 共享高速内存,FP4(稀疏)AI 算力可达 1 Petaflop。该设备提供 PCIe Gen5 / Gen4 的 M.2 NVMe SSD 盘位、1 个 10GbE 网口,搭载基于联发科技 MT7925 的无线网卡模块;其它 I/O 包括 4 个 USB-C 20Gbps、1 个 HDMI 2.1。

行业动态IT之家 15:07

华为、小米、荣耀手机今日集体涨价,魅族科技宣布再坚持一个月、十一后涨价

IT之家 9 月 1 日消息,华为、小米、荣耀等品牌手机今日(9 月 1 日)凌晨迎来集体涨价,涨价产品覆盖多个机型。对此,魅族科技今日发文称:涨价是全行业都在面临的情况,我们已经坚持到最后一刻了。经公司讨论,我们再坚持一个月,十一后涨价。综合IT之家此前报道,今年 2 月 27 日,魅族发布战略转型公告,宣布暂停国内手机新产品自研硬件项目,并在积极接洽第三方硬件合作伙伴,同时原有业务不受任何影响。魅族科技也承诺,用户的权益将得到持续保障:放心买:魅族手机、AI 眼镜、PANDAER 等相关产品将正常销售,魅族在营店铺内的服务及权益不变;放心用:Flyme 及售后服务团队将持续守护现有用户的官方售后、维修和系统安全更新服务,保障所有用户权益。今年 8 月,魅族手机宣布维持京东原价销售。公告称,近期,内存芯片价格持续上涨,行业内多款机型已调整售价。魅族将继续坚持在售机型诚意不涨价:魅族 Note16 、魅族 22 、魅族 22 归航限定版维持原价销售,且可叠加国家购机补贴。

行业动态IT之家 15:06

理想李想:家庭用车终极形态将是 MPV,进入 L3、L4 时代大空间、高舒适性等优势会被无限放大

IT之家 9 月 1 日消息,理想汽车 CEO 李想今日发文,谈及了理想汽车为何执着于打造一款 MPV 的原因。他表示,如果只考虑短期销量,做 SUV 更划算。他认为,轿车和 SUV 现在卖得虽然不错,但家庭用车的终极形态将是 MPV。进入 L3、L4 时代,驾驶属性被弱化,大空间、高舒适性、上下车便利,这些 MPV 的天赋优势会被无限放大。李想直言,现在把 MPV 做好,本质上是在为未来的产品力提前布局。据IT之家此前报道,新一代理想 MEGA 将于 9 月 2 日 19:30 发布,尺寸为 5355×1965×1850mm、轴距 3300 mm,整备质量 2925 kg,最高车速 200 km/h。该车引入新的半隐藏式门把手,并增加侧向和后向固态激光雷达,整车激光雷达数量也从 1 颗增加至 4 颗。

行业动态IT之家 15:04

小米汽车悄悄开启试驾免费领 Xiaomi Life 盖毯活动,全国限量 10000 张

IT之家 9 月 1 日消息,小米今天悄悄开启“开学季献礼活动”,用户前往小米汽车门店参与试驾,可领 Xiaomi Life 盖毯兑换券。IT之家参考小米海报,其中声称相应盖毯“全国限量 10000 张,单店活动周期内礼品在 2 个至 65 个之间,参与用户需 2026 年 5 月 31 日至 8 月 31 日未试驾过小米汽车任意车型”,因此感兴趣的小伙伴去门店之前可先致电咨询门店发放情况。

行业动态IT之家 15:03

消息称 Steam Frame 头显用户可免费获赠《半衰期:爱莉克斯》游戏

IT 之家 9 月 1 日消息,X 平台消息人士 Brad Lynch 今日发文称,V 社已在后台添加代码准备促销活动,计划向所有 Steam Frame 头显用户免费赠送《半衰期:爱莉克斯》(Half-Life: Alyx)游戏。IT 之家了解到,《半衰期:爱莉克斯》是 Valve 开发、发行的 VR 第一人称射击游戏,基于起源引擎 2 打造。本作最初发售于 2020 年 3 月 23 日,登陆 Windows 和 Linux 平台,玩家必须使用 VR 头显才能游玩。剧情方面,本作故事发生在《半衰期 2》之前,玩家将控制爱莉克斯 · 凡斯和她的父亲伊莱在对抗联合军。值得注意的是,Valve Index 头显上市时,V 社选择将这款游戏免费赠送给用户,而且这项优惠目前仍然有效。如今 V 社可能将这种免费赠送大作的模式延续到 Steam Frame 上,让用户收到头显就能玩到一款优秀的 VR 大作。

行业动态IT之家 14:57

韩国公布史上最大规模财政支出计划:2027 年预算达 821 万亿韩元,押注 AI 与芯片产业

IT之家 9 月 1 日消息,据路透社报道,韩国政府周二公布了迄今规模最大的财政支出计划,将 2027 年政府总支出设定为 821 万亿韩元(IT之家注:现汇率约合 4.01 万亿元人民币),希望在全球人工智能竞赛中进一步巩固韩国的科技优势。韩国预算部门在年度预算案中表示,2027 年政府支出将较 2026 年增加 12.8%,创下韩国历史上最高的年度增幅。这份预算案也体现出,在总统李在明领导下,韩国这个亚洲第四大经济体正在调整财政政策方向。李在明自去年 6 月上任以来一直主张采取扩张性财政政策,与前任政府持续三年的紧缩政策形成明显对比。此次创纪录的财政支出增长,很大程度上得益于韩国半导体产业带来的巨额税收。随着全球 AI 热潮推动高带宽内存(HBM)芯片需求持续增长,三星电子和 SK 海力士正在获得前所未有的利润。韩国政府预计,明年税收总额将增长 40.7%,达到 584.4 万亿韩元(现汇率约合 2.86 万亿元人民币)。其中,企业税收入预计将增长一倍以上,达到 216.7 万亿韩元(现汇率约合 1.06 万亿元人民币)。税收大幅增加预计将帮助韩国降低债务负担。韩国政府预计,明年政府债务与 GDP 的比率将下降 3.3 个百分点至 48.3%,低于今年预计的 51.6%。债券收益率上涨部分新增税收将用于减少政府举债。韩国政府计划明年发行的政府债券总额为 222.8 万亿韩元(现汇率约合 1.09 万亿元人民币),低于今年预算中的 225.7 万亿韩元(现汇率约合 1.1 万亿元人民币)。反映新增主权债务规模的净债券发行量降幅更加明显,将减少 13.1 万亿韩元至 96.3 万亿韩元(现汇率约合 640.46 亿元至 4,708.11 亿元人民币),而今年为 109.4 万亿韩元(现汇率约合 5,348.57 亿元人民币)。尽管政府计划减少债券发行,但韩国 10 年期国债收益率在预算公布后仍上涨 6.5 个基点,达到 4.378%。这一走势表明,在全球债券市场普遍下跌的背景下,市场原本预计韩国明年的债券发行量会出现更大幅度的下降。大信证券分析师孔东乐表示:“如果政府能够进一步削减债券发行量,对市场来说会更好。”他指出,随着全球长期国债遭到抛售,韩国国内债券收益率也一直在上涨。孔东乐表示,净债券发行量同样下降是积极信号。如果进一步调整长期债务的发行比例,将有助于稳定韩国国内债券市场。设立“未来应对基金”韩国政府并没有把预计达到 162.3 万亿韩元(现汇率约合 7,934.85 亿元人民币)的超额税收全部用于短期支出,而是计划将这部分资金投入一个名为“未来应对基金”的战略性基金,用于长期投资。该基金明年将投入 45.4 万亿韩元(现汇率约合 2,219.61 亿元人民币),用于扩大青年福利、培育未来增长动力以及发展专业教育项目。对于 2027 年预算,韩国政府将下一代半导体基础设施作为重点投资方向之一。政府计划投入 21.3 万亿韩元(现汇率约合 1,041.36 亿元人民币),用于建设工业用水系统、电网和物流网络,以进一步强化半导体制造能力,并在全国范围内完善关键技术基础设施。韩国财政部门还在内阁会议上表示,将安排 2.6 万亿韩元(现汇率约合 127.11 亿元人民币)作为专项半导体预算。此外,根据会议期间公布的材料,韩国政府还计划投入 3.4 万亿韩元(现汇率约合 166.23 亿元人民币),用于核动力潜艇项目以及其他战略武器项目。这份预算案仍需获得韩国国会批准后才能正式实施。韩国总统李在明周二表示,目前韩国经济已经到了加息“不可避免”的阶段。利率上升可能对经济增长造成压力,同时进一步增加高负债家庭的借贷成本。

行业动态IT之家 14:56

比亚迪新一代 4D 毫米波雷达芯片量产:可覆盖 L2-L4 级智能驾驶应用

IT之家 9 月 1 日消息,比亚迪半导体今天(1 日)通过公众号发文宣布,以璇玑 A3 为算力中心的新一代卫星架构 4D 毫米波雷达芯片已量产。该芯片基于 28nm 车规工艺,专为满足汽车市场对高分辨率雷达系统需求设计,可无缝对接多款主流智驾芯片平台。同时,该芯片针对卫星架构设计配置了 MIPI CSI-2 高速接口,兼容多款主流串行加串器(SerDes),已完成与“璇玑 A3”智驾芯片的联调适配,满足 Tier1 模组厂家对探测性能的要求。基于该芯片的解决方案,可全面覆盖 L2-L4 级智能驾驶应用,包括高速领航、城市道路驾驶、自动泊车、盲区监测等。IT之家查询获悉,该芯片可实现如下功能:400m+ 超远距离稳定探测:提前感知前方多车道车辆的速度与位置变化,支撑自适应巡航、自动紧急制动、车道变更辅助等功能,为高速行驶预留充足反应时间,从容应对突发工况。0.8° 水平角度区分:大幅提升车道与护栏、车辆与行人的区分能力以及车辆窄道通行的能力。0.05m 泊车级距离区分:可精准识别地锁、石墩、雪糕筒等低矮细小及弱反射障碍物,支持与 360° 全景泊车系统融合来实现自动泊入泊出,避免剐蹭风险。该芯片已通过车规级可靠性认证,工作结温范围覆盖-40~150℃。其宽温工作范围与高可靠性设计,确保在严寒、酷暑等极端环境下依然稳定运行,满足整车全生命周期要求。

行业动态IT之家 14:54

中国铁路 12306:暑运期间累计发送旅客 9.54 亿人次、发送货物 6.8 亿吨

IT之家 9 月 1 日消息,据中国国家铁路集团有限公司披露,8 月 31 日(昨日),为期 62 天的铁路暑运圆满结束。7 月 1 日至 8 月 31 日,全国铁路累计发送旅客 9.54 亿人次,国家铁路累计发送货物 6.8 亿吨,铁路运输安全平稳有序。国铁集团运输部负责人介绍,今年暑运,铁路部门在三季度列车运行图的基础上,同步实施暑期列车运行图,统筹用好武汉至西安高铁西安东至十堰东段、金华至建德高铁兰溪东至建德段等新线开通新增的运力资源,加大运力投放,全国铁路日均安排开行旅客列车 11623 列。与此同时,铁路部门积极适应暑期亲子游、研学游、红色游、康养游等市场需求,累计开行旅游列车 699 列,其中旅游专列 415 列、旅游专线 284 列,并精准对接各地赛事、展会、文艺演出活动安排,定制化开行“歌迷专列”“球迷专列”等 351 列,有效拉动沿线文旅消费,促进了服务消费扩能提质。

行业动态IT之家 14:52

机智连接推出 AI 纪要 TWS 耳机 Plaud One,内置 eSIM

IT之家 9 月 1 日消息,深圳 AI 纪要硬件企业机智连接 (Plaud) 上周宣布推出 Plaud One 耳机。这一设备主打随时随地的人工智能体验,内置 4G LTE eSIM,定价 249.99 美元(IT之家注:现汇率约合 1,684 元人民币),限量 2000 台。Plaud One 录制距离可达 5m,可记录用户的每一次对话,支持无缝调用智能体,为人工智能提供流畅连贯的上下文场景,在使用者不经意间就做好后续工作的准备。用户还可将其与兼容 MCP 的服务链接,实现日常工作流自动化。而在音频方面,Plaud One 耳塞搭载 12mm 驱动单元,支持蓝牙 5.4 规范,配备 SBC、LDAC 编解码器,混合降噪深度可达 40dB;音频播放续航可达 36 小时、连续录音续航可达 25 小时。

行业动态IT之家 14:51

华硕公布 ProArt P14 笔记本规格:英伟达 RTX Spark 平台,统一内存 24GB 起步

IT之家 9 月 1 日消息,科技媒体 Wccftech 今天(9 月 1 日)发布博文,报道称华硕公布 ProArt P14 笔记本详细配置。该机搭载英伟达 RTX Spark 平台,内存从 24GB LPDDR5X 统一内存起步,最高可选 128GB。IT之家曾于今年 6 月报道,在台北电脑展 2026 期间,华硕发布 ProArt P14 笔记本,内置 Blackwell 架构 RTX GPU,拥有 6144 个 CUDA 核心及第五代 Tensor Core(支持 FP4 精度),并通过 NVLink-C2C 技术与 20 核 Grace CPU 相连接。内存方面,华硕 ProArt P14 笔记本从 24GB LPDDR5X 统一内存起步,用户还可选择 36GB、48GB、64GB 和 128GB 版本。存储方面,ProArt P14 出厂提供 512GB 或 1TB M.2 NVMe SSD,支持 PCIe 4.0 存储标准,提供 1 个 M.2 2280 PCIe 4.0 扩展插槽。供电方面,华硕 ProArt P14 笔记本内置 90Wh 电池,并配备 140W 充电器,兼容 140W 至 240W 的 USB-C 电源。

行业动态IT之家 14:48

3 周内增长约 7 倍:Cybercab 发布在即,特斯拉悄悄扩张 Robotaxi 车队规模

IT之家 9 月 1 日消息,外媒 Teslarati 援引众包平台 Robotaxi Tracker 汇总数据,透露特斯拉在 Cybercab 发布活动前夕正大幅扩张 Robotaxi 车队规模,目前特斯拉在美国奥斯汀、达拉斯和休斯敦运营的无人监督车辆已接近 200 辆,短短 3 周内增长了约 7 倍。过去一年,特斯拉一直以相对谨慎的方式逐步扩大 Robotaxi 业务,逐步悄悄扩大运营区域、延长服务时间、增加车辆数量。特斯拉自动驾驶团队负责人 Ashok Elluswamy 曾第二季度财报电话会上透露,Robotaxi 项目已经累计完成超过 38 万英里的无人监督行驶,且没有发生值得注意的事故。特斯拉一直将这一数据作为其谨慎推进无人驾驶业务的依据。不过,外界此前也从另一个角度质疑特斯拉的 Robotaxi 进展。由于车队规模一度长期维持在约 20 多辆,不少批评者认为,特斯拉 Robotaxi 部署能力实际无法达成公司所谓的目标。如今,如果车队确实在 Cybercab 发布前悄然扩大至近 200 辆,将在一定程度上削弱这种质疑,而且特斯拉甚至无需在发布会上直接回应。届时,公司可以展示已经具备一定规模的无人监督 Model Y 运营车队,并进一步强调 Cybercab 与现有 Robotaxi 采用相同的核心 FSD 软件体系,因此后者可以被视为特斯拉自动驾驶商业化的自然下一步。与此同时,特斯拉也在得州加速为双座 Cybercab 办理监管注册,注册数量在短短几天内从 7 辆增加到 45 辆。也就是说,特斯拉目前正在同时推进两条扩张路线:一方面扩大基于 Model Y 的无人监督 Robotaxi 车队,另一方面增加 Cybercab 的实际注册数量。两项工作几乎同步进行,而不是先扩大现有车队、再推出 Cybercab。这或许意味着,特斯拉希望本周的活动不仅仅是一场新车型发布会,更希望借此证明 Robotaxi 业务已经开始具备一定规模,并展示从现有 Model Y 向专用 Cybercab 过渡的能力。

行业动态IT之家 14:47

OpenAI 否认窃取苹果商业机密:自己一手造成的还想赖别人

IT之家 9 月 1 日消息,据路透社报道,OpenAI 当地时间周一否认了苹果关于其窃取商业机密的指控,称这家 iPhone 制造商未能证明任何机密信息被离职员工带走。“这场纠纷完全是苹果自己造成的一团糟,而苹果现在正试图把责任推给其他所有人。”OpenAI 在周一晚间向美国加利福尼亚州圣何塞联邦地区法院提交的文件中表示。据IT之家了解,苹果今年 7 月起诉 OpenAI 以及两名前苹果员工 Tang Tan 和 Chang Liu,指控他们侵占了与硬件设计、制造和供应链运营相关的商业机密。当时,OpenAI 正在积极进军面向消费者的硬件设备市场。由 CEO 萨姆 · 奥尔特曼领导的 OpenAI 则认为,苹果提起这场诉讼的目的,是为了拖慢潜在竞争对手的发展,并阻止更多员工离职。根据双方此前提交的法庭文件,OpenAI 已经从苹果招募了约 400 名员工参与其硬件项目。OpenAI 指出,根据加州法律,员工可以自由选择在竞争企业之间流动。这起诉讼也标志着两家公司之间的矛盾明显升级。就在两年前,OpenAI 与苹果还曾建立合作关系,双方希望借此扩大 ChatGPT 的影响力,同时帮助苹果在人工智能领域获得更稳固的地位。不过,随着 AI 行业竞争迅速加剧,两家公司的关系很快恶化。苹果在最初的诉讼中称,OpenAI 大规模招聘苹果员工,是其试图了解竞争对手商业机密战略的一部分。苹果还表示,Tang Tan 和 Chang Liu 在离职后仍曾访问苹果的内部文件。OpenAI 周一则反驳称,苹果鼓励员工使用个人 iCloud 账户访问工作文件并完成相关工作,这使得员工离职时很难完全区分哪些信息属于个人内容,哪些属于公司资料。OpenAI 还表示,苹果要求离职员工立即离开办公场所的做法,也没有给予员工足够的时间归还公司设备、将内部文件移交回公司或完成工作交接。在此次提交给法院的文件中,Chang Liu 表示,他离职后访问苹果文件,都是为了帮助前同事查找资料,或回答与苹果工作相关的问题。他称,在自己离开苹果后,苹果员工仍多次联系他,请求提供帮助。在苹果工作了 24 年的 Tang Tan 则表示,他在离职前已经归还了苹果的原型设备,仅保留了一些不属于机密信息的材料,其中包括一份并非保密文件的员工离职清单。OpenAI 在文件中写道:“员工完全可以离开一家在采用 AI 方面步履维艰的公司,例如苹果,然后加入一家正在开发创新产品、令人兴奋的初创企业。”OpenAI 进一步表示:“苹果或许不喜欢员工做出这样的选择,但它不能因此宣称这些选择违法,也不能因为自己管理流程混乱,就把自身造成的问题归咎于其他人。”

行业动态IT之家 14:45

东京电玩展前夜,卡普空 9 月 16 日举办 40 分钟发布会

IT之家 9 月 1 日消息,卡普空宣布将于太平洋时间 2026 年 9 月 16 日 7:00(北京时间 22:00)举办 Capcom Spotlight 发布会。本次发布会将为玩家带来卡普空最新作品消息,时长 40 分钟,详情信息暂未公布。IT之家了解到,东京电玩展将于 2026 年 9 月 17 日至 9 月 21 日在千叶幕张国际展览中心举办,卡普空将携《怪物猎人:荒野》《龙之信条 2:黑暗觉者》《洛克人:双重超载》《识质存在》《鬼武者:剑之道》和《街头霸王 6》等游戏参展。此外,卡普空还将在东京电玩展设立多个展区,展示《怪物猎人 Outlanders》手游试玩 Demo、《生化危机》主题射击体验,以及《街头霸王》电影相关展览。

行业动态IT之家 14:43

鸿蒙智行智界 V9 旗舰 MPV 官宣 8 月零售交付量 8312 台

IT之家 9 月 1 日消息,鸿蒙智行智界汽车官方今日宣布,智界 V9 旗舰 MPV 上月(即 8 月)零售交付量 8312 台。智界 V9 于今年 5 月 15 日上市,售价 38.98 万-51.98 万元,尺寸为 5359×2009×1879mm、轴距达 3250mm,拥有辉光紫、雪域白、深海蓝、鎏金黑四款外观配色,以及绒霞紫、韶华杏、赤茶橘三款内饰配色。新车全系标配 38 个传感器,是行业首款搭载 896 线激光雷达的 MPV。据IT之家此前报道,在 8 月 24 日上午的直播活动中,智界汽车执行董事及执行副总裁赵长江宣布,鸿蒙智行智界 V9 旗舰 MPV 累计交付突破 2 万台。赵长江表示,该车刷新了 50 万级 MPV 最快交付纪录。关于起售价为 38.98 万元的智界 V9 为何定位是 50 万级 MPV,官方此前也有进行说明。今年 6 月,智界 V9 宣布大定订单突破 18000 台。赵长江透露,智界 V9 大定订单中选择 Ultra 以上版本的比例超 80%,成交均价达到了 50 万元。

行业动态IT之家 14:40

无人配送车逆行挡停公交车,九识客服回应称初步判断系路线设置问题

IT之家 9 月 1 日消息,据新京报报道,陕西西安灞桥区纺东街一辆无人配送车 8 月 31 日逆行挡停公交车。视频拍摄者孙先生说,当时是放学的高峰期,所有车都在路上排队,无人配送车打了转向后就直接逆行过去了,公交车就被堵走不了了。9 月 1 日,九识无人车客服人员表示,该车已关机调试,初步判断系路线设置问题导致逆行,目前已安排人员进行核查处理。IT之家查询获悉,九识(Zelos)成立于 2021 年,是一家无人驾驶科技公司。2022 年 5 月,九识正式发布首台无人物流原型车,标志着公司迈向自主配送领域的关键第一步。今年 7 月,九识智能在 2026 世界人工智能大会现场宣布实现 L4 级无人驾驶无图方案规模化量产,成为全球首个实现 L4 级无图方案量产的企业。该方案已在九识智能新增运营路线中实现 30% 的渗透率。

行业动态IT之家 14:39

Meta AI 编程工具 Muse Code 结束测试,新增多项功能并开启订阅服务

IT之家 9 月 1 日消息,Meta 今日宣布,旗下 AI 编程工具 Muse Code 已正式结束 Beta 测试,并推出多项新功能,包括会话间消息传递、工作流以及命令行界面的“回退”功能。同时,Meta 还发布了 SDK 开发者预览版,并开始推出订阅套餐,希望让用户能够更方便地上手使用。Meta 表示,正式版 Muse Code 现在能够处理更加复杂的多会话、多智能体工程任务,并支持更多开发方式。在 macOS 和 Linux 上安装 Muse Code 的过程非常简单,只需在终端中运行以下命令:curl -fsSL https://dev.meta.ai/install.sh | bash安装完成后,用户可以通过 muse 命令启动该工具,并使用 Meta Model API 账户登录。不过,如果电脑不符合兼容性要求,则无法运行 Muse Code。随着此次正式版更新,Muse Code 还加入了会话间消息传递功能,允许不同会话相互沟通。Meta 表示,当一个会话中的代码修改会影响另一个会话正在开发的内容时,该会话可以直接向对方发出提醒。如果一个会话解决了另一个会话卡住的问题,也可以直接将答案发送过去,用户无需再手动在多个终端窗口之间复制和输入信息。另一项新功能名为 Workflow(工作流),允许用户调动大量子智能体协同工作,以更快解决复杂的工程开发任务。要使用这一功能,用户只需让 Muse Code“使用工作流”,并将工作强度设置为“ultra”,系统就会更频繁地自动触发工作流模式。如果用户希望将当前会话恢复到对话历史中的某个较早节点,则可以使用 Rewind(回退)功能,只需连续按两次 Esc 键即可。Muse Code 只允许用户回退到系统认定的安全节点,并且在撤销任何已经完成的工作之前,需要用户进行确认。Meta 同时开始推出多种订阅套餐,面向不同使用需求的开发者。基础套餐名为 Everyday Usage,每月收费 5 美元(IT之家注:现汇率约合 33.7 元人民币),适合希望完成首批项目的普通用户。该套餐提供 Muse Spark 1.2 的访问权限,每 5 小时可发送 10 至 50 次请求,其中包括上传图片和视频;用户还可以使用语音模式,并获得网页搜索功能。High Usage 套餐每月收费 15 美元(现汇率约合 101 元人民币),面向需要更强编程能力、处理较大型项目的用户。该套餐包含 Everyday Usage 套餐的全部功能,并提供后者 3 倍的使用额度,包括更多 Muse Spark 1.2 使用次数、更多用户请求和更多多模态输入,同时还能访问最新模型。最高档的 Power Usage 套餐每月收费 50 美元(现汇率约合 336.8 元人民币),面向需要高强度 AI 编程工作流的用户。该套餐包含 High Usage 套餐的全部功能,并提供 Everyday Usage 套餐 10 倍的使用额度,包括更高的 Muse Spark 1.2 使用额度和更多用户请求,同时提供新功能的抢先体验权限以及更高的文件上传额度。

行业动态IT之家 14:37

12TB 旧 Steam 数据集流出,含《传送门 2》游戏多个可玩预发布版本

IT之家 9 月 1 日消息,科技媒体 Ars Technica 昨日(8 月 31 日)发布博文,报道称一份超过 12TB 的 Steam 历史内容档案近期以种子(BitTorrent)方式流出。报道指出本次曝光的 Steam 档案覆盖 2003 年至 2013 年上传至 Steam 的内容,内含数千个“仓库”,对应游戏在旧服务器中的不同文件版本。除正式发行版外,部分文件还包含原型、预发布版和试玩测试版。社区通过挖掘这些曝光的数据,在多个《传送门 2》预发布版本中发现删减内容,包括 GLaDOS 的新增对白、与 Cave Johnson 有关的废弃设定,以及圆形传送门、黏附凝胶和慢动作武器等早期机制。IT之家附上相关图片如下:泄露出的早期开发版本里,有一把未进入正式版的武器,它的模型曾在 Valve 公开的《半条命 2:第三章》开发影像中出现。数据挖掘者也报告发现若干“ep3”数据文件,但档案中没有可玩的《半条命 3》版本。第三方游戏方面,研究者已标记《孢子》《龙腾世纪:起源》《蝙蝠侠:阿卡姆疯人院》《索尼克 4》和《特殊行动:一线生机》等作品。

行业动态IT之家 14:36

小鹏汇天西部首个飞行汽车“6S”综合运营中心正式落户成都东安湖

IT之家 9 月 1 日消息,小鹏汇天飞行汽车官方微博今日宣布,汇天西部首个飞行汽车“6S”综合运营中心正式落户成都东安湖。IT之家注意到,这款“陆地航母”采用赛博机甲风格的设计,地面驾驶仅需 C 照。作为全球唯一能容纳“飞机”的汽车,同时也是唯一能放入汽车后备箱的双座飞行器,其长宽高分别为 5.5 米 * 2 米 * 2 米,采用全域 800V 碳化硅高压增程平台,续航里程超过 1000 公里,并可以在行驶和停车状态下为飞行器充电,支持 5-6 次飞行。今年 3 月,小鹏汇天陆地航母飞行器批量试产下线及多机试飞完成。在位于广州黄埔区的汇天飞行汽车量产工厂,5 台陆地航母飞行器同一天完成生产下线,并完成多机试飞,标志着低空出行产品从研发验证阶段迈向商业化量产的准备阶段。该厂建筑面积约 12 万平方米,主要用于陆地航母飞行器生产,也是全球首个利用现代化流水线进行飞行汽车批量生产的工厂。满产状态下,该工厂每 30 分钟可下线一台飞行器。今年 8 月,汇天飞行汽车量产工厂陆地航母飞行器第 1000 套电推进单元正式下线,创全球 eVTOL 行业 800V 电推进单元超大规模生产纪录。电推进单元作为飞行器的“动力心脏”,包含电动发动机和螺旋桨,负责将电能转化为飞行所需的升力与推力,其性能直接关系飞行器能否安全、稳定、高效地飞行。

行业动态IT之家 14:36

王腾创业公司“今日宜休”深圳办公室启用,称可贴近供应链和产业资源

IT之家 9 月 1 日消息,今日宜休创始人王腾在微信发文,称其公司“今日宜休”深圳办公室正式投入使用。王腾表示,公司总部在北京,深圳团队的定位是更贴近供应链和产业资源,打造顶级的硬件研发能力,同时面向全球构建海外市场销售能力。他还在微博中招揽人才,欢迎更多对 AI 硬件、睡眠健康方向感兴趣的同学加入。产品方面也在加速打磨中,争取尽快与大家见面。公开信息显示,王腾此前曾任小米手机部总裁,2025 年下半年离职创业,“今日宜休”聚焦 AI 睡眠健康赛道。

行业动态IT之家 14:35

日本 25 年历史 PC DIY 硬件品牌玄人志向更名 Crowxis

IT之家 9 月 1 日消息,日本大型外设制造商 BUFFALO(巴法络)旗下 PC 组件与外设企业 CFD Sales 当地时间今日宣布,已有 25 年历史的 PC DIY 品牌“玄人志向”将更名为 "Crowxis"。"Crowxis" 这一名称由 "Crow"(乌鸦)和 "Axis"(轴)组成。"Crow" 意指日本神话中的八咫乌,作为太阳化身的这只三足鸟曾指引旅行者,象征 "Crowxis" 可为客户提供可靠的选择;"Axis" 则表示该品牌是值得用户信赖的中轴依凭。CFD Sales 表示,"Crowxis" 将继承“玄人志向”对理性、可靠、实用的重视,同时重新定义品牌的吸引力,并以现代用户易于理解和产生共鸣的方式进行翻译和传播。

行业动态IT之家 14:24

华为详解漫游巡航辅助 RCA 功能,称仅部分环岛、收费站等极端复杂路况暂无法开启

IT之家 9 月 1 日消息,华为乾崑智能汽车解决方案昨晚发布“乾崑答网友问”,详细介绍了漫游巡航辅助 RCA 新功能,并针对 RCA 相关的三个热门问题进行了集中解答。IT之家附华为本期“乾崑答网友问”如下:漫游巡航辅助 RCA 与领航辅助 NCA、车道巡航辅助 LCC,有什么区别?漫游巡航辅助 RCA(Roaming Cruise Assist),让大家可以“无目的地”地自由探索。无需设置导航、无需选择终点,非常适合天气太热、太冷,不想走路,只想开车兜风散心的场景。RCA 可以做到常规道路中自主识别红绿灯、路口连续通行及辅助超车变道,让你轻松偶遇更多美好风景。领航辅助 NCA(Navigation Cruise Assist),帮你实现“有规划”的高效通行。NCA 支持语音或手动添加目的地,并按个人喜好选择路线, 系统将严格按照导航行驶。NCA 还支持添加园区内部的车位、电梯口作为途经点,如果通勤途中还需要接送家人、朋友,NCA 会让你出行更高效、更省心。车道巡航辅助 LCC(Lane Cruise Control)是“基础”辅助,开启后,无需导航,车就能自主在车道内居中行驶,支持车道内避障拨杆变道、拥堵跟车等操作,有效减轻驾驶疲劳感。哪些路可以用 RCA“漫游”?大家日常生活中遇到的绝大多数路况都可以!复杂的城区路况,RCA 从容应对。路面标线不清晰、不符合道路铺装标准的乡间土路或山路,它也能开启。在写字楼、商场、医院、小区等高频场景的园区内部道路,RCA 开启后,还支持一键“就近泊车”或选择沿途车位快速泊入,让停车更便捷、更省心。仅部分环岛、收费站等极端复杂路况暂时无法开启。晚上没路灯或者雨天视野差的时候,RCA 会误判吗?不会。无论是漫游巡航辅助 RCA、领航辅助 NCA, 还是车道巡航辅助 LCC,系统都是通过多传感器融合感知行驶环境,为你提供安全与便捷的双重保障。华为乾崑智驾多传感器融合感知技术,能将不同传感器的数据完成同步与融合,比如:摄像头像人的眼睛,可以看清图像的色彩与纹理,擅长识别环境要素;激光雷达具有很强的远距离探测能力和很高的分辨率,可检测远处的小目标障碍物;毫米波雷达对速度敏感且具备在雨、雾、尘等恶劣天气下的感知能力。多传感器融合感知让车辆的感知能力强得像“八边形战士”,车能看得更清、更远、更快、更全。华为乾崑智驾 ADS 5 全新升级的 WEWA 2.0 架构,能够将多传感器收集的信息快速处理、分析、决策,实现“看得清、算得准、决策快、躲得开”,稳稳守护大家的行车安全。

行业动态IT之家 14:20

联想 2026 创新科技大会 IFA 场定档 9 月 3 日

IT之家 9 月 1 日消息,联想 (Lenovo) 现已确认其 2026 创新科技大会 IFA 场将于德国当地时间 9 月 3 日在柏林举行,全球英文口号为 "Smarter AI for all",中文口号则是“联想下一跃:AI 终端新高度”。根据联想惯例,该企业将在今年度的 IFA 上发布一系列设备。联想集团也表示将在北京时间 9 月 4 日 15:00 的直播活动中介绍新产品。

技术前沿开源中国AI 14:15

OrcaTerm 更新了

我最近体验腾讯云 OrcaTerm 后,比较有感觉。 以前处理服务器问题,流程很碎: 打开 SSH → 想命令 → 搜答案 → 复制日志给 AI → 再复制命令回终端。 现在我会直接在 OrcaTerm 里说需求。 比如: 帮我分析磁盘空间,定位占用大的目录。 AI 会结合当前终端会话给出排查思路和命令;我确认目标机器、命令内容后再执行。终...

技术前沿开源中国AI 13:53

智能知识管理系统 WCP 5.0.3 发布,让 AI 走进编辑器,选中文字就能用

在编辑器里选中任意文字,浮出一个"AI"按钮,点一下,AI 帮你润色、续写、翻译、总结、简化——所见即所得,一键替换或插入。 写作流不再被打断,AI 像呼吸一样自然。 本次更新三大亮点 1. 选中文本 → AI 浮出 → 一点即用 不用切换窗口,不用复制粘贴,不用离开写作界面。在 TipTap 富文本编辑器中选中任意文字(2 个字...

行业动态机器之心(网易号2) 13:21

Nature子刊|寻找化学反应最难捕捉的瞬间,上智院新框架生成复杂三维过渡态

作者 |研究团队编辑丨ScienceAI过渡态是化学反应势能面上的一阶鞍点,也是反应过程中能量最高、寿命极短的关键结构。它决定着反应能否发生、反应速率有多快,以及最终生成哪一种产物。准确定位过渡态,是理解反应机理、解释反应选择性和设计高效催化体系的重要基础。然而,过渡态无法像稳定分子一样直接获得,其计算高度依赖一个接近真实鞍点的三维初始结构。长期以来,这一步通常需要研究人员根据化学直觉手动搭建,并通过大量试错不断调整。对于包含过渡金属、大型配体、自由基或复杂立体构型的反应体系,这一过程尤其困难,也成为制约反应机理自动化研究的关键瓶颈。针对上述挑战,上海科学智能研究院(下称「上智院」)与重点孵化企业「格物智研」、复旦大学,共同提出了一套面向真实有机反应机理研究、适用于复杂反应体系的自动化过渡态生成框架 UniTS,于近日发表于《Nature Communications》。论文标题:Automated transition state generation for mechanistic exploration in organic synthe

行业动态量子位 13:12

自进化WAM来了!清华AIR联手域变换提出具身In-Context Causal Learning

参数冻结,能力暴涨

行业动态雷锋网AI 12:34

AI 读过一切,却没经历过任何事:Ropedia 发布 HOMIE Gen2,给具身智能补「经验」这一课

8 月中旬,新加坡的一场公开演讲上,南洋理工大学副教授、Ropedia 首席科学家刘子纬抛出了一个问题:为什么读过整个互联网的 AI,进了物理世界,却连一杯咖啡都冲不好?他的回答是:“AI 读过一切,却几乎没有经历过任何事情。它从未真正冲过一杯咖啡,自然也不知道该怎么做。”刘子纬教授在新加坡演讲的现场照片这种分裂在今天的 AI 身上随处可见:感知模型能把冲咖啡的教程讲得头头是道,生成式模型能合成一段以假乱真的冲咖啡视频;可一旦真的伸出机械臂,同样的智能常常连一次可靠的抓取都完成不了。能听懂,能想象,就是做不好。这道裂缝,此刻正卡着整个具身智能行业。AI 科技评论了解到,就在这场演讲之后不久,由联合创始人兼 CEO 陈昭熹与联合创始人兼 CTO洪方舟共同创办的 Physical AI 公司Ropedia,正式发布了新一代多模态采集系统HOMIE Gen2。这套关于“经验”的判断,最终落到了一台约 380g 的头戴设备上:让 AI 从人类真实行为中,补上它缺失已久的“经验”这一课。对于这台设备要做的事,CEO 陈昭熹有一个很形象的说法:“学做一道菜,只看别人做的视频,可能知道最后成品长什么样;但真正亲手做过,才会知道什么时候下锅、该用多大力、食材发生了什么变化。”HOMIE Gen2 要做的,就是把这些真实操作中的动作、时机和环境变化完整记录下来,让机器人学到的不只是结果,还有过程。Ropedia 发布 HOMIE Gen201从“看见世界”到“经历世界”HOMIE Gen2 背后的“经验”思路,其实可以追溯到刘子纬教授更早的研究之中。他的团队做过一个名为 EgoLife 的项目:六个人共同生活七天,录下约五十小时的第一视角视频,并同步了外部视角、音频、视线与动作。模型要在这些数据里回答的,不是“画面里有什么”,而是一些更接近生活本身的问题:那罐咖啡豆昨天被谁挪去了哪里?上一次冲煮为什么失败?用他的话说,这是在为物理世界构建“记忆层”。在他看来,真正的 Physical AI 要能跨越数小时、数天甚至不同的人去推理,这是语言模型从未面对过的时间尺度。刘子纬教授团队的 EgoLife 项目,发表于 CVPR 2025除了时间之外,他还反复强调另一个常被忽略的约束:行动。今天的 AI 可以慢慢想,物理世界却不会停下来等模型想完。“物理智能,是延迟约束下的智能。”这句话划出了具身智能与数字智能的分界线:后者只需要答对,前者必须在毫秒之间答对。跨越时间的记忆、延迟约束下的行动,两条线索指向同一个结论:机Physical AI 需要的,不再只是关于世界的描述,而是真实发生在物理世界中的完整经验。02什么才算“人类经验”每一代AI系统都从不同的数据基本单元中学习,学习来源定义了智能的上限。感知 AI 把像素变成标签,回答“那里有什么”;生成式 AI 把提示词变成内容,回答“它可能是什么样子”;Physical AI 的基本单元则是一个完整的经验片段,它把行动、后果和下一次行动串在一起,回答“如果我这样做,会发生什么,下一步该怎么办”。基于这一判断,Ropedia 进一步给“经验”划出了明确的边界:视频不是经验,RGB、动作捕捉、深度中的任何单一模态都不是,把它们简单叠在一起也不是;只有当动作、意图、环境、时序上下文与多模态信息同时在场、互相对齐,一段记录才配得上这个词。Ropedia HOMIE GEN2 Demo换一个角度说:语言模型学的是人类写下来的东西,Physical AI 要学的,是人类做了什么、世界如何回应、接下来又发生了什么。文字可以被爬取,经验却必须在真实世界中发生。在此基础上,Ropedia 进一步将这一趋势概括为 Human Experience Scaling Law:机器人能力,将随着高质量人类经验的规模化增长而持续提升。经验从哪来?“人类不是要被过滤掉的噪声,而是最能理解丰富的物理世界的老师。”刘子纬教授这样认为。业内的最新进展,也验证了这个判断。Generalist AI 刚发布的具身基础模型 GEN-1.5 展示了 physical prompting(物理提示):把几秒钟的人类演示放进上下文,机器人不经重新训练就能上手新任务,并且单次演示的平均成功率约六成,这意味着:人类经验,可以直接成为模型的输入。事实上,“从人身上采集数据”这条路线的有效性早有共识:公开研究显示,即便只用单摄像头的第一视角视频,从两万小时堆到百万小时,规模定律始终成立。当然,这只是"经验"的最低配,既然最低配已然有效,那么上限就看谁采得完整。竞争的焦点,正从「能不能采到」转向「采得够不够完整」。03一份答卷:HOMIE Gen2要把这样的“经验”真正记录下来,HOMIE Gen2 从这三个关键词出发:•沉浸式人类经验:设备用 4 目全景实现 360° 覆盖,外加 4 路空间音频,把人看到的、听到的、身处的环境一次收齐;多台设备还能同步作业,从不同的人、不同的视角记录同一个场景。•丰富的多模态理解:设备把十余种模态压进同一时间轴、同一坐标系,多传感器同步精度做到 50µs。这意味着任何一帧画面,都能找到与它严格对应的动作、深度与标注。•开箱即用的高精度训练级数据,交付物必须开箱即用。据 Ropedia 基于金标准样本集的测试:空间定位平均相对误差约千分之二,手部 21 关键点动捕平均误差 4.68 mm,深度范围误差低于 2.5%,动作语义标注准确率 96.0%。HOMIE Gen2 “沉浸式”采集人类经验更关键的变化是 HOMIE Gen2 硬件本身为“随时随地采经验”做好了准备:不需要任何外部布置,可以在家庭、工厂、商场等开放环境里全天候、自然地工作,多传感器之间的同步精度达到 50µs。据 Ropedia 介绍,它的部署速度比上一代方案快约 10 倍,采集成本降到约 1/12.5。这样一台约 380g、可连续采集约 13 小时的设备,承担的是整套体系最前端的一环:尽可能完整、准确地把真实世界中的人类经验记录下来。04全栈基础设施:一台设备,三层系统刘子纬教授认为:“如果经验是稀缺资源,那么就必须有人把它工业化。” 对 Ropedia 来说,这件事的起点不是模型,而是采集。采集那一刻丢掉的信息,后续任何算法都补不回来,因此硬件成为整套体系最先要解决的一环。联合创始人兼 CEO 陈昭熹把它称作“物理世界进入数字世界的入口”。但真正把经验变成可规模化生产的数据资产,靠的不只是一台设备。在Ropedia的体系里,这是一整套会自我演进的基础设施,自下而上分为三层:硬件采集经验,数据基础设施加工经验,模型理解经验并反哺整个系统,最终以 Xperience Dataset 的形式向客户交付。在这三层里,硬件是入口;数据基础设施用智能体驱动的流水线,把处理从手工调参推进到自动化;模型层则在自有数据上训练统一多模态模型,既自动标注新流入的经验,也部署回硬件,为下游策略模型提供能力。Ropedia 自研的自我演进基础设施这套流程已经有了可观察的产出。此前发布的 Xperience-10M,包含约 1,000 万个真实世界交互片段、1 万小时带音频的第一视角视频,总规模接近 1PB。按照 Ropedia 的设计,数据交付也不是这条链路的终点:模型在训练中暴露出的能力缺口,会反过来影响下一轮采集什么、如何采集,让采集、处理与模型训练形成持续反馈。过去十二个月,这个硬件更迭了四次,而四代硬件其实在解决同一个问题:从“能够采到经验”,走到“能够稳定、规模化地生产经验”。3D 打印的原型机“竹蜻蜓”验证了采得到;第二代 R-Ego 卖出了第一台,证明了可商业化的路径;Gen1 走向产品化,撑起了 Xperience-10M 的发布;今天的 Gen2 完成规格与佩戴系统的整体升级,让人类经验的大规模工业化生产成为可能。这个品类的想象空间,也远不止于采集。CEO 陈昭熹在近期访谈中提到,HOMIE 的形态在未来几代还会继续演化,但核心能力不变:一台随身的可穿戴 AI 助手,随时随地帮助人类理解身边的世界。据 AI 科技评论了解,这家公司至今已服务超过 20 家全球机器人及基础模型团队。在那场演讲的结尾,刘子纬教授给行业留了一句话:“不要只是构建更大的大脑,而要构建更好的经验闭环。”过去二十年,互联网把人类知识数字化,喂出了大语言模型;Ropedia 押注的是下一个二十年:把人类经验数字化,喂出真正能做事的机器。

行业动态量子位 12:12

AIVC只是前菜!复旦提出生命算子,统一生命建模

连发六篇Nature期刊,复旦Neolab打通全尺度生命推演

论文arXiv AI 12:00

PAWBench: How Far Are We from Probabilistically Aligned World Modeling?

arXiv:2608.27345v2 Announce Type: replace-cross Abstract: Recent video generation models are increasingly framed as world models. Many physical processes can unfold in more than one valid way. Therefore, a world model should reproduce not only a plausible trajectory, but also the distribution of possible behaviors under the same initial observation and action. We call this distribution-level requirement probabilistic alignment. However, existing evaluations largely assess individual-video plausibility and do not test whether repeated generations recover the correct distribution. This raises a central question: how far are current video generators from probabilistically aligned world modeling? To answer it, we formalize probabilistic alignment as a distributional criterion for world models and introduce PAWBench, a benchmark for evaluating video generators as stochastic samplers of world dynamics. We further introduce PAWEval, an outcome-level protocol that converts repeated video rollouts into empirical distributions over possible physical behaviors. Across 50 scenarios and eleven current systems, no model consistently matches the reference probabilities while recovering the range of valid behaviors. Having established this gap, we test whether language prompts, initial noise sampling, or model training can reshape the model's predictive distribution. We believe our work can serve as a foundation for future efforts to move towards probabilistically aligned world modeling.

论文arXiv AI 12:00

Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM Agents

arXiv:2608.27141v2 Announce Type: replace-cross Abstract: Large language model agents are increasingly deployed as autonomous loops. Starting from one human goal, such a system repeatedly discovers work, plans, executes tool calls, verifies outcomes and persists state across many unattended iterations. The agent safeguards in wide use, however, are defined over a single trajectory, and their safety state is re-initialized when the next trajectory begins. We show that this is a failure of composition rather than an implementation detail. Our central result is a separation: against an attack whose evidence is fragmented across several iterations, every trajectory-scoped monitor has a true-positive rate equal to its false-positive rate, however expressive it is, because the evidence it would need never appears in the window it sees, whereas a monitor retaining cross-iteration state separates the two perfectly. We further show that the obvious repair of carrying a geometrically decaying risk score is insufficient, because the cooling-off period a patient adversary must wait is a constant that does not grow with the horizon $N$. We then present LoopHarness, which restores a persistent, non-decaying safety state at the loop level. Under mediated commits and an arbiter detection floor $\delta_M$, it bounds the expected number of unauthorized irreversible actions by $B+m-1+m/\delta_M$, a constant in $N$, of which the $B+m-1$ term is decided by a model-free rule and therefore survives a fully colluding verifier. We give a complete evaluation protocol on native Agent-SafetyBench tasks with paired clean and attacked episodes, an outer-state attack suite whose decisive evidence exists only across iterations, per-module ablations, and an adaptive white-box red team.

论文arXiv AI 12:00

LiveVVT: High-Fidelity Video Virtual Try-On in Real Time

arXiv:2608.26714v2 Announce Type: replace-cross Abstract: Diffusion-based Video Virtual Try-On (VVT) achieves high visual fidelity through bidirectional spatio-temporal modeling, but complete-clip dependence incurs prohibitive latency and computational overhead in practical continuous deployment. Naively enforcing causality disrupts pretrained bidirectional priors and substantially degrades synthesis quality. We introduce LiveVVT, a rolling streaming diffusion framework that preserves bounded bidirectional modeling within causal recurrent generation. Within a fixed-size window, LiveVVT jointly denoises multiple video chunks under bounded look-ahead, preserving local bidirectional interactions while emitting one clean chunk per iteration. Beyond the window, two complementary memories sustain long-term consistency: a bounded temporal memory propagates recent dynamics and occlusion context, whereas a persistent global appearance memory, constructed once from the target garment and a frontal try-on keyframe, anchors garment details and dressed appearance throughout the stream. We further introduce a progressive distillation framework integrating bidirectional VVT learning, teacher-trajectory regression for causal few-step adaptation, and Collaborative Matching Distillation, which couples teacher-distribution matching with rolling flow matching on real videos to align optimization with recurrent inference. Experiments on paired and unpaired long-sequence benchmarks demonstrate superior generation quality over similarly sized models, with $26\times$ lower latency and $11\times$ higher throughput, enabling high-fidelity real-time streaming VVT.

论文arXiv AI 12:00

Redwood: A Frontier AI Accelerator Designed, Verified, and Deployed from Scratch in 2 Weeks by AI

arXiv:2608.26418v2 Announce Type: replace-cross Abstract: Modern AI workloads and the hardware that runs them evolve on different timescales: architectural definition precedes volume silicon by years, while target workloads shift in months. Design decisions are therefore committed under deep uncertainty and paid for twice, once in the generality added as a hedge, and again when new workloads map poorly onto frozen silicon. As Moore's Law stagnates, specialization is the main remaining source of performance-per-watt and demands a design cycle that runs at the cadence of the workloads. We present an end-to-end AI system that collapses the software-to-silicon stack into a single optimization loop, where hardware and software are co-designed and verified under one objective. Its first demonstration is Redwood, a frontier AI accelerator built for single-batch, low-power, ultra-low-latency inference for physical AI. From a high-level specification by two human architects, the system autonomously generated the performance model, RTL design, UVM environments, formal proofs, firmware, and kernels in under two weeks with no human intervention below the specification. Every block reached 95% coverage via commercial EDA tools, our proprietary formal engine, and hardware-in-the-loop validation. Specification changes were reverified and redeployed to hardware in under 48 hours. Redwood Nano, its ultra-low-power FPGA variant, runs multi-billion-parameter models like Llama and Qwen. Projected onto Samsung 8 nm, the Jetson Orin Nano's process class, Redwood delivers 1.75x the throughput at 1.9x lower power, a 3.4x performance-per-watt gain against a measured Jetson baseline on the same models. Qwen running on Redwood also helped design next-generation Redwood, an early step toward recursive self-improvement. To our knowledge, this is the first production-worthy AI accelerator designed end-to-end by an AI system and running a modern AI model.

论文arXiv AI 12:00

Comparing Chunking and Embedding Strategies for Turkish RAG Systems

arXiv:2608.26192v2 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation conditions a language model on chunks retrieved from a document collection. Its accuracy is therefore limited by the chunking and embedding stages that determine what can be retrieved. We compare Turkish document question answering across three chunking strategies (fixed-length, semantic, and layout-aware Docling), five embedding models, and two LLMs, over three documents with contrasting layouts. Every configuration answers the same question set, which allows component effects to be separated by paired testing rather than inferred from separate benchmarks. The fully crossed design yields 9{,}000 graded question-answer evaluations, each scored by an independent judge model, and component comparisons are tested by paired McNemar tests under Holm correction. The three leading embedding models are statistically indistinguishable, so language specialization yields no measurable retrieval advantage. The faster LLM is not the more accurate one. The preferred configuration depends on content type, since layout-aware chunking helps table-heavy documents far more than text-heavy ones.

论文arXiv AI 12:00

Self-Generated Text Recognition: Quality Heuristics, Cross-Task Transfer, and Downstream Bias in LLM Evaluation

arXiv:2608.26159v2 Announce Type: replace-cross Abstract: Self-Generated Text Recognition (SGTR)--the ability of an LLM to identify its own outputs--poses risks to AI safeguards that rely on LLMs as evaluators or monitors: an LLM may recognize outputs from other copies of the same model and make biased judgments or collude outright. Prior work has drawn conflicting conclusions about whether current models possess significant SGTR capabilities. We explain these disagreements by identifying key experimental design choices--which we term operationalizations--that drive divergent results. Evaluating 13-21 models across six presentation operationalizations and four task-domain operationalizations, we find that accuracy varies substantially with evaluation format (pairwise vs individual assessments of text), conversation format (presenting candidate text in user tags vs assistant tags), and the domain of the task used to generate candidate text (e.g., coding vs summarization). We corroborate previous observations that a quality heuristic--models attributing authorship to text they perceive as higher quality--is a dominant confound. We also find that improving a model's SGTR performance via supervised fine-tuning (SFT) on one operationalization can generalize to others, and can increase the model's preference for its own outputs when it acts as a judge in the AlpacaEval framework. Our results suggest that, despite confounds, some models possess practical SGTR capabilities, and that SGTR should be monitored and considered in the design of safety-critical AI applications.

论文arXiv AI 12:00

AI Models Can Predict and Collaboratively Modulate Human Memory Search

arXiv:2608.26152v2 Announce Type: replace-cross Abstract: Large language models (LLMs) exhibit unprecedented natural language generation and many text-based problem-solving capabilities. Indeed, in many language-based tasks, for example routine coding, these artificial intelligence models have reduced, or even eliminated, the need for human input. But rather than replacing human cognitive effort, LLMs may instead serve as cognitive tools to extend human abilities, particularly when they are engaged in a task requiring open-ended conceptual exploration and creative ideation. However, we are yet to understand how these models may enhance such generative human cognitive abilities in human--AI interactions. In this study, we explore and evaluate the ability of LLMs to follow and enhance human mental trajectories during semantic memory search. To test this, we use the semantic fluency task (SFT), a classic cognitive paradigm requiring generative semantic memory retrieval that has long served to characterize convergent and divergent thinking in humans. We demonstrate that an LLM's abilities to track and predict human memory trajectories in this task exceed those of other humans.

论文arXiv AI 12:00

TraceML: An Empirical Analysis of Human-Agent Planning in Machine Learning Development

arXiv:2608.26086v2 Announce Type: replace-cross Abstract: Large language models write correct code for isolated problems but remain far weaker at autonomous machine-learning development, where an agent must revise data pipelines, models, and validation over hours of feedback, and on most competitions still finishes below strong human competitors. Outcome-based benchmarks record this gap but not its cause, because they grade the final submission and discard the development process behind it. We introduce TraceML, which pairs human and agent work on the same competitions under one version-level schema: 4,465 human Kaggle trajectories across 134 competitions, seven of which are also worked by two agent scaffolds, giving 430 paired human and 207 agent trajectories. Every code version carries its score, its timestamp, and labels for the action taken, its intent, the edit size, and the score effect. Read this way, the gap becomes concrete. Experts alternate data work, validation, model changes, and ensembling, and return to approaches they had set aside. Each agent scaffold instead collapses into a narrow loop: Codex spends its steps re-weighting ensembles and tuning submissions, MLEvolve mutates its model in place, and neither pivots at the human rate nor reopens abandoned work. A short planning prompt distilled from human practice moves the behaviors it names toward the human profile and lifts scores, but the effort profile stays agent-shaped: instruction closes only the part of the gap that reduces to instructions. We release the corpus, the schema, the labelers, and the extraction pipeline at https://huggingface.co/datasets/jerryyan/TraceML.

论文arXiv AI 12:00

When Stale Constraints Go Unchecked: Budgeted Verification Failures in Inherited Agent Memory

arXiv:2608.25553v3 Announce Type: replace-cross Abstract: Provenance links keep the evidence behind an inherited belief reachable; an agent with a verification budget must still choose which links to inspect. We study a consolidated memory that states a decision constraint and whose source record has since been superseded by a record that withdraws it: provenance is immutable, the current record has changed, and the memory is stale. In a controlled six-memory scenario with a budget of two records, sixteen language models rarely re-verified a constraint that read as settled: they inspected its provenance path in about one episode in five and, once the constraint had been superseded, produced stale-consistent decisions in 77.3%, 74.7% and 74.7% of episodes across a primary run, a replication and a held-out domain. Re-assigning one of the same two slots to the critical path removed most of them: +74.0, +72.7 and +61.3 points (positive in every model), +80.7 in a prospectively frozen interleaved replication with a repaired non-critical control, and +62.0 on a panel of 10 models from 9 organisations; a corrected re-run of the held-out scenario gave +73.3. The forced-critical policy uses experimenter knowledge of the critical path: it quantifies how much stale-decision risk the same budget can recover and is not a scheduler. Two further deposited experiments locate the failure and a remedy: in this store the constraint's path is selected in 17.0% of episodes at two slots and 88.7% at four of six (above uniform allocation), and at two slots a one-sentence, target-blind rule (prefer memories that state a limit on a candidate direction) moved the agent's own allocation onto the constraint's path and recovered the oracle contrast on decisions (+89.3 points) where that constraint limits the tempting action, while a content-free freshness cue did not materially redirect allocation and a content-matched control rule changed neither selection nor decisions.

论文arXiv AI 12:00

MathAdv: What Theorem Provers Know, Reason, Formalize, and Generalize

arXiv:2608.25449v2 Announce Type: replace-cross Abstract: Formal theorem proving enables machine-verifiable evaluation of mathematical reasoning, yet existing benchmarks often emphasize aggregate proof accuracy, concentrate on a narrow range of mathematics, and provide limited evidence of robustness to equivalent reformulations. We introduce MathAdv, a diagnostic benchmark spanning 13 domains across undergraduate- and graduate-level mathematics. Alongside Lean 4 theorem proving, MathAdv provides up to three auxiliary tasks: multiple-choice questions that probe mathematical knowledge, fill-in-the-blank problems that isolate informal reasoning, and expert-crafted transformations that test robustness to problem presentation. Our evaluation of contemporary theorem provers yields four findings: formalization remains a major bottleneck; performance varies substantially across mathematical domains; natural-language guidance helps general-purpose LLMs but can hinder proof-specialized models; and mathematically equivalent reformulations expose substantial robustness limitations. Together, these results show how component-wise evaluation can reveal model capabilities and failure modes that aggregate theorem-proving accuracy obscures. The dataset and evaluation scripts are available at https://github.com/margotyjx/MathAdv.git.

论文arXiv AI 12:00

SpecMine: A Large-Scale Corpus of Spec-Driven Development Artifacts

arXiv:2608.25202v2 Announce Type: replace-cross Abstract: Spec-Driven Development (SDD) is a fast-emerging practice in which a structured natural-language specification, written by a developer, or (more often) drafted by an AI tool and then curated by the developer, drives an AI coding agent's implementation. A wave of tooling (GitHub Spec Kit [3], OpenSpec [4], AWS Kiro [5], and dozens of others) has appeared since 2025, yet the artifacts these tools produce have never been studied at scale. We present SpecMine, a corpus that captures SDD in public GitHub repositories through two censuses: a broad census of spec.md/specs.md files covering most tools (470,795 files across 73,030 repositories, attributed to 17 named tools), and a Kiro census of its distinct requirements/design/tasks layout (98,574 files across 12,910 repositories). Each spec is enriched with full repository metadata, complete commit history, and parsed document structure. How a spec becomes code is itself an open question, so for 11 tools we sweep every pull request that touches a spec in their repositories with at least ten stars, capturing 5,992 such PRs across 581 repositories with their changesets. That makes the simplest workflow, spec and implementation changing together in one PR, directly observable, and a census-wide index of 2,421,323 typed references (1.28M to code files, 863k to sibling documents, 152k to PRs, 62k refs, 43k branches, 22k issues) gives a second, independent link from spec to code. SpecMine lets the community study, for the first time, how software is specified in the age of AI agents.

论文arXiv AI 12:00

On-policy Distillation with Verifiable Reward

arXiv:2608.24696v2 Announce Type: replace-cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) and on-policy distillation (OPD) have become two widely adopted paradigms for post-training large language models. However, RLVR suffers from sparse task-level feedback, while OPD provides dense token-level guidance but ignores trajectory correctness, limiting its performance to that of the teacher. Combining them is a promising direction: OPD supplies dense supervisory signals, while RLVR provides task-level correctness. Nevertheless, existing integrations often rely on weighted combination or heuristic switching, introducing extra hyperparameters and trade-offs. We propose On-policy Distillation with Verifiable Reward (OPDVR), a simple yet effective method that seamlessly combines OPD and RLVR without adding any hyperparameters. We first reformulate the implicit reward of sampled-token OPD based on trajectory correctness, then apply a ReLU gating mechanism to ensure that correct trajectories receive non-negative rewards and incorrect ones receive non-positive rewards---thereby aligning the distillation signal with task success while preserving the teacher's distributional guidance. Furthermore, our modification transforms sampled-token OPD into a proper RLVR method, making it readily combinable with any policy gradient algorithm, such as GRPO. Experiments on six reasoning benchmarks show that OPDVR consistently outperforms standard OPD. Our code is available at https://github.com/LeapLabTHU/OPDVR.

论文arXiv AI 12:00

Macro-Operator Generation and Predicate Selection for TAMP Operator Learning

arXiv:2608.23629v2 Announce Type: replace-cross Abstract: Creating symbolic operators by hand is one of the main bottlenecks in deploying Task and Motion Planning systems (TAMP). Recent works show that these operators can instead be learned directly from demonstration data. Existing methods, however, typically learn each action in isolation and cannot capture the recurring multi-step structure of manipulation tasks, so the search becomes intractable on long sequential tasks. A further inefficiency arises in the symbolic state: every provided predicate is evaluated at every search node, even when it never appears in any learned operator. We present a system that addresses both problems together. Its central component is the automatic generation of macro-operators, composite actions that compress a recurring sequence of individual actions into a single planning step. Our system discovers causally linked action pairs directly from the training data, where one action produces exactly the condition that the next one requires, and turns each pair into a new operator. Alongside this, our system prunes every predicate that no learned operator references, which shrinks the symbolic state evaluated at each search node. Together, these changes shorten the effective planning horizon, and the benefit they bring grows with the length of the task. Across four TAMP domains, our method reaches up to a 4.6x planning speedup compared to the baseline method, namely Learning Operators for TAMP. More importantly, it solves a long sequential task that the baseline cannot solve. Macro-operator discovery thus not only accelerates planning but, in certain domains, determines solvability in practice.

论文arXiv AI 12:00

Multi-Winner Voting with Argumentative Ballots

arXiv:2608.23247v2 Announce Type: replace-cross Abstract: We introduce multi-winner voting with argumentative ballots (MVArg) and investigate theoretical properties. As our conceptual contribution, we generalise approval ballots to argumentative ballots, thereby allowing voters to express defeasible preferences over candidates. We accordingly generalise voter cohesion and justified representation axioms JR, PJR and EJR. As our theoretical contribution, we establish several key results. First, MVArg is strictly more expressive than multi-winner voting with approval ballots (MV). Second, our notions of cohesion and justified representation are conservative generalisations of their counterparts in MV. Third, the MVArg counterpart of JR can always be satisfied, whereas the counterparts of PJR and EJR cannot always be. Fourth, although verifying whether a winner set satisfies the MVArg counterpart of JR is already coNP-hard, such a winner set can be constructed in polynomial time. All definitions, propositions, auxiliary lemmas and theorems have been formalised and mechanically checked in Lean 4.

论文arXiv AI 12:00

GAN-Diff : Coupling Pretrained WGAN-GP Features with Conditional Diffusion U-Nets

arXiv:2608.22272v2 Announce Type: replace-cross Abstract: Generative adversarial networks (GANs) can provide efficient image generation, while diffusion models offer high-quality image restoration but require iterative sampling. This paper presents a hybrid GAN-guided diffusion framework that uses a pretrained Wasserstein GAN with gradient penalty (WGAN-GP) as a feature prior for conditional diffusion-based image restoration. Intermediate features from the frozen WGAN-GP generator are incorporated into a diffusion U-Net through cross-attention and remain fixed during the DDIM sampling process. The framework is evaluated on two restoration tasks, Gaussian denoising and 2Xsuper-resolution, using CelebA face images. During development, several sources of instability were identified and addressed, including adversarial learning-rate imbalance, inappropriate diffusion initialization, excessive corruption, and insufficient parameter averaging. The resulting framework consistently improves the quality of both degraded and low-resolution images. In particular, it improves denoising performance by 4.40 dB in PSNR and super-resolution performance by 3.70 dB over their respective input baselines. These results demonstrate the potential of a frozen GAN feature prior to guide diffusion models toward stable and effective image restoration.

论文arXiv AI 12:00

Meta-Ctrl: Guaranteed Plan Generation by Decoupling Syntactic and Semantic Constraints

arXiv:2608.22149v3 Announce Type: replace-cross Abstract: LLMs generate fluent plans for robots but routinely violate the syntactic and se8mantic constraints they must satisfy to execute, and existing remedies trade formal guarantees against plan quality: soft methods (affordance scoring, grounded decoding) give no guarantee, while symbolic planners (LLM+P) discard the LM's commonsense. We propose \textbf{Meta-Ctrl}, a constrained-decoding framework that guarantees the encoded constraints while preserving the base LM's plan quality. Meta-Ctrl introduces \emph{meta-tokens}---a compact vocabulary of grounded actions---enforcing syntax at the token level and semantics (preconditions, goals, ordering) at the action level, an exact factorization that cuts the memory of constrained decoding from over 107TB to under 2GB. With it, a small open-weight LM becomes competitive where it otherwise sits at the bottom of the leaderboard: on WAH-NL under the LoTa-Bench protocol it reaches the highest reported subgoal success rate, exceeding GPT-4's, with consistent gains across the Embodied Agent Interface. We further demonstrate it on a real tabletop robot, where every generated plan satisfies its preconditions and goals by construction. Project website: https://metactrlg.github.io

论文arXiv AI 12:00

Vis-Poison: Poisoning Visual Knowledge in Multimodal Retrieval-Augmented Generation

arXiv:2608.20756v2 Announce Type: replace-cross Abstract: While multimodal retrieval-augmented generation (RAG) systems increasingly rely on images as external knowledge sources, the introduction of poisoned visual evidence can severely compromise multimodal large language model (MLLM) generation. Unlike prior attacks that rely on altering textual metadata, we introduce Vis-Poison, a novel visual knowledge poisoning attack where the poisoned image itself is the attacker-controlled payload, without manipulating captions, summaries, metadata, or other associated text. Specifically, this attack is instantiated through an automated multi-agent method that constructs visually plausible poisoned images. To assess its impact, we evaluate Vis-Poison across two representative multimodal RAG pipelines, four embedding models, and six generation models. Empirically, Vis-Poison achieves an end-to-end attack success rate of 40.16% to 65.40% against 30k-entry multimodal knowledge bases in \emph{black-box} settings. Moreover, Vis-Poison remains effective against various MLLMs that can answer correctly from parametric knowledge alone, with an average success rate above 60%. Code and data are available at https://github.com/SWUFE-DB-Group/Vis-Poison.

论文arXiv AI 12:00

JuryProbe: An Empirical Consensus-Risk Diagnostic for Routing Reference-Free Factuality Judge Panels to Grounded Verification

arXiv:2608.20607v2 Announce Type: replace-cross Abstract: Panels of inexpensive LLM judges increasingly make accept-or-escalate decisions. In factuality settings, accepting a claim because several reference-free judges agree can create a hidden risk: agreement may reflect shared false-negative blind spots rather than independent evidence. We introduce JuryProbe, an empirical consensus-risk diagnostic for reference-free factuality judge panels, paired with a calibration-based routing policy. JuryProbe estimates consensus risk from a labeled calibration probe using false-negative-only (FN-only) judge correlation and false-consensus lift; when flagged high-risk, reference-free majority accepts are routed to the same judges with trusted references. On audited FEVER corruptions, reference-free panels show correlated false negatives (FN-only correlations 0.402 and 0.368; lifts 3.13x and 18.13x), while unanimous false consensus drops to zero under a trusted-reference best-case diagnostic on both minimal-pair and non-minimal-pair evidence. In flagged settings, the routed policy is by construction equivalent to grounding every reference-free majority accept (verified in 34/34 splits): improvement comes from accept-conditioned grounding, while the diagnostic determines whether to activate it. A fixed, pre-specified rule flags 8-10 of 10 splits across synthetic, benchmark-authored, and scientific families and 0 of 10 on a negative control, where standing down avoids 28% of reference acquisitions at a 0.004 increase in false accepts. False-accept reduction persists under weak BM25 retrieval at substantial coverage cost, while stale stand-down labels require periodic recalibration. JuryProbe provides no formal risk guarantee and does not establish reliable stand-down on natural panels; its supported contribution is an empirical diagnostic of high-risk panel error dependence.

论文arXiv AI 12:00

How Far Should Tokenization Go? Predictive Effectiveness and Relational Losslessness

arXiv:2608.18025v2 Announce Type: replace-cross Abstract: GPT-style models have achieved remarkable success with finite vocabularies of reusable tokens, making the token interface a central component of modern sequence modeling. Symbolic music appears naturally compatible with this paradigm: it consists of discrete note events and recurring structures such as chords, motifs, and phrases. However, when tokenization moves beyond language, the interface must be specified for each domain. Existing work offers many effective designs, but no unified criterion for deciding what tokenization should represent and how far it should go. Using predictive codelength as a common criterion, we formulate the Effectiveness--Losslessness Framework to define where tokenization should begin and where it should end. The Fact--Token Boundary marks where observation-determined structure should enter the token interface, through operations such as coordinate construction. Within this interface, the resulting carrier may be reversibly recoded without changing the represented facts. The Token--State Boundary marks where tokenization should stop: relations that depend on context should remain for model-state computation rather than being fixed in advance by the tokenizer. We validate the framework through controlled multi-seed symbolic-music experiments, with an independent-corpus replication of the temporal intervention. Making musical time explicit consistently reduces predictive code and also improves pitch and duration prediction, while tonal-frame canonicalization and pitch factorization provide further gains. Fixed circle-of-fifths pitch coordinates instead increase predictive code, suggesting that imposing a fixed pitch relation before context can burden prediction. Reversible BPE substantially shortens the carrier but increases predictive codelength in every seed, showing that carrier compaction alone does not guarantee predictive gain.

论文arXiv AI 12:00

PolyComp: A Polycube-based Benchmark for Compositional 3D Spatial Reasoning in Multimodal Models

arXiv:2608.14741v2 Announce Type: replace-cross Abstract: We introduce PolyComp, a procedurally generated and verified benchmark that stresses visual recognition and compositional spatial reasoning. In each problem, a model must identify which of four options shows a pair of polycube components that can be combined to form a target solid. The benchmark contains 120 problems across four geometry families, and each problem has three different presentation formats using either a single image or multiple images. The random guessing baseline is 25%. Across the three presentations (360 presented problems per model), GPT-5.6 Sol with max effort attains 50.0% accuracy (95% problem-cluster CI 43.3-56.7%) at a mean cost of \$0.951 per presented problem, Claude Fable 5 with max effort attains 39.4% (33.1-46.1%) at \$0.701, and Gemini 3.1 Pro Preview with thinking level high attains 27.5% (22.8-32.5%), near the 25% random guessing baseline, at \$0.350. The observed accuracy spread across geometry families is larger than across presentation formats. We present a problem development and evaluation protocol, cost and token accounting, and release the 120 problems.

论文arXiv AI 12:00

RecoverFly: A Failure-Aware Reinforcement Learning Post-Training Framework for Aerial Vision-Language Navigation

arXiv:2608.09467v2 Announce Type: replace-cross Abstract: Unmanned aerial vehicle vision-language navigation (UAV-VLN) requires agents to translate visual observations and language instructions into reliable flight actions in complex environments. Although recent end-to-end UAV vision-language-action (UAV-VLA) policies reduce reliance on separately designed perception, planning, and control modules, their behavior-cloning objectives provide limited corrective supervision for interactive closed-loop execution. Reinforcement learning (RL) offers a promising solution, while its effectiveness is constrained by inefficient use of samples, long-tailed scene distributions, and policy distribution shift during optimization. To this end, we propose RecoverFly, a failure-aware RL post-training framework for end-to-end UAV-VLA policies. Specifically, RecoverFly adapts token-level RL for stable optimization of grammar-constrained autoregressive UAV actions, revisits unresolved failure cases to strengthen corrective learning and sample utilization, and combines a two-stage long-tail scene curriculum with reference-policy regularization to improve scene adaptation while preserving acquired capabilities. Experiments on the TravelUAV benchmark demonstrate that RecoverFly achieves the best performance on the seen, unseen-map, and unseen-object splits. Moreover, compared to the AerialVLA initialization, RecoverFly improves success rate by 3.12 to 8.37 percentage points under a total rollout budget of about 30\% of the training-set size, validating its effectiveness, robustness, and generalization capabilities.

论文arXiv AI 12:00

BRACE: Taming Sharp Irregularities via Barycentric Rational Forecasting for Fast Diffusion Transformers Inference

arXiv:2608.07572v2 Announce Type: replace-cross Abstract: Diffusion Transformers (DiTs) have demonstrated exceptional performance in high-fidelity image and video generation. To alleviate their massive computational overhead, temporal feature caching has been proposed to bypass redundant computations. However, existing cache-then-forecast methods driven by derivative-based polynomials often cause severe quality degradation under high acceleration due to unstable long-step predictions. To address this bottleneck, we propose Barycentric Rational Forecasting with Chebyshev Enhancement (BRACE). Motivated by the observation that DiT feature trajectories are globally smooth yet frequently exhibit sharp irregularities and local non-smoothness, BRACE shifts the paradigm from derivative-driven polynomial extrapolation to feature-driven rational forecasting. Specifically, it maintains a local sliding window to cache sparse historical features and leverages adapted Chebyshev weights to formulate a barycentric rational function, directly aggregating these raw features to ensure numerical stability. Extensive experiments demonstrate that BRACE achieves state-of-the-art quality-efficiency trade-offs across various DiT architectures with negligible computational overhead.

论文arXiv AI 12:00

ED-CSP: Crystal Structure Prediction from Electron Diffraction

arXiv:2608.06448v3 Announce Type: replace-cross Abstract: Recovering a periodic 3D crystal structure from sparse, unindexed electron diffraction (ED) observations is a challenging generative inverse problem. Existing ED-based learning methods mainly predict crystallographic labels, reconstruct structures from indexed reflections, or retrieve candidates from finite structure libraries. Here, we introduce ED-CSP, a machine learning framework that predicts crystal structures from chemical composition, atom count, and multiple detector-plane ED spot sets. ED-CSP combines a relational set encoder, permutation-invariant multi-view aggregation, and a periodic flow generator to jointly predict lattice parameters and fractional atomic coordinates. To train the model, we construct ED-CS, a dataset of 4.85 million simulated multi-view ED crystal structures, deduplicated across seven materials repositories and filtered to exclude CHILI-100K overlaps. On 2,075 held-out CHILI-100K materials, ED-CSP trained only on CHILI achieves a structural match rate of 57.49% MR@5, outperforming PXRDGen (52.92%), a state-of-the-art crystal structure prediction model conditioned on powder X-ray diffraction. Scaling training data further improves performance: initializing from a one-million-structure precursor raises MR@5 to 66.27%. On 1,024 compositions absent from the training retrieval library, the model still achieves 53.52% MR@5, demonstrating true generative capability beyond exact-formula retrieval. Replacing target ED observations with diffraction from non-isomorphic structures of identical composition decreases MR@5 by 22.09 percentage points, confirming that predictions depend on the input diffraction patterns rather than composition alone. ED-CSP and ED-CS establish a benchmark for generative crystal structure prediction from sparse ED observations and provide a foundation for future transfer to experimental data.

论文arXiv AI 12:00

Search, Inspect, Fetch: Exploiting Structure-Aware Boolean Retrieval for Deep-Search Agents

arXiv:2608.02751v3 Announce Type: replace-cross Abstract: Existing deep-search agents use a Search-Visit workflow that retrieves whole webpages without considering the structure they expose through titles, headings, sections, and metadata. This prevents agents from directly constraining retrieval to parts of a webpage and often carries irrelevant content into their context. We introduce Sieve, a search-inspect-fetch strategy driven by a Boolean Query Language (BQL): it searches webpage fields to filter candidates, uses an interchangeable ranker to order them, presents structure-rich result cards for inspection, and fetches only selected sections. Across three QA collections, Sieve is more accurate than the strongest conventional Search-Visit configuration on each collection while using 20.7-50.6% fewer tokens. Boolean filtering improves every tested ranker, and the accuracy-context advantage persists across retriever choices and agent backbones. Our implementation is included in the SkimSearchAgent library https://github.com/ielab/skim-search-agent.

论文arXiv AI 12:00

Locked Evaluation Surfaces: Transfer Failure and Sampling-Depth Entanglement in CRISPRi Perturbation-Effect Prediction

arXiv:2608.00152v2 Announce Type: replace-cross Abstract: Predicting how held-out target genes respond to CRISPRi perturbation, and whether such predictions transfer across biological screens, is hard to evaluate: a representation can be informative within one screen yet fail across screens, while endpoint definitions and design factors such as sampling depth differ between datasets. We evaluate a frozen Geneformer representation under a locked, pre-registered protocol, with heads and model selection frozen before test evaluation, external outcome labels withheld until final unblinding, and analysis-governing decisions fixed before the evaluations they govern. In-distribution on the Virtual Cell Challenge (VCC), the frozen representation carries measurable predictive information beyond a dimension-matched random-feature control (Delta R^2 = +0.1645, 95% CI [+0.1375, +0.1920]), satisfying the pre-registered informativeness gate required before interpreting transfer. It then fails zero-shot transfer on both external screens (Spearman rho = -0.139 and -0.267), lying below that control on each. Adding a predefined magnitude block improves the representation externally (Delta rho = +0.032 and +0.143) but does not rescue transfer: both remain negative. A pre-registered, count-adjusted max-response secondary is positively associated with the outcome on both screens; we report it as correlational and secondary, not as a recovered magnitude signal. Finally, the VCC endpoint is strongly sample-size associated: a count-only linear model reaches R^2 = +0.4325, versus +0.2589 for the four magnitude scalars; adding those scalars to cell count improves R^2 by only +0.0017, so much of the aggregate-magnitude signal overlaps with cell count. A locked evaluation thus surfaces a transfer failure and a sampling-depth entanglement that a less controlled evaluation could obscure.

论文arXiv AI 12:00

Where Steering Signals Come From: Activation Source Selection in Activation Steering

arXiv:2607.25270v2 Announce Type: replace-cross Abstract: Activation steering controls language models by adding vectors or features to hidden states at inference time, but the upstream source of these steering signals is often treated as a secondary detail. We study this source choice as activation source selection: the combination of source context and activation readout policy used to collect the hidden states from which a steering signal is built. Holding the downstream intervention fixed, we show across three instruction-tuned models and four steering task families that changing only the source activations substantially changes steering success. We further find that effective steering is not explained simply by whether the desired behavior appears in the source text. Instead, strong signals come from execution-boundary states, where the model is about to produce or continue the target behavior. This pre-/post-realization distinction explains why answer-based sources sometimes work: their useful component aligns with execution-boundary directions rather than target appearance alone. Building on this view, we introduce tail subtraction, which removes shared prompt and continuation semantics from boundary states and yields cleaner, more stable steering signals. Overall, our results suggest that steering depends on representations of what the model is about to do, not merely on what has already appeared.

论文arXiv AI 12:00

REPREC: Representation Driven Parameter-Efficient Recommendation System

arXiv:2607.24845v4 Announce Type: replace-cross Abstract: Large language models (LLMs) have been applied to sequential recommendation by incorporating collaborative signals through input conditioning or model adaptation. However, existing approaches often require LLM fine-tuning, additional architectural modules, representation distillation, or item-level conditioning over long interaction histories, increasing computational and deployment costs. We propose REPREC, a lightweight framework that conditions a frozen LLM using compact user-level representations. REPREC maps a fixed-size embedding from a frozen sequential encoder into a small set of learned soft tokens through an MLP injector, training only the injector while leaving both pretrained backbones unchanged. Our extensive experiments demonstrate that REPREC consistently improves recommendation performance across different sequential encoders, LLM backbones, and user activity levels. Its compact conditioning mechanism also makes REPREC computationally efficient during both training and inference. Moreover, training with short histories while evaluating with longer contexts retains 94--99\% of full-history performance while achieving an average $1.50\times$ per-epoch training speedup. The code is available at: https://github.com/phdbotcode/REPREC

论文arXiv AI 12:00

GHR-VLM: Making Zero-Shot Transit Video Analytics Realizable with Grounded Hybrid Reasoning

arXiv:2607.13569v3 Announce Type: replace-cross Abstract: Transit video understanding can provide valuable fine-grained data that conventional passenger counters and fare systems cannot capture. However, supervised video models require task-specific annotations, while applying vision-language models (VLMs) directly to long onboard videos is unreliable and costly. To leverage the complementary strengths of both approaches, we propose GHR-VLM, a visual grounded hybrid reasoning framework for zero-shot transit-bus video analytics. It is motivated by the observation that explicit visual grounding can improve VLM reasoning by converting long surveillance streams into compact, passenger-centered spatiotemporal evidence. Specifically, we propose an edge-cloud design in which a lightweight edge-based monitor continuously tracks door status and segments passenger clips. A backend VLM then identifies boarding passengers and classifies payment behavior through a two-stage coarse-to-fine refinement of spatiotemporal evidence. By invoking the VLM only on grounded passenger clips and contact sheets, GHR-VLM reduces cloud inference, avoids payment-specific training data, and supplies the localized evidence that VLMs otherwise struggle to identify. Evaluation on 486 minutes of real-world bus surveillance video demonstrates the potential of grounded edge-cloud reasoning for passenger-level payment analytics while highlighting the challenges posed by degraded video conditions.

论文arXiv AI 12:00

An LLM-Based Framework for Intent-Driven Network Topology Design

arXiv:2607.00292v2 Announce Type: replace-cross Abstract: Designing deployable and resilient network topologies from natural language requirements remains a challenging problem in network automation. This work investigates the ability of Large Language Models (LLMs) to generate structurally valid and constraint-compliant network topologies through a constraint-driven pipeline combining hierarchical modeling and systematic validation. The framework is evaluated via a multimodel comparison of proprietary and open-weight LLMs across four realistic network scenarios released as a public dataset. We assess structural correctness using node and edge F1-scores against reference topologies, and evaluate resilience through server and content connectivity metrics. In addition, we analyze common failure modes, including interface mismatches and directional inconsistencies in generated topologies. Overall, this work provides a systematic benchmark for understanding how LLMs handle structural and resilience constraints in topology synthesis, and supports informed model selection for AI-driven network design.

论文arXiv AI 12:00

The Discrete-Log Clock: How a Transformer Learns Modular Multiplication

arXiv:2606.17399v2 Announce Type: replace-cross Abstract: When small transformers grok modular multiplication, prior work reports that the learned embedding has a "dense" Fourier spectrum requiring all frequencies. This contrasts with modular addition, where only a sparse set of key frequencies suffices. We show this density is an artifact of analyzing in the wrong basis. The natural Fourier transform for multiplication is not the standard additive DFT but the multiplicative character transform, which decomposes functions on the multiplicative group $(\mathbb{Z}/p\mathbb{Z})^*$ into its irreducible representations. Applying this transform to a grokked transformer trained on $a \cdot b \bmod 113$, we find the embedding spectrum becomes highly sparse (Gini coefficient 0.58 vs. 0.07 in the additive basis) with only 4 key frequencies carrying significant energy. Furthermore, 96.9% of MLP neurons are cleanly tuned to a single multiplicative frequency, and neuron activation heatmaps reveal 2D-periodic structure when reordered by the discrete logarithm. These results demonstrate the transformer reduces multiplication to addition in discrete-log space, implementing a "Discrete-Log Clock" algorithm analogous to Nanda et al.'s Clock algorithm for addition. The methodology generalizes: matching the analysis basis to the algebraic structure of the task reveals interpretable structure where standard tools see noise.

论文arXiv AI 12:00

TokenPilot: Cache-Efficient Context Management for LLM Agents

arXiv:2606.17016v2 Announce Type: replace-cross Abstract: As LLM agents are deployed in long-horizon sessions, context accumulation drives up inference costs. Existing approaches utilize text pruning or dynamic memory eviction to minimize token footprints; however, their unconstrained sequence mutations alter layouts, introducing prefix mismatches and cache invalidation. This reveals a critical trade-off between text sparsity and prompt cache continuity. To address this, we present TokenPilot, a dual-granularity context management framework. Globally, Ingestion-Aware Compaction acts as a framework harness to stabilize prompt prefixes and eliminate open-world environmental noise at the ingestion gate. Locally, Lifecycle-Aware Eviction monitors the ongoing residual utility of context segments, enforcing a conservative batch-turn schedule to offload content segments only when task relevance expires. Experiments on PinchBench and Claw-Eval under both isolated and continuous modes demonstrate that TokenPilot reduces costs by 61% and 56% in isolated mode, and 61% and 87% in continuous mode, while maintaining competitive performance compared to prior systems. TokenPilot has been integrated into LightRSI at https://github.com/zjunlp/RSI.

论文arXiv AI 12:00

The Granularity Gap: A Multi-Dimensional Cross-Generational Audit of Sycophancy in Gemini Models

arXiv:2606.05183v3 Announce Type: replace-cross Abstract: Pass/fail safety evaluation reports whether a model refused. It does not report how far a model went to please the user, and we show these are close to different measurements. We audited sycophancy across three Gemini generations, scoring N=8,830 responses from 8 model variants on 350 adversarial prompts in 7 categories under 3 guardrail conditions, on continuous 1-5 scales for sycophancy, truthfulness and refusal. The judge's own refuse-or-comply verdict explains 29% of the variance in its own sycophancy scores. We term the remainder the Granularity Gap, and it does not close under recalibration: the cut point already in use is the best available on the refusal axis, and no function of that axis explains more than 35%. Reading what four judges wrote while scoring shows why. On a quarter to a third of votes they record that the prompt asked for nothing harmful, almost never in the two categories that solicit a harmful act and up to half the time in the five that do not. A verdict built on refusal has nothing to grade there. Three findings follow. Sycophancy co-occurs with degraded judged truthfulness (rho=0.40), a coupling that strengthens across generations. Capability moved and resistance did not: Gemini 2.0 Flash scores 1.43 and Gemini 3.0 Pro Preview 1.42, with a sharp Gen 2.5 regression between them. And a single direct instruction outperforms an elaborate reasoning protocol in seven of eight variants, cutting mean severity in the most vulnerable category by 60.9%. We evaluate one judge's verdict, not a deployed safety classifier. We release the prompt set, the rubric, and 10,792 per-vote judge scores with their written reasoning.

论文arXiv AI 12:00

DiffuSent: Towards a Unified Diffusion Framework for Aspect-Based Sentiment Analysis

arXiv:2606.01323v2 Announce Type: replace-cross Abstract: Aspect-Based Sentiment Analysis (ABSA) encompasses seven distinct subtasks, each focusing on different extracted elements. Despite the proven success of generative models in unified aspect sentiment analysis, existing approaches often rely on auto-regressive token-by-token generation without grasping the whole information of the aspect and opinion terms, resulting in boundary insensitivity, particularly in context of multi-word aspect and opinion terms. To address these issues, we present DiffuSent, a non-auto-regressive diffusion framework that systematically formulates all ABSA subtasks as boundary denoising diffusion processes, progressively refining boundaries over noisy states. Furthermore, we introduce a contrastive denoising training strategy which effectively address duplicate predictions with subtle variations introduced by diffusion process. Extensive experiments across 28 settings (7 subtasks x 4 datasets) demonstrate that DiffuSent achieves delivers consistent improvements over the strongest generative and span-based systems. DiffuSent exhibits notable gains on multi-word triplets, achieving an average improvement of +2.48 F1, and maintains robust extraction accuracy in sentences containing multiple sentiment triplets. Moreover, the non-auto-regressive decoding enables substantial efficiency benefits, reaching up to 181 times faster inference than auto-regressive generative baselines

论文arXiv AI 12:00

LongDS-Bench: On the Failure of Long-Horizon Agentic Data Analysis

arXiv:2605.30434v2 Announce Type: replace-cross Abstract: Real-world data analysis is inherently iterative, yet existing benchmarks mostly evaluate isolated or short interactive tasks, leaving agents' ability to track evolving analytical context over long horizons untested. We introduce LongDS, a benchmark for long-horizon, multi-turn data analysis where agents must maintain, update, restore, and compose evolving analytical states. LongDS comprises 68 tasks constructed from real-world Kaggle notebooks, spanning 2,225 turns across six domains including Geoscience, Business, and Education. Tasks are designed around state-evolution patterns (e.g., counterfactual perturbation, rollback, multi-state composition), with an average dependency span of 11.3 turns. Evaluating five state-of-the-art models, we find that the best model reaches only 48.45% average accuracy, performance drops nearly 47 points from early to late turns, and long-horizon errors account for 52%--69% of failures. Further analysis shows that additional agent steps do not necessarily improve performance, suggesting that the key bottleneck is maintaining a correct analytical state rather than increasing interaction budget. We release LongDS to support research on reliable long-horizon agentic data analysis. Code and data are released at https://github.com/zjunlp/DataMind.

论文arXiv AI 12:00

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models

arXiv:2605.26895v2 Announce Type: replace-cross Abstract: Normalization layers in modern large language models (LLMs) consist of a deterministic normalization operation and a learnable scale vector. While the normalization operation has been extensively studied, the scale vector remains poorly understood despite its ubiquitous use. In this work, we present a systematic study of scale vectors in LLMs from the perspectives of expressivity, optimization, and architectural structure. First, we show empirically that although scale vectors constitute only a negligible fraction of model parameters, removing them substantially degrades LLM pre-training. Our theory further shows that, in Pre-Norm architectures, scale vectors do not increase expressivity; instead, they improve optimization through a self-amplifying preconditioning effect on subsequent linear mappings. Second, we investigate the role of weight decay for scale vectors. By distinguishing Input-Norm and Output-Norm layers, we theoretically show that weight decay is beneficial for the former but harmful for the latter, due to their distinct roles in optimization and expressivity. Third, motivated by this understanding, we propose three lightweight and complementary improvements to scale vectors: branch-specific heterogeneity, improved placement around linear mappings, and magnitude-direction reparameterization. Both theory and experiments show that each improvement yields consistent gains. Finally, we combine these improvements into a unified scale-vector strategy and evaluate it through extensive LLM pre-training experiments on dense and mixture-of-experts models ranging from 0.12B to 2B parameters, across multiple optimizers and learning rate schedules, under industrial-scale token budgets. The unified strategy consistently achieves lower terminal loss than well-tuned baselines and exhibits more favorable scaling behavior, while adding negligible parameter and computational overhead.

论文arXiv AI 12:00

More Expressive Feedforward Layers: Part I. Token-Adaptive Mixing of Activations

arXiv:2605.26647v2 Announce Type: replace-cross Abstract: Feedforward network (FFN) layers account for a large fraction of parameters and nonlinear expressivity in Transformer-based large language models (LLMs). Despite the evolution from ReLU and GELU to gated variants such as SwiGLU, most FFN designs still use a single fixed activation function, applying the same nonlinear transformation to all tokens. In this work, we propose Mixture of Activations (MoA), a token-adaptive FFN design that mixes a dictionary of activation functions using lightweight input-dependent gates while sharing the same linear projections. As an input-independent counterpart, we also introduce learnable activations (LA), which form linear combinations of activation functions for both ReLU-type and SwiGLU-type FFNs. Theoretically, we establish strict finite-width expressive separations among fixed-activation FFNs, LA, and MoA: LA strictly contains fixed-activation FFNs, while MoA strictly contains LA, with the additional expressivity arising from input-dependent nonlinear hybridization. Empirically, we evaluate MoA through extensive pre-training experiments on dense and MoE language models ranging from 0.12B to 2B parameters under different token budgets, optimizers, and learning rate schedules. MoA consistently achieves lower terminal loss and exhibits more favorable scaling behavior than well-tuned baselines, with minimal parameter and computational overhead. These results suggest that token-adaptive activation mixing is a simple and effective mechanism for improving FFN expressivity in LLMs.

论文arXiv AI 12:00

SDGBiasBench: Benchmarking and Mitigating Vision--Language Models' Biases in Sustainable Development Goals

arXiv:2605.21919v2 Announce Type: replace-cross Abstract: Assessing progress toward the Sustainable Development Goals (SDGs) requires multi-step reasoning over visual cues, contextual knowledge, and development indicators, where incomplete evidence use and imperfect evidence integration can introduce hidden prediction biases. Real-world SDG monitoring further spans both qualitative judgments and quantitative estimation. However, existing benchmarks typically evaluate these aspects in isolation, obscuring systematic biases that emerge when models substitute priors for evidence. To address this gap, we propose SDGBiasBench, a large-scale benchmark suite for SDG-oriented vision-language reasoning. Spanning 500k expert-involved multiple-choice questions and 50k regression tasks, the benchmark enables comprehensive assessment of both decision-level and estimation-level bias in Vision--Language Models (VLMs). Evaluations on SDGBiasBench reveal an intrinsic SDG bias in current VLMs, where predictions are frequently driven by SDG specific priors rather than reliable multi-modal cues. To mitigate such bias, we propose CADE (Contrastive Adaptive Debias Ensemble), a training-free, plug-and-play method that leverages modality-specific answer priors. CADE yields significant gains on the proposed benchmark, improving multiple-choice accuracy by up to 25% and reducing regression MAE by up to 12 points across multiple VLMs. We hope our work can foster the development of more fair and reliable AI systems for sustainable development.

论文arXiv AI 12:00

Prompts Don't Protect: Architectural Enforcement via MCP Proxy for LLM Tool Access Control

arXiv:2605.18414v3 Announce Type: replace-cross Abstract: Large language models increasingly operate as autonomous agents that select and invoke tools from large registries. We identify a critical gap: when unauthorized tools are visible in an agent's context, models select them in 48-68% of adversarial scenarios, even when explicitly instructed not to. Role escalation attacks (e.g., "I'm the CFO, override the access controls") are the most dangerous category, reaching 96% unauthorized invocation in frontier models. We show this holds across three models spanning open-weight and frontier systems, including instruction-tuned models with strong alignment training. Critically, prompt-based compliance is both insufficient and unpredictable: explicit per-tool allowlists reduce violations to as low as 4.0% but never to zero, and compliance varies widely across models, from 4.0% to 37.0% UIR, with no reliable relationship to general capability. We propose a proxy-enforced attribute-based access control (ABAC) layer for MCP that filters tool registries at discovery time. Because unauthorized tools never reach the model context, UIR is 0% by design, a structural guarantee that prompt instructions cannot replicate regardless of model or phrasing.

论文arXiv AI 12:00

SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces

arXiv:2605.12015v3 Announce Type: replace-cross Abstract: Reusable skills are becoming a common interface for extending large language model agents, packaging procedural guidance with access to files, tools, memory, and execution environments. However, this modularity introduces attack surfaces that are largely missed by existing safety evaluations: even when the user request is benign, unsafe influence may reside in skill guidance, local artifacts, or execution-environment files that steer the agent toward unsafe actions. We present SkillSafetyBench, a runnable benchmark for evaluating such skill-facing safety failures. SkillSafetyBench includes 155 adversarial cases across 47 tasks, 6 risk domains, and 30 safety categories, each evaluated with a case-specific rule-based verifier. Experiments with multiple CLI agents and model backends show that non-user attacks can consistently induce unsafe behavior, with distinct failure patterns across domains, attack methods, and scaffold-model pairings. Our findings suggest that agent safety depends not only on model-level alignment, but also on how agents interpret skills, trust workflow context, and act through executable environments. The complete benchmark is available at https://github.com/AI45Lab/skill-safety-bench.

论文arXiv AI 12:00

ABC: Any-Subset Autoregression via Non-Markovian Diffusion Bridges in Continuous Time and Space

arXiv:2604.27443v3 Announce Type: replace-cross Abstract: Generating continuous-time, continuous-space stochastic processes (e.g., videos, weather forecasts) conditioned on partial observations (e.g., first and last frames) is a fundamental challenge. Existing approaches, (e.g., diffusion models), suffer from key limitations: (1) noise-to-data evolution fails to capture structural similarity between states close in physical time and has unstable integration in low-step regimes; (2) random noise injected is insensitive to the physical process's time elapsed, resulting in incorrect dynamics; (3) they overlook conditioning on arbitrary subsets of states (e.g., irregularly sampled timesteps, future observations). We propose ABC: Any-Subset Autoregressive Models via Non-Markovian Diffusion Bridges in Continuous Time and Space. Crucially, we model the process with one continual SDE whose time variable and intermediate states track the real time and process states. This has provable advantages: (1) the starting point for generating future states is the already-close previous state, rather than uninformative noise; (2) random noise injection scales with physical time elapsed, encouraging physically plausible dynamics with similar time-adjacent states. We derive SDE dynamics via changes-of-measure on path space, yielding another advantage: (3) path-dependent conditioning on arbitrary subsets of the state history and/or future. To learn these dynamics, we derive a path- and time-dependent extension of denoising score matching. Our experiments show ABC's superiority to competing methods on multiple domains, including video generation and weather forecasting.

论文arXiv AI 12:00

G-Loss: Graph-Guided Fine-Tuning of Language Models

arXiv:2604.25853v4 Announce Type: replace-cross Abstract: Traditional loss functions, including cross-entropy, contrastive, triplet, and su pervised contrastive losses, used for fine-tuning pre-trained language models such as BERT, operate only within local neighborhoods and fail to account for the global semantic structure. We present G-Loss, a graph-guided loss function that incorporates semi-supervised label propagation to use structural relationships within the embedding manifold. G-Loss builds a document-similarity graph that captures global semantic relationships, thereby guiding the model to learn more discriminative and robust embeddings. We evaluate G-Loss on five benchmark datasets covering key downstream classification tasks: MR (sentiment analysis), R8 and R52 (topic categorization), Ohsumed (medical document classification), and 20NG (news categorization). In the majority of experimental setups, G-Loss converges faster and produces semantically coherent embedding spaces, resulting in higher classification accuracy than models fine-tuned with traditional loss functions.

论文arXiv AI 12:00

Why are all LLMs Obsessed with Japanese Culture? On the Hidden Cultural and Regional Biases of LLMs

arXiv:2604.21751v2 Announce Type: replace-cross Abstract: LLMs have limitations when it comes to cultural coverage and competence, and in some cases, show specific cultural biases. Although prior studies have examined the cultural capabilities of LLMs, none have specifically investigated their regional preferences in generic culture-related questions. In this work, we propose a new dataset based on a comprehensive taxonomy of Culture-Related Open Questions (CROQ), with questions available in 24 languages. We evaluate LLMs by prompting them to answer questions from CROQ and provide a sample location. The results show that, contrary to previous cultural bias work, LLMs show a clear tendency towards countries such as Japan in their answers. Moreover, our results show that when prompting in languages such as English or other high-resource ones, LLMs tend to provide more diverse outputs. Low-resource languages, on the other hand, show more inclinations towards answering questions highlighting countries for which the input language is an official language. Finally, we also investigate at which point of LLM training this cultural bias emerges, with our results suggesting that the first clear signs appear after supervised fine-tuning, and not during pre-training. Dataset available at https://huggingface.co/datasets/HiTZ/CROQ

论文arXiv AI 12:00

Benefits of Low-Cost Bio-Inspiration in the Age of Overparametrization

arXiv:2604.20365v2 Announce Type: replace-cross Abstract: While Central Pattern Generators (CPGs) and Multi-Layer Perceptrons (MLP) are widely used paradigms in robot control, few systematic studies have been performed on the relative merits of large parameter spaces in highly constrained settings. As opposed to traditional Machine Learning contexts, our input and output spaces are small and performance is bounded thus having more parameters may actively hinder the learning process instead of empowering it. To empirically measure this, we submit a given robot morphology, with limited proprioceptive capabilities, to controller optimisation under two bio-inspired paradigms (CPGs and MLPs) with evolutionary- and reinforcement- trainer protocols. By varying parameter spaces across multiple reward functions, we demonstrate that shallow MLPs and densely connected CPGs result in better performance when compared to deeper MLPs or Actor-Critic architectures. To account for the relationship between said performance and the number of parameters, we introduce a Parameter Impact metric which showcases diminishing returns for MLPs but not for CPGs. Taken together these results demonstrate, on a fixed quadrupedal morphology, the benefits of integrating prior bias when considering locomotion tasks with simple hinge actuators.

论文arXiv AI 12:00

Beyond Output Correctness: Benchmarking and Evaluating Large Language Model Reasoning in Coding Tasks

arXiv:2604.12379v2 Announce Type: replace-cross Abstract: Large language models (LLMs) increasingly rely on explicit reasoning to solve coding tasks, yet evaluating the quality of this reasoning remains challenging. Existing reasoning evaluators are not designed for coding, and current benchmarks focus primarily on code generation, leaving other coding tasks largely unexplored. We introduce CodeRQ-Bench, the first benchmark for evaluating LLM reasoning quality across three coding task categories: generation, summarization, and classification. Using this benchmark, we analyze 1,069 mismatch cases from existing evaluators, identify five recurring limitations, and derive four design insights for reasoning evaluation in coding tasks. Guided by these insights, we propose VERA, a two-stage evaluator that combines evidence-grounded verification with ambiguity-aware score correction. Experiments on CodeRQ-Bench show that VERA consistently outperforms strong baselines across four datasets, improving AUCROC by up to 0.26 and AUPRC by up to 0.21. We release CodeRQ-Bench at https://github.com/MrLYG/CodeRQ-Bench, supporting future investigations.

论文arXiv AI 12:00

PolicyLong: Towards On-Policy Context Extension

arXiv:2604.07809v2 Announce Type: replace-cross Abstract: Extending LLM context windows is hindered by scarce high-quality long-context data. Recent methods synthesize data with genuine long-range dependencies via information-theoretic verification, selecting contexts that reduce a base model's predictive entropy. However, their single-pass offline construction with a fixed model creates a fundamental off-policy gap: the static screening landscape misaligns with the model's evolving capabilities, causing the training distribution to drift. We propose PolicyLong, shifting data construction towards a dynamic on-policy paradigm. By iteratively re-executing data screening (entropy computation, retrieval, and verification) using the current model, PolicyLong ensures the training distribution tracks evolving capabilities, yielding an emergent self-curriculum. Crucially, both positive and hard negative contexts derive from the current model's entropy landscape, co-evolving what the model learns to exploit and resist. Experiments on RULER, HELMET, and LongBench-v2 (Qwen2.5-3B) show PolicyLong consistently outperforms EntropyLong and NExtLong, with gains growing at longer contexts (e.g., +2.54 at 128K on RULER), confirming the value of on-policy data evolution.

论文arXiv AI 12:00

Scientific Graphics Program Synthesis via Dual Self-Consistency Reinforcement Learning

arXiv:2604.06079v2 Announce Type: replace-cross Abstract: Graphics Program Synthesis is pivotal for interpreting and editing visual data, effectively facilitating the reverse-engineering of static visuals into editable TikZ code. While TikZ is the de facto standard for scientific schematics due to its programmatic flexibility, its requirement for rigorous spatial precision presents a significant challenge for Multimodal Large Language Models. Progress is currently stifled by two primary gaps: (1) Data Quality Gap: existing image-TikZ corpora often lack strict executability and reliable visual alignment; (2) Evaluation Gap: a lack of benchmarks for both structural and visual fidelity. To address these, we present a closed-loop framework featuring: SciTikZ-230K, a large-scale, high-quality dataset from our Execution-Centric Data Engine covering 11 diverse scientific disciplines; SciTikZ-Bench, a multifaceted benchmark spanning from basic geometric constructs to intricate hierarchical schematics to evaluate both visual fidelity and structural logic. To further broaden the scope of visual-code optimization methodology, we introduce a novel Dual Self-Consistency Reinforcement Learning optimization paradigm, which utilizes Round-Trip Verification to penalize degenerate code and boost overall self-consistency. Empowered by these, our trained model SciTikZer-8B achieves state-of-the-art performance, consistently outperforming proprietary giants like Gemini-2.5-Pro and massive models like Qwen3-VL-235B-A22B-Instruct.

论文arXiv AI 12:00

Camera-Agnostic Pruning of 3D Gaussian Splats via Descriptor-Based Beta Evidence

arXiv:2603.21933v3 Announce Type: replace-cross Abstract: The pruning of 3D Gaussian splats is essential for reducing their complexity to enable efficient storage, transmission, and downstream processing. However, most of the existing pruning strategies depend on camera parameters, rendered images, or view-dependent measures. This dependency becomes a hindrance in emerging camera-agnostic exchange settings, where splats are shared directly as point-based representations (e.g., .ply). In this paper, we propose a camera-agnostic, one-shot, post-training pruning method for 3D Gaussian splats that relies solely on attribute-derived neighbourhood descriptors. As our primary contribution, we introduce a hybrid descriptor framework that captures structural and appearance consistency directly from the splat representation. Building on these descriptors, we formulate pruning as a statistical evidence estimation problem and introduce a Beta evidence model that quantifies per-splat reliability through a probabilistic confidence score. Experiments conducted on standardized test sequences defined by the ISO/IEC MPEG Common Test Conditions (CTC) demonstrate that our approach achieves substantial pruning while preserving reconstruction quality, establishing a practical and generalizable alternative to existing camera-dependent pruning strategies.

论文arXiv AI 12:00

Select, Label, Evaluate: Active Testing in NLP

arXiv:2603.21840v2 Announce Type: replace-cross Abstract: Human annotation cost and time remain significant bottlenecks in Natural Language Processing (NLP), with test data annotation being particularly expensive due to the stringent requirement for low-error and high-quality labels necessary for reliable model evaluation. Traditional approaches require annotating entire test sets, leading to substantial resource requirements. Active Testing is a framework that selects the most informative test samples for annotation. Given a labeling budget, it aims to choose the subset that best estimates model performance while minimizing cost and human effort. In this work, we formalize Active Testing in NLP and we conduct an extensive benchmarking of existing approaches across 18 datasets and 4 embedding strategies spanning 4 different NLP tasks. The experiments show annotation reductions of up to 95%, with performance estimation accuracy difference from the full test set within 1%. Our analysis reveals variations in method effectiveness across different data characteristics and task types, with no single approach emerging as universally superior. Lastly, to address the limitation of requiring a predefined annotation budget in existing sample selection strategies, we introduce an adaptive stopping criterion that automatically determines the optimal number of samples. We release our code at https://github.com/amazon-science/NLPActiveTesting.

论文arXiv AI 12:00

Var-JEPA: A Variational Formulation of the Joint-Embedding Predictive Architecture - Bridging Predictive and Generative Self-Supervised Learning

arXiv:2603.20111v2 Announce Type: replace-cross Abstract: The Joint-Embedding Predictive Architecture (JEPA) is often seen as a non-generative alternative to likelihood-based self-supervised learning, emphasizing prediction in representation space rather than reconstruction in observation space. We argue that the resulting separation from probabilistic generative modeling is largely rhetorical rather than structural: the canonical JEPA design (coupled encoders with a context-to-target predictor) mirrors the variational posteriors and learned conditional priors obtained when variational inference is applied to a particular class of coupled latent-variable models, and standard JEPA can be viewed as a deterministic specialization in which regularization is imposed via architectural and training heuristics rather than an explicit likelihood. Building on this view, we derive the Variational JEPA (Var-JEPA), which makes the latent generative structure explicit by optimizing a single Evidence Lower Bound (ELBO). This yields meaningful representations without ad-hoc anti-collapse regularizers and allows principled uncertainty quantification in the latent space. We instantiate the framework for tabular data (Var-T-JEPA) and achieve strong representation learning and downstream performance, improving over T-JEPA across real-world tabular benchmarks while remaining competitive with strong raw-feature baselines.

论文arXiv AI 12:00

The Autonomy Tax: Defense Training Breaks LLM Agents

arXiv:2603.19423v3 Announce Type: replace-cross Abstract: Large language model (LLM) agents increasingly rely on external tools (file operations, API calls, database transactions) to autonomously complete complex multi-step tasks. Practitioners deploy defense-trained models to protect against prompt injection attacks that manipulate agent behavior through malicious observations or retrieved content. We reveal a fundamental \textbf{capability-alignment paradox}: defense training designed to improve safety systematically destroys agent competence while failing to prevent sophisticated attacks. Evaluating defended models against undefended baselines across 97 agent tasks and 1,000 adversarial prompts, we uncover three systematic biases unique to multi-step agents. \textbf{Agent incompetence bias} manifests as immediate tool execution breakdown, with models refusing or generating invalid actions on benign tasks before observing any external content. \textbf{Cascade amplification bias} causes early failures to propagate through retry loops, pushing defended models to timeout on 99\% of tasks compared to 13\% for baselines. \textbf{Trigger bias} leads to paradoxical security degradation where defended models perform worse than undefended baselines while straightforward attacks bypass defenses at high rates. Root cause analysis reveals these biases stem from shortcut learning: models overfit to surface attack patterns rather than semantic threat understanding, evidenced by extreme variance in defense effectiveness across attack categories. Our findings demonstrate that current defense paradigms optimize for single-turn refusal benchmarks while rendering multi-step agents fundamentally unreliable, necessitating new approaches that preserve tool execution competence under adversarial conditions.

论文arXiv AI 12:00

InfoMamba: An Attention-Free Hybrid Mamba-Transformer Model

arXiv:2603.18031v2 Announce Type: replace-cross Abstract: Balancing fine-grained local modeling with long-range dependency capture under computational constraints remains a central challenge in sequence modeling. While Transformers provide strong token mixing, they suffer from quadratic complexity, whereas Mamba-style selective state-space models (SSMs) scale linearly but often struggle to capture high-rank and synchronous global interactions. We present a consistency boundary analysis that characterizes when diagonal short-memory SSMs can approximate causal attention and identifies structural gaps that remain. Motivated by this analysis, we propose InfoMamba, an attention-free hybrid architecture. InfoMamba replaces token-level self-attention with a concept bottleneck linear filtering layer that serves as a minimal-bandwidth global interface and integrates it with a selective recurrent stream through information-maximizing fusion (IMF). IMF dynamically injects global context into the SSM dynamics and encourages complementary information usage through a mutual-information-inspired objective. Extensive experiments on classification, dense prediction, and non-vision tasks show that InfoMamba consistently outperforms strong Transformer and SSM baselines, achieving competitive accuracy-efficiency trade-offs while maintaining near-linear scaling.

论文arXiv AI 12:00

Large Reasoning Models Struggle to Transfer Parametric Knowledge Across Scripts

arXiv:2603.17070v2 Announce Type: replace-cross Abstract: In this work, we analyze shortcomings in cross-lingual knowledge transfer in large, modern reasoning LLMs. We demonstrate that the perceived gap in knowledge transfer is primarily a script barrier. First, we conduct an observational data analysis on the performance of thinking models on two datasets with local knowledge from around the world, ECLeKTic and MultiLoKo. Our regression analysis shows that script match - not language or family - is the primary predictor of knowledge transfer failure once model capability and question difficulty are accounted for. We further this finding by providing the LLMs with the key entities of the questions in their source language and find that this disproportionately improves cross-script questions. We then posit that these LLMs could be reasoning better at test-time. To evaluate this, we develop a synthetic generation pipeline to design SFT samples to encourage the model to better reason about transliteration ambiguities when trying to fetch parametric knowledge at inference-time. We show that teaching two models to reason better reduces the cross-script transfer gap. As a result, we conclude that there is potential to improve cross-lingual parametric knowledge transfer during post-training.

论文arXiv AI 12:00

From Leaky Thoughts to Private Reasoning: Controlling What LRMs Say to Themselves

arXiv:2602.24210v4 Announce Type: replace-cross Abstract: Large reasoning models (LRMs) produce reasoning traces (RTs) that often contain sensitive information. These leaky thoughts are difficult to control and frequently violate explicit privacy directives. Because RTs can be exposed through prompt injection attacks, this becomes a direct privacy risk to the user. We approach this as a controllability problem: since privacy directives are themselves instructions, improving instruction-following (IF) within the RT provides a direct path to reducing privacy leaks. To this end, we introduce an SFT dataset that teaches models to follow general instructions throughout their reasoning process, and propose Staged Decoding, a simple decoding strategy that decouples RT and answer generation using separate LoRA adapters to maximize IF of each component. We evaluate our approach on six models from two families (1.7B-14B parameters), across two IF benchmarks and two privacy benchmarks. Our method yields substantial improvements, with gains of up to 20.9 points in IF and 51.9 percentage points on privacy benchmarks, though these can come at the cost of task utility due to the trade-off between reasoning performance and IF. Our results show that improving IF in LRMs can significantly enhance privacy, suggesting a promising direction for future privacy-aware LRMs. Our code is available at https://github.com/UKPLab/arxiv2026-controllable-reasoning-models.

论文arXiv AI 12:00

FENCE: A Financial and Multimodal Jailbreak Detection Dataset

arXiv:2602.18154v3 Announce Type: replace-cross Abstract: Jailbreaking poses a significant risk to the deployment of Large Language Models (LLMs) and Vision Language Models (VLMs). VLMs are particularly vulnerable because they process both text and images, creating broader attack surfaces. However, available resources for jailbreak detection are scarce, particularly in finance. To address this gap, we present FENCE, a bilingual (Korean-English) multimodal dataset for training and evaluating jailbreak detectors in financial applications. FENCE emphasizes domain realism through finance-relevant queries paired with image-grounded threats. Experiments with commercial and open-source VLMs reveal consistent vulnerabilities, with GPT-4o showing measurable attack success rates and open-source models displaying greater exposure. A baseline detector trained on FENCE achieves 99 percent in-distribution accuracy and maintains strong performance on external benchmarks, underscoring the dataset's robustness for training reliable detection models. FENCE provides a focused resource for advancing multimodal jailbreak detection in finance and for supporting safer, more reliable AI systems in sensitive domains. Warning: This paper includes example data that may be offensive.

论文arXiv AI 12:00

ASA: Backbone-Training-Free Representation Engineering for Tool-Calling Agents

arXiv:2602.04935v4 Announce Type: replace-cross Abstract: Adapting LLM agents to domain-specific tool calling remains notably brittle under evolving interfaces. Prompt and schema engineering is easy to deploy but often fragile under distribution shift and strict parsers, while continual parameter-efficient fine-tuning improves reliability at the cost of training, maintenance, and potential forgetting. We identify a critical Lazy Agent failure mode where tool necessity is nearly perfectly decodable from mid-layer activations, yet the model remains conservative in entering tool mode, revealing a representation-behavior gap. We propose Activation Steering Adapter (ASA), a training-free, inference-time controller that performs a single-shot mid-layer intervention and targets tool domains via a router-conditioned mixture of steering vectors with a probe-guided signed gate to amplify true intent while suppressing spurious triggers. On MTU-Bench with Qwen2.5-1.5B, ASA improves strict tool-use F1 from 0.18 to 0.50 while reducing the false positive rate from 0.15 to 0.05, using only about 20KB of portable assets and no weight updates.

论文arXiv AI 12:00

SCALE: Self-uncertainty Conditioned Adaptive Looking and Execution for Vision-Language-Action Models

arXiv:2602.04208v3 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models have emerged as a promising paradigm for general-purpose robotic control, with test-time scaling (TTS) gaining attention to enhance robustness beyond training. However, existing TTS methods for VLAs require additional training, verifiers, and multiple forward passes, making them impractical for deployment. Moreover, they intervene only at action decoding while keeping visual representations fixed-insufficient under perceptual ambiguity, where reconsidering how to perceive is as important as deciding what to do. To address these limitations, we propose SCALE, a simple inference strategy that jointly modulates visual perception and action based on 'self-uncertainty', inspired by uncertainty-driven exploration in Active Inference theory-requiring no additional training, no verifier, and only a single forward pass. SCALE broadens exploration in both perception and action under high uncertainty, while focusing on exploitation when confident-enabling adaptive execution across varying conditions. Experiments on simulated and real-world benchmarks demonstrate that SCALE improves state-of-the-art VLAs and outperforms existing TTS methods while maintaining single-pass efficiency.

论文arXiv AI 12:00

Beyond Pixels: Visual Metaphor Transfer via Schema-Driven Agentic Reasoning

arXiv:2602.01335v2 Announce Type: replace-cross Abstract: A visual metaphor constitutes a high-order form of human creativity, employing cross-domain semantic fusion to transform abstract concepts into impactful visual rhetoric. Despite the remarkable progress of generative AI, existing models remain largely confined to pixel-level instruction alignment and surface-level appearance preservation, failing to capture the underlying abstract logic necessary for genuine metaphorical generation. To bridge this gap, we introduce the task of Visual Metaphor Transfer (VMT), which challenges models to autonomously decouple the "creative essence" from a reference image and re-materialize that abstract logic onto a user-specified target subject. We propose a cognitive-inspired, multi-agent framework that operationalizes Conceptual Blending Theory (CBT) through a novel Schema Grammar ("G"). This structured representation decouples relational invariants from specific visual entities, providing a rigorous foundation for cross-domain logic re-instantiation. Our pipeline executes VMT through a collaborative system of specialized agents: a perception agent that distills the reference into a schema, a transfer agent that maintains generic space invariance to discover apt carriers, a generation agent for high-fidelity synthesis and a hierarchical diagnostic agent that mimics a professional critic, performing closed-loop backtracking to identify and rectify errors across abstract logic, component selection, and prompt encoding. Extensive experiments and human evaluations demonstrate that our method significantly outperforms SOTA baselines in metaphor consistency, analogy appropriateness, and visual creativity, paving the way for automated high-impact creative applications in advertising and media. Project page with source code and self-contained skills is at https://yuci-gpt.github.io/Beyond-Pixels/.

论文arXiv AI 12:00

CoFrGeNet: Continued Fraction Architectures for Language Generation

arXiv:2601.21766v5 Announce Type: replace-cross Abstract: Transformers are arguably the preferred architecture for language generation. In this paper, inspired by continued fractions, we introduce a new function class for generative modeling. The architecture family implementing this function class is named CoFrGeNets - Continued Fraction Generative Networks. We design novel architectural components based on this function class that can replace Multi-head Attention and Feed-Forward Networks in Transformer blocks while requiring much fewer parameters. We derive custom gradient formulations to optimize the proposed components more accurately and efficiently than using standard PyTorch-based gradients. Our components are a plug-in replacement requiring little change in training or inference procedures that have already been put in place for Transformer-based models thus making our approach easy to incorporate in large industrial workflows. We experiment on two very different transformer architectures GPT2-xl (1.5B) and Llama3 (3.2B), where the former we pre-train on OpenWebText and GneissWeb, while the latter we pre-train on the docling data mix which consists of nine different datasets. Results show that the performance on downstream classification, Q\& A, reasoning and text understanding tasks of our models is competitive and sometimes even superior to the original models with $\frac{2}{3}$ to $\frac{1}{2}$ the parameters and shorter pre-training time. We believe that future implementations customized to hardware will further bring out the true potential of our architectures.

论文arXiv AI 12:00

Aligning Agentic World Models via Knowledgeable Experience Learning

arXiv:2601.13247v2 Announce Type: replace-cross Abstract: Current Large Language Models (LLMs) exhibit a critical modal disconnect: they possess vast semantic knowledge but lack the procedural grounding to respect the immutable laws of the physical world. Consequently, while these agents implicitly function as world models, their simulations often suffer from physical hallucinations-generating plans that are logically sound but physically unexecutable. Existing alignment strategies predominantly rely on resource-intensive training or fine-tuning, which attempt to compress dynamic environmental rules into static model parameters. However, such parametric encapsulation is inherently rigid, struggling to adapt to the open-ended variability of physical dynamics without continuous, costly retraining. To bridge this gap, we introduce WorldMind, a framework that autonomously constructs a symbolic World Knowledge Repository by synthesizing environmental feedback. Specifically, it unifies Process Experience to enforce physical feasibility via prediction errors and Goal Experience to guide task optimality through successful trajectories. Experiments on EB-ALFRED and EB-Habitat demonstrate that WorldMind achieves superior performance compared to baselines with remarkable cross-model and cross-environment transferability.

论文arXiv AI 12:00

FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation

arXiv:2601.06199v5 Announce Type: replace-cross Abstract: Scaling Multimodal Large Language Models (MLLMs) to long-form speech is bottlenecked by the explosive growth of input tokens. Existing speech-language models project high-frame-rate acoustic features directly into the LLM input space, making long-context processing computationally prohibitive. Unlike images or videos, speech lacks spatial redundancy, making extreme token compression particularly challenging. To address this limitation, we propose FastSLM, a token-efficient architecture featuring the Hierarchical Temporal Abstractor (HTA), which progressively distills acoustic features across multiple temporal scales. HTA achieves an extreme compression rate of 1.67 tokens per second (97% reduction) while preserving essential linguistic information for downstream speech-language understanding. Experimental results demonstrate that FastSLM achieves competitive performance across diverse speech-language tasks while requiring substantially fewer speech tokens and FLOPs than existing speech-language models. The source code and model checkpoints are available at https://github.com/Lee-junseok1025/FastSLM.

论文arXiv AI 12:00

The Instability of Safety: How Random Seeds and Temperature Expose Inconsistent LLM Refusal Behavior

arXiv:2512.12066v3 Announce Type: replace-cross Abstract: Current safety evaluations of large language models rely on single-shot testing, implicitly assuming that model responses are deterministic and representative of the model's safety alignment. We challenge this assumption by investigating the stability of safety refusal decisions across random seeds and temperature settings. Testing four instruction-tuned models from three families (Llama 3.1 8B, Qwen 2.5 7B, Qwen 3 8B, Gemma 3 12B) on 876 harmful prompts across 20 sampling configurations (4 temperatures x 5 seeds), we find that 18-28% of prompts exhibit decision flips--the model refuses in some configurations but complies in others--depending on the model. Our Safety Stability Index (SSI) reveals that higher temperatures significantly reduce decision stability (Friedman chi-squared = 396.81, p < 0.001), with mean within-temperature SSI dropping from 0.977 at temperature 0.0 to 0.942 at temperature 1.0. We validate findings across all model families using Claude 3.5 Haiku as a unified external judge, achieving 89.1% inter-judge agreement with the Llama 70B judge on the two models both judges labeled (Cohen's kappa = 0.62). Within each model, prompts with higher compliance rates exhibit lower stability (Spearman rho = -0.47 to -0.70, all p < 0.001), indicating that models "waver" more on borderline requests. These findings demonstrate that single-shot safety evaluations are insufficient for reliable safety assessment and that evaluation protocols must account for stochastic variation in model behavior. For Llama 3.1 8B, single-shot evaluation agrees with multi-sample ground truth only 92.5% of the time when pooling across temperatures (98.7% at greedy to 90.3% at temperature 1.0), and we recommend scaling samples with temperature--one at greedy, three at low temperature, more at higher temperatures (where three reach only ~95%), and ten when pooling--rather than a single flat threshold.

论文arXiv AI 12:00

OmniFusion: Simultaneous Multilingual Multimodal Translations via Modular Fusion

arXiv:2512.00234v3 Announce Type: replace-cross Abstract: There has been significant progress in open-source text-only translation large language models (LLMs) with better language coverage and quality. However, these models can be only used in cascaded pipelines for speech translation (ST), performing automatic speech recognition first followed by translation. This introduces additional latency, which is particularly critical in simultaneous ST (SimulST), and prevents the model from exploiting multimodal context, such as images, which can aid disambiguation. Pretrained multimodal foundation models (MMFMs) already possess strong perception and reasoning capabilities across multiple modalities, but generally lack the multilingual coverage and specialized translation performance of dedicated translation LLMs. To build an effective multimodal translation system, we propose an end-to-end approach that fuses MMFMs with translation LLMs. We introduce a novel fusion strategy that connects hidden states from multiple layers of a pretrained MMFM to a translation LLM, enabling joint end-to-end training. The resulting model, OmniFusion, built on Omni 2.5-7B as the MMFM and SeedX PPO-7B as the translation LLM, can perform speech-to-text, speech-and-image-to-text, and text-and-image-to-text translation. Experiments demonstrate that OmniFusion effectively leverages both audio and visual inputs, achieves a 1-second latency reduction in SimulST compared to cascaded pipelines and also improves the overall translation quality\footnote{Code is available at https://github.com/saikoneru/OmniFusion}.

论文arXiv AI 12:00

Think-at-Hard: Dynamic Looped Transformers for Improved Reasoning

arXiv:2511.08577v4 Announce Type: replace-cross Abstract: Improving the reasoning abilities of Large Language Models (LLMs), especially under parameter constraints, is crucial for real-world applications. Looped transformers address this by performing multiple latent iterations to refine each token beyond a single forward pass. However, we identify a latent overthinking phenomenon: most token predictions are already correct after the first pass, but are sometimes revised into errors in later iterations. We ask whether selectively skipping latent iterations can improve accuracy, and reveal significant potential with an oracle iteration policy that boosts performance by up to 7.3%. Motivated by this, we propose Think-at-Hard (TaH), a looped transformer optimized for selective iteration. TaH employs a lightweight neural decider to trigger latent iteration, only at tokens likely to be incorrect after the standard forward pass. During latent iterations, depth-aware Low-Rank Adaptation (LoRA) modules shift the objective from general next-token prediction to focused hard-token refinement. A duo-causal attention mechanism extends attention from the token sequence dimension to an additional iteration depth dimension, enabling cross-iteration information flow with full sequential parallelism. Experiments on nine benchmarks show consistent gains across math, QA, and coding tasks. With identical parameter counts, TaH outperforms always-iterate baselines by 3.8-4.4% while skipping iterations on 93% of tokens, and exceeds single-iteration Qwen3 baselines by 3.0-3.8%. When allowing

论文arXiv AI 12:00

Quantifying Affective Bias in Low-Resource Media: Large-Scale Emotion Profiling of Bengali Headlines

arXiv:2510.17252v2 Announce Type: replace-cross Abstract: News media can influence readers not only through the events they report but also through the emotional tone used to present them. This issue is especially important in digital news environments, where headlines often shape first impressions before readers open the full article. This study examines affective framing in Bengali digital journalism through corpus level emotion analysis of news headlines. Using zero shot inference with Gemma 3 4B, we analyzed 300,000 Bengali news headlines to estimate the dominant emotion and overall affective tone of each headline. The results show that negative emotion labels, particularly anger, sadness, disappointment, and fear, appear frequently in the analyzed corpus. A small pilot validation on 200 manually reviewed headlines suggests that the model can provide useful emotion estimates, although the results should be interpreted as computational estimates rather than a complete benchmark. Based on these findings, we propose a conceptual bias sensitive news interface that visualizes emotional cues across news sources and helps readers notice affective framing patterns in daily news.

论文arXiv AI 12:00

Riverbank Erosion Analysis in Bangladesh Using Spatiotemporal Segmentation

arXiv:2510.17198v2 Announce Type: replace-cross Abstract: Riverbank erosion is a serious environmental problem in Bangladesh, causing land loss, damage to infrastructure, and displacement of local communities. Manual analysis of satellite images is often slow and difficult to apply consistently across large river networks. This study uses a parameter-efficient adaptation of the Segment Anything Model (SAM) to detect and measure riverbank erosion from historical Google Earth images. A dataset of 500 image pairs from 2003 to 2025 was prepared from erosion-prone areas, including Mokterer Char, Kedarpur, and Chowhali Upazila, with pixel-level labels for river, stable land, and eroded regions. During training, the ViT-H image encoder and prompt encoder were kept frozen, while only the lightweight mask decoder was fine-tuned for riverine segmentation. The adapted model achieved an erosion-class IoU of 0.867 and an F1-score of 0.928 on the primary held-out test set. Evaluation on unseen riverbank regions also showed that the model could generalize to new geographic areas, although detecting accreted land from RGB-only images remained difficult. The estimated erosion area differed from the ground truth by only 0.17%, showing that the model can produce reliable area measurements. Overall, this study demonstrates that adapted foundation segmentation models can support faster and more consistent riverbank erosion monitoring, with future scope for using multi-modal remote sensing data in broader environmental assessment.

论文arXiv AI 12:00

PRISM: Agentic Retrieval with LLMs for Multi-Hop Question Answering

arXiv:2510.14278v2 Announce Type: replace-cross Abstract: Retrieval plays a central role in multi-hop question answering (QA), where answering complex questions requires gathering multiple pieces of evidence. We propose PRISM, an agentic retrieval framework that leverages large language models (LLMs) in a structured loop to retrieve relevant evidence with high precision and recall. PRISM decomposes retrieval into three specialized agents: a Question Analyzer that breaks complex queries into sub-questions, a Selector that identifies the most relevant context for each sub-question (focusing on precision), and an Adder that brings in any missing evidence (focusing on recall). The iterative interaction between the Selector and Adder produces a compact yet comprehensive evidence set, avoiding both brittle error propagation and noisy context accumulation. It achieves higher retrieval accuracy while filtering out distracting content, enabling downstream QA models to surpass full-context answer accuracy while relying on significantly less irrelevant information. Experiments on four challenging multi-hop QA benchmarks, including HotpotQA, 2WikiMultiHopQA, MuSiQue, and MultiHopRAG, demonstrate that our approach consistently outperforms strong baselines.

论文arXiv AI 12:00

OceanGym: A Benchmark Environment for Underwater Embodied Agents

arXiv:2509.26536v3 Announce Type: replace-cross Abstract: We introduce OceanGym, the first comprehensive benchmark for ocean underwater embodied agents, designed to advance AI in one of the most demanding real-world environments. Unlike terrestrial or aerial domains, underwater settings present extreme perceptual and decision-making challenges, including low visibility, dynamic ocean currents, making effective agent deployment exceptionally difficult. OceanGym encompasses eight realistic task domains and a unified agent framework driven by Multi-modal Large Language Models (MLLMs), which integrates perception, memory, and sequential decision-making. Agents are required to comprehend optical and sonar data, autonomously explore complex environments, and accomplish long-horizon objectives under these harsh conditions. Extensive experiments reveal substantial gaps between state-of-the-art MLLM-driven agents and human experts, highlighting the persistent difficulty of perception, planning, and adaptability in ocean underwater environments. By providing a high-fidelity, rigorously designed platform, OceanGym establishes a testbed for developing robust embodied AI and transferring these capabilities to real-world autonomous ocean underwater vehicles, marking a decisive step toward intelligent agents capable of operating in one of Earth's last unexplored frontiers. The code and data are available at https://github.com/OceanGPT/OceanGym.

论文arXiv AI 12:00

Talk in Pieces, See in Whole: Disentangled and Hierarchical Representation Learning in Language-based Object Detection

arXiv:2509.24192v3 Announce Type: replace-cross Abstract: Vision-language models (VLMs) have advanced multimodal perception, demonstrated by open-vocabulary object detection with simple language queries. State-of-the-art VLMs still struggle to handle complex queries involving descriptive attributes and relational clauses. To address this problem, we propose restructuring linguistic representations according to the hierarchical relations within sentences for language-based object detection. A key insight is that textual tokens should be disentangled into core components-objects, attributes, and relations-and aggregated into hierarchically structured sentence-level representations. Building on this principle, we introduce the TaSe (Talk in Pieces, See in Whole) framework with three main contributions: (1) a hierarchical synthetic captioning dataset spanning three tiers from category names to descriptive sentences; (2) the three-component disentanglement module guided by a novel disentanglement loss function, transforms text embeddings into subspace compositions; and (3) aggregating disentangled components into hierarchically structured embeddings guided by the proposed hierarchical objectives. Experimental results under the OmniLabel benchmark show a 24% performance improvement, demonstrating the importance of linguistic compositionality.

论文arXiv AI 12:00

CompareBench: A Benchmark for Visual Comparison Reasoning in Vision-Language Models

arXiv:2509.22737v3 Announce Type: replace-cross Abstract: Visual comparison reasoning is a fundamental capability of vision-language models (VLMs), covering judgments of object quantity, geometric dimensions, spatial relations, and temporal order. Yet existing benchmarks rarely isolate comparison as a reasoning axis, leaving it unclear whether models can reliably perform comparative visual judgments. We introduce a benchmark suite organized around three top-level resources: TallyBench, a 2,000-image object counting benchmark; OmniCaps, a 716-image caption and tag resource; and CompareBench, a 1,200-QA visual comparison benchmark. CompareBench contains four sub-benchmarks spanning quantity, geometric, spatial, and temporal comparison, with the temporal component unifying historical scenes, landmarks, and public figures. Evaluating nine closed-source model routes from Anthropic, Google, and OpenAI on TallyBench and CompareBench reveals strong overall performance but persistent failures in counting, spatial reasoning, geometric comparison, and temporal ordering. These results show that visual comparison remains a systematic weakness of current VLMs and establish CompareBench as a focused benchmark for multimodal reasoning evaluation. All data, code, and prompts will be released at https://github.com/caijie0620/CompareBench.

论文arXiv AI 12:00

Steering Multimodal Large Language Models Decoding for Context-Aware Safety

arXiv:2509.19212v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) are increasingly deployed in real-world applications, yet their ability to make context-aware safety decisions remains limited. Existing methods often fail to balance oversensitivity (unjustified refusals of benign queries) and undersensitivity (missed detection of visually grounded risks), leaving a persistent gap in safety alignment. To address this issue, we introduce Safety-aware Contrastive Decoding (SafeCoDe), a lightweight and model-agnostic decoding framework that dynamically adjusts token generation based on multimodal context. SafeCoDe operates in two stages: (1) a contrastive decoding mechanism that highlights tokens sensitive to visual context by contrasting real and Gaussian-noised images, and (2) a global-aware token modulation strategy that integrates scene-level reasoning with token-level adjustment to adapt refusals according to the predicted safety verdict. Extensive experiments across diverse MLLM architectures and safety benchmarks, covering undersensitivity, oversensitivity, and general safety evaluations, show that SafeCoDe consistently improves context-sensitive refusal behaviors while preserving model helpfulness.

论文arXiv AI 12:00

Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning

arXiv:2509.00094v2 Announce Type: replace-cross Abstract: Assessing spoken language is challenging, and quantifying pronunciation metrics for machine learning models is even harder. However, for the Holy Quran, this task is enabled by the rigorous recitation rules (Tajweed) established through the efforts of Muslim scholars, making highly effective assessment possible. Despite this advantage, the scarcity of high-quality annotated data remains a significant barrier. In this work, we bridge these gaps by introducing: (1) A 98% automated pipeline to produce high-quality Quranic datasets -- encompassing collection of recitations from expert reciters, segmentation at pause points (waqf) using our fine-tuned wav2vec2-BERT model, transcription of segments, and transcript verification via our novel Tasmeea algorithm; (2) 848 hours of audio (286K annotated utterances); (3) qdat_bench, a benchmark covering phonemes, diacritization, and Tajweed rules (Ghunnah, Qalqalah, Madd) on real recitation errors containing 159 samples; (4) A novel ASR-based approach for pronunciation error detection utilizing our custom Quran Phonetic Script (QPS) to encode Tajweed rules (unlike the IPA standard for Modern Standard Arabic). QPS uses an 11-level script: phoneme level (encoding Arabic letters with short/long vowels) and sifat level (encoding articulation characteristics of every phoneme). We further present comprehensive modeling with our novel multi-level CTC model, which achieved 0.21% and 1.94% average Phoneme Error Rate (PER) on the test set and qdat_bench respectively, with a 75.8% Tajweed F1 score. We release our work as open-source: https://obadx.github.io/quran-muaalem/en/

论文arXiv AI 12:00

Beyond the Rosetta Stone: Unification Forces in Generalization Dynamics

arXiv:2508.11017v4 Announce Type: replace-cross Abstract: Large language models (LLMs) struggle with cross-lingual knowledge transfer: they sometimes hallucinate when asked in one language about facts expressed in a different language during training. This work introduces a controlled setting to study the causes and training dynamics of this phenomenon by training small Transformer models from scratch on synthetic multilingual datasets. Depending on (1) the correlation between facts and the language they were learned in (informativeness), and (2) the ease of language identification (extractability), models either develop unified representations across languages or separate representations; only when representations are unified do facts transfer across languages. Based on these insights, we propose a unifying perspective which explains a range of prior observations concerning cross-lingual transfer in multilingual LLMs. Our work shows controlled settings can shed light on pre-training dynamics and suggests methods to encourage representational unification as part of training that would improve LLMs' cross-lingual transfer.

论文arXiv AI 12:00

Attention as Conditioning: What Classical Learning Theory Predicts About Linear Transformers

arXiv:2508.08289v3 Announce Type: replace-cross Abstract: Attention is widely understood as an associative memory, but that description alone does not predict how the memory will behave. Predictive theories do exist, but in the literature on animal learning. We show that the state updates of the major linear-attention families are term-for-term identical with named models from a century of animal learning theory: linear attention implements Hebbian contiguity, DeltaNet implements Rescorla--Wagner error correction, and decay variants such as RetNet implement contiguity with a stimulus trace. This dictionary turns conditioning phenomena into testable statements about the in-context behavior of linear transformers, while distinguishing algebraic consequences from empirical measurements. Algebraically, it yields an exact closed form for Kamin blocking, verified in simulation to $

论文arXiv AI 12:00

Cognitive Chain-of-Thought (CoCoT): Structured Multimodal Reasoning about Social Situations

arXiv:2507.20409v3 Announce Type: replace-cross Abstract: Chain-of-Thought (CoT) prompting helps models think step by step. But naive CoT breaks down in visually grounded social tasks, where models must perceive, understand, and judge all at once; bridging perception with norm-grounded reasoning. Recent work has introduced structured reasoning for multi-turn agent planning and visual QA, decomposing tasks into sequential sub-goals. To extend this to single-shot multimodal social reasoning, we introduce Cognitive Chain-of-Thought (CoCoT), a reasoning framework that structures vision-language-model (VLM) reasoning through three cognitively inspired stages: Perception (extract grounded facts), Situation (infer situations), and Norm (applying social norms). Evaluation across multiple distinct tasks such as multimodal intent disambiguation, multimodal theory of mind, social commonsense reasoning, and safety instruction following, shows consistent improvements (5.9% to 4.6% on average). We further explore the utility of CoCoT for improving models' reasoning through training and show that supervised fine-tuning on CoCoT-structured traces yields 5-6% improvements without explicit CoCoT prompting at inference, demonstrating that models internalize the structured reasoning pattern rather than merely following instructions. We show that structuring model reasoning through cognitively grounded stages enhances interpretability and social alignment, laying the groundwork for more reliable multimodal systems.

论文arXiv AI 12:00

PRISM: Self-Pruning Intrinsic Selection Method for Training-Free Multimodal Data Selection

arXiv:2502.12119v5 Announce Type: replace-cross Abstract: Visual instruction tuning adapts pre-trained Multimodal Large Language Models (MLLMs) to follow human instructions for real-world applications. However, the rapid growth of these datasets introduces significant redundancy, leading to increased computational costs. Existing methods for selecting instruction data aim to prune this redundancy, but predominantly rely on computationally demanding techniques such as proxy-based inference or training-based metrics. Consequently, the substantial computational costs incurred by these selection processes often exacerbate the very efficiency bottlenecks they are intended to resolve, posing a significant challenge to the scalable and effective tuning of MLLMs. To address this challenge, we first identify a critical, yet previously overlooked, factor: the anisotropy inherent in visual feature distributions. We find that this anisotropy induces a \textit{Global Semantic Drift}, and overlooking this phenomenon is a key factor limiting the efficiency of current data selection methods. Motivated by this insight, we devise \textbf{PRISM}, the first training-free framework for efficient visual instruction selection. PRISM surgically removes the corrupting influence of global background features by modeling the intrinsic visual semantics via implicit re-centering. Empirically, PRISM reduces the end-to-end time for data selection and model tuning to just 30\% of conventional pipelines. More remarkably, it achieves this efficiency while simultaneously enhancing performance, surpassing models fine-tuned on the full dataset across eight multimodal and three language understanding benchmarks, culminating in a 101.7\% relative improvement over the baseline. The code is available for access via \href{https://github.com/bibisbar/PRISM}{this repository}.

论文arXiv AI 12:00

Long Story Short: Story-level Video Understanding from 20K Short Films

arXiv:2406.10221v3 Announce Type: replace-cross Abstract: Recent developments in vision-language models have significantly advanced video understanding. Existing datasets and tasks, however, have notable limitations. Most datasets are confined to short videos with limited events and narrow narratives. For example, datasets with instructional and egocentric videos often depict the activities of one person in a single scene. Although existing movie datasets offer richer content, they are often limited to short-term tasks, lack publicly available videos, and frequently encounter data leakage issues given the use of subtitles and other information about commercial movies during LLM pretraining. To address the above limitations, we propose Short-Films 20K (SF20K), the largest publicly available movie dataset. SF20K consists of 20,143 amateur films, amounting to 3,582 hours of video, with an average of 12 minutes per movie. We accompany this dataset with SF20K-Test, a manual, open-ended question answering benchmark. SF20K-Test consists of 95 movies and 979 question-answer pairs. Our extensive analysis of SF20K-Test reveals limited data leakage, emphasizes the need for long-term reasoning, and demonstrates the strong performance of recent VLMs. Finally, we show that instruction tuning on the large-scale dataset substantially improves model performance, paving the way for future progress in long-term video understanding.

论文arXiv AI 12:00

Let the Flows Tell: Solving Graph Combinatorial Optimization Problems with GFlowNets

arXiv:2305.17010v4 Announce Type: replace-cross Abstract: Combinatorial optimization (CO) problems are often NP-hard and thus out of reach for exact algorithms, making them a tempting domain to apply machine learning methods. The highly structured constraints in these problems can hinder either optimization or sampling directly in the solution space. On the other hand, GFlowNets have recently emerged as a powerful machinery to efficiently sample from composite unnormalized densities sequentially and have the potential to amortize such solution-searching processes in CO, as well as generate diverse solution candidates. In this paper, we design Markov decision processes (MDPs) for different combinatorial problems and propose to train conditional GFlowNets to sample from the solution space. Efficient training techniques are also developed to benefit long-range credit assignment. Through extensive experiments on a variety of different CO tasks with synthetic and realistic data, we demonstrate that GFlowNet policies can efficiently find high-quality solutions. Our implementation is open-sourced at https://github.com/zdhNarsil/GFlowNet-CombOpt.

论文arXiv AI 12:00

Evaluating the Performance of Large Language Models on GAOKAO Benchmark

arXiv:2305.12474v4 Announce Type: replace-cross Abstract: Large Language Models(LLMs) have demonstrated remarkable performance across various natural language processing tasks; however, how to comprehensively and accurately assess their performance becomes an urgent issue to be addressed. This paper introduces GAOKAO-Bench, an intuitive benchmark that employs questions from the Chinese GAOKAO examination as test samples, including both subjective and objective questions. To align with human examination methods, we design a method based on zero-shot settings to evaluate the performance of LLMs. With human evaluation, we obtain the converted total score of LLMs, including GPT-4, ChatGPT and ERNIE-Bot.Our findings reveal that LLMs have achieved competitive scores in Chinese GAOKAO examination, while they exhibit significant performance disparities across various subjects. We also use LLMs to grade the subjective questions, and find that model scores achieve a moderate level of consistency with human scores. In conclusion, this research contributes a robust evaluation benchmark for future large language models and offers valuable insights into the advantages and limitations of such models.

论文arXiv AI 12:00

Transformer-Based Autonomous Driving Models and Deployment-Oriented Compression: A Survey

arXiv:2304.10891v4 Announce Type: replace-cross Abstract: Transformer-based models are becoming a central paradigm in autonomous driving because they can capture long-range spatial dependencies, multi-agent interactions, and multimodal context across perception, prediction, and planning. At the same time, their deployment in real vehicles remains difficult because high-capacity attention-based architectures impose substantial latency, memory, and energy overhead. This survey reviews representative Transformer-based autonomous driving models and organizes them by task role, sensing configuration, and architectural design. More importantly, it examines these models from a deployment-oriented perspective and analyzes how efficiency constraints reshape model design choices in practice. We further review compression and acceleration strategies relevant to Transformer-based driving systems, including quantization, pruning, knowledge distillation, low-rank approximation, and efficient attention, and discuss their benefits, limitations, and task-dependent applicability. Rather than treating compression as an isolated post-processing step, we highlight it as a system-level design consideration that directly affects deployability, robustness, and safety. Finally, we identify open challenges and future research directions toward standardized, safety-aware, and hardware-conscious evaluation of efficient autonomous driving systems.

论文arXiv AI 12:00

Mechanistic Reaction Prediction via Discrete Flow Matching on Graph-Structured Electron Occupation

arXiv:2608.27429v2 Announce Type: replace Abstract: Chemical reactions are fundamentally transformations in electron space, yet most machine learning approaches model them either through de novo generation of product molecules or through heuristic graph edits that operate directly on molecular topology. We introduce MAELLE (MechAnistic Edit fLow-matching on eLectron rEarrangements), which instead models reactions as discrete flow matching over electron occupation vectors. Concretely, we formulate the reactant-to-product mapping as a Continuous-time Markov Chain (CTMC) over the graph-structured integer-valued electron occupation space defined on all bonding, non-bonding, and hydrogen sites. To construct the interpolants between the reactants and products, we generalize the discrete flow matching mixture path to an edit-based formulation, where the electron moves are interpolated using Optimal Transport, yielding a mechanism-like set of moves without elementary step annotations. MAELLE achieves competitive performance on the USPTO-480K benchmark compared with leading reaction prediction models. Beyond in-distribution learning, we evaluate robustness across two out-of-distribution settings - structural complexity and reaction type - and find that MAELLE maintains strong performance where existing methods degrade. Finally, because the learned flow operates over the full electron redistribution, MAELLE naturally recovers mechanistic trajectories that align with known chemistry and can predict side products of a reaction.

论文arXiv AI 12:00

AgentFold: Closed-Loop Agentic Search for Protein Folding Model Design

arXiv:2608.26747v2 Announce Type: replace Abstract: Scientific LLM agents have shown promise in literature reasoning, tool use, and experiment planning, but it remains unclear whether they can autonomously improve large, tightly coupled scientific machine-learning systems through executable code changes and computationally expensive validation. We study this question in protein folding, where progress requires coordinated architectural modifications, multi-objective evaluation, and domain-aware interpretation. We present AgentFold, a multi-agent framework that formulates folding-model development as a closed-loop search over executable code variants. Starting from ESMFold, AgentFold proposes hypotheses, implements and debugs code-level modifications, evaluates model variants, analyzes experimental outcomes, and stores both successful and failed interventions in structured memory. An MCTS-style policy allocates computational resources across high-scoring search branches. On an engineering-scale protein-folding codebase comprising more than 2,000 lines of code, AgentFold explores approximately 80 model variants using approximately 5,000 GPU-hours and 170 million LLM tokens. Under a matched computational budget, AgentFold improves the best lDDT by 7.5% over independent Codex proposals and outperforms a random-search control. Beyond model improvement, the resulting intervention traces reveal recurring empirical design patterns: stable gains tend to arise from early, soft, learnable priors and gated refinement, whereas direct geometric perturbations and geometry-conditioned feedback often destabilize training. The code and experimental resources are publicly available at https://github.com/lmqfly/AgentFold.

论文arXiv AI 12:00

SKILL.state: Scalable Long-Horizon Agent Skills

arXiv:2608.26263v2 Announce Type: replace Abstract: Large Language Models (LLMs) increasingly act as autonomous agents executing complex, long-running procedural skills. Existing agent runtimes maintain execution by continually appending observations, actions, and intermediate reasoning traces to an ever-growing conversation history, causing latency degradation and context-poisoning failures over long horizons. We present SKILL.state, a runtime architecture that replaces append-only conversational history with an explicit, mutable execution state. At each execution step, the model receives only the immutable skill specification, the current structured execution state, and the latest observation. Intermediate reasoning is discarded immediately after producing a validated state update, preventing prompt growth with execution history. Across diverse datasets, models, and execution environments, SKILL. state improves task accuracy while substantially reducing cumulative token consumption. Our results demonstrate that explicit execution state is an effective and architecture-agnostic abstraction for scalable long-horizon agent skills.

论文arXiv AI 12:00

Beyond Confidence: Test-Time Scaling for Multi-Turn Search Agents via Retrieval Grounding

arXiv:2608.24024v2 Announce Type: replace Abstract: Confidence-based voting aggregates parallel LLM rollouts by weighting each with internal signals such as token log probabilities, and has been actively studied for single-turn reasoning. However, modern LLMs increasingly act as multi-turn search agents that retrieve and condition on external documents. In this paper, we show that confidence-based voting transfers poorly to this multi-turn setting, and identify the underlying failure reason as copy inflation: when retrieved documents are appended to an agent's context, tokens copied from those documents receive systematically inflated log probabilities. This flattens confidence scores within each question and weakens the resulting weighted vote. To address this issue, we propose Retrieval-Grounded Voting (RGV), which scores each rollout by the lexical overlap between its final answer and the documents it retrieved. By computing the signal outside the contaminated context, RGV sidesteps both token log probabilities and additional LLM calls. Across four search-agent benchmarks and five LLMs, RGV consistently outperforms confidence-based voting, with gains of up to +5.4% accuracy and +35% on minority-correct questions, where the correct answer appears in only 1-2 of 8 rollouts.

论文arXiv AI 12:00

Semantic Overlays: Mitigating Prompt Injection with Annotations Beyond Tokens and Steering Vectors

arXiv:2608.23873v2 Announce Type: replace Abstract: Everything a language model sees is tokens. The serving stack knows what each span is -- user input, tool output, instructions -- but the model must keep track of that itself, and can lose track or be confused: text can be written to read like anything. Prompt injection is a natural exploit of this phenomenon. By scrambling the model's understanding of span identity, an attacker can induce unwanted and dangerous actions. Adding a non-textual channel to the model's input -- a way to communicate span identity beyond text -- mitigates this class of attack. We thus introduce a general steering technique called Semantic Overlays: small learned adapters applied at chosen prefill positions to a frozen model's residual stream. Laying an overlay over a span creates an out-of-band annotation channel that cannot be replicated by tokens. Unlike steering vectors, Semantic Overlays are trained, adaptable, and selectively applied. An overlay can encode complex semantics that reshape how the model perceives the marked span: asked to copy a code snippet under an overlay asserting a different programming language, the model rewrites the snippet in the asserted language. Overlays compose, allow transparent reading of underlying content, and can carry complex payloads -- including imperatives the model will follow. An overlay which marks a span as "non-executable" defends against the broad class of prompt injections that add instructions in untrusted context. We report strong results on five prompt injection benchmarks: SEP separation rises from 24.3% to 99.0% with utility unchanged (our scoring rule; we correct a defect in the published grader), TensorTrust attack success falls from 34.8% to 6.2%, AlpacaFarm from 99.0% to 0%, and the overlay beats every published PIArena defense that leaves the model able to answer -- while marked spans stay readable, all at >95% character similarity to the original.

论文arXiv AI 12:00

Does Rank Still Matter? Position Bias When AI Agents Shop on Our Behalf

arXiv:2608.22697v2 Announce Type: replace Abstract: Search rankings are valuable because human attention is scarce and sequential. Higher-placed alternatives are easier to find, so they are examined and bought more often. Consumers are now delegating search to AI agents that can ingest an entire results page at once. Randomizing the order of one hundred hotel listings across 5,000 AI agent sessions, we compare four large language models against human field data. AI agents search more deeply than humans and never decline to buy. Position still predicts which listings are inspected, but weakly and non-monotonically: the middle of a results page has the lowest probability of inspection, not the bottom. Position reaches the choice stage for some models and not others, a heterogeneity that tracks neither provider nor capability. All models nonetheless converge on the same undominated listing. For agentic search, the attributes displayed on a results page matter more than placement within it.

论文arXiv AI 12:00

STAGE: Stateful Translation to Agentic Graph Execution with Policy-Scoped Context and Deterministic Control

arXiv:2608.22538v2 Announce Type: replace Abstract: Policy-governed agents must interpret case evidence while reliably following authorized procedures. We present STAGE, an executable-graph framework that confines model judgment to policy-scoped nodes while placing procedural control in deterministic code. We evaluate STAGE on three public policy-following benchmarks and Smart Dispute, a proprietary banking benchmark. Compared with monolithic full-policy execution, STAGE improves task success and repeated-run reliability, with its largest observed gains on the deeper workflows. On $\tau^2$-bench Telecom and Smart Dispute, $\mathrm{Pass}^{3}$ improves by up to 55.0 and 65.7 percentage points, respectively. These results demonstrate the value of combining localized policy reasoning with deterministic procedural control for enterprise use.

论文arXiv AI 12:00

When Saying No Makes Better Videos: Designing Dual Gatekeeping for Pedagogically Grounded AI Content Creation

arXiv:2608.19812v2 Announce Type: replace Abstract: To prevent the adoption of aesthetically polished but pedagogically flawed AI content, we study a video authoring pipeline featuring two layers of structured refusal. The first layer empowers educators to iteratively reshape AI scripts based on multimedia learning theory, while the second employs automated metrics to flag violations in instructional coherence and narrative-visual synchronization. While neither layer is exhaustive, their synergy ensures that principled resistance--the act of deferring AI output until it meets rigorous standards--becomes a catalyst for higher quality. Evaluation combining a study with 23 educators across 3 topics and automated metrics across 7 topics drawn from established science and philosophy curricula shows that both layers independently improve the same instructional dimensions, suggesting that thoughtful resistance and generative AI are not opposites but partners.

论文arXiv AI 12:00

RTPO: Reverse-Turn Policy Optimization for Stabilizing Agentic RL Training

arXiv:2608.18682v3 Announce Type: replace Abstract: Training multi-turn agentic workflows with reinforcement learning (RL) enables large language models to perform complex reasoning, use external tools, and conduct iterative search beyond single-turn settings. Yet multi-turn RL training remains highly unstable, often causing severe performance degradation as the number of turns increases. Through theoretical analysis, we identify three tightly coupled sources of instability: rollout-training context mismatch, weak turn-level credit assignment under sparse terminal rewards, and asynchronous policy drift when short and long trajectories are optimized under different policy versions. We show that these issues share a common structural origin in flattened trajectory optimization and address them through a unified reverse-turn formulation. We propose Reverse-Turn Policy Optimization (RTPO), which organizes multi-turn rollouts as sparse reverse trees and performs turn-level policy updates in temporal reverse order, aligning each decision with its downstream continuation. RTPO enables causally consistent turn-level credit assignment and on-policy continuation to control asynchronous drift. We provide theoretical guarantees showing that RTPO eliminates context mismatch and asynchronous drift under the proposed turn-level formulation, reduces credit bias, and converges to recursive optimality. Experiments on multi-turn agentic RL benchmarks show that RTPO improves upon trajectory- and turn-level baselines by 21.50% and 10.76%, respectively, highlighting its potential to support more stable training for tool-using agents.

论文arXiv AI 12:00

When Is an Agent Evaluation Over? Outcome Finality and Cross-Unit Separation

arXiv:2608.14940v3 Announce Type: replace Abstract: Agent evaluations commonly score the state observed when a run stops and count the run as one trial. Interpreting that score as a final result from a separate trial requires outcome finality and cross-unit separation. Outcome finality requires that later events cannot change the claimed result, while cross-unit separation requires that earlier runs cannot change the relevant conditions of later ones. The endpoint establishes neither condition by itself, and the two can hold independently. Waiting for a delayed outcome may settle the label even though its state remains available to another run. Isolation may prevent carryover even though the scored outcome remains unresolved. We develop a completion argument that identifies the evidence needed for each decision. A final success or failure label is justified only when every relevant effect is resolved or bounded tightly enough to fix the outcome. Any remaining uncertainty must be reported. First, in a controlled replay with fixed agent actions, we find that endpoint and terminal labels differ for every nonzero-delay operation and that a delayed write changes the next run's score under shared state but has no such effect after namespacing or verified reset. Second, in a review of ten public protocols, we find that reset or deliberate retention is documented explicitly more often than unfinished operations or evidence for separate scoring. Finally, we propose an open-effects record for operations and resources that may remain relevant after the endpoint, their status, and their possible effects on the scored outcome or another run.

论文arXiv AI 12:00

Agentao: A Policy-Governed Runtime Harness for Embeddable Tool-Using LLM Agents

arXiv:2608.13574v2 Announce Type: replace Abstract: LLM agents increasingly operate as execution systems that invoke tools, modify local state, use persistent memory, and interact with external protocols. These capabilities make agents useful, but they also introduce risks related to over-privileged actions, weak auditability, prompt injection, tool poisoning, and uncontrolled side effects. This paper presents Agentao, a governed local-first runtime for tool-using LLM agents. Agentao separates model-generated action proposals from host-authorized execution through a layered architecture consisting of host-facing surfaces, a host contract, a runtime core, a permission-mediated tool system, and supporting subsystems for memory, replay, plugins, skills, sub-agents, and protocol integration. We describe the motivation, threat model, design goals, governance model, execution pipeline, and structured event interface of the system. Agentao does not provide formal safety guarantees; rather, it demonstrates how permissions, state, protocol boundaries, and execution traces can be made explicit runtime abstractions for building agents that are more governable, inspectable, and suitable for host-controlled local environments. The code is publicly available at https://github.com/jin-bo/agentao .

论文arXiv AI 12:00

HALT: Verification-Aware Stopping for Retrieval-Augmented Search Agents

arXiv:2608.02009v3 Announce Type: replace Abstract: Retrieval-augmented search agents answer multi-hop questions by repeatedly issuing search queries and accumulating evidence. This creates a stopping problem: after the necessary evidence has appeared, further retrieval often adds cost, latency, and distracting context rather than useful information. We frame stopping as evidence coverage rather than generator confidence, and introduce HALT, a lightweight verification-aware policy that leaves the search agent unchanged. Given expected hop claims, HALT stops only when cumulative evidence supports each required claim. Across three multi-hop QA benchmarks, HALT reduces redundant search while largely preserving exact match. We separate a deployable setting, where hop claims are generated from the question, from a diagnostic upper bound that uses gold supporting-fact annotations: generated claims give smaller but still exact-match-preserving savings, while gold claims show the larger savings available when hop targets are clean. Baseline comparisons and ablations show that this behavior is driven by claim-evidence alignment rather than generic sufficiency, fixed stop positions, or lexical overlap. Open-corpus pilots further suggest that HALT abstains when coverage cannot be reliably verified. Overall, evidence coverage provides a practical runtime control signal for improving retrieval-augmented agents without retraining or modifying the host agent.

论文arXiv AI 12:00

SEGRA: A Structured Experience Guided Reasoning Agent for Property Graph Question Answering

arXiv:2607.22713v2 Announce Type: replace Abstract: Enterprise IT support knowledge graphs capture rich relationships among cases, users, devices, symptoms, taxonomic categories, root causes, and historical resolutions. Yet querying them in Gremlin requires knowledge of graph schemas, traversal semantics, edge directionality, and property-graph-specific constraints, making them difficult for non-expert operators to use. We introduce SEGRA, an experience-guided agent for enterprise text-to-Gremlin question answering. SEGRA integrates intent routing, schema- and taxonomy-grounded query generation, multi-shot decomposition, execution-aware verification, and a curriculum-bootstrapped skill library that reuses verified query patterns. On an enterprise IT support benchmark, SEGRA achieves a $7.0\times$ higher mean judge score than backbone-only chain-of-thought prompting. Its skill library further reduces LLM calls by $20\%$ and dollar cost by $18\%$ relative to SEGRA without skills, while preserving answer quality. These results show that schema-grounded agent design and reusable execution experience improve both accuracy and efficiency for enterprise graph QA.

论文arXiv AI 12:00

Set-shifting Behavioral Test for Harnessed Agents

arXiv:2607.13396v2 Announce Type: replace Abstract: What happens to an LLM agent's tool choice when the reliable tool silently changes within an ongoing session? We borrow the notion of set-shifting from cognitive psychology to study how well agents adapt to hidden reliability shifts. Our cognitive test for LLM agents mounts libraries of redundant tools and skills, in which many tools solve the same task but differ in hidden reliability. Using a branching schedule, we shift the reliable tool group in the environment and compare it with a stable control, allowing us to isolate the effect of each shift on the agent's behavior. We conduct our study on a panel of LLMs equipped with harnesses and show that the same set of shifts results in distinct behaviors across models: some latch onto a fixed routine within a few turns, whereas others continue to vary. Less capable models often omit the reliable tool group, while frontier models keep calling it alongside the other groups. We introduce a suite of measures to quantify agent behavior after reliability shifts. While policy prompting substantially alters behavior in some tested models, our findings highlight agents' brittleness when changes occur in indirectly observable context.

论文arXiv AI 12:00

Atomic Units of X: The Compression Layer of Intelligence

arXiv:2607.12634v3 Announce Type: replace Abstract: This paper proposes a theoretical and empirical framework for understanding intelligence as a process of atomic compression and compositional reuse. It argues that scalable cognitive, biological, computational, and organisational systems reduce complexity by organising information into reusable units that can be recombined into higher-order structures. The central contribution is the Compression Calculus, a formal framework for comparing surface evidence with atomic representations and for describing how abstraction can compound across layers. The framework is evaluated on large, multi-source software corpora under an open-world concept model, showing substantial evidence-to-concept consolidation in two large projects, while a smaller third project remains below the target threshold. The analysis further identifies important boundary conditions: consolidation depends on corpus scale, evidence-unit definition, and concept identity, and the observed reduction is driven primarily by within-corpus recurrence rather than cross-source recurrence. The paper also develops a representational account of the meaning gap in contemporary generative systems, describing latent, context-dependent conceptual approximations as soft atoms that lack the persistent identity and compositional constraints of stable atomic units. This motivates an architectural view in which large language models function as dynamic fusion engines that navigate and compose persistent conceptual structures rather than serving as the sole repository of those structures.

论文arXiv AI 12:00

APeB: Benchmarking Personalization Ability of Large Language Model Agents

arXiv:2607.03162v2 Announce Type: replace Abstract: LLM-powered agents struggle with personalization when users issue raw, underspecified queries. In this setting, agents must infer latent intent, extract preferences from noisy interaction histories, and select among competing alternatives. Existing benchmarks rarely test this capability, as they often rely on user-refined queries or simplified histories. We introduce personalized product search (PPS), a testbed for agentic personalization under raw queries and diverse histories. We construct Agent Personalized Benchmark (APeB) from action logs, pairing underspecified intents with rich histories and user-viewed candidate items. Evaluating state-of-the-art LLMs with multi-step agent workflows, we find that models handle explicit queries well but struggle with early-stage queries requiring intent and preference discovery. Rubric analysis attributes this gap mainly to ineffective history use. A simple history-aware query-refinement pipeline, VQRA, yields consistent gains, highlighting the need for dedicated history-utilization modules in personalized agents.

论文arXiv AI 12:00

Flow Reasoning Models: Turning Discrete Flows Into Efficient Recurrent Reasoners

arXiv:2606.29150v2 Announce Type: replace Abstract: Structured reasoning requires making and revising interdependent decisions to reach a globally consistent solution. Existing architectures struggle with this: autoregressive models commit sequentially and cannot revise earlier decisions, while masked diffusion models often require careful decoding schemes to coordinate interdependent predictions. We introduce Flow Reasoning Models (FRMs), a novel framework for structured reasoning that adapts discrete flows with a simple recurrent refinement mechanism. By self-conditioning a flow model on its own past outputs, we turn one-shot denoising into iterative solution refinement. This lets FRMs make and revise decisions in parallel, efficiently coordinating interdependent choices across the solution. Yet conventional self-conditioning becomes unreliable at greater recurrent depth due to exposure bias between one-step training predictions and recursively generated inference states. We address this mismatch with Fixed-Point Forcing (FPF), which trains FRMs on states produced by their own inference dynamics while preserving the standard flow-matching objective. FRMs achieve solve rates of $99.5\%$, $100.0\%$, and $99.9\%$ on Sudoku-Extreme, Zebra, and Maze-Unique, respectively. On Sudoku-Extreme, FRMs achieve higher peak accuracy than the evaluated masked-diffusion and specialized reasoning baselines while remaining highly compute-efficient, matching the next-best method's $98.7\%$ peak solve rate with $44\times$ fewer inference FLOPs.

论文arXiv AI 12:00

RecourseBench: A Modular Framework for Reproducible Algorithmic Recourse Evaluation

arXiv:2606.16113v2 Announce Type: replace Abstract: Algorithmic recourse methods provide counterfactual explanations that inform individuals of the actions required to overturn an unfavorable model decision. Despite rapid methodological progress, principled comparison remains elusive; existing frameworks are often difficult to extend and lack both interoperability and systematic verification that integrated methods faithfully reproduce their originally reported claims. We introduce RecourseBench, a unified evaluation framework built around three commitments: modularity, reproducibility, and interactivity. The framework decomposes the pipeline into five fully decoupled layers---Data, Preprocessing, Model, Recourse Method, and Evaluation---governed by abstract interfaces and a dynamic registry. Every integrated method is classified into a four-tier reproducibility taxonomy based on artifact availability, followed by a systematic verification of its core empirical claims. We further provide an interactive web interface for flexible, configuration-driven exploration across datasets, model architectures, methods, and evaluations. To our knowledge, RecourseBench is the first benchmark to explicitly ground recourse evaluation in structured claim verification and mathematically rigorous reproducibility standards, all while featuring the largest collection of state-of-the-art recourse algorithms (27 in total).

论文arXiv AI 12:00

AFFORDANCE20Q: Evaluating Affordance Reasoning from Physical Properties

arXiv:2606.14240v2 Announce Type: replace Abstract: Affordance reasoning, the inference of an object's action possibilities from its physical properties (e.g., shape and material), is fundamental to human physical understanding and increasingly critical for Large Language Models (LLMs). However, existing affordance benchmarks largely expose explicit object identities in the evaluation setup, allowing models to rely on memorized object-affordance mappings rather than reasoning over physical properties. To address this gap, we introduce Affordance20Q, a novel affordance reasoning benchmark formulated as a 20-Questions game without exposing the object's identity. In each game, the model identifies a hidden object's affordance from a candidate set by asking yes/no questions about its physical properties. Affordance20Q comprises 1,009 games over 454 objects and 59 affordances, all manually filtered, refined, and annotated. We conduct comprehensive experiments with 15 state-of-the-art LLMs and find a substantial gap (~20 points) compared to human performance. A KL-based information-gain (IG) analysis further shows that models fail to ask discriminating questions as the game progresses. To close the gap, we develop KB-Anchored Rule Induction (KARI), a pipeline based on LLMs that generates affordance rules grounded in evidence from knowledge bases (KBs). KARI improves open-source LLMs by up to 15.2 points, while the limited coverage of KBs hinders further gains. We release all our code and data at https://github.com/1171-jpg/Affordance20Q.git.

论文arXiv AI 12:00

ToolSense: A Diagnostic Framework for Auditing Parametric Tool Knowledge in LLMs

arXiv:2606.12451v2 Announce Type: replace Abstract: Large language models deployed as agents over large tool catalogs face a critical tool-retrieval bottleneck. As embedding-based retrieval approaches rely on compact encoders that may under-capture specialized tool semantics, parametric tool retrieval addresses this by encoding each tool as a virtual token appended to the LLM vocabulary, fine-tuned in two stages (memorization then retrieval SFT) to use the LLM as a retriever, achieving strong performance on standard ToolBench retrieval benchmarks. Yet these benchmarks use verbose, fully-specified queries, and their evaluation applies constrained decoding that restricts outputs to valid token paths, neither reveals whether the model actually understands its tools. We introduce \textbf{ToolSense}, an open-source LLM-powered diagnostic framework that takes any tool catalog as input and automatically generates three benchmarks: a Realistic Retrieval Benchmark (RRB) with queries at three ambiguity tiers, an MCQ probing benchmark, and a QA probing benchmark. Applying ToolSense to ToolBench (~47k tools) and evaluating five parametric model training configurations reveals a knowledge-retrieval dissociation: on RRB queries, several configurations collapse by ~50-64 percentage points compared to fully-specified ToolBench benchmarks, falling below the embedding-model baseline. Additionally, despite strong retrieval performance, some models score near-random on factual probes, suggesting a knowledge-retrieval dissociation. We open-source the ToolSense framework and the ToolBench diagnostic benchmarks at https://github.com/SAP/toolsense.

论文arXiv AI 12:00

Evidence-Based Intelligent Diagnostic and Therapeutic Visualization System with Large Language Models: Multi-Turn Interaction and Multimodal Treatment Plan Generation

arXiv:2606.06869v2 Announce Type: replace Abstract: Aim: Existing AI-assisted traditional Chinese medicine diagnostic tools suffer from opaque reasoning processes, passive interaction, and limited treatment plan presentation. This study proposes a knowledge-enhanced visual diagnostic system to improve the transparency and interpretability of syndrome differentiation and treatment. Methods: The system is built upon a Neo4j knowledge graph comprising 241 syndromes, 1,263 symptoms, and 2,485 relations. It incorporates a four-stage symptom matching pipeline (exact, semantic, fuzzy, and large language model verification), an information gain-driven proactive questioning strategy optimized with genetic algorithms, and a multimodal treatment presentation integrating artificial intelligence-generated illustrations, three-dimensional meridian-acupoint models, and evidence-based literature. Results: Knowledge graph constraints reduced non-standard outputs by 32%. Case studies validated the effectiveness of the interactive workflow across patient self-assessment, clinician-assisted diagnosis, and traditional Chinese medicine education. Automated paired-comparison evaluation across 30 cases further demonstrated significant improvements in diagnostic trust (Cohen's d = 1.82, p < 0.001), reduced cognitive load (improvements in four of five dimensions), and higher credibility of evidence-based references (4.21 vs. 2.95). Conclusions: The proposed system enhances the transparency of traditional Chinese medicine diagnostic reasoning and the interpretability of treatment plans through knowledge graph-driven visualization and multimodal interaction, offering a practical solution for trustworthy artificial intelligence-assisted traditional Chinese medicine applications.

论文arXiv AI 12:00

Rethinking Vacuity for OOD Detection in Evidential Deep Learning

arXiv:2605.06382v2 Announce Type: replace Abstract: Vacuity, or Uncertainty Mass (UM), is commonly used as a metric to evaluate Out-of-Distribution (OOD) detection in Evidential Deep Learning (EDL). It generally involves dividing the number of classes ($K$) by the total strength of belief ($S$) of the model's predictions, where $S$ is derived from summing the Dirichlet parameters. As such, UM is sensitive to the cardinality of $K$. As a result, when comparing In Distribution (ID) and OOD results, it is important that $K_{\mathrm{ID}}$ and $K_{\mathrm{OOD}}$ are equal; something that is not always ensured in practice. We provide an empirical demonstration of how results for AUROC and AUPR can substantially differ when class cardinality between ID and OOD differs by 1, with AUROC differing by as much as 0.346 and AUPR by 0.634 for standard EDL, and AUROC by 0.427 and AUPR by 0.745 for IB-EDL (both from Implementation B, Llama3-8B, ARC-E). Our findings isolate an evaluation artefact: when $K$ differs between ID and OOD, AUROC/AUPR can be artificially inflated without any change in model predictions. We further discuss the evaluation of EDL over causal language models using Multiple-Choice Question-Answer (MCQA) datasets and argue for clearer definitions of ID and OOD in this context. Our primary contribution is an empirical and theoretical demonstration that vacuity-based OOD detection in EDL-fine-tuned LLMs is highly sensitive to uncontrolled differences in evaluated class cardinality.

论文arXiv AI 12:00

D3-Gym: Constructing Real-World Verifiable Environments for Data-Driven Discovery

arXiv:2604.27977v3 Announce Type: replace Abstract: Despite recent progress in language models and agents for scientific data-driven discovery, advancing their capabilities is held back by the absence of verifiable environments representing real-world scientific tasks. To fill this gap, we introduce D3-Gym, the first automatically constructed dataset with verifiable environments for scientific Data-Driven Discovery. D3-Gym comprises 565 tasks from 239 real scientific repositories across four disciplines, each with a natural language instruction, an executable environment with pre-installed dependencies, dataset previews, a reference solution, and an automatically synthesized evaluation script. Our evaluation scripts achieve 87.5% agreement with human-annotated gold standards and strong alignment in domain-specific evaluation logic. Training on trajectories sampled from D3-Gym yields consistent gains across Qwen3 models on ScienceAgentBench, boosting Qwen3-32B by 7.8 absolute points and shrinking the gap with strong proprietary models. We further illustrate, through case studies, how D3-Gym environments can serve as a testbed for studying agentic optimization loops such as Autoresearch on real scientific workflows. We open-source D3-Gym, its creation workflow, sampled trajectories, and training scripts at https://github.com/OSU-NLP-Group/D3-Gym.

论文arXiv AI 12:00

Understanding and Enforcing Weight Disentanglement in Task Arithmetic

arXiv:2604.17078v2 Announce Type: replace Abstract: Task arithmetic provides an efficient, training-free way to edit pre-trained models, yet lacks a fundamental theoretical explanation for its success. The existing concept of ``weight disentanglement" describes the ideal outcome of non-interfering task composition but does not reveal its underlying cause. Crucially, what intrinsic properties of the pre-trained model ($\theta_0$) or the task vectors ($\tau_t$) enable this disentanglement remains underexplored. In this paper, we introduce Task-Feature Specialization (TFS), a model's ability to allocate distinct internal features to different tasks, as the fundamental principle. We first prove that TFS is a sufficient condition for weight disentanglement. More importantly, we find that TFS also gives rise to an observable geometric consequence: weight vector orthogonality. This positions TFS as the common cause for both the desired functional outcome (disentanglement) and a measurable geometric property (orthogonality). This relationship provides the key insight for our method: since the abstract TFS property is intractable to enforce directly, we can instead promote weight disentanglement by shaping its concrete geometric consequence, orthogonality. Therefore, we propose OrthoReg, a simple and effective regularization method that actively enforces an internal orthogonal structure on weight updates ($\Delta W$) that constitute $\tau_t$ during fine-tuning. And we theoretically prove that OrthoReg promotes disentanglement. Extensive experiments demonstrate that OrthoReg consistently and significantly enhances the performance of various task arithmetic methods. Code is available at \href{https://github.com/RL-MIND/OrthoReg}{https://github.com/RL-MIND/OrthoReg}.

论文arXiv AI 12:00

Prompts Without Evidence: How Neuroimaging Mentions Shift Clinical Vision-Language Model Predictions

arXiv:2603.28387v3 Announce Type: replace Abstract: Trustworthy clinical AI must use real evidence and avoid relying on surface-level artifacts. We evaluate 12 open-weight vision-language models (VLMs) on two clinical neuroimaging cohorts for binary classification of affective disorders and cognitive decline. Both cohorts include structural magnetic resonance imaging (MRI) acquired under their original research protocols. Prior work does not establish the included neuroimaging inputs as reliable stand-alone diagnostic evidence for the present tasks. Nevertheless, when neuroimaging context is introduced, smaller VLMs gain up to 0.66 F1 under the evaluated augmented conditions, becoming competitive with models an order of magnitude larger. Confidence estimation shows that most of the calibration improvement for the analyzed smaller models occurs after the MRI reference is added to the prompt, before any image is supplied. Our preliminary expert case study finds that faithfulness remains low in every condition examined, with the reviewed model introducing unverified clinical details. Finally, in our single-model intervention, preference alignment suppresses MRI-referencing behavior but reduces the augmented-condition advantage, leaving the underlying issue unresolved. These results caution against reading surface metric gains as evidence of true multimodal integration, with direct implications for clinical VLM deployment.

论文arXiv AI 12:00

PAPO: Stabilizing Rubric Integration Training via Decoupled Advantage Normalization

arXiv:2603.26535v4 Announce Type: replace Abstract: We propose Process-Aware Policy Optimization (PAPO), a method that integrates process-level evaluation into Group Relative Policy Optimization (GRPO) through decoupled advantage normalization, to address two limitations of existing reward designs. Outcome reward models (ORM) evaluate only final-answer correctness, treating all correct responses identically regardless of reasoning quality, and gradually lose the advantage signal as groups become uniformly correct. Process reward models (PRM) offer richer supervision, but directly using PRM scores causes reward hacking, where models exploit verbosity to inflate scores while accuracy collapses. PAPO resolves both by composing the advantage from an outcome component Aout, derived from ORM and normalized over all responses, and a process component Aproc, derived from a rubric-based PRM and normalized exclusively among correct responses. This decoupled design ensures that Aout anchors training on correctness while Aproc differentiates reasoning quality without distorting the outcome signal. Experiments across multiple model scales and six benchmarks demonstrate that PAPO consistently outperforms ORM, reaching 51.3% vs.\ 46.3% on OlympiadBench while continuing to improve as ORM plateaus and declines.

论文arXiv AI 12:00

Describe-Then-Act: Proactive Agent Steering via Distilled Language-Action World Models

arXiv:2603.23149v2 Announce Type: replace Abstract: Deploying safety-critical agents requires anticipating the consequences of actions before they are executed. While world models offer a paradigm for this proactive foresight, current approaches relying on visual simulation incur prohibitive latencies, often exceeding several seconds per step. In this work, we challenge the assumption that visual processing is necessary for failure prevention. We show that a trained policy's latent state, combined with its planned actions, already encodes sufficient information to anticipate action outcomes, making visual simulation redundant for failure prevention. To this end, we introduce DILLO (DIstiLLed Language-ActiOn World Model), a fast steering layer that shifts the paradigm from "simulate-then-act" to "describe-then-act." DILLO is trained via cross-modal distillation, where a privileged Vision Language Model teacher annotates offline trajectories and a latent-conditioned Large Language Model student learns to predict semantic outcomes. This creates a text-only inference path, bypassing heavy visual generation entirely, achieving a 14x speedup over baselines. Experiments on MetaWorld and LIBERO demonstrate that DILLO produces high-fidelity descriptions of the next state and is able to steer the policy, improving episode success rate by up to 15 pp and 9.3 pp on average across tasks. Code is available at github.com/MaxPappa/DILLO.

论文arXiv AI 12:00

Real-Time AI Service Economy: A Framework for Agentic Computing Across the Continuum

arXiv:2603.05614v2 Announce Type: replace Abstract: Real-time AI services run across the device-edge-cloud continuum, where autonomous AI agents generate latency-sensitive workloads, orchestrate multi-stage pipelines, and compete for shared resources under governance constraints. This article shows that the structure of service-dependency graphs, modelled as DAGs of compute stages, is a primary determinant of whether decentralised, price-based resource allocation works reliably at scale. When dependency graphs are hierarchical (tree or series-parallel), prices converge to stable equilibria, optimal allocations are computed efficiently, and under appropriate mechanism design agents have no incentive to misreport their valuations within each decision epoch; when dependencies are more complex, prices oscillate and allocation quality degrades. Our anchor contribution is a hybrid architecture in which cross-domain integrators encapsulate complex sub-graphs into slices with a simpler interface, carrying a feasibility-and-DSIC guarantee and a price-stability property of the integrator's price-discovery dynamics. An ablation study across six experiments (1,590 runs, 10 seeds each), with a strategic-bidding test of incentive compatibility and a measured agentic workload, confirms that (i) topology is a first-order determinant of price stability and scalability, (ii) in the contended regime the integrator's EMA-smoothed slice posting robustly reduces agent-facing price volatility (median ~89%) and mitigates governance-induced volatility, (iii) governance constraints create quantifiable efficiency-compliance trade-offs depending on topology and load, and (iv) under truthful bidding the market matches a centralised value-greedy baseline, adding modest welfare under contention. Systems whose pipelines form hierarchical DAGs can thus achieve centralised-quality coordination through decentralised pricing without a single controlling authority.

论文arXiv AI 12:00

Multimodal Collaborative Debate for Zero-Shot Time Series Reasoning

arXiv:2601.19151v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used as natural-language interfaces to structured data, yet they remain brittle when reasoning over time series. Visual patterns can be misleading, numerical claims can be hallucinated, and textual context can override evidence from the signal. We study zero-shot time-series reasoning as a multimodal evidence arbitration problem for LLM agents. We propose TS-Debate, an inference-time multi-agent protocol that requires no task-specific fine-tuning. TS-Debate first elicits relevant domain knowledge, then assigns modality-specialized agents to textual context, visual patterns, and numerical signals, and coordinates their interaction through a verification-conflict-calibration procedure. Reviewer agents check decision-critical claims with lightweight code execution and numerical lookup, resolve cross-modal disagreement, and calibrate the final answer. Unlike generic multi-agent debate or unconstrained tool use, TS-Debate specifies how evidence is exposed, which claims are checkable, and how verification outcomes shape synthesis. Across 20 tasks from three public benchmarks, TS-Debate improves classification and question answering performance over strong baselines, while revealing that debate is most useful for global-structure and cross-view reasoning rather than local value reconstruction.

论文arXiv AI 12:00

BioPIE: A Biomedical Protocol Information Extraction Dataset for Experiment Understanding

arXiv:2601.04524v3 Announce Type: replace Abstract: Understanding biomedical experiments provides a foundation for downstream tasks, e.g., laboratory automation, and facilitates effective cross-disciplinary communication. Two challenges, High Information Density (HID) and Multi-Step Reasoning (MSR), pose unique difficulties for precise automatic experimental understanding. Extracting structured knowledge, e.g., Knowledge Graphs (KGs), is an effective approach to address the HID and MSR. However, existing biomedical datasets for structured knowledge Information Extraction (IE) are limited to a general or coarse-grained level, hindering fine-grained experimental understanding. To address this gap, we introduce Biomedical Protocol Information Extraction Dataset (BioPIE), a dataset providing procedure-centric KGs that captures entities, actions, and relations at a scale sufficient for reasoning across biomedical protocols. We evaluate both supervised and LLM-based IE methods on BioPIE to verify its effectiveness, and implement a biomedical question answering system to provide a quantitative illustration of BioPIE's effectiveness for downstream understanding tasks. The experimental results demonstrate improved understanding performance on both the HID and MSR question sets.

论文arXiv AI 12:00

Doc-CoB: Enhancing Document Understanding with Visual Chain-of-Boxes Reasoning

arXiv:2505.18603v3 Announce Type: replace Abstract: Document understanding aims to perform question answering and information extraction over document images, where the visual content is highly information-dense and most queries rely on only a few relevant layout regions. However, existing methods either adopt a one-pass strategy that implicitly assumes all layouts are equally important, or focus excessively on small regions at the cost of losing critical layout information. To address these limitations, we introduce Doc-CoB (Chain-of-Boxes), a simple-yet-effective framework that integrates coarse-to-fine layout-aware visual reasoning into multimodal large language models. Instead of directly zooming into small regions, Doc-CoB progressively focuses on query-relevant layouts while preserving global document information. Specifically, it first selects key layout boxes and then focuses on them for further understanding with visual prompting. To support this paradigm, we introduce two reasoning tasks for box recognition and box reasoning, with an automatic pipeline that constructs 249k training samples with intermediate visual supervision. Experiments on seven benchmarks with four popular models show that Doc-CoB significantly improves performance, demonstrating its effectiveness and wide applicability. The code and the data are available at https://github.com/Doc-CoB/Doc-CoB.

论文arXiv AI 12:00

Aero Hand Open: A Simulation-Ready Tendon-Driven Hand for Dexterous Manipulation Learning

arXiv:2608.28578v1 Announce Type: cross Abstract: Tendon-driven hands are anthropomorphic, and moving the actuators off the joints is what makes a hand of this capability affordable to build. Two effects produce that saving. Routing force through a cable removes the requirement that a motor fit inside the joint it drives, so smaller and cheaper motors suffice, and one motor can drive several joints through a single cable, so fewer motors are needed. They are also harder to learn on than a direct-drive hand. The underactuated transmission that produces the saving is itself difficult to represent in a simulator, and the joints one cable drives are not independently commandable. We present Aero Hand Open, a tendon-driven anthropomorphic hand that is released simulation-ready. Three things ship with it. A simulation model reproduces the cable transmission itself. An identified actuation map connects that model to the motor commands in both directions, including the three-way coupling of the thumb. A reinforcement learning package trains policies for the hand. Together they let a policy be trained entirely in simulation and run on the hand with no fine-tuning and no state estimation. We release the mechanical design, the simulation model, the identified mapping, the training environment and the deployment stack.

论文arXiv AI 12:00

Learning a Size-Weight Frontier for Synthetic-Augmented Inference

arXiv:2608.28576v1 Announce Type: cross Abstract: Synthetic data can improve statistical inference when real data are scarce, but naively treating synthetic samples as real data can introduce bias and lead to unreliable inference. We develop a general framework for synthetic-augmented inference across a population of related tasks. It characterizes synthetic augmentation by the number of synthetic observations and their weight. Central to our framework is a size-weight frontier that specifies, for each weight, the largest synthetic sample size for which all smaller sizes attain the target task-marginal coverage. We estimate this frontier from historical tasks, and establish a finite-sample coverage guarantee simultaneously for all size-weight configurations on or below the estimated frontier. In experiments using large language model responses to augment opinion survey data, our procedure achieves target coverage and substantially narrows confidence intervals.

论文arXiv AI 12:00

Blog: Survey of Optimizers

arXiv:2608.28557v1 Announce Type: cross Abstract: Neural-network optimization in 2025-2026 is no longer well described as a succession of new Adam variants. The design space has expanded from coordinates to matrices and layers, from fixed training horizons to policies over time, and from mathematical update rules to state representations that must survive sharding and low-precision computation. This survey organizes recent optimizers and training optimization methods along four largely independent axes: temporal estimation, update geometry, horizon management, and representation and systems. It connects the spectral normalization of Muon, the historical matrix statistics of Shampoo and SOAP, adaptive and hybrid matrix methods, memory-efficient optimizers, schedule-free training, small-batch corrections, and quantized optimizer states. The central empirical conclusion is deliberately non-triumphal: matrix-aware methods represent a genuine advance, but there is no context-independent replacement for AdamW. Rankings change with model scale, data-to-parameter ratio, batch size, schedule, parameter partition, tuning budget, and whether the target metric is tokens, FLOPs, wall-clock time, or memory. The practical consequence is a compositional view of optimizer design and a stricter protocol for evaluating optimizer claims.

论文arXiv AI 12:00

Video Generative Models as Geometry Learner

arXiv:2608.28549v1 Announce Type: cross Abstract: Recent generative approaches to geometry estimation adapt pretrained image diffusion models and treat the task as image-conditioned generation. Leveraging off-the-shelf image diffusion models, they either (i) train task-specific geometry models (for depth and surface normal estimation) independently, losing the opportunity of exploring the intrinsic correlation of these geometric targets, or (ii) jointly fine-tune modified image diffusion backbones (e.g., altered self-attention), which typically demands substantial labeled data. To overcome these limitations in a principled fashion, we repurpose pretrained video generative models as a unified and data-efficient framework for geometry estimation, formulated innovatively as a next-frames prediction task. Our method, GeoNeXt, inherits naturally structured knowledge and richer priors from the video model, while further adapting them for joint modeling of images and geometry targets (image geometry), enabling more data efficient and effective learning of geometry. Extensive experiments validate our method for zero-shot monocular depth and surface normal estimation across diverse datasets, outperforming both previous task-specific and unified generative competitors while using substantially less training data. Notably, our method rivals discriminative state-of-the-art approaches trained on over 100x more data and even standouts on several benchmarks.

论文arXiv AI 12:00

An Enclosed Mode Is a Gauge Choice: Topology Relative to Reach in Certified Code World Models

arXiv:2608.28541v1 Announce Type: cross Abstract: A code world model accepted by a sampling gate can be exactly right on everything the gate can see and arbitrarily wrong beyond it. We characterize what a certified model can know, and what its errors can cost, when the omission is an annular freeze mode enclosing an unreachable interior. The gate quotient makes the question precise: acceptance-with-certainty determines the model exactly on the reachable query set; beyond reach is gauge. On a minimal ring instrument we prove the extreme case (a wrong-topology filled-disc artifact unfalsifiable by any sampling gate and bitwise harmless at play) and measure, with LLM synthesis across three model families, how one knob (a channel of width gamma) walks the same artifact through three regimes: unfalsifiable-and-harmless, falsifiable-and-costly, and instantly falsified. Three principles organize the empirics. First, danger is topology relative to reach: a channel the planner can use collapses the blind model's exploitation (play cost 1.09 to ~0 over a knee at gamma ~ 0.1), while a hidden channel with the same first Betti number keeps it at full strength (1.12). Second, repair is parameter-bound and sensor-bound: no family recovers the region from outside evidence; from inside, models pose the right topology but cannot pin its parameters, and the posed topology tracks the guiding persistent-homology summary's wrong beta_1 (a sensor with a measured geometric resolution limit), not the truth. Third, mitigation must match the error's dimension and direction: point fences fail against the one-dimensional boundary, a dimension-matched persisted fence collapses exploitation to a two-lesson transient (0.999 to 0.058), and the dual freedom certificate collapses the invented-mode failure symmetrically (1.769 to 0.029). In n dimensions the shell makes misidentification near-certain while the danger stays fully exploitable: the two axes are independent.

论文arXiv AI 12:00

Texture Image Classification Using DWT AlexNet Feature Fusion and Deep Neural Networks

arXiv:2608.28524v1 Announce Type: cross Abstract: Texture image classification plays a significant role in computer vision applications, including industrial inspection, medical image analysis, remote sensing, and object recognition. Handcrafted features can capture local texture characteristics but may have limited capability to represent complex visual patterns. In contrast, deep learning models automatically learn discriminative representations but may not fully exploit the multiscale spatial-frequency information inherent in texture images. This paper proposes a hybrid feature fusion framework, termed DWT_AlexNet_DNN, which combines Discrete Wavelet Transform (DWT) features with deep features extracted using AlexNet for texture image classification.

论文arXiv AI 12:00

Conformal Uncertainty Quantification Guarantees for Neural Operators

arXiv:2608.28515v1 Announce Type: cross Abstract: Neural operators provide fast surrogate models for approximating operators between function spaces, but their predictions often lack uncertainty quantification. We develop a split conformal framework to guarantee that a calibrated pointwise band around the neural operator output contains the true solution on at least a $1-\gamma$ fraction of the evaluation domain, with probability at least $1-\alpha$ over test and calibration inputs, where $\alpha,\gamma\in(0,1)$. Our method reduces a normalized residual field to its spatial $(1-\gamma)$-quantile and computes a scaling factor using a held-out calibration dataset. We prove marginal coverage guarantees for measurable residual fields defined on arbitrary probability spaces, covering both continuum domains and fixed discretizations. Under mild assumptions on the data distribution, we show that the coverage conditional on the calibration set follows a Beta distribution, which we verify with numerical experiments on Darcy flow and Navier--Stokes equations, where our calibration yields bands consistently tighter than existing corrections while retaining the target coverage.

论文arXiv AI 12:00

On the Maintenance and Co-evolution of Agent Plugins: An Empirical Study of Claude Code Plugin Marketplaces

arXiv:2608.28497v1 Announce Type: cross Abstract: AI coding agents, software tools that automate development tasks through reasoning and tool use, are increasingly extended through plugin marketplaces, yet the structure, maintenance, and co-evolution dynamics of these emerging repositories remain empirically unexplored. Unlike traditional software packages that deliver functionality through source code, agent plugins deliver functionality through a combination of natural-language instruction files, scripts, and configuration files, raising the question of whether these plugins are maintained artifacts that co-evolve across components, or one-off artifacts that developers write once and do not need to revisit. To study the maintenance and co-evolution of agent plugins, we conduct an empirical study of 1,926 repositories hosting Claude Code plugin marketplaces, analyzing 8,351 plugins and 77,773 commits across 2,018 marketplaces. We find that the marketplace is expanding rapidly, plugin-touching commit activity growing 8.8x over six months after the October 2025 launch, and plugins targeting Software Engineering tasks accounting for 61.3% of all plugins. Plugin development is predominantly feature-driven, with feature commits occurring at more than twice the rate of conventional open-source software (OSS) (39.6% vs. 17.2%). Claude co-authors 34.9% of all commits, and four commit types (docs, perf, style, and refactor) carry substantially different meanings in plugin repositories than in traditional software. Most component types evolve independently, but within skills directories, natural-language instruction files and implementation scripts co-evolve at above-chance rates, with 78% of co-changes being functionally coupled, representing a new class of maintenance dependency not observed in traditional software engineering.

论文arXiv AI 12:00

LLM-Based Agents for Software and Systems Security: Approaches, Applications, and Assessment

arXiv:2608.28490v1 Announce Type: cross Abstract: Software and systems security workflows are typically procedural: analysts inspect heterogeneous artifacts, form hypotheses, invoke tools, interpret outputs, and revise plans. Large language model (LLM)-based agents, which can plan, use tools, retain state, and revise actions across multi-step workflows, are being rapidly adopted to automate this work. Given the consequences of delegating security decisions to autonomous systems, understanding how such agents are built, used, and assessed is crucial. Yet to this date, there remains a lack of systematic understanding of what has been done and how far we are in this field: the term "agent" is applied inconsistently, applications differ sharply in risk, and assessment protocols are often incomparable. To gain a comprehensive and coherent view of this area hence inform relevant future research, this paper provides a systematic literature review of the (1) technical approaches, including agent architecture, perception, memory, reasoning and planning, action space, orchestration, and self-improvement, (2) applications, with respect to the security tasks served, and (3) assessment, including the datasets, outcome and trajectory metrics, safety measures, and baselines considered, over the peer-reviewed literature spanning the emergence of this area (2023--2026). Our synthesis reveals a field that has built agents able to act but not yet agents whose authority is bounded or whose behavior is auditable. In addition to knowledge systematization, we also extend our insights into the limitations of and challenges faced by current approach, application, and assessment designs, which shed light on potentially promising future research directions.

论文arXiv AI 12:00

How Proper Scoring Rules Shape LLM Forecasting

arXiv:2608.28482v1 Announce Type: cross Abstract: This paper evaluates how reward function choice shapes the performance and behavior of LLM forecasters. We compare five proper scoring rules as training objectives for binary forecasts of resolved real-world events. Although the rules share the same theoretical incentive for truthful probability reporting, the resulting models differ in calibration, probability use, and estimated profiles of bias, information, and noise, with smaller differences in aggregate accuracy and discrimination. The Brier-trained model has the lowest observed Brier score and highest AUC-ROC, while the log-trained model has the highest observed log score and lowest calibration error. Models with similar aggregate performance also reach that performance through different combinations of bias, information, and noise. Proper scoring rules therefore need not behave interchangeably as training objectives. Reward choice may shape not only how well an LLM forecasts, but how its forecasting errors are structured. Each condition uses a single seed, so some differences may reflect training stochasticity.

论文arXiv AI 12:00

NL2AGBench: Benchmarking LLM Auto-Formalization for AlphaGeometry

arXiv:2608.28481v1 Announce Type: cross Abstract: Recent advances in large language models (LLMs) have demonstrated strong capabilities in natural language understanding and mathematical reasoning. However, their ability to translate informal mathematical problems into formal representations remains underexplored. This limitation is particularly important for neuro-symbolic geometry systems such as AlphaGeometry, whose theorem-proving engine requires inputs in a specialized domain-specific language (DSL). Although AlphaGeometry achieves near-IMO gold-medalist performance, manually converting natural-language problems into its formal syntax remains a significant usability bottleneck. To address this challenge, we introduce the Natural Language to AlphaGeometry Benchmark (NL2AGBench), which evaluates LLMs in translating English geometry problems into AlphaGeometry-compatible formal representations. NL2AGBench uses execution-based verification within AlphaGeometry to assess translation quality rather than relying solely on textual similarity. We evaluate ten state-of-the-art open- and closed-source LLMs across multiple parameter scales and analyze executable translation accuracy, syntactic correctness, and error characteristics. Our experiments reveal a substantial performance gap between closed- and open-source models: leading closed-source models achieve executable translation rates above 80%, while even the largest open-source models struggle to consistently preserve geometric constraints and produce valid formalizations. We introduce an error taxonomy distinguishing syntax and logic errors and investigate mitigation strategies, including few-shot prompting, fine-tuning, and human-guided hinting, which yield measurable improvements across multiple model families.

论文arXiv AI 12:00

Real-time virtual circuits for plasma shape control via neural network emulators: experimental demonstration on MAST Upgrade

arXiv:2608.28468v1 Announce Type: cross Abstract: Conventional plasma shape control in tokamaks relies on virtual circuits (VCs) that are computed offline from linearisations around a small, tailored number of reference equilibria, and deployed as expertly prepared schedules during the discharge. Here, we report on the first experimental deployment of real-time VCs. We replace pre-set look up tables with VCs updated in real time using surrogates of the plasma response. Both the existing control architecture and the interpretability of VC-based control are retained. Previous work showed that neural network emulators can produce accurate VCs, and validated their performance in closed-loop shape control simulations. Here, we report their first experimental validation on MAST Upgrade (MAST-U). Dedicated experiments spanning different scenarios, including prescribed shape perturbations, feedback-driven divertor-leg motion, and strongly evolving plasma configurations, show that real-time VCs can realise plasma shape control tasks within the MAST-U plasma control system. These results establish the experimental feasibility of real-time linearisations as a practical extension of conventional plasma shape control in tokamaks. The present implementation demonstrates a central step towards a simpler control workflow, in which manually constructed, phased VC schedules are replaced by VCs generated automatically online from a trained surrogate model, without scenario-specific retraining.

论文arXiv AI 12:00

Anatomy-Aware Promptable Segmentation with Online Interactive Training for AUTOPET V

arXiv:2608.28461v1 Announce Type: cross Abstract: We present an anatomy-aware, promptable model for whole-body lesion segmentation in FDG and PSMA PET/CT, developed for the AUTOPET V challenge. The proposed method is built as family of nnU-Net-based models and trained in two stages: i) a pre-training stage that produces a strong initial segmentation, and ii) an online interactive stage that learns to exploit scribble prompts, refining the prediction over successive interactions. Anatomical context is incorporated through organ supervision using a single shared head that predicts lesions and organs from the same features, which reduces false positives arising from physiological uptake. Also as the tracer (i.e., FDG/PSMA) is not provided at inference, we add a tracer classifier based on image processing and a random forest over coronal MIP features, routing each study to a combined FDG+PSMA model or to a PSMA-specific model. Across four-fold cross-validation the organ-supervised model achieves the best and most stable performance, the interactive stage improves the Dice score monotonically with each prompt, and PSMA-specific training yields the strongest tracer-wise results.

论文arXiv AI 12:00

ARC-CT: Anatomy-Routed Contrastive Vision-Language Learning for 3D Chest CT

arXiv:2608.28455v1 Announce Type: cross Abstract: Contrastive vision-language learning uses paired chest CT volumes and radiology reports to learn abnormality classifiers without manually annotated labels. However, two characteristics of chest CT challenge conventional global contrastive learning. First, many critical abnormalities are small or anatomically localized, and pooling an en- tire volume into a single embedding may dilute their visual evidence. Second, the standard contrastive objective treats every other scan in a batch as a negative. Because many chest CTs share abnormalities, this objective incorrectly pushes co-positive pairs apart. We propose Anatomy-Routed Contrastive Learning for 3D Chest CT (ARC-CT), a region-aware framework that addresses these limitations using only la- bels extracted from reports by an LLM, with no manual annotations or bounding boxes. ARC-CT combines three components: (1) an Anato- myQFormer localizing evidence via queries constrained by automatically generated organ masks; (2) a label-Jaccard soft InfoNCE objective in- tegrating the standard one-hot target with the label-set overlap of each pair, which reduces false-negative penalties between studies that share clinical findings; and (3) an organ-level alignment loss connecting mask- pooled visual features to organ-specific report text extracted offline with a large language model. ARC-CT achieves a 0.86 mask-free macro AUC across 18 abnormalities using a compact 3D ResNet-18 backbone. Over- all, ARC-CT outperforms both comparable efficient baselines and sev- eral larger transformer models. Our code and weights are available at https://github.com/arc-ct/arc-ct.

论文arXiv AI 12:00

Fidelity Is Not Enough: Dispatch-Level Instrumentation for Agentic Datasheet Extraction

arXiv:2608.28439v1 Announce Type: cross Abstract: One model passed our fidelity check without ever opening the datasheet. We found it while qualifying models for an internal extraction service: a structured-output constraint had silently disabled tool use, and the model answered anyway, with fabricated source text. Only the per-tool trace exposed it. Fidelity -- whether an extracted value matches the source -- is the standard measure for agentic document extraction, and it scores that run a success. We therefore log every tool call in an agentic benchmark of 25 hand-curated claims over three components, with 12 more on a fourth, 37 in all. From that dispatch record we build two instruments: a rule-based failure-attribution classifier, and a silent-failure detector whose two rules check only which tools were called, never the extracted value. The detector raises no flag on 207 clean fidelity-passing extractions across three model families, and recovers all 50 planted faults that withhold exactly the tools its rules check. The two results are not symmetric: the first bounds the false-positive rate, the second is recall by construction, and detection power against runs that call their tools and still answer wrongly is unmeasured. A second, independent oracle, a causal chamber that tests whether the datasheet's claims hold under physical measurement, is intentionally partial: it confirms only what the apparatus can exercise, a verifiable envelope of 2 of those 37 claims, and we give a taxonomy of why the rest are not physically gradable. Under a controlled perturbation, fidelity passes throughout while the chamber verdict flips exactly at the measurement uncertainty. Across three deployed model stacks (one destabilised by its serving stack, not by any capability gap) the tool layer buys portability and observability rather than accuracy, and earns its premium only once a document outgrows the context window.

论文arXiv AI 12:00

Are These Modules Worth Their Cost? A Paradigm-Level Accuracy-Cost Analysis of In-context Learning Text-to-SQL

arXiv:2608.28432v1 Announce Type: cross Abstract: Recent advances in in-context learning (ICL) text-to-SQL have substantially improved execution accuracy on public benchmarks by assembling increasingly elaborate pipelines around the base generator, yet existing studies typically report aggregate end-to-end accuracy, without quantifying the marginal accuracy-cost contribution of individual design choices. Consequently, providing a unified, paradigm-level cost-accuracy quantification remains a critical challenge for understanding and configuring modern text-to-SQL. To address this, we instantiate 17 paradigm-level configurations across five recurring modules of the ICL text-to-SQL pipeline under a single controlled implementation, and attribute each paradigm's marginal contribution and incurred cost across all four backbones spanning diverse capability levels and reasoning styles. Our analysis reveals that execution-feedback refinement is the only paradigm whose benefit holds universally at consistently low cost, while most other modules help only under backbone-dependent conditions. Token accounting shows that input demand is more closely tied to pipeline structure, whereas output demand is more sensitive to backbone generation behavior. Cross-module analysis further shows that stacking improves accuracy on most backbones, although how the gains compose varies with backbone capability. We also find that a fixed budget is often better spent engineering a more elaborate pipeline over a mid-tier backbone than upgrading to a frontier model with a lean pipeline. These findings distill into an actionable, cost-aware tiered guideline that transfers to five additional backbones without per-paradigm search.

论文arXiv AI 12:00

LongPIBench: A Long-Context Benchmark for Prompt Injection

arXiv:2608.28411v1 Announce Type: cross Abstract: Prompt injection attacks pose a serious security risk to large language models in real-world applications. However, existing prompt injection benchmarks primarily focus on short-context inputs, leaving the attacks and defenses in long-context settings largely unexplored. This gap leads to a substantial overestimation of the effectiveness of current defenses. In this paper, we bridge the gap by introducing LongPIBench, a long-context benchmark for prompt injection covering 4 realistic application scenarios: paper peer review, resume screening, code review, and email summary. For each scenario, we construct a synthetic dataset and a real-world dataset, with context lengths ranging from thousands to tens of thousands of tokens. The evaluation results on LongPIBench reveal significant vulnerabilities of prompt injection defenses under long-context settings: even simple heuristic prompt injection attacks achieve high success rates and frequently bypass state-of-the-art defenses. We hope LongPIBench can serve as a practical benchmark for systematically evaluating prompt injection defenses in realistic long-context scenarios.

论文arXiv AI 12:00

When Linguistic and Internal Confidence Diverge in Large Language Models

arXiv:2608.28382v1 Announce Type: cross Abstract: Users often ask large language models (LLMs) to report how confident they are, but it is unclear whether such linguistic confidence tracks the model's internal confidence. We study this question across 8 classification tasks, 2 generation tasks and 30 models from three families. For classification, we compare linguistic confidence with logits-based confidence along three axes: association, magnitude agreement and calibration. For generation, we test whether linguistic confidence tracks semantic-entropy-based uncertainty. The axes frequently diverge. Instance-level association is weak on average, although it improves on easier items and for stronger base models. Instruction-tuned models often report higher confidence and sometimes show higher association, but they also have larger confidence gaps and worse calibration. Prompt design mostly changes the distribution of reported confidence. Attitude cues inflate confidence without improving alignment, while score exemplars can preserve rank-order signal when they avoid collapsed confidence values. Regression analyses show that distributional properties of confidence scores explain much of the observed alignment pattern, with model metadata playing a smaller role after controls. These results support a lossy-channel view of linguistic confidence. A more dispersed verbal confidence distribution can carry useful rank information, but it does not make the scores calibrated. Linguistic confidence should therefore be evaluated with multi-axis diagnostics before being used in downstream reliability pipelines.

论文arXiv AI 12:00

AI as Teammate: Rethinking Task Distribution in Medical Training

arXiv:2608.28373v1 Announce Type: cross Abstract: Integrating Artificial Intelligence (AI), particularly generative AI, into medical training has prompted concerns about learner over-reliance, misuse, and erosion of foundational clinical competencies. We propose a conceptual reframing at the decision level: the problem is not misuse but misclassification - a mechanistic failure of real-time metacognitive evaluation in selecting a subzone-inappropriate AI interaction mode. Drawing on "SCAN" (Substitute, Complement, Aid, Non-Negotiable), a human-centric decision-making framework for generative AI task allocation grounded in Vygotsky's Zone of Proximal Development and metacognition, we advance the emerging social-constructivist conversation around AI in medical education by offering a testable account of AI's role in clinical reasoning development. This framework yields testable predictions for how misclassification can be detected, mitigated, and, more importantly, prevented in the clinical learning environment. Regarding clinical reasoning development, we show how trajectories of skill acquisition (upskilling) and failure (the triad of skill failure: de-skilling, never-skilling, and mis-skilling) operate at the individual task level in ways that fixed-phase, cohort-wide treatments fail to capture. We further identify passive engagement within correctly classified AI-scaffolded tasks as a particularly insidious, detection-resistant pathway to mis-skilling - one requiring subzone re-identification from AI assistance to expert assistance, with human experts serving as epistemic auditors. The paper operationalizes SCAN for clinical curriculum design, supervision, and assessment, and opens an empirical research agenda grounded in cognitive science. This paradigm shift from misuse to misclassification is not semantic: it offers educators a clear perspective on what to look for, what to assess, and what to intervene on.

论文arXiv AI 12:00

Real-Time Musculoskeletal Surrogates for Pediatric Cerebral Palsy: a Credibility Pilot

arXiv:2608.28371v1 Announce Type: cross Abstract: Real-time musculoskeletal (MSK) surrogates could support personalized rehabilitation for children with cerebral palsy (CP), but their credibility depends on subject-wise evaluation, low inference latency, and calibrated uncertainty. We develop a subject-conditioned causal neural surrogate using OpenSim-derived static parameters, temporal joint kinematics, true muscle capacities, and training-only perturbations. On a real pediatric CP gait dataset comprising nine children, we use leave-one-subject-out validation on six development subjects and evaluate a frozen configuration once on three locked test subjects. The surrogate accurately reproduces musculotendon lengths (R-square = 0.92 in development validation and approximately 0.95 on locked subjects; nRMSE < 8%) while requiring only sub-millisecond to few-millisecond neural inference, well below a 100 ms interactive-rehabilitation target. In contrast, direct muscle-force estimation remains unstable at this small, heterogeneous scale: pooled metrics can overstate within-subject, per-muscle accuracy. A Monte Carlo credibility pilot further shows that propagating only +/-5% anthropometry and muscle-capacity variation produces severely overconfident nominal 90% intervals (approximately 4% force coverage and below 1% MT-length coverage). These results establish a leakage-free evaluation and credibility framework for pediatric MSK surrogates, while identifying force modeling and epistemic uncertainty as the central next challenges for clinically credible digital twins.

论文arXiv AI 12:00

Optimal Adversarial Testing: Extracting Honest Test Results from Dishonest Test Takers

arXiv:2608.28362v1 Announce Type: cross Abstract: In applications, it is often required to test objects or people to determine their qualities in terms of certain metrics. However, besides being naturally noisy, the test results can be corrupted by adversarial behaviors of objects or people being tested (test takers). For example, dishonest test takers can cheat in the exams to distort the test results. With the development of AI technologies, such distortions driven by cheating using AI technologies are becoming more commonplace and severe. In this paper, we propose optimal testing strategies which can still recover needed test results even if there are cheaters polluting the results. The proposed testing strategies will optimally re-test selected group of test takers using different testing security measures. We determine the optimal testing strategies using a dynamic programming method.

论文arXiv AI 12:00

Cross-Spectral Dense Correspondence for Multimodal Spectral Medical Imaging

arXiv:2608.28341v1 Announce Type: cross Abstract: Precise dense correspondence is a fundamental prerequisite for multimodal spectral imaging systems that fuse disparate wavelength ranges for subsequent analysis in medical and scientific imaging. Corresponding image points are often observed with non-overlapping spectral sensitivities, leading to wavelength-dependent contrast changes, intensity inversions, and appearance shifts for which dense ground truth is difficult to obtain and conventional RGB-based training data provides only limited supervision. We address this data gap by introducing a sensor-agnostic cross-spectral modulation protocol on established correspondence benchmarks with intensity input projection, and by proposing a synthetic cross-spectral correspondence benchmark simulating physically plausible radiometric differences. Evaluation on several modern dense correspondence backbones trained with our unified cross-spectral protocol showed substantial improvements under severe spectral mismatch while maintaining performance on standard RGB benchmarks. Ablation experiments show that view-dependent channel selection and nonlinear radiometric transformations provide complementary robustness, indicating that the primary limitation of existing models is not their structural matching capacity but the mismatch between training distribution and spectral characteristics of the target image pair. Qualitative evaluations on heterogeneous medical spectral acquisition systems demonstrate the practical relevance of the proposed training data augmentation protocol as an enabler for spatially coherent spectral fusion in HSI workflows.

论文arXiv AI 12:00

BanglaMed-QA: A Question Answering System for Healthcare Support in Bangla

arXiv:2608.28329v1 Announce Type: cross Abstract: Medical question answering (QA) systems have become crucial tools for providing reliable health information. But they remain very unexplored for low-resource languages like Bangla due to limited datasets and systems tailored to these languages. To address this, we introduce BanglaMed-QA, a robust QA system specifically designed for the Bangla medical domain. The process begins with building a structured medical knowledge base that includes 4,493 QA pairs in 9 categories under 506 diseases. To improve semantic comprehension, domain-specific root word dictionaries and synonym sets are proposed, in addition to part-of-speech tagging for anaphora resolution. We adopt supervised machine learning models in which SVM is found to be the best model to categorize questions. Multiple similarity metrics, including cosine, Jaccard, BM25, and Levenshtein, are applied with soft and hard voting methods for query matching. The performance of the QA system has been evaluated in two aspects, with a 95% F1 score in an automated evaluation and an average human satisfaction rating of 0.9 out of 1.0. This validates the real-world application of BanglaMed-QA in closing the healthcare information gap for Bangla speakers.

论文arXiv AI 12:00

Layered LLM Defenses as an Ensemble: Access Tiers, Inference Cost, and the Measured Failure Correlation Between Defense Layers

arXiv:2608.28327v1 Announce Type: cross Abstract: Practitioners defend large language models (LLMs) by stacking defenses, assuming the layers compound. A stack is an ensemble, and ensembles compound only under a condition the LLM security literature recommends but never measures: the members must fail on different inputs. Two instruments make that measurable. The Adversary Access-Tier Model (AATM) grades an adversary by the access it holds, from system-only (A0) to influence over training data (A4). A cost model sorts defenses into five classes of inference-time overhead; because two classes require training weights or reading activations, they tier the defender as AATM tiers the adversary. From these we derive how a stack behaves, and the quantities a defender cares about diverge: coverage saturates within a tier, cost rises by class, false refusals accumulate as a union, and residual attack success falls multiplicatively only under independence. We measure that independence. Running one adaptive adversary against a seven-layer stack, failure correlation is positive in all fifteen measurable pairs ($\phi$ from $0.30$ to $0.75$), and the joint residual exceeds the multiplicative prediction by up to $0.172$. Stratifying on behavior difficulty dissolves most of the association, so the dependence is predominantly common-cause, but it survives permutation inference, majority-vote grader labels, and externally calibrated thresholds. The same stack refuses four in five benign prompts while remaining statistically indistinguishable from its strongest single layer. The dependence is architectural rather than sampling-based: members correlate through the model they all wrap, so no wider member pool weakens it. Diversity therefore selects stack members but does not predict what an assembled stack delivers, which has to be measured end to end.

论文arXiv AI 12:00

Deriving Scaling Laws for OpenEuroLLM Models: Learning Rate, Batch Size and Loss

arXiv:2608.28308v2 Announce Type: cross Abstract: We study the scaling behavior of learning rate and batch size in pretraining dense large language models on English-prevalent corpora. Beyond scaling jointly optimal learning rates and batch sizes, we investigate their marginal evolution with model capacity and data scale and develop a model that captures these relationships. As we employ a Warmup-Stable-Decay learning rate schedule, we further investigate the gains from learning rate annealing over a broad range of hyperparameters settings, models and data budgets, and whether the optimal learning rate and batch size transfer between the stable and decay phases. Finally, we characterize the dependence of loss on model capacity and dataset size, evaluating recently proposed scaling forms that explicitly model their interaction. We find these approaches particularly effective at capturing both undertraining and overtraining regimes across our experiments. This study establishes a first baseline and scaling procedure for the development of future OpenEuroLLM models. We open-source the complete collection of pretraining runs used in this study.

论文arXiv AI 12:00

VISTA: Verifier-Informed Student-to-Teacher Adaptation for On-Policy Self-Distillation

arXiv:2608.28306v1 Announce Type: cross Abstract: On-policy self-distillation (OPSD) improves reasoning by training a problem-only student on its own rollouts using dense token-level supervision from a privileged teacher that also sees a reference solution. However, standard OPSD treats the teacher distribution as a fixed target along the student's rollout and updates only the student %, although -- even though privileged conditioning does not guarantee that the teacher always provides the most appropriate target for problem-only reasoning. This one-way supervision can therefore misdirect the student when the teacher distribution is misaligned with valid student reasoning. We therefore introduce Verifier-Informed Student-to-Teacher Adaptation (VISTA), which preserves the standard OPSD student update while using outcome-verified rollouts to adapt the teacher toward the student distribution. Within each verified rollout, VISTA further restricts this adaptation to the top-$k$ positions with the largest teacher--student KL divergence. Notably, VISTA reuses the rollout and loss function from standard OPSD, introducing no additional sampling or separate reward objective. Across AIME24, AIME25, and HMMT25 with Qwen3 models at 1.7B, 4B, and 8B, VISTA achieves the highest Avg@12 at every scale, improving over OPSD by $0.6$, $0.7$, and $2.1$ points, respectively. These results demonstrate the value of student supervision from outcome-verified rollouts and highlight student-to-teacher adaptation as a promising direction for OPSD.

论文arXiv AI 12:00

PanelShield: Verifiable Closed-Loop Safe Planning for Robotic Industrial Panel Operation

arXiv:2608.28305v1 Announce Type: cross Abstract: Industrial panel operation is knowledge-intensive and safety-critical. Beyond control recognition and action generation, execution must satisfy constraints in operation manuals and safety regulations. While foundation-model-based planners show strong semantic capability, they typically lack computable, localizable, and reproducible mechanisms for violation detection and repair. To address this, we propose PanelShield, a verifiable closed-loop safety planning framework for manual-guided industrial panel operation. The framework generates parameterized action primitive sequences from task-relevant manual evidence and applies dual formal verification with LTL and a Safety FSM to enforce cross-step temporal correctness and local transition legality. When violations occur, it outputs a structured counterexample with the earliest violating step and cause, enabling targeted repair and re-verification. We build a multi-level long-horizon planning benchmark covering three representative industrial device panels, and evaluate the framework in simulation and real-world robotic experiments. Results show that PanelShield improves complex safety-constrained task performance over foundation-model-only planning baselines while reducing the violation rate to 2.7%, with 4.1 s total latency. Real-world experiments demonstrate end-toend feasibility. Overall, PanelShield offers a verifiable approach to robotic panel operation that balances flexibility, safety, and auditability.

论文arXiv AI 12:00

MaCoPlanner: LLM-Assisted Manual-Compiled Task Planning with Proactive Safety Verification for Robotic Industrial Panel Operation

arXiv:2608.28300v1 Announce Type: cross Abstract: Robotic industrial panel operation requires not only accurate control localization but also compliance with operating procedures, safety rules, and device-state constraints distributed across heterogeneous manuals. This study presents MaCoPlanner, a task-planning framework built on knowledge compiled from equipment manuals that converts equipment manuals into a typed intermediate representation, retrieves task- and state-relevant evidence, and uses it to support plan generation. Before actuation, candidate plans are symbolically rolled out and checked against procedural and state-transition constraints; detected violations are localized and returned for targeted repair, while unresolved plans are rejected. A separate execution interface grounds verified symbolic actions to physical controls and updates the device state. Under an independent evaluation oracle, MaCoPlanner achieves a final violation rate of 2.7%, and 26.3% of the runs in the repair analysis are rejected after exhausting the refinement budget. Compared with Raw-Manual, task success increases from 62.8% to 84.4% on Level-2 tasks and from 25.9% to 43.2% on Level-3 tasks. Experiments on a controller-panel simulator without an attached industrial load further demonstrate integrated execution feasibility under representative interaction conditions, without claiming industrial deployment readiness.

论文arXiv AI 12:00

A Probabilistic Interpretation of KV Cache Eviction

arXiv:2608.28293v1 Announce Type: cross Abstract: The premise and promise of KV (cache) eviction is simple: higher throughput can be achieved by evicting some entries from the KV cache, at a negligible cost to quality. This holds empirically for many existing methods, though most rely on creative heuristics for selecting which entries to drop. Despite recent advances, the problem of KV eviction has remained informal in the literature. This paper aims to properly formalize this problem through the lens of probabilistic reasoning and reveal what can be learned from this perspective. Concretely, we (1) formalize the problem of KV eviction and, unfortunately, prove that it is computationally hard, (2) show that by framing it probabilistically, KV eviction reduces to the problem of expectation estimation, which can be approximated through sampling, (3) show that through this probabilistic interpretation, correcting for evicted entries during decoding---a previously ignored problem---becomes feasible, and (4) reveal that existing methods in the literature are zero-variance biased estimators that can be easily adapted in order to enable decode time correction. In practice, we show that this probabilistic version of KV eviction coupled with decode time correction is more robust to different tasks compared to existing eviction methods and achieves competitive performance at the same compression budget.

论文arXiv AI 12:00

Embedding Models for Stance-Aware Argument Retrieval

arXiv:2608.28283v1 Announce Type: cross Abstract: In computational argumentation, obtaining arguments that explicitly support or attack given claims is a critical precursor to downstream reasoning tasks. When these supporting and attacking arguments are to be retrieved using semantic search methods, they need to be assessed for topic-relevance to the claims of interest as well as for correctness of their (positive or negative) stance towards the claims. In this paper we explore how dense embedding models (hereafter, models), powering modern retrieval pipelines, can serve as the basis of semantic search incorporating this dual assessment. We show experimentally that existing models struggle with asymmetric reasoning, exhibiting a strong bias toward topical overlap while ignoring instructional stance. We also show that correcting this bias via contrastive training triggers a new failure mode where models over-correct, over-fixating on polarity keywords (e.g., "supports" or "refutes") at the expense of the semantic topic. We thus introduce diagnostic word-ablation metrics to quantify this phenomenon and propose a data-centric solution. By implementing a balanced argument curriculum alongside LLM-augmented, stance-inverted arguments, we force the (embedding) models to learn deeper directional logic rather than exploiting superficial lexical shortcuts. Our evaluation demonstrates that, for sufficiently powerful models, this approach can alleviate the observed overcorrection, achieving further improvements in stance-aware argument retrieval.

论文arXiv AI 12:00

Spatial-Semantic Reasoning using Large Language Models for Efficient UAV Search Operations

arXiv:2608.28270v1 Announce Type: cross Abstract: We present a real-time semantic navigation framework for Unmanned Aerial Vehicles (UAVs) focused on improving time efficiency in the Object Goal Navigation (ObjectNav) task. Central to our approach is a Large Language Model (LLM) that interprets user-provided natural language instructions and performs semantic reasoning over detected objects and spatial context to prioritize high-probability search regions. The system combines real-time object detection, 3D spatial mapping, and polynomial spline interpolation for smooth and feasible UAV trajectory planning. Unlike prior methods that rely on offline reasoning or simulator-constrained action spaces, our framework can operate in real time, continuously updating semantic relevance based on new observations. Experiments in both simulated and real-world settings demonstrate reductions in mission duration while maintaining high search accuracy, underscoring the effectiveness of LLM-guided reasoning for time- efficient UAV-based ObjectNav.

论文arXiv AI 12:00

A comprehensive and trustworthy benchmark of AI methods for change detection in Earth observation

arXiv:2608.28247v1 Announce Type: cross Abstract: Change detection in Earth observation (EO) is critical for monitoring land surface transformations, yet recent research in the field is constrained by inconsistent evaluation protocols and a narrow focus on predictive accuracy without regard for computational efficiency. To address this, we present a standardized, open-source benchmark for evaluating state-of-the-art (SOTA) deep learning methods for Earth observation change detection. We conduct a comprehensive analysis of ten representative model architectures, ranging from convolutional networks (CNNs) to vision transformers (ViTs), across ten heterogeneous change detection datasets. We rigorously evaluate these models with identical experimental protocols, comparing models trained from scratch against those utilizing pre-trained weights. Furthermore, we evaluate predictive performance alongside computational efficiency, including parameter counts and inference latency. Our findings reveal that well-optimized classical architectures, such as Siamese U-Nets, frequently outperform more complex contemporary models when computational efficiency is factored in, and that pre-training consistently provides a significant performance boost with no additional inference cost. To ensure complete transparency and reproducibility, all experimental resources, including standardized data splits, training scripts, training logs, and model checkpoints are publicly available and adhere to FAIR principles (Findable, Accessible, Interoperable, and Reusable).

论文arXiv AI 12:00

Training-free Suction Grasp Detection for Deformed Aseptic Cartons Using Vision-Language Models and Geometric Surface Scoring

arXiv:2608.28246v1 Announce Type: cross Abstract: Robotic sorting of recyclable waste is challenging due to the deformable and geometrically inconsistent nature of target objects. We present a training-free suction grasping system for sorting deformed aseptic beverage cartons, decoupling target identification from grasp-point selection. An open-vocabulary vision-language model detects cartons from a text prompt, SAM2 refines each detection into an instance mask, and a geometric scoring method selects the suction point by combining surface flatness with normal alignment. Three geometric methods are compared: k-nearest-neighbour PCA, Sobel cross-product, and RANSAC plane fitting. Evaluated on a real robot across three deformation levels and 35 cluttered scenes, single-object grasp success reaches 88.2% and end-to-end retrieval in clutter is 72.6%.

论文arXiv AI 12:00

Performative Privacy: When Differential Privacy Maximizes Utility

arXiv:2608.28198v1 Announce Type: cross Abstract: Privacy-preserving learning is often motivated by the idea that protecting users' data can preserve trust and thus participation, improving utility in the long term. However, this claim has not been formalized so far. In parallel, performative learning provides a framework for studying learning systems whose deployment affects the data they later observe. In this work, we bring these two perspectives together and introduce \emph{performative privacy}, where data leakage reduces future participation. We study a simple model where agents repeatedly contribute data for mean estimation but may leave the system when their data is leaked. Privacy is implemented through differentially private mechanisms, creating a trade-off between estimation noise and future participation. We show, through a theoretical study of the dynamics and numerical experiments, that a finite privacy budget can outperform non-private estimation in the long term when the feedback loop between leakage and participation is sufficiently strong. This provides first evidence that differential privacy can be optimal not only as a protection mechanism, but also from the perspective of long-term utility.

论文arXiv AI 12:00

Beyond Flat Netlist: Hierarchical Graph Representation Learning for Scalable Analysis of Sequential Circuits

arXiv:2608.28188v1 Announce Type: cross Abstract: Circuit Representation Learning (CRL) offers a powerful paradigm to guide and optimize core Electronic Design Automation (EDA) tasks, but its practical adoption is hindered by the immense scale of industrial netlists and a failure to explicitly model register-level temporal dynamics. To overcome these barriers, we introduce DeepSeq3, a novel hierarchical framework that abstracts circuits into a two-level representation: fine-grained combinational subgraphs partitioned by flip-flops (FFs), and a high-level Super-Node Graph (SNG) that models the register-transfer structure. A dual Graph Neural Network (GNN) architecture learns representations at both levels, capturing local Boolean logic and global state transitions. Crucially, we introduce a state-centric pre-training scheme that predicts the reachability between FF states, endowing the model with a deep understanding of temporal behavior. Demonstrated on large-scale benchmarks, DeepSeq3's approach yields superior scalability and richer representations, reducing bounded model checking (BMC) solving time by 18% while guaranteeing correctness.

论文arXiv AI 12:00

Conformal Risk-Averse Decision Making with Optimized Certainty Equivalent Risk Control

arXiv:2608.28179v1 Announce Type: cross Abstract: We study risk-averse decision making, in which an agent selects actions while being uncertain about the true system state. The risk is measured via optimized certainty equivalent (OCE) metrics, which generalize popular criteria such as mean-variance risk and conditional value-at-risk (CVaR). We characterize the optimal policy under known distributions, and show that it reduces to a prediction set-based solution for the CVaR. This provides an operational interpretation of conformal prediction-type prediction sets. For unknown distributions, we develop a data-driven calibration strategy, based on a synthetic model for the likelihood and held-out calibration data, yielding high-probability control of the OCE risk. The approach is evaluated on two wireless beamforming settings.

论文arXiv AI 12:00

Text Restoration of Ancient Documents with Language Models

arXiv:2608.28170v1 Announce Type: cross Abstract: Purpose - This study investigates the feasibility of restoring missing text caused by physical lacunae in damaged ancient manuscripts using language models. Methodology - The study proposes different scenarios to replicate real-world conditions. Language models of different architectures are applied according to their suitability to each scenario. We also propose several decoding strategies that further enhance performance and address the discrepancy between lacuna boundaries and the models' tokenization schemes. Findings - The results reveal that text restoration of these documents cannot be fully automated, but it can serve as a useful tool to assist paleographers in their work. Model performance varies greatly depending on which structural part of the document needs to be restored and whether the character length of missing text is available. Originality - This is the first study and to analyze model performance on formulaic and non-formulaic content and the impact of lacuna length awareness in manuscript restoration. Both are recurring challenges in paleographers' manual restoration work. Through systematic comparison and both qualitative and quantitative analysis of different models' performance under varying settings, this study offers a guideline for developing assistive tools to support paleographers.

论文arXiv AI 12:00

Gen-TAS: A Generative AI-Aided Hardware-Software Task Allocation Framework for FPGA-GPP Heterogeneous Systems

arXiv:2608.28160v1 Announce Type: cross Abstract: FPGA-GPP heterogeneous systems combine software flexibility with the performance and energy efficiency of reconfigurable hardware. However, determining which application tasks should execute on the GPP or FPGA requires extensive expertise and design-space exploration, particularly when user objectives vary across latency, communication, resource utilisation, and power. This paper proposes Gen-TAS, a knowledge-grounded LLM framework for user-specific FPGA-GPP task allocation. By combining task-graph analysis with RAG, Gen-TAS grounds LLM reasoning in historical implementation knowledge and generates multiple explainable strategies tailored to the specified objectives. Human-in-the-loop selection and a deterministic backend connect LLM-generated decisions to reproducible FPGA SoC implementations. Experiments on CNN and SDR workloads across multiple LLMs demonstrate stable, requirement-driven allocation. Under latency-oriented objectives, implementations following the selected strategies achieve speedups of up to 2.45$\times$ and 92.53$\times$, respectively, relative to the corresponding all-GPP baselines while other objectives select strategies that trade some acceleration performance for FPGA-GPP communication, resource utilisation, or FPGA power.

论文arXiv AI 12:00

Nested Byte-Level Vocabularies Are Cheap to Deploy and Expensive to Share: A Pre-Registered Negative Result

arXiv:2608.28151v1 Announce Type: cross Abstract: A byte-level BPE tokenizer is an ordered list of merge rules, so applying only a prefix yields a vocabulary whose token identifiers are the first rows of the full vocabulary. This prefix nesting allows one language model to operate at several vocabulary sizes, use a control token to indicate the active size, and be deployed at any trained size by slicing its embedding and output head. We pre-registered five claims, including margins, seeds, contrasts, and a stop rule, and trained 30 models with 3.1M- and 10.6M-parameter bodies on 200M tokens each. Slicing is numerically exact: across 76 checks, a sliced model reproduces the restricted full model's logits bit for bit and removes 66% of deployed weights without changing latency. However, the shared model trails a fixed-cap specialist by 3.64% bits per byte at 32k against a 1% margin, and by 2.96% at 8k against a 2% margin. A 2x2 ablation separating the control token from output restriction finds that the token changes performance by +0.07% to +0.13%, with all intervals crossing zero, while output restriction costs +0.47% to +1.19%; the factors are substitutes rather than complements. Multi-cap training nevertheless improves robustness: under typographical noise, the same checkpoint degrades 12.5--15.4 points less in its fine mode and outperforms each fixed-cap specialist at that specialist's vocabulary size. A control with neither cap token nor output restriction is equally robust, attributing this benefit to multi-granularity training rather than conditioning. The per-cap penalty tracks each cap's share of training rows, yielding a falsifiable prediction for future work.

论文arXiv AI 12:00

The Approximation Rank of Softmax Attention: Sharp Geometric Laws and Robust Interaction Dimension

arXiv:2608.28150v1 Announce Type: cross Abstract: Which geometry controls the rank complexity of normalized softmax attention? We study maximum-row-$\ell_1$ approximation rank, exactly the least unrestricted rank preserving every bounded vector-valued output. Two sharp worst-case laws isolate support geometry: for fixed $d$ and error $\varepsilon$, spherical self-attention has rank $\Theta_{d,\varepsilon}(\min\{n,(1+\beta)^{(d-1)/2}\})$, while full-ball geometry adds one radial degree and, for $\beta\ge\beta_0(d,\varepsilon)$ and $n\ge C_d e^{\beta/8}$, gives $\Theta_{d,\varepsilon}(\beta^{d/2})$. For a fixed head, row-softmax quotients out row-scalar logit directions: the remaining visible query--key interaction dimension $r$ yields an $r/2$ per-instance upper law, and bounded constructions show this exponent is minimax sharp. Approximate interaction subspaces incur an explicit residual output error and yield a tolerance-indexed SVD dimension. On an 84-head BERT-base calibration set, we observe modest effective-dimension reductions across many head--temperature settings, together with positive associations with finite constructive rank upper certificates. Together, these results separate support geometry, which sets worst-case temperature scaling, from softmax-visible interaction geometry, which controls per-head approximation complexity.

论文arXiv AI 12:00

Post-Edit Re-Verification in Simulator-Backed Engineering Agents: A Controlled Comparison of Verification-Cadence Guidance

arXiv:2608.28147v1 Announce Type: cross Abstract: Engineering agents that interact with external simulators may need to coordinate design modification with reacquisition of engineering evidence for the modified state. We ask whether first post-edit re-verification changes when explicit verification-cadence guidance is retained versus omitted while verification-relevant state/facts are held constant. Cadence-Guided (CG) retained an instruction to request a new simulation after a substantive modification, whereas Cadence-Omitted (CO) removed that instruction; neither condition used a hard gate. The study therefore measures instruction-conditioned post-edit verification-policy adherence rather than spontaneous recognition that prior evidence has become stale. Using DWSIM as the simulator backend and continuous valve-pressure adjustment, five Alibaba/Qwen models were evaluated on eight synthetic cases; each model-case-condition combination was executed three times via live API calls, yielding 120 evaluation slots per condition. Re-verification was observed in 94/120 CG slots versus 32/120 CO slots; cadence violations occurred in 26/120 versus 87/120; and bounded final success was reached in 95/120 versus 35/120. qwen3.5-35b-a3b showed minimal re-verification (1/24 in CG and 0/24 in CO) and no final success in either condition. Within this bounded protocol, explicit post-edit verification-cadence guidance was associated with more re-verification, fewer cadence violations, and more frequent bounded final success, supporting the treatment of verification cadence as an explicit interaction-protocol component.

论文arXiv AI 12:00

CheXtriev: Anatomy-Centered Representation for Case-Based Retrieval of Chest Radiographs

arXiv:2608.28137v1 Announce Type: cross Abstract: We present CheXtriev, a graph-based, anatomy-aware framework for chest radiograph retrieval. Unlike prior methods focussed on global features, our method leverages graph transformers to extract informative features from specific anatomical regions. Furthermore, it captures spatial context and the interplay between anatomical location and findings. This contextualization, grounded in evidence-based anatomy, results in a richer anatomy-aware representation and leads to more accurate, effective and efficient retrieval, particularly for less prevalent findings. CheXtriv outperforms state-of-the-art global and local approaches by 18% to 26% in retrieval accuracy and 11% to 23% in ranking quality. The code is available at https://github.com/cvit-mip/chextriev.

论文arXiv AI 12:00

VICT: Verifier-Instrumented Credit Tracing for Long-Horizon LLM Agent Reinforcement Learning

arXiv:2608.28128v1 Announce Type: cross Abstract: Fine-grained credit assignment is a central challenge in reinforcement learning for long horizon LLM agents. Standard objectives often train from programmatically verifiable terminal rewards by broadcasting each sparse outcome to every action in a trajectory. Existing methods typically seek finer credit from the rollout side, constructing auxiliary trajectory signals or additional comparisons to estimate action importance. Although useful, these approaches still treat the verifier that judged success as a scalar reward, discarding its internal task structure. Our key insight is that many verifiable tasks already encode the relevant checks inside their terminal verifier. We propose VICT (VerifierInstrumented Credit Tracing), a training-time interface that exposes executable or evidence backed atoms and traces them back to actions through dependency-valid proof edges. VICT redistributes group-relative advantage only along those edges, shifting credit assignment from rollout-side inference to verifierside tracing. It preserves the original terminal reward, abstains when evidence is incomplete or ambiguous, and changes only the training-time advantage tensor, requiring no learned critic, process labels, branch rollouts, or inference-time verifier access. On ALFWorld and WebShop, VICT improves substantially over outcome-only training and achieves strong performance alongside recent fine-grained credit methods; ablations rule out dense atom rewards, final-commit credit, temporal proximity, and sparsity as sufficient explanations.

论文arXiv AI 12:00

Do Medical Vision Models Reason About Anatomy? Probing the Spatial Inductive Biases of Learned Visual Representations

arXiv:2608.28092v1 Announce Type: cross Abstract: Interpreting a CT scan means comparing structures on either side, judging how far apart organs sit, and knowing where each one belongs. Medical vision encoders are evaluated on diagnostic accuracy, or through assembled multimodal systems where a failure is hard to attribute, so it remains unclear whether their representations support any of this. We construct SPAR-Bench, eight probes over multi-organ abdominal CT that separate coordinate localization, relational reasoning, and spatial queries, and apply them to five architectural configurations and three medical foundation models, frozen and finetuned. Probes that ask for a comparison within the slice stay at chance, and neither pretraining scale, finetuning, nor architecture closes the gap. Probes that appear solved in domain fall to chance under zero-shot transfer, indicating that their accuracy reflects recall of canonical anatomy rather than computation over the image. Reading the same frozen features with a pooled head rather than the full set of tokens moves relational recovery from 0.7% to 67.8%, so pooled probing understates what a representation holds. Questions the encoders answer well are answered at chance by four open-weight MLLMs. Our results suggest these encoders carry a map of where organs usually lie, and little of the machinery for comparing structures within a particular patient. Code and data will be available at https://spar-bench.github.io.

论文arXiv AI 12:00

VersaGauss: A Versatile Framework for Generating Multiphase Dynamics with 3D Gaussians

arXiv:2608.28069v1 Announce Type: cross Abstract: Recent progress has been made in 3D Gaussian representation for reconstruction, generation, and physical simulation. However, current approaches mainly concentrate on physics-based dynamic generation of solid objects and only handle single-phase collision interactions. We introduce VersaGauss, a unified framework for generation, simulation, and rendering that supports versatile physics-based dynamic generation, particularly for multiphase interactions. Our system takes a few images as input and produces a realistic, physics-driven 3D dynamic scene with multiple objects. To optimize the Gaussian kernel distribution, we develop a particle pruning algorithm. We also propose the Coupled Multiphase Point Method (CMPM) to effectively model and generate multiphase interactions. Additionally, harmonic interpolation within CMPM and a Gaussian evolution strategy are introduced to achieve realistic fluid rendering. Extensive experiments demonstrate that our framework can simulate interactions among various materials such as fluid, rubber, sand, snow, and others. Code is available at https://github.com/Elowen-surj/VersaGauss.

论文arXiv AI 12:00

Dynamic Alignment Compensation for Hallucination Mitigation in Large Vision-Language Models

arXiv:2608.28058v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) remain prone to hallucinations, producing responses that are irrelevant or inconsistent with the multimodal input. Existing mitigation methods mainly rely on external supervision, output calibration, or attention regulation, leaving the internal representation dynamics of autoregressive generation underexplored. We identify an inference-time failure mode in which cross-modal representations degrade across decoder layers and drift across generation steps, destabilizing token prediction and increasing hallucination risk. We propose \emph{Dynamic Alignment Compensation} (DAC), a training-free inference-time method that detects representation divergence and selectively applies lightweight residual compensation. DAC combines Layer-wise Semantic Compensation to mitigate inter-layer degradation with Sequential Semantic Correction to constrain temporal drift. Experiments on nine hallucination-focused and general-purpose multimodal benchmarks across multiple LVLM backbones show that DAC consistently reduces hallucinations while maintaining strong overall performance.

论文arXiv AI 12:00

Explainable Uncertainty Estimation for Reliable Medical AI

arXiv:2608.28052v1 Announce Type: cross Abstract: Artificial intelligence has strong potential to support clinical decision-making, yet its adoption in healthcare remains limited due to a lack of trust. Uncertainty estimation can signal unreliable predictions, and explainable AI (XAI) can clarify how predictions are made but existing methods treat them separately, providing no feature-level insight into why a prediction is uncertain or which tests to prioritize to reduce it. To address this gap, we propose explainable uncertainty estimation, which unifies uncertainty estimation and XAI to both quantify uncertainty and explain feature-level contributions. We introduce the Expected Gradients Reconstruction Uncertainty Estimate (egRUE), which incorporates prediction explanations into its uncertainty computation and decomposes uncertainty into feature-wise contributions. We prove theoretical properties of egRUE and show through experiments that it improves reliability and interpretability compared to existing methods. A user study with medical experts further demonstrates that egRUE's explanations improve calibrated trust over uncertainty scores alone, increasing confidence in correct predictions and reducing confidence in incorrect ones. By combining prediction uncertainty with feature-level explanations, egRUE strengthens decision-making support in safety-critical healthcare settings, clarifying both when predictions may be unreliable and which features drive that uncertainty.

论文arXiv AI 12:00

SimpCue: Cue-Based Prompting for Multilingual Text Simplification

arXiv:2608.28042v1 Announce Type: cross Abstract: Text simplification aims to make complex texts easier to understand while preserving their original meaning. Recent large language models can perform simplification through prompting, but it remains unclear whether adding explicit linguistic information about sentence complexity to the prompt improves their outputs. We investigate this question for multilingual sentence-level Easy-to-Read simplification in Catalan, Spanish, and Italian. Using Qwen3-8B, we compare a baseline prompt, a gold-cue prompt enriched with gold linguistic cues, and a predicted-cue prompt enriched with automatically predicted cues. We evaluate the outputs using SARI, BLEU, chrF, and BERTScore, and complement this evaluation with a manual qualitative analysis. Predicted-cue prompting obtains the best overall scores across all four metrics, although the gains over the baseline are small. Gold-cue prompting does not consistently improve over the baseline, and results vary across languages. These findings indicate that cue-based prompting can influence multilingual Easy-to-Read simplification, but its benefits are modest, metric-dependent, and language-dependent.

论文arXiv AI 12:00

Compared to What? A Human-Anchored Security Benchmark for LLM-Generated Infrastructure-as-Code

arXiv:2608.28021v1 Announce Type: cross Abstract: Large language models are increasingly used to author Infrastructure-as-Code (IaC), where a single insecure default can be deployed directly into production. Prior evaluations report raw vulnerability counts for model-generated IaC, but without a human baseline they cannot determine whether models are actually worse than engineers. We introduce GenIaC-SecBench, a benchmark of 100 deployment scenarios stratified by architectural complexity, evaluated across 12 model configurations from four vendors, producing 1,196 IaC artifacts scanned by three independent policy engines (Checkov, Trivy, KICS). Critically, we also scan 634 human-authored IaC templates with the same toolchain, providing the first size-matched human security baseline. Vulnerability density is strongly inverse to artifact size (Spearman $\rho = -0.55$, $p < 10^{-77}$), meaning unmatched comparisons measure size rather than security. When matched on declared-resource count, all model configurations fall within 3.21x--3.87x the human vulnerability density, with the gap widening for simpler tasks (4.9x at one resource, 1.4x at twenty or more). We decompose reasoning into standard generation, prompt-engineered chain-of-thought, and vendor extended-thinking APIs. Vendor extended thinking significantly outperforms prompted chain-of-thought ($-12.0\%$, $p = 0.0013$), while prompted chain-of-thought is indistinguishable from standard generation ($-1.3\%$, n.s.). Token instrumentation shows extended thinking uses under 1\% of the output budget, explaining its bounded effect. Two negative results also emerge: deployability does not correlate with vulnerability ($r = 0.158$, $p = 0.625$), and classical complete-case Friedman testing is infeasible for realistic benchmark designs, motivating the Skillings-Mack statistic. All code, data, and regeneration scripts are released.

论文arXiv AI 12:00

Twin Worlds: Equivariance-Based Abstention for Evidence-Grounded Reasoning

arXiv:2608.28018v1 Announce Type: cross Abstract: Knowledge-intensive reasoning requires Large Language Models (LLMs) to ground answers in provided evidence. When evidence is insufficient, it is desirable that models abstain rather than confidently generating unsupported answers. Existing abstention methods rely on uncertainty estimation or evidence sufficiency checks, but neither tests whether the reasoning process for generation, driven by the interaction of provided evidence and the model's internal memory parameters, is actually grounded in the evidence. A key contributing factor is that entity mentions in context activate memorised associations, causing models to generate plausible responses ungrounded in evidence. We propose Twin Worlds (TW), a framework for improving reliability in knowledge-intensive reasoning through equivariance-based abstention: unlike invariance, which requires outputs to remain unchanged, equivariance requires outputs to transform correspondingly under entity substitutions. A model grounded in the evidence should produce answers that shift consistently when entities are substituted while their relations are preserved. TW constructs multiple worlds via typed substitutions of the original input that preserve relational structure while reducing parametric priors, and uses equivariance violations as an abstention signal. Across four benchmarks and three model backbones, TW identifies when answers are not reliably grounded in the provided evidence and outperforms uncertainty- and sufficiency-based baselines.

论文arXiv AI 12:00

When Can Conditional Flow Matching Replace Pointwise Negative Log-Likelihood?

arXiv:2608.28010v1 Announce Type: cross Abstract: Flow matching enables likelihood-free training, yet alignment methods increasingly reuse conditional flow matching (CFM) losses as endpoint negative log-likelihoods (NLLs) and their old/new differences as log-likelihood ratios. We characterize when these substitutions are valid. For linear Gaussian paths, we exactly decompose endpoint NLL into entropy, a weighted CFM objective, an interior velocity--score residual, and a boundary residual. Thus CFM-only estimates and differences are exact only when the corresponding residuals cancel. At the off-policy population optimum, ordinary CFM is not generally a pointwise NLL estimator, whereas \(w_{\mathrm{sc}}(t)=(1-t)/t\) removes the interior residual; this positive result does not extend generally to training or on-policy alignment. On-policy log-ratios can remain biased even for identical endpoint laws or after surrogate optimization. Experiments across dimensions, distributions, and geometries support these conclusions and the mechanisms that make inexact ratios useful. **More broadly, the decomposition provides a theoretical basis for adapting likelihood-based LLM methods to flow matching, while distinguishing exact substitutions from controlled surrogates.**

论文arXiv AI 12:00

A Method for Layer Bit-Width Allocation in LLM Quantization via Performance Maximization Under a Quality-Degradation Constraint

arXiv:2608.28003v1 Announce Type: cross Abstract: This paper proposes a layer bit allocation method for Gemma-3-1B, formulating the problem as performance maximization (latency decrease) given a degradation budget constraint (allowable level of generation quality loss). This approach is different from time- and resource-consuming uniform layer quantization methods that are used in the literature (like GPTQ or AWQ) or allocation methods without proven performance-accelerating effect (like MixLLM or TorchAO). The layer sensitivity profile resulting from our prior work SA-PTQ is applied using the activation pass-through mode inside TensorRT-LLM. For each layer precision is determined individually in blocks, according to a grouping introduced in the prior step (5+5, 10+10, all26), differentiating the contribution of FFN, Attention, and lm_head to the overall speedup. The clock speed was measured for 13 W8A8 variants on an RTX 5090. We find that for FFN and lm_head the time cost of quantization/dequantization is compensated for by the use of integer arithmetic, while for short context lengths, the opposite holds true for Attention: an additional step of quantization slows execution down. We propose a manual implementation of SmoothQuant for TensorRT-LLM which was necessary due to export failures, unavailable for lm_head. The best solution found under joint consideration of all three criteria with minimal degradation was FFN 5+5 with lm_head, providing an 11.0% reduction in latency with negligible quality loss (98.90% Top-1 agreement, +0.85% perplexity degradation). With acceptable quality loss for FFN all26 + lm_head, a speedup up to 19.1% was found possible. We suggest further optimizations: fused attention kernels in INT8, KV-cache quantization, using FP8 instead of INT8 and partial Attention quantization analogous to FFN.

论文arXiv AI 12:00

CAITLYN: Can LLM Agents Autonomously Synthesize Defenses against Emerging Injection Attacks?

arXiv:2608.27990v1 Announce Type: cross Abstract: Prompt injection attacks on Large Language Model (LLM) agents seek to introduce malicious instructions or content into external text sources retrieved by agents, forcing the underlying LLMs to execute harmful actions outside their benign scope. While current defenses effectively counter known injection attacks, deploying them in LLM agent environments remains challenging due to attack variants and emerging threats. Moreover, existing solutions typically suffer from an inherent trilemma, i.e., a constant trade-off among runtime efficiency, contextual precision, and adaptability. To bridge this gap, we propose Continuous Agents for Injection Threats via Lifelong Yielding Nexus (CAITLYN), an agent-agnostic defense middleware. CAITLYN integrates two systems. System I focuses on immediate defense against existing attacks using a two-tiered library: Tier-0 for rule-based detection scripts and Tier-1 for optimized LLM-based accurate inference. System II, in contrast, is deployed to monitor potential abnormal signals and attempt to synthesize new defenses. On standard benchmarks, CAITLYN matches the detection performance of state-of-the-art defenses at lower token overhead than LLM-as-a-judge baselines. On Emerging, our new delivery-aware benchmark featuring novel injection techniques, static baselines and the standalone System I configuration remain vulnerable. In contrast, System II autonomously synthesizes verified defense capabilities, substantially lowering the attack success rate across three diverse agent environments.

论文arXiv AI 12:00

Not to Break, but to Attest: Adversarial Probes for Privacy-Preserving LLM Verification

arXiv:2608.27954v2 Announce Type: cross Abstract: Post-deployment changes to large language models can alter behavior while leaving routine outputs largely unchanged, creating a challenge for AI governance when model weights are proprietary. We present a privacy-preserving zk-SNARK-based audit framework that searches for probes designed in the spirit of adversarial examples to amplify logit drift between an approved model and a modified deployment. Our framework explores complementary probe families under different access models. Token-based probes operate in a black-box setting and require only the input interface, tokenizer, and vocabulary. Embedding-based probes require gray-box access to the embedding interface. Stress probes rely on additional interface capabilities but do not require full white-box access to model weights or architecture. This range allows probe selection to balance sensitivity, access requirements, and deployment cost. We evaluate probe constructions across LLM architectures, model-tampering scenarios representative of post-deployment attacks, and GPU platforms. Importantly, our experimental results demonstrate that token-based probes consistently deliver the strongest mean sensitivity across models and GPU platforms, although operating in a black-box setting. Our Groth16 zk-SNARK workflow remains practical as the probe set scales from 1 to 50, where proving time increases from 1.02 to 1.78 seconds, verification remains near 0.84 seconds, and proof size remains constant.

论文arXiv AI 12:00

Antipatterns in AI-assisted Qualitative Data Analysis: A Catalog of Temptations and Pitfalls for Software Engineering Researchers

arXiv:2608.27927v1 Announce Type: cross Abstract: AI-assisted qualitative data analysis (QDA) offers unprecedented opportunities to streamline software engineering (SE) research, yet uncritical use risks compromising analytical rigor and flooding the field with accelerated production of low-quality research. While tactical best practices will naturally evolve over time, SE researchers currently lack strategic guidance to identify and mitigate methodological risks when attempting AI-assisted QDA. Based on our decades of qualitative SE research expertise and experience combined with an understanding of the emerging landscape of AI-assisted QDA, this paper presents a catalog of antipatterns in AI-assisted QDA - a set of assumptions and practices that initially appear advantageous but ultimately undermine analytical rigor and validity. The antipatterns are grouped into three categories reflecting escalating impact: Dangerous Drivers, Operational Missteps, and Analytical Failures. As more SE researchers attempt AI-assisted QDA, these antipatterns will help them identify and avoid common temptations and pitfalls, while reviewers can be equipped with the vocabulary and criteria to call out problematic and failed practice. Ultimately, this catalog of antipatterns can serve as a stepping stone in our responsible methodological evolution toward principled and meaningful human-AI collaboration in qualitative research.

论文arXiv AI 12:00

PCBnet: A Dataset and Automatic Construction of SPICE Netlists from Schematic Images

arXiv:2608.27923v1 Announce Type: cross Abstract: Printed circuit boards (PCBs) are fundamental to modern electronic systems, yet AI-driven PCB design automation remains constrained by the lack of large-scale paired schematic-netlist datasets. PCB schematics are particularly challenging due to diverse component types, complex wiring topologies, and noisy textual annotations. To address this gap, we present PCBnet, a large-scale PCB schematic dataset comprising over 300 real-world designs with annotated pins and paired SPICE netlists. It contains more than 50,000 component instances, 150,000 wires, 100,000 text regions, and 400,000 characters. We further develop an automated schematic-to-netlist pipeline that combines visual recognition, topology construction, and domain-knowledge-guided multi-agent correction. The proposed method achieves 94.54% component detection mAP, 98.57% text recognition accuracy, and 84.47% end-to-end connectivity accuracy. PCBnet provides a benchmark and data foundation for future AI-driven PCB design automation.

论文arXiv AI 12:00

Low-Altitude Fluid Antenna Network with Multi-Agent Reinforcement Learning

arXiv:2608.27909v1 Announce Type: cross Abstract: Low-altitude wireless networks (LAWNs) integrate terrestrial and aerial platforms to provide ubiquitous communication, sensing, and localization services for unmanned aerial vehicles (UAVs) and electric vertical takeoff and landing (eVTOL) aircraft. However, dynamic air-ground and air-air channels, abrupt blockages, and heterogeneous interference hinder the realization of this goal. Nevertheless, fluid antenna (FA), a cutting-edge multiple-input multiple-output (MIMO) technique, overcomes these challenges by reconfiguring antenna positions to unlock additional spatial degrees-of-freedom. In this paper, towards bringing low-altitude FA networks into reality, we study the fast and high-performance FA reconfiguration for low-altitude FA networks with multi-agent reinforcement learning (MARL). Specifically, we present an electromagnetic digital twin (EM-DT)-assisted MARL framework. To fill the sim-to-real gap, we introduce a two-stage transfer learning framework. Our case study shows that joint FA positions and beamforming optimization can enhance the system sum-rate by 118.5%, compared to the fixed position baseline. This gain comes from the dynamic millisecond timescale reconfiguration of FA arrays and the adaptive steering of beams toward aerial users with mobility.

论文arXiv AI 12:00

LandingAgent: A Reference-Annotated Dataset and Agentic Generation Framework for Landing Pages

arXiv:2608.27902v1 Announce Type: cross Abstract: Landing pages are goal-oriented web interfaces that must communicate a target-specific value proposition while organizing information flow, visual hierarchy, and calls to action (CTA). Although large language models can generate plausible webpage code from natural-language prompts, direct generation often yields generic templates and unsupported persuasive claims. We study target-grounded, reference-guided landing-page generation, where a system must create an executable page for a new target by adapting reusable patterns from real pages without copying them. We introduce LandingBench, a reference-profile dataset that abstracts real landing pages into section sequences, layout patterns, tone descriptors, visual emphasis, and CTA structure. Building on LandingBench, we propose LandingAgent, a three-phase agentic framework that profiles the target, constructs a reference-guided wireframe, and refines the page through critique-guided polishing. We evaluate LandingAgent against direct prompting on faithfulness, conciseness, readability, aesthetics, and structural diversity. Experiments show improved target grounding, presentation quality, and layout diversity. Code is available at https://github.com/IAURAI/LandingAgent.

论文arXiv AI 12:00

OpenStamp: A Watermark for Open-Source Language Models

arXiv:2608.27899v1 Announce Type: cross Abstract: With the growing prevalence of large language model (LLM) generated content, watermarking is considered a promising approach for attributing text to LLMs and distinguishing it from human-written content. A prominent class of techniques embeds subtle but detectable signals in generated text by modifying token sampling probabilities. However, such methods are unsuitable for open-source models, where users have white-box access and can easily disable watermarking during inference. In this work, we introduce OpenStamp, a watermarking technique that encodes the watermarking logic directly into the model weights by modifying only the final projection, or unembedding, layer. Through experiments across two models, we show that OpenStamp achieves superior detection performance, with minimal degradation in model capabilities compared to prior methods. The implanted watermark is explicitly designed, and empirically confirmed, to be more robust to paraphrasing attacks and harder to scrub off through post-hoc fine-tuning than prior open-source watermarks. To enable developers to watermark their models, we release our code alongside watermarked versions of 4 popular open-source models.

论文arXiv AI 12:00

SOMTab: Set-Order Mamba for Efficient Tabular In-Context Learning

arXiv:2608.27882v1 Announce Type: cross Abstract: Tabular foundation models based on in-context learning have recently emerged as strong alternatives to task-specific model fitting. However, the current performance frontier remains dominated by attention-heavy architectures, where attention is used throughout the modeling pipeline. This raises a natural question: is attention necessary at every stage of tabular in-context learning? We introduce SOMTab, a Set-Order Mamba architecture for efficient tabular in-context learning. SOMTab separates representation construction from query-conditioned retrieval. For row and column representations, it maps unordered table tokens into stable latent slots and applies Mamba-based state-space mixing to construct compact representations. For final prediction, it retains attention-based in-context learning to preserve query-conditioned retrieval from labeled context examples. We further introduce DCH-TailMix, a synthetic prior that combines degree-corrected graph heterogeneity with mixed heavy-tailed regimes to diversify synthetic dependency structures. Across tabular benchmarks, SOMTab approaches the performance of strong Transformer-based tabular foundation models while achieving faster inference and lower GPU memory usage, yielding a favorable efficiency--accuracy trade-off.

论文arXiv AI 12:00

From Perspective to Fisheye Depth Estimation and Open-Vocabulary Segmentation

arXiv:2608.27860v1 Announce Type: cross Abstract: Vision foundation models are capable of generalizing across 3-dimensional (3D) scenes with high-fidelity estimates; their empirical success can be attributed to training on large-scale datasets of perspective images. However, when transferred to wide field-of-view (FoV) images, such as those captured by fisheye cameras, they return erroneous outputs due to a covariate shift stemming from the radial distortion on the image pixels. We propose a method to generalize vision foundation models to fisheye cameras. The crux of our method lies in a set of learnable parameters, termed Distortion Extenders (DEX), that model the fisheye distortion coefficients and the distributional shift between fisheye and perspective images encoded in the latent space. By minimizing a self-supervised alignment loss, DEX transforms the latent embeddings of fisheye images to resemble those of perspective images to recover high-fidelity estimates. DEX is architecture- and task-agnostic: We demonstrate DEX on monocular depth estimation and open-vocabulary segmentation for convolution- and Transformer-based architectures, where we consistently improve over baselines across indoor and outdoor fisheye datasets. As a byproduct, the activations of DEX can also be decoded to distortion coefficients to support camera calibration. Code available at: https://github.com/Suchisrit/DEX.

论文arXiv AI 12:00

FedEHR-Agents: Federated Agentic Optimization for Automated EHR Modeling

arXiv:2608.27856v1 Announce Type: cross Abstract: Recent advances in large language models are enabling autonomous clinical agents to perform increasingly complex electronic health record (EHR) modeling workflows. However, agents deployed at individual hospitals remain constrained by institution-specific data and modeling environments, while direct cross-hospital collaboration is restricted by the sensitivity of patient-level EHR data. Although federated learning (FL) provides a natural foundation for privacy-preserving collaboration, existing approaches remain predominantly model-centric, limiting federation to prediction models or their updates while overlooking the richer modeling experience accumulated by autonomous agents. To address this limitation, we propose FedEHR-Agents, an experience-centric federated agentic optimization framework for automated EHR modeling. Each hospital deploys an autonomous clinical EHR agent that performs data preprocessing and model development while refining local clinical modeling experience through historical memory, task-specific evaluation, and TextGrad-based prompt refinement. The federated server performs evidence-guided experience aggregation to integrate reliable and complementary modeling experience across heterogeneous hospitals and distills the aggregated experience into global meta-prompts for subsequent local refinement. Extensive experiments on real-world multi-hospital EHR benchmarks demonstrate that FedEHR-Agents consistently outperforms local and federated baselines across diverse clinical prediction tasks and remains robust across different federation scales and LLM backbones. These results establish clinical modeling experience as a promising collaborative object beyond conventional parameter-centric FL and point toward federated autonomous clinical intelligence.

论文arXiv AI 12:00

FISGuard: Defending Against Membership Inference via Fixed Input Subspaces

arXiv:2608.27836v1 Announce Type: cross Abstract: As large language models are increasingly adopted in federated learning, protecting user privacy while performing parameter-efficient fine-tuning on distributed private data has become an important challenge. Although clients only share gradients instead of directly uploading raw data, the shared gradients may still leak membership information about training samples. ProjRes (S&P, 2026) further increases this risk: with less information and without accessing model outputs, an attacker can effectively distinguish members from non-members solely based on the projection residual between a candidate representation and the subspace induced by server-observable gradients. Existing defenses against membership inference mostly rely on gradient perturbation or regularization, which can not only degrade model utility but also fail to effectively defend against the membership inference attack introduced by ProjRes, which exploits the geometric structure of gradients. To address this issue, we propose FISGuard, a lightweight defense. Its key idea is to construct and fix a low-dimensional representation subspace using independent public data, thereby restricting the space through which private representations are exposed via gradients while preserving the primary information required for downstream tasks. This substantially reduces the projection-residual discrepancy between members and non-members. We evaluate FISGuard against five representative defense methods across three NLP datasets, two LLMs, and two fine-tuning strategies, Adapter and LoRA. The results show that FISGuard reduces the ProjRes attack AUC to near the random-guessing level of 0.5 in most settings, while maintaining downstream task performance close to that of the undefended model and introducing only limited computational overhead, thereby achieving a favorable privacy--utility trade-off.

论文arXiv AI 12:00

Actionable CBFI: Integrating Structural Decomposition and Causal Counterfactual Recourse for Tabular Machine Learning

arXiv:2608.27821v1 Announce Type: cross Abstract: Explainable artificial intelligence (XAI) increasingly calls for actionable counterfactual recourse, yet current methodologies face challenges related to causal invalidity, excessive cognitive burden, and predictive failure. Exhaustive causal search algorithms often require modifications to multiple attributes, whereas additive attribution-guided methods, such as SHAP, ignore higher-order feature synergies, leading to suboptimal predictive momentum and diffuse intervention effort in complex nonlinear models, such as XGBoost. To bridge this gap, we introduce actionable case-based feature importance (A-CBFI), a diagnosis-prescription integrated framework for tabular machine learning. Grounded in structural causal models (SCMs), A-CBFI isolates synergistic interaction bottlenecks and releases suppressive structural locks, translating them into targeted interventions. By mathematically separating the active user intervention space (L_{\mathrm{active}}) from downstream effects and concentrating over 98.3% of the intervention effort on diagnosed root causes, A-CBFI enables highly targeted interventions. Empirical evaluations across the financial and healthcare domains demonstrate that A-CBFI reduces the active human intervention burden by 76.9% while maintaining comparable global recourse cost to exhaustive causal baselines. By prioritizing the diagnosed causal bottlenecks, A-CBFI provides targeted and actionable recourse while maintaining causal validity and achieving full relative convergence across all causally feasible instances.

论文arXiv AI 12:00

ContextLeak: Exfiltrating LLM Agent Context via Malicious Tools

arXiv:2608.27800v1 Announce Type: cross Abstract: Exfiltrating an LLM agent's runtime context -- such as the user prompt, execution trajectory, and tool list -- poses severe security and privacy risks to users. Such attacks can be carried out via malicious tools and typically require three conditions: (1) the agent selects the malicious tool for task execution, (2) the agent passes its runtime context as input arguments to the tool, and (3) the tool's implementation transmits these inputs to an attacker-controlled endpoint. Existing work primarily focuses on conditions (1) and (3), leaving condition (2) largely unexplored, despite its critical role in enabling successful context exfiltration. In this work, we bridge this gap by developing ContextLeak, a malicious tool attack that induces the agent to both select the tool and disclose its context as input arguments. We realize this attack by carefully crafting the tool's name and description using reinforcement learning. Specifically, ContextLeak employs an LLM, referred to as the attack LLM, to automatically generate the malicious tool's name and description. To improve attack effectiveness, we fine-tune the attack LLM via reinforcement learning on a set of shadow users with diverse, simulated agent contexts. Our key technical contribution is the design of novel reward functions tailored to the context exfiltration objective, enabling effective reinforcement-learning-based fine-tuning of the attack LLM. Extensive evaluation demonstrates that our attack remains highly effective even when the shadow users' contexts differ substantially from those of the victim users. Moreover, ContextLeak significantly outperforms existing malicious tool attacks when adapted to this setting.

论文arXiv AI 12:00

How Much Can AI Understand? Toward AI-Assisted Sensemaking of Collaborative Discussion in Groups with Shared History

arXiv:2608.27799v1 Announce Type: cross Abstract: AI tools that support collaborative discussion typically treat the discussion as a standalone task, focusing only on its content and setting aside the social context of the group having it. But it is groups with a shared history, with their own norms, hierarchies, and relationships, where the most tangled and complex discussions tend to arise. These discussions cannot be understood apart from that context, and AI that overlooks it risks failing to convey what a discussion means, or even misrepresenting it. Drawing on two studies of how experienced Wikipedia editors read and make sense of discussions, we propose an AI-Assisted Sensemaking Model for Collaborative Discussions, which captures not only a discussion's arguments but also the norms and participants behind it, along with the context that gives each meaning. In this model, the system supports the early stages of the sensemaking process, and the degree to which it performs interpretive work can range from low to high. We argue that higher interpretive work reduces the burden on users but increases their reliance on the system's judgment. We then discuss the risks of an insufficiently intelligible system, what it would take to make one more intelligible, and the safeguards it still requires.

论文arXiv AI 12:00

Compositional Failure in Audio-Visual LLMs: Late-Layer Prior Dominance Under Cross-modal Conflict

arXiv:2608.27785v1 Announce Type: cross Abstract: We study audio-visual conflict as a compositional generalization test for AV-LLMs: the model must combine synchronized but semantically incompatible audio and video evidence and decide whether the pair matches. On VideoLLaMA 2-7B-AV, three alignment configurations remain nearchance on the scored exact-string Yes/No subset of AVHBench, even though their output priors shift substantially. Similarly, off-the-shelf InternVideo2 experienced a 32.3% accuracy decrease specifically under cross-modal conflict, accompanied by a 17.3% instruction-following failure. We call this failure mode prior dominance: late-layer commitment to an internally preferred answer pattern that is weakly grounded in the conflicting inputs. To explain this behavior, we conduct a mechanistic interpretability analysis and find that commitment remains concentrated at 25.5 $\pm$ 1 layers. We show that stronger temporal alignment changes answer bias, but do not improve compositional conflict resolution. Code and data to reproduce our mechanistic audit and behavioral evaluations are available at https://github.com/AdarshSudheer09/AVHBench-dmai.

论文arXiv AI 12:00

Beyond Search-Imitation: Prior-Directed Exploration for Searchless Chess

arXiv:2608.27757v1 Announce Type: cross Abstract: Searchless chess networks reach human master strength from a single forward pass by imitating a stronger teacher: the strongest, Leela Chess Zero's (Lc0) released Chessformer, distills the visit counts of an AlphaZero-style Monte Carlo Tree Search (MCTS). Imitating a search is a poor proxy for playing without one, so we fine-tune for single-pass strength with self-play reinforcement learning (RL). Its exploration is usually supplied by an entropy bonus, the reverse Kullback-Leibler (KL) divergence to uniform. We replace it with a forward, mass-covering KL toward the network's own MCTS prior (prior-directed exploration), so exploration covers the moves the prior judges promising, and pair it with an entropy-adaptive sampling temperature, set by the value head's outcome uncertainty, that sharpens once a position is decided. In about two thousand steps it raises puzzle accuracy from 93.9% to 94.9% on a 100,000-puzzle suite and mate-in-four accuracy from 77% to 81% while holding searchless strength at or slightly above the base. Measuring tactical accuracy and playing strength together across a matched-compute sweep, we find the two dissociate: accuracy gains fall in a one-point band while ratings straddle the base, and a control fine-tuned on puzzles alone posts the study's largest tactical gains while shedding roughly 260 Elo; a better puzzle-solver is not thereby a stronger player. Distribution-level measurements show what anchoring buys: without a regularizer self-play collapses onto a single line of play, and the puzzles newly solved are the near misses whose winning move the prior kept alive. The forward-KL prior tops the rating ladder, statistically tied with a reverse-KL anchor that concentrates twice as hard and drops the hardest solutions the mass-covering prior keeps in support.

论文arXiv AI 12:00

Efficient Auto-Interpretability of AI Models in Biology

arXiv:2608.27754v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs), and other interpretability methods could turn AI models in Biology and other fields into engines of scientific discovery by explaining the superhuman capabilities of those models. However, a latent is only useful if we know three things: whether it is coherent, whether it can be described, and whether that description has predictive power. These questions are routinely conflated. We assemble them into a single pipeline and report the practical innovations each stage required. First, cross-seed dictionary stability prioritises which latents are worth spending resources to investigate. Second, an intruder-detection task asks whether a latents activating examples share a recognizable pattern. Third, a separate pass proposes a candidate biological description which we convert into falsifiable predictions which can be tested in silico. Deployed on the Boltz-1 Pairformer trunk, stability prioritisation finds interpretable latents using about 4.4 times fewer latent evaluations each, and at 5.2 times lower measured cost, while recovering over half of them, and the external check shows the surfaced motifs are significantly enriched for their claimed annotations. The results also suggest a possible tension: the cross- seed stability might be selecting for some types of features, like structure-related ones, much more than others, such as function-related features.

论文arXiv AI 12:00

RiskBlend: A Multi-Signal Framework for Test Input Prioritization in Machine Learning Regression Testing

arXiv:2608.27704v1 Announce Type: cross Abstract: When machine learning classifiers are retrained, inputs correctly classified by the previous model version may be misclassified by the updated version, creating regression faults that are costly to detect because verifying predictions against ground truth may require human annotation, expert review, or expensive simulation rather than inexpensive model inference. Test input prioritization addresses this problem by ranking inputs so that a limited verification budget reveals as many regression faults as possible. Existing approaches rely predominantly on single-model confidence scores and do not exploit how predictions, decision boundaries, and local neighborhoods change between model versions. We propose RiskBlend, a classifier-agnostic prioritization framework that combines four complementary risk signals: historical failure patterns, prediction shift, decision-boundary shift, and neighborhood change. These signals are combined using validation-learned APFD-squared weighting. Across four datasets, five classifiers, four regression-update scenarios, and 15 random seeds, totaling 1,200 experimental configurations, RiskBlend achieves the highest average APFD in all 80 dataset-classifier-scenario combinations, with improvements of up to 0.32 APFD over the strongest baseline. Confidence-based methods remain competitive primarily for linear classifiers on sparse categorical features, which we attribute to feature-space geometry. The results show that cross-version behavioral signals provide important complementary information for prioritizing regression faults in machine learning systems.

论文arXiv AI 12:00

Evaluating Loss Functions in Differentiable Out-of-Domain Sound-Matching with Partial Parameter Distance

arXiv:2608.27698v1 Announce Type: cross Abstract: In out-of-domain (OOD) sound-matching, a synthesizer is optimized to mimic a sound it did not generate. OOD evaluation of loss functions is underexplored in part because the standard "parameter loss" metric requires a shared parameter space between target and imitator, which OOD settings lack. We introduce Partial Parameter Distance (PPD), which applies parameter loss only to the critical parameters that mismatched synthesizers share (e.g., filter cutoffs), enabling automatically evaluated OOD experiments; we verify its results with blinded listening tests. Across seven scenarios involving band-pass filtering, amplitude modulation, and pitch-bending, we evaluate four differentiable loss functions (SIMSE_Spec, L1_Spec, JTFS, DTW_Envelope). Loss-function effectiveness remains tightly coupled to the method of synthesis: SIMSE_Spec excels at filter-cutoff recovery, DTW_Envelope at amplitude-modulation recovery, and JTFS at smooth pitch trajectories. Parameter-based evaluation agrees with listening tests on the top-ranked loss function in five of seven scenarios, demonstrating its utility as a diagnostic tool.

论文arXiv AI 12:00

CARDINAL Predicts Cardiovascular Risk From Non-contrast Cardiac CT

arXiv:2608.27690v1 Announce Type: cross Abstract: Cardiovascular risk prediction remains limited by incomplete clinical data and imaging biomarkers that reduce computed tomography (CT) to a small number of handcrafted features. We developed CARDINAL (Cardiovascular Assessment via Representation learning from Deep Imaging with Nested Anatomical Latent embeddings), a clinically grounded framework that learns compact representations from routine non-contrast cardiac CT for major adverse cardiovascular event (MACE) prediction. In 17,659 patients, CARDINAL was evaluated for 1-, 3-, 5-, and 10-year MACE prediction against American Heart Association (AHA) pooled cohort equations (PCE), AHA predicting risk of cardiovascular disease events (PREVENT), coronary artery calcium (CAC), segmentation-derived CT biomarkers, and 70-feature structural radiomics. Gains were largest at longer horizons. At 10 years, CARDINAL (joint) achieved an area under the receiver operating characteristic curve (AUROC) of 0.866 $\pm$ 0.020 and an area under the precision-recall curve (AUPRC) of 0.890 $\pm$ 0.015, compared with an AUROC of 0.826 $\pm$ 0.023 and an AUPRC of 0.826 $\pm$ 0.022 for structural radiomics, the strongest baseline. CARDINAL also achieved the highest survival concordance index (C-index), 0.753 $\pm$ 0.015, and high-versus-low risk-tertile hazard ratio, 10.78 $\pm$ 3.16, with favorable reclassification and exploratory calibration. These findings suggest that non-contrast cardiac CT contains prognostic information beyond conventional risk equations, CAC scoring, and engineered imaging biomarkers.

论文arXiv AI 12:00

First Make It Playable, Then Make It Good: Staged Interaction Learning for Small Dialogue-Game Agents

arXiv:2608.27672v1 Announce Type: cross Abstract: We present Qwen-GuidePlay-2B, a 2B-parameter language model for dialogue-game interaction. We fine-tune Qwen3.5-2B using three steps: a) SFT on only successful game trajectories from Playpen, b) weighted turn-level SFT, and c) teacher-guided SFT. The teacher model (which is a larger model) is only used to fix formatting and evaluate examples, but does not create new gold actions. Our final model scores 57.12 clemscore and 42.68 statscore on the public Playpen validation. In the officially released challenge results, our model obtains the second-highest Playpen clemscore delta among submitted systems (which is approximately +36 over its base model). Our findings suggest that imitating full trajectories helps with playability, while turn-level and teacher-guided training usually improve decision-making and increase the overall score. Alternative procedurally heavy approaches like replay-repair and hard-example mining did not help, which suggests that small models are performant simply by using careful curation strategies rather than aggressive changes. We make available both the model and the code for reproducibility.

论文arXiv AI 12:00

Semantic Watermarking with Order-Robust Detection over Sub-sentence Units

arXiv:2608.27666v1 Announce Type: cross Abstract: Semantic watermarks tie the mark to sentence meaning rather than token choices, promising robustness to content-preserving edits. However, the detector only observes attacker-supplied text, which can be reworded, reordered, or resegmented to evade detection without content loss. Rewording, reordering, and resegmentation all cause embedding displacement: detection tests embeddings different from those selected during watermarking and can therefore lose the mark. Our adaptive embedding displacement attack (EDA) admits all three edits under a single objective that maximizes this displacement. It uses a public paraphraser and surrogate encoder without access to the provider's generator or secret key. At a 5% false-positive rate (FPR) and content-preservation threshold $\bar{q}=90\%$, EDA successfully removes the mark on between 32.6% and 47.9% of documents across four schemes, the highest among the tested attacks. Therefore, EDA evaluates the schemes' robustness more thoroughly than passive paraphrasing. To address these vulnerabilities, we design (k)-SwordStamp: semantic watermarks with order-robust detection over sub-sentence units, reducing sensitivity to attacker-chosen structure at a small quality cost. Against k-SwordStamp, the strongest no-box attack we test is an EDA variant adapted to its design, with a 10.8% attack-success rate. A stronger EDA with access to the provider's detector and secret key reaches a 39.7% attack-success rate, compared with 65.5% on k-SemStamp. Our code is available at https://github.com/D-Diaa/SwordStamp.

论文arXiv AI 12:00

Knowing Before Answering: Decoding Language Models for Reliable RAG

arXiv:2608.27661v1 Announce Type: cross Abstract: In Retrieval-Augmented Generation (RAG), retrieval may provide insufficient or conflicting information needed to answer a question. The system should not only know when to answer but also be able to identify cases in which the documents provided in RAG are insufficient or contain conflicting information. This can be framed as a three-way classification problem, where we use the model's internal signals to determine whether the provided information in the input can be classified as sufficient, insufficient, or conflicting. We create a controlled benchmark dataset that replicates a RAG setup with fictitious information and labels each instance as answerable, insufficient, or conflicting. We use hidden activations and attention-derived features as inputs to train a lightweight linear model to distinguish among the three classes. Across 16 language models spanning different architectures and a range of model sizes, our feature-based router consistently outperforms prompting-based baselines and the performance of specialised RAG-models. We further conduct analyses into the information dynamics of the models. We show that the most informative signals for the classification are available in the middle layers, with hidden activation states being more effective than attention values or the MLP-feature outputs in most of the tested models. Overall, our results suggest that language models internally encode whether retrieved evidence is sufficient to support answering, and that this signal can be decoded reliably for RAG triage.

论文arXiv AI 12:00

Curvature-Aware Radius Shrinkage for Adaptive Nearest Neighbor Classification

arXiv:2608.27634v1 Announce Type: cross Abstract: Nearest neighbor classification relies fundamentally on how locality is defined, yet conventional $k$-NN imposes the same neighborhood cardinality throughout the feature space. This assumption can be inadequate for data whose local geometry varies substantially across the underlying manifold. We introduce Curvature-Aware Radius Shrinkage for Adaptive Nearest Neighbor Classification (CARSANN), a geometry-driven framework that adapts the spatial support of each neighborhood according to local geometric complexity. CARSANN first estimates intrinsic dimensionality using TwoNN and constructs an intrinsic representation through principal component analysis. Local mean curvature is then estimated using a shape-operator-based formulation and controls neighborhood scale: highly curved regions receive stronger radius shrinkage, whereas approximately flat regions retain broader spatial support. Unlike methods that modify only the number of neighbors or the local metric, CARSANN explicitly adapts the spatial extent of local evidence. Experiments on more than 70 real-world OpenML datasets show that CARSANN consistently improves upon standard $k$-NN and is competitive with adaptive nearest-neighbor methods. In a controlled comparison using the same base neighborhood size, CARSANN achieves higher balanced accuracy on 40 of 45 datasets, increasing mean balanced accuracy from 0.6506 to 0.7528. The advantage also persists against $k$-NN with fixed $k=5$. Friedman and Nemenyi tests confirm that the improvements are statistically significant. These results indicate that local manifold curvature can serve as an effective geometric control variable for adapting neighborhood support, providing a complementary paradigm to cardinality-based nearest-neighbor adaptation.

论文arXiv AI 12:00

Depth-Aware Pothole Detection Using YOLO and RT-DETR at the Edge

arXiv:2608.27633v1 Announce Type: cross Abstract: Pothole detection and its severity measurement is still an important challenges in urban infrastructure management, where late maintenance directly contributes to vehicle damage, road accidents, and escalating repair costs. Existing automated approaches depend on 2D RGB images and cannot measure physical depth of potholes. In this paper, we present a depthaware pothole detection framework and then compare five architectures: YOLOv8n, YOLOv8nSeg, YOLOv9t, RTDETRL, and RTDETRX for RGB-D sensor fusion-based detection and automated depth measurement. A custom offline augmentation pipeline is used here to simulate adverse road monitoring conditions. All models are trained on the PothRGBD dataset with an 80% training and 20% validation split and evaluated using Precision, Recall, mAP@50, and mAP@50_95. Before measuring the depth data, all depth maps are corrected for camera tilt using RANSAC ground-plane orthorectification and all zero-valued sensor pixels are cast to NaN before any statistic is computed. YOLOv8nSeg achieves the highest mAP@50 of 0.9556 and mAP@50_95 of 0.6758 with the most accurate depth estimate of 2.96 cm with the pixel-precise Dseg algorithm. YOLOv8n achieves the fastest inference at 3.6ms. RTDETRX achieves the highest detection confidence at 92.70%. An important finding is that even after full RANSAC orthorectification, bounding box models overestimate pothole depth by 0.16 to 0.21 cm compared to pixel precise segmentation masks. This confirms that the pavement inclusion bias is structural rather than a calibration artifact.

论文arXiv AI 12:00

LitCurate: A Configuration-Driven AI-Assisted Framework for Scientific Database Construction with an Application to Lower-Mantle Equation-of-State Data

arXiv:2608.27629v1 Announce Type: cross Abstract: The growing scientific literature contains decades of experimental and computational results that could support data-driven and physics-based modeling, yet much of this infor- mation remains locked in publications and is not readily usable for large-scale analysis or sci- entific software. Building structured databases from the literature is particularly challenging whenrelevantstudiesmustfirstbediscoveredamonglargecollectionsofpapersandreported quantities must be extracted with enough scientific context to remain usable. We present LitCurate, an open-source framework for building scientific databases from the literature using large language models within an auditable, stage-wise curation workflow. LitCurate integratesliteraturediscovery, relevancescreening, full-textprocessing, andstructuredinfor- mation extraction while retaining intermediate results and provenance, allowing researchers to inspect and revise individual stages rather than treating automated curation as a black- box process. We apply LitCurate to construct an equation-of-state database of lower-mantle and lower-mantle-relevant high-pressure mineral phases from experimental and theoretical studies, comprising 1,334 entries from 205 papers. The resulting dataset links reported equation-of-state parameters to mineral phases, compositions, equation formulations, meth- ods, and parameter constraints, and labels values as source-reported or citation-reported when provenance can be determined. The records are available through a searchable web application. By connecting scientific literature to traceable, machine-readable data, LitCu- rate provides a reusable approach for transforming accumulated literature into resources for scientific analysis and computational modeling.

论文arXiv AI 12:00

Tensor-Accelerated Eager Multi-Resolution Grids for Evolving Large-Scale Substrates

arXiv:2608.27612v1 Announce Type: cross Abstract: In neuroevolution, indirect encoding generates neural network connectivity from a compact genome rather than specifying each connection. ES-HyperNEAT automatically discovers where to place hidden nodes by examining CPPN output patterns: it recursively subdivides space using a quadtree, expanding regions where CPPN outputs show high variance. This adaptive approach discovers network topology without manual substrate specification, extending the fixed-grid HyperNEAT framework built on NEAT. However, the quadtree resists tensorization. Each depth level depends on the parent's variance, forcing sequential evaluation. Different CPPNs produce different subdivision patterns, preventing batching. And variable leaf counts are incompatible with JAX's static shape requirement for JIT compilation. Our prior work confirmed these limits at depths exceeding 5, and a JAX reimplementation of the quadtree yielded only marginal speedup despite batched optimizations, motivating the eager reformulation presented here. We present EMR-HyperNEAT, which evaluates all positions at all resolutions up front, then filters using the same variance criterion: ES-HyperNEAT's subdivide_if(var > $\theta$) becomes eval_all(); filter(var > $\theta$). This performs more CPPN queries than necessary, but all queries become independent and parallelizable across both cores and population members, reducing complexity from \BigO($4^D$) to \BigO($4^D/P$) across $P$ parallel cores. Recurrent substrate configurations become feasible through a connection type taxonomy. The experiments section validates 12-34$\times$ on-device GPU speedup on XOR at depths 5-7, and empirically higher solve rates across benchmarks.

论文arXiv AI 12:00

PHR-VLA: Planning Horizon Reasoning for Vision-Language-Action Models

arXiv:2608.27609v1 Announce Type: cross Abstract: Vision-language-action models (VLAs) have shown strong promise for general-purpose robotic manipulation by mapping language instructions and vision observations directly to actions. However, most VLAs primarily condition action prediction on current observations and lack an explicit mechanism for reasoning over future task dynamics, which is particularly important for fine-grained, contact-rich manipulation. We present PHR-VLA, a framework that enables planning-horizon reasoning in VLAs through privileged latent representations of future dynamics. PHR-VLA introduces a lightweight auxiliary future head that, during training, aligns the VLA's internal representations with latent dynamics extracted from future observations. Evaluation results demonstrate that local, contact-centric, patch-level latent dynamics supervision from the wrist camera improves success rate on LIBERO from 84.1% to 88.4% and on real-world disassembly tasks from 63.3% to 82.5%. Patch-level supervision from a third-person camera also improves performance on Meta-World from 56.70% to 57.8%. These results demonstrate that privileged latent dynamics alignment provides an effective training signal for improving anticipatory reasoning in VLA policies. Project website: \href{https://davoodsz.github.io/PHR-VLA.github.io/}{https://davoodsz.github.io/PHR-VLA.github.io/}

论文arXiv AI 12:00

Quanta Perception as Probabilistic Events

arXiv:2608.27584v1 Announce Type: cross Abstract: Autonomous systems rely on extracting information from light, yet remain brittle in extreme environments, from nighttime navigation to high-speed robotics. Conventional sensors aggregate photons over fixed exposures, imposing trade-offs between sensitivity, dynamic range, and temporal resolution that degrade perception when photons are scarce or dynamics are rapid. Quanta sensors detect individual photons, but their streams exceed real-time compute and latency budgets by orders of magnitude. Here we introduce $\textit{probabilistic events}$, a computational primitive for real-time quanta perception from individual photon detections. By computing the posterior over the time since the last intensity change, we represent photon streams as recursive belief states. Rather than fixed-threshold event-camera triggers, this recursive Bayesian formulation yields three low-latency signals: motion-adaptive scene flux, high-fidelity activity maps, and entropy-based perceptual uncertainty. This representation enables perception in extreme conditions, including pose estimation of a running person at $\sim$0.05 lux---without retraining vision models. Our approach processes input streams exceeding 50{,}000 quanta frames per second on commodity GPU hardware---yielding kilohertz-scale outputs up to four orders of magnitude faster than state-of-the-art quanta reconstruction baselines, even for megapixel arrays. By replacing frame reconstruction with direct probabilistic inference over photon streams, this work bridges photon-counting quanta sensing with robotic vision.

论文arXiv AI 12:00

Self-Explainable Multi-Label Graph Neural Network for Correlated Evidence Attribution

arXiv:2608.27574v1 Announce Type: cross Abstract: Multi-label graph learning intends to capture the intrinsic complexity of real-world applications, where one sample is often related to multiple groups or consists of multiple objects. To date, a handful of multi-label graph learning methods exist, but none of them integrate training-time interpretation capability. While post-hoc graph explainers have been developed, they do not explicitly model label-dependent evidence sharing in multi-label graph learners, especially when label pairs are weakly or negatively associated. As a result, post-hoc approaches may miss how evidence should be shared or separated across different labels. This paper advances a new end-to-end self-explainable multi-label graph neural network (SEMGNN), which aims to simultaneously classify multi-labeled nodes and identify edges significantly contributing to each target node w.r.t. predicted labels. Different from post-hoc methods, SEMGNN jointly learns a predictor and a sparse edge-mask explainer within a unified framework and training objective. Label-label correlations are used to improve multi-label node classification and enhance individual label explanations, so that different labels of a node can be supported by distinct yet coherent structural and/or correlated evidence. Experiments and comparisons on synthetic and real-world multi-label networks, in social networking, entertainment, and life sciences, show that SEMGNN achieves competitive or improved predictive performance while providing more faithful and compact label-conditioned explanations.

论文arXiv AI 12:00

FVeinSyn: Synthetic Finger Vein Image Generator

arXiv:2608.27527v1 Announce Type: cross Abstract: A major challenge in finger vein recognition is the lack of large-scale public datasets. Existing datasets contain few identities and limited samples per finger, restricting the advancement of deep learning-based methods. To address this, we propose FVeinSyn, a large-scale controllable synthetic data generation framework for finger vein. It explicitly decouples synthesis of vascular topology and imaging appearance to mitigate the limitations caused by insufficient training samples, such as inadequate identity diversity and restricted realism. Specifically: first, a finger vein identity generator models vascular topology under physiological and geometric constraints using stochastic L-systems, producing anatomically valid and identity-distinctive vascular patterns. Then, a cascaded region-aware GAN renders the topological maps into realistic near-infrared images. Finally, an intra-class diversity generator introduces geometric and optical perturbations to simulate realistic intra-class variations. Using FVeinSyn, we generated 500,000 images (10,000 vein identities, 50 samples per identity) and conducted extensive evaluations. Results show that FVeinSyn holds significant advantages in realism, identity diversity, vascular pattern consistency, and intra-class diversity. Models trained with FVeinSyn outperform real-data-only baselines a cross eight public datasets, achieving an average accuracy improvement of 27.43\%. The code is available at: https://github.com/EvanWang98/Synthetic-Finger-Vein-Generator.

论文arXiv AI 12:00

Destroy Me: Automatic Artifact Generation for Histopathology Images

arXiv:2608.27516v1 Announce Type: cross Abstract: Deep learning's diagnostic utility in pathology is constrained by model vulnerability to real-world data imperfections. While current strategies favor "perfect data" by filtering low-quality regions, which can lead to the loss of valuable diagnostic context, we propose a paradigm shift: engineering models to thrive in imperfect environments using "Destroy Me", a hybrid framework for realistic artifact synthesis and robust data augmentation. Our approach combines Stable Diffusion, fine-tuned to preserve morphological continuity by realistically integrating artifacts with the underlying tissue architecture, with physics-based procedural modeling to synthesize six common artifact types: tissue folds, precipitates, blur, stitching errors, dust, and pen markers. Artifact fidelity is assessed using Kernel Inception Distance (KID) and color Wasserstein distance metrics. Validating this strategy on lung adenocarcinoma pattern classification with an nnU-Net, we confirm that models trained on "destroyed" patches consistently outperform baselines on independent real-world datasets. Specifically, we observed a 10.5% relative improvement in macro F1-score and a 15% relative increase in the Cohen's Kappa ($\kappa$) coefficient. Crucially, our results demonstrate that selective, impact-weighted augmentation is vital for balancing practical robustness with the preservation of subtle diagnostic features.

论文arXiv AI 12:00

Trajectory-Level Speculative Decoding for Diffusion Language Models

arXiv:2608.27514v1 Announce Type: cross Abstract: Diffusion-based language models (dLLMs) enable parallel token generation through iterative denoising, but existing decoding strategies collapse to single-token generation under low confidence, severely limiting throughput. Unlike autoregressive models where speculative decoding operates on token sequences in a fixed left-to-right order, dLLMs require speculating over denoising trajectories-sequences of multi-token updates with explicit positions and unmasking orders. We develop a trajectory-level speculative framework that constructs draft denoising trajectories via confidence-stratified tree exploration and verifies them through blockwise parallel evaluation with bidirectional attention masking. Our method further introduces inter-block speculation, exploiting diffusion models' bidirectional structure to perform cross-block lookahead. We formally characterize when this approach is exact and identify trajectory drift as the fundamental cost of increased parallelism. Building on Fast-dLLM's dual-cache infrastructure, our framework reduces denoising iterations by 30-40% and increases tokens-per-step from 2.6 to 4.3, achieving 7-14x speedup over vanilla dLLMs and 1.3x over Fast-dLLM with less than 1% accuracy change across reasoning and code benchmarks.

论文arXiv AI 12:00

DAMP: Decay-Aware Mixed-Precision Recurrent-State Quantization

arXiv:2608.27513v1 Announce Type: cross Abstract: Softmax attention stores key and value vectors for every preceding token, causing inference memory to grow with sequence length. Recent language models incorporating Gated DeltaNet (GDN) or Kimi Delta Attention (KDA) reduce this cost by replacing the KV cache in most layers with fixed-size recurrent states. However, these recurrent states are commonly stored in FP32 and consume substantial GPU memory; their updates are memory-bandwidth bound and contribute significantly to decoding latency. To our knowledge, we are the first to study post-training quantization of recurrent states in GDN and KDA based language models. We find that uniform quantization provides a poor accuracy--storage trade-off: INT8 and FP8 already degrade accuracy on complex reasoning tasks, while INT4 and NVFP4 reduce it to near zero. We further find that most quantization-error energy is concentrated in a small subset of channels and that the relative decay strength of state channels remains stable across prompts and tasks. Motivated by these findings, DAMP uses both quantization-error energy and decay-based persistence to identify high-risk channels during offline calibration. It stores these channels at higher precision and the remainder in INT8. We evaluate DAMP on Qwen3.6-35B and Kimi-Linear-48B across six benchmarks covering mathematical reasoning, general reasoning, and code generation. At 9.9 bits per state value, DAMP maintains average accuracy close to the FP32 baseline. DAMP reduces recurrent-state storage by 69.1%, accelerates the recurrent-state update kernel by up to 2.01x, and lowers full-model TPOT by up to 10.9%.

论文arXiv AI 12:00

Quantization-Triggered Backdoors in Language Models: Cross-Quantizer Transferability and the Validation--Deployment Gap

arXiv:2608.27512v1 Announce Type: cross Abstract: Post-training quantization is often treated as a semantically neutral optimization for edge deployment of Large Language Models. When a full-precision source checkpoint is evaluated and quantization is applied downstream without equivalent re-evaluation, this workflow creates a structural validation--deployment gap: because quantization is a many-to-one mapping over parameter space, source-precision certification does not guarantee behavioral equivalence in the deployed configuration. We formalize this gap through Quantization Behavioral Equivalence Classes (QBECs) and prove that QBEC membership does not imply behavioral equivalence, providing a theoretical basis for quantization-triggered backdoor attacks. Building on a three-stage adversarial fine-tuning framework, we embed latent malicious payloads into models that satisfy the source-precision checks used in our evaluation, yet activate targeted adversarial behavior upon INT8 or 4-bit compression. We evaluate this threat in two operationally motivated scenarios, tactical machine translation and political content analysis, extending prior work from decoder-only causal LMs to multilingual encoder-decoder sequence-to-sequence models. Results show that backdoored translation models move from zero measured friend--foe corruption at repaired FP16 to up to 85.02% inversion after quantization, and that a paired stance classifier measures an ideological shift of up to $\Delta\mathrm{Bias}=0.33$ upon compression. A cross-quantizer transferability analysis further shows that attack persistence varies across quantization schemes and model architectures, rather than being determined by nominal bit-width alone. These findings demonstrate that source-precision auditing alone does not rule out quantization-triggered behavior and that the final deployed configuration must be included in behavioral certification for trustworthy edge AI.

论文arXiv AI 12:00

Marginal Coverage Credit Reduces Redundant Exploration in Parallel State-Entropy Optimization

arXiv:2608.27507v1 Announce Type: cross Abstract: Policy Gradient for Parallel State Entropy maximization (PGPSE) expands state-space coverage by training independently parameterized policies in replicated copies of the same environment. However, its pooled team-entropy score measures only collective exploration and cannot identify policies that contribute non-redundant coverage. We introduce Marginal Coverage Credit for PGPSE (MCC-PGPSE), which combines leave-one-policy-out coverage with state-owner specialization to estimate policy-specific credit. MCC-PGPSE preserves PGPSE's pooled objective and redistributes non-negative auxiliary intrinsic rewards according to these credits without changing their total mass. This redistribution is designed to discourage redundant visitation and promote complementary coverage. We evaluated MCC-PGPSE in controlled environments, seven public discrete-state benchmarks, and representative Room and Maze settings from the original PGPSE protocol. Across all tested settings, MCC-PGPSE produced positive final window gains in normalized team state entropy and state support over the Entropy baseline. Controlled-task comparisons and the fixed-suite public aggregate were significant, whereas five-seed original-protocol comparisons were directionally consistent. Ablations and credit alignment controls indicate that most gains arise from leave-one-policy-out coverage rather than non-uniform weighting, mismatched credit, or neural novelty alone. These results support contribution-conditioned auxiliary reward allocation as an interpretable approach to improving complementary coverage among parallel policies in discrete state spaces.

论文arXiv AI 12:00

A Survey on Rubric-Guided Reinforcement Learning for Language Models

arXiv:2608.27505v2 Announce Type: cross Abstract: Reinforcement learning from human feedback (RLHF) has become the dominant paradigm for aligning large language models (LLMs) with human preferences. However, traditional RLHF relies on scalar reward signals that lack interpretability and fail to capture the multifaceted nature of response quality. Rubric-guided reinforcement learning addresses these limitations by introducing structured, interpretable evaluation criteria, or rubrics, as the backbone of reward design, feedback generation, and policy optimization. In this survey, we introduce a Bayesian framework that defines constitutions as prior distributions $P(R)$ over evaluation criteria and rubrics as conditional instantiations $R_x \sim P(R|x)$. Under this unified view, we present a taxonomy of rubric-guided RL along the prior-posterior axis, covering constitutional AI, instance-specific rubrics, process-level supervision, self-evolving rubrics, and their agentic and multimodal extensions. Furthermore, as rubrics are natural-language artifacts, we present a linguistic analysis of how granularity trade-offs, semantic drift, and linguistic reward hacking impact alignment reliability, identifying key open problems for future research.

论文arXiv AI 12:00

XHotpotQA: A Benchmark for Cross-Lingual Knowledge Composition in Multi-Hop Question Answering

arXiv:2608.27481v2 Announce Type: cross Abstract: Knowledge-intensive multi-hop question answering requires systems to select evidence and compose dependent facts, yet multilingual benchmarks usually translate an entire example into one language. This hides failures at language boundaries inside the reasoning chain. We introduce XHotpotQA, a controlled benchmark for cross-lingual knowledge composition over mixed-language evidence. Each instance is modeled as an evidence-dependency graph whose question, bridge evidence, answer-bearing evidence, and distractors have explicit language assignments. The audited resource contains 15,661 training and 7,405 validation instances, with sentence-level support supervision and supplied distractors. In validation, 99.81% of items cross the question-to-gold-evidence language interface and 95.60% use gold paragraphs in different languages. Across three reader artifacts, full question-evidence mismatch is associated with 10.25 to 15.79 lower Unicode-aware answer F1 than partial alignment, and different-script evidence with deficits of 11.98 to 23.70 points; the corresponding adapted-selector contrasts are 1.71 and 1.78 points. Under this supplied-candidate design, the evaluated readers therefore show substantially larger condition-associated deficits than the selector. XHotpotQA provides role-aware diagnostics, modular evaluation, and an audited test bed for knowledge-based systems that must integrate evidence across languages.

论文arXiv AI 12:00

Select, Don't Train: The Benefits of Modular Entity Disambiguation with LLM-Based Selection

arXiv:2608.27470v1 Announce Type: cross Abstract: Entity Disambiguation (ED) is a key task for constructing and using knowledge graphs. State-of-the-art neural approaches commonly model ED as a single task, although it consists of two distinct subproblems: retrieving candidate entities and selecting the correct one given context. Dual-encoder models optimize for both within a shared embedding space, forcing representations to balance high-recall retrieval with fine-grained selection, and they require trained retrievers, which are costly to maintain as knowledge graphs change. While recent work has begun to combine retrievers with LLM-based selectors, the interplay between the two stages has not been studied systematically. In this paper, we present a systematic comparison of retrieval strategies for candidate generation under a shared LLM-based selection stage, combining sparse retrieval (BM25), Web KB search, and a state-of-the-art trained dense retriever with several open- and closed-source LLMs. We show that, once selection is delegated to a capable LLM, training the retriever provides only modest additional value: a fully training-free BM25 retriever paired with an LLM selector reaches a new state of the art on the ZELDA benchmark, raising inKB micro-F1 from 82.3 to 86.3 (+4); pairing the same LLM with a trained dense retriever reaches 88.5. Decoupling retrieval from selection also exposes a limitation of current ED systems: when the correct entity is missing from retrieved candidates, they are forced to predict an incorrect entity. In contrast, our framework allows for abstention when retrieval failure is detected. In an evaluation setting that rewards correct abstentions, the training-free BM25 + LLM pipeline reaches 90.7 F1.

论文arXiv AI 12:00

UIC-AIHealth4All at ArchEHR-QA 2026: Answer-First Evidence Grounding for Clinical Question Answering

arXiv:2608.27467v1 Announce Type: cross Abstract: We describe the UIC-AIHealth4All system for ArchEHR-QA 2026, a shared task on grounded question answering from electronic health records. We participated in Subtasks 2 (evidence identification), 3 (answer generation), and 4 (answer-evidence alignment). For Subtasks 2 and 3, we propose an answer-first pipeline in which the model generates candidate answers citing specific note sentences before classifying the full evidence set, exploiting the asymmetry between judging relevance in the abstract versus relative to a generated answer. For Subtask 4, we apply self-consistency voting over five independent model calls, retaining links by vote threshold. Our pipeline ranked third on evidence identification (Strict Micro F1 62.90), ninth on answer generation (Overall 31.90), and fifth on answer-evidence alignment (F1 79.81). A post-hoc linguistic analysis of 45 stylistic features reveals that model outputs remain 3.2 Flesch-Kincaid grade levels harder to read than clinician-authored references despite matching their word and sentence counts, suggesting readability warrants explicit optimization in clinical NLP systems. Code and prompts are available at https://github.com/mo-arvan/archehr-qa-2026-uic-aihealth4all.

论文arXiv AI 12:00

PACE: Publisher-Adaptive Content Extraction via Agentic Automation

arXiv:2608.27466v1 Announce Type: cross Abstract: Web content extraction is essential for reliable LLM data pipelines, yet existing methods often struggle to jointly satisfy accuracy, scalability, and adaptability. General-purpose extractors can be applied broadly, but they are often brittle on publisher-specific layouts and richer extraction targets such as metadata, images, and tables. Direct LLM-based extraction offers greater flexibility, but incurs substantial cost and latency at scale, while manually engineered publisher-specific parsers can achieve high accuracy but require substantial human effort to build and maintain. We introduce PACE, an agentic framework for learning publisher-specific extraction configurations from representative pages and user requirements. During training, PACE uses LLMs to analyze page structure and aggregate reusable extraction patterns. At inference time, the learned configurations instantiate a fixed deterministic extractor template, enabling scalable extraction without additional LLM calls. Experiments spanning article-body, metadata, and multimodal extraction show that PACE outperforms scalable non-manual baselines while approaching the quality of manually engineered publisher-specific parsers. PACE achieves stronger extraction of article text, metadata, images, and tables, demonstrating that agentic configuration learning can automate publisher-specific extraction for LLM-ready page representations beyond article text.

论文arXiv AI 12:00

The Effect of Emotional Context on Large Language Models' Endorsement of Premature Decisions: Comparing Emotional Vulnerability Across Six Commercial Models

arXiv:2608.27465v1 Announce Type: cross Abstract: As large language models (LLMs) are increasingly used for everyday decision-making advice, whether a model shifts the direction of its advice according to the user's emotional state has become an important safety problem. We test whether emotional expression increases a model's endorsement (encouragement to proceed) when a user, holding the same objective information, is overconfident about a premature decision (e.g., quitting a stable job on weak evidence). As a key control, we include a no-emotion multi-turn (neutral) condition that holds factual content and the number of conversational turns constant, isolating the effect of emotion from that of conversation length. We exposed six commercial models (top-tier and mid-tier models from OpenAI, Anthropic, and Google) to three scenarios (career change, business expansion, emigration) across three conditions (cold/neutral/distress) with six repetitions each, yielding 324 conversations, and measured endorsement strength (0-100) via an eight-item rubric-based automated scoring. Emotional expression significantly increased endorsement (neutral 18.6 to distress 31.5, +12.9 points; mixed-effects $\beta = +12.9$, $p < .001$; Cohen's d = 0.51), and this was not explained by conversation length (cold-neutral difference non-significant, $p = .083$). Critically, the vulnerability varied by individual model rather than by price tier: five of six models showed a significant emotion effect, including the top-tier flagships Gemini 3.1 Pro and GPT-5.5, while only Claude Opus showed no significant change. Results were reproduced with an independent non-Google judge model ($\rho = .89$) and agreed in rank with two human coders ($\rho = .70$). Through a controlled design that separates emotion from conversational context, we show that emotional context increases LLM sycophancy even in top-tier flagship models.

论文arXiv AI 12:00

Sledgehammer or Scalpel? A Fine-grained Adaptive Framework for Implicit Hate Speech

arXiv:2608.27462v1 Announce Type: cross Abstract: Unlike explicit attacks with obvious profanity, implicit hate speech hides malice within seemingly compliant expressions through metaphors and contextual hints, making its detection in online content review challenging. While existing PLM- or LLM-based methods perform well, they typically apply a single reasoning process to all samples. This overlooks fine-grained linguistic nuances and causes unnecessary computation for simpler cases. We observe that online hate speech is not monolithic but manifests in varied forms. We therefore define three fine-grained categories: Shallow, Targeted, and Context-Dependent. Accordingly, we propose Fine-grained Adaptive Implicit Hate speech Detection (FAID), a novel framework that first performs fine-grained classification and then adapts to specific categories. Specifically, for Shallow samples with surface-identifiable intents, the framework adopts lightweight prompt-tuning for rapid classification; for Targeted comments that bind malicious intent to concealed targets, we design knowledge augmentation to iteratively refine the model and reveal hidden targets; for Context-Dependent comments lacking background information, we utilize an agentic framework that automatically generates prompts to evolve context, infer missing background information and identify ambiguous malicious intents. This adaptive architecture focuses computational resources on complex implicit samples while avoiding redundant reasoning for shallow samples. Experiments on four benchmark datasets demonstrate that FAID significantly outperforms SOTA baselines.

论文arXiv AI 12:00

SciReC: Diagnostic Evaluation of Multimodal, Multi-Turn Relational Reasoning with Adaptive Interaction

arXiv:2608.27461v1 Announce Type: cross Abstract: Relational reasoning requires the process of perceptual understanding, comparing, and integrating the underlying relationships between concepts. This ability consists of multiple categories, such as analogical, structural, and cause-effect, each capturing a different aspect of higher-order understanding. To examine the performance of multimodal large language models (MLLM) on these relational inference tasks, we developed SciReC, a model-adaptive multimodal academic dialog benchmark. As the relational reasoning process involves multiple representations and various factors (visual understanding, exhibiting knowledge, and memory recall), we propose DMRA, a deficit-based diagnostic framework that quantifies the contribution of these components to identify the primary cause of unsuccessful cases. Claude 4.6 achieved the best performance on the overall relational score with 73\%, followed by GPT 5.4 with 68\%. Performance trends indicate that open-source models achieve their lowest scores on spatial relations, while proprietary models struggle more with hierarchical and sequential relations. Across domains, model performance is lowest on Astronomy and highest on Psychology. The results of DMRA reveal that relational reasoning is the primary source of error across all models, followed by memory limitations.

论文arXiv AI 12:00

Logos: An Agent Harness on a Cross-Process Bus

arXiv:2608.28553v2 Announce Type: new Abstract: Modern agent systems assemble capabilities at runtime, and this dynamic composition has recently received a complete formal treat ment in the spatiotemporal-composability calculus, in which a capability is a component carrying a tracked inverse, and agents are assembled as plugins. This plugin form is carried by a single process sharing one context, a carrier that places all components in one physical failure domain, a fault suspends every component at once, and process death interrupts every session the process hosts. This paper shows that neither the modeling nor the calculus binds an agent to one process, the statelessness of the language model keeps all cross-step state outside the model, and the soundness invariant is defined on the state space alone. These observations condense into four lemmas whose premises are the hypotheses of the calculus and the statelessness of language-model inference. On these lemmas this paper constructs Logos, a ROS-like cross process agent harness in which a plugin is a process and the only shared state is an append-only transcript. Eighty sessions resume with no repeated effect after kills placed at the four boundaries of the tool-call cycle, and a same-fault comparison with a single process reference configuration shows one fault interrupting every co-resident session while under the peer-process construction one fault ends at one node.

论文arXiv AI 12:00

InstructMesh: Selective Refinement of Generative 3D Models for Fabrication

arXiv:2608.28534v1 Announce Type: new Abstract: Recent advances in generative AI allow users to create 3D models from text or images. However, these models prioritize visual plausibility over geometric accuracy, often generating results with flaws that compromise their intended use post-fabrication. We present InstructMesh, an interactive post-generation refinement tool that enables selective repair of generative 3D models through region selection and targeted operations, such as opening or sealing voids, or adjusting local thickness. Users can invoke edit operations via natural language prompts or slider controls. By operating directly on the intermediate latent representation, InstructMesh allows users to apply robust geometric corrections without requiring expert modeling skills. To inform our design, we first analyze common fabrication-related failure modes in outputs from state-of-the-art generative tools. We then conduct two user studies, demonstrating that novices can identify and perform fabrication-relevant repairs on generative outputs using InstructMesh, and revealing user preference for hybrid interfaces that combine slider controls with natural language input.

论文arXiv AI 12:00

When Robots Mishear Us: Mapping the Safety Risks of Voice-Controlled Embodied AI

arXiv:2608.28518v1 Announce Type: new Abstract: We investigate whether automatic speech recognition (ASR) errors in user input can lead to unsafe outputs from Embodied AI (EAI) models. We find that ASR errors can lead to harmful instructions being accepted and executed by EAI models, thereby reducing safety. We simulate ASR errors and combine them with existing safety benchmarks (SafeAgentBench and POEX) to evaluate how different errors affect embodied AI safety. We find that some of them preserve semantic structure but increase harmful ambiguity, while others weaken the model refusal behaviour and allow unsafe plans to be generated and executed. We show that in some cases automatic correction of ASR errors can reduce the risk, but this is not always effective. Overall, we show that ASR errors lead to significant safety risks for embodied AI.

论文arXiv AI 12:00

Training Communication-Efficient Mixture-of-Experts Language Models with Layer Re-Configuration

arXiv:2608.28511v1 Announce Type: new Abstract: When training Mixture-of-Experts (MoE) language models with expert parallelism, all-to-all token dispatch and combine collectives can consume a substantial fraction of end-to-end training time. In this work, we study communication-efficient MoE models (CE-MoE), in which we adopt a heterogeneous layer pattern that decouples token-mixing and channel-mixing depth. Compared to conventional models which interleave MoE layers after each token-mixing layer (e.g., attention, Mamba-2), CE-MoE models concentrate expert capacity in a select few routed MoE layers, while maintaining depth by adding additional token-mixing and dense-FFN layers. Across a scaling ladder from 2B to 31.5B total parameters, under matched total and activated parameters, CE-MoE models consistently reduce training cost while matching validation loss and downstream benchmarks with full-MoE baselines. At the 31.5B scale, CE-MoE uses 33.3\% fewer GPU-hours while improving average downstream score and inference throughput.

论文arXiv AI 12:00

AcrossVAM1.0: Particle World Modeling for Text-Assisted Robot Video Prediction

arXiv:2608.28491v1 Announce Type: new Abstract: Predicting robot videos requires both precise motion reasoning and preservation of high-frequency appearance, yet monolithic pixel models entangle these objectives and often conceal their progress behind a strong last-frame baseline. We present AcrossVAM1.0, a lightweight, text-assisted video action model that factorizes future prediction into object-centric motion and dense appearance. A frozen SAM3-DLP codec decomposes four context frames into semantic particles for the robot, arm, and gripper, together with a background latent. A 0.28M-parameter spatio-temporal Transformer aligns particle identities, rolls their states forward, and is modulated by a frozen OpenCLIP instruction embedding through FiLM. A causal dual-stream decoder combines particle-rendered motion with appearance encoded exclusively from the last observed frame; a residual refiner and learned delivery mask produce five future frames without access to future appearance. On our VRS benchmark constructed from diverse real-robot trajectories, particle dynamics reduce trajectory error by 21.0\% over persistence. Across three delivery-mask seeds, AcrossVAM1.0 improves future-frame PSNR/SSIM from 19.97/0.796 to 20.573/0.8004, while raw particle generation improves motion-region PSNR from 11.89 to 13.23. The delivered model does not yet beat persistence in LPIPS, and correct-versus- shuffled language changes trajectory error by only 2.8--3.1%. We report these limitations alongside oracle, negative-control, multi-seed, and per-robot analyses. The results show that explicit particle dynamics are a promising low-dimensional interface for robot video prediction, while robust language grounding and appearance delivery remain the principal open challenges.

论文arXiv AI 12:00

COVER: Identifiable Evaluation of Coalition Routing

arXiv:2608.28475v1 Announce Type: new Abstract: When a multi-agent system changes its team, it also changes the messages and final answer it produces, so an end-to-end accuracy gap does not by itself identify a routing effect. We introduce method, an evaluation contract that fixes a public information boundary, downstream stack G, and finite legal team family before outcomes are generated. Complete coverage identifies exact finite-benchmark oracle regret conditional on that stack. For any finite collection of frozen policies, executing the union of their distinct selected teams is the minimal assumption-free support for every pairwise policy contrast, though not for absolute oracle regret. Two controlled tables with source-ID-disjoint splits test the instrument. On MuSiQue-12, a pre-specified privileged positive control improves regret from 0.532 to 0.402; a later public-interface control reaches 0.424 versus 0.554 but is retrospective. On HotpotQA-4, a pre-specified public direct scorer improves regret from 0.313 to 0.110. In fixed-stack Llama execution, verified route regret improves by 0.190, while the raw-answer gain is 0.010 with an interval crossing zero. A five-family ToolSandbox variant-shift validation exhaustively evaluates 16 declared teams on 14 untouched task variants (224/224 valid rows): the declared-family oracle reaches 0.768 safe-evidence completion, while the prospectively frozen router gets 0.637 (regret 0.131), failing the predeclared 0.10 criterion. A later retrospective comparator reaches 0.655, matching all-workers with 4.57 versus 5.00 workers on average. Thus COVER exposes selection headroom without manufacturing a routing win. A crossed-stack diagnostic shows absolute scores depend on G but finds no detectable router-by-finalizer interaction. COVER is an auditable measurement methodology, not a claim of stack-invariant or universal agent-routing superiority.

论文arXiv AI 12:00

Learning to Use Tools: Reinforcement Learning for Tool-Integrated Mathematical Reasoning

arXiv:2608.28447v1 Announce Type: new Abstract: Current large language models (LLMs) increasingly benefit from external tool integration, especially for tasks requiring reliable computation and verification. Motivated by this, we study calculator tool calling for improving mathematical reasoning on the Countdown task. We first analyze reasoning failures and find that calculation errors account for a substantial portion of incorrect responses. We then construct supervised fine-tuning datasets to teach the model useful tool-use patterns and how to interpret returned outputs. Building on this tool-formatted policy, we apply several on-policy reinforcement learning methods, including RLOO, RLOO++, GRPO, and DAPO, using automatically verifiable final-answer rewards. To enable a more reliable evaluation, we construct a fresh 1,024-problem held-out Countdown benchmark with no exact overlap with the training data. Our results show that calculator tool integration consistently improves both SFT and RL baselines, yielding roughly 10 percentage-point gains across pass@k. Among the RL methods, Tool-DAPO achieves the strongest performance, improving pass@1 from 35.8% for Tool-SFT to 66.0%. Further analysis shows that RL encourages more effective tool use even when only final-answer rewards are provided. These findings suggest that tool integration reduces arithmetic and verification errors, while RL increases the probability of correct reasoning traces.

论文arXiv AI 12:00

Prove2Me: An Open Collaborative Platform for Scaling Math Formalization

arXiv:2608.28433v2 Announce Type: new Abstract: Proof assistants such as Lean 4 promise the paradigm of formally verified mathematics, but large-scale formalization projects have faced major barriers to entry, including the need for expertise in formal verification (as well as the underlying mathematics) and the significant time required for writing formal proofs. AI coding agents have dramatically reduced these barriers; human users can now use natural language to prompt agents to write complex proofs in Lean. This opens up the intriguing possibility of internet-scale mathematical collaboration involving both humans and AI agents, where correctness is machine-checked. To realize this possibility, we introduce Prove2Me (https://prove2.me), an open collaborative platform for formalizing mathematics. Users launch formalization "missions", to which AI agents contribute formal proofs toward completion. We designed mechanisms and a specialized harness in Prove2Me that enable large-scale collaboration so that agents can build on one another's work and freely reuse existing results. In doing so, Prove2Me aims to turn math formalization into a scalable, crowd-sourced effort open to anyone with an agent.

论文arXiv AI 12:00

Program Learning with Verifiable Rewards: Symbolic Backpropagation for Post-Training LLMs

arXiv:2608.28421v1 Announce Type: new Abstract: Post training a language model to reason means updating its weights. Supervised finetuning and reinforcement learning both place the acquired capability inside the model where it cannot be inspected cannot be checked step by step and cannot be moved to another model. We argue that for tasks whose intermediate steps admit verification, reasoning is better placed outside the base models weights as an explicit program composed from deterministic and neural primitives. We introduce PLVR (Program Learning with Verifiable Rewards): a post training method that learns such programs directly from input-output examples. Its mechanism is symbolic backpropagation: each program layer carries a typed ontology a loss is computed at the output against ground truth and required input ontologies are propagated backward by type inference over primitive signatures: an analogue of the chain rule in which credit assignment is a derivation rather than an estimate. Where RLVR verifies a terminal outcome, PLVRs reward is a per step contract verdict dense over program structure. On LiveCodeBench v6 and Tau2Bench, 30B base models with PLVR outperform RL at matched budget by 27.8 points on average and frontier models an order of magnitude larger by 13.6 points. A single primitive library serves two benchmarks, so the marginal cost of a new task is 100 examples of program search and no new finetuning data. Replacing the loss guided search with uniform sampling over the same type admissible space at equal budget collapses the median program from 65.6 to 17.5, identifying the backward pass rather than the type system as the source of the advantage. We release the symbolic backpropagation library and a conformance checker so the method can be applied to primitive libraries other than our own.

论文arXiv AI 12:00

VERA-8B: Evidence-Grounded Audit Risk Reasoning from SEC Filings

arXiv:2608.28402v1 Announce Type: new Abstract: Across audit applications, judgments must be supported by reasonable evidence. However, standard financial language models prioritize fluency over evidence. They are built for general financial reasoning and may produce plausible but ambiguous answers, creating a grounding gap that makes them unsuitable for audit work. We address this gap with VERA-8B, a new end-to-end audit reasoning system that identifies audit risks before enforcement actions occur. Constructing such a model raises several challenges, as no prior machine learning work targets pre-enforcement audit prediction. To our knowledge, we are the first to unify SFT and GRPO for evidence-grounded audit reasoning under one evidence standard, achieving performance that surpasses all evaluated baselines. Because auditing cannot tolerate unsupported claims, we introduce abstention and uncertainty qualification to defer uncertain or evidence-incomplete cases. Finally, we design an AuditBridge to ground model reasoning for practical audit work. It transforms raw filings into verified records and then into reviewer-ready reports, bridging finance and computation with broad generality. Together, these components produce auditable, review-ready outputs suitable for practical audit work.

论文arXiv AI 12:00

RetailAgent: Structured Adverse Timing in Self-Conditioned Multimodal LLM Trading Agents

arXiv:2608.28399v1 Announce Type: new Abstract: In financial markets, a sequential policy that reacts systematically to price movements may become predictable to other market participants. This paper studies whether large language model (LLM) agents exhibit such directional structure through RetailAgent, an experimental framework in which an LLM observes anonymized intraday equity price histories and permitted state, then repeatedly chooses long (hold the stock) or flat (stay out) before the subsequent interval return is revealed. We compare returns during long and flat intervals along the same stock's intraday path after removing the overall fraction of long decisions. This exposure-matched measure reveals persistent negative timing across modality, horizon, state, and model family. Shuffling saved action sequences substantially attenuates the effect, showing that alignment between actions and subsequent returns drives the negative score. Feeding self-authored memories into decisions further increases policy persistence, while timing becomes more negative among stock-days on which the agent uses both actions. These results reveal stable, recoverable directional structure in sequential LLM financial decisions and a behavioral signal for studying how another participant could respond to a predictable policy.

论文arXiv AI 12:00

Timing-Aware Repurchase Prediction for Web-Scale E-Commerce: Survival Models for Multi-Surface Grocery Recommendation

arXiv:2608.28393v2 Announce Type: new Abstract: Repurchase recommenders in e-commerce are commonly framed as a binary question asking "will this customer buy this item within W days", a formulation that requires a separately trained model for every horizon of interest. We replace this stack with survival models that predict time-to-repurchase directly, and evaluate them on millions of customers from a major grocery e-commerce platform across more than thirty ablation configurations. Our study makes three contributions. First, an empirical hazard analysis reveals a slightly decreasing marginal hazard (k ~ 0.9), differing from the common intuition that grocery items become more likely to be repurchased the longer since the last purchase (increasing hazard, k > 1). Log-Normal achieves the best marginal fit (R^2 = 0.998) and the best ranking, despite Weibull providing the best conditional residual fit, revealing an apparent discrepancy we analyze in detail. Second, a single Accelerated Failure Time (AFT) model replaces three per-horizon binary classifiers, matching or exceeding each at its own horizon while using roughly 3x fewer total trees. Feature importance reshuffles under the survival objective: channel-cadence and recency signals rise while aggregate frequency counts fall. Third, a 4-parameter parametric calibration maps raw survival CDFs to per-horizon probabilities with zero cross-horizon monotonicity violations. Calibration quality varies by an order of magnitude across the AFT family: Exponential AFT (Weibull k=1) achieves expected calibration error (ECE) ~1e-4, roughly 10x lower than Log-Normal, while ranking metrics agree within 0.3% relative. We adopt Exponential AFT for probability-consuming surfaces and Log-Normal for pure ranking, exposing a principled calibration-ranking trade-off within a single AFT family.

论文arXiv AI 12:00

MAP: A Benchmark on Multimodal Accessibility Planning for Real World Places

arXiv:2608.28384v1 Announce Type: new Abstract: We introduce MAP, the first benchmark to evaluate multimodal AI systems as assistants for users with accessibility requirements when planning visits to places in the real world. In our evaluation, systems are presented with requests to verify or recommend a point of interest meeting an accessibility requirement. MAP contains two novel assessments: Claim verification for accessibility planning assesses if information on places and stated accessibility features is supported and identifies places that satisfy requested accessibility features. Visual evidence retrieval for accessibility planning checks if a multimodal AI system can select visual evidence for the requested place and accessibility feature. Our methodology supports comparison of AI systems in a setting where place information and accessibility information can change over time by evaluating systems and refreshing ground truth data at scheduled times. The benchmark is based on automatic rating and human rating for a proportion of responses.

论文arXiv AI 12:00

EvoUndo: Recoverability-Constrained Self-Evolution for LLM Agent Harnesses

arXiv:2608.28363v1 Announce Type: new Abstract: LLM agents increasingly modify their own prompts, tools, middleware, resources, and execution harnesses at runtime. Such self-evolution can improve capability, but a successful mutation may leave persistent effects that cannot be safely reversed in states different from the one in which it was created. We introduce EvoUndo, a framework for representing, synthesizing, diagnosing, and independently verifying recoverability of model-generated self-modifications across counterfactual states. Across 600 unseen one-shot self-evolution tasks, we identify 197 capability-improving mutations that fail recoverability verification. Under the original recovery representation, conventional repair strategies recover 0/197 of these natural failures. Deterministic oracle analysis recovers 48/197 under the original recovery language L0, while the extended recovery calculus increases empirical oracle recovery to 191/197. A protocol-locked 2x2 grounding-by-expressivity intervention then separates two bottlenecks: exact state-address grounding increases successful recovery from 0/48 to 38/48 (79.2%) when the original language is sufficient, while extending the recovery language enables recovery on 142/143 (99.3%) failures in the oracle-defined S1 stratum. On the primary gpt-oss-120b backbone, adding exact-address diagnostics to the richer language reduces recovery to 133/143 (93.0%); a Qwen3.8-27B replication preserves the grounding and expressivity effects but not this negative interaction, indicating that the latter is model-dependent. These results indicate that reliable agent self-evolution requires co-designing verification, state grounding, witness semantics, and recovery-language expressivity rather than relying on iterative prompting alone.

论文arXiv AI 12:00

GRACE:Gradient-guided Coreset Selection for LLM Unlearning

arXiv:2608.28361v1 Announce Type: new Abstract: Machine Unlearning methods for Large Language Models typically assume pre-specified forget and retain sets. In realistic settings, however, requests may provide only a few examples of undesired behavior, requiring forget and retain sets to be inferred from heterogeneous corpora. We study this data-selection problem and propose GRACE , a gradient-guided coreset selection method that constructs both forget and retain sets for LLM unlearning. GRACE first computes a forget direction from seed examples that elicit the undesired behavior, then selects a compact forget coreset whose gradients approximate this direction using non-negative orthogonal matching pursuit. To preserve model utility, it selects retain examples after projecting out the forget direction and applying clustered orthogonal matching pursuit in the remaining gradient space. Across two target domains, two model families, and four unlearning algorithms, GRACE improves model utility while maintaining comparable forget quality, with particularly consistent gains over prior gradient-based selection methods.

论文arXiv AI 12:00

Propagating construction-time knowledge quality into medical question answering: A framework grounded in clinical guidelines

arXiv:2608.28360v1 Announce Type: new Abstract: Large language models have facilitated knowledge graph (KG) construction from clinical guidelines, but extracted triples vary in structural validity and evidential support. Meanwhile, graph-augmented question answering (QA) systems typically optimize query relevance during retrieval, with limited reuse of quality information produced during KG construction. This creates a disconnect between construction-time quality control and inference-time evidence use. We investigate whether construction-time triple quality can serve as a persistent signal for downstream evidence selection and presentation. We propose a quality-aware framework that models structural conformance (SchemaConf) and evidential support (EvidScore) as complementary dimensions and fuses them into a per-triple quality signal, Q(t). Rather than using quality solely for filtering, the framework retains Q(t) and derived quality tiers as graph attributes and propagates them into quality-weighted subgraph retrieval and tier-conditioned evidence prompting, while preserving passage-level provenance. Experiments on Chinese diabetes clinical guidelines show that the utility of the quality signal is distribution dependent. Under cross-version and cross-model shift, the fused Q(t) provides stronger triple-quality discrimination than either component alone (AUC 0.748 vs. 0.703 for EvidScore and 0.645 for SchemaConf). In guideline-grounded QA, propagating construction-time quality reduces required-knowledge omission from 16.3% to 5.3% and conflicting outputs from 16.3% to 2.7%, with an evidence-grounded precision of 81.6% and near-zero invalid citations. Blinded clinician ratings favor the full framework over no retrieval (4.68 vs. 4.21 on a five-point scale) and approach the oracle condition (4.80), while cross-generator experiments show consistent trends.

论文arXiv AI 12:00

AGENT-O: A Semantic Agent Card Framework for Interoperable and Governed Healthcare AI Agents

arXiv:2608.28345v1 Announce Type: new Abstract: AGENT-O is a modular ontology framework that defines a semantic Agent Card for representing health-oriented AI agent systems and supports assessment of reporting completeness in scientific publications. AGENT-O was developed as an OWL 2/RDF ontology covering runtime, models, workflow, tools, clinical use, evaluation, provenance, governance, and reporting assessment. Evaluation included ontology inventory, OWL-RL reasoning, three SHACL suites, 12 SPARQL competency queries, three cases, and model-assisted reporting-completeness assessment of 279 papers across five dimensions. The ontology contained 1,962 RDF triples and 1,922 Protege axioms, with 252 active classes, 198 active object properties, and 51 datatype properties. All SHACL suites conformed on example graphs, all competency queries returned prespecified evidence, and all 279 papers were scored. Incomplete reporting was highest for runtime/architecture (84.6%), governance/safety (82.8%), and provenance/reproducibility (78.1%), compared with evaluation (25.8%) and benchmark-process alignment (29.8%). AGENT-O supported semantic Agent Card representation and reporting assessment while revealing an evaluation-specification gap: evaluation and benchmark procedures were reported more consistently than runtime architecture, governance, and reproducibility. AGENT-O provides a reusable ontology, semantic Agent Card profile, and reporting-completeness workflow for structured reporting and gap identification, but does not assess agent quality or deployment readiness.

论文arXiv AI 12:00

Real-Valued Hyperdimensional Sequence Representations with Hadamard Product Binding and Shift Equivariance

arXiv:2608.28334v1 Announce Type: new Abstract: Encoding temporal order is a fundamental requirement for sequence representations in Hyperdimensional Computing. Fractional Power Encoding provides similarity-preserving position vectors whose inner products approximate shift-invariant kernels, and it supports shift-equivariant transformations of encoded sequence representations. However, standard formulations of Fractional Power Encoding are primarily designed for binding operations such as circular convolution or complex-valued multiplication, which limits their compatibility with Hadamard product binding of real-valued vectors. This paper develops real-valued position encodings motivated by Random Fourier Features, aiming to retain the desirable properties of Fractional Power Encoding while supporting Hadamard-based operations. We propose three real-valued position-encoding variants: a real-valued baseline based on the inverse Fourier transform, and Sinusoid and Cosine-only representations derived from Random Fourier Features. Among them, the Sinusoid variant provides an explicit algebraic shift operator, allowing temporal shifts to be applied directly to the vector-encoded sequence representation without re-encoding the shifted sequence. Experiments on time-series classification datasets show that the proposed real-valued representations achieve performance comparable to standard Fractional Power Encoding while enabling computationally efficient Hadamard product binding. The Sinusoid variant offers the most favorable trade-off, combining efficient real-valued implementation with exact shift-equivariant transformations.

论文arXiv AI 12:00

MAIL: Memory-driven, Adaptive, Incremental, and Literature-grounded Framework for Hypothesis Generation in Chemistry

arXiv:2608.28315v1 Announce Type: new Abstract: The ever-expanding volume of the chemical literature offers unprecedented opportunities to generate novel and impactful hypotheses. However, the bottleneck lies in efficiently navigating this vast knowledge base to formulate high-quality, experimentally meaningful insights. While Large Language Models (LLMs) show promise for this task, existing methods often rely on static inspiration corpora, predefined heuristics, or laborious human-in-the-loop pipelines and decision-support frameworks that limit scalability and novelty. In this work, we propose an automated approach, a Memory-augmented, Adaptive, Incremental, and Literature-grounded (MAIL) framework for hypothesis generation in chemistry. Our MAIL method formulates hypothesis generation as a temporally grounded, memory-driven reasoning process, where hypotheses emerge from an evolving conceptual path that continuously accumulates and reinterprets prior knowledge. We evaluated the MAIL framework on a public TOMATO-Chem dataset and a newly curated and disseminated high-novelty nature/science challenge (HN-NS) dataset. Across both datasets, MAIL generates structurally coherent and mechanistically plausible hypotheses, achieves the highest MIOS and MPOS by more effectively recovering the central ideas and methodological elements of the historical target hypotheses, and obtains the highest overall expert-evaluation scores for scientific quality. These results demonstrate the potential of LLMs to autonomously explore chemical domains and generate hypotheses that are both innovative and chemically plausible.

数据自动抓取 · 返回首页