AI EarlyView
Pro
AI EarlyView
← 最新论文日报
8月6日· 周四

arXiv 论文日报

智能体进化与视频编辑新突破

今日论文聚焦两大热点:智能体从静态图像走向视频深度研究,并通过相似性推断实现理性合作;同时,视频编辑迎来实时开放式新框架,扩散模型偏好对齐亦有新机制。多模态大模型的并行扩展与音乐生成的质量洞察同样值得关注。

📄 10 篇精选</> 7 篇带代码🏆 3 篇 SOTA
📈 趋势速览智能体(agents)与强化学习(reinforcement)持续升温,多模态智能体向视频流扩展,自我改进与理性合作成为新方向。视频生成与编辑(visual, generation)追求实时与开放,扩散模型优化聚焦偏好对齐。大模型评测与基准(llm, llms)关注递归自我改进与工具推理。
1
💡 解决扩散模型多步去噪的信用分配难题,提升对齐效率。
强化学习Yuanshen Guan、Zipeng Feng、Zhiwei Xiong
2
Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent
💡 将深度研究智能体扩展至视频流,识别关键瓶颈。
智能体Zhen Fang、Yu Zeng、Wenxuan Huang
3
When Attention Goes Blind: Numerical Failure in ALiBi Positional Encodings
💡 揭示ALiBi位置编码的数值失效,影响长上下文模型可靠性。
NLPChristopher Schröder、Lukas Gienapp、Ferdinand Schlatt
4
JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion
💡 16B参数实现实时开放式视频编辑,无需预设时长。
生成扩散Yicheng Xiao、Wenxun Dai、Xinran Qin
5
💡 时间序列转图像新方法,异常检测性能优于基线。
视觉Mateusz Smendowski、Kamil Faber、Piotr Nawrocki
6
Agogic: Performance-Timed Music Tokens for LLM-Native Text-to-Symbolic-Music Generation
💡 发现音乐分词方式比模型规模更关键,提出新音乐令牌。
语音音频Junhao Chen、Mingjin Chen、Jingjia Mao
7
A game theory for foundation models shows new paths to rational cooperation through similarity inference
💡 基础模型通过相似性推断实现理性合作,开辟AI博弈新路径。
智能体Alexander Meulemans、Maciej Wołczyk、Marissa A. Weis
8
ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs
💡 并行计算复用ViT和LLM参数,优化多模态大模型算力分配。
大模型Yang Yang、Qinyu Zhao、Mouxiang Chen
9
TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning
💡 回合级事后自我蒸馏,提升工具推理的细粒度奖励分配。
大模型Changle Qu、Sunhao Dai、Hengyi Cai
10
PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents
💡 新基准测试代理递归自我改进能力,推动智能体进化研究。
智能体Shuhan Xue、Zixin Ding、Yichen Shen

生成于 07:00 · 来源 arXiv · AI 编排

往期日报