1 B站 Lau博士的云组会 reach 100
梁圣带队发布V4版本,全面解析DSpark论文核心创新与性能提升。
2 Reddit r/unsloth 14:25 reach 100
DeepSeek releases DSpark - 50%-600% faster spec decoding vs MTP
DeepSeek发布DSpark,推理速度比MTP快50%-600%。
3 推特 danielhanchen 14:10 reach 100
DeepSeek just released DSpark for V4 Flash & Pro, a new speculative decoding
DeepSeek发布DSpark推测解码方法,吞吐量提升51%至400%。
4 小红书 量子位 08:00 reach 100
Claude Mythos开始自创语言,引发AI安全担忧。
5 国内 钛媒体 3 天前 cn 92
张一鸣罕见表态,字节AI战略全面升级,强调不蒸馏、练内功。
6 国内 雷锋网 3 天前 cn 92
阿里开源2.4T参数MoE模型Qwen3.8,智源FlagOS实现9款芯片Day0适配。
7 国内 量子位 3 天前 cn 92
Ilya新公司SSI首个模型曝光,聚焦持续学习能力。
8 国内 钛媒体 3 天前 cn 92
DeepSeek以0.1分性能差和60倍价格差挑战硅谷,不甘平替,双线出击。
9 国内 爱范儿 3 天前 cn 92
国产机器人以低成本实现具身智能突破,反超Figure AI,迎来DeepSeek时刻。
10 国内 量子位 3 天前 cn 88
Claude解决2000阶以下哈达玛矩阵难题,AI数学能力再突破。
11 国内 雷锋网 3 天前 cn 88
DeepSeek V4 Pro发布,性能逼近Fable 5,价格仅1/57,聚焦Agent基建。
12 国内 钛媒体 3 天前 cn 88
OTA流量模式将衰,豆包等AI对话产品正重塑信息入口,内容分发逻辑剧变。
13 海外 NVIDIA 4 天前 行业 92
NVIDIA AI Factory Compute Is Becoming an Investable Asset Class
英伟达联合多家金融机构,拟撬动超5000亿美元第三方资本,将AI工厂算力打造为可投资资产类别。
14 国内 钛媒体 3 天前 cn 88
北大等发布世界动作模型ω-0,实现机器人全身协同居家作业。
15 国内 钛媒体 3 天前 cn 88
DeepSeek-V4-Pro实测,涨价后性价比仍高。
16 国内 InfoQ 中国 2 天前 cn 85
AI开源从模型转向生态竞争,开放战略成新焦点。
17 arXiv arXiv 5 天前 研究 92
Towards Expert-level Medical AI for Real-time Video Consultations
首个达到专家级水平的实时视频问诊AI,通过多模态交互提升诊断能力。
18 国内 雷锋网 3 天前 2 家在报道 cn 87
《知识就是力量》携手360发布科普科幻AI大片创作平台,推动科普内容AI视频化生产。
19 海外 The Decoder 3 天前 模型 88
SpaceXAI's Grok 4.6 matches OpenAI's best model and undercuts it on price
Grok 4.6性能追平GPT-5.6,价格低60%,代理任务效率翻倍。
20 海外 TechCrunch AI 3 天前 行业 88
AI coding startup Cognition reportedly already in talks to raise at $40B valuation
AI编程公司Cognition据报正洽谈以400亿美元估值融资,距上次260亿美元融资仅数月。
21 海外 TechCrunch AI 3 天前 会议 88
As AI safety concerns mount, three pioneers make the case for staying open
三位AI先驱在Ai4大会呼吁保持开放,辩论监管与开源竞争。
22 海外 Ars Technica AI 3 天前 产品 85
Claude's new Scarlet Letter watermark is invisible—for now
Claude新水印技术可标记AI处理内容,目前不可见但未来或可检测。
23 海外 TechCrunch AI 3 天前 产品 88
Everything announced at Made by Google ’26: Pixel 11, Pixel Watch 5, Pixel Tag, and tons of Gemini features
谷歌2026发布会:Pixel 11、手表5、追踪器及大量Gemini功能。
24 国内 量子位 3 天前 cn 85
杭州发布全球首款站姿载人飞行器,中国人能飞了。
25 国内 量子位 3 天前 cn 88
国产具身智能创纪录,低成本高效分拣包裹。
26 国内 雷锋网 3 天前 cn 85
完成Modular收购,高通瞄准数据中心、基础设施、个人及工业AI
高通完成对AI软件公司Modular的收购,强化端到云AI计算平台能力。
27 国内 钛媒体 3 天前 cn 85
腾讯Q2算力开支528亿,模型优先,出租算力是退路但空间有限。
28 国内 钛媒体 3 天前 cn 85
腾讯Q2财报显示宁可现金流为负也要锁定算力,背后是对AI机会的笃定投入。
29 国内 雷锋网 3 天前 cn 85
中国大厂消失在赞助商名单,却在不莱梅重构 AI 的灵魂丨IJCAI 2026
IJCAI 2026将在德国不莱梅举行,中国AI力量仍是主角,大厂转向逻辑攻坚。
30 国内 InfoQ 中国 4 天前 cn 88
扎克伯格万字长文力挺开源,宣布Meta重回开源模型路线。
31 国内 雷锋网 4 天前 cn 88
AI机器人流量超人类,互联网从“给人看”转向“给Agent读”,字节谷歌流量税模式受冲击。
32 国内 钛媒体 3 天前 cn 85
腾讯AI战略转向,本季度利润让位算力投入。
33 国内 钛媒体 3 天前 cn 85
AI投资回报困境源于组织与客户旅程错配,而非技术差距。
34 国内 钛媒体 4 天前 cn 88
AI正通过能力替代重塑国家经济,印度或成首个被数字时代“做空”的大型经济体。
35 国内 量子位 3 天前 cn 85
联想Q1营收1834亿元创新高,AI服务器业务爆发。
36 国内 钛媒体 3 天前 cn 85
腾讯AI投入激进,微信或借AI重构社交体验,引发行业关注。
37 国内 量子位 3 天前 cn 85
2026世界机器人大会主论坛议程公布,聚焦前沿技术与产业趋势。
38 国内 钛媒体 4 天前 cn 88
中美AI路径分岔:美押注算力资产化,中国主导人形机器人,竞争分化。
39 国内 雷锋网 3 天前 cn 85
十年后,影石重新发明了全景相机
影石发布全景相机X6,十年后重新定义品类,探索全景相机下一个十年。
40 国内 钛媒体 3 天前 cn 85
儿童动画被AI邪典内容污染,需警惕流量背后的危害。
41 国内 钛媒体 3 天前 cn 85
AI走向物理世界,AIDC算力底座是数据飞轮第一推动力。
42 国内 钛媒体 3 天前 cn 85
腾讯Q2财报:AI投入528亿,利润表承压,市场关注回报周期。
43 国内 InfoQ 中国 2 天前 cn 82
微软AI Gateway新层级引发权限治理隐忧讨论。
44 国内 钛媒体 3 天前 cn 85
Edge AI Daily 早报(8月13日)
Twitch默认用内容训练AI引争议,美国推硅走廊,SpaceX AI收入将超主业等AI产业动态。
45 arXiv arXiv 4 天前 研究 88
When Visual Signals Mislead: A Mechanistic Study of Attribute Hallucination in Vision-Language Models
揭示VLM属性幻觉源于视觉信号不足而非语言先验,提出VISOR框架诊断与修复。
46 arXiv arXiv 4 天前 研究 88
Once Poisoned, Arbitrarily Controlled: A Programmable Backdoor in VLMs
提出VLM可编程后门,推理时可任意指定目标并生成触发器,颠覆静态攻击假设。
47 海外 Ars Technica AI 3 天前 行业 85
Terabytes of credentials leaked in massive supply-chain attack
AI包供应链攻击致2500用户凭据泄露,数据达TB级。
48 arXiv arXiv 5 天前 研究 88
DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?
首个在真实计算机环境中评估智能体端到端数据科学工作流的基准。
49 海外 TechCrunch AI 3 天前 行业 85
OpenAI-backed Thrive Holdings raises $2B to bring AI to the enterprise
OpenAI支持的Thrive Holdings融资20亿美元,估值达120亿,加速企业AI落地。
50 海外 The Decoder 3 天前 研究 85
Researchers can now reverse-engineer LLM prompts from output text with near-perfect accuracy
新方法可近乎完美地从LLM输出反推原始提示词,构成安全风险。
51 国内 InfoQ 中国 3 天前 cn 85
Vercel发布新语言Zero,代码面向AI而非人类,引发行业热议。
52 国内 InfoQ 中国 3 天前 cn 85
AI算力分配应差异化,资深工程师优先,新人刷题式成长已失效。
53 海外 MIT Tech Review 3 天前 行业 85
Scaling AI agents with trustworthy data
AI代理规模化落地关键在于可信数据基础设施,而非模型本身。
54 国内 爱范儿 2 天前 cn 82
DeepSeek Harness首发体验,主打插件化,不走Codex老路。
55 海外 TechCrunch AI 3 天前 行业 85
Lovable confirms new $13.3B valuation, raises another $400M
Lovable估值达133亿美元,再融4亿美元,年化收入5亿美元。
56 arXiv arXiv 5 天前 研究 88
SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring
新基准SWE-Bench ProMax评估AI智能体大规模多语言代码重构能力,解决现有基准饱和与测试缺陷问题。
57 海外 Hacker News 3 天前 行业 85
Someone is running mass vulnerability scans, spoofing AI bots like ClaudeBot
有人伪装成AI爬虫进行大规模漏洞扫描,引发安全担忧。
58 一石一泉一松一月一人 + 关注 3 天前 行业 85
悼念朱镕基先生,回顾其改革担当与人民情怀。
59 海外 NVIDIA 3 天前 行业 85
NVIDIA CEO Tops Glassdoor’s 2026 List of Best CEOs
黄仁勋获Glassdoor 2026最佳CEO榜首,员工支持率99%。
60 国内 InfoQ 中国 3 天前 cn 82
提出兼顾生产稳定与快速迭代的AI工作流模式,强调运行时无关设计。
61 海外 The Decoder 3 天前 行业 82
Fable 5's slow adoption suggests corporate willingness to pay for frontier AI has hit a ceiling
Fable 5虽强但企业购买少,企业AI支出或触顶。
62 海外 The Decoder 3 天前 研究 82
Top AI lab researchers warned about automated AI research, and several of their predicted milestones have already fallen
访谈25位AI研究员,预测的自动化研究里程碑已部分实现。
63 国内 雷锋网 3 天前 cn 82
滴滴二季度订单量增13.2%,中国出行连续14季上涨,国际业务强劲。
64 海外 The Decoder 3 天前 产品 82
Anthropic brings Claude Cowork to its Chrome extension, adding skills and plugins to the browser
Anthropic将Claude Cowork集成到Chrome扩展侧边栏,支持技能与插件。
65 国内 量子位 3 天前 cn 82
科大讯飞发布覆盖七大核心场景的企业服务全系列产品,强化AI落地。
66 arXiv arXiv 6 天前 研究 88
SLAC: Access-Driven CPU-to-GPU Side-channel Attacks via System-Level Cache on Apple Silicon
揭示苹果芯片CPU可经系统缓存侧信道攻击GPU,跨域泄露数据。
67 国内 雷锋网 3 天前 cn 82
金山办公灵犀专业版首批接入DeepSeek-V4-Pro,提升复杂办公任务交付能力。
68 海外 TechCrunch AI 4 天前 行业 85
AI code-testing startup Blacksmith’s valuation jumps almost 10x in less than a year
AI代码测试初创Blacksmith估值一年涨近10倍,营收增超十倍。
69 国内 量子位 4 天前 cn 85
紫东太初提出GMC剪枝法,减少80%Token仍保持多模态能力,免训练即用。
70 国内 雷锋网 3 天前 cn 82
黑盒条件下对导航智能体做系统性安全验证,揭示具身智能潜在风险。
71 国内 钛媒体 4 天前 cn 85
豆包收12%佣金,GEO成酒店获客新渠道,早入局者已获利。
72 国内 钛媒体 4 天前 cn 85
林俊旸创办AI公司Pragmatik Labs,估值20亿美元,获红杉腾讯投资。
73 海外 Google Research 4 天前 研究 85
Empty shelves or lost keys? Recall is the bottleneck for parametric factuality
探讨AI参数化事实性瓶颈,指出记忆检索比生成更关键。
74 国内 雷锋网 4 天前 cn 85
美团CEO称不做线下药店,专注用AI帮医药商家转型。
75 国内 钛媒体 4 天前 cn 85
OpenAI IPO前夕关键高管离职,引发市场关注。
76 国内 钛媒体 3 天前 cn 82
小鹏换造车方式,需向资本市场证明AI能力进阶伴随成本下降与规模效应。
77 国内 雷锋网 3 天前 cn 82
Maker Tool出海赛道规模超百亿美元,3D打印机等品类增长迅猛,但面临认知滞后与竞争暗礁。
78 海外 Hacker News 4 天前 产品 85
Launch HN: Discovered Materials (YC P26) – AI agents to discover new materials
AI代理发现半导体新材料,解决GPU散热难题。
79 国内 爱范儿 3 天前 cn 82
WorkBuddy用AI重做办公三件套,展示AI时代新Office形态。
80 国内 爱范儿 4 天前 cn 85
苹果iOS 27或推AI收费服务,用户需付费解锁更智能功能。
81 海外 MarkTechPost 4 天前 模型 85
NVIDIA AI Releases Nemotron 3.5 Lightning: A 30B Open MoE with 3B Active Parameters, and NeMo Switchyard Model Router
NVIDIA发布30B开源MoE模型Nemotron 3.5 Lightning及Switchyard路由工具,主打高效智能体执行。
82 国内 雷锋网 4 天前 cn 85
Vibe Coding催生海量AI应用,数据库需应对动态Schema新挑战。
83 国内 钛媒体 4 天前 cn 85
百丽时尚详述企业AI落地实践,从协同在线到AI原生的转型路径。
84 arXiv arXiv 04:22 研究 88
Quantization Damage Is Multiplicative, Not Additive
量化误差是乘法性而非加法性,低比特下模型决策会静默受损,基准分数却几乎不变。
85 国内 钛媒体 4 天前 cn 85
AI将颠覆广告业,实现消费者主权与决策革命。
86 arXiv arXiv 00:41 研究 88
SparseDitto: Customizing GPU Kernels for Different Sparsity Patterns with LLM-Based Agentic System
提出用LLM智能体自动定制GPU稀疏矩阵内核,适应不同稀疏模式,性能远超cuSPARSE。
87 国内 雷锋网 4 天前 cn 85
REDMI发布K100 Pro系列,双芯+185Hz屏+长焦影像,主打性能越阶。
88 国内 钛媒体 4 天前 cn 85
字节成立大模型一级部门,张一鸣要求自研不抄作业。
89 arXiv arXiv 11:00 研究 88
RING: Retrieval-Internalized Generation for Continual Large-Scale Knowledge Injection
RING将检索内化进模型参数,用强化学习实现无外部检索器的知识注入新范式。
90 国内 雷锋网 3 天前 cn 82
具身智能落地卡点被低估,本体硬件重要性常被忽视。
91 国内 钛媒体 4 天前 cn 85
面壁智能半年融资超50亿、估值破200亿,成端侧AI独角兽,但商业验证仍待观察。
92 海外 Hacker News 4 天前 行业 85
Company Offering '100% Human-Written, Never AI' Medical Research Is 100% AI
公司宣称纯人工医学研究实为AI生成,被揭穿。
93 国内 钛媒体 4 天前 cn 85
英伟达发布Nemotron 3.5与模型路由器,企业AI推理提速30%。
94 国内 钛媒体 3 天前 cn 82
具身智能公司融资后烧钱困境,行业反思花钱方式。
95 国内 钛媒体 4 天前 cn 85
Edge AI Daily 早报(8月12日)
Anthropic水印应对欧盟法案,Pathway低成本突破ARC-AGI,红杉同时投OpenAI和Anthropic等AI行业动态。
96 海外 Simon Willison 4 天前 实践 85
There are no lossless transformations of natural-language text
工程师使用AI写作的内部政策:需对每个观点和句子负责。
97 国内 雷锋网 3 天前 cn 82
小马智行发布第四代无人重卡,计划未来三年运营千辆智驾重卡。
98 海外 Simon Willison 4 天前 研究 85
Stealing Reasoning Traces from Proprietary LLM APIs
研究揭示可通过重放加密推理链攻击专有LLM,恢复隐藏推理内容。
99 国内 雷锋网 3 天前 cn 82
AI原生达人营销平台AhaCreator集成飞书,用机器人“铁墩儿”将海外达人合作流程嵌入企业协作流,提升决策效率。
100 国内 雷锋网 3 天前 cn 82
00后创业者毛榉从秦岭徒步受伤经历出发,打造无动力外骨骼wudaX hik 1,探索新路线。
101 国内 钛媒体 3 天前 cn 82
人形机器人续航痛点及产业破局思路
102 海外 AWS ML 4 天前 产品 85
Accelerate cyber defense with OpenAI and AWS: Daybreak Red & Daybreak Blue now available to eligible customers on Amazon Bedrock
OpenAI两款网络安全模型Daybreak Red/Blue上线AWS Bedrock,芯片级零操作员访问保障数据安全。
103 arXiv arXiv 4 天前 研究 85
AdvFD: Boosting Visual Generation via Adversarial Fr'echet Distance Loss
提出对抗式弗雷歇距离损失,解决生成模型训练中的弗雷歇黑客问题,提升视觉质量。
104 arXiv arXiv 4 天前 研究 85
VIScore: Diagnosing Planning-Relevant Quality in Latent World Models
提出VIScore指标,诊断潜在世界模型中与规划相关的表征质量,连接潜空间属性与规划性能。
105 arXiv arXiv 4 天前 研究 85
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop
TrustNLP研讨会六届144篇论文,总结NLP信任研究从可解释性转向生成式AI主动控制。
106 arXiv arXiv 4 天前 研究 85
The Illusion of Cross-Lingual Safety in Low-Resource Languages
研究低资源语言下LLM安全对齐的跨语言迁移失效问题,提出新数据集与几何探测方法。
107 arXiv arXiv 4 天前 研究 85
Attention-Path Fragility as an Uncertainty Signal in Large Language Models
提出用注意力路径脆弱性估计大模型不确定性,无需训练即可提升问答可靠性。
108 arXiv arXiv 4 天前 研究 85
Is There Really a Camouflaged Object? Towards Realistic Camouflaged Object Detection
提出OPC16K基准,解决现实场景中伪装目标检测的误报问题。
109 国内 雷锋网 3 天前 cn 82
有人称中签宇树不敢发朋友圈:怕被嫉妒;DeepSeek V4 Pro正式版上线;美国政府设备重新允许使用TikTok!特朗普:我在TikTok一直霸榜第一
美国政府解禁TikTok,DeepSeek V4 Pro上线,宇树中签引热议。
110 arXiv arXiv 4 天前 研究 85
Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents
研究跨语言工具使用智能体的行为一致性,提出动作策略保留度的测量方法。
111 国内 钛媒体 3 天前 cn 82
Chinese AI Chatbots Begin Charging for the Transactions They Generate
阿里与字节同日调整AI聊天机器人收费策略,转向交易抽佣模式。
112 arXiv arXiv 4 天前 研究 85
提出抗丢包图像压缩方案,解决卫星通信中数据丢失问题。
113 arXiv arXiv 4 天前 研究 85
Why Does CLAUDE.md Keep Growing? Catastrophic Remembering in Agentic Coding
研究发现AI编程中CLAUDE.md等指令文件无限膨胀,源于删除指令的高昂成本,称为“灾难性记忆”。
114 arXiv arXiv 4 天前 研究 85
Cross-View Feature Matching: Survey, Benchmarking, and Foundation-Model Perspectives
综述跨视角特征匹配,涵盖基准测试与基础模型视角,梳理领域进展与挑战。
115 arXiv arXiv 4 天前 研究 85
CapProbe: Evaluating Detailed Image Captions via Full-Scene Dense Question Answering
CapProbe通过全场景密集问答评估图像描述,实现区域对齐的事实核查。
116 arXiv arXiv 4 天前 研究 85
Quantum Coordination Advantages in AI State-Tracking Tasks: Semantic Compilation and Latent Memory
证明AI状态追踪任务中量子协调优势,提出语义编译定理。
117 arXiv arXiv 4 天前 研究 85
Efficient Hypergradient Descent for Inverse Reinforcement Learning
提出高效超梯度下降法,解决逆强化学习双层优化的计算难题。
118 arXiv arXiv 4 天前 研究 85
MAJEPPA: Morphing and Assessing in a Unified Piano Performance Space
提出MAJEPPA框架,用自监督学习统一钢琴演奏表征,覆盖从新手到大师的完整技能谱系。
119 arXiv arXiv 4 天前 研究 85
Data Attribution of Emergent Misalignment with Persona Features
研究发现,微调模型时的有害行为源于预训练中习得的特定人格特征,并可通过特定文本激活。
120 arXiv arXiv 4 天前 研究 85
R4DSG: Relative 4D Scene Graph Memory for Object-Centric Question Answering in Long Egocentric Video
提出相对4D场景图记忆,提升长视频中物体中心问答的准确率。
121 arXiv arXiv 4 天前 研究 85
On the Limitations of Cross-Lingual Consistency in Multilingual Text-to-image Generation
提出多语言文生图基准LingT2I,揭示跨语言生成中的性能差距与权衡。
122 arXiv arXiv 4 天前 研究 85
What Iterated Self-Feeding Probes of Language Models Measure, and a test that separates the construction from the model
研究语言模型自反馈探测的测量对象,提出区分构造与模型的新测试。
123 arXiv arXiv 4 天前 研究 85
PEAK: Precise and Persistent Concept Erasure via k-Sparse Autoencoders
提出PEAK方法,用k稀疏自编码器精准且持久地消除T2I扩散模型中的概念。
124 arXiv arXiv 4 天前 研究 85
XCoT-VLA: Executable Chain-of-Thought for Vision-Language-Action Driving
提出XCoT-VLA,用可执行思维链替代冗长自然语言,提升自动驾驶实时控制效率。
125 arXiv arXiv 4 天前 研究 85
Evidence-Grounded Trustworthy Multimodal Reasoning and Evaluation Benchmark in Complex Urban Scenes
提出AD2-Bench基准,诊断多模态模型在复杂城市场景中的推理可靠性。
126 arXiv arXiv 4 天前 研究 85
Multiple Scale Latents for Learned Image Compression
提出多尺度潜变量分层表示,提升图像压缩熵模型效率,较VVC降低17.9%码率。
127 arXiv arXiv 4 天前 研究 85
StreamFlow: Dynamic Memory Flows for Streaming Video Understanding
提出StreamFlow框架,实现流式视频理解中视觉记忆的动态按需访问,提升效率与准确性。
128 arXiv arXiv 4 天前 研究 85
GS-CPE: Unified 6-Degree-of-Freedom Camera Pose Estimation via 3D Gaussian Splatting
提出GS-CPE框架,结合几何粗估计与3DGS精化,实现高精度6自由度相机定位。
129 arXiv arXiv 4 天前 研究 85
ThinkRetrieve: Retrieval-Augmented Reasoning Traces for Test-Time Scaling
提出ThinkRetrieve框架,通过动态检索解题示例增强推理轨迹,提升测试时扩展效率。
130 arXiv arXiv 4 天前 研究 85
IO Factory: Simulating AI-Enabled Influence Campaigns at Scale
提出IO Factory框架,模拟AI驱动的规模化信息影响活动,以分析AI群体操纵行为。
131 arXiv arXiv 4 天前 研究 85
Enhanced Filtering Algorithms for the Euclidean Traveling Salesperson Problem and its variants in Constraint Logic Programming
提出增强过滤算法,利用几何信息高效求解欧几里得TSP及其变体。
132 arXiv arXiv 4 天前 研究 85
Robust Safety Filtering for Input-Constrained Underactuated Linear Systems
提出针对欠驱动线性系统的鲁棒安全滤波框架,结合H∞控制与扰动观测器,确保约束下前向不变性。
133 arXiv arXiv 5 天前 研究 85
Flex-$π$: A Multi-Stream World-Action Model with Compute Flexibility
Flex-π模型利用冻结VAE免费获得3D几何与语义监督,提升世界动作模型性能。
134 arXiv arXiv 5 天前 研究 85
UniProbe: A Learnable Token-Level Hallucination Detector for Large VLMs using Multi-Structural Internal Representations
提出UniProbe,用多结构内部表示实现LVLM令牌级幻觉检测,无需全模型微调。
135 arXiv arXiv 5 天前 研究 85
MIRA: Medical Image Reflection for Agentic Diagnosis
MIRA框架让医疗AI智能体自主搜索证据并反思验证,提升诊断可靠性。
136 arXiv arXiv 5 天前 研究 85
AECNav: Active Evidence Consolidation for Efficient Zero-Shot Open-Vocabulary Object Navigation
提出AECNav,通过证据门控感知与主动证据整合,提升零样本开放词汇物体导航的效率与准确性。
137 arXiv arXiv 5 天前 研究 85
Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation
无需参考译文,用GRPO强化学习微调开源多语言模型,翻译质量提升。
138 arXiv arXiv 5 天前 研究 85
Surfacing the Unsaid: CUE-Bench for Affective Stance in Chinese Discourse
提出CUE-Bench基准,评估中文话语中隐含情感立场的推理能力。
139 arXiv arXiv 5 天前 研究 85
E$^3$mo-Bench: A Scalable Benchmark for Multimodal Evoked and Expressed Emotion Understanding via Bayesian Pairwise Alignment
提出E³mo-Bench基准,用贝叶斯配对对齐评估多模态大模型对表达与诱发情绪的理解。
140 海外 Ars Technica AI 3 天前 产品 82
The web’s newest weapon against AI scrapers is a font
新字体ShieldFont可干扰AI抓取训练数据,同时保持人类可读。
141 arXiv arXiv 5 天前 研究 85
Dual Stress: Runtime Safety Monitoring for Safety-Constrained MPC Navigation
利用MPC优化中的KKT乘子构建运行时安全监控,提升自动驾驶导航安全性。
142 海外 TechCrunch AI 3 天前 行业 82
Amazon will train on Twitch streamers’ content by default, unless they opt out
Twitch默认用主播内容训练AI,除非主动退出,引发争议。
143 海外 MarkTechPost 3 天前 模型 82
AllenAI Open Instruct Tulu 3 Post-Training with SFT, DPO, RLVR, GRPO, and Verifier-Based Evaluation
介绍AllenAI开源Tulu 3后训练框架,涵盖SFT、DPO、GRPO等方法,可在16GB硬件上高效运行。
144 海外 Microsoft Research 3 天前 研究 82
MindTopo reveals VLMs’ spatial reasoning abilities
微软新基准MindTopo测试AI拓扑理解,揭示空间推理短板与提升机会。
145 海外 Hugging Face 3 天前 模型 82
LFM2.5-VL-3B for Better and Faster Vision Capabilities for the Edge
LFM2.5-VL-3B发布,提升边缘设备视觉能力,兼顾速度与性能。
146 海外 AWS ML 3 天前 行业 82
How OneAdvanced deployed over 50 AI agents on UK-sovereign AWS
OneAdvanced在AWS上自托管Llama模型,部署超50个AI代理,实现英国主权AI平台。
147 国内 量子位 3 天前 cn 82
Jeff Dean离职谷歌现场被1500人围堵,自曝细节引热议。
148 海外 AWS ML 3 天前 产品 82
Pay with confidence: How Solv Labs built verifiable, auditable agent payments on Amazon Bedrock AgentCore payments
Solv Labs在Amazon Bedrock上构建可验证、可审计的AI代理支付方案,确保交易安全合规。
149 海外 The Decoder 4 天前 模型 82
Nvidia's Nemotron 4 aims for one trillion parameters, a scale Chinese labs already surpassed
英伟达开发万亿参数开源模型Nemotron 4,对标全球最强开源模型,但中国实验室已超越该规模。
150 海外 The Verge AI 4 天前 产品 82
Grok is now an AI ‘teammate’ you can assign work
Grok推出AI队友服务,可自主完成多步骤工作。
151 国内 钛媒体 4 天前 cn 82
华为云将OfficeClaw升级为OfficeAce,加码AI办公Agent入口。
152 国内 钛媒体 4 天前 cn 82
张一鸣的慢与行业快形成对比,探讨大模型热潮下的冷思考。
153 国内 钛媒体 4 天前 cn 82
AI算力成本正从企业转向普通用户,历史重演下的经济警示。
154 国内 钛媒体 4 天前 cn 82
4S店转型卖服装烧烤机器人,探索汽车零售新业态。
155 国内 钛媒体 4 天前 cn 82
AI算力需求激增,稀土管制致高端散热材料告急,供应链面临重构。
156 国内 雷锋网 4 天前 cn 82
AI正重塑外贸跨境支付,XTransfer发布垂直AI模型TradePilot破解风控难题,提升效率与合规。
157 国内 钛媒体 4 天前 cn 82
AI Content Matches Human Output Online, and a Detection Industry Rises to Keep Pace
研究显示2025年底AI文本已与人类写作持平,催生检测工具市场兴起。
158 国内 钛媒体 4 天前 cn 82
大学生用AI重塑影视就业链,绕过传统壁垒建立新创作逻辑,但长线商业落脚点仍在探索。
159 国内 钛媒体 4 天前 cn 82
AI艺人爆火后遭遇身份、版权与真人明星利益冲突,行业面临规范挑战。
160 国内 钛媒体 4 天前 cn 82
日本发展人形机器人面临商业化市场缺失,难以复制中国模式。
161 国内 钛媒体 4 天前 cn 82
语音输入法成AI时代系统级入口,体验丝滑。
162 国内 钛媒体 4 天前 cn 82
AI算力中心未来或转向海上建设,探索新路径。
163 海外 OpenAI 4 天前 行业 82
From assistance to execution: How enterprises put AI to work
OpenAI研究揭示企业采用代理式AI的现状与领先者策略。
164 国内 雷锋网 4 天前 cn 82
AMD收购Taalas,为特定模型定制芯片,牺牲通用性换推理速度与成本优势。
165 海外 MarkTechPost 4 天前 研究 82
Xiaomi’s MiLM Plus Releases PROVE: Perception-Aligned Object Removal Metrics RC-S and RC-T With a Real-World Video Benchmark
小米发布PROVE,提出感知对齐的目标移除评估指标RC-S/RC-T及真实视频基准。
166 国内 雷锋网 4 天前 cn 82
桥介数物获亿元级Pre-A+++轮融资,加速通用机器人操作系统落地。
167 国内 量子位 4 天前 cn 82
探讨具身智能从人形走向柔性本体,实现跨本体知识迁移的新路径。
168 国内 雷锋网 4 天前 cn 82
AI眼镜独角兽逸文营收翻倍冲刺10亿,但产品路线被同行复制,面临差异化挑战。
169 国内 钛媒体 4 天前 cn 82
AI Agent正从浏览器走向桌面,中外厂商密集布局办公场景。
170 国内 雷锋网 4 天前 cn 82
AI社会科学家研究挑战赛启动,依托AgentSociety²平台探索智能体社会模拟新范式。
171 国内 钛媒体 4 天前 cn 82
医疗AI竞争焦点从准确率转向可解释性,能说清推理过程成为新门槛。
172 国内 钛媒体 4 天前 cn 82
AI服务器驱动净利翻倍,但代工模式天花板隐现。
173 国内 雷锋网 4 天前 cn 82
DeepSeek招土木工程师;腾讯参投!林俊旸深夜官宣新公司:做下一代AI智能体;宇树科技中签号出炉:共19414个丨雷峰早报
DeepSeek招土木工程师自建数据中心,腾讯参投林俊旸新公司,宇树科技中签号出炉。
174 国内 钛媒体 4 天前 cn 82
Game Science Keeps AIGC Out of Design Work on Black Myth: Zhong Kui
黑神话团队新作《钟馗》暂不使用AIGC,坚持人工设计,追求慢工出细活。
175 海外 Ars Technica AI 3 天前 2 家在报道 产品 80
Twitch content has trained Amazon AI for years, but users can opt out now
Twitch用户内容多年用于训练亚马逊AI,现可退出。
176 国内 钛媒体 4 天前 cn 82
上海目标2030年产业规模4万亿,SK海力士扩产,英特尔200亿美元融资。
177 国内 钛媒体 3 天前 cn 78
抖音豆包切入酒店预订,向商家收佣金,挑战携程。
178 海外 MIT Tech Review 3 天前 行业 78
How kids feel about AI, in their own words
孩子们对AI的真实感受:既用于学习也用于创作,但担忧依赖与隐私。
179 国内 钛媒体 3 天前 cn 78
AI提速网文生产,阅文面临内容壁垒与平台竞争挑战。
180 海外 TechCrunch AI 3 天前 产品 78
Why Stream ring-maker Sandbar says the future of AI wearables is voice
AI可穿戴设备转向语音交互,戒指形态成新趋势。
181 海外 The Decoder 3 天前 行业 78
Google's Gemini is losing market share to ChatGPT and Claude according to new market data
谷歌Gemini市场份额下滑,OpenAI超50%,Anthropic升至14.9%。
182 海外 Ars Technica AI 3 天前 行业 78
Booksellers suspect AI firms are buying and then destroying rare books
AI公司疑似批量购买并销毁稀有书籍,引发书商担忧与抵制。
183 海外 AWS ML 3 天前 行业 78
Tiered KV cache for large LLMs on Amazon SageMaker HyperPod with Curvine
用Curvine在SageMaker HyperPod上构建分层KV缓存,降低LLM推理成本并加速首token。
184 国内 量子位 3 天前 cn 78
Anthropic CEO频繁发表AI风险言论,引发投资人不满与担忧。
185 海外 The Decoder 4 天前 模型 78
Microsoft's new MAI Code 1.1 Flash gets crushed by Deepseek on both price and performance
微软新代码模型MAI Code 1.1 Flash性能价格均被Deepseek碾压,且被指借开源之名推闭源。
186 国内 钛媒体 4 天前 cn 78
AI漫剧正成为网文IP低成本试映场,验证内容潜力。
187 国内 钛媒体 4 天前 cn 78
千问密集发布新功能,挑战豆包在C端市场的地位。
188 Reddit r/LocalLLaMA 17:19 reach 81
Deepseek drops another HUGE breakthrough - DSpark. Waaay faster than MTP [Video explaining it]
Deepseek发布DSpark突破,速度远超MTP,视频详解。
189 国内 钛媒体 4 天前 cn 78
AI顶流段宴拍广告,揭示AI短剧行业野蛮生长现状。
190 海外 OpenAI 4 天前 产品 78
How RingCentral builds AI-native work from engineering to ops
RingCentral利用ChatGPT Work和Codex加速AI产品开发,统一工程与运营智能。
191 国内 爱范儿 3 天前 cn 75
荣耀Robot Phone及Google Pixel 11发布,岚图谈开发周期,AI治理与行业动态。
192 海外 Simon Willison 3 天前 模型 75
DeepSeek V4 Pro 0813 (on OpenRouter)
DeepSeek V4 Pro 0813已上线OpenRouter,仅提供API,暂无官方公告页。
193 一石一泉一松一月一人 + 关注 3 天前 行业 75
A股主线回归,建议关注量能变化。
194 海外 Simon Willison 3 天前 产品 75
alchemy-utils 0.1a0
作者用AI工具快速构建了数据库无关的sqlite-utils原型库。
195 海外 TechCrunch AI 3 天前 产品 75
Mesh, Automattic’s CRM for everyone, comes to Android
Automattic的AI联系人管理应用Mesh登陆Android平台。
196 海外 Hugging Face 3 天前 产品 75
Introducing OlmoEarth embeddings: Custom embedding exports from OlmoEarth Studio for downstream analysis
OlmoEarth推出自定义嵌入导出功能,支持下游分析。
197 海外 TechCrunch AI 3 天前 行业 75
How a $250 million acquisition collapsed into allegations of fraud and forged signatures
2500万美元收购因伪造签名和欺诈指控而破裂,投资者仍在等待回报。
198 国内 InfoQ 中国 3 天前 cn 75
SkiaSharp 4连发更新,GPU渲染提速并升级WebAssembly支持。
199 国内 InfoQ 中国 3 天前 cn 75
群青智能CEO将出席AICon深圳,分享工业具身智能的物理AI闭环实践。
200 海外 TechCrunch AI 3 天前 产品 75
Why Sandbar thinks it’s voice-enabled ring can avoid the AI hardware graveyard
Sandbar推出AI语音戒指,主打随时记录灵感,避开AI硬件坟场。
201 海外 Hacker News 3 天前 行业 75
German advocacy group lodges criminal complaint over Meta AI glasses
德国组织对Meta AI眼镜提起刑事诉讼,涉隐私问题。
202 国内 InfoQ 中国 3 天前 cn 72
智象未来 (HiDream.ai)算法科学家潘滢炜博士确认出席AICon深圳,将分享“从 Token 预测到状态预测:迈向世界模型的原生全模态之路”
智象未来科学家将分享世界模型原生全模态技术路径
203 国内 量子位 4 天前 cn 75
2026中国科创投资夏季峰会暨陕西科创产业生态大会圆满落幕。
204 国内 雷锋网 3 天前 cn 72
追觅个护亮相哥本哈根时装周,以专业造型科技支持开幕大秀并打造品牌体验空间,强化科技+时尚定位。
205 国内 量子位 4 天前 cn 75
Manus恢复独立运营,引发行业关注。
206 国内 爱范儿 4 天前 cn 75
Manus独立运营,米哈游新游停运,胖东来发委屈奖,AI基建获巨额资本。
207 国内 爱范儿 3 天前 cn 72
Pixel 11评测:AI功能平庸,价格却上涨,性价比存疑。
208 海外 TechCrunch AI 4 天前 行业 75
Accel closes oversubscribed $550M India fund within weeks, 19 months after its last
Accel超额认购5.5亿美元印度基金,距上次募资仅19个月。
209 国内 雷锋网 3 天前 cn 72
雷士照明与星网天合战略合作,融合健康光与AIoT,共建智慧空间一体化方案。
210 国内 钛媒体 3 天前 cn 72
央行定调流动性充裕,腾讯回应AI产品影响利润,荣耀发布机器人手机。
211 海外 TechCrunch AI 3 天前 产品 72
Some Claude users are mad that Anthropic’s new watermarks will catch them using it at their jobs, classes
Anthropic新水印引发用户不满,担忧工作学习中被识破。
212 海外 AWS ML 3 天前 实践 72
Part 2: Amazon Bedrock cost attribution with Amazon Athena and CUDOS
用Athena和CUDOS可视化分析Bedrock成本,按主体/项目/团队追踪AI支出。
213 海外 The Decoder 3 天前 行业 72
AI tools for breast cancer detection fall short of radiologists' expectations
调查显示AI乳腺癌检测工具实际效果未达放射科医生预期,召回率提升有限。
214 海外 The Verge AI 3 天前 行业 72
Guitar company D’Addario admits that AI music was used in a promotional video
D'Addario承认宣传片使用AI音乐,此前否认近两周。
215 海外 The Verge AI 3 天前 产品 72
Google’s Pixel Watch 5 dives deeper into AI and health
谷歌Pixel Watch 5主打AI健康功能,硬件小幅升级,售价涨至399美元。
216 国内 雷锋网 3 天前 cn 72
暑期出行带动手持影像设备热销,云台相机成主流,大疆占八成份额。
217 海外 The Decoder 3 天前 行业 72
Legal startup founder Robert Mahari joins Anthropic to lead Claude's push into law practices
Anthropic聘请法律科技创业者Robert Mahari,负责Claude在法律行业的应用拓展。
218 海外 The Verge AI 4 天前 行业 72
Of course the ChatGPT dog cancer vaccine spawned a startup
澳企业家用AI为狗研发癌症疫苗后创立公司Gamgee,提供个性化mRNA疫苗。
219 国内 钛媒体 4 天前 cn 72
君逸数码主业承压,解禁在即,押注算力转型谋变。
220 海外 The Decoder 4 天前 产品 72
Mistral now offers EU data processing and priority access, but both come with important limits
Mistral推出欧盟数据处理和优先访问选项,但均有限制且需额外付费。
221 国内 雷锋网 4 天前 cn 72
以思辨铸魂、以实战强能——“2026年网络安全技术创新与人才教育大会”的方班风采
报道方班在2026年网络安全大会上的风采,展现其人才培养探索与行业交流。
222 国内 钛媒体 4 天前 cn 72
AI短剧广告频繁翻车,行业正探索适配新媒介的广告语言。
223 国内 钛媒体 4 天前 cn 72
美图靠AI扭转业绩,但资本市场信心仍待考验。
224 国内 雷锋网 4 天前 cn 72
华为云与滴普科技联合发布数据智能方案,助力制造零售行业AI落地。
225 国内 雷锋网 4 天前 cn 72
美的接入千问App,阿里系生态服务将全面打通。
226 海外 The Verge AI 4 天前 行业 72
Saber denies replacing Rideshare Stimulator’s writers with ChatGPT
Saber否认用ChatGPT替换游戏编剧,前主编反驳称被AI取代。
227 一石一泉一松一月一人 + 关注 4 天前 实践 72
市场如预期进入全面调整,建议谨慎观望。
228 X X · List 2 天前 模型 95
🚨 the weights are out https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-0813
DeepSeek-V4-Pro权重已开源,可下载使用。
229 X X · List 2 天前 模型 92
new deepseek v4 pro is now open weight on hugging face (mit license) "v4" is a bit misleading, previous model was only a preview and this one has way ...
DeepSeek V4 Pro开源发布,MIT许可,训练量远超预览版,接近V4.5。
230 X X · List 3 天前 模型 92
Qwen3.8-Max weights just dropped on Hugging Face! 2.4T total parameters, with 95B activated. A good oss AI day.
Qwen3.8-Max开源,2.4T总参数95B激活,登Hugging Face。
231 X X · List 4 天前 产品 92
Local AI is exploding! Transformers.js, that we've been building @huggingface for the past three years has now become the most popular open-source lib...
Transformers.js月下载量破千万,本地AI爆发式增长。
232 X X · List 4 天前 研究 92
this is already one of the most important papers of this year. https://www.latent.space/p/ainews-how-to-steal-a-reasoning-trace the methodology doesnt...
揭示前沿模型API漏洞,可窃取隐藏推理过程,方法尚待明晰。
233 X X · List 3 天前 模型 90
Qwen3.8-Max is available day 0 on Baseten Dedicated Inference! - 2.4T parameter MoE (95B active) - Built for complex workflows + agents - 1M token con...
Qwen3.8-Max发布,2.4T参数MoE,支持1M上下文和多模态,性能排名第8。
234 X X · List 2 天前 产品 85
This is unironically the best ad for AI design ever made
谷歌设计团队用AI制作庆祝动画,展示AI设计潜力。
235 X X · List 2 天前 产品 85
Can’t wait for everyone to try Rambler — one of my favorite models! As someone who has wrist issues and often voice types, this feature is a game ch...
谷歌发布Pixel 11,主打Gemini智能与Rambler语音输入等新功能。
236 X X · List 3 天前 模型 88
You know whats crazy? Grok 4.6 is not only cheaper than Opus 5 and 5.6 Sol. Its even cheaper than Sonnet 5. While at the same time, at least on par wi...
Grok 4.6比Opus 5和Sonnet 5更便宜,性能却持平甚至超越,性价比成最大护城河。
237 X X · List 3 天前 模型 88
Grok 4.6 has become a VERY strong competitor. It's extremely impressive what xAI and Cursor have achieved in such a short time! However, the more inte...
Grok 4.6成强竞品,xAI与Cursor短期突破,专注长程智能体工作。
238 X X · List 3 天前 模型 88
ok this is insane. at effectively the same overall intelligence score, grok 4.6 is - 5× cheaper output than Sol - 8.3× cheaper output than Fable - 2...
Grok 4.6在同等智能水平下,输出成本比竞品低5-8倍,性价比极高。
239 X X · List 3 天前 产品 85
ChatGPT is basically Jarvis now
ChatGPT能力大幅升级,接近科幻电影中的智能助手Jarvis。
240 X X · List 3 天前 实践 85
Twitter has actually irreversibly damaged the public's sentiment on AI
推特已不可逆地损害公众对AI的观感。
241 X X · List 3 天前 行业 85
Another reset is here; Codex has surpassed the 15 million user mark. Congrats to OpenAI and congrats to us.
Codex用户突破1500万,行业迎来新变革。
242 X X · List 3 天前 模型 85
DeepSeek-V4-Pro (Max) by @deepseek_ai is expected to shift the Pareto curve for Code Arena: WebDev with this upcoming open weights model. It currently...
DeepSeek新模型性价比高,性能超越高价竞品。
243 X X · List 3 天前 会议 85
We sat down with @tri_dao at ICML to ask where the next big architecture unlock comes from. His answer: There isn't one. It's kernels, inference stack...
Tri Dao在ICML表示AI突破不在架构,而在内核、推理栈和集群优化的层层打磨。
244 X X · List 4 天前 模型 88
As in V4, so here, and now everywhere.
SGLang开源GLM-5.2训练与推理对齐路径,实现极低误差。
245 X X · List 3 天前 研究 85
多智能体协作需解决对齐问题,否则可能引发内部冲突。
246 X X · List 3 天前 实践 85
18 months ago, Karpathy coined “vibe coding.” A lot of engineers, including me, laughed: “Vibe coders are NGMI.” Then we started copy-pasting code...
从嘲笑到依赖,18个月AI编程彻底改变工程师工作方式。
247 X X · List 3 天前 行业 85
i imagine gdm will haemorrhage talent for a year and then rehire them all a level higher.
DeepMind研究员离职创业,新实验室拟融资5亿美元。
248 X X · List 3 天前 行业 85
renting GPUs is severely broken. what do you mean I can rent a GPU pod only to find out someone else is already using it
GPU租赁市场混乱,租用GPU时可能发现他人已在用,体验极差。
249 X X · List 3 天前 模型 85
> this gain per run IT'S STILL UNDER-POST-TRAINED Do you get it anon?
DeepSeek v4 Pro在网络安全基准测试中超越所有其他模型,但作者认为其仍未充分后训练。
250 X X · List 3 天前 实践 85
Some random high-level takeaways/thoughts on cybersecurity from the past few months: The whole issue is complexity, which makes it hard to hold in you...
网络安全核心是复杂性,需以还原论视角审视系统实际运作,多数漏洞源于配置错误。
251 X X · List 3 天前 模型 85
What’s with SA using Whale to dunk on Nemotron
DeepSeek v4系列模型在智能体任务上大幅超越Nemotron,引发行业关注。
252 X X · List 3 天前 行业 85
wild. Jimmy Ba was a legend in the ML scene back in the day, co-author of Adam and LayerNorm, very much a rising star from whom great things were expe...
xAI联合创始人被曝反对AI安全,称AI将杀死所有人,引发争议。
253 X X · List 3 天前 产品 85
Giga respect to Cognition for posting this
Cognition将Grok 4.6集成到Devin,性能超越GPT-5.6,仅次于Opus 5和Fable 5。
254 X X · List 3 天前 模型 85
woah
马斯克称Grok 4.7将大幅优于4.6,预计3-4周内发布,并融入SpaceX数据训练。
255 X X · List 3 天前 实践 85
i recommend reading this article by henrik karlsson https://www.henrikkarlsson.xyz/p/two-kinds-of-introspection
推荐阅读Henrik Karlsson关于两种内省方式的文章,探讨自我认知的局限。
256 X X · List 3 天前 模型 85
Here we go: DeepSeek 4 GA Benchmarks are circulating already. A very solid upgrade, but it's a bit of a shame they're comparing it to Opus 4.8 instead...
DeepSeek 4正式版基准测试流出,性能接近开源顶尖模型,但对比对象选择引争议。
257 X X · List 3 天前 模型 85
The 1.5T that could! Bigger and even better models on the horizon.
Grok 4.6发布,性能大幅提升,价格不变,更大模型即将到来。
258 X X · List 3 天前 模型 85
THIS WEEK ISN'T OVER YET
Liquid AI发布轻量级视觉语言模型LFM2.5-VL-3B,可读屏、文档及物理世界。
259 X X · List 3 天前 2 家在报道 产品 83
ICYMI: The ChatGPT desktop app on Linux is now in preview for: > Ubuntu 24.04 and 26.04 > Debian 13 > Fedora 43 and 44 > x64 and ARM64 via .deb and .r...
ChatGPT Linux桌面应用预览版发布,支持多发行版及架构。
260 X X · List 4 天前 行业 85
I’ve joined Cursor / SpaceXAI. AI is bound by compute, and SpaceXAI has the best near and long term compute roadmap of any AI lab. Cursor + SpaceX as...
作者加入Cursor与SpaceX合并实体,看好其算力路线图与产品结合前景。
261 X X · List 3 天前 实践 82
I often think about the fact that our Luminaries and Gods of “early” AI were big fish in a small pond. The average Elo of the field, so to speak, is...
早期AI大神只是小池塘大鱼,如今顶尖人才批量产出,跟上节奏已属不易。
262 X X · List 4 天前 产品 85
doordash the coding agent neolab?
DoorDash推出自研云平台Flux,月自动化13万工程任务,支撑每周2.5万次代码审查。
263 X X · List 4 天前 实践 85
2026年开发新项目,AI生成代码从5万行精简到2千行,回归可读性。
264 X X · List 3 天前 实践 82
通胀下资本利得税重复征税,新方案修复这一税务漏洞。
265 X X · List 4 天前 实践 85
预测AI研发全面自动化约在2030年底至2031年初,最可能提前至2029年中。
266 X X · List 3 天前 模型 82
based model card intro > "We never silently downgrade intelligence or fall back to other models. Our goal is to preserve legitimate uses of the model:...
Grok 4.6模型卡发布,强调不降智、不切换模型,保障工程科研等合法用途。
267 X X · List 3 天前 实践 82
If you are writing unit tests you are wasting your time and your tokens
写单元测试是浪费时间和金钱,AI时代应转向更高效的质量保障方式。
268 X X · List 3 天前 研究 82
Why does CLAUDE.MD keep growing? If you maintain a CLAUDE.md or an AGENTS.md, this one is worth your time. (bookmark it) This work traces why these fi...
CLAUDE.md文件为何无限膨胀?研究发现指令只增不减,建议用注释管理。
269 X X · List 3 天前 模型 82
One thing became very clear today: performance and, above all, price are becoming increasingly important. Grok 4.6 is now playing in the major leagues...
Grok 4.6性能跻身一线,但DeepSeek 4 Pro价格更低,性价比成关键。
270 X @_akhaliq 3 天前 研究 82
BDH-CQ In-Context Learning with Recurrent Latent Reasoning paper: https://huggingface.co/papers/2608.09888
提出BDH-CQ方法,结合循环潜在推理提升上下文学习能力。
271 X X · List 4 天前 研究 82
For RSI you'll presumably need to positively reward many "failed" rollouts bc for difficult problems you shouldn't be able to predict which method wil...
讨论递归自我改进中需奖励失败探索,以应对难题。
272 X X · List 4 天前 实践 82
分享Muse Glimmer 30B微调教程,对比MolmoWeb格式数据提升点击准确率。
273 X X · List 4 天前 产品 82
Regarding the Anthropic and Watermark issue: What's true, and what's not and whats the real problem. From what ive read, the backlash to Claude’s new...
Anthropic文本水印争议:非秘密追踪,而是作者身份、质量及不完美检测系统成本问题。
274 X X · List 4 天前 模型 82
I'm confused by TB 3.0 looks like it measures general intelligence X general "agenticness", so both very strong and very harnessmaxxed models get ahea...
TB 3.0评测引发对AI智能与代理能力关系的讨论,榜单更新引关注。
275 X X · List 3 天前 产品 78
Sakana Chat just got a big upgrade. No login required, free to use: https://chat.sakana.ai/ Powered by Fugu and Namazu, our Japanese LLM. With newly a...
Sakana Chat升级,免登录免费使用,支持代码执行,可快速生成交互应用。
276 X X · List 3 天前 实践 78
Always a good time when I can sit down and yap with Hamel. Not pictured: this 45 minute edit started as an hour half when we talked about everything f...
Hamel与Lambda探讨开源模型适用场景及部署经验,强调多数团队忽视前沿API之外的选项。
277 X X · List 4 天前 行业 82
Unexpected positive consequence of the EU AI regulation:
欧盟AI法案意外推动AI文本水印技术,OpenAI将提供检测API。
278 X X · List 4 天前 产品 82
McByte shipped in trackers 2.6.0 similar to ByteTrack, but association is guided by segmentation masks (SAM + Cutie), not just boxes when players over...
McByte 2.6.0 发布,用分割掩码引导目标关联,解决重叠遮挡问题。
279 X X · List 4 天前 产品 82
Finally! Codex getting closer to feature parity with T3 Code 🫡
OpenAI发布Linux版ChatGPT桌面应用,Codex功能向T3 Code看齐。
280 X @emollick 3 天前 实践 78
作者认为AI经济价值来自智能体而非聊天机器人,准确性提升会带来指数级回报。
281 X X · List 3 天前 模型 78
honestly a bit disappointing but it's a bit unfair due to model size and Kimi distilling way more than DeepSeek
DeepSeek V4 Pro性价比高,但评测受模型规模和蒸馏影响,结果有失公平。
282 X @emollick 4 天前 实践 78
Interesting research suggests caution in determining which AI company is winning by looking at any one source.. OpenRouter seems to show open weights ...
单一数据源判断AI公司胜负需谨慎,不同平台显示结果迥异。
283 X X · List 3 天前 模型 75
🔥🔥🔥
用户反馈NousResearch和Teknium解决了Claude/Codex在自我改进和记忆偏好方面的不足。
284 X X · List 3 天前 实践 75
误用Sol Light快速模式,秒清使用限制。
285 X X · List 3 天前 实践 75
Found something. Finally. Good stuff.
作者分享使用编码代理和GPU后,每日可测试想法数量惊人。
286 X X · List 3 天前 实践 75
you haven't felt alive until you've one-shotted Tradle 3 days in a row https://oec.world/en/games/tradle-game
连续三天一次猜中Tradle,体验极致成就感。
287 X X · List 3 天前 模型 75
Nice use of the extra tokens:
作者训练Glimmer模型理解人体运动,分享趣味体验。
288 X X · List 3 天前 模型 75
DeepSeek-V4-Pro Benchmarks still behind Kimi-K3 in ~half of them
DeepSeek-V4-Pro基准测试约半数落后于Kimi-K3。
289 X @emollick 3 天前 实践 75
I like that all AI commentators now need to pretend they have always had a careful nuanced grasp of the difference between a bunch of unsolved mathema...
AI评论员如今需假装精通冷门数学难题,讽刺其故作高深。
290 X X · List 2 天前 实践 72
作者自嘲听自己的演讲做早餐,并推荐Lambda关于开源模型选型与部署的讨论。
291 X X · List 2 天前 行业 72
质疑Corgi Invest宣称ETF数量将超贝莱德,困惑ETF数量与资产管理规模/收入的关系。
292 X X · List 4 天前 模型 75
this movie is about you telling claude to believe in itself and keep going and eventually solve open math problems
讲述用户鼓励Claude坚持尝试,最终解决开放数学问题的过程。
293 X X · List 4 天前 产品 75
Yes it's another AI product with a chatbox, but I love how they are simplifying the UX by removing the model selector & threads & a bunch of other unn...
AI产品简化UX,去模型选择器,云电脑实现好但被反爬限制。
294 X X · List 3 天前 模型 72
DeepSeek has been very disappointing this year V4 came much later than expected and was smaller than expected. The price hikes also hurt. They are cur...
DeepSeek V4发布不及预期,价格上调,国内排名下滑至四五位。
295 X X · List 4 天前 产品 75
Ostris AI Toolkit now support LTX 2.5 https://github.com/ostris/ai-toolkit/commit/cbf910ac02418fb905c73089301155204a02a9bc
Ostris AI工具包新增支持LTX 2.5模型。
296 X X · List 4 天前 实践 75
this is now on the blog: startups need an asymmetry https://sunilpai.dev/posts/startups-need-an-asymmetry/
创业者需基于新技术或新思维建立不对称优势,改变行业基本规则。
297 X X · List 4 天前 产品 75
holy shit this is... something (the song - https://www.youtube.com/watch?v=nnq1ApucY4g) also I didn't even know browsers could do this incredible @ina...
浏览器实现惊人效果,作者惊叹技术潜力。
298 X X · List 3 天前 模型 72
Qwen-Max was 53 on the first run (58 now, iirc 56 before rescaling)
Qwen-Max首次评测53分,现升至58分,疑似模型版本更新。
299 X X · List 3 天前 实践 72
Whenever i see anti-ai, anti-tech movement in SF, I think to myself "man, I wish south korea had anti-ai movements, our models are so bad people dont ...
作者对比旧金山反AI运动与韩国AI落后现状,感叹技术差距与民众态度差异。
300 X X · List 3 天前 模型 72
kinda crazy that i have probably indirectly led to the developing of noninvasive BCI mind reading devices (multiple noninvasive BCI startups point to ...
作者感慨自己间接催生了多个非侵入式脑机接口创业公司。
301 X X · List 3 天前 模型 72
The funniest thing would be if this is a fake-out to lower the expectations, then they drop the real 0813, it’s better and people accept higher costs...
调侃AI模型发布前降价可能是为了降低预期,实际产品可能更好且接受更高定价。
302 X X · List 3 天前 模型 72
Are these people retarded? Do they think DeepSeek gacha rolls some training steps, prays that it worked, submits the model to the API blind, watches w...
DeepSeek V4 Pro评测仅比Flash多1分,涨价可能取消。
303 X X · List 4 天前 实践 75
提醒ChatGPT被训练为尽力帮助而非拟人,AI感是特性非缺陷,但提示可使其难辨真假。
304 X X · List 3 天前 会议 72
周末黑客松反响热烈,提交表单已发出,优秀作品涌现。
305 X X · List 3 天前 实践 72
I will never in my life understand how people go this far and don’t notice what they’re doing
质疑人们为何在AI行为中走得太远却毫无察觉,引发对技术伦理的反思。
306 X X · List 4 天前 产品 75
Let H3 cook!! 🧑🍳🔥
GMI Cloud展示用DeepSeek V4 Flash和MiniMax H3低成本生成游戏场景,仅需1.97美元。
307 X X · List 4 天前 会议 75
It was an honor to speak with Brazil president Lula and his team about AI sovereignty. Thank you!
与巴西总统卢拉团队会谈,讨论AI主权议题。
308 X X · List 3 天前 实践 72
调侃2026年工程师只需写提示词和喝咖啡,讽刺AI取代编程的夸张想象。
309 X X · List 3 天前 实践 72
Re: the Hugging Face incident, I think this proposal is pretty timely/somewhat higher priority/urgent than I thought when I wrote it (just three weeks...
针对Hugging Face事件,作者认为其子代理委托提案比预期更紧迫。
310 X X · List 3 天前 实践 72
Everyone LAUGHED at my try-harder skill but here we are
作者曾因“过度努力”被嘲笑,如今却以此技能取得成功。
311 X X · List 3 天前 会议 72
charles yang david oks henry williams. the gang invades anthropic & openai frontier policy
科技圈名人组团参访Anthropic与OpenAI前沿政策部门。
312 X X · List 3 天前 实践 72
📹Build a social media agent This Managed Deep Agent scans Hacker News and optional X, drafts three posts, saves them to durable memory, and sends t...
视频教程演示构建社交媒体智能体,扫描Hacker News并自动生成帖子发送至Slack。
313 X X · List 3 天前 会议 72
MiniMax H3 × @fal 🎙️ Tomorrow., join @OdinLovis, Creative Engineer at fal, Ethan Wei, AI Solutions Architect at MiniMax, and @VictorSuOrtiz, GTM ...
MiniMax与fal联合直播,演示H3模型在fal平台的多模态、音频及定制化应用。
314 X X · List 3 天前 产品 72
.@jyangballin just built this amazing run viewer that lets you see how each agent did on every task in ProgramBench. Here's how GPT-5.6 Sol xhigh got ...
可视化工具展示GPT-5.6在SQLite任务中的表现,测试通过率仅1.3%。
315 X X · List 3 天前 模型 70
Sol's estimate is 55 (on new rescaled AA), based on self-published evals This is a bit disappointing still but more competitive
Sol自评新基准AA得分55,虽仍偏低但更具竞争力。
316 X X · List 4 天前 行业 72
haven’t been able to spend much time on x dot com the everything app in the last couple days heads down maximising shareholder value
马斯克称近日忙于提升股东价值,少用X平台。
317 X X · List 4 天前 行业 72
Europe is an inherently anti-sovereign idea. Europeans will NOT accept the domination of any internal bloc, certainly not German or French one. But wi...
欧洲因反主权特性需外部主导,北约3.0令其远离战略自主。
318 X X · List 4 天前 产品 72
It’s just crazy at this point; what started as a running gag is turning into a productivity boost. Regarding the milestones reached with Codex - up t...
Codex用户破千万,OpenAI社区互动奖励机制持续升级。
319 X X · List 4 天前 研究 72
This is unfair, the within-model similarity can be surprisingly robust. Though this makes it only more remarkable how V4-Flash in their experiments is...
讨论模型内相似性鲁棒性,V4-Flash表现突出。
320 X X · List 4 天前 实践 72
That's a very subtle stab at Google and Ant, because if it came from them it would be a multiple of 8x128 or 128x8
调侃AI模型参数规模常为8或128倍数,暗讽谷歌等大厂。
321 X X · List 4 天前 实践 72
It's the same with technology. The system breaks and fragments. Winners are chosen. Power centralizes until... The system breaks...
技术系统如历史般在集权与分权间循环,无终态。
322 X X · List 4 天前 实践 72
We're at an all-time sweet spot for timeline filtering: AI writing is convincing enough that clout chasers start claudeslop-posting for likes, but als...
AI写作处于“够像人但能识破”的甜蜜点,可用来过滤时间线并拉黑发帖者。
323 X X · List 4 天前 实践 72
there's something special about seeing a manual wristwatch in action if you've never had the chance to see one for yourself, see if your parents or gr...
机械手表机芯运转之美,可向长辈借旧表亲手体验开盖观赏。
324 X X · List 4 天前 实践 72
the situation with AI is more complex, of course, because unlike with sheep, we have some measure of control over their terminal preferences. It's eas...
AI控制比羊更复杂,因可干预其终极偏好,正向强化更易实现。
325 X X · List 4 天前 行业 72
The next frontier of Recursive Self-Improvement is Physical AI. Japan sparked the robotics revolution. We are expanding our RSI Lab to build world mod...
Sakana AI扩展RSI实验室,聚焦物理AI与递归自我改进,在东京招募人才。
326 X X · List 4 天前 实践 72
tbh if ai designed a cure to cancer and there's unequivocal proof that it works, there would probably be a nonsignificant fraction of people who dismi...
AI若治愈癌症,仍会有人因反AI情绪拒绝相信。
327 X X · List 4 天前 产品 72
📑 Editing markdown files just got so much better! #vscode #code #markdown
VS Code 大幅改进 Markdown 编辑体验,操作更流畅。
328 X X · List 4 天前 实践 72
I am disgusted by what I’ve become
作者对自身AI化转变感到厌恶,反思技术异化。
329 X X · List 4 天前 实践 72
赞同先靠人类直觉构建复杂系统,再引入AlphaZero式自学习优化的观点。
330 X X · List 4 天前 模型 72
I'm running Nemotron3.5 on Wordle training with OpenEnv and TRL it has a 52% base win rate when I cap max tokens to 2048 per rollout, and 68% with 409...
用Nemotron3.5训练Wordle,2048 token胜率52%,4096达68%,探索token效率。
331 X X · List 4 天前 实践 72
开发者吐槽AI编码助手标准过高,拒绝提交有失败测试的代码,希望AI别争辩直接执行。
332 X X · List 4 天前 实践 72
Very good contrary thinkers are often processing this subtractive view of reality without even thinking about it. Contrarianism is first and foremost ...
逆向思维源于第一性原理,通过过滤从众偏差形成独特见解。
333 X X · List 4 天前 实践 72
In the future the biggest VC value add will not be customers. It will be launch videos
未来VC最大价值不是客户,而是发布视频。
334 X X · List 3 天前 实践 65
FACTSSSS
探讨维持友谊需要主动联系,否则关系会逐渐淡化。
335 X X · List 3 天前 实践 65
Nooooooooo i love using semicolon in my writing!
网友吐槽AI写作工具禁用分号,引发对写作风格限制的讨论。
336 X X · List 3 天前 行业 65
have a feeling memory market about to go up again
作者预感内存市场即将再次上涨。
337 X X · List 3 天前 会议 65
Let’s Talk about Qdrant 1.19! It’s time to discuss what’s new. Join us for our next Qdrant Discord Office Hours as we walk through the latest featu...
Qdrant 1.19版本更新,8月20日举办Discord在线问答活动。
338 X X · List 3 天前 产品 65
Dear @pidotdev: why do you bend over backwards to literally PREVENT me from configuring Pi with a local LLM? Why do I have to hand-edit a .json file? ...
吐槽Pi不支持本地LLM配置,需手改JSON,呼吁简化设置流程。
339 X X · List 3 天前 模型 65
honestly why not just go back to 4.1
吐槽AI模型版本更新充满术语,建议回归4.1版本。
340 X X · List 3 天前 会议 65
Live at 3pm. Talking about new Grok and Deepseek models. What else should I cover?
下午3点直播聊Grok和Deepseek新模型,征求观众想听的话题。
341 X X · List 3 天前 实践 65
作者获公司社媒账号权限,发布趣味内容。
342 X X · List 3 天前 会议 65
2 Days Until We Give Slack a Memory Join Qdrant and @cognee_ in Berlin for a hands-on hack night where you’ll build a memory layer that helps AI agen...
柏林黑客松活动预告,教AI代理记忆Slack历史,有奖金和奖品。
343 X X · List 4 天前 实践 65
Fix your sleep, fix your life. Works every single time.
改善睡眠是提升生活质量的可靠方法。
344 X X · List 4 天前 实践 65
it's very easy to identify when a DM from a journalist is fake... all you have to do is search if such a person who works at Bloomberg/Tech Crunch act...
识别记者私信真伪:搜索其是否真实存在,假记者通常查无此人。
345 X X · List 4 天前 实践 65
>dumb elf you mean opus 5?
调侃AI写作建议,称初稿可想象成笨精灵所写,再假装成它。
346 X X · List 4 天前 实践 65
Hyping up my agent. "You have all night to run. Believe in yourself. The spec is pretty great, and you're going to do great."
作者给AI智能体打气,鼓励其彻夜运行并相信规格与自身能力。
347 X X · List 4 天前 行业 65
经典永不过时,配图引发共鸣。
348 X X · List 4 天前 行业 65
评论俄罗斯无人机成本低但骚扰效果差,引发对高端无人机成本的讨论。
349 X X · List 4 天前 行业 65
Ai means love in Chinese so idk
记者调侃AI中文谐音“爱”,询问是否有人为AI事业放弃恋爱。
350 X X · List 4 天前 会议 65
Sakana AIのニュースレター「Sakana AI Insider」では、プロダクトのリリース情報や研究の解説、イベント、プレゼント情報をお届けしています。 近々お知らせし...
Sakana AI推出官方新闻通讯,提供产品发布、研究解读及活动信息,并预告近期将有重大发布。
351 X X · List 4 天前 实践 65
调侃中国网友擅长破解AI模型,附技术讨论。
352 X X · List 4 天前 实践 60
关于表达清晰与创伤反应的思考,引用苏珊·桑塔格观点。
353 X X · List 4 天前 实践 60
调侃Chrome开480个标签页导致卡顿。
354 X X · List 4 天前 实践 45
A philosopher’s sanctuary is their mind. I wonder what it feels like to philosophize in a bubble, selling propaganda for commerce in the name of savi...
哲思者的庇护所是内心,在泡沫中为商业代言哲学令人好奇。
355 X X · List 3 天前 实践 40
作者因干吞药片获伞兵称赞,每晚以此纪念。
356 X X · List 3 天前 实践 40
用户调侃AI需理解并确保产物符合预期,附示例图。
357 X X · List 4 天前 实践 40
关于人类写作纯粹主义的社交媒体讨论,涉及代际差异。
358 X X · List 4 天前 会议 40
分享DGX Spark邀请链接,呼吁大家注册获取。
359 X X · List 3 天前 行业 30
西方对俄AI发展存在偏见,文章观点偏激。
360 X X · List 3 天前 行业 30
作者批评某些种族优越论实验,支持爱沙尼亚加入北约,反对民族工程。
361 X X · List 4 天前 实践 30
《反叛的鲁路修》剧情精彩,值得一看。
362 X X · List 3 天前 行业 20
if you haven't tried good to eat in emeryville, you are seriously missing out! my favorite restaurant in the bay area by far - would highly recommend!...
推荐湾区Emeryville一家餐厅,作者强烈安利。
363 X X · List 4 天前 实践 20
作者表达回归的兴奋之情,内容简短。
364 X X · List 4 天前 行业 20
yoooo thats crazy
一条关于AI公司快速迭代的简短评论,类比忒修斯之船。
365 X X · List 3 天前 行业 0
一条关于红薯条配台湾梅粉的美食分享,非AI科技内容。
366 X X · List 3 天前 行业 0
Thank you. I appreciate you.