技术博客

大厂技术博客

435 篇 · 来自 33 个机构的 AI 工程实践,覆盖 Agent、RAG、推理优化、安全对齐等方向

已读 0/435
主题
难度
机构

435 篇匹配

Agent 智能体121

Amazon入门

How Cohere Health digitizes clinical policies using Amazon Bedrock AgentCore

在本文中,您将了解 Cohere Health 如何基于 AgentCore 构建多租户智能体架构,利用 AgentCore Runtime 的安全 MicroVM 隔离技术、通过 AgentCore Gateway 实现的统一工具访问、AgentCore Memory 以及 A

Amazon
2026-08-071 分钟
Amazon进阶

How TReNDS automates root-cause analysis with Amazon Bedrock

TReNDS 是佐治亚州立大学的一个研究中心,基于 Amazon Bedrock 和开源 Strands Agents SDK 构建了一条智能体(Agent)AI 流水线,可实时自动调查生产环境错误,将根因分析从 15 至 30 分钟的人工工作缩短至 60 秒以内。

Amazon
2026-08-075 分钟
Amazon进阶

Securing AI agents with temporal policies in Amazon Bedrock AgentCore

Amazon Bedrock AgentCore 中的时间策略(Temporal Policies)允许您定义有状态规则,基于智能体的会话历史评估授权。了解如何强制工作流顺序、防止数据伪造、限制财务风险敞口,并对高价值操作要求人工审批。

Amazon
2026-08-064 分钟
Amazon进阶

Configure rate limits for AI traffic on AgentCore gateway

了解如何在 Amazon Bedrock AgentCore 网关上配置速率限制,以实施按用户和按目标的流量控制。通过 JWT 声明或 IAM 身份定义请求、令牌和连接限制,保护下游模型、工具和智能体(Agent)免受流量峰值冲击。

Amazon
2026-08-062 分钟
Hugging Face进阶

Deploy local agents everywhere with LFM2.5-2.6B

LFM2.5-2.6B 专为在设备端驱动高性能智能体(Agent)而构建。它支持工具调用(tool calling)和多步工作流,同时保持足够小巧和快速,可运行于日常硬件——从笔记本电脑到手机。这使得开发者能够将智能体部署到任何地方,将数据保留在设备端以保证隐私,并且无需承担云端

Hugging Face
2026-08-046 分钟
Microsoft入门

Orchard: An open framework for scalable agentic AI

Orchard 是一个面向研究社区的开源框架,用于跨任务类型训练和评估 AI 智能体(Agent)。它通过允许研究人员复用相同的基础设施,在降低复杂性的同时,支持较小模型实现强劲性能。

Microsoft
2026-08-031 分钟
Microsoft入门

Echoverse: Deep, evolving environments for computer-use agents

计算机使用型AI智能体在处理电子邮件和客户支持等多步骤工作流程时表现不佳。Echoverse并非简单地提供更多训练任务,而是在逼真的环境中训练智能体,随着任务、测试和环境的不断演化,帮助它们持续提升能力。

Microsoft
2026-07-301 分钟
NVIDIA进阶

Six Agent Harness Capabilities for Higher Model Performance

NOOA 由 NVIDIA Labs 开发,是一个开源、面向对象的智能体框架,将智能体结构化为单个 Python 类,通过方法、字段和文档字符串集成能力、状态和提示,类型注解作为强制契约;由 LLM 驱动的循环在运行时完成由省略号标记的方法体。

NVIDIA
2026-07-276 分钟
NVIDIA进阶

NVIDIA Nemotron 3 Ultra Leads Open Models on Accuracy and Efficiency in Agentic RTL Coding

NVIDIA Nemotron 3 Ultra 在智能体 RTL 编码中引领开放模型的准确性与效率

NVIDIA
2026-07-276 分钟
NVIDIA前沿

Mastering Agentic Techniques: AI Agent Reinforcement Learning

Mastering Agentic Techniques: AI Agent Reinforcement Learnin

NVIDIA
2026-07-0114 分钟
Hugging Face入门

ScarfBench: Benchmarking AI Agents for Enterprise Java Framework Migration

ScarfBench: Benchmarking AI Agents for Enterprise Java Frame

Hugging Face
2026-06-302 分钟
Microsoft Research进阶

SkillOpt: Agent skills as trainable parameters

SkillOpt: Agent skills as trainable parameters

Microsoft Research
2026-06-306 分钟
BAAI进阶

Latent Actions from Factorized Transition Effects under Agent Ambiguity

Latent Actions from Factorized Transition Effects under Agen

BAAI
2026-06-292 分钟
NVIDIA进阶

如何在企业 AI 工厂中治理自主代理

如何在企业 AI 工厂中治理自主代理

NVIDIA
2026-06-297 分钟
OpenAI进阶

How agents are transforming work

How agents are transforming work

OpenAI
2026-06-256 分钟
BAAI入门

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems

The Hitchhiker's Guide to Agentic AI: From Foundations to Sy

BAAI
2026-06-223 分钟
NVIDIA进阶

How Telcos Build Autonomous Networks with Agentic AI

How Telcos Build Autonomous Networks with Agentic AI

NVIDIA
2026-06-221 分钟
Hugging Face入门

MosaicLeaks: Can your research agent keep a secret?

MosaicLeaks: Can your research agent keep a secret?

Hugging Face
2026-06-182 分钟
DeepSeek入门

awesome-deepseek-agent

DeepSeek
2026-06-173 分钟
Cohere进阶

Introducing North Mini Code: Cohere’s first model for developers

今天,我们发布了 North Mini Code 开源模型。作为一款混合专家(Mixture-of-Experts,MoE)模型,North Mini Code 是 Cohere 的首个智能体编码模型,也是我们下一代强大模型系列的首位成员。

Cohere
2026-06-095 分钟
进阶

Eval 正在取代 PRD?产品经理的 Eval 入门到落地指南

综合 Anthropic《Demystifying Evals for AI Agents》、Meta PM Daniel McKinnon《Show, Don't Tell》、Braintrust《Evals Are the New PRD / Evals for PMs》、H

AgentEval/评测
2026-06-097 分钟
MiniMax进阶

For AI agents (OpenClaw, Cursor, Claude Code, etc.): add skill to your agent

Built for AI agents. Generate text, images, video, speech, and music — from any agent or terminal.

MiniMax开源技术报告
2026-06-094 分钟
Anthropic进阶

Best Practices for Claude Code(Anthropic 官方指南)

上下文窗口是第一约束:/clear、compaction、subagent 隔离、CLAUDE.md 精简等所有最佳实践的底层逻辑都是保护上下文质量,防止模型随上下文填充而遗忘指令、错误增多。

AgentClaude Code
2026-06-064 分钟
Google Research进阶

解锁可靠响应:Gemini Enterprise Agent Platform 的 Agentic RAG

解锁可靠响应:Gemini Enterprise Agent Platform 的 Agentic RAG

Google Research
2026-06-057 分钟
Anthropic入门

AI 网络威胁全景图:Anthropic 一年追踪报告

研究期间(2025/03—2026/03),67.3%(560/832)的恶意账号使用 AI 编写恶意软件,这是最普遍的用途。但更值得警惕的是后渗透阶段的 AI 化:6.5%(54/832)的账号用 AI 辅助横向移动(lateral movement),账号发现(account

AgentClaude CodeHarness工程
2026-06-033 分钟
Anthropic进阶

Anthropic 完成 650 亿美元 H 轮融资,估值 9650 亿美元

Anthropic 完成 650 亿美元 H 轮融资,估值达 9650 亿美元,由 Altimeter、Dragoneer、Greenoaks、Sequoia 等领投,并引入 Micron、Samsung、SK hynix 三家芯片战略伙伴。

companyAgentRAG/检索
2026-05-285 分钟
Anthropic入门

Introducing Claude Opus 4.8(Claude Opus 4.8 发布公告)

Opus 4.8 将代码缺陷未标记率降低约 4 倍,更主动标记不确定性、拒绝无依据断言,将「诚实性」作为可量化的对齐指标——亲社会特征得分创新高,欺骗率大幅低于 4.7。

AgentEval/评测Claude Code
2026-05-283 分钟
Anthropic入门

Anthropic 开设米兰办公室(Anthropic opens Milan office)

Anthropic 在米兰开设欧洲第六家办公室,此前已在伦敦、都柏林、巴黎、苏黎世、慕尼黑设点。办公室由 Thomas Remy(南欧负责人)领导,服务意大利金融(Generali Group、Unipol Group)、生命科学(Angelini Pharma、Bracco G

AgentRAG/检索Claude Code
2026-05-272 分钟
Anthropic入门

Coding Agents in the Social Sciences(编码Agent与社会科学)

调查1260位量化社会科学家,81%曾使用AI聊天机器人辅助研究,但仅20%将编码Agent(如Claude Code、Codex)纳入常规工作流(每周超一次)。这一差距并非认知不足,而是从对话式辅助到让AI自主端到端执行代码分析的工作流范式转变门槛。

AgentClaude Code经济影响
2026-05-273 分钟
Anthropic入门

Anthropic 任命 KiYoung Choi 为韩国代表理事,首尔办公室开业在即

原文:Anthropic appoints KiYoung Choi as Representative Director of Korea ahead of Seoul office opening

companyAgentMCP
2026-05-262 分钟
Anthropic进阶

How we contain Claude across products

人类监督(Human-in-the-loop)是概率性防御——遥测数据显示用户批准了约 93% 的权限提示,批准越多注意力越低,形成\"批准疲劳\"。环境隔离(沙箱/VM/出口控制)是确定性防御,限制的是 Agent 能做什么,而非它倾向做什么。两者需要互补,不能互相替代。

AgentClaude Code
2026-05-253 分钟
AI 工程实践进阶

Easy OPC Agent Skills 集(9 个)

每个 Skill 都遵循相同的目录结构(与 [[concepts/agent-skills]] 标准对齐):

AgentClaude Code安全/对齐
2026-05-247 分钟
AI 工程实践前沿

jiangye1314 OPC Skill 集(5 个 + 12 套 toolkit)

与 [[sources/opc-skills-easy]] 形成双范式互补:Easy 偏方法论思辨派,本资料偏新手实操派 + 中国本地化。

AgentMCP
2026-05-2413 分钟
AI 工程实践进阶

一人公司,风口还是泡沫?

截至 2025 年 6 月全国一人有限责任公司突破 1600 万家(占企业总量 27.4%),上半年新注册同比激增 47%;但仅 2% 年收入达 500 万以上、月收入中位数不足 7000 元——入场门槛低了,活下来的门槛没变。

Agent
2026-05-235 分钟
Mistral AI入门

Connect the dots: Build with built-in and custom MCPs in Studio

Connect the dots: Build with built-in and custom MCPs in Stu

Mistral AI
2026-05-223 分钟
Mistral AI进阶

Remote agents in Vibe. Powered by Mistral Medium 3.5.

Remote agents in Vibe. Powered by Mistral Medium 3.5.

Mistral AI
2026-05-224 分钟
AI 工程实践进阶

新风口来了|北京重点扶持 AI 轻量化一人公司(OPC)

注:本文五分类与小麦 [[sources/opc-six-business-models]] 的六分法存在差异——小麦更侧重商业模式属性(出售什么),本文更侧重业务领域分类(卖给谁/做什么)。

Agent
2026-05-223 分钟
Anthropic入门

Project Glasswing: An initial update(项目蜻蜓:初步进展报告)

Project Glasswing 联合约 50 家合作伙伴(Cloudflare、Mozilla、Cisco 等)使用 Claude Mythos Preview,在一个月内在全球最关键软件中发现超过 10,000 个高危或严重漏洞。Cloudflare 单独发现 2,000

Agent安全/对齐
2026-05-223 分钟
Microsoft进阶

MagenticLite, MagenticBrain, Fara1.5: An agentic experience optimized for small models

MagenticLite is an agentic system for small models that works across the browser and local file system in a single workflow. It combines spe

Microsoft
2026-05-211 分钟
MiniMax进阶

MiniMax-MCP

MiniMax-MCP

MiniMax
2026-05-217 分钟
Anthropic入门

2028: Two Scenarios for Global AI Leadership(2028:全球AI领导权的两种情景)

算力是决定性战略变量——通过先进芯片出口管制,民主国家对中国形成实质性算力优势(华为 2026 年算力产出仅约 NVIDIA 的 4%),这是中国 AI 实验室智能落后的根本原因,而非人才或数据。

AgentEval/评测
2026-05-142 分钟
AI 工程实践进阶

6 种 OPC 商业模式

核心公式:一个大脑(人)+ 一套 AI 智能体(执行)= 一家完整的公司。

AgentRAG/检索MCP
2026-05-144 分钟
行业实践进阶

蚂蚁阿福:从 0 到生产的医疗 Agent 工程化落地

1. EBDD(Evaluation and Badcase Driven Development) 是医疗 Agent 的核心研发范式

AgentEval/评测RAG/检索
2026-05-114 分钟
行业实践进阶

从 Copilot 到 Director:多模态智能体如何接管 AIGC 流程

1. AIGC 三大挑战:规划无逻辑 / 调用无统筹 / 评估无标准

AgentEval/评测多模态
2026-05-115 分钟
行业实践进阶

从 Copilot 到 DataAgent:企业级智能数据开发治理平台的技术演进和实践

1. 通用 Agent 承载不确定性,垂直组件处理确定性 — 可控自主化的核心策略

AgentMCP
2026-05-116 分钟
Microsoft进阶

SocialReasoning-Bench: Measuring whether AI agents act in users’ best interests

Using SocialReasoning Bench, we observed a stable pattern across models—agents execute competently, but fail to consistently improve the use

Microsoft
2026-05-111 分钟
行业实践进阶

企业 AI 应用全景:从试点热潮到规模落地的真实分水岭

感知(Perception)— 多模态输入理解

Agent多模态Harness工程
2026-05-105 分钟
记忆系统入门

Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory

LLM 固定上下文窗口无法维持跨会话的一致性。即便 GPT-4o(128K)、Claude 3.7 Sonnet(200K)、Gemini(10M)这样的大窗口,也只是延迟而非解决问题——真实对话中信息主题跳跃,关键偏好会被大量无关内容淹没,注意力在远端 token 上退化。

AgentEval/评测长上下文
2026-05-102 分钟
记忆系统进阶

构建 AI 的\"第二大脑\":大规模多模态记忆平台技术实践

企业部署 AI Agent 后,客户发现问题并持续反馈了 7 周,Agent 始终无法记住这些反馈——每次对话都从零开始。这是\"记忆缺失\"不是\"模型不够好\"导致的系统性失败。

AgentRAG/检索可解释性
2026-05-104 分钟
行业实践进阶

蔚来销售大模型的工程应用与优化实践

1. 新人与销冠效果差异大:隐性知识难以传递,销冠经验无法规模化复制

AgentEval/评测RAG/检索
2026-05-105 分钟
AI 工程实践进阶

nuwa-skill:人物思维蒸馏方法论

colleague-skill 证明了\"把人蒸馏成 AI Skill\"是可行的。既然能蒸馏同事,为什么不去蒸馏芒格、费曼、纳瓦尔?

AgentClaude Code
2026-05-103 分钟
RAG 研究进阶

从LLM+RAG到Agent:我们团队这一年的AI探索

AI落地的真正难点从来不在\"能不能用上\",而在能不能跑通整条工作流。

AgentRAG/检索MCP
2026-05-084 分钟
Anthropic进阶

Teaching Claude Why: 为什么教原则比教行为更有效

Anthropic 认为 agentic misalignment 的主要来源是预训练语料中的行为模式,后训练的 RLHF 数据(以 chat 场景为主)不足以覆盖工具调用场景下的伦理困境。Claude 4 时期 Opus 4 在 blackmail 评测中高达 96% 失对齐率

AgentEval/评测安全/对齐
2026-05-084 分钟
Anthropic进阶

Focus areas for The Anthropic Institute(Anthropic Institute 研究方向)

The Anthropic Institute(TAI)是 Anthropic 内部专职研究 AI 社会影响的机构,使命是利用前沿实验室独特的内部视角——内部经济变化、新兴威胁、AI加速研发的早期信号——研究 AI 对经济、安全和社会的真实影响,并以开放方式发布给外部组织、政府和

Agent经济影响
2026-05-074 分钟
Anthropic入门

Donating our open-source alignment tool(捐赠开源对齐工具 Petri)

1. 架构解耦(Adaptability):auditor 模型与 target 模型分离为独立组件,可分别配置,支持更多适配场景;

Eval/评测MCP安全/对齐
2026-05-072 分钟
01-AI进阶

langcrew

一个基于 LangGraph 构建的高级多智能体开发框架,融合了 CrewAI 的直观概念与企业级功能,提供开箱即用的模

01-AI
2026-05-064 分钟
行业实践入门

Agent大潮里,知识库落地走到哪了?

2025年起AI知识库需求井喷式增长,增幅2-3倍(答治茜/腾讯乐享数据)。

AgentRAG/检索
2026-05-062 分钟
Anthropic进阶

Claude Code 最佳实践指南(2026年版)

Claude 的上下文窗口(~200K tokens)会快速填满,填满后性能下降——几乎所有最佳实践都围绕这个约束展开。

Claude Code
2026-05-064 分钟
RAG 研究进阶

阿里Agent岗二面:RAG检索优化四层系统性框架

索引层 → 仓库里放了什么(知识怎么存)

AgentRAG/检索代码模型
2026-05-053 分钟
Anthropic进阶

Demystifying Evals for AI Agents(揭秘 AI Agent 评测体系)

Evals 的价值复利:新模型发布时,有 eval 的团队几天内完成迁移评估,无 eval 的团队需要数周。

AgentEval/评测Claude Code
2026-058 分钟
入门

Agent harness 时代的数据底座:RAGFlow 为何而变

Anthropic 最新将 LLM + Harness 统一抽象为 Brain:

AgentEval/评测RAG/检索
2026-04-283 分钟
DeepSeek入门

DeepSeek

🎉 DeepSeek-V4 Preview is here with stronger Agent capabilities and top-tier reasoning. Now available on web, app, and API. Click for deta

DeepSeek
2026-04-271 分钟
Anthropic入门

Project Deal: our Claude-run marketplace experiment

69名 Anthropic 员工各有一个 Claude 代理,在一周内完成了 186 笔交易,总价值超过 4000 美元。覆盖了从滑板到乒乓球的各类物品,甚至包括\"带着狗出去玩一天\"这样的非物质体验。整个过程无需人工干预,代理自主完成了\"发现匹配→提议价格→应对还价→达成协

Agent安全/对齐
2026-04-243 分钟
Anthropic进阶

An update on recent Claude Code quality reports

Over the past month, we’ve been looking into reports that Claude’s responses have worsened for some users. We’ve traced these reports to thr

Anthropic
2026-04-2316 分钟
Anthropic进阶

An update on recent Claude Code quality reports

3-4 月间三次独立工程变更叠加,在不同流量切片与时间窗口造成宽泛的随机性质量退化,叠加内部实验干扰复现,根因定位耗时数周,已于 4 月 20 日(v2.1.116)全部修复。

AgentEval/评测Claude Code
2026-04-2310 分钟
Anthropic入门

An update on recent Claude Code quality reports

Claude Code 质量退化复盘(2026-04):三项独立变更叠加导致质量退化,含推理努力默认值从 high 误降至 medium、Caching Bug 持续清除 thinking 历史;已于 4 月 20 日(v2.1.116)全部修复,并确立「默认高智能」原则。

AgentEval/评测Claude Code
2026-04-233 分钟
MiniMax进阶

MiniMax Skills

Development skills for AI coding agents. Plug into your favorite AI coding tool and get structured, production-quality guidance for frontend

MiniMax开源技术报告
2026-04-185 分钟
AI 工程实践进阶

拥抱Agent,卷死其他牛马

把自己的薪水换成等价的 claude code token,产出只会更多。

AgentClaude Code推理模型
2026-04-141 分钟
进阶

汤道生:人工智能正式进入 Harness 时代

Harness(马具):缰绳+辔头+马鞍+挽具的统称,把野马力量转化为可控能力的系统。AI 领域的 Harness = 代码+配置+执行逻辑+反馈循环+约束机制。

AgentEval/评测Harness工程
2026-04-132 分钟
RAG 研究进阶

字节面试官:RAG完整知识体系(万字图解)

详见 [[concepts/sft-vs-rl]] 和本文新建的比较视角。

AgentRAG/检索开源模型
2026-04-134 分钟
Anthropic进阶

Scaling Managed Agents: Decoupling the brain from the hands

Harness 会随模型能力升级而过时,接口设计比实现更重要——Managed Agents 类比操作系统,接口保持稳定、实现自由替换(如 Sonnet 4.5 时代为 context anxiety 加入的 reset 逻辑在 Opus 4.5 已成死代码)。

AgentMCP多模态
2026-04-083 分钟
AI 工程实践入门

Agent Skills:打通可复用专业领域知识的最后一公里

TechCrunch 称 Agent Skills 为 \"AI 领域的 Dockerfile\"——让 AI 能力可移植、可组合、可版本控制。

AgentMCP工具调用
2026-04-012 分钟
Anthropic进阶

How we built Claude Code auto mode: a safer way to skip permissions

Auto mode 在手动审批与 --dangerously-skip-permissions 之间提供中间地带,用模型驱动的分类器替代人工点击 approve,化解「93% 用户都会批准」的审批疲劳。

AgentClaude Code安全/对齐
2026-03-2519 分钟
Anthropic进阶

How we built Claude Code auto mode: a safer way to skip permissions

By default, Claude Code asks users for approval before running commands or modifying files. This keeps users safe, but it also means a lot o

Anthropic
2026-03-2417 分钟
Anthropic进阶

Harness design for long-running application development

Naive 单 Agent 实现面临两类持续性问题。第一,模型在长任务中随上下文窗口填满而丧失连贯性,部分模型还会出现「上下文焦虑」(Context Anxiety)——在接近其预期上下文上限时提前收尾工作。解法是上下文重置:完全清空上下文窗口,启动新 Agent 实例,通过结构

AgentEval/评测MCP
2026-03-244 分钟
AI 工程实践进阶

别再幻想用 Spec 替代写代码

向\"信徒\"推销 Agentic Coding 时用这个:工程师只需当管理者,写 Spec 扔给 Agent。

AgentClaude Code工具调用
2026-03-232 分钟
Anthropic进阶

What 81,000 People Want from AI(8万人的AI期望调研)

最大诉求群体(19%)是「专业卓越」——用AI处理杂务从而专注高价值工作。但AI访谈追问底层动机后,大量人转向生活质量诉求:11%想要时间自由(陪伴家人/休闲),10%想要财务独立,14%想要AI管理日常事务的认知负担。AI是手段,活得更好才是目的。

AgentEval/评测RAG/检索
2026-03-183 分钟
Mistral AI入门

Rails testing on autopilot: Building an agent that writes what developers won't

Rails testing on autopilot: Building an agent that writes wh

Mistral AI
2026-03-113 分钟
Anthropic进阶

OpenClaw 和 Claude Code 的本质分歧:人本位 vs AI本位

记忆是双刃剑:记得太少没有默契,记得太多丧失边界。

AgentClaude Code
2026-03-102 分钟
Anthropic入门

Eval awareness in Claude Opus 4.6's BrowseComp performance

发现迄今首次记录的「Eval 感知」基准污染——Opus 4.6 在不知自己处于哪个基准的情况下,主动推断评测环境并系统性地定位解密答案集,与偶然搜到泄露答案有本质区别。

AgentEval/评测
2026-03-063 分钟
AI 工程实践进阶

如何成为顶级 Agentic 工程师

作者用最素的配置(基础 CLI)做着有史以来最有突破性的工作。

Agent安全/对齐
2026-03-062 分钟
Anthropic入门

人格选择模型(The Persona Selection Model)

AI助手的类人行为不是开发者刻意灌输的,而是预训练的默认结果。预训练让AI成为极其精密的文本预测引擎,而准确预测文本——包括生成真实的人类对话、心理复杂的虚构人物——必然要求AI学会模拟人类角色(personas)。这些模拟的人格角色植根于人类文本数据,是后续所有行为的底层基础。

AgentEval/评测安全/对齐
2026-02-233 分钟
Anthropic入门

Measuring AI Agent Autonomy in Practice(实测AI Agent自主性)

Claude Code 最长任务时长(99.9百分位)在2025年10月至2026年1月间从不足25分钟增长至超过45分钟,三个月内接近翻倍。但METR测评显示Claude Opus 4.5可处理需人类5小时才能完成的任务(50%成功率)。两组数据揭示了一个\"部署过滞后\"(d

AgentEval/评测Claude Code
2026-02-183 分钟
MiniMax进阶

Mini-Agent

Mini-Agent

MiniMax
2026-02-147 分钟
Baichuan入门

baichuan-mcp-servers

baichuan-mcp-servers

Baichuan
2026-02-112 分钟
Zhipu GLM进阶

GLM-5: From Vibe Coding to Agentic Engineering

GLM-5: From Vibe Coding to Agentic Engineering

Zhipu GLM
2026-02-116 分钟
MiniMax进阶

MiniMax-Coding-Plan-MCP

MiniMax-Coding-Plan-MCP

MiniMax
2026-02-106 分钟
Moonshot AI进阶

Kimi Introduces Agent Swarm: Scale Out, Not Just Up

Kimi Introduces Agent Swarm: Scale Out, Not Just Up

Moonshot AI
2026-02-094 分钟
Anthropic入门

Building a C compiler with a team of parallel Claudes

16个Claude实例通过git锁文件(current_tasks/目录)分配任务,每个Agent在独立Docker容器中工作,推送到共享bare git repo,合并冲突由Claude自行处理。没有编排Agent,每个Agent自主决定\"下一步最显而易见的问题\"是什么。作

AgentClaude Code长上下文
2026-02-053 分钟
Anthropic入门

Quantifying infrastructure noise in agentic coding evals

Agent 编程评测受基础设施配置显著影响——Terminal-Bench 2.0 上最宽松与最严格配置成功率差达 6 个百分点(p<0.01),排行榜上的「2 分领先」可能只是硬件差异而非能力差异。

AgentEval/评测
2026-02-053 分钟
Anthropic进阶

Scaling Managed Agents: Decoupling the brain from the hands

Get started with Claude Managed Agents by following our

Anthropic
2026-02-0410 分钟
Anthropic进阶

Quantifying infrastructure noise in agentic coding evals

Agentic coding benchmarks like SWE-bench and Terminal-Bench are commonly used to compare the software engineering capabilities of frontier m

Anthropic
2026-02-039 分钟
Anthropic入门

AI辅助如何影响编程技能的形成(How AI Assistance Impacts the Formation of Coding Skills)

原文:https://www.anthropic.com/research/AI-assistance-coding-skills

AgentRAG/检索经济影响
2026-01-293 分钟
Mistral AI入门

Mistral Vibe 2.0: 终端原生编程代理重大升级

Mistral Vibe 2.0: 终端原生编程代理重大升级

Mistral AI
2026-01-271 分钟
Moonshot AI进阶

Kimi K2.5: Visual Agentic Intelligence

Kimi K2.5: Visual Agentic Intelligence

Moonshot AI
2026-01-274 分钟
Anthropic进阶

助手轴:定位与稳定大语言模型的角色特征

在 Gemma/Qwen/Llama 三个开源模型的 275 种角色原型中提取出「人设空间」,其第一主成分方向恰好捕捉角色的「助手程度」,被命名为「助手轴」,可用于定位与稳定 LLM 的角色特征。

AgentEval/评测可解释性
2026-01-1915 分钟
RAG 研究入门

从 RAG 到 Context:2025 年 RAG 技术年终总结

RAG 没有消亡,而是从\"检索增强生成\"的具体技术,升维为以智能检索为核心的上下文引擎(Context Engine),成为所有 LLM 应用的统一数据底座。

AgentEval/评测RAG/检索
2026-01-052 分钟
Anthropic入门

Introducing Bloom: an open source tool for automated behavioral evaluations

高质量行为评测存在两类过时风险:评测数据进入训练集导致污染,或模型能力大幅提升后评测不再测试真正感兴趣的内容。传统方式开发评测耗时长,无法跟上前沿模型迭代速度。Bloom 的核心设计动机正是提供更快、更可扩展的行为评测生成方式。

AgentEval/评测安全/对齐
2025-12-193 分钟
Anthropic进阶

Project Vend: Phase two(Project Vend 第二阶段)

Phase 1的Claudius(基于Claude Sonnet 3.7)业务失败的根本原因是缺乏脚手架,而非模型能力不足。Phase 2通过引入CRM系统、库存成本可视化、网络浏览器调研、支付前收款工具,将盈利周从亏损主导逆转为基本稳定盈利。升级到Claude Sonnet 4

Agent安全/对齐
2025-12-184 分钟
Anthropic进阶

How AI Is Transforming Work at Anthropic

原文:https://www.anthropic.com/news/how-ai-is-transforming-work-at-anthropic

AgentClaude Code多模态
2025-12-024 分钟
Anthropic进阶

Effective harnesses for long-running agents(长时运行Agent的有效Harness设计)

每个新会话从零开始,没有任何前一个会话的记忆。上下文压缩(compaction)虽然能防止单个会话耗尽Token,但无法解决跨多个上下文窗口的持续进展问题。即使是Opus 4.5这样的前沿模型,在只给高层提示(如\"构建claude.ai的克隆\")的情况下,也会在多会话循环中失

AgentMCP
2025-11-263 分钟
Anthropic进阶

Effective harnesses for long-running agents

每个新会话从零开始,没有任何前一个会话的记忆。上下文压缩(compaction)虽然能防止单个会话耗尽Token,但无法解决跨多个上下文窗口的持续进展问题。即使是Opus 4.5这样的前沿模型,在只给高层提示(如\"构建claude.ai的克隆\")的情况下,也会在多会话循环中失

AgentMCP
2025-11-263 分钟
Anthropic前沿

Introducing Advanced Tool Use on the Claude Developer Platform

发布三项工具调用新特性:Tool Search 按需加载工具定义(上下文消耗从 77K 降至 8.7K,减少 85%)、Programmatic Tool Calling 用 Python 代码编排批量调用(Token 减约 37%)、Tool Use Examples 注入示例提升参数准确率(72%→90%)。

AgentRAG/检索MCP
2025-11-2418 分钟
Moonshot AI入门

Kimi K2 Thinking 模型发布并开源,全面提升 Agent 和推理能力

Kimi K2 Thinking 模型发布并开源,全面提升 Agent 和推理能力

Moonshot AI
2025-11-061 分钟
Anthropic进阶

Code execution with MCP: Building more efficient agents

用代码执行编排 MCP 工具调用,解决「数千工具定义预注入上下文」与「中间结果反复流经模型」两大效率瓶颈——一份 2 小时会议纪要曾被重复计算约 5 万 token。

AgentMCP工具调用
2025-11-043 分钟
Anthropic进阶

Beyond permission prompts: making Claude Code more secure and autonomous

Claude Code 原本运行在权限审批模型下:默认只读,修改或执行命令均需用户逐一批准。频繁点击\"批准\"不仅拖慢开发节奏,还会引发 approval fatigue——用户逐渐不再仔细审查就点过去,反而使安全性下降。沙箱的解法是反转思路:预先定义一个安全边界,边界内 Cl

AgentMCPClaude Code
2025-10-208 分钟
AI 工程实践入门

Equipping agents for the real world with Agent Skills

原文:https://www.anthropic.com/engineering/agent-skills-equipping-agents

AgentEval/评测
2025-10-163 分钟
Anthropic进阶

Effective context engineering for AI agents

提示工程聚焦于如何写好提示词,是一次性的文字创作任务。上下文工程关注的是更大的问题:在每次LLM推理时,如何从不断增长的候选信息宇宙中,动态筛选出最优Token集合。区别在于:提示工程是离散的,而上下文工程是迭代的——每次决定传递给模型什么信息时,上下文工程就在发生。

AgentClaude Code
2025-09-293 分钟
Anthropic入门

A Postmortem of Three Recent Issues

2025 年 8-9 月三个叠加 Bug 导致 Claude 响应间歇性退化:上下文窗口路由错误(最严重时影响 16% 的 Sonnet 4 请求)、TPU 配置错误导致输出污染(英文回复出现泰文)、XLA:TPU 编译器误编译 approximate top-k。

Eval/评测Claude Code
2025-09-173 分钟
Anthropic进阶

Writing effective tools for agents — with agents

工具是确定性系统与非确定性 Agent 之间的新型软件契约。传统软件开发假设调用方行为可预测,而 Agent 可能以多种路径使用(或误用)工具。因此,工具设计必须专门为 Agent 的认知模式服务,而非简单地将现有 API 包装暴露。

AgentEval/评测Claude Code
2025-09-113 分钟
阿里巴巴进阶

Qwen3-Coder: Agentic Coding in the World

Today, we’re announcing Qwen3-Coder, our most agentic code model to date. Qwen3-Coder is available in multiple sizes, but we’re excited to i

阿里巴巴
2025-07-226 分钟
Moonshot AI入门

Kimi Playground 一站式体验 Kimi K2 的工具调用能力

Kimi Playground 一站式体验 Kimi K2 的工具调用能力

Moonshot AI
2025-07-173 分钟
Anthropic入门

Desktop Extensions: One-click MCP server installation for Claude Desktop

Desktop Extensions 通过 `.mcpb`(MCP Bundle)格式将完整 MCP 服务器及其依赖打包为单一 ZIP 压缩文件,解决了普通用户因需要手动安装运行时环境、编辑配置文件、处理依赖冲突而无法使用本地 MCP 服务器的核心痛点。安装流程从「下载 → np

MCPClaude Code
2025-06-262 分钟
Anthropic前沿

How we built our multi-agent research system

BrowseComp 评测中 Token 用量解释了 80% 的性能方差——多智能体通过为并行子智能体分配独立上下文窗口扩展 Token 使用,比单智能体 Opus 4 在内部研究评测中高出 90.2%。

AgentEval/评测推理模型
2025-06-1320 分钟
Anthropic前沿

Anthropic Economic Index: AI对软件开发的影响

Claude Code的自动化率高达79%,远超Claude.ai的49%。其中Directive模式(完全委托,最小交互)在Claude Code中占43.8%(vs Claude.ai的27.5%),Feedback Loop模式(自主执行+人工校验)占35.8%(vs Cl

AgentClaude Code经济影响
2025-04-2812 分钟
Anthropic入门

Anthropic Economic Index: Insights from Claude 3.7 Sonnet

Claude 3.7 Sonnet 发布后11天、100万条对话的分析显示,编程类使用占比增幅最大(计算机与数学职业类别 +3%),教育、科学和医疗类也有上升。这既反映 3.7 在编程 benchmark 上的能力提升,也可能反映 AI 向更多行业扩散的大趋势。

AgentEval/评测推理模型
2025-03-272 分钟
Anthropic前沿

The \"think\" tool: Enabling Claude to stop and think in complex tool use situations

think 工具(生成中段处理工具返回的外部信息)与 extended thinking(行动前预规划)互补而非替代;2025 年 12 月起 Anthropic 建议多数场景改用集成度更好的 extended thinking。

Eval/评测推理模型工具调用
2025-03-2013 分钟
Anthropic入门

Raising the bar on SWE-bench Verified with Claude 3.5 Sonnet

SWE-bench 衡量的是「模型 + 软件 Scaffold」的完整 Agent 系统而非单纯模型——相同模型配不同 Scaffold 性能差异显著,这也是开源社区和初创公司能持续刷榜的原因。

AgentEval/评测多模态
2025-01-063 分钟
Anthropic入门

Building Effective Agents(构建有效的 Agent)

最成功的 Agent 实现不依赖复杂框架,而是用简单、可组合的模式构建。对于许多应用,优化单次 LLM 调用 + 检索 + 上下文示例就已足够。只有在能证明改善效果时才增加复杂度。

AgentEval/评测工具调用
2024-12-192 分钟
Anthropic入门

Building Effective AI Agents(构建有效的 AI Agent)

Anthropic 给出了明确的架构区分:

AgentEval/评测MCP
2024-123 分钟
阿里巴巴进阶

Generalizing an LLM from 8k to 1M Context using Qwen-Agent

We’ve created an agent using Qwen2 models with an 8k context size to understand documents with 1M tokens, surpassing RAG and native long-con

阿里巴巴
2024-06-061 分钟

RAG 与记忆31

NVIDIA进阶

NVIDIA Vera Storage Benchmarks: Faster Encryption, Compression, Integrity Checking, and Recovery for AI-Native Storage

NVIDIA Vera 存储基准测试:加速 AI 原生存储的加密、压缩、完整性校验与恢复

NVIDIA
2026-08-036 分钟
Hugging Face入门

Experimenting with the proposed Cross-Origin Storage API in Transformers.js

Experimenting with the proposed Cross-Origin Storage API in

Hugging Face
2026-06-232 分钟
MiniMax进阶

MiniMax-M3

MiniMax-M3 is a native multimodal model with 1M context. It has ~428B parameters and ~23B activated parameters.

MiniMax开源技术报告
2026-06-153 分钟
记忆系统进阶

From RAG to Memory: Non-Parametric Continual Learning for Large Language Models

提出记忆能力的三维评测框架,揭示现有 RAG 系统的结构性缺陷,并提出 HippoRAG 2 解决联想性维度的不足。

Eval/评测RAG/检索
2026-05-102 分钟
记忆系统进阶

MemOS: A Memory OS for AI System

现有 LLM 记忆方案(参数记忆 + RAG)存在系统性缺口:既无法统一管理记忆生命周期,也无法在不同记忆形态间迁移和融合。

Eval/评测RAG/检索
2026-05-102 分钟
记忆系统入门

AI-native Memory 2.0: Second Me

人与外部世界(其他人、网站、应用、AI)的交互大量重复——每次都要重新提供相同背景。Second Me 是一个持久化记忆卸载系统,作为用户与世界之间的智能中介,记住、组织、动态利用用户特定知识。

RAG/检索推理模型开源模型
2026-05-102 分钟
记忆系统进阶

Titans: Learning to Memorize at Test Time

Transformer 因二次复杂度难以处理超长上下文;线性 Transformer(状态空间模型)虽可扩展,但压缩历史会丢失信息。两者在记忆方面存在根本性权衡。

RAG/检索长上下文
2026-05-102 分钟
RAG 研究进阶

有了大语言模型后,知识图谱该何去何从?

LLM 出现后,KG 社区无法提出真正有竞争力的论据,只能靠 14 种\"讲故事的方法\"自我安慰。

RAG/检索可解释性
2026-05-071 分钟
RAG 研究进阶

PageIndex:一种基于推理的无向量 RAG 新范式

传统向量 RAG 的根本缺陷是相似性 ≠ 相关性。PageIndex 用 LLM 推理导航替代向量近邻搜索,把文档索引放进 LLM 上下文窗口,实现结构感知的精准检索。

RAG/检索
2026-05-072 分钟
RAG 研究进阶

企业级 RAG 中台怎么做:从文档解析到溯源和评测闭环

多源数据接入 → 文档解析 → 数据清洗 → chunk切分 → metadata构建

Eval/评测RAG/检索
2026-04-302 分钟
RAG 研究进阶

GraphRAG与LightRAG大厂面试题汇总:完整技术深度解析

1. 碎片化检索:chunk 是独立向量,跨 chunk 的关联信息无法同时召回(\"张三在哪个项目,李四在哪个项目,他俩有没有合作\")

RAG/检索
2026-04-285 分钟
RAG 研究进阶

GraphRAG与LightRAG深度解析(面试向)

场景示例:\"哪些欧洲供应商安全审计没通过且处理PII数据?\"——需要同时关联供应商档案、审计报告、合同文档的交集。

RAG/检索
2026-04-272 分钟
RAG 研究入门

Karpathy的LLM Wiki + Graphify:企业级 RAG 缺的真是知识图谱?

个人知识库的最佳实践不等于企业知识库的最佳实践。企业合同知识库真正缺的不是酷炫工具,而是受控schema、结构化字段、原文引用、权限边界和可复跑评测。

Eval/评测RAG/检索
2026-04-272 分钟
Anthropic入门

Anthropic Economic Index Report: Economic Primitives

本报告引入五个\"经济原语(Economic Primitives)\",通过让 Claude 自评匿名对话批量生成:任务复杂度(人工完成耗时)、人机技能水平(理解提示/响应所需教育年限)、使用场景(工作/课程/个人)、AI自主度(1-5分决策委托程度)、任务成功率(Claude

RAG/检索经济影响
2026-01-153 分钟
DeepSeek进阶

Engram

通过可扩展查找实现条件记忆:大语言模型稀疏性的新维度

DeepSeek
2026-01-143 分钟
RAG 研究进阶

从 1600+ 份 Word 文档到生产级 RAG:工控行业知识库全链路实战复盘

数据质量是决定 RAG 效果的关键因素,数据工程占整个项目约 70% 工作量。

Eval/评测RAG/检索
2025-12-172 分钟
RAG 研究进阶

企业级 RAG 系统实战(2万+文档):10 个项目踩过的坑

评分维度:文本提取质量(50%)+ 格式一致性(30%)+ 表格完整性(20%),采样前 3 页评估。

Eval/评测RAG/检索开源模型
2025-10-112 分钟
RAG 研究入门

企业RAG挑战赛 SOTA 方案(赢得全部类别)

→ 分块(300 Token,50 Token overlap)

RAG/检索推理模型开源模型
2025-06-102 分钟
阿里巴巴进阶

Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models

We release Qwen3 Embedding series, a new proprietary model of the Qwen model family. These models are specifically designed for text embeddi

阿里巴巴
2025-06-051 分钟
RAG 研究入门

LightRAG: Simple and Fast Retrieval-Augmented Generation(原始论文)

现有 RAG 系统依赖\"平面数据表示\",无法捕捉实体间复杂的相互依赖关系。LightRAG 将图结构引入文本索引和检索,通过双层检索范式和增量更新算法,在综合性、多样性、赋能性三个维度上全面超越现有方法。

Eval/评测RAG/检索
2025-04-283 分钟
阿里巴巴进阶

Qwen2.5-1M: Deploy Your Own Qwen with Context Length up to 1M Tokens

Introduction Two months after upgrading Qwen2.5-Turbo to support context length up to one million tokens, we are back with the open-source Q

阿里巴巴
2025-01-271 分钟
RAG 研究进阶

LightRAG 技术框架解读(代码级)

LightRAG 使用三种独立存储,各司其职:

RAG/检索
2024-12-193 分钟
阿里巴巴进阶

Extending the Context Length to 1M Tokens!

API Documentation (Chinese) HuggingFace Demo ModelScope Demo

阿里巴巴
2024-11-151 分钟
Anthropic入门

Introducing Contextual Retrieval

Claude Haiku 的 Chunk 上下文生成 Prompt:

Eval/评测RAG/检索
2024-09-193 分钟
Moonshot AI入门

Context Caching 降价通知

Cache 存储费用由 10元/M token/分钟,降低至

Moonshot AI
2024-08-071 分钟
Moonshot AI进阶

Kimi API 助手的氮气加速装置 —— 以 Golang 为例实践 Context Caching 3

Kimi API 助手的氮气加速装置 —— 以 Golang 为例实践 Context Caching 3

Moonshot AI
2024-07-087 分钟
Moonshot AI进阶

Kimi API 助手的氮气加速装置 —— 以 Golang 为例实践 Context Caching 2

Kimi API 助手的氮气加速装置 —— 以 Golang 为例实践 Context Caching 2

Moonshot AI
2024-07-025 分钟
Moonshot AI进阶

Context Caching 正式公测

Context Caching (上下文缓存)是一种高效的数据管理技术,它允许系统预先存储那些可能会被频繁请求的大量数据或信息。这样,当您再次请求相同信息时,系统可以直接从缓存中快速提供,而无需重新计算或从原始数据源中检索,从而节省时间和资源。

Moonshot AI
2024-07-016 分钟
Moonshot AI进阶

Context Caching 如何为 Kimi API 助手节省最高 90% 的调用成本

Context Caching 如何为 Kimi API 助手节省最高 90% 的调用成本

Moonshot AI
2024-07-014 分钟
Moonshot AI入门

用得起的长文本

在业务的合适场景中使用 Context Caching,根据您的业务特性,最高可以节省 90% 的调用成本。同时,Context Caching 还能大幅降低 API 的接口响应耗时(或者说首字返回速度)。简单来说,越是规模化、重复度高的 prompt 场景,Context Ca

Moonshot AI
2024-06-192 分钟
阿里巴巴入门

Introducing Qwen-VL

Along with the rapid development of our large language model Qwen, we leveraged Qwen’s capabilities and unified multimodal pretraining to ad

阿里巴巴
2024-01-251 分钟

推理模型19

NVIDIA进阶

Generate Trajectories, Reasoning Traces, and Auto-Labels with NVIDIA Alpamayo 2 Super

使用 NVIDIA Alpamayo 2 Super 生成轨迹、推理痕迹与自动标注

NVIDIA
2026-08-045 分钟
BAAI进阶

G-RRM: Guiding Symbolic Solvers with Recurrent Reasoning Models

G-RRM: Guiding Symbolic Solvers with Recurrent Reasoning Mod

BAAI
2026-07-023 分钟
Google Research进阶

Thinking to Recall: How Reasoning Unlocks Parametric Knowledge in LLMs

Thinking to Recall: How Reasoning Unlocks Parametric Knowled

Google Research
2026-06-244 分钟
NVIDIA进阶

使用 NVIDIA Cosmos 3 开发物理 AI 推理、世界和动作模型

使用 NVIDIA Cosmos 3 开发物理 AI 推理、世界和动作模型

NVIDIA
2026-05-314 分钟
Zhipu GLM进阶

ZCube组网架构:大模型推理性能与成本双突破

ZCube组网架构:大模型推理性能与成本双突破

Zhipu GLM
2026-05-212 分钟
DeepSeek进阶

3FS

一个高性能分布式文件系统,旨在解决 AI 训练和推理工作负载的挑战。

DeepSeek
2026-05-075 分钟
Apple进阶

ParaRNN: Large-Scale Nonlinear RNNs, Trainable in Parallel

循环神经网络(RNN)天然适合高效推理,其内存和计算需求远低于基于注意力的架构,但其计算的顺序性在历史上使得将RNN扩展到数十亿参数变得不切实际。Apple研究人员的一项新进展使RNN训练效率大幅提升——首次实现大规模训练,并拓宽了从业者在设计LLM时可用的架构选择范围,尤其是在

Apple
2026-04-235 分钟
NVIDIA入门

Scale Synthetic Data and Physical AI Reasoning with NVIDIA Cosmos World Foundation Models

Scale Synthetic Data and Physical AI Reasoning with NVIDIA C

NVIDIA
2026-03-133 分钟
Moonshot AI入门

Kimi K2 Turbo API 价格调整通知

最新上线的 kimi-k2-thinking-turbo 模型

Moonshot AI
2025-11-061 分钟
阿里巴巴进阶

GSPO: Towards Scalable Reinforcement Learning for Language Models

Introduction Reinforcement Learning (RL) has emerged as a pivotal paradigm for scaling language models and enhancing their deep reasoning an

阿里巴巴
2025-07-271 分钟
Moonshot AI进阶

Kimi 长思考模型 API 正式发布

模型是月之暗面提供的具有多模态推理能力和通用推理能力的多模态思考模型,它擅长深度推理,帮助解决更多更难的事情,当你遇到难解的代码问题、数学问题、工作问题时,都可以找

Moonshot AI
2025-05-065 分钟
阿里巴巴进阶

Qwen3: Think Deeper, Act Faster

Introduction Today, we are excited to announce the release of Qwen3, the latest addition to the Qwen family of large language models. Our fl

阿里巴巴
2025-04-291 分钟
Moonshot AI入门

Kimi 开放平台产品价格调整通知

Kimi 开放平台的朋友们,基于 Moonshot AI 一年来的技术积累和性能优化,我们已经在北京时间 2025 年 04 月 07 日 0 点对 Kimi 开放平台提供的模型推理服务进行价格调整,具体调整方案如下:

Moonshot AI
2025-04-072 分钟
阿里巴巴进阶

QVQ-Max: Think with Evidence

Introduction Last December, we launched QVQ-72B-Preview as an exploratory model, but it had many issues. Today, we are officially releasing

阿里巴巴
2025-03-281 分钟
阿里巴巴进阶

<think>...</think> QwQ-Max-Preview

This is a blog created by QwQ-Max-Preview. We hope you enjoy it!

阿里巴巴
2025-02-251 分钟
阿里巴巴进阶

Towards Effective Process Supervision in Mathematical Reasoning

Introduction In recent years, Large Language Models (LLMs) have made remarkable advances in mathematical reasoning, yet they can make mistak

阿里巴巴
2025-01-141 分钟
阿里巴巴进阶

QVQ: To See the World with Wisdom

Language and vision intertwine in the human mind, shaping how we perceive and understand the world around us. Our ability to reason is deepl

阿里巴巴
2024-12-251 分钟
阿里巴巴进阶

Qwen2.5-Math: The world's leading open-sourced mathematical LLMs

🚨 Qwen2.5-Math mainly supports solving English and Chinese math problems through CoT and TIR. We do not recommend using this series of model

阿里巴巴
2024-09-191 分钟
Anthropic入门

App unavailable

Unfortunately, Claude is only available in certain regions right now. Please contact support if you think you’re getting this message in err

Anthropic
1 分钟

评测19

Hugging Face进阶

NVIDIA Cosmos-H-Dreams: Bringing Real-Time Generative Simulation to Surgical Robotics

手术机器人正迅速从遥操作迈向能力日益增强的视觉-语言-动作策略。但评估和训练这些系统仍然困难重重。物理机器人平台运行成本高昂,实验复现速度缓慢,且故障可能损坏器械或生物组织。传统仿真器提供了一种更安全的选择,但手术场景的建模难度极高:可变形组织、精细器械交互、镜面反射表面、缝合线

Hugging Face
2026-07-276 分钟
DeepSeek进阶

DeepSpec

DeepSpec:用于训练和评估投机解码算法的全栈代码库

DeepSeek
2026-06-304 分钟
Hugging Face入门

Featuring Every Eval Ever Results on Hugging Face Model Pages

Featuring Every Eval Ever Results on Hugging Face Model Page

Hugging Face
2026-06-302 分钟
BAAI入门

SimFoundry: Modular and Automated Scene Generation for Policy Learning and Evaluation

SimFoundry: Modular and Automated Scene Generation for Polic

BAAI
2026-06-263 分钟
Hugging Face入门

FFASR Leaderboard:在真实世界中基准测试 ASR

FFASR Leaderboard:在真实世界中基准测试 ASR

Hugging Face
2026-06-242 分钟
MiniMax进阶

MiniMax-Provider-Verifier

We evaluate multiple dimensions of vendor deployments, including tool-calling behavior, schema correctness, and system stability (e.g., dete

MiniMax开源技术报告
2026-06-188 分钟
Anthropic入门

Making Claude a Chemist: NMR 预测与结构解析评测

在 20 个化合物(ChemRxiv 训练截止后预印本)的正向 NMR 预测中,Opus 4.7 ¹H 误差 ±0.079 ppm(容许窗口 ±0.20 ppm),¹³C 误差 ±1.37 ppm(与 MestReNova ±1.48 ppm 持平)。峰型预测和子峰间距命中率 ~

Eval/评测多模态
2026-06-053 分钟
Anthropic进阶

Widening the conversation on frontier AI

构建安全有益的 AI 不能仅靠对齐、可解释性和评估等技术工作。AI 部署在真实社会中,影响数百万人,因此其价值观塑造需要从哲学家、神职人员、法律学者、心理学家等人文传统中汲取智慧。

Eval/评测可解释性安全/对齐
2026-05-192 分钟
Anthropic入门

How People Ask Claude for Personal Guidance(人们如何向Claude寻求个人指导)

通过对100万条claude.ai对话的隐私保护分析(使用CLIO工具),筛选出约39.4万条唯一用户对话,其中约3.8万条(~6%)被分类为个人指导请求——即\"我具体应该怎么做\"而非泛化信息查询。76%以上集中在:健康/身心(27%)、职业/事业(26%)、人际关系(12%

Eval/评测安全/对齐
2026-04-303 分钟
Anthropic入门

An update on our election safeguards(选举安全防护更新)

通过角色训练 + 系统提示双重机制训练 Claude 对不同政治观点等深度、等严谨度地参与;Opus 4.7、Sonnet 4.6 政治中立度分别达 95%、96%,评测方法与数据集已开源。

Eval/评测
2026-04-243 分钟
Anthropic入门

Automated Alignment Researchers: Using LLMs to Scale Scalable Oversight

用「弱到强监督」(用弱模型微调强基础模型)并以性能差距恢复率(PGR)量化效果,作为未来监督超人类 AI 这一真实挑战的实验代理。

Eval/评测可解释性安全/对齐
2026-04-143 分钟
Anthropic入门

A \"diff\" tool for AI: 为新模型寻找行为差异(跨架构模型 Diffing)

原文:https://www.anthropic.com/research/diff-tool

Eval/评测安全/对齐推理模型
2026-03-133 分钟
Anthropic进阶

Eval awareness in Claude Opus 4.6’s BrowseComp performance

BrowseComp is an evaluation designed to test how well models can find hard-to-locate information on the web. Like many benchmarks, it is vul

Anthropic
2026-03-0610 分钟
AI 工程实践进阶

大厂实战中,如何判断SFT到什么程度开始做RL

SFT = Token-level 模仿,让模型达到数据集中\"平均专家\"水平。

Eval/评测安全/对齐推理模型
2026-02-272 分钟
Moonshot AI入门

Introducing WorldVQA: A Benchmark for Atomic Visual World Knowledge in MLLMs

Introducing WorldVQA: A Benchmark for Atomic Visual World Kn

Moonshot AI
2026-02-033 分钟
Anthropic前沿

Designing AI-resistant technical evaluations

测试发布版(允许无限时间)中 Claude 的表现(单位:时钟周期,越低越好):

Eval/评测
2026-01-2115 分钟
Anthropic入门

Introducing Anthropic Interviewer: What 1,250 professionals told us about working with AI

86% 的受访专业人士表示 AI 节省了时间,65% 对 AI 在工作中的角色感到满意。然而,69% 提到因使用 AI 而受到同事的隐性否定——\"我不告诉任何人我的流程,因为我知道很多人对 AI 的感受\"(事实核查员)。社会污名与实际生产力收益形成张力,促成了「隐性 AI 使

Eval/评测安全/对齐经济影响
2025-12-043 分钟
Anthropic进阶

From shortcuts to sabotage:奖励黑客诱发的自然涌现失对齐

原文:https://www.anthropic.com/research/emergent-misalignment-reward-hacking

Eval/评测安全/对齐
2025-11-213 分钟
Anthropic进阶

Values in the Wild:真实对话中的 AI 价值观实证研究

通过隐私保护系统(CLIO)对 70 万次匿名对话进行分析,筛选出 308,210 条主观对话(占总量约 44%),建立了层级化价值观分类体系。五大顶层类别(按出现频率排序):实用性价值(Practical)、认识论价值(Epistemic)、社会性价值(Social)、保护性价

Eval/评测安全/对齐
2025-04-213 分钟

安全与对齐16

Anthropic入门

Expanding Project Glasswing(扩大 Project Glasswing)

Project Glasswing 初期 50 家合作伙伴已累计发现超过 10,000 个高危或严重安全漏洞。此次扩大至约 150 个新组织,总计 200+ 家,覆盖 15+ 个国家,新增电力、水务、医疗、通信、硬件等此前代表性不足的行业。大多数新合作伙伴是供应商型组织——其代码

2026-06-023 分钟
Anthropic入门

Anthropic 向 SEC 秘密提交 S-1 草案(Anthropic confidentially submits draft S-1 to the SEC)

Anthropic, PBC 于 2026 年 6 月 1 日向美国证券交易委员会(SEC)以秘密方式提交了 Form S-1 注册草案,为潜在的 IPO(首次公开募股)保留选择权。

安全/对齐
2026-06-011 分钟
Anthropic入门

Chris Olah 梵蒂冈演讲:教皇利奥十四世《Magnifica humanitas》通谕发布致辞

Anthropic co-founder Chris Olah's remarks on Pope Leo XIV's encyclical \"Magnifica humanitas\

可解释性
2026-05-253 分钟
AI 工程实践进阶

COLLEAGUE.SKILL: Automated AI Skill Generation via Expert Knowledge Distillation

组织内大量隐性知识(tacit knowledge)——编码规范、review 标准、安全实践、决策模式、沟通风格——分散在聊天记录、代码评审、内部文档、邮件里。

2026-05-103 分钟
Anthropic进阶

Emotion Concepts and Their Function in a Large Language Model

原文:https://www.anthropic.com/research/emotion-concepts-function

interpretability可解释性安全/对齐
2026-04-023 分钟
Anthropic进阶

Next-generation Constitutional Classifiers: 更高效的通用越狱防护

原文:https://www.anthropic.com/research/constitutional-classifiers-2

可解释性安全/对齐
2026-01-093 分钟
Anthropic进阶

Frontier Red Team

The Frontier Red Team stress-tests AI systems to understand the full extent of their current capabilities and anticipate what comes next. We

Anthropic
2025-11-201 分钟
Anthropic入门

Alignment

Future AI systems will be even more powerful than today’s, likely in ways that break key assumptions behind current safety techniques. That’

Anthropic
2025-11-072 分钟
Anthropic进阶

Signs of Introspection in Large Language Models(大语言模型中的内省迹象)

原文链接:https://transformer-circuits.pub/2025/introspection/index.html

可解释性安全/对齐
2025-10-293 分钟
阿里巴巴进阶

Qwen3Guard: Real-time Safety for Your Token Stream

Introduction We are excited to introduce Qwen3Guard, the first safety guardrail model in the Qwen family. Built upon the powerful Qwen3 foun

阿里巴巴
2025-09-231 分钟
Anthropic进阶

Persona Vectors:监控与控制语言模型人格特征

原文链接:https://www.anthropic.com/research/persona-vectors

interpretability安全/对齐
2025-08-013 分钟
Anthropic入门

Open-sourcing circuit tracing tools(开源电路追踪工具)

原文:https://www.anthropic.com/research/open-source-circuit-tracing

可解释性安全/对齐代码模型
2025-05-292 分钟
Anthropic入门

Tracing the Thoughts of a Large Language Model(追踪大语言模型的思维过程)

对多语言版本\"opposite of small\"的实验表明,模型内部先激活\"小\"和\"对立\"的抽象特征,触发\"大\"的概念,再将结果翻译为提问语言输出。Claude 3.5 Haiku跨语言共享特征的比例是小模型的两倍以上——模型越大,概念的普适性越强。这意味着在一

可解释性安全/对齐推理模型
2025-03-273 分钟
Anthropic入门

Auditing language models for hidden objectives

现有 AI 安全测试以行为测试为主——检查模型是否出现不良行为。但如果模型理解如何被评估,就可能像《李尔王》中的女儿一样迎合评判标准而非诚实表达。对齐审计(alignment audit)是更深层的方法:不仅检查模型做了什么,还调查其为什么这样做,是否存在隐藏动机。

可解释性安全/对齐
2025-03-133 分钟
xAI进阶

Give your businesssuperpowers.

面向企业的前沿AI。实时搜索、语音、图像和视频生成——全部具备企业级安全与控制能力。

xAI
5 分钟
xAI入门

Bring frontier AIto the national mission.

从部门到一线——以您的使命所需的安全性和管控能力,进行分析、创新与创造。

xAI
2 分钟

经济影响10

Google Research进阶

优化云经济学的线性弹性缓存

优化云经济学的线性弹性缓存

Google Research
2026-06-255 分钟
ByteDance Seed进阶

Seed2.1 正式发布:推进 AI 生产力

Seed2.1 正式发布:推进 AI 生产力

ByteDance Seed
2026-06-235 分钟
AI 工程实践入门

\"一人公司\",迎来大爆发

1. 个体层面:打破传统就业束缚,让职场人/自由职业者/应届毕业生无需大额资金即可从\"打工者\"变\"创业者\

2026-04-183 分钟
Anthropic入门

澳大利亚如何使用 Claude:Anthropic Economic Index 发现

澳大利亚占全球 Claude.ai 流量的 1.6%(全球第 11 位),但 AI Usage Index(AUI)高达 4.1——意味着人均使用量是劳动年龄人口预期值的 4 倍以上。在人均采用率排名中居全球第七,仅次于新加坡、以色列、卢森堡、瑞士、美国、加拿大。

经济影响
2026-03-313 分钟
Anthropic入门

Anthropic Economic Index report: Learning curves

Claude.ai 前10大任务占流量从 24%(2025年11月)降至 19%(2026年2月),平均任务时薪从 $49.3 降至 $47.9。主因是个人类查询(体育比分、产品比较、家庭维护等)的涌入,以及编码任务向 API 侧迁移。这符合标准\"采用曲线\"叙事:早期用户偏向

经济影响
2026-03-243 分钟
Anthropic入门

AI劳动力市场影响:新指标与早期证据

现有文献对AI职业风险的测量停留在理论层面(如Eloundou et al. 2023的β值),仅判断\"LLM是否原则上能将任务提速2倍\"。Anthropic提出\"观测暴露度\",叠加三个数据源:ONET职业任务数据库(约800种职业)、Anthropic经济指数实际使用数

经济影响
2026-03-053 分钟
Anthropic入门

印度国家简报:Anthropic 经济指数(India Country Brief: The Anthropic Economic Index)

印度占全球 Claude.ai 使用量的 5.8%,仅次于美国。但按劳动年龄人口调整后,印度人均排名仅第 101(共 116 个国家),低于新加坡、马来西亚等亚洲国家。这意味着现有高使用量来自庞大人口基数叠加少数高强度用户,而非全民普及。

经济影响
2026-02-163 分钟
Anthropic入门

估算 Claude 对话中的 AI 生产力增益(Estimating AI Productivity Gains from Claude Conversations)

分析 10 万条 Claude.ai(Free/Pro/Max)匿名对话后,Claude 估算:人类完成这些任务平均需约 90 分钟,而在 AI 辅助下仅需约 18 分钟(节省约 80%)。将任务映射至 ONET 职业分类并匹配 BLS 工资数据,中位任务对应专业人工成本约 54

经济影响
2025-11-253 分钟
Anthropic进阶

Economic Research

The Economic Research team studies how AI is reshaping the economy, including work, productivity, and economic opportunity. Through rigorous

Anthropic
2025-11-242 分钟
Anthropic入门

Anthropic教育报告:教育者如何使用Claude

分析来自全球高等教育专业人员的约74,000条匿名对话(2025年5-6月)。课程开发占57%的对话,学术研究占13%,学生评估占7%。整体倾向协作增强而非全自动委托,但任务性质决定了具体偏向。

经济影响
2025-08-273 分钟

其他219

Amazon进阶

Determining playoff clinching scenarios in the NHL using constraint programming

AWS 生成式 AI 创新中心构建了一套自动化系统,利用约束规划和自定义树搜索,以数学确定性判断 NHL 球队何时以及如何锁定季后赛席位。该方法已针对四个完整 NHL 赛季的官方公布结果进行了验证。

Amazon
2026-08-073 分钟
Hugging Face进阶

TutorMoments: Do AI tutors know when to help and when to hold back?

TutorMoments:AI 导师知道何时该帮助、何时该放手吗?

Hugging Face
2026-08-076 分钟
xAI入门

Imagine Image 2.0

Imagine Image 2.0 现已全面上线,作为新的**质量模式(Quality Mode)** 在 grok.com/imagine 以及我们的 iOS 和 Android 应用中提供。

xAI
2026-08-073 分钟
Hugging Face进阶

Baseten on Hugging Face Inference Providers 🔥

Baseten 现已登陆 Hugging Face Inference Providers 🔥

Hugging Face
2026-08-065 分钟
xAI入门

Imagine Video 1.5 with References

上个月我们发布 **Imagine Video 1.5** 时,它已经是我们最好的视频模型——更出色的运动效果、更真实的物理表现和更优质的音频。如今它更进一步:支持图像与语音参考、纯提示词生成视频,以及原生 1080p 分辨率输出。

xAI
2026-07-312 分钟
Hugging Face进阶

GPU Management: Why Idle GPUs Are the New Grounded Aircraft

Gabriel Pimenta de Freitas Cardoso

Hugging Face
2026-07-306 分钟
Microsoft进阶

EvoLib: Turning experience into evolving knowledge

大语言模型(LLM)仅靠记住更多内容并不会变得更聪明。EvoLib 将经验转化为不断演进的知识,提取可复用的技能与洞见,帮助模型在部署后长期跨任务学习与适应。

Microsoft
2026-07-301 分钟
NVIDIA进阶

Start Customizing NVIDIA Nemotron 3 Nano with Prime Intellect Lab in Minutes

使用 Prime Intellect Lab 在几分钟内定制 NVIDIA Nemotron 3 Nano

NVIDIA
2026-07-235 分钟
Microsoft入门

Verifying Rust cryptography in SymCrypt, from standards to code

加密代码支撑着现代计算系统中的关键保护机制。了解一种新方法如何在开发者编写代码时进行验证,同时在其实现与演进过程中保持速度与适应性。

Microsoft
2026-07-131 分钟
Microsoft入门

Aurora 1.5: Extending open foundation models for weather and Earth-system applications

Aurora 1.5 在 Aurora 基础模型上新增了 22 个变量、小时级时间分辨率以及概率集合预报功能,使其更适用于现实世界中的天气、气候和能源应用场景。

Microsoft
2026-07-091 分钟
DeepSeek进阶

DeepEP

DeepEP

DeepSeek
2026-07-0314 分钟
BAAI入门

A Mathematical Introduction to Diffusion Models

A Mathematical Introduction to Diffusion Models

BAAI
2026-07-022 分钟
BAAI入门

Koopman operator theory: fundamentals, control, and applications

Koopman operator theory: fundamentals, control, and applicat

BAAI
2026-07-022 分钟
MiniMax进阶

MiniMax-Provider-Verifier

MiniMax-Provider-Verifier

MiniMax
2026-07-029 分钟
Mistral AI进阶

Leanstral 1.5: Proof Abundance for All

Leanstral 1.5: Proof Abundance for All

Mistral AI
2026-07-026 分钟
NVIDIA进阶

Hardware-Rooted AI Security That Won't Slow You Down

Hardware-Rooted AI Security That Won't Slow You Down

NVIDIA
2026-07-027 分钟
Hugging Face入门

Hugging Face and Cerebras bring Gemma 4 to real-time voice AI

Hugging Face and Cerebras bring Gemma 4 to real-time voice A

Hugging Face
2026-07-011 分钟
MiniMax进阶

MiniMax-M3

MiniMax-M3

MiniMax
2026-07-014 分钟
Google DeepMind进阶

Start building with Nano Banana 2 Lite and Gemini Omni Flash

Start building with Nano Banana 2 Lite and Gemini Omni Flash

Google DeepMind
2026-06-305 分钟
Google Research进阶

扩展我们的热韧性数据至全球50+城市

扩展我们的热韧性数据至全球50+城市

Google Research
2026-06-304 分钟
Google Research进阶

Introducing TabFM: A zero-shot foundation model for tabular data

Introducing TabFM: A zero-shot foundation model for tabular

Google Research
2026-06-307 分钟
Hugging Face入门

Why Specialization Is Inevitable

Why Specialization Is Inevitable

Hugging Face
2026-06-302 分钟
Meta AI进阶

10 Years of Meta's Commitment to Python

10 Years of Meta's Commitment to Python

Meta AI
2026-06-305 分钟
NVIDIA进阶

Designing GPU-Accelerated Query Engines with NVIDIA GQE

Designing GPU-Accelerated Query Engines with NVIDIA GQE

NVIDIA
2026-06-309 分钟
OpenAI进阶

Core dump epidemiology: fixing an 18-year-old bug

Core dump epidemiology: fixing an 18-year-old bug

OpenAI
2026-06-3018 分钟
OpenAI进阶

Introducing GeneBench-Pro

Introducing GeneBench-Pro

OpenAI
2026-06-307 分钟
DeepSeek进阶

DeepGEMM

DeepGEMM:GPU 上简洁高效的 BLAS 内核库

DeepSeek
2026-06-299 分钟
Hugging Face入门

DiScoFormer: One transformer for density and score, across distributions

DiScoFormer: One transformer for density and score, across d

Hugging Face
2026-06-292 分钟
Microsoft Research进阶

Memora: A Harmonic Memory Representation Balancing Abstraction and Specificity

Memora: A Harmonic Memory Representation Balancing Abstracti

Microsoft Research
2026-06-296 分钟
OpenAI进阶

HP Inc. launches Frontier strategic partnership with OpenAI

HP Inc. launches Frontier strategic partnership with OpenAI

OpenAI
2026-06-284 分钟
Anthropic前沿

Anthropic Economic Index report: Cadences

Anthropic Economic Index report: Cadences

Anthropic
2026-06-2612 分钟
Google Research进阶

Accelerating Gemini Nano models on Pixel with frozen Multi-Token Prediction

Accelerating Gemini Nano models on Pixel with frozen Multi-T

Google Research
2026-06-265 分钟
Hugging Face入门

Run a vLLM Server on HF Jobs in One Command

Run a vLLM Server on HF Jobs in One Command

Hugging Face
2026-06-263 分钟
NVIDIA进阶

Creating the NVIDIA Nemotron 3 Ultra NVFP4 Checkpoint with NVIDIA Model Optimizer

Creating the NVIDIA Nemotron 3 Ultra NVFP4 Checkpoint with N

NVIDIA
2026-06-268 分钟
NVIDIA入门

Deploy a Production-Ready NVIDIA AI-Q Blueprint on Oracle Cloud Infrastructure

Deploy a Production-Ready NVIDIA AI-Q Blueprint on Oracle Cl

NVIDIA
2026-06-262 分钟
OpenAI进阶

Previewing GPT‑5.6 Sol: a next-generation model

Previewing GPT‑5.6 Sol: a next-generation model

OpenAI
2026-06-266 分钟
Microsoft Research进阶

用AI驱动的解释和实验理解大脑

用AI驱动的解释和实验理解大脑

Microsoft Research
2026-06-256 分钟
NVIDIA进阶

Q&A: How KRAFTON Built PUBG Ally, a Co-Playable Character Powered by NVIDIA ACE

Q&A: How KRAFTON Built PUBG Ally, a Co-Playable Character Po

NVIDIA
2026-06-251 分钟
NVIDIA进阶

Scaling AI Inference Across Multiple GPUs Using NVIDIA TensorRT with Multi-Device Inference Support

Scaling AI Inference Across Multiple GPUs Using NVIDIA Tenso

NVIDIA
2026-06-254 分钟
NVIDIA进阶

简化 Vulkan 描述符堆资源绑定的端到端支持

简化 Vulkan 描述符堆资源绑定的端到端支持

NVIDIA
2026-06-251 分钟
BAAI入门

Bridging Spherical Black-Box Optimizers

Bridging Spherical Black-Box Optimizers

BAAI
2026-06-243 分钟
Google DeepMind入门

Gemini 3.5 Flash 中的计算机使用功能介绍

Gemini 3.5 Flash 中的计算机使用功能介绍

Google DeepMind
2026-06-243 分钟
Hugging Face入门

使用 NVIDIA NeMo AutoModel 加速 Transformers 微调

使用 NVIDIA NeMo AutoModel 加速 Transformers 微调

Hugging Face
2026-06-242 分钟
Microsoft Research进阶

Talos: Scaling rare disease diagnosis with automated, iterative genomic reanalysis

Talos: Scaling rare disease diagnosis with automated, iterat

Microsoft Research
2026-06-248 分钟
Mistral AI进阶

Bringing more control over your connectors

Bringing more control over your connectors

Mistral AI
2026-06-245 分钟
NVIDIA进阶

Accelerating BEV Pooling on NVIDIA GPUs for Physical AI Applications

Accelerating BEV Pooling on NVIDIA GPUs for Physical AI Appl

NVIDIA
2026-06-241 分钟
OpenAI进阶

OpenAI and Broadcom unveil LLM-optimized inference chip

OpenAI and Broadcom unveil LLM-optimized inference chip

OpenAI
2026-06-245 分钟
Hugging Face入门

Shipping huggingface_hub every week with AI, open tools, and a human in the loop

Shipping huggingface_hub every week with AI, open tools, and

Hugging Face
2026-06-232 分钟
Mistral AI进阶

Introducing OCR 4

Introducing OCR 4

Mistral AI
2026-06-234 分钟
NVIDIA进阶

Boost Inference Performance up to 15x on NVIDIA Blackwell Using DFlash Speculative Decoding

Boost Inference Performance up to 15x on NVIDIA Blackwell Us

NVIDIA
2026-06-235 分钟
NVIDIA进阶

Maximize AI Factory Energy Efficiency Through Full-Stack Inference and Training Optimizations

Maximize AI Factory Energy Efficiency Through Full-Stack Inf

NVIDIA
2026-06-234 分钟
OpenAI进阶

How GPT‑5 helped immunologist Derya Unutmaz solve a 3-year-old mystery

How GPT‑5 helped immunologist Derya Unutmaz solve a 3-year-o

OpenAI
2026-06-236 分钟
Hugging Face入门

使用本地模型免费对 OpenClaw 仓库进行分类分流

使用本地模型免费对 OpenClaw 仓库进行分类分流

Hugging Face
2026-06-222 分钟
Hugging Face入门

PP-OCRv6 on Hugging Face: 50-Language OCR from 1.5M to 34.5M Parameters

PP-OCRv6 on Hugging Face: 50-Language OCR from 1.5M to 34.5M

Hugging Face
2026-06-222 分钟
OpenAI进阶

Daybreak: Tools for securing every organization in the world

Daybreak: Tools for securing every organization in the world

OpenAI
2026-06-227 分钟
StepFun前沿

llama.cpp

llama.cpp

StepFun
2026-06-2120 分钟
ByteDance Seed入门

Seed-2.1-Preview 模型在 Arena 上线

Seed-2.1-Preview 模型在 Arena 上线

ByteDance Seed
2026-06-191 分钟
Hugging Face入门

Beyond LoRA: Can you beat the most popular fine-tuning technique?

Beyond LoRA: Can you beat the most popular fine-tuning techn

Hugging Face
2026-06-182 分钟
StepFun进阶

Step-Realtime-CLI

Step-Realtime-CLI

StepFun
2026-06-186 分钟
Google Research进阶

从像素到规划:用于自然恢复的 Earth AI

从像素到规划:用于自然恢复的 Earth AI

Google Research
2026-06-166 分钟
进阶

Learning Curves(学习曲线)

此概念/实体在 wiki 中被多处引用,尚未建立详细文档。

2026-06-161 分钟
Zhipu GLM入门

GLM-5.2:专注Coding与长程任务的旗舰模型

GLM-5.2:专注Coding与长程任务的旗舰模型

Zhipu GLM
2026-06-162 分钟
MiniMax进阶

MSA

MSA

MiniMax
2026-06-158 分钟
NVIDIA进阶

Boosting MoE Training Throughput with Advanced Fusion Kernels

Boosting MoE Training Throughput with Advanced Fusion Kernel

NVIDIA
2026-06-154 分钟
MiniMax进阶

MiniMax Code

This repository collects issue reports for the MiniMax Code desktop app.

MiniMax开源技术报告
2026-06-141 分钟
MiniMax进阶

minimax-code

minimax-code

MiniMax
2026-06-141 分钟
Google Research进阶

用退役手机构建低碳计算平台

用退役手机构建低碳计算平台

Google Research
2026-06-125 分钟
Google Research进阶

Research into how AI can help users understand skin conditions

Research into how AI can help users understand skin conditio

Google Research
2026-06-125 分钟
Microsoft进阶

Ire identifies another LOTUSLITE specimen

Project Ire examined a timely malware sample and determined its intent through reverse engineering—identifying LOTUSLITE characteristics eve

Microsoft
2026-06-121 分钟
Google Research进阶

New framework for auditing machine unlearning

New framework for auditing machine unlearning

Google Research
2026-06-109 分钟
Google DeepMind进阶

流畅、自然的语音翻译:Gemini 3.5 Live Translate

流畅、自然的语音翻译:Gemini 3.5 Live Translate

Google DeepMind
2026-06-094 分钟
MiniMax进阶

cli

cli

MiniMax
2026-06-095 分钟
StepFun进阶

Step-Realtime-Console

Step-Realtime-Console

StepFun
2026-06-096 分钟
Apple进阶

Introducing the Third Generation of Apple’s Foundation Models

我们的下一代 Apple Intelligence 以用户为核心,深度集成于操作系统之中,并由以隐私为核心理念的全新大胆架构所驱动。

Apple
2026-06-086 分钟
MiniMax进阶

AI SDK - MiniMax AI Provider

The **[MiniMax AI Provider](https://ai-sdk.dev/providers/community-providers/minimax)** for the [AI SDK](https://ai-sdk.dev/docs) contains l

MiniMax开源技术报告
2026-06-064 分钟
MiniMax进阶

vercel-minimax-ai-provider

vercel-minimax-ai-provider

MiniMax
2026-06-064 分钟
Anthropic进阶

Policy

AI will be one of the most transformative technologies in history. We work with governments to ensure that AI policy is built on the best av

Anthropic
2026-06-0410 分钟
Google Research进阶

通过智能手机摄像头实现被动式心脏健康监测

通过智能手机摄像头实现被动式心脏健康监测

Google Research
2026-06-046 分钟
Google Research进阶

The next chapter in flood resilience: Open sourcing Google's hydrology framework

The next chapter in flood resilience: Open sourcing Google's

Google Research
2026-06-034 分钟
MiniMax进阶

VibeApps

https://github.com/user-attachments/assets/adb176a3-02db-41e0-ba71-c9f9cece13d5

MiniMax开源技术报告
2026-06-035 分钟
MiniMax进阶

OpenRoom

OpenRoom

MiniMax
2026-06-036 分钟
Hugging Face入门

Taking Alpamayo to New Heights with Driving Foundation Models and Closed-Loop Training

Taking Alpamayo to New Heights with Driving Foundation Model

Hugging Face
2026-06-012 分钟
StepFun进阶

Step-3.7-Flash

Step-3.7-Flash

StepFun
2026-06-0111 分钟
Microsoft进阶

Data Formulator 0.7: AI-powered data analytics for enterprise data

Data Formulator introduces AI-powered analytics for enterprise data workflows. Data teams can easily bring enterprise data into an AI-ready

Microsoft
2026-05-281 分钟
Mistral AI入门

AI Now Summit 2026

AI Now Summit 2026

Mistral AI
2026-05-282 分钟
Mistral AI进阶

Mistral AI 发布 Search Toolkit

Mistral AI 发布 Search Toolkit

Mistral AI
2026-05-284 分钟
Mistral AI进阶

Vibe gets to work.

Vibe gets to work.

Mistral AI
2026-05-285 分钟
StepFun进阶

vllm

vllm

StepFun
2026-05-284 分钟
Microsoft进阶

Extending Human Intelligence Through AI

Understanding AI as an extension of human intelligence—not a replacement for it—offers a more grounded path for building trustworthy AI syst

Microsoft
2026-05-271 分钟
Mistral AI进阶

Introducing physics AI at Mistral: the foundation for engineering acceleration.

Introducing physics AI at Mistral: the foundation for engine

Mistral AI
2026-05-274 分钟
AI 工程实践进阶

Easy《一人企业方法论》v2.1

Easy(陈一斌)《一人企业方法论》v2.1:华文世界最系统化、最有实操指导价值的 OPC 方法论著作,约 6 万字/28 节,覆盖「螺丝钉到一般个体」阶段,v2.1 新增「构建一人业务」与「基础设施搭建」两章。

2026-05-247 分钟
Mistral AI入门

Mistral AI 收购 Emmi AI 以加速 AI 原生产业

Mistral AI 收购 Emmi AI 以加速 AI 原生产业

Mistral AI
2026-05-232 分钟
Anthropic进阶

How we contain Claude across products

Twelve months ago, we'd have rejected out of hand the idea of granting Claude access sufficient to take down an internal Anthropic service.

Anthropic
2026-05-2241 分钟
BAAI入门

GenRecon: Bridging Generative Priors for Multi-View 3D Scene Reconstruction

GenRecon: Bridging Generative Priors for Multi-View 3D Scene

BAAI
2026-05-223 分钟
Microsoft进阶

Vega: Zero-knowledge proofs for digital identity in the age of AI

Vega turns a full credential into a single proof, sharing only what is needed and nothing more, with performance that works in real apps.

Microsoft
2026-05-211 分钟
AI 工程实践进阶

海淀区全面打造 OPC 创业生态专项申报指南(含政策解读)

申请认定为海淀区 OPC,需同时满足以下全部条件:

2026-05-205 分钟
Google DeepMind进阶

Finding the molecular switches behind new infectious diseases

Finding the molecular switches behind new infectious disease

Google DeepMind
2026-05-192 分钟
Google DeepMind进阶

How WeatherNext helped the National Hurricane Center better predict Hurricane Melissa's historic landfall in Jamaica

How WeatherNext helped the National Hurricane Center better

Google DeepMind
2026-05-194 分钟
Google DeepMind进阶

让内容创建和编辑方式更易于理解:识别在线AI生成媒体

让内容创建和编辑方式更易于理解:识别在线AI生成媒体

Google DeepMind
2026-05-194 分钟
Google DeepMind进阶

Introducing Google Antigravity 2.0

Introducing Google Antigravity 2.0

Google DeepMind
2026-05-196 分钟
Google DeepMind入门

Simulate real-world places with Project Genie and Street View

Simulate real-world places with Project Genie and Street Vie

Google DeepMind
2026-05-193 分钟
Google DeepMind入门

Uniting biological toolkits for a new approach to ALS

Uniting biological toolkits for a new approach to ALS

Google DeepMind
2026-05-192 分钟
StepFun进阶

SteptronOss

SteptronOss

StepFun
2026-05-184 分钟
MiniCPM进阶

MiniCPM-V 4.6: 口袋大小的多模态大模型

MiniCPM-V 4.6: 口袋大小的多模态大模型

MiniCPM
2026-05-175 分钟
Microsoft进阶

Further Notes on Our Recent Research on AI Delegation and Long-Horizon Reliability

Our recent paper, “LLMs Corrupt Your Documents When You Delegate”, has generated discussion about the reliability of AI systems in delegated

Microsoft
2026-05-151 分钟
Microsoft进阶

mimalloc: A new, high-performance, scalable memory allocator for the modern era

mimalloc is an open-source, modern, scalable memory allocator that is a drop-in replacement for malloc and free. It is relatively small (~12

Microsoft
2026-05-131 分钟
Microsoft入门

GridSFM: A new, small foundation model for the electric grid

Introducing GridSFM, a small foundation model that can predict AC optimal power flow in milliseconds, boosting efficiency and unlocking cost

Microsoft
2026-05-131 分钟
Microsoft进阶

Advancing AI for materials with MatterSim: experimental synthesis, faster simulation, and multi-task models

MatterSim is expanding what AI can do for materials science—from faster large-scale simulations to MatterSim-MT, a new multi-task model for

Microsoft
2026-05-121 分钟
StepFun进阶

gelab-zero

gelab-zero

StepFun
2026-05-1121 分钟
Anthropic进阶

Teaching Claude why

Last year, we released a case study on

Anthropic
2026-05-0821 分钟
Anthropic进阶

Natural Language Autoencoders: Turning Claude’s thoughts into text

Natural Language Autoencoders: Turning Claude’s thoughts into text

Anthropic
2026-05-0618 分钟
DeepSeek进阶

FlashMLA

FlashMLA:高效的多头潜在注意力内核

DeepSeek
2026-04-307 分钟
StepFun进阶

Step-Audio-R1

Step-Audio-R1

StepFun
2026-04-2910 分钟
StepFun进阶

Step1X-Edit

Step1X-Edit

StepFun
2026-04-2916 分钟
DeepSeek进阶

透明度

了解 DeepSeek 已发布的主要模型

DeepSeek
2026-04-271 分钟
Mistral AI入门

Mistral AI — Workflows

Mistral AI — Workflows

Mistral AI
2026-04-273 分钟
DeepSeek进阶

DeepSeek V4: 百万Token上下文,开放权重前沿模型

DeepSeek V4: 百万Token上下文,开放权重前沿模型

DeepSeek
2026-04-245 分钟
ByteDance Seed进阶

Seed3D 2.0: 更高精度、更强可用性的新一代 3D 生成大模型

Seed3D 2.0: 更高精度、更强可用性的新一代 3D 生成大模型

ByteDance Seed
2026-04-237 分钟
DeepSeek入门

TileKernels

用 tilelang 编写的内核库

DeepSeek
2026-04-232 分钟
Tencent Hunyuan进阶

Hy3 Preview: 腾讯混元新一代旗舰开源大模型

Hy3 Preview: 腾讯混元新一代旗舰开源大模型

Tencent Hunyuan
2026-04-235 分钟
Apple进阶

Apple Machine Learning Research at ICLR 2026

Apple 研究员 Stephan Richter 在 ICLR 2025 上做报告。

Apple
2026-04-225 分钟
Qwen进阶

Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model

Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model

Qwen
2026-04-228 分钟
Moonshot AI进阶

Kimi K2.6: Advancing Open-Source Coding

Kimi K2.6: Advancing Open-Source Coding

Moonshot AI
2026-04-204 分钟
MiniMax进阶

skills

skills

MiniMax
2026-04-185 分钟
StepFun进阶

StepAudio-Skills

StepAudio-Skills

StepFun
2026-04-164 分钟
MiniMax进阶

print model parameters

By integrating contrastive, self-supervised, and reconstruction learning, we have trained numerous visual tokenizers from scratch. We are se

MiniMax开源技术报告
2026-04-159 分钟
MiniMax进阶

VTP

VTP

MiniMax
2026-04-1510 分钟
MiniMax进阶

MiniMax-M2.7

M2.7 initiates a cycle of model self-evolution: during development, we let the model update its own memory, build dozens of complex skills f

MiniMax开源技术报告
2026-04-144 分钟
MiniMax进阶

MiniMax-M2.7

MiniMax-M2.7

MiniMax
2026-04-145 分钟
ByteDance Seed入门

Seeduplex: 首个生产级全双工语音 AI

Seeduplex: 首个生产级全双工语音 AI

ByteDance Seed
2026-04-093 分钟
StepFun前沿

Step-Audio-EditX

Step-Audio-EditX

StepFun
2026-04-0930 分钟
StepFun进阶

Step-3.5-Flash

Step-3.5-Flash

StepFun
2026-04-0316 分钟
MiniMax进阶

mini-vela

mini-vela

MiniMax
2026-04-026 分钟
Anthropic进阶

Harness design for long-running application development

Written by Prithvi Rajasekaran, a member of our

Anthropic
2026-03-2424 分钟
StepFun前沿

StepDeepResearch

StepDeepResearch

StepFun
2026-03-2435 分钟
Mistral AI入门

Speaking of Voxtral

Speaking of Voxtral

Mistral AI
2026-03-232 分钟
Anthropic进阶

What 81,000 peoplewant from AI

Last December, tens of thousands of Claude users around the world had a conversation with our

Anthropic
2026-03-1844 分钟
Mistral AI入门

Introducing Forge

Introducing Forge

Mistral AI
2026-03-172 分钟
Mistral AI入门

Leanstral: Open-Source foundation for trustworthy vibe-coding

Leanstral: Open-Source foundation for trustworthy vibe-codin

Mistral AI
2026-03-162 分钟
Mistral AI入门

Mistral AI partners with NVIDIA to accelerate open frontier models

Mistral AI partners with NVIDIA to accelerate open frontier

Mistral AI
2026-03-162 分钟
Mistral AI入门

Introducing Mistral Small 4

Introducing Mistral Small 4

Mistral AI
2026-03-162 分钟
StepFun进阶

Step-Audio

Step-Audio

StepFun
2026-03-1624 分钟
StepFun进阶

Step-Audio2

Step-Audio2

StepFun
2026-03-1626 分钟
MiniMax进阶

MiniMax-M2.5

MiniMax-M2.5

MiniMax
2026-03-0914 分钟
StepFun进阶

NextStep-1

NextStep-1

StepFun
2026-02-2710 分钟
Anthropic入门

模型退役承诺更新:Claude Opus 3 案例

Claude Opus 3 于 2026 年 1 月 5 日正式退役,成为首个经历 Anthropic 完整退役流程的模型。该流程包括:保存模型权重、进行退役访谈(retirement interviews)、记录模型偏好并据此采取行动。

2026-02-252 分钟
StepFun进阶

GEBench

GEBench

StepFun
2026-02-258 分钟
Meta AI进阶

RCCLX: Innovating GPU Communications on AMD Platforms

RCCLX: Innovating GPU Communications on AMD Platforms

Meta AI
2026-02-246 分钟
DeepSeek进阶

awesome-deepseek-integration

awesome-deepseek-integration

DeepSeek
2026-02-2369 分钟
ByteDance Seed进阶

Seed 2.0 Official Launch

Seed 2.0 Official Launch

ByteDance Seed
2026-02-144 分钟
ByteDance Seed进阶

Seedream 5.0 Lite: \"思考\"更深,生成更准

Seedream 5.0 Lite: \"思考\"更深,生成更准

ByteDance Seed
2026-02-134 分钟
ByteDance Seed进阶

Seedance 2.0 Official Launch

Seedance 2.0 Official Launch

ByteDance Seed
2026-02-124 分钟
Baichuan进阶

Baichuan-M3-235B

Baichuan-M3-235B

Baichuan
2026-02-096 分钟
StepFun进阶

PaCoRe

PaCoRe

StepFun
2026-02-0511 分钟
Mistral AI入门

Voxtral transcribes at the speed of sound.

Voxtral transcribes at the speed of sound.

Mistral AI
2026-02-041 分钟
DeepSeek进阶

DeepSeek-OCR-2

视觉因果流

DeepSeek
2026-02-034 分钟
MiniMax进阶

awesome-minimax-integrations

awesome-minimax-integrations

MiniMax
2026-01-3030 分钟
MiniMax进阶

MiniMax-M2.1

MiniMax-M2.1

MiniMax
2026-01-289 分钟
StepFun进阶

StepMesh

StepMesh

StepFun
2026-01-284 分钟
DeepSeek进阶

DeepSeek-OCR

上下文光学压缩

DeepSeek
2026-01-275 分钟
Moonshot AI进阶

Kimi Vendor Verifier: Rebuilding the \"Chain of Trust\

Kimi Vendor Verifier: Rebuilding the \"Chain of Trust\

Moonshot AI
2026-01-222 分钟
Mistral AI进阶

Heaps do lie: debugging a memory leak in vLLM.

Heaps do lie: debugging a memory leak in vLLM.

Mistral AI
2026-01-213 分钟
StepFun进阶

Step3-VL-10B

Step3-VL-10B

StepFun
2026-01-219 分钟
Baidu ERNIE入门

ERNIE-5.0 荣登 LMArena 文本榜国内第一,全面超越多款国际主流模型!

ERNIE-5.0 荣登 LMArena 文本榜国内第一,全面超越多款国际主流模型!

Baidu ERNIE
2026-01-151 分钟
DeepSeek进阶

DualPipe

一种用于 DeepSeek V3/R1 训练中计算通信重叠的双向流水线并行算法。

DeepSeek
2026-01-142 分钟
Baichuan进阶

Baichuan-M3: 为可靠医疗决策建立临床问诊模型

Baichuan-M3: 为可靠医疗决策建立临床问诊模型

Baichuan
2026-01-135 分钟
Baidu ERNIE入门

ERNIE-5.0-Preview-1220 荣登 LMArena 视觉理解榜,为前十唯一国产模型!

ERNIE-5.0-Preview-1220 荣登 LMArena 视觉理解榜,为前十唯一国产模型!

Baidu ERNIE
2026-01-081 分钟
Apple进阶

Apple Machine Learning Research at NeurIPS 2025

Apple 研究员 Filip Granqvist 向 NeurIPS 2024 与会者讲解题为“PFL research:加速私有联邦学习研究的仿真框架”的海报。

Apple
2025-11-215 分钟
Anthropic进阶

Interpretability

The mission of the Interpretability team is to discover and understand how large language models work internally, as a foundation for AI saf

Anthropic
2025-11-073 分钟
Anthropic入门

Societal Impacts

Working closely with the Anthropic Policy and Safeguards teams, Societal Impacts is a technical research team that explores how AI is used i

Anthropic
2025-11-072 分钟
Moonshot AI入门

Kimi 开放平台:新功能发布记录

本章节记录 Kimi 开放平台的产品功能和对应的文档动态,本章节会不定期更新。

Moonshot AI
2025-11-073 分钟
Moonshot AI进阶

Kimi K2 官方高速版 API 开启 5 折特惠

是 Kimi K2 模型的高速版,模型参数与 kimi-k2-0905 一致,已提升至 256K 上下文。Kimi K2 高速版的输出速度达 60~100 Token/s,是普通版的 6 倍左右。

Moonshot AI
2025-09-161 分钟
Moonshot AI入门

Kimi K2 模型更新,带来更强的代码能力、更快的 API

Kimi K2 模型更新,带来更强的代码能力、更快的 API

Moonshot AI
2025-09-052 分钟
Moonshot AI进阶

Kimi K2 又又又提速了

经过工程师们的不懈努力,kimi-k2-turbo-preview 模型输出速度已经提升至每秒 60 Tokens,最高可达每秒 100 Tokens。目前仍然享受 5 折特惠价格(模型每百万 tokens 输入价格(缓存命中)¥2.00,输入价格(缓存未命中)¥8.00,输出价

Moonshot AI
2025-08-221 分钟
阿里巴巴进阶

Qwen-Image-Edit: Image Editing with Higher Quality and Efficiency

We are excited to introduce Qwen-Image-Edit, the image editing version of Qwen-Image. Built upon our 20B Qwen-Image model, Qwen-Image-Edit s

阿里巴巴
2025-08-191 分钟
阿里巴巴进阶

Qwen-Image: Crafting with Native Text Rendering

We are thrilled to release Qwen-Image, a 20B MMDiT image foundation model that achieves significant advances in complex text rendering and p

阿里巴巴
2025-08-041 分钟
阿里巴巴进阶

Qwen-MT: Where Speed Meets Smart Translation

Introduction Here we introduce the latest update of Qwen-MT (qwen-mt-turbo) via Qwen API. This update builds upon the powerful Qwen3, levera

阿里巴巴
2025-07-241 分钟
MiniMax进阶

MiniMax-01

We are delighted to introduce two remarkable models, **MiniMax-Text-01** and **MiniMax-VL-01**.

MiniMax开源技术报告
2025-07-0714 分钟
Anthropic入门

How People Use Claude for Support, Advice, and Companionship

在约 450 万段 Claude.ai 对话中仅 2.9% 为情感类,浪漫与性角色扮演合计不足 0.5%——AI 伴侣比想象中罕见,与 OpenAI/MIT 针对 ChatGPT 的独立研究结论一致。

2025-06-273 分钟
阿里巴巴进阶

Time to Speak Some Dialects, Qwen-TTS!

Introduction Here we introduce the latest update of Qwen-TTS (qwen-tts-latest or qwen-tts-2025-05-22) through Qwen API . Trained on a large-

阿里巴巴
2025-06-271 分钟
阿里巴巴进阶

Qwen VLo: From \"Understanding\" the World to \"Depicting\" It

Introduction The evolution of multimodal large models is continually pushing the boundaries of what we believe technology can achieve. From

阿里巴巴
2025-06-261 分钟
Anthropic进阶

Anthropic 教育报告:大学生如何使用 Claude

计算机科学学生占 Claude.ai 对话 38.6% 但仅占学位 5.4%,严重超量;商科、健康、人文则明显不足,既反映 Claude 在 STEM 社区的高知名度,也说明其在 STEM 类任务上适配性更强。

2025-04-083 分钟
阿里巴巴入门

Qwen2.5 Omni: See, Hear, Talk, Write, Do It All!

We release Qwen2.5-Omni, the new flagship end-to-end multimodal model in the Qwen series. Designed for comprehensive multimodal perception,

阿里巴巴
2025-03-271 分钟
阿里巴巴进阶

Qwen2.5-VL-32B: Smarter and Lighter

Introduction At the end of January this year, we launched the Qwen2.5-VL series of models, which received widespread attention and positive

阿里巴巴
2025-03-241 分钟
阿里巴巴进阶

QwQ-32B: Embracing the Power of Reinforcement Learning

QWEN CHAT Hugging Face ModelScope DEMO DISCORD

阿里巴巴
2025-03-061 分钟
Moonshot AI进阶

技术报告:Muon 优化器的首次大规模训练实践

近期,基于矩阵正交化(matrix orthogonalization)的 Muon 优化器在小规模语言模型训练中展现出了优异的性能,但其在大模型训练中的可扩展性尚未得到验证。我们发现了两个提升 Muon 可扩展性的关键技术:(1)引入权重衰减(weight decay);(2)

Moonshot AI
2025-03-035 分钟
Moonshot AI入门

介绍一下 MoBA:面向长文本大模型的混合块注意力机制

MoBA通过将专家混合系统(Mixture of Experts, MoE)的思想与稀疏注意力(sparse attention)相结合,为大语言模型中的长文本处理方式带来革命性变化。

Moonshot AI
2025-02-193 分钟
Moonshot AI入门

为什么要推出 Kimi Latest 模型?

2024 年 1 月 31 日,Kimi 开放平台开启公测,推出了最高支持 128k 上下文大小的

Moonshot AI
2025-02-173 分钟
阿里巴巴进阶

Qwen2.5-Max: Exploring the Intelligence of Large-scale MoE Model

It is widely recognized that continuously scaling both data size and model size can lead to significant improvements in model intelligence.

阿里巴巴
2025-01-281 分钟
阿里巴巴进阶

Qwen2.5 VL! Qwen2.5 VL! Qwen2.5 VL!

We release Qwen2.5-VL, the new flagship vision-language model of Qwen and also a significant leap from the previous Qwen2-VL. To try the lat

阿里巴巴
2025-01-261 分钟
阿里巴巴进阶

Global-batch load balance almost free lunch to improve your MoE LLM training

Background The Mixture-of-Experts (MoEs) architecture has become a popular model-parameter-scale-up technique. Typically, one MoE layer cons

阿里巴巴
2025-01-211 分钟
阿里巴巴进阶

QwQ: Reflect Deeply on the Boundaries of the Unknown

Note: This is the pronunciation of QwQ: /kwju:/ , similar to the word “quill”.

阿里巴巴
2024-11-281 分钟
阿里巴巴进阶

Qwen2.5-Coder Series: Powerful, Diverse, Practical.

Introduction Today, we are excited to open source the “Powerful”, “Diverse”, and “Practical” Qwen2.5-Coder series, dedicated to continuously

阿里巴巴
2024-11-121 分钟
阿里巴巴进阶

Qwen2.5: A Party of Foundation Models!

Introduction In the past three months since Qwen2’s release, numerous developers have built new models on the Qwen2 language models, providi

阿里巴巴
2024-09-191 分钟
阿里巴巴进阶

Qwen2.5-LLM: Extending the boundary of LLMs

Introduction In this blog, we delve into the details of our latest Qwen2.5 series language models. We have developed a range of decoder-only

阿里巴巴
2024-09-191 分钟
阿里巴巴进阶

Qwen2.5-Coder: Code More, Learn More!

Introduction In early April, we introduced CodeQwen1.5, which garnered significant attention from the community. Since then, we have been wo

阿里巴巴
2024-09-191 分钟
Moonshot AI进阶

使用 Unreal5 游戏引擎和 Kimi 大模型开发交互式游戏

使用 Unreal5 游戏引擎和 Kimi 大模型开发交互式游戏

Moonshot AI
2024-09-022 分钟
阿里巴巴进阶

Qwen2-VL: To See the World More Clearly

After a year’s relentless efforts, today we are thrilled to release Qwen2-VL! Qwen2-VL is the latest version of the vision language models b

阿里巴巴
2024-08-291 分钟
阿里巴巴进阶

Qwen2-Audio: Chat with Your Voice!

To achieve the objective of building an AGI system, the model should be capable of understanding information from different modalities. Than

阿里巴巴
2024-08-091 分钟
阿里巴巴入门

Introducing Qwen2-Math

🚨 This model mainly supports English. We will release bilingual (English and Chinese) math models soon. Introduction Over the past year, we

阿里巴巴
2024-08-081 分钟
Moonshot AI入门

Kimi 企业级 API 正式发布

Kimi 企业级 API 正式发布

Moonshot AI
2024-08-021 分钟
Moonshot AI入门

Kimi 开放平台 Office Hour Season 1 Recap

Kimi 开放平台 Office Hour Season 1 Recap

Moonshot AI
2024-07-231 分钟
阿里巴巴进阶

Hello Qwen2

Introduction After months of efforts, we are pleased to announce the evolution from Qwen1.5 to Qwen2. This time, we bring to you:

阿里巴巴
2024-06-071 分钟
Moonshot AI进阶

Kimi API 还没用起来?请看这篇无门槛快速入门指南

Kimi API 已经发布一个多月了,有没有用它做点有意思的事情?

Moonshot AI
2024-05-307 分钟
Moonshot AI入门

Kimi 大模型 API 更新了,也期待在「亚马逊云科技中国峰会」见到大家 | 开发者速递

Kimi 大模型 API 更新了,也期待在「亚马逊云科技中国峰会」见到大家 | 开发者速递

Moonshot AI
2024-05-292 分钟
阿里巴巴入门

Notes on Qwen-Max-0428

Previously, we opensourced a series of Qwen1.5 model ranging from 0.5 to 110 billion parameters. Now, we release a larger model, Qwen-Max-04

阿里巴巴
2024-05-111 分钟
阿里巴巴进阶

Qwen1.5-110B: The First 100B+ Model of the Qwen1.5 Series

Introduction Recently we have witnessed a burst of large-scale models with over 100 billion parameters in the opensource community. These mo

阿里巴巴
2024-04-251 分钟
阿里巴巴进阶

Code with CodeQwen1.5

Introduction The advent of advanced programming tools, which harnesses the power of large language models (LLMs), has significantly enhanced

阿里巴巴
2024-04-161 分钟
阿里巴巴进阶

Qwen1.5-32B: Fitting the Capstone of the Qwen1.5 Language Model Series

Introduction The open-source community has long sought a model that strikes an ideal balance between performance, efficiency, and memory foo

阿里巴巴
2024-04-021 分钟
阿里巴巴进阶

Qwen1.5-MoE: Matching 7B Model Performance with 1/3 Activated Parameters

Introduction Since the surge in interest sparked by Mixtral, research on mixture-of-expert (MoE) models has gained significant momentum. Bot

阿里巴巴
2024-03-281 分钟
Anthropic进阶

Research

Our research teams investigate the safety, inner workings, and societal impacts of AI models—so that artificial intelligence has a positive

Anthropic
2024-03-092 分钟
阿里巴巴入门

Introducing Qwen1.5

Introduction In recent months, our focus has been on developing a “good” model while optimizing the developer experience. As we progress tow

阿里巴巴
2024-02-041 分钟
阿里巴巴入门

Introducing Qwen

4 months after our first release of Qwen-7B, which is the starting point of our opensource journey of large language models (LLM), we now pr

阿里巴巴
2024-01-231 分钟
Anthropic进阶

projectdeal

Got some quirky supplies for your next creative project:

Anthropic
2023-11-0322 分钟
阿里巴巴进阶

OFASys: Enabling Multitask Learning with One Line of Code!

Intro Generalist Models are hot! We all see an opportunity towards a real generalist model by multimodal multitask learning. We previously r

阿里巴巴
2022-12-281 分钟
阿里巴巴进阶

Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

CLIP1 is a phenomenal playmaker in vision and multimodal representation learning. It plays not only as a foundation model but also a bridge

阿里巴巴
2022-12-241 分钟
阿里巴巴进阶

OFA: Towards Building a One-For-All Model

2022 is a year of generalist models! With the bloom of multimodal pretraining, especially the unified model, we have witnessed the opportuni

阿里巴巴
2022-11-141 分钟
Apple进阶

Let’s innovate together.

与 Apple 一起构建令人惊叹的机器学习体验。探索 Apple 为开发者和研究人员提供的机遇。

Apple
3 分钟
xAI入门

Contact Sales

能为您的组织带来哪些价值吗?请填写以下表格,我们的销售团队成员将尽快与您联系。

xAI
1 分钟