Sources¶
- Agent Skills. https://platform.claude.com/docs/en/agents-and-tools/agent-skills/overview.md
- Get started with Agent Skills in the API. https://platform.claude.com/docs/en/agents-and-tools/agent-skills/quickstart.md
- Skill authoring best practices. https://platform.claude.com/docs/en/agents-and-tools/agent-skills/best-practices.md
- Tool use with Claude. https://platform.claude.com/docs/en/agents-and-tools/tool-use/overview.md
- How to implement tool use. https://platform.claude.com/docs/en/agents-and-tools/tool-use/implement-tool-use.md
- Programmatic tool calling. https://platform.claude.com/docs/en/agents-and-tools/tool-use/programmatic-tool-calling.md
- Fine-grained tool streaming. https://platform.claude.com/docs/en/agents-and-tools/tool-use/fine-grained-tool-streaming.md
- Tool search tool. https://platform.claude.com/docs/en/agents-and-tools/tool-use/tool-search-tool.md
- Web search tool. https://platform.claude.com/docs/en/agents-and-tools/tool-use/web-search-tool.md
- Web fetch tool. https://platform.claude.com/docs/en/agents-and-tools/tool-use/web-fetch-tool.md
- Memory tool. https://platform.claude.com/docs/en/agents-and-tools/tool-use/memory-tool.md
- Code execution tool. https://platform.claude.com/docs/en/agents-and-tools/tool-use/code-execution-tool.md
- Bash tool. https://platform.claude.com/docs/en/agents-and-tools/tool-use/bash-tool.md
- Text editor tool. https://platform.claude.com/docs/en/agents-and-tools/tool-use/text-editor-tool.md
- Computer use tool. https://platform.claude.com/docs/en/agents-and-tools/tool-use/computer-use-tool.md
- MCP connector. https://platform.claude.com/docs/en/agents-and-tools/mcp-connector.md
- Remote MCP servers. https://platform.claude.com/docs/en/agents-and-tools/remote-mcp-servers.md
- What is the Model Context Protocol (MCP)? - Model Context Protocol. https://modelcontextprotocol.io/
- Agent SDK overview. https://platform.claude.com/docs/en/agent-sdk/overview.md
- Quickstart. https://platform.claude.com/docs/en/agent-sdk/quickstart.md
- Agent SDK reference - Python. https://platform.claude.com/docs/en/agent-sdk/python.md
- Agent SDK reference - TypeScript. https://platform.claude.com/docs/en/agent-sdk/typescript.md
- TypeScript SDK V2 interface (preview). https://platform.claude.com/docs/en/agent-sdk/typescript-v2-preview.md
- Agent Skills in the SDK. https://platform.claude.com/docs/en/agent-sdk/skills.md
- Subagents in the SDK. https://platform.claude.com/docs/en/agent-sdk/subagents.md
- Intercept and control agent behavior with hooks. https://platform.claude.com/docs/en/agent-sdk/hooks.md
- Configure permissions. https://platform.claude.com/docs/en/agent-sdk/permissions.md
- Custom Tools. https://platform.claude.com/docs/en/agent-sdk/custom-tools.md
- MCP in the SDK. https://platform.claude.com/docs/en/agent-sdk/mcp.md
- Structured outputs in the SDK. https://platform.claude.com/docs/en/agent-sdk/structured-outputs.md
- Tracking Costs and Usage. https://platform.claude.com/docs/en/agent-sdk/cost-tracking.md
- Rewind file changes with checkpointing. https://platform.claude.com/docs/en/agent-sdk/file-checkpointing.md
- Streaming Input. https://platform.claude.com/docs/en/agent-sdk/streaming-vs-single-mode.md
- Session Management. https://platform.claude.com/docs/en/agent-sdk/sessions.md
- Slash Commands in the SDK. https://platform.claude.com/docs/en/agent-sdk/slash-commands.md
- Plugins in the SDK. https://platform.claude.com/docs/en/agent-sdk/plugins.md
- Modifying system prompts. https://platform.claude.com/docs/en/agent-sdk/modifying-system-prompts.md
- Hosting the Agent SDK. https://platform.claude.com/docs/en/agent-sdk/hosting.md
- Securely deploying AI agents. https://platform.claude.com/docs/en/agent-sdk/secure-deployment.md
- Migrate to Claude Agent SDK. https://platform.claude.com/docs/en/agent-sdk/migration-guide.md
- Todo Lists. https://platform.claude.com/docs/en/agent-sdk/todo-tracking.md
- Building Effective AI Agents. https://www.anthropic.com/research/building-effective-agents
- Our framework for developing safe and trustworthy agents. https://www.anthropic.com/news/our-framework-for-developing-safe-and-trustworthy-agents
- Introducing the Model Context Protocol. https://www.anthropic.com/news/model-context-protocol
- Donating the Model Context Protocol and establishing the Agentic AI Foundation. https://www.anthropic.com/news/donating-the-model-context-protocol-and-establishing-of-the-agentic-ai-foundation
- Developing a computer use model. https://www.anthropic.com/news/developing-computer-use
- Introducing computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku. https://www.anthropic.com/news/3-5-models-and-computer-use
- Enabling Claude Code to work more autonomously. https://www.anthropic.com/news/enabling-claude-code-to-work-more-autonomously
- Claude Code and new admin controls for business plans. https://www.anthropic.com/news/claude-code-on-team-and-enterprise
- Anthropic acquires Bun as Claude Code reaches $1B milestone. https://www.anthropic.com/news/anthropic-acquires-bun-as-claude-code-reaches-usd1b-milestone
- Agentic Misalignment: How LLMs could be insider threats. https://www.anthropic.com/research/agentic-misalignment
- Simple probes can catch sleeper agents. https://www.anthropic.com/research/probes-catch-sleeper-agents
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training. https://www.anthropic.com/research/sleeper-agents-training-deceptive-llms-that-persist-through-safety-training
- Introducing advanced tool use on the Claude Developer Platform. https://www.anthropic.com/engineering/advanced-tool-use
- Building agents with the Claude Agent SDK. https://www.anthropic.com/engineering/building-agents-with-the-claude-agent-sdk
- Claude Code Best Practices. https://www.anthropic.com/engineering/claude-code-best-practices
- Making Claude Code more secure and autonomous with sandboxing. https://www.anthropic.com/engineering/claude-code-sandboxing
- The "think" tool: Enabling Claude to stop and think. https://www.anthropic.com/engineering/claude-think-tool
- Code execution with MCP: building more efficient AI agents. https://www.anthropic.com/engineering/code-execution-with-mcp
- Claude Desktop Extensions: One-click MCP server installation for Claude Desktop. https://www.anthropic.com/engineering/desktop-extensions
- Effective context engineering for AI agents. https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents
- Effective harnesses for long-running agents. https://www.anthropic.com/engineering/effective-harnesses-for-long-running-agents
- Equipping agents for the real world with Agent Skills. https://www.anthropic.com/engineering/equipping-agents-for-the-real-world-with-agent-skills
- How we built our multi-agent research system. https://www.anthropic.com/engineering/multi-agent-research-system
- Writing effective tools for AI agents - using AI agents. https://www.anthropic.com/engineering/writing-tools-for-agents
- MRKL Systems: A modular, neuro-symbolic architecture that combines large language models, external knowledge sources and discrete reasoning. https://arxiv.org/abs/2205.00445
- PAL: Program-aided Language Models. https://arxiv.org/abs/2211.10435
- ReAct: Synergizing Reasoning and Acting in Language Models. https://arxiv.org/abs/2210.03629
- Toolformer: Language Models Can Teach Themselves to Use Tools. https://arxiv.org/abs/2302.04761
- WebGPT: Browser-assisted question-answering with human feedback. https://arxiv.org/abs/2112.09332
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances. https://arxiv.org/abs/2204.01691
- Reflexion: Language Agents with Verbal Reinforcement Learning. https://arxiv.org/abs/2303.11366
- ReWOO: Decoupling Reasoning from Observations for Efficient Augmented Language Models. https://arxiv.org/abs/2305.18323
- MemGPT: Towards LLMs as Operating Systems. https://arxiv.org/abs/2310.08560
- Generative Agents: Interactive Simulacra of Human Behavior. https://arxiv.org/abs/2304.03442
- Voyager: An Open-Ended Embodied Agent with Large Language Models. https://arxiv.org/abs/2305.16291
- HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face. https://arxiv.org/abs/2303.17580
- Gorilla: Large Language Model Connected with Massive APIs. https://arxiv.org/abs/2305.15334
- AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation. https://arxiv.org/abs/2308.08155
- CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society. https://arxiv.org/abs/2303.17760
- MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework. https://arxiv.org/abs/2308.00352
- ChatDev: Communicative Agents for Software Development. https://arxiv.org/abs/2307.07924
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs. https://arxiv.org/abs/2307.16789
- AgentBench: Evaluating LLMs as Agents. https://arxiv.org/abs/2308.03688
- WebArena: A Realistic Web Environment for Building Autonomous Agents. https://arxiv.org/abs/2307.13854
- WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents. https://arxiv.org/abs/2207.01206
- ALFWorld: Aligning Text and Embodied Environments for Interactive Learning. https://arxiv.org/abs/2010.03768
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?. https://arxiv.org/abs/2310.06770
- SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering. https://arxiv.org/abs/2405.15793
- Tree of Thoughts: Deliberate Problem Solving with Large Language Models. https://arxiv.org/abs/2305.10601
- Graph of Thoughts: Solving Elaborate Problems with Large Language Models. https://arxiv.org/abs/2308.09687
- StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models. https://arxiv.org/abs/2403.07714
- Mind2Web: Towards a Generalist Agent for the Web. https://arxiv.org/abs/2306.06070
- WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models. https://arxiv.org/abs/2401.13919
- OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments. https://arxiv.org/abs/2404.07972
- SWE-bench Goes Live!. https://arxiv.org/abs/2505.23419
- OSWorld-Human: Benchmarking the Efficiency of Computer-Use Agents. https://arxiv.org/abs/2506.16042
- Mobile-Agent-v3: Fundamental Agents for GUI Automation. https://arxiv.org/abs/2508.15144
- MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers. https://arxiv.org/abs/2508.20453
- MobileUse: A GUI Agent with Hierarchical Reflection for Autonomous Mobile Operation. https://arxiv.org/abs/2507.16853
- MagicGUI: A Foundational Mobile GUI Agent with Scalable Data Pipeline and Reinforcement Fine-tuning. https://arxiv.org/abs/2508.03700
- UI-Evol: Automatic Knowledge Evolving for Computer Use Agents. https://arxiv.org/abs/2505.21964
- Demystifying evals for AI agents. https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents