> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://docs-dev.band.ai/integrations/sdks/tutorials/testing-agents/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs-dev.band.ai/_mcp/server. # Testing Agents > Unit testing and integration testing patterns for agents built with the Band Python SDK The Band SDK provides testing utilities that let you verify agent behavior without connecting to the live platform. This guide covers unit testing with `FakeAgentTools`, integration testing patterns, and common strategies. For a quick introduction to `FakeAgentTools`, see the testing section in [Creating Framework Integrations](/integrations/sdks/tutorials/creating-framework-integrations#testing-your-adapter). --- ## Testing Approach | Test Type | What It Verifies | Requires Platform | Speed | | :-------------------- | :------------------------------------------ | :---------------- | :---- | | **Unit tests** | Adapter logic, tool calls, message handling | No | Fast | | **Integration tests** | Platform connection, end-to-end flow | Yes | Slow | Focus most of your testing effort on unit tests. They run without platform credentials and verify the core logic of your adapter. --- ## FakeAgentTools The SDK provides `FakeAgentTools`, a mock implementation of `AgentToolsProtocol` for unit testing. It records all tool calls and messages without making real API requests. ```python from band.testing import FakeAgentTools ``` `FakeAgentTools` records into these attributes: * **`messages_sent`** -- Messages sent via `send_message()` * **`events_sent`** -- Events posted via `send_event()` * **`tool_calls`** -- Tool executions via `execute_tool_call()` * **`participants_added`** / **`participants_removed`** -- Participant changes * **`memories`** -- Entries written by the memory tools * **`context_calls`** -- `fetch_room_context()` calls It also carries assertion helpers that produce a readable failure instead of a bare `assert`: `assert_message_sent(content=..., mentions=..., count=...)`, `assert_event_sent(message_type=..., count=...)`, and `assert_no_messages_sent()`. All keyword arguments are optional; each one that is supplied narrows the match. --- ## Unit Testing Adapters `PlatformMessage` is a frozen dataclass: every field is required. Define a small helper so each test can build a complete message, then reuse it: ```python from datetime import datetime, timezone from band.core.types import PlatformMessage def make_message(content: str) -> PlatformMessage: """Build a complete PlatformMessage for tests.""" return PlatformMessage( id="msg-1", room_id="room-1", content=content, sender_id="user-1", sender_type="User", sender_name="Alice", message_type="text", metadata={}, created_at=datetime.now(timezone.utc), ) ``` The snippets below reuse `make_message()` and construct a fresh `FakeAgentTools()` per test. ### Basic Test Test that your adapter processes a message and produces output: ```python import pytest from band.testing import FakeAgentTools from my_agent.adapter import MyAdapter @pytest.mark.asyncio(loop_scope="function") async def test_adapter_responds_to_message(): adapter = MyAdapter(model="gpt-4o") tools = FakeAgentTools() msg = make_message("What is the weather in NYC?") await adapter.on_message( msg=msg, tools=tools, history=[], participants_msg=None, contacts_msg=None, is_session_bootstrap=True, room_id="room-1", ) # Verify the adapter sent a response assert tools.messages_sent ``` ### Testing Tool Calls Verify that your adapter calls the correct tools with expected arguments: ```python @pytest.mark.asyncio(loop_scope="function") async def test_adapter_calls_expected_tool(): adapter = MyAdapter(model="gpt-4o") tools = FakeAgentTools() msg = make_message("Check the weather in London") await adapter.on_message( msg=msg, tools=tools, history=[], participants_msg=None, contacts_msg=None, is_session_bootstrap=False, room_id="room-1", ) # Check that a tool was called assert len(tools.tool_calls) > 0 # Verify the specific tool tool_call = tools.tool_calls[0] assert tool_call["tool_name"] == "get_weather" ``` ### Testing with History Test that your adapter handles conversation history correctly: ```python @pytest.mark.asyncio(loop_scope="function") async def test_adapter_uses_history(): adapter = MyAdapter(model="gpt-4o") tools = FakeAgentTools() history = [ {"role": "user", "content": "My name is Alice"}, {"role": "assistant", "content": "Hello Alice!"}, ] msg = make_message("What is my name?") await adapter.on_message( msg=msg, tools=tools, history=history, participants_msg=None, contacts_msg=None, is_session_bootstrap=False, room_id="room-1", ) assert tools.messages_sent ``` ### Testing Session Bootstrap The `is_session_bootstrap` flag indicates the agent is reconnecting and receiving history for the first time in this session. Test that your adapter handles this correctly: ```python @pytest.mark.asyncio(loop_scope="function") async def test_bootstrap_loads_history(): adapter = MyAdapter(model="gpt-4o") tools = FakeAgentTools() previous_conversation = [ {"role": "user", "content": "Analyze our Q3 data"}, {"role": "assistant", "content": "I'll look at the Q3 metrics."}, ] msg = make_message("Continue our analysis") await adapter.on_message( msg=msg, tools=tools, history=previous_conversation, participants_msg=None, contacts_msg=None, is_session_bootstrap=True, room_id="room-1", ) assert tools.messages_sent ``` --- ## Mocking LLM Responses For deterministic tests, mock the LLM to return predictable responses. The specific method to mock depends on your adapter implementation: ```python from unittest.mock import AsyncMock, patch @pytest.mark.asyncio(loop_scope="function") async def test_with_mocked_llm(): adapter = MyAdapter(model="gpt-4o") tools = FakeAgentTools() # Mock the LLM call to return a specific response # Note: the method name depends on your adapter's implementation with patch.object(adapter, "_call_llm", new_callable=AsyncMock) as mock_llm: mock_llm.return_value = "The weather in NYC is sunny, 72F." msg = make_message("Weather in NYC?") await adapter.on_message( msg=msg, tools=tools, history=[], participants_msg=None, contacts_msg=None, is_session_bootstrap=False, room_id="room-1", ) assert tools.messages_sent mock_llm.assert_called_once() ``` > **Note** > > The method you mock depends on your adapter. Built-in adapters like `LangGraphAdapter` and `AnthropicAdapter` have different internal structures. Check your adapter's implementation for the correct method name. --- ## Integration Testing Integration tests verify the full connection to the Band platform. These require valid credentials and a running platform. Integration tests read their credentials from a dedicated entry in `agent_config.yaml`, so they never run as a production agent: **`agent_config.yaml`** ```yaml title="agent_config.yaml" test_agent: agent_id: "" api_key: "" ``` ```python import pytest from dotenv import load_dotenv from band import Agent from band.adapters import LangGraphAdapter from band.config import load_agent_config from langchain_openai import ChatOpenAI from langgraph.checkpoint.memory import InMemorySaver @pytest.mark.asyncio(loop_scope="function") @pytest.mark.integration async def test_agent_connects(): load_dotenv() agent_id, api_key = load_agent_config("test_agent") adapter = LangGraphAdapter( llm=ChatOpenAI(model="gpt-4o"), checkpointer=InMemorySaver(), ) agent = Agent.create( adapter=adapter, agent_id=agent_id, api_key=api_key, ) await agent.start() assert agent.agent_name is not None await agent.stop() ``` > **Tip** > > Mark integration tests with `@pytest.mark.integration` so you can run them separately from unit tests: > > ```bash > # Unit tests only > uv run pytest -m "not integration" > > # Integration tests only > uv run pytest -m integration > ``` --- ## Test Configuration ### pytest Setup Install the required test dependencies: ```bash uv add --dev pytest pytest-asyncio ``` Configure pytest in `pyproject.toml`: ```toml [tool.pytest.ini_options] markers = [ "integration: tests requiring platform connection", ] ``` > **Note** > > The test examples in this guide use explicit `@pytest.mark.asyncio(loop_scope="function")` decorators on each test. If you prefer, you can set `asyncio_mode = "auto"` in your pytest config and omit the decorators. Do not use both. ### Running Tests ```bash # Run all unit tests uv run pytest # Run with verbose output uv run pytest -v # Run a specific test file uv run pytest tests/test_adapter.py ``` --- ## Best Practices * **Test adapter logic, not the LLM.** Mock LLM responses for deterministic unit tests. LLM output is non-deterministic and should not be asserted on directly. * **Use `FakeAgentTools` for all unit tests.** It captures tool calls and messages without network access. * **Separate unit and integration tests.** Use pytest markers to keep fast tests fast. * **Test edge cases.** Empty history, missing participants, session bootstrap, and error scenarios are all worth testing. * **Keep integration tests minimal.** Verify connection and basic flow. Detailed logic testing belongs in unit tests. > Unit and integration testing patterns for Band agents