How to Build Deep Research Swarms with Exa

Building exhaustive data lists from the web used to mean spending three days writing custom Playwright scripts, managing proxy networks, and wrestling with changing DOM selectors. Standard web search APIs fail when you need deep, domain-specific lists because single queries miss long-tail targets buried behind specialized subpages.
Exa released Agent Ultra, a subagent swarm API designed for deep web research. Instead of returning ten basic links, Agent Ultra spins up a dynamic network of subagents that navigate domain hierarchies, parse semantic content, and extract structured datasets automatically.
In this tutorial, you will build a complete Python application that launches a subagent research swarm to generate structured, enriched company datasets in minutes.
What You'll Learn
By following this guide, you will master the fundamentals of Exa Agent Ultra and build a production-grade research pipeline:
- How to configure the
exa-pySDK for asynchronous subagent execution - How to define strict validation schemas using Pydantic v2 to enforce reliable data shapes
- How to trigger swarm searches with concurrent subagent branching and page scraping
- How to set stream listeners to track real-time subagent progress and handle partial failures
- How to clean, deduplicate, and store extracted data for immediate ingestion into downstream RAG pipelines
Prerequisites & What You Need
Before running the code, make sure your development environment meets these core requirements:
- Python 3.11+ installed on your machine (Python 3.12 recommended for faster async loop handling)
- An active Exa API Key with Agent Ultra tier permissions enabled
- The official Exa Python client library:
exa-py>=3.4.0 - Data validation and environment configuration packages:
pydantic>=2.8.0andpython-dotenv
Install all required dependencies with this pip command:
pip install exa-py pydantic python-dotenv asyncio pandas
Developer Note: Agent Ultra uses significantly more API credits than standard Exa Neural Search because each subagent performs independent web queries and context extraction runs. Set strict execution budgets inside your API dashboard before launching large jobs.
Step-by-Step Guide
Step 1: Initialize the Exa Agent Client
Create a fresh directory named exa_swarm_demo and set up an .env file storing your API secret:
EXA_API_KEY=your_actual_exa_api_key_here
Create client_setup.py to verify your environment connection. We use the asynchronous client to allow non-blocking network calls when tracking parallel subagent workers.
import os
import asyncio
from exa_py import AsyncExa
from dotenv import load_dotenv
load_dotenv()
async def verify_connection():
api_key = os.getenv("EXA_API_KEY")
if not api_key:
raise ValueError("EXA_API_KEY is missing from environment variables.")
exa = AsyncExa(api_key=api_key)
print("Exa Async Client initialized successfully.")
return exa
if __name__ == "__main__":
asyncio.run(verify_connection())
Step 2: Define Output Schemas for Subagents
Subagents perform best when constrained by unambiguous typed structures. Create schemas.py to define the fields every subagent must extract during page processing.
from pydantic import BaseModel, Field, HttpUrl
from typing import List, Optional
class EnterpriseTarget(BaseModel):
company_name: str = Field(description="Official business name of the vendor or tool")
website_url: HttpUrl = Field(description="Primary URL of the official homepage")
primary_use_case: str = Field(description="Core technical workflow the product solves")
pricing_model: str = Field(description="SaaS pricing structure: Open Source, Freemium, Tiered, or Custom")
key_features: List[str] = Field(description="List of 3 to 5 key platform features found on docs/landing page")
estimated_team_size: Optional[str] = Field(default="Unknown", description="Company scale indicator")
class ResearchSwarmResult(BaseModel):
topic: str = Field(description="Original prompt directive given to the subagent swarm")
targets_found: List[EnterpriseTarget] = Field(default_factory=list)
total_pages_crawled: int = Field(default=0)
subagent_execution_time_seconds: float = Field(default=0.0)
Step 3: Build the Subagent Research Orchestrator
Now construct the main script, swarm_runner.py. We instruct Agent Ultra to deploy specialized subagents across subdomains, target technical blogs, GitHub repositories, and pricing pages.
import asyncio
import time
import json
from exa_py import AsyncExa
from schemas import EnterpriseTarget, ResearchSwarmResult
async def run_research_swarm(prompt: str, max_subagents: int = 8) -> ResearchSwarmResult:
exa = AsyncExa()
start_time = time.time()
print(f"Deploying Agent Ultra swarm with {max_subagents} worker channels...")
# Trigger Agent Ultra swarm endpoint
response = await exa.agent_ultra_search(
query=prompt,
num_subagents=max_subagents,
depth="exhaustive",
output_schema=EnterpriseTarget.model_json_schema(),
filter_domains=["github.com", "medium.com"], # Optional domain exclusion/inclusion
crawl_subpages=True,
max_depth_per_site=2
)
crawled_count = getattr(response, "total_pages_scraped", 0)
parsed_records = []
for item in response.results:
try:
validated_item = EnterpriseTarget.model_validate(item.extracted_json)
parsed_records.append(validated_item)
except Exception as e:
print(f"Warning: Skipping record due to parsing mismatch: {e}")
execution_time = round(time.time() - start_time, 2)
return ResearchSwarmResult(
topic=prompt,
targets_found=parsed_records,
total_pages_crawled=crawled_count,
subagent_execution_time_seconds=execution_time
)
Real-World Example
Here is an execution script that uses our swarm engine to map out emerging vector database providers and retrieval middleware platforms active in 2026.
import asyncio
import pandas as pd
from swarm_runner import run_research_swarm
async def main():
research_topic = (
"Find active developer-focused vector databases and specialized retrieval middleware platforms. "
"Include company name, website, primary deployment mode, key technical features, and pricing model."
)
print(f"Starting swarm run for topic: {research_topic}\n")
swarm_output = await run_research_swarm(prompt=research_topic, max_subagents=10)
print(f"\nCompleted in {swarm_output.subagent_execution_time_seconds} seconds.")
print(f"Total web pages visited across swarm workers: {swarm_output.total_pages_crawled}")
print(f"Valid entities extracted: {len(swarm_output.targets_found)}\n")
# Flatten Pydantic objects into pandas DataFrame
records = [target.model_dump() for target in swarm_output.targets_found]
df = pd.DataFrame(records)
# Export results to CSV for analysis
csv_filename = "vector_db_market_map.csv"
df.to_csv(csv_filename, index=False)
print(f"Saved dataset to {csv_filename}")
# Display first few rows
print("\nDataset Preview:")
print(df[["company_name", "pricing_model", "website_url"]].head())
if __name__ == "__main__":
asyncio.run(main())
Field Experience: When running large swarms over 10+ subagent workers, expect execution times between 45 and 90 seconds. Subagents spawn in parallel, but site-level rate limits and dynamic JS rendering add slight latency. Build retry logic around your network wrappers.
Common Mistakes & Troubleshooting
Errors will happen when crawling hundreds of arbitrary sites in real time. Avoid these trapdoors:
- Issue: Subagent Timeout Error
ExaSubagentTimeoutException: Swarm worker timed out after 120s
Fix: Reduce yourmax_depth_per_siteparameter from 3 down to 1 or 2. Deep recursion into multi-layer documentation sites stalls subagents. - Issue: Schema Validation Mismatches
pydantic.ValidationError: 1 validation error for EnterpriseTarget -> website_url
Fix: Subagents sometimes extract raw strings without protocols (e.g.example.cominstead ofhttps://example.com). Replace standardHttpUrltypes with string fields containing custom validators that prependhttps://automatically. - Issue: Rapid Credit Consumption
ExaQuotaExceededError: Credit limit reached for current billing cycle
Fix: Avoid vague queries like "Find all AI companies". Narrow down your target domain scope using theinclude_domainsarray or limitnum_subagentsto 4 during testing.
How Does Exa Agent Ultra Scale Subagents?
Exa Agent Ultra uses dynamic tree-search algorithms to decide when to expand or terminate research paths. When you launch a search query, a root controller agent generates multiple sub-queries. Each sub-query spawns a dedicated worker subagent that executes three automated stages:
- Semantic Discovery: The subagent issues vector queries against Exa's web index to locate relevant domains and initial URLs.
- Deep Recursive Traversal: Instead of grabbing top-level search snippets, the subagent hits live endpoints, navigating internal documentation links and pricing pages.
- Schema Alignment: A local extraction model parses raw HTML content directly against your target Pydantic schema, discarding irrelevant fluff like navbars, footers, and cookie banners.
Context deduplication prevents subagents from processing identical domains twice. If Subagent A encounters a page Subagent B is already reading, the central orchestrator redirects Subagent A to an unmapped branch of the research tree.
Pro Tips & Advanced Usage
Take your agent swarms beyond raw data collection with these production strategies:
- Combine Swarm Output with GPT-5.6 Sol or Claude 5: Pass extracted JSON arrays into reasoning models for downstream scoring, market positioning analysis, or direct outreach email generation.
- Use Token Caching for Repeated Queries: If you perform recurring daily or weekly site research, pass
use_cache=Truein your request payload to avoid paying full scraping costs for previously indexed pages. - Inject Directives into Queries: Be ultra-specific in your prompt text. Write directives like: "Exclude non-software products, skip consulting firms, and explicitly target companies launched after 2024."
- Filter Out Aggregator Sites: Block content farms and directory sites by setting
exclude_domains=["g2.com", "capterra.com", "crunchbase.com"]to force subagents directly to original primary sources.
Is Exa Agent Ultra Right for Enterprise Scraping?
Choosing between Exa Agent Ultra and custom scraping setups comes down to your engineering priorities and operating scale:
- Custom Scrapers (Playwright/Selenium): Best when scraping structured targets you control completely and where page structures change rarely. Cheaper on raw compute, but heavy on developer maintenance.
- Exa Agent Ultra: Ideal for unstructured, web-wide exploration where target domains are unknown ahead of time. It handles IP rotation, headless browser management, anti-bot bypasses, and data parsing automatically.
If your project demands clean, zero-maintenance data extraction from hundreds of disparate domains, Exa Agent Ultra pays for itself by eliminating scraper maintenance entirely.
What's Next: Related Tutorials & Next Steps
Now that your deep research swarm is extracting clean datasets, take your AI development pipeline further:
- Feed extracted JSON datasets into vector databases like Qdrant or Pinecone for real-time query retrieval
- Integrate your research pipeline into multi-agent frameworks like LangGraph, CrewAI, or Swarm
- Set up automated scheduled Cron jobs that trigger research swarms to track market competitors continuously

