MCP Server Dataset Builder
@wanghaisheng
About MCP Server Dataset Builder
No overview available yet
Config
Add this server to your MCP-compatible client using the configuration below.
{
"mcpServers": {
"mcp-server-dataset": {
"command": "python",
"args": [
"extract_mcp_servers.py"
]
}
}
}Tools
No tools detected
We auto-extract tools from the README. The maintainer can list them under a ## Tools heading to populate this section.
Overview
What is MCP Server Dataset Builder?
MCP Server Dataset Builder is a tool that automatically collects, categorizes, and updates information about MCP (Model Context Protocol) servers from multiple sources. It merges data from the curated awesome‑mcp‑servers repository and GitHub search, then outputs a daily CSV file. It is designed for developers and researchers who need an up‑to‑date, structured catalog of MCP servers.
How to use MCP Server Dataset Builder?
The dataset is automatically updated daily via a GitHub Actions workflow; no manual steps are required. To run locally, install the dependencies (pip install -r requirements.txt) then execute python extract_mcp_servers.py or python daily.py. A manual workflow run can be triggered from the Actions tab, with optional customization of search keywords, minimum stars, and minimum forks.
Key features of MCP Server Dataset Builder
- Combines data from curated lists and GitHub search
- Automatically assigns categories based on repository content
- Detects programming languages and frameworks used
- Adds emoji indicators for quick visual identification
- Updates daily to keep the dataset current
- Preserves historical data while adding new entries
Use cases of MCP Server Dataset Builder
- Building and maintaining a comprehensive registry of MCP servers
- Discovering new MCP repositories through automated GitHub search
- Analyzing the MCP ecosystem by category, tech stack, or popularity
- Keeping a personal or team dataset synchronized with the latest servers
- Integrating MCP server metadata into dashboards or monitoring tools
FAQ from MCP Server Dataset Builder
What data sources does MCP Server Dataset Builder use?
It extracts data from the awesome‑mcp‑servers GitHub repository and performs its own GitHub search using MCP‑related keywords. Both sources are merged and deduplicated.
How often is the dataset updated?
The dataset is automatically updated daily via a GitHub Actions workflow. No manual intervention is required, but you can trigger a manual run at any time.
Can I customize the search criteria?
Yes. When running the workflow manually, you can set keywords, minimum stars, and minimum forks. Locally, these are controlled via environment variables (KEYWORDS_ENV, MIN_STARS, MIN_FORKS).
What fields are included in the CSV output?
The CSV contains: name, description, html_url, stars, forks, keywords, category, techstack, and emojis. These fields provide a structured overview of each repository.
Does the tool require a GitHub API token?
Yes, for the GitHub search to work, a GITHUB_TOKEN environment variable must be set. The token is used for authentication when querying the GitHub API.
Frequently asked questions
What data sources does MCP Server Dataset Builder use?
It extracts data from the awesome‑mcp‑servers GitHub repository and performs its own GitHub search using MCP‑related keywords. Both sources are merged and deduplicated.
How often is the dataset updated?
The dataset is automatically updated daily via a GitHub Actions workflow. No manual intervention is required, but you can trigger a manual run at any time.
Can I customize the search criteria?
Yes. When running the workflow manually, you can set keywords, minimum stars, and minimum forks. Locally, these are controlled via environment variables (KEYWORDS_ENV, MIN_STARS, MIN_FORKS).
What fields are included in the CSV output?
The CSV contains: name, description, html_url, stars, forks, keywords, category, techstack, and emojis. These fields provide a structured overview of each repository.
Does the tool require a GitHub API token?
Yes, for the GitHub search to work, a GITHUB_TOKEN environment variable must be set. The token is used for authentication when querying the GitHub API.
Basic information
More Data & Analytics MCP servers

Kinetic Pricing
Kinetic PricingKinetic Pricing lets you run pricing research with your own customers and manage the work around each decision. Create Van Westendorp, Gabor-Granger, MaxDiff, and choice-based conjoint studies; preview and launch surveys
CoinLobster
CoinLobsterThe only MCP server with live whale trades across 15 exchanges plus on-chain DEX flow. Smart Money Radar, real liquidations and outcome-scored signals. No API key required.
xverum mcp
Xverum-LLCFind and enrich the right people from 750M professional profiles. Search by role, seniority, skills, industry, and location in plain English. Pull profiles with work history, education, and seniority, then see who's like

Subtext
Subtext by FullstorySession replay, built for agents. Subtext is agentic session review: it captures production sessions of your app and connects them to your coding agent — Claude Code, Cursor, Codex, Devin, your own harness — so it can

Waqi - AI Privacy Layer
ajprolificWaqi — hosted MCP server that redacts PII from your business tools before the AI sees them. Audit log included.
Comments