Reference

MCP server

@promptflow/mcp-server (v2.0.0) exposes the engine to MCP clients such as GitHub Copilot agent mode, Claude, and custom agents, over stdio. It is local, read-only, needs no API key, and makes no network calls.

Configure

Build it from the repository (npm publishing is pending), then point your client at dist/bin.js:

shell
pnpm --filter @promptflow/engine build && pnpm --filter @promptflow/mcp-server build
.vscode/mcp.json
{
  "servers": {
    "promptflow": {
      "type": "stdio",
      "command": "node",
      "args": ["/absolute/path/to/Prompt/packages/mcp-server/dist/bin.js"]
    }
  }
}
Claude Desktop / generic client
{
  "mcpServers": {
    "promptflow": { "command": "node", "args": ["/absolute/path/to/Prompt/packages/mcp-server/dist/bin.js"] }
  }
}

Tools (14)

ToolDescription
optimize_promptOptimize a prompt with deterministic, verified transforms (duplicates, filler, optional sections). Every candidate is checked to preserve all sentences, requirements, code, and variables; the original is kept if nothing is better. Returns the selected prompt, token savings, quality score, and what changed.
optimize_prompt_pipelineRun the optimization pipeline and return per-stage results plus token and estimated cost metrics, in the result shape used by the PromptFlow VS Code extension.
compress_promptRemove exact duplicate sentences and extra whitespace without rewriting any wording.
validate_promptCheck a prompt for missing objective or output format, ambiguity, conflicting instructions, duplicates, prompt-injection patterns, secrets, and undelimited variables. isValid is false only when a critical issue exists.
estimate_tokensCount tokens and estimate input cost for a model. "exact" is true only for OpenAI models with a loaded tokenizer; other providers are approximate. Costs are estimates from a reference price table.
analyze_promptDetect intent, task type, technical domain, requested output, requirements, and entities (languages, frameworks, file paths, identifiers).
extract_entitiesExtract languages, frameworks, libraries, cloud platforms, services, technologies, file paths, identifiers, URLs, and commands.
summarize_contextExtractive summary: the first N unique, non-empty sentences of a context block, in order. Does not rank or rewrite.
score_promptScore prompt quality 0-100: 100 minus a penalty per issue (critical 30, warning 10, info 3; capped at 50 with any critical issue), with per-dimension scores and suggestions.
evaluate_promptFull quality report: overall score, grade, eight quality dimensions with the checks that failed, and every finding with why it matters and how to fix it.
compare_promptsCompare two prompts: token and estimated cost delta, quality delta per dimension, and which findings were resolved or introduced.
scan_prompt_securityPrompt-injection risk (0-100, signature-based, with threat ids), sensitive-data counts by category, undelimited template variables, and destructive operations. Never returns detected secret values.
scan_sensitive_dataScan text for credentials, cloud secrets, and PII with local deterministic detectors. Returns offsets and per-category counts, never the sensitive values.
sanitize_promptReplace detected sensitive data with placeholders such as [REDACTED_API_KEY]. Returns the sanitized text and counts, never the original values. Use before sending a prompt to an external model.

All tools are annotated read-only and idempotent. The first eleven keep the exact result fields they had in the PromptFlow AI extension 1.5.x, so existing agent configurations keep working.

Resources and prompts

  • promptflow://rules: the rule catalogue, dimensions, and scoring penalties.
  • promptflow://models: models, token encodings, and reference prices.
  • Prompt review-prompt: asks the assistant to evaluate and optimize a prompt, with the prompt wrapped in tags as data.

Security

  • Arguments are validated with a schema. Text is capped at 100,000 characters by default (PROMPTFLOW_MCP_MAX_PROMPT_CHARS, hard ceiling 2,000,000).
  • Bad arguments return tool errors (isError) rather than protocol errors, so agents can recover.
  • No tool returns detected secret values: sensitive-data results carry categories and offsets only.
  • With PROMPTFLOW_MCP_LOG=debug, the server logs tool name, input size, outcome, and duration to stderr, never content.