io.github.hidai25/evalview-mcp
Regression testing for AI agents. Golden baselines, CI/CD, LangGraph, CrewAI, OpenAI, Claude.
io.github.hidai25/evalview-mcp · Repository · version 0.6.0 · 128 stars · listed from registry
Install
The sweep installs into a clean prefix, but this is the command a user would run:
uvx evalview
Results by platform
Linux
MCP handshake failed2026-08-16 · probed over pypi2 recorded runs, oldest first.
handshake failed: MCP error -32000: Connection closed
--- stderr (tail) ---
…s.
New here? Start with:
demo See it work in 30 seconds
init Detect your agent and create first tests
Regression Gating — did my agent change?
snapshot Capture current behavior as baseline
snapshot list List saved baselines
snapshot show <name> Inspect a baseline
snapshot delete <name> Remove a baseline
check Compare against baseline — catch regressions
Evaluation — how good is my agent?
generate Auto-generate tests from live probing
run Execute tests, score with LLM judge
Build Your Test Suite:
capture --agent <url> Record real traffic as tests
capture --multi-turn Record a conversation as a multi-turn test
import <log_file> Convert production logs into EvalView tests
expand Generate test variations with LLM
compare Compare two agent endpoints on the same suite
Explore & Learn:
chat Interactive AI assistant for eval guidance
gym Practice agent eval patterns
Reports:
replay Open a trajectory diff for one test
visualize Generate a visual HTML report from results
trends Performance trends over time
Production:
watch Re-run checks on file change (local dev)
monitor Continuous regression detection (+ Slack alerts)
CI/CD:
ci comment Post results to a GitHub PR
badge Generate shields.io status badge
init --ci Generate GitHub Actions workflow
Advanced:
login Connect to EvalView Cloud
logout Disconnect from EvalView Cloud
whoami Show current cloud login status
feedback Open a pre-filled GitHub issue
skill Test Claude Code skills
trace Trace LLM calls in scripts
traces Query stored trace data
expand Generate test variations with LLM
Options:
--version Show the version and exit.
--help Show this message and exit.
Commands:
badge Generate a shields.io-compatible badge for your README.
capture 🎯 Capture real traffic as tests — tests from real usage, not...
chat Interactive chat interface for EvalView.
check Check current behavior against snapshot baseline.
ci CI/CD integration commands.
compare Run the same tests against two agent endpoints and compare...
demo Live regression demo — see EvalView catch and auto-heal...
expand Expand test cases into variations using LLM.
feedback Send feedback, report a bug, or request a feature.
generate Generate a draft regression suite from live agent probing.
import Convert production logs into EvalView test cases.
init Initialize EvalView in the current directory.
login Connect to EvalView Cloud.
logout Disconnect from EvalView Cloud.
mcp Manage MCP contracts (detect external server interface drift).
monitor Continuously check for regressions with optional Slack alerts.
openclaw OpenClaw integration — install skills and manage the...
replay Replay a test and show full trajectory diff vs baseline.
run Run test cases against the agent.
skill Commands for testing Claude Code skills.
snapshot Run tests and snapshot passing results as baseline.
trace Trace LLM calls in any Python script.
traces Query and manage local trace storage.
trends Show performance trends over time.
visualize Generate a visual HTML report, optionally comparing multiple...
watch Watch for file changes and re-run regression checks.
whoami Show current cloud login status.
macOS
MCP handshake failed2026-08-16 · probed over pypi2 recorded runs, oldest first.
handshake failed: MCP error -32000: Connection closed
--- stderr (tail) ---
…s.
New here? Start with:
demo See it work in 30 seconds
init Detect your agent and create first tests
Regression Gating — did my agent change?
snapshot Capture current behavior as baseline
snapshot list List saved baselines
snapshot show <name> Inspect a baseline
snapshot delete <name> Remove a baseline
check Compare against baseline — catch regressions
Evaluation — how good is my agent?
generate Auto-generate tests from live probing
run Execute tests, score with LLM judge
Build Your Test Suite:
capture --agent <url> Record real traffic as tests
capture --multi-turn Record a conversation as a multi-turn test
import <log_file> Convert production logs into EvalView tests
expand Generate test variations with LLM
compare Compare two agent endpoints on the same suite
Explore & Learn:
chat Interactive AI assistant for eval guidance
gym Practice agent eval patterns
Reports:
replay Open a trajectory diff for one test
visualize Generate a visual HTML report from results
trends Performance trends over time
Production:
watch Re-run checks on file change (local dev)
monitor Continuous regression detection (+ Slack alerts)
CI/CD:
ci comment Post results to a GitHub PR
badge Generate shields.io status badge
init --ci Generate GitHub Actions workflow
Advanced:
login Connect to EvalView Cloud
logout Disconnect from EvalView Cloud
whoami Show current cloud login status
feedback Open a pre-filled GitHub issue
skill Test Claude Code skills
trace Trace LLM calls in scripts
traces Query stored trace data
expand Generate test variations with LLM
Options:
--version Show the version and exit.
--help Show this message and exit.
Commands:
badge Generate a shields.io-compatible badge for your README.
capture 🎯 Capture real traffic as tests — tests from real usage, not...
chat Interactive chat interface for EvalView.
check Check current behavior against snapshot baseline.
ci CI/CD integration commands.
compare Run the same tests against two agent endpoints and compare...
demo Live regression demo — see EvalView catch and auto-heal...
expand Expand test cases into variations using LLM.
feedback Send feedback, report a bug, or request a feature.
generate Generate a draft regression suite from live agent probing.
import Convert production logs into EvalView test cases.
init Initialize EvalView in the current directory.
login Connect to EvalView Cloud.
logout Disconnect from EvalView Cloud.
mcp Manage MCP contracts (detect external server interface drift).
monitor Continuously check for regressions with optional Slack alerts.
openclaw OpenClaw integration — install skills and manage the...
replay Replay a test and show full trajectory diff vs baseline.
run Run test cases against the agent.
skill Commands for testing Claude Code skills.
snapshot Run tests and snapshot passing results as baseline.
trace Trace LLM calls in any Python script.
traces Query and manage local trace storage.
trends Show performance trends over time.
visualize Generate a visual HTML report, optionally comparing multiple...
watch Watch for file changes and re-run regression checks.
whoami Show current cloud login status.
Windows
MCP handshake failed2026-08-16 · probed over pypi2 recorded runs, oldest first.
handshake failed: MCP error -32000: Connection closed
--- stderr (tail) ---
…nit Detect your agent and create first tests
Regression Gating � did my agent change?
snapshot Capture current behavior as baseline
snapshot list List saved baselines
snapshot show <name> Inspect a baseline
snapshot delete <name> Remove a baseline
check Compare against baseline � catch regressions
Evaluation � how good is my agent?
generate Auto-generate tests from live probing
run Execute tests, score with LLM judge
Build Your Test Suite:
capture --agent <url> Record real traffic as tests
capture --multi-turn Record a conversation as a multi-turn test
import <log_file> Convert production logs into EvalView tests
expand Generate test variations with LLM
compare Compare two agent endpoints on the same suite
Explore & Learn:
chat Interactive AI assistant for eval guidance
gym Practice agent eval patterns
Reports:
replay Open a trajectory diff for one test
visualize Generate a visual HTML report from results
trends Performance trends over time
Production:
watch Re-run checks on file change (local dev)
monitor Continuous regression detection (+ Slack alerts)
CI/CD:
ci comment Post results to a GitHub PR
badge Generate shields.io status badge
init --ci Generate GitHub Actions workflow
Advanced:
login Connect to EvalView Cloud
logout Disconnect from EvalView Cloud
whoami Show current cloud login status
feedback Open a pre-filled GitHub issue
skill Test Claude Code skills
trace Trace LLM calls in scripts
traces Query stored trace data
expand Generate test variations with LLM
Options:
--version Show the version and exit.
--help Show this message and exit.
Commands:
badge Generate a shields.io-compatible badge for your README.
capture \U0001f3af Capture real traffic as tests � tests from real usage, not...
chat Interactive chat interface for EvalView.
check Check current behavior against snapshot baseline.
ci CI/CD integration commands.
compare Run the same tests against two agent endpoints and compare...
demo Live regression demo � see EvalView catch and auto-heal...
expand Expand test cases into variations using LLM.
feedback Send feedback, report a bug, or request a feature.
generate Generate a draft regression suite from live agent probing.
import Convert production logs into EvalView test cases.
init Initialize EvalView in the current directory.
login Connect to EvalView Cloud.
logout Disconnect from EvalView Cloud.
mcp Manage MCP contracts (detect external server interface drift).
monitor Continuously check for regressions with optional Slack alerts.
openclaw OpenClaw integration � install skills and manage the...
replay Replay a test and show full trajectory diff vs baseline.
run Run test cases against the agent.
skill Commands for testing Claude Code skills.
snapshot Run tests and snapshot passing results as baseline.
trace Trace LLM calls in any Python script.
traces Query and manage local trace storage.
trends Show performance trends over time.
visualize Generate a visual HTML report, optionally comparing multiple...
watch Watch for file changes and re-run regression checks.
whoami Show current cloud login status.
Badge
Paste this into the project README to show the current result:
[](https://doesitinstall.com/s/io.github.hidai25__evalview-mcp.html)