does-it-install?

← All servers

io.github.hidai25/evalview-mcp

Regression testing for AI agents. Golden baselines, CI/CD, LangGraph, CrewAI, OpenAI, Claude.

io.github.hidai25/evalview-mcp · Repository · version 0.6.0 · 128 stars · listed from registry

Install

The sweep installs into a clean prefix, but this is the command a user would run:

uvx evalview

Results by platform

Linux

MCP handshake failed2026-08-16 · probed over pypi

2 recorded runs, oldest first.

handshake failed: MCP error -32000: Connection closed
--- stderr (tail) ---
…s.

  New here? Start with:
    demo                    See it work in 30 seconds
    init                    Detect your agent and create first tests

  Regression Gating — did my agent change?
    snapshot                Capture current behavior as baseline
    snapshot list            List saved baselines
    snapshot show <name>    Inspect a baseline
    snapshot delete <name>  Remove a baseline
    check                   Compare against baseline — catch regressions

  Evaluation — how good is my agent?
    generate                Auto-generate tests from live probing
    run                     Execute tests, score with LLM judge

  Build Your Test Suite:
    capture --agent <url>   Record real traffic as tests
    capture --multi-turn    Record a conversation as a multi-turn test
    import <log_file>       Convert production logs into EvalView tests
    expand                  Generate test variations with LLM
    compare                 Compare two agent endpoints on the same suite

  Explore & Learn:
    chat                    Interactive AI assistant for eval guidance
    gym                     Practice agent eval patterns

  Reports:
    replay                  Open a trajectory diff for one test
    visualize               Generate a visual HTML report from results
    trends                  Performance trends over time

  Production:
    watch                   Re-run checks on file change (local dev)
    monitor                 Continuous regression detection (+ Slack alerts)

  CI/CD:
    ci comment              Post results to a GitHub PR
    badge                   Generate shields.io status badge
    init --ci               Generate GitHub Actions workflow

  Advanced:
    login                   Connect to EvalView Cloud
    logout                  Disconnect from EvalView Cloud
    whoami                  Show current cloud login status
    feedback                Open a pre-filled GitHub issue
    skill                   Test Claude Code skills
    trace                   Trace LLM calls in scripts
    traces                  Query stored trace data
    expand                  Generate test variations with LLM

Options:
  --version  Show the version and exit.
  --help     Show this message and exit.

Commands:
  badge      Generate a shields.io-compatible badge for your README.
  capture    🎯 Capture real traffic as tests — tests from real usage, not...
  chat       Interactive chat interface for EvalView.
  check      Check current behavior against snapshot baseline.
  ci         CI/CD integration commands.
  compare    Run the same tests against two agent endpoints and compare...
  demo       Live regression demo — see EvalView catch and auto-heal...
  expand     Expand test cases into variations using LLM.
  feedback   Send feedback, report a bug, or request a feature.
  generate   Generate a draft regression suite from live agent probing.
  import     Convert production logs into EvalView test cases.
  init       Initialize EvalView in the current directory.
  login      Connect to EvalView Cloud.
  logout     Disconnect from EvalView Cloud.
  mcp        Manage MCP contracts (detect external server interface drift).
  monitor    Continuously check for regressions with optional Slack alerts.
  openclaw   OpenClaw integration — install skills and manage the...
  replay     Replay a test and show full trajectory diff vs baseline.
  run        Run test cases against the agent.
  skill      Commands for testing Claude Code skills.
  snapshot   Run tests and snapshot passing results as baseline.
  trace      Trace LLM calls in any Python script.
  traces     Query and manage local trace storage.
  trends     Show performance trends over time.
  visualize  Generate a visual HTML report, optionally comparing multiple...
  watch      Watch for file changes and re-run regression checks.
  whoami     Show current cloud login status.

macOS

MCP handshake failed2026-08-16 · probed over pypi

2 recorded runs, oldest first.

handshake failed: MCP error -32000: Connection closed
--- stderr (tail) ---
…s.

  New here? Start with:
    demo                    See it work in 30 seconds
    init                    Detect your agent and create first tests

  Regression Gating — did my agent change?
    snapshot                Capture current behavior as baseline
    snapshot list            List saved baselines
    snapshot show <name>    Inspect a baseline
    snapshot delete <name>  Remove a baseline
    check                   Compare against baseline — catch regressions

  Evaluation — how good is my agent?
    generate                Auto-generate tests from live probing
    run                     Execute tests, score with LLM judge

  Build Your Test Suite:
    capture --agent <url>   Record real traffic as tests
    capture --multi-turn    Record a conversation as a multi-turn test
    import <log_file>       Convert production logs into EvalView tests
    expand                  Generate test variations with LLM
    compare                 Compare two agent endpoints on the same suite

  Explore & Learn:
    chat                    Interactive AI assistant for eval guidance
    gym                     Practice agent eval patterns

  Reports:
    replay                  Open a trajectory diff for one test
    visualize               Generate a visual HTML report from results
    trends                  Performance trends over time

  Production:
    watch                   Re-run checks on file change (local dev)
    monitor                 Continuous regression detection (+ Slack alerts)

  CI/CD:
    ci comment              Post results to a GitHub PR
    badge                   Generate shields.io status badge
    init --ci               Generate GitHub Actions workflow

  Advanced:
    login                   Connect to EvalView Cloud
    logout                  Disconnect from EvalView Cloud
    whoami                  Show current cloud login status
    feedback                Open a pre-filled GitHub issue
    skill                   Test Claude Code skills
    trace                   Trace LLM calls in scripts
    traces                  Query stored trace data
    expand                  Generate test variations with LLM

Options:
  --version  Show the version and exit.
  --help     Show this message and exit.

Commands:
  badge      Generate a shields.io-compatible badge for your README.
  capture    🎯 Capture real traffic as tests — tests from real usage, not...
  chat       Interactive chat interface for EvalView.
  check      Check current behavior against snapshot baseline.
  ci         CI/CD integration commands.
  compare    Run the same tests against two agent endpoints and compare...
  demo       Live regression demo — see EvalView catch and auto-heal...
  expand     Expand test cases into variations using LLM.
  feedback   Send feedback, report a bug, or request a feature.
  generate   Generate a draft regression suite from live agent probing.
  import     Convert production logs into EvalView test cases.
  init       Initialize EvalView in the current directory.
  login      Connect to EvalView Cloud.
  logout     Disconnect from EvalView Cloud.
  mcp        Manage MCP contracts (detect external server interface drift).
  monitor    Continuously check for regressions with optional Slack alerts.
  openclaw   OpenClaw integration — install skills and manage the...
  replay     Replay a test and show full trajectory diff vs baseline.
  run        Run test cases against the agent.
  skill      Commands for testing Claude Code skills.
  snapshot   Run tests and snapshot passing results as baseline.
  trace      Trace LLM calls in any Python script.
  traces     Query and manage local trace storage.
  trends     Show performance trends over time.
  visualize  Generate a visual HTML report, optionally comparing multiple...
  watch      Watch for file changes and re-run regression checks.
  whoami     Show current cloud login status.

Windows

MCP handshake failed2026-08-16 · probed over pypi

2 recorded runs, oldest first.

handshake failed: MCP error -32000: Connection closed
--- stderr (tail) ---
…nit                    Detect your agent and create first tests

  Regression Gating � did my agent change?
    snapshot                Capture current behavior as baseline
    snapshot list            List saved baselines
    snapshot show <name>    Inspect a baseline
    snapshot delete <name>  Remove a baseline
    check                   Compare against baseline � catch regressions

  Evaluation � how good is my agent?
    generate                Auto-generate tests from live probing
    run                     Execute tests, score with LLM judge

  Build Your Test Suite:
    capture --agent <url>   Record real traffic as tests
    capture --multi-turn    Record a conversation as a multi-turn test
    import <log_file>       Convert production logs into EvalView tests
    expand                  Generate test variations with LLM
    compare                 Compare two agent endpoints on the same suite

  Explore & Learn:
    chat                    Interactive AI assistant for eval guidance
    gym                     Practice agent eval patterns

  Reports:
    replay                  Open a trajectory diff for one test
    visualize               Generate a visual HTML report from results
    trends                  Performance trends over time

  Production:
    watch                   Re-run checks on file change (local dev)
    monitor                 Continuous regression detection (+ Slack alerts)

  CI/CD:
    ci comment              Post results to a GitHub PR
    badge                   Generate shields.io status badge
    init --ci               Generate GitHub Actions workflow

  Advanced:
    login                   Connect to EvalView Cloud
    logout                  Disconnect from EvalView Cloud
    whoami                  Show current cloud login status
    feedback                Open a pre-filled GitHub issue
    skill                   Test Claude Code skills
    trace                   Trace LLM calls in scripts
    traces                  Query stored trace data
    expand                  Generate test variations with LLM

Options:
  --version  Show the version and exit.
  --help     Show this message and exit.

Commands:
  badge      Generate a shields.io-compatible badge for your README.
  capture    \U0001f3af Capture real traffic as tests � tests from real usage, not...
  chat       Interactive chat interface for EvalView.
  check      Check current behavior against snapshot baseline.
  ci         CI/CD integration commands.
  compare    Run the same tests against two agent endpoints and compare...
  demo       Live regression demo � see EvalView catch and auto-heal...
  expand     Expand test cases into variations using LLM.
  feedback   Send feedback, report a bug, or request a feature.
  generate   Generate a draft regression suite from live agent probing.
  import     Convert production logs into EvalView test cases.
  init       Initialize EvalView in the current directory.
  login      Connect to EvalView Cloud.
  logout     Disconnect from EvalView Cloud.
  mcp        Manage MCP contracts (detect external server interface drift).
  monitor    Continuously check for regressions with optional Slack alerts.
  openclaw   OpenClaw integration � install skills and manage the...
  replay     Replay a test and show full trajectory diff vs baseline.
  run        Run test cases against the agent.
  skill      Commands for testing Claude Code skills.
  snapshot   Run tests and snapshot passing results as baseline.
  trace      Trace LLM calls in any Python script.
  traces     Query and manage local trace storage.
  trends     Show performance trends over time.
  visualize  Generate a visual HTML report, optionally comparing multiple...
  watch      Watch for file changes and re-run regression checks.
  whoami     Show current cloud login status.

Badge

Paste this into the project README to show the current result:

[![does it install](https://img.shields.io/endpoint?url=https://doesitinstall.com/badge/io.github.hidai25__evalview-mcp.json)](https://doesitinstall.com/s/io.github.hidai25__evalview-mcp.html)