> ## Documentation Index
> Fetch the complete documentation index at: https://mcpjam-mintlify-docs-update-pr-3812-1786323532738.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Compare a run against a baseline

> Compare this run against a baseline run: per-case status (`regressed`, `fixed`, `new_case`, `removed_case`, `changed`), per-scorer pass-rate and mean deltas from the evaluation contract, and whether the evaluation config changed. Omit `baseRunId` to compare against the nearest earlier **completed** run in the same suite. Returns `404` with `details.reason` = `BASELINE_NOT_FOUND` when there is no comparable predecessor — that is an incomplete comparison, not a failing one. A scorer whose `definitionChanged` is `true` was graded by a different definition on each side, so its delta is not a regression.



## OpenAPI

````yaml /reference/openapi.json get /projects/{projectId}/eval-runs/{runId}/compare
openapi: 3.1.0
info:
  title: MCPJam API
  version: 1.0.0-preview
  description: >-
    Programmatic access to MCP servers saved in your MCPJam projects — live
    diagnostics (validate, inspect, export) and operations: call tools, render
    prompts, run eval suites asynchronously and poll their results, and import
    OAuth tokens.


    **The API is in preview**: the surface may change while we finish the
    design. Error `code` values are stable; error `message` strings are not.
    Write clients that ignore unknown response fields.
  contact:
    name: MCPJam
    url: https://github.com/MCPJam/inspector/issues
servers:
  - url: https://app.mcpjam.com/api/v1
    description: Hosted MCPJam
security:
  - bearerAuth: []
tags:
  - name: Clients
    description: >-
      Clients — the named, reusable configurations that define how MCPJam
      connects to and talks to your MCP servers. The original `/hosts` paths
      remain as deprecated, ID-only compatibility aliases with their original
      DTOs and their original (tokenless) write contracts; every alias response
      carries `Deprecation: true`. New integrations should use `/clients`.
  - name: Environments
    description: >-
      Project environments: named, live-editable execution bundles (one host, an
      optional standalone server group, optionally pinned skills and plugin
      versions) that eval suites and journeys run against. Distinct from Sandbox
      images, which are Computer base images. Reads require project membership;
      every write requires project admin.
  - name: Plugins
    description: >-
      Agent Plugins imported into a project — read-only inventory and version
      detail.
  - name: Skills
    description: >-
      Cloud Skills: authored SKILL.md files stored in a project. Read-only here.
      Environments pin skills by id (`skillSelection.skillIds`) and eval runs
      pin them with `--compose-skill`, so this surface exists to give an
      unattended caller those ids; authoring is an app flow behind a beta gate.
  - name: Sandbox images
    description: >-
      Custom Computer images: a digest-pinned Dockerfile built into an immutable
      image your project's computers boot from.
  - name: Server diagnostics
    description: Connect-level health checks against a saved MCP server.
  - name: Primitives
    description: 'The server''s MCP primitives: tools, prompts, and resources.'
  - name: Export
    description: Full-server snapshots for diffing and CI.
  - name: Execution
    description: 'Run the server''s primitives: call tools, render prompts.'
  - name: Eval runs
    description: >-
      Asynchronous eval suite runs: create with 202, poll status, iterations,
      and traces.
  - name: Conformance runs
    description: >-
      Ingest MCP spec-conformance results from the SDK/CLI into project-owned
      history. Distinct from Eval runs (authored LLM cases) and from directory
      readiness.
  - name: Server connections
    description: >-
      Connect an MCP server URL to a project, authorizing in a browser when the
      server requires it.
  - name: OAuth
    description: 'Bring-your-own OAuth: import externally obtained tokens for a server.'
  - name: Scenarios
    description: >-
      Read-only access to the scenarios published from a project: listing,
      settings, attached servers, and share links.
  - name: Catalog
    description: >-
      Discover the resources the other routes operate on: your account,
      projects, servers, eval suites, and chat sessions.
  - name: Tunnels
    description: >-
      Relay tunnels that expose local MCP servers through a public URL,
      registered as first-class project servers (the `mcpjam cloud tunnel` CLI
      flow).
  - name: Agent
    description: >-
      Headless agent turns over the public API: send a message history, the
      server runs one assistant turn with project-scoped workspace tools (eval
      reads + suite creation) on a pinned hosted model, and returns the reply
      plus created-resource references.
  - name: Swarms
    description: >-
      Personas, journeys and swarm containers — the authoring half of Swarms —
      plus the model-backed generation that drafts them.
  - name: Swarm runs
    description: >-
      Launching journeys and reading what they produced. Launching SPENDS — see
      the per-operation notes.
  - name: Swarm insights
    description: >-
      What a swarm run revealed. The scorecard and findings are deterministic
      and free; requesting wave insights runs models and draws on your shared
      daily ledger.
  - name: User testing
    description: >-
      Publishing an environment for real visitors, and controlling who can reach
      it. Several of these NARROW access and take effect immediately.
  - name: Directory readiness
    description: >-
      Grade a saved server against a publisher's listing requirements:
      Anthropic's connector directory or OpenAI's plugin directory. Reported as
      lane status and coverage, never as a numeric score, and excluded from
      `pooledConformanceScore`. Deterministic grading is free; model-backed
      experience observations are an explicit opt-in that consumes MCPJam
      credits and can never decide a verdict.
  - name: Registry
    description: >-
      Search the scraped MCP directories (Claude, ChatGPT, and any future
      source), list curated/org registry cards, and install them into a project.
      Install writes a `servers` row and provenance — it does not open a live
      session. There is no catalog-uninstall route: delete the project server
      instead. Directory reads require a bearer (including minted guest tokens)
      but do not materialize a user. Card/connection reads and all writes are
      authed-non-guest.
paths:
  /projects/{projectId}/eval-runs/{runId}/compare:
    get:
      tags:
        - Eval runs
      summary: Compare a run against a baseline
      description: >-
        Compare this run against a baseline run: per-case status (`regressed`,
        `fixed`, `new_case`, `removed_case`, `changed`), per-scorer pass-rate
        and mean deltas from the evaluation contract, and whether the evaluation
        config changed. Omit `baseRunId` to compare against the nearest earlier
        **completed** run in the same suite. Returns `404` with `details.reason`
        = `BASELINE_NOT_FOUND` when there is no comparable predecessor — that is
        an incomplete comparison, not a failing one. A scorer whose
        `definitionChanged` is `true` was graded by a different definition on
        each side, so its delta is not a regression.
      operationId: compareEvalRun
      parameters:
        - $ref: '#/components/parameters/projectId'
        - $ref: '#/components/parameters/runId'
        - name: baseRunId
          in: query
          required: false
          description: >-
            Run ID to compare against. Omit to use the nearest earlier completed
            run in the same suite. Mutually exclusive with baseCommitSha.
          schema:
            type: string
        - name: baseCommitSha
          in: query
          required: false
          description: >-
            Source commit SHA to compare against, resolved to the completed run
            in this suite recorded against it. Mutually exclusive with
            baseRunId; sending both is a 400. A SHA matching no completed run is
            the ordinary 404 with details.reason = BASELINE_NOT_FOUND - an
            incomplete comparison, not a regression.
          schema:
            type: string
      responses:
        '200':
          description: The comparison.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/EvalRunCompare'
        '401':
          $ref: '#/components/responses/Unauthorized'
        '403':
          $ref: '#/components/responses/Forbidden'
        '404':
          $ref: '#/components/responses/NotFound'
        '429':
          $ref: '#/components/responses/RateLimited'
        '500':
          $ref: '#/components/responses/InternalError'
components:
  parameters:
    projectId:
      name: projectId
      in: path
      required: true
      description: ID of the hosted project that contains the server.
      schema:
        type: string
    runId:
      name: runId
      in: path
      required: true
      description: Eval run ID, as returned by `POST /eval-runs`.
      schema:
        type: string
  schemas:
    EvalRunCompare:
      type: object
      properties:
        suite:
          type: object
          properties:
            id:
              type: string
            name:
              type: string
          required:
            - id
            - name
        baseline:
          type: object
          properties:
            policy:
              type: string
              enum:
                - previous_completed
                - previous_completed_same_environment
                - run
                - commit_sha
              description: >-
                How the baseline was chosen. `commit_sha` means it was resolved
                from the requested source SHA.
            baseRunId:
              type: string
            baseCommitSha:
              type: string
              description: >-
                The pinned source SHA, echoed back for the `commit_sha` policy
                only.
            matchCount:
              type: integer
              description: >-
                Present ONLY when uniqueness could not be established - the SHA
                matched several eligible runs, or the bounded lookup saturated.
                Absent means the match was unambiguous; do not default it to 1.
            matchCountTruncated:
              type: boolean
              description: >-
                `matchCount` is a FLOOR, not a total - including when it reads
                1. Read it together with `matchCount`; a count without this flag
                would assert a uniqueness nobody checked.
          required:
            - policy
            - baseRunId
        baseRun:
          $ref: '#/components/schemas/EvalRunCompareSide'
        compareRun:
          $ref: '#/components/schemas/EvalRunCompareSide'
        passSummary:
          description: >-
            Run-summary counters. Named `passSummary`, not `scores`, so it
            cannot be confused with `scoreContract` — the two answer different
            questions.
          type: object
          properties:
            passRatePercent:
              $ref: '#/components/schemas/NumericDiff'
            total:
              $ref: '#/components/schemas/NumericDiff'
            passed:
              $ref: '#/components/schemas/NumericDiff'
            failed:
              $ref: '#/components/schemas/NumericDiff'
          required:
            - passRatePercent
            - total
            - passed
            - failed
        metrics:
          type: object
          properties:
            wallDurationMs:
              $ref: '#/components/schemas/NumericDiff'
            totalTokens:
              $ref: '#/components/schemas/NumericDiff'
            estimatedCostUsd:
              $ref: '#/components/schemas/NumericDiff'
          required:
            - wallDurationMs
            - totalTokens
            - estimatedCostUsd
        skills:
          type: object
          nullable: true
          description: >-
            Which skills changed between the two runs — the configuration
            attribution that usually explains the case-level differences. `null`
            when NEITHER run recorded pinned skills (an empty section would
            instead claim no skills were involved). Absent when the deployment
            predates skill attribution, so clients must tolerate all three
            states.
          properties:
            base:
              $ref: '#/components/schemas/EvalRunCompareSkillsSide'
            compare:
              $ref: '#/components/schemas/EvalRunCompareSkillsSide'
            changes:
              type: array
              description: >-
                Added, removed and changed skills only, changed first. Unchanged
                skills are counted, not listed.
              items:
                $ref: '#/components/schemas/EvalRunCompareSkillChange'
            unchangedCount:
              type: integer
          required:
            - base
            - compare
            - changes
            - unchangedCount
        scoreContract:
          type: object
          properties:
            base:
              $ref: '#/components/schemas/ScoreContractSide'
            compare:
              $ref: '#/components/schemas/ScoreContractSide'
            evaluationConfigChanged:
              type: boolean
            scorers:
              type: array
              items:
                $ref: '#/components/schemas/ScoreContractScorer'
          required:
            - base
            - compare
            - evaluationConfigChanged
            - scorers
        cases:
          type: array
          items:
            $ref: '#/components/schemas/EvalRunCompareCase'
      required:
        - suite
        - baseline
        - baseRun
        - compareRun
        - passSummary
        - metrics
        - scoreContract
        - cases
    EvalRunCompareSide:
      type: object
      properties:
        id:
          type: string
        runNumber:
          type: integer
        result:
          type: string
        createdAt:
          type: integer
        completedAt:
          type:
            - integer
            - 'null'
        summary:
          oneOf:
            - type: object
              properties:
                total:
                  type: integer
                passed:
                  type: integer
                failed:
                  type: integer
                passRate:
                  type: number
              required:
                - total
                - passed
                - failed
                - passRate
            - type: 'null'
        environment:
          type: object
          properties:
            id:
              type: string
            name:
              type:
                - string
                - 'null'
          required:
            - id
            - name
        effectiveModelId:
          type: string
          description: Model the compared run actually executed with.
        modelSource:
          type: string
          enum:
            - client_default
            - override
      required:
        - id
        - runNumber
        - result
        - createdAt
        - completedAt
        - summary
    NumericDiff:
      description: >-
        A base/compare pair with its delta. Rate-valued instances are FRACTIONS
        unless the field name ends in `Percent`.
      type: object
      properties:
        base:
          type:
            - number
            - 'null'
        compare:
          type:
            - number
            - 'null'
        delta:
          type:
            - number
            - 'null'
        percentDelta:
          type:
            - number
            - 'null'
      required:
        - base
        - compare
        - delta
        - percentDelta
    EvalRunCompareSkillsSide:
      type: object
      properties:
        excluded:
          type: boolean
          description: >-
            This run deliberately ran with skills disabled — distinct from a run
            that simply pinned none.
        count:
          type: integer
          description: How many skills this run pinned.
      required:
        - excluded
        - count
    EvalRunCompareSkillChange:
      type: object
      properties:
        key:
          type: string
          description: Stable match key; opaque, safe for list keys and dedupe.
        name:
          type: string
        modelRef:
          type: string
          description: Namespaced runtime address for a plugin-channel skill.
        channels:
          type: array
          items:
            type: string
            enum:
              - host
              - environment
              - plugin
              - mcp-server
        kind:
          type: string
          enum:
            - added
            - removed
            - changed
        renamedFrom:
          type: string
          description: >-
            Present when the skill was renamed between the runs; it is still
            matched as ONE skill by its logical id.
        base:
          $ref: '#/components/schemas/EvalRunCompareSkillSide'
        compare:
          $ref: '#/components/schemas/EvalRunCompareSkillSide'
        versionDelta:
          type: string
          description: >-
            Human-readable revision move (`v3 → v4`), present only when BOTH
            sides recorded a number. A change with no delta is a real content
            change whose revisions are unknown, not a smaller change.
      required:
        - key
        - name
        - channels
        - kind
    ScoreContractSide:
      type: object
      properties:
        evaluationConfigHash:
          type:
            - string
            - 'null'
        scoreIntegrity:
          type:
            - string
            - 'null'
          enum:
            - valid
            - invalid
            - null
          description: >-
            `null` means no verdict was produced. A gate must treat it exactly
            like `invalid` — absent evidence is not valid evidence.
        scoredIterations:
          type: integer
        quarantinedIterations:
          type: integer
          description: >-
            Iterations with at least one row that failed to verify at ingest.
            Counted here, and excluded from every rate.
      required:
        - evaluationConfigHash
        - scoreIntegrity
        - scoredIterations
        - quarantinedIterations
    ScoreContractScorer:
      type: object
      properties:
        scorerId:
          type: string
        gating:
          type: boolean
        deterministic:
          type: boolean
        definitionChanged:
          type: boolean
          description: >-
            The same scorer id was graded under a different definition hash on
            each side. Its delta is NOT a regression — the two runs did not
            measure the same thing.
        passRate:
          $ref: '#/components/schemas/NumericDiff'
        meanValue:
          $ref: '#/components/schemas/NumericDiff'
        errorCount:
          type: object
          properties:
            base:
              type: integer
            compare:
              type: integer
          required:
            - base
            - compare
      required:
        - scorerId
        - gating
        - deterministic
        - definitionChanged
        - passRate
        - meanValue
        - errorCount
    EvalRunCompareCase:
      type: object
      properties:
        caseKey:
          type: string
        title:
          type: string
        status:
          type: string
          enum:
            - unchanged_passed
            - unchanged_failed
            - regressed
            - fixed
            - new_case
            - removed_case
            - changed
        configChanged:
          type: boolean
          description: The scenario's own config (prompt, steps, expectations) changed.
        evaluationConfigChanged:
          type: boolean
          description: This case's evaluation config changed.
        scoreDeltas:
          type: array
          items:
            $ref: '#/components/schemas/CaseScoreDelta'
        base:
          $ref: '#/components/schemas/EvalRunCompareCaseSide'
        compare:
          $ref: '#/components/schemas/EvalRunCompareCaseSide'
      required:
        - caseKey
        - title
        - status
        - configChanged
        - evaluationConfigChanged
        - scoreDeltas
        - base
        - compare
    Error:
      type: object
      required:
        - code
        - message
      properties:
        code:
          type: string
          description: >-
            Stable, machine-readable error code. New codes may be added over
            time; treat unknown codes as non-retryable failures unless the HTTP
            status says otherwise.
          enum:
            - UNAUTHORIZED
            - FORBIDDEN
            - NOT_FOUND
            - CONFLICT
            - VALIDATION_ERROR
            - RATE_LIMITED
            - FEATURE_NOT_SUPPORTED
            - SERVER_UNREACHABLE
            - TIMEOUT
            - OAUTH_REQUIRED
            - INTERNAL_ERROR
        message:
          type: string
          description: >-
            Human-readable description. May change between releases — don't
            match on it.
        details:
          type: object
          description: Optional, unstructured context bag.
          additionalProperties: true
    EvalRunCompareSkillSide:
      type: object
      description: One skill's identity on one side of the comparison.
      properties:
        contentHash:
          type: string
        aggregateHash:
          type: string
          description: >-
            Complete-artifact hash; present only when supporting files diverge
            it from `contentHash`.
        versionNumber:
          type: integer
          description: Authored-skill revision, when the run recorded one.
        serverSkillVersionNumber:
          type: integer
          description: MCP-captured revision, when the run recorded one.
      required:
        - contentHash
    CaseScoreDelta:
      type: object
      properties:
        scorerId:
          type: string
        gating:
          type: boolean
        deterministic:
          type: boolean
        definitionChanged:
          type: boolean
        base:
          oneOf:
            - $ref: '#/components/schemas/CaseScoreSide'
            - type: 'null'
        compare:
          oneOf:
            - $ref: '#/components/schemas/CaseScoreSide'
            - type: 'null'
        value:
          $ref: '#/components/schemas/NumericDiff'
      required:
        - scorerId
        - gating
        - deterministic
        - definitionChanged
        - base
        - compare
        - value
    EvalRunCompareCaseSide:
      type: object
      properties:
        outcome:
          type: string
          enum:
            - passed
            - failed
            - absent
        iterationIds:
          type: array
          items:
            type: string
        representativeIterationId:
          type:
            - string
            - 'null'
        error:
          type:
            - string
            - 'null'
      required:
        - outcome
        - iterationIds
        - representativeIterationId
        - error
    CaseScoreSide:
      type: object
      properties:
        status:
          type: string
          enum:
            - scored
            - error
            - skipped
            - not_applicable
        value:
          type:
            - number
            - 'null'
        passed:
          type:
            - boolean
            - 'null'
      required:
        - status
        - value
        - passed
  responses:
    Unauthorized:
      description: >-
        Missing, invalid, revoked, or orphaned key (`UNAUTHORIZED`) — or the
        **target MCP server** needs an OAuth grant (`OAUTH_REQUIRED`), which is
        a property of the server, not your key.
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/Error'
          examples:
            badKey:
              summary: Invalid or revoked key
              value:
                code: UNAUTHORIZED
                message: Invalid API key
            oauthRequired:
              summary: Target server needs an OAuth grant
              value:
                code: OAUTH_REQUIRED
                message: Server requires OAuth authorization
    Forbidden:
      description: Key is valid but not allowed to do this.
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/Error'
          example:
            code: FORBIDDEN
            message: You do not have access to this project
    NotFound:
      description: Unknown project, server, or resource.
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/Error'
          example:
            code: NOT_FOUND
            message: Server not found
    RateLimited:
      description: >-
        Per-key rate limit exceeded (60 requests/minute sustained, bursts up to
        10). Honor `Retry-After` and back off with jitter.
      headers:
        Retry-After:
          description: Seconds to wait before retrying.
          schema:
            type: integer
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/Error'
          example:
            code: RATE_LIMITED
            message: API key rate limit exceeded. Slow down and retry.
    InternalError:
      description: Something failed on MCPJam's side.
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/Error'
          example:
            code: INTERNAL_ERROR
            message: Unexpected internal error
  securitySchemes:
    bearerAuth:
      type: http
      scheme: bearer
      description: >-
        MCPJam API key (`sk_…`). Create one at [Settings → API
        keys](https://app.mcpjam.com/settings/api-keys). Guest sessions cannot
        use the API, and API keys cannot manage other API keys.

````