Fit score agent

Designed customizable AI fit scoring around user-defined criteria, visible reasoning, and clear prioritization—turning a traditionally black-box category into a system sales teams could understand and act on.

A missing feature with a trust problem

Seam was losing deals because it did not offer account scoring. But customers were not asking for another version of 6sense or Demandbase. Their core frustration with existing products was that a score appeared without showing how it had been derived.

The opportunity was not simply to add scoring. It was to design a system users could trust enough to use in real prioritization decisions.

Letting users describe fit in their own language

Early concepts used structured response chips. They made the setup feel controlled, but also forced nuanced sales strategy into predefined categories. I replaced them with open text fields so users could paste directly from an ICP, sales playbook, or strategy document.

The final setup centered on three questions: what should the agent evaluate, what makes a good fit, and what makes a bad fit. The framing was simple enough to start quickly while leaving room for company-specific judgment.

Replacing fake precision with a useful scale

A zero-to-one-hundred score implied a level of precision the model could not meaningfully support. For a sales rep, the difference between 66 and 67 did not change what to do next.

I introduced an A-to-D scale instead. The broader bands made the result easier to interpret, compare, and translate into an account-prioritization workflow.

Designing explanations for scanning

We debated whether to show the model’s raw text output or transform it into a structured score card. I chose the score card because sales teams needed to scan the decisive factors, not read an unedited model response.

The structured breakdown kept the reasoning visible while making it easier to compare accounts and identify the evidence behind each grade.

Validating before spending at scale

After launch, we added backtesting so users could try their scoring logic on a small sample before committing credits to a full run. This gave teams a fast feedback loop for refining ambiguous criteria and catching unexpected interpretations.

Impact and what I learned

Fit Score became a core part of the product:

  • 100% adoption across all paying organizations
  • About 80% of total credit usage, making it one of the largest revenue-generating features
  • Competitors copied the approach within months of its September 2025 launch

Explainability cannot be added after a black-box decision. It has to shape the full interaction: how users express their logic, how the system communicates progress, how results are grouped, and how the evidence can be inspected.

The most useful AI score was not the one that looked most precise. It was the one a team could define, challenge, and confidently turn into action.