Contact Us

What We Build - AI Lead Qualification

Rule-based lead scoring works until the rules run out of signal. Company size, industry code and form fields tell you very little about whether a company has the problem you solve — which is why so many scoring models end up mildly correlated with nothing in particular.

AI qualification reads what a company actually says about itself: its website, its job postings, its filings, and the text of the enquiry someone submitted. That is a much richer signal, and it is exactly the kind of input a language model handles better than a rule.

6 min read5 sectionsWhat We Build

What you'll take away

  • AI beats rules when the signal is unstructured text. It does not beat rules on data you already have in structured form.
  • Data quality first. AI on a messy CRM produces confident, well-formatted, wrong answers at scale.
  • An evaluation harness is mandatory. Without one you cannot prove the model beats the rule it replaced.
  • Scoring decides routing and priority. The human still decides what to do about it.

What we build

Fit scoring from unstructured signals
Company website copy, job postings, published filings, technology signals and news read and scored against your written ICP — rather than inferred from a size bracket and an industry code.
Enquiry scoring
The text of an inbound enquiry scored on problem fit, budget signals and urgency. Especially valuable for services businesses, where senior time is the constraint being protected.
Behavioural scoring
Product usage, site behaviour and engagement combined with fit into a single score that drives sequence, priority and routing.
Routing on the score
High-fit leads to the right owner immediately with a context package; low-fit into nurture or a self-serve path. The score has to change what happens or it is decoration.
Account research briefs
A generated one-page summary of why a company scored as it did, so the rep starts the conversation informed rather than trusting a number.
The evaluation harness
A held-out set of historical leads with known outcomes, so you can measure whether the model actually predicts better than your existing rule — before it goes live and continuously afterwards.

How we work

  1. Define the target honestly

    What are we predicting — a meeting, a qualified opportunity, a closed win? Each produces a different model. Most failed scoring projects skipped this and optimised for whatever was easiest to label.

  2. Assemble the historical set

    Past leads with known outcomes, split into training and held-out evaluation. If your CRM history is not good enough to support this, that is the first thing we fix — and we will tell you that before taking the project.

  3. Establish the baseline

    Measure how well your current rule or manual triage performs on the held-out set. Without a baseline, "the AI works" is an unfalsifiable claim.

  4. Build and evaluate

    Start with the simplest approach that could work, measure against the baseline, and add complexity only where it demonstrably earns its place. Frequently a well-designed prompt over enriched text beats a trained model and is far easier to maintain.

  5. Deploy with a fallback

    Scores written to the CRM with the reasoning attached, routing driven by the score, and an automatic fallback to the previous rule if the model is unavailable or confidence is low.

  6. Monitor for drift

    Continuous measurement of prediction quality against actual outcomes, with alerting when performance degrades. Models drift as the business changes; unmonitored, they degrade silently.

What you get

A scoring service in production
Running against new records in real time, writing scores and reasoning back to the CRM where your team already works.
Routing driven by the score
High-fit leads reaching the right person immediately with context, low-fit routed away from expensive human time.
An evaluation report
Measured performance against your previous rule on held-out historical data. If it does not beat the baseline, we say so.
Monitoring and drift alerting
Ongoing measurement against actual outcomes, so degradation surfaces before it costs you a quarter of misrouted leads.
Documentation and a fallback path
How the model works, what it uses, how to retrain it, and what happens automatically when it is unavailable.

Where this works best

Table 01
AI qualification is not universally better. These are the cases where it clearly is.
SituationWhy AI helpsTypical result
Services firm, partners take every enquiry callEnquiry text carries fit and urgency signal that no form field capturesSenior time concentrated on viable conversations
High inbound volume, thin qualificationReading company context at scale is not humanly feasibleReps work a prioritised queue instead of arrival order
ICP defined by problem, not firmographicsStructured data cannot express "companies with this specific problem"Scoring that reflects the actual ICP for the first time
Product-led with a sales overlayCombining usage behaviour with company context needs both signal typesSales calls the self-serve accounts worth calling
Low volume, long cycles, few data pointsNot enough history to train or evaluate anythingWe would recommend improving the rule instead

Why our approach differs

We measure before claiming
Every engagement includes an evaluation against your existing baseline on held-out data. Most AI scoring projects skip this, which is why so many quietly stop being used.
Simplest thing that works
Often a well-designed prompt over enriched text outperforms a trained model and is far cheaper to maintain. We are not attached to complexity.
Reasoning, not just a number
Every score comes with why. A rep who can see the reasoning trusts and uses the score; a bare number gets ignored within a month.
Fallback by design
If the model is unavailable or unsure, routing falls back to the previous rule automatically. Nothing sits unrouted because an API had a bad afternoon.
We will tell you not to do it
If your data volume or history cannot support evaluation, AI scoring is not the right project. We say that during the diagnostic rather than after invoicing.

Frequently asked questions

How is AI lead qualification different from lead scoring?

Traditional lead scoring uses structured fields — size, industry, form answers — with manually assigned weights. AI qualification reads unstructured text such as website copy, job postings and enquiry wording, which carries far more signal about whether a company has the problem you solve.

How much historical data do we need?

Enough labelled outcomes to evaluate meaningfully — as a rough guide, several hundred leads with known results. Below that, we would recommend improving your rules and instrumentation first, and we will say so during the diagnostic.

Will the model make mistakes?

Yes, as does your current rule and your current manual triage. The question is whether it makes fewer, which is exactly what the evaluation harness measures. We deploy only when it beats the baseline on held-out data.

Does this replace human qualification?

No. It decides priority and routing so that human attention goes to the right conversations. The discovery call, the judgement and the decision to pursue remain human — that is where the actual qualification happens.

What about data protection?

Models run against data you control, with configurable providers including options that keep processing within the EU. Which company data is sent where is an explicit design decision made with you, not a default we impose.
Measured, not asserted

Scoring that provably beats the rule it replaced

We build AI qualification on your data with an evaluation harness against your current baseline. If it does not outperform what you already have, we will tell you.