What We Build - AI Lead Qualification
Rule-based lead scoring works until the rules run out of signal. Company size, industry code and form fields tell you very little about whether a company has the problem you solve — which is why so many scoring models end up mildly correlated with nothing in particular.
AI qualification reads what a company actually says about itself: its website, its job postings, its filings, and the text of the enquiry someone submitted. That is a much richer signal, and it is exactly the kind of input a language model handles better than a rule.
6 min read5 sectionsWhat We Build
What you'll take away
- AI beats rules when the signal is unstructured text. It does not beat rules on data you already have in structured form.
- Data quality first. AI on a messy CRM produces confident, well-formatted, wrong answers at scale.
- An evaluation harness is mandatory. Without one you cannot prove the model beats the rule it replaced.
- Scoring decides routing and priority. The human still decides what to do about it.
What we build
- Fit scoring from unstructured signals
- Company website copy, job postings, published filings, technology signals and news read and scored against your written ICP — rather than inferred from a size bracket and an industry code.
- Enquiry scoring
- The text of an inbound enquiry scored on problem fit, budget signals and urgency. Especially valuable for services businesses, where senior time is the constraint being protected.
- Behavioural scoring
- Product usage, site behaviour and engagement combined with fit into a single score that drives sequence, priority and routing.
- Routing on the score
- High-fit leads to the right owner immediately with a context package; low-fit into nurture or a self-serve path. The score has to change what happens or it is decoration.
- Account research briefs
- A generated one-page summary of why a company scored as it did, so the rep starts the conversation informed rather than trusting a number.
- The evaluation harness
- A held-out set of historical leads with known outcomes, so you can measure whether the model actually predicts better than your existing rule — before it goes live and continuously afterwards.
How we work
Define the target honestly
What are we predicting — a meeting, a qualified opportunity, a closed win? Each produces a different model. Most failed scoring projects skipped this and optimised for whatever was easiest to label.
Assemble the historical set
Past leads with known outcomes, split into training and held-out evaluation. If your CRM history is not good enough to support this, that is the first thing we fix — and we will tell you that before taking the project.
Establish the baseline
Measure how well your current rule or manual triage performs on the held-out set. Without a baseline, "the AI works" is an unfalsifiable claim.
Build and evaluate
Start with the simplest approach that could work, measure against the baseline, and add complexity only where it demonstrably earns its place. Frequently a well-designed prompt over enriched text beats a trained model and is far easier to maintain.
Deploy with a fallback
Scores written to the CRM with the reasoning attached, routing driven by the score, and an automatic fallback to the previous rule if the model is unavailable or confidence is low.
Monitor for drift
Continuous measurement of prediction quality against actual outcomes, with alerting when performance degrades. Models drift as the business changes; unmonitored, they degrade silently.
What you get
- A scoring service in production
- Running against new records in real time, writing scores and reasoning back to the CRM where your team already works.
- Routing driven by the score
- High-fit leads reaching the right person immediately with context, low-fit routed away from expensive human time.
- An evaluation report
- Measured performance against your previous rule on held-out historical data. If it does not beat the baseline, we say so.
- Monitoring and drift alerting
- Ongoing measurement against actual outcomes, so degradation surfaces before it costs you a quarter of misrouted leads.
- Documentation and a fallback path
- How the model works, what it uses, how to retrain it, and what happens automatically when it is unavailable.
Where this works best
| Situation | Why AI helps | Typical result |
|---|---|---|
| Services firm, partners take every enquiry call | Enquiry text carries fit and urgency signal that no form field captures | Senior time concentrated on viable conversations |
| High inbound volume, thin qualification | Reading company context at scale is not humanly feasible | Reps work a prioritised queue instead of arrival order |
| ICP defined by problem, not firmographics | Structured data cannot express "companies with this specific problem" | Scoring that reflects the actual ICP for the first time |
| Product-led with a sales overlay | Combining usage behaviour with company context needs both signal types | Sales calls the self-serve accounts worth calling |
| Low volume, long cycles, few data points | Not enough history to train or evaluate anything | We would recommend improving the rule instead |
Why our approach differs
- We measure before claiming
- Every engagement includes an evaluation against your existing baseline on held-out data. Most AI scoring projects skip this, which is why so many quietly stop being used.
- Simplest thing that works
- Often a well-designed prompt over enriched text outperforms a trained model and is far cheaper to maintain. We are not attached to complexity.
- Reasoning, not just a number
- Every score comes with why. A rep who can see the reasoning trusts and uses the score; a bare number gets ignored within a month.
- Fallback by design
- If the model is unavailable or unsure, routing falls back to the previous rule automatically. Nothing sits unrouted because an API had a bad afternoon.
- We will tell you not to do it
- If your data volume or history cannot support evaluation, AI scoring is not the right project. We say that during the diagnostic rather than after invoicing.
Frequently asked questions
How is AI lead qualification different from lead scoring?
- Traditional lead scoring uses structured fields — size, industry, form answers — with manually assigned weights. AI qualification reads unstructured text such as website copy, job postings and enquiry wording, which carries far more signal about whether a company has the problem you solve.
How much historical data do we need?
- Enough labelled outcomes to evaluate meaningfully — as a rough guide, several hundred leads with known results. Below that, we would recommend improving your rules and instrumentation first, and we will say so during the diagnostic.
Will the model make mistakes?
- Yes, as does your current rule and your current manual triage. The question is whether it makes fewer, which is exactly what the evaluation harness measures. We deploy only when it beats the baseline on held-out data.
Does this replace human qualification?
- No. It decides priority and routing so that human attention goes to the right conversations. The discovery call, the judgement and the decision to pursue remain human — that is where the actual qualification happens.
What about data protection?
- Models run against data you control, with configurable providers including options that keep processing within the EU. Which company data is sent where is an explicit design decision made with you, not a default we impose.
Scoring that provably beats the rule it replaced
We build AI qualification on your data with an evaluation harness against your current baseline. If it does not outperform what you already have, we will tell you.