Blog / Article

Building PoD®: Rethinking financial risk modelling

By Jonathan Davies

Company Watch is, at its core, a data company. We collect, clean, and structure an enormous volume of information on UK businesses, covering everything from audited financial accounts to court judgments to director-level histories to statutory filing behaviour. That breadth of data is the foundation of what we do.

So when I started thinking seriously about what a modern insolvency prediction model should look like, the question was less about what data we could access and more about whether we were genuinely using all of it. The answer, if I’m honest, was no. And that felt like a problem worth solving.

A graphic of the Company Watch platform, demonstrating its capabilities.

PoD® in action on the Company Watch platform.

The case for a new model

The H-Score® has been the backbone of Company Watch’s risk scoring for decades, and it performs well for what it was designed to do: take a company’s audited financial accounts and derive a structured risk assessment from them. For a regression-based model built on financial statement analysis, it is rigorous and well-understood.

The limitation is structural rather than methodological. Financial accounts are filed annually, and in practice the data reaching the model can be anywhere from 12 to 18 months old by the time it is incorporated. That time lag is inherent to the H-Score® by design, and for many use cases it is perfectly acceptable.

However, insolvency is rarely a tidy annual event. The signals that precede financial distress often emerge continuously and incrementally: a winding-up petition filed by a creditor, an accumulation of County Court Judgments, a director with a track record of association with insolvent entities, persistent delays in statutory filing obligations. These are real predictive signals, and none of them are captured in a model that reads only from the balance sheet.

What does PoD® measure?

PoD® (Probability of Distress) produces a single output: the estimated probability that a given company will enter a state of financial distress within the next 12 months. Internally this is calculated as a value between 0 and 1, and surfaced in the UI as a percentage.

The choice to output a calibrated probability rather than a score band or risk grade is deliberate. A probability is semantically unambiguous: an internal value of 0.73, shown as 73% in the UI, means the model assigns a 73% likelihood of distress within the forecast horizon. That is directly interpretable and directly actionable in a way that a graded category is not.

The thinking behind PoD®

Defining the target variable

Precision on what constitutes “distress” matters a great deal when constructing a labelled training dataset, and we were careful here.

A company is classified as having entered distress if it subsequently enters administration, receivership, a Company Voluntary Arrangement (CVA), a Creditor Voluntary Liquidation, or a Compulsory Liquidation. These events share a common characteristic: they each indicate that a company has encountered genuine financial difficulty it cannot resolve through normal operations.

We exclude strike-offs, and Members’ Voluntary Liquidations from the distress definition. These events do not carry the same implication. A solvent company can choose to dissolve itself, and treating that as a distress event would introduce significant noise into the training labels. The model is trained specifically to identify financially-driven failure, and the target variable reflects that. 

Feature engineering: Using the full dataset

The feature set for PoD® extends considerably beyond what is available from financial accounts alone. The key additional data categories include:

Group structure and ownership. Parent, subsidiary, shareholder and PSC relationships materially change the interpretation of a set of accounts. A subsidiary can look adequately capitalised on a standalone basis while depending entirely on intra-group funding, and a sound trading company can be dragged down by a distressed parent. Resolving the group allows the model to treat these as connected rather than independent risks. Ownership changes are informative in their own right: an acquisition, a change of ultimate parent, or a rapid succession pattern in which activity is moved between related entities all carry signal.

Mortgages and charges. Registered charges are filed at Companies House and Company Watch’s charges information reflects those filings, updating as charge forms are lodged. The predictive content lies less in the existence of security than in its composition and timing: how many charges are outstanding rather than satisfied, how recently new security has been granted, and what kind of lender holds it. A new charge in favour of an invoice finance or asset-based lender, particularly where none existed previously, is often a sign that cheaper facilities are no longer available.

Winding-up petitions. A winding-up petition is a creditor’s formal application to the court to have a company compulsorily liquidated. The threshold for filing is meaningful: it represents a creditor who has exhausted other remedies. The presence and timing of petitions carries strong predictive weight.

County Court Judgments (CCJs). Both the volume of CCJs registered against a company and their trajectory over time are informative. An increasing rate of CCJ accumulation is a materially different signal from a stable historical level.

Creditor and debtor exposure. Two related datasets sit here. The first is creditor detail drawn from statements of affairs filed in insolvency proceedings, which breaks liabilities into six categories — secured, preferential, unsecured, associated company, international and other — and so reveals the ranking of debt as well as its size. This is what allows the observation that a company can appear to carry very little trade debt while being materially exposed through bank borrowing or intercompany loans. The second is distressed debtor exposure, identified by matching a company’s trade debtors against businesses in administration, in liquidation, or showing high distress indicators. This is a contagion signal: the failure of a significant customer can create real financial pressure for the supplier where the exposure is large, and it is visible before it appears anywhere in the supplier’s own accounts. 

Statutory filing behaviour. Consistent late filing of statutory obligations, particularly confirmation statements and accounts, is a documented precursor to financial distress. The mechanism is intuitive: companies under operational or financial strain tend to deprioritise compliance.

Director history. This is one of the more technically distinctive elements of the feature set. Company Watch’s Enhanced Directorships capability resolves director identities across company records to construct individual-level histories. This allows the model to incorporate, as a feature, whether a company’s directors have been previously associated with insolvent entities. The signal is strong, and it is not available to any model working from financial accounts alone.

All of these inputs are updated on a continuous basis, which means the model score reflects current conditions rather than the state of the business at the last filing date.

The modelling framework: XGBoost

PoD® is implemented using XGBoost, a gradient boosted decision tree algorithm that has demonstrated consistently strong performance on structured tabular data across a wide range of supervised learning applications.

The mechanics are worth explaining briefly. A single decision tree partitions the feature space through a sequence of binary splits, recursively dividing the data until it reaches leaf nodes that carry a prediction. The limitation of a single tree is its tendency to overfit: it can memorise the training data without generalising well to unseen observations.

Gradient boosting addresses this by building an ensemble of trees sequentially. Each successive tree is fit to the residuals of the current ensemble, meaning it is specifically targeted at correcting the errors that remain. The final prediction is the aggregated output of the full ensemble, which in practice yields substantially lower variance than any single tree and a much better fit to the underlying data-generating process.

XGBoost is an efficient and scalable implementation of this framework with strong regularisation capabilities. Given the volume of companies we score and the frequency at which scores are updated, computational efficiency is not a trivial consideration. XGBoost handles this well. 

Transparency via Shapley Decomposition

The interpretability of model outputs was a design constraint I took seriously throughout development. A model that produces a probability without any explanation of what is driving it is difficult for end users to interrogate, and in a risk context that is a meaningful limitation. Practitioners need to understand not just what the model is saying but why.

To address this, we compute Shapley values for each scored company at inference time. Shapley values, drawn from cooperative game theory, provide a principled method for attributing a model’s output to its input features. For each company, we decompose the predicted probability into the positive and negative contributions of each feature group, giving the user a clear view of which factors are elevating or suppressing the distress probability.

These contributions are surfaced at a category level rather than at the raw feature level. This serves two purposes: it makes the output more interpretable for non-technical users, and it provides a degree of protection for the proprietary data and feature engineering that underpins the model. The output is less granular than the coefficient-level decomposition available from a regression model like the H-Score®, but it gives a meaningful and honest steer on what is driving each individual prediction.

PoD® and the H-Score®: Complementary tools

The H-Score® provides a rigorous, regression-based financial risk assessment grounded in audited accounts data. For users who want a structured, fully decomposable score derived from statutory financial data, it remains a well-validated option.

PoD® takes a different approach. It incorporates a broader feature set updated continuously, applies a machine learning framework that can capture non-linear relationships and feature interactions that a linear model will miss, and outputs a calibrated probability rather than a score index. The incremental predictive power comes directly from the additional data: more signals, updated more frequently, processed through a more expressive model architecture.

The two scores are complementary. They are answering related but distinct questions, using different data and different methods, and for many users there is value in both.

The launch and what comes next

PoD® officially launches on 21st October 2026, following an extensive beta period that put the model in front of experienced practitioners across a range of use cases and industries.

The beta period had a straightforward objective: to gather structured feedback from real users on a model that had been developed and tested internally. No amount of internal validation fully replicates the conditions of production use, and there were always going to be usability questions, edge cases, and interpretability issues that would only surface once the model was being used by practitioners across a range of workflows.

The beta gave us the opportunity to identify and address those issues ahead of full release. As a result, the version of PoD® now available is the most robust and trustworthy model we can produce.

On a longer horizon, the model will be retrained annually, which creates a natural mechanism for incorporating additional data and refining the feature set over time. There is also scope to evaluate alternative modelling frameworks as the field continues to develop. For now, though, the focus is on supporting users as they integrate PoD® into their work and on learning from how they engage with it.

PoD® Beta Insights: What 64 Credit and Risk Professionals Found

If you already use Company Watch, PoD® sits alongside the H-Score® on the companies you are scoring today. The most useful first exercise is to run it across your existing ledger and see where the new probability rating adds additional value to our risk assessments.

Jonathan Davies
Jonathan Davies
Lead Data Scientist
Jonathan Davies is Lead Data Scientist at Company Watch and the architect of the PoD® model.