Skip to content

Some model review findings matter less in real terms

In a review we might identify a variable that could be acting as a proxy for a protected attribute. On face value, it’s worth raising, because the variable is in the model, the correlation is plausible, and someone needs to look at it.

Then we get into the data, and the field is populated for a small set of customers or carries little weight in the model. On its own this doesn't say much. What matters is whether these customers have different outcomes to others. If not, the proxy is there but with no real impact, or impact limited enough that it may not justify immediate effort. Effort that could be directed at more important things, in the short term.

For example, a model review might find vehicle type acting as a proxy for disability. The range of vehicle types includes vehicles with accessibility modifications, like modified driving controls. This type of issue could come up in a vehicle finance model review, or in a motor claims model review.

But when we look at the data, there's only a handful of loans or claims involved, and those are not treated unfavourably. The system won’t really pick up on them, even if vehicle type is fed in. There simply isn't enough of it for the machine learning models to learn from; the field is available to the models, but effectively unused.

This doesn't mean the original finding was wrong; we need to raise it. But until we know how often it happens and how much it moves the decision, we don't really know whether the problem is real or theoretical. Real needs action now, theoretical could wait for the next cycle with monitoring in the meantime.

This is especially important when models are not changed very often, and immediate action means an off-cycle change in the model. If the issue is theoretical, we could delay the change in the model until the next proper cycle, rather than trying to deal with it straight away and creating way more immediate effort than we really need to.

So when we find a potential problem in the models or rules of an algorithmic system, we can ask a couple of questions before deciding what action to take. How often does this actually happen, and are the affected customers getting different outcomes? If they aren't, do we need to change it now, or monitor and fix it in the next cycle? 

 


Disclaimer: The info in this article is not legal advice. It may not be relevant to your circumstances. It was written for specific contexts within banks and insurers, may not apply to other contexts, and may not be relevant to other types of organisations.