Quality Data is better than breadth of data

Pebbles FS
June 30, 2026


There is a common misconception that more data is always better. More data can be better, but what matters is depth, not breadth. For example, it is better to have a predictive model running over 100K policies than 10K (depth). That is not the same as saying it is better to have 100 variables than 10 variables (breadth). Exploring endless variables rapidly becomes counter-productive.

There is a relationship between the depth and the breadth of data, and it is the depth of data that is key. As the breadth of data being used increases, then the depth needs to increase exponentially, a constraint many practitioners overlook.

A modern algorithm can sift through a sprawling number of variables, but relative to actual data depth, it will lead to overfitting. The effect is magnified when variables contain missing or erroneous values. Gradient Boosting Machines (GBMs) excel at identifying discrete risk groups and can technically cope with incomplete variables or poor depth to breadth ratios, but they overfit and create misleading outcomes that drive adverse selection.

In the words of George Box (Professor Emeritus of Statistics, University of Wisconsin-Madison) “All models are wrong, but some are useful.” The pricing team’s role is to keep models useful by using relevant and complete data. They need to explain what is driving outcomes, rather than just throwing data at the ‘wall’ to see what sticks and not being able to communicate outcomes.

There is also a hidden operational cost to storing, processing, cleaning and transforming raw data, which ties up valuable human and system resources. Some clients have full teams that are tasked with maintaining vehicle files.

In our view, data used by analytical teams should not require remediation. When purchasing external data, it should be usable “out of the box”.

Pebbles data enrichment products are engineered from the ground up to bridge the gap between datasets and modern machine learning (ML) models.

Built by insurance practitioners who understand live pricing workflows, our postcode and vehicle datasets arrive complete, model-ready, and designed to eliminate statistical noise.

Elevate pricing models

At the core of insurance is risk selection, which is fundamental to building a sustainable insurance business. Risk selection is the execution of risk appetite. In simple terms, risk selection defines the business you want to write, while pricing and product determine your ability to win that business.

Contrary to popular opinion, there is not a price for every risk. Attempting to price every risk will result in a negative underwriting result. To build a pricing model that drives profits, it is important to separate two concepts:

Relying on pricing models for segments outside your core risk appetite, or on the fringes where you are trading with poor conversion will, without exception, lead to significant losses. This is especially true when pricing models rely on data that only describes the segments you currently write, or is a limited sample size relative to your competitors. This is where quality external data helps.

Pebbles datasets is granular, complete, and with regularly updated peril risk scores and a market premium scores across the UK population. It supports profitable growth by:

Optimise footprint expansion with strategic risk selection

Managing loss ratio performance is as much about the risks you decline or refer as the business you write. In high-volume distribution channels, automated underwriting decisions are essential, and referrals rarely exist. In other channels, referrals remain a critical part of risk selection.

Pebbles market and risk scores act as a reliable filter at the point of quote. Our ground-up machine learning-derived scores help underwriters implement sophisticated risk selection rules:

Article content
Pebbles Theft Risk Score distributed by postcode


The result is better loss ratios. By having risk selection rules working in tandem with pricing models, you can materially improve performance. Not every risk has a price. Every risk within your risk appetite should have a price, provided you can recognise where pricing models are extrapolating and tighten your risk selection appropriately.

In summary

Feed pricing models with granular, engineered, complete features and you will see an immediate uplift in pricing performance, while fine-tuning risk selection rules for profitable growth.

Why Pebbles data

Building internal teams of pricing and data scientists to manually clean data, map public accident statistics and reverse-engineer market prices takes time, experience and resources every year. Large insurers have even spun up dedicated data enrichment teams. Pebbles have already done this work. For a fraction of the cost, you can access clean, ready-to-use data that is regularly maintained and continuously enhanced. Enabling teams to focus on using data to driver better outcomes.

The next step

Contact us to organise a trial period to assess Pebbles data, and be live within weeks.

Back to Articles