About us

We turn complex datainto clear decisions.

ARST Data started as a used-vehicle valuation platform for 13 European markets. Every night we process close to 3M second-hand listings, of which ~1.4M are live at any time, and turn them, with our own multi-level XGBoost engine (L2/L3/L4 depending on how much data exists per model), into appraisals calibrated against real sales. No hype: every number we publish can be audited.

13

Active EU markets

~3M/day

Listings processed

60M+

Historical observations

Mission

Institutional-grade intelligence,accessible to any team.

Historically, predictive models and institutional-grade data have been reserved for large corporations with in-house data science teams. We want to change that.

We focus on specific sectors we know well (electric mobility, insurance, and macro across Europe and Latin America) and build tools that work right away, with no extra infrastructure.

01

Vertical specialization

We are not a generic vendor. Every model and every dataset is built for a specific use case in sectors we know well: automotive, insurance, and macro across Europe and Latin America.

02

Quantitative rigor

Every model goes through cross-validation, backtesting, and source auditing. We do not publish signals we cannot defend with numbers.

03

Up-to-date data

Automated pipelines that continuously aggregate and normalize data. A signal that arrives late has already lost its value.

04

No conflicts of interest

We are independent. No product distributors, no hidden incentives. Our only goal is for you to make better decisions.

Methodology

How we buildour models and datasets.

From data ingestion to the final prediction, every step is designed to guarantee statistical rigor and reproducibility.

01

Primary and secondary sources

We combine our own data, gathered from primary market sources, with macroeconomic feeds from public bodies: the ECB, INE, SUSEP and Banco de España.

02

Normalization and cleaning

Deduplication, outlier detection, and normalization of model names, trims, and equipment. A machine learning model is only as good as the data that feeds it.

03

Proprietary multi-level engine

Gradient boosting (XGBoost) organized in three levels — L2 by segment, L3 by make and model, L4 by exact trim — depending on how much real data exists. Conformal prediction calibrates the confidence interval on every appraisal.

04

Validation and backtesting

Strict temporal holdout. Models are evaluated on data they have never seen, to ensure generalization is real and not the product of overfitting.

Let’s work together

Got ause case?

Tell us about your specific problem. If we can help with data, models, or analysis, we’ll tell you straight away. We respond within 24 hours.