An index rulebook may describe a universe, filters, rankings, weighting rules, constraints and rebalancing schedules. Yet every one of these elements depends on data definitions and conventions that can materially affect the outcome. Two strategies may both be described as "European quality indices" or "AI thematic baskets" and still behave very differently, not because their formulas differ, but because the underlying data do.
The methodology begins before the formula
Many methodological decisions are embedded in the data itself. Consider a value strategy based on free-cash-flow yield or earning-to-price ratios. The result depends on accounting definitions, fiscal periods, treatment of negative values, currency conversions and the timing of reported information.
The same applies to thematic investing. Which companies are considered exposed to a theme? Is inclusion based on revenues, business descriptions, proprietary classifications or external datasets? How are diversified firms treated, and what level of exposure is considered meaningful?
These questions extend to ESG metrics, carbon data, liquidity measures, analyst forecasts and alternative datasets. They also determine the economic meaning of a strategy: liquidity data influences investability, market-capitalization and free-float data define the opportunity set, while classifications and fundamental datasets shape sector, thematic and factor exposures.
Data definitions are not implementation details beneath the methodology. They are among its most important assumptions.
Accurate point-in-time data is fundamental
The second dimension is time. A strategy should be evaluated using only the information that was actually available when each historical decision was made.
Using today's company classifications in a historical simulation can distort a thematic strategy. Using revised fundamentals instead of the figures available at the rebalance date can artificially improve a factor backtest. These are examples of look-ahead bias: the strategy unknowingly benefits from future information. It is a common pitfall and often more subtle than it first appears.
This is why point-in-time data is essential. For an index-grade methodology, it is not enough to know the value eventually associated with a company. One must know what information the methodology was allowed to observe at each selection or rebalancing date. Without this discipline, historical results can provide a misleading picture of how the strategy actually behaved.
Proprietary data is becoming a strategic asset
The rise of self-service indexing makes the data issue even more important. Asset managers, asset owners and index providers increasingly want to build methodologies using their own information, whether internal scores, proprietary classifications, exclusion lists or thematic research.
Historically, integrating such datasets into a benchmark often required a dedicated development project. In a self-service environment, proprietary data should instead become a native component of the methodology, usable in universe construction, selection, weighting, constraints and analysis.
This changes the nature of customization. Institutions are no longer limited to adjusting parameters within a vendor's predefined framework. They can transform their own research and intellectual property into systematic and operational investment methodologies.
Reproducibility requires a data context
Once data becomes part of the methodology, reproducibility requires more than preserving a list of index rules.
To reproduce a simulation or a live calculation, it is also necessary to preserve the relevant data context, including dataset versions, classifications, mappings and processing conventions. This becomes particularly important when methodologies evolve, datasets are corrected or historical results need to be reviewed months later.
An institutional index platform should therefore treat the data environment itself as a traceable and versioned object linked to every simulation and calculation.
What this means in Folio
This principle is central to the architecture of MCFT Folio. The platform is built around a unified data abstraction capable of combining and versioning market data, fundamentals, classifications, corporate actions, derived indicators and user-provided datasets within a single framework.
Client datasets can be uploaded in point-in-time form and used directly throughout the methodology pipeline. Internal scores, classifications, inclusion lists and proprietary indicators can feed filters, weighting rules, constraints and analytical workflows.
Behind these workflows, Folio relies on versioned DataSessions, structured snapshots of the financial-data universe designed for both high-speed simulation and reproducibility. Each simulation is linked to an identifiable data context, while the daily processing pipeline manages ingestion, normalization, quality controls and operational updates. The objective is simple: methodologies should not only be configurable, but also supported by data that is integrated, inspectable and reproducible.
The deeper implication
In the self-service indexing era, institutions will not compete solely through better formulas. They will also compete through better definitions, better classifications, stronger proprietary datasets and greater control over how these inputs become systematic investment methodologies.
The future index-technology stack is therefore more than a calculation engine. It is a framework through which data, investment beliefs and methodological rules are transformed into transparent and operational strategies. Because in systematic investing, the methodology is never independent from the data.
The MCFT Founders - Romain Charlassier, Jonathan Klein and Lucas Mouilleron

