Dataset Cleansing for Financial Forward Curves

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current financial instrument trading systems face inefficiencies due to incomplete and inaccurate data in forward curves, leading to faulty results and system slowdowns, particularly when dealing with missing or erroneous data for specific contract positions.

Innovation Solution

A dataset cleansing method that identifies and removes anomalies using historical patterns influenced by external factors like temporal, meteorological, and system factors, generating missing data elements through interpolation and extrapolation techniques to create complete and consistent datasets for forward curves.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If actual trades data is used to derive forward curves, then the data reflects real market conditions, but missing or faulty data occurs when no trades exist for certain contract positions

Engineering Contradiction:
Improvedata accuracyVSAvoidmissing data
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent uses statistical models and historical patterns as intermediary mechanisms to generate missing forward curve data. When actual trade data is unavailable for specific contract positions, the system employs regression analysis, time series modeling, and other statistical techniques to interpolate and extrapolate reasonable values based on available data from similar contracts and historical trends, thereby maintaining data completeness without compromising reliability

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system performs preliminary data validation and anomaly detection before forward curve derivation. By pre-identifying missing or faulty data points and preparing statistical supplementation methods in advance, the system prevents processing errors and ensures continuous operation even when trade data is incomplete

Inventive Principle:
Principle #10Preliminary action

2Productivity

If erroneous data from atypical trades or system errors is processed, then all available data is utilized, but faulty results and system slowdowns occur

Engineering Contradiction:
Improvesystem efficiencyVSAvoiddata quality
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent implements feedback mechanisms through statistical validation and anomaly detection systems. The system continuously monitors forward curve data for inconsistencies, compares derived values against historical patterns and statistical thresholds, and automatically identifies erroneous data points. This feedback loop enables the system to detect and correct data quality issues in real-time, preventing faulty results while maintaining processing efficiency

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system extracts and removes erroneous data points from the processing pipeline through statistical outlier detection and validation rules. By identifying and excluding atypical trades and system errors before they propagate through the forward curve derivation process, the system maintains high data quality without sacrificing overall productivity

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS10380690B2Dataset cleansing
Publication Date: 2019.08.13 CHICAGO MERCANTILE EXCHANGE INC
  • US10380690B2 patent drawing
  • US10380690B2 patent drawing
  • US10380690B2 patent drawing

AI summary

Datasets may be characterized by patterns. The patterns may be caused or otherwise influenced by external factors, such as temporal, meteorological, and/or system factors. The external factors, as well as the patterns which result in the data values of the dataset because of the external factors, may provide for techniques used to account for missing data elements, outlier data elements and/or otherwise cleanse the dataset. New elements may be generated to provide for the missing data elements, and derivative datasets may be generated based on one or more cleansed datasets.