Synthetic Data Risk Management Method for Sparse Datasets

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current risk management methods face challenges with sparse data problems, leading to inefficient database construction, high computational costs, and inconsistencies in Monte Carlo and Bootstrap simulations, which hinder accurate risk assessment and decision-making.

Innovation Solution

A risk management method utilizing synthetic data that replicates experimental data with 99% statistical confidence, allowing for flexible construction of risk scenarios and estimation of consistent risk sensitivities, avoiding asymptotic bias by modifying sample size and econometric parameters, and integrating sampling design, operations research, and decision theory.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If Monte Carlo and Bootstrap simulations are used to increase sample size, then the number of observations is increased, but computational time and processing power requirements increase significantly

Engineering Contradiction:
Improvesample sizeVSAvoidcomputational time
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

The patent creates synthetic copies of observed data through simulation models that replicate the statistical properties and relationships of the original dataset. These synthetic datasets can be generated rapidly without requiring extensive computational resources, thus providing large sample sizes without the proportional increase in computational time associated with traditional Monte Carlo methods

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent performs preliminary analysis of the observed data to identify key statistical properties, relationships, and patterns before generating synthetic data. By pre-characterizing the data structure and using this information to guide synthetic data generation, the method avoids the need for repeated complex simulations, thereby reducing computational time while maintaining sample size

Inventive Principle:
Principle #10Preliminary action

2Reliability

If traditional sampling methods are used to ensure representativeness, then data quality is improved, but costs and time requirements increase

Engineering Contradiction:
Improvedata representativenessVSAvoidsampling design complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent creates synthetic copies of the observed dataset that preserve the statistical properties, relationships, and variability patterns of the original data. These synthetic datasets serve as representative surrogates that can be used for analysis without requiring complex sampling designs or additional fieldwork, thereby maintaining data reliability while reducing operational complexity

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent transforms the original observed data into synthetic data by applying parameter-based transformations that preserve key statistical properties. By changing the representation format from raw observed values to synthetically generated values with matched distributions and relationships, the method maintains representativeness while simplifying the sampling process

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If asymptotic estimators are used with large sample sizes, then estimator consistency is improved, but bias problems arise due to interactions between individual-effects and time frequencies

Engineering Contradiction:
Improveestimator consistencyVSAvoidasymptotic bias
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent generates synthetic data with controlled parameter settings that allow researchers to study estimator performance under specific conditions. By systematically varying parameters such as sample size, time frequency, and individual effects while maintaining consistent data generation processes, the method enables precise measurement of estimator properties without the asymptotic bias that arises in traditional large-sample approaches

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent creates synthetic replicas of the data generation process that preserve the underlying relationships and structures. These synthetic datasets allow for repeated estimation under identical conditions, enabling researchers to assess estimator consistency and bias without the confounding effects that arise when simply increasing sample size in observational studies

Inventive Principle:
Principle #26Copying

Data Source

PatentUS10296983B2How to model risk on your farm
Publication Date: 2019.05.21 CARBAJAL DE NOVA CAROLINA
  • US10296983B2 patent drawing
  • US10296983B2 patent drawing
  • US10296983B2 patent drawing

AI summary

This is an invention implying a risk management method based on synthetic data. Synthetic data are random variables that take their mean and standard deviation from experimental data. It reproduces the experimental data underlying statistic information with a statistical confidence of 99%. Experimental data could represent financial variables i.e., prices and quantities. This method implements alternative recursive bias approach that achieve convergence with related estimations already published. Also, this method is decomposable allowing synthetic data and econometric models parameter values and different risk scenarios alternatives in order to measure the system responses under different controls.