Data Validation Using Encode Values for Integrity

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data-monitoring systems are unable to validate the integrity of large datasets, such as those with billions of records, as they lack the capability to detect inherent abnormalities in semantic content, making it infeasible to manually inspect each record for data integrity.

Innovation Solution

A data monitoring system that uses encode values generated from previous datasets to validate new or updated datasets through machine learning models, specifically autoencoders, to detect anomalies and ensure data integrity before publication.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual inspection of each data record is performed to ensure data integrity, then measurement precision is improved, but productivity deteriorates

Engineering Contradiction:
Improvedata integrity validationVSAvoiddata processing speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent introduces encode values as an intermediary representation of dataset characteristics. Instead of manually inspecting individual data records, the system generates encode values that capture essential dataset properties and uses these intermediaries to validate data integrity, thereby maintaining measurement precision while avoiding the productivity loss of manual inspection

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces the mechanical process of manual data record inspection with an automated computational system. Machine learning models process the encode values to detect anomalies and validate data integrity, substituting human manual verification with automated algorithms that achieve both high precision and scalability

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If machine learning models are used to validate datasets, then productivity is improved, but device complexity worsens

Engineering Contradiction:
Improvedata validation speedVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent extracts essential dataset characteristics into compact encode values, separating the complex data validation task from the raw data itself. This extraction allows machine learning models to work with simplified representations, improving productivity while managing system complexity by focusing computational resources on essential features rather than entire datasets

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS11797565B2Data validation using encode values
Publication Date: 2023.10.24 PAYPAL INC
  • US11797565B2 patent drawing
  • US11797565B2 patent drawing
  • US11797565B2 patent drawing

AI summary

Techniques are disclosed relating to data validation using encode values. In various embodiments, a data monitoring system may retrieve a plurality of datasets from a live database at a non-production datacenter. The data monitoring system may perform encoding operations on one or more of the plurality of datasets to generate encode values that correspond to the plurality of datasets. The data monitoring system may then retrieve an updated dataset, for example from an experimental database at the non-production datacenter, and perform validation operations to validate one or more characteristics of the updated dataset. For example, in some embodiments, the data monitoring system may retrieve the encode values corresponding to the plurality of datasets and use the encode values to validate the updated dataset. The data monitoring system may then generate a validation output indicative of a result of the validation operations.