Structured Data Summary Objects for Machine Learning Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine learning systems face challenges in performing low-latency interactive computations on large datasets, requiring efficient methods to analyze and generate inferences quickly to support decision-making in real-time applications.

Innovation Solution

A computing system that generates structured datasets by performing operations on a first dataset to produce multiple second datasets, creating summary objects that characterize differences between groups, and using these objects to inform a machine learning system for analyzing operations and generating data analysis models indicating expected outcomes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional machine learning systems process large datasets directly, then analysis accuracy is maintained, but processing time increases and latency is high

Engineering Contradiction:
Improveanalysis accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments large datasets into multiple structured datasets with consistent data structures. Each structured dataset is further divided into groups that can be processed independently. This segmentation allows the machine learning system to process smaller, manageable units while maintaining overall analysis accuracy, thereby reducing processing time and latency.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extracts essential characteristics and differences from groups of data elements to create summary objects. These summary objects capture the essential information needed for analysis while being significantly smaller than the original datasets. This extraction process reduces the computational burden on the machine learning system while preserving the key information needed for accurate analysis.

Inventive Principle:
Principle #2Taking out (Extraction)

2Measurement precision

If complex operations are performed on large datasets, then detailed analysis is achieved, but processor utilization increases and computational efficiency decreases

Engineering Contradiction:
Improveanalysis detailVSAvoidcomputational efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent extracts only the essential differences and characteristics from data groups into summary objects, rather than processing entire datasets. This extraction maintains the analytical detail needed for accurate insights while dramatically reducing the computational workload on processors, thereby improving computational efficiency.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent creates simplified copies of data groups in the form of summary objects that have consistent data structures. These summary objects serve as efficient representations that can be processed quickly by the machine learning system while preserving the essential information needed for detailed analysis.

Inventive Principle:
Principle #26Copying

3Ease of operation

If summary objects with consistent data structures are generated, then machine learning processing is simplified, but additional preprocessing steps are required

Engineering Contradiction:
Improveprocessing simplicityVSAvoidpreprocessing complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent segments datasets into structured groups with consistent data structures, which simplifies the subsequent machine learning processing. While this segmentation requires preprocessing steps, the structured organization makes the overall system easier to operate and maintain, as the consistent structure allows the machine learning model to process data more efficiently.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transforms raw data into summary objects with consistent data structures by changing parameters such as data organization, grouping criteria, and representation format. This parameter transformation, while requiring preprocessing, ultimately simplifies the machine learning operation by providing uniformly structured input data that the model can process efficiently.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS10846616B1System and method for enhanced characterization of structured data for machine learning
Publication Date: 2020.11.24 IQVIA INC
  • US10846616B1 patent drawing
  • US10846616B1 patent drawing
  • US10846616B1 patent drawing

AI summary

A computer-implemented method includes a computing system having a database that stores multiple datasets and that accesses the database to perform operations on a first dataset to produce multiple second datasets. The system determines a relationship between the first dataset and each second dataset of the multiple second datasets. The system also determines a relationship between respective groups of the first dataset and determines a relationship between respective groups of each second dataset. The system generates summary objects based, in part, on the determined relationships between respective groups of the first and second datasets. The system includes a machine learning system that uses the respective summary objects to analyze the performed operations that produced the multiple second datasets. Based on the analyzed performed operations, the machine learning system generates a data analysis model that indicates sequences of operations for achieving particular desired data analysis outcomes.