Predictive Data Clustering Using Merged Population and Location Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing predictive data analysis solutions face efficiency and reliability challenges in performing predictive operations on candidate entities.

Innovation Solution

A computer-implemented method that merges population data with ancillary data from multiple datasets, generates location data, associates it with external domain data, determines distances between entities, and generates clusters to train a prediction machine learning model for ranking candidate entities.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional predictive data analysis solutions are used, then predictive operations can be performed, but efficiency and reliability are insufficient

Engineering Contradiction:
Improveefficiency of predictive operationsVSAvoidreliability of predictive operations
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent merges population data with ancillary data from multiple datasets (first ancillary dataset and second ancillary dataset) to create a comprehensive data structure. This combination of data sources improves both efficiency and reliability by leveraging diverse data types for more accurate predictions while reducing computational operations through integrated data processing.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent segments the data processing into distinct components: population data extraction, ancillary data integration, location data generation, and clustering. This segmentation allows each component to be optimized independently, improving overall efficiency while maintaining reliability through structured processing of complex predictive operations.

Inventive Principle:
Principle #1Segmentation

2Reliability

If more data is used to improve prediction accuracy, then reliability improves, but computational operations and training data entries increase

Engineering Contradiction:
Improveaccuracy of predictive modelVSAvoidnumber of computational operations
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

By merging population data with ancillary data from multiple datasets into a unified structure, the patent achieves improved prediction accuracy without proportionally increasing computational operations. The integrated data structure allows the system to process and analyze data more efficiently, reducing the number of separate computational steps required.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent performs preliminary data processing steps including extraction of population data, integration of ancillary data, and generation of location data before the main predictive modeling process. This preliminary action prepares the data in advance, reducing the computational burden during actual prediction operations while maintaining high accuracy.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250028980A1Machine learning techniques for feature prediction based on clustering using ancillary and location data
Publication Date: 2025.01.23 UNITEDHEALTH GROUP INC
  • US20250028980A1 patent drawing
  • US20250028980A1 patent drawing
  • US20250028980A1 patent drawing

AI summary

Various embodiments of the present disclosure provide methods, apparatus, systems, computing devices, computing entities, and/or the like for a predictive data analysis system that is configured to rank one or more candidate entities. A machine learning model is trained to rank the one or more candidate entities for initiating the performance of one or more prediction-based actions based on one or more sets of a plurality of clusters generated based on population data merged with ancillary data, and an association of location data with external domain data. The plurality of clusters is generated by generating embeddings for one or more features associated with a plurality of entities selected for clustering and determining a similarity score for entity pairs selected from the plurality of entities based on a distance function and the embeddings.