Predictive Data Clustering Using Merged Population and Location Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing predictive data analysis solutions face efficiency and reliability challenges in performing predictive operations on candidate entities.
Innovation Solution
A computer-implemented method that merges population data with ancillary data from multiple datasets, generates location data, associates it with external domain data, determines distances between entities, and generates clusters to train a prediction machine learning model for ranking candidate entities.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If traditional predictive data analysis solutions are used, then predictive operations can be performed, but efficiency and reliability are insufficient
Solution Approach 1:
The patent merges population data with ancillary data from multiple datasets (first ancillary dataset and second ancillary dataset) to create a comprehensive data structure. This combination of data sources improves both efficiency and reliability by leveraging diverse data types for more accurate predictions while reducing computational operations through integrated data processing.
Solution Approach 2:
The patent segments the data processing into distinct components: population data extraction, ancillary data integration, location data generation, and clustering. This segmentation allows each component to be optimized independently, improving overall efficiency while maintaining reliability through structured processing of complex predictive operations.
2Reliability
If more data is used to improve prediction accuracy, then reliability improves, but computational operations and training data entries increase
Solution Approach 1:
By merging population data with ancillary data from multiple datasets into a unified structure, the patent achieves improved prediction accuracy without proportionally increasing computational operations. The integrated data structure allows the system to process and analyze data more efficiently, reducing the number of separate computational steps required.
Solution Approach 2:
The patent performs preliminary data processing steps including extraction of population data, integration of ancillary data, and generation of location data before the main predictive modeling process. This preliminary action prepares the data in advance, reducing the computational burden during actual prediction operations while maintaining high accuracy.
Data Source
AI summary
Various embodiments of the present disclosure provide methods, apparatus, systems, computing devices, computing entities, and/or the like for a predictive data analysis system that is configured to rank one or more candidate entities. A machine learning model is trained to rank the one or more candidate entities for initiating the performance of one or more prediction-based actions based on one or more sets of a plurality of clusters generated based on population data merged with ancillary data, and an association of location data with external domain data. The plurality of clusters is generated by generating embeddings for one or more features associated with a plurality of entities selected for clustering and determining a similarity score for entity pairs selected from the plurality of entities based on a distance function and the embeddings.


