Persistent Member Identifier Generation for Variant Testing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Data from disparate sources with different identifier formats poses challenges in identification and aggregation, leading to difficulties in performing meaningful variant testing and data analysis, such as preventing leakage in AB tests and timely identifying health risks.
Innovation Solution
The system generates persistent member identifiers and employs identifier matching algorithms to cross-reference and aggregate data from various sources, using fuzzy mapping and logic mapping to match and tag data with member identifiers, allowing for automatic data aggregation and notification of relevant health interventions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If data from disparate sources with different identifier formats is aggregated, then data completeness and analysis capability are improved, but identification accuracy and data matching reliability deteriorate
Solution Approach 1:
The patent introduces standardized identifiers as an intermediary layer between disparate data sources. These standardized identifiers act as mediators that translate and harmonize different identifier formats from various sources, enabling accurate data matching without direct comparison of incompatible formats. This resolves the contradiction by maintaining identification accuracy while aggregating diverse data.
Solution Approach 2:
The system transforms identifier parameters from their original heterogeneous formats into a standardized format. By changing the parameter representation of identifiers across different data sources to a common standard, the system enables reliable data aggregation while preserving identification accuracy through consistent parameter structures.
2Productivity
If identifier matching algorithms are used to aggregate data, then data aggregation capability is improved, but system complexity increases
Solution Approach 1:
The patent segments the identifier matching process into distinct modular components: identifier extraction, format normalization, matching algorithm application, and result validation. This segmentation allows each component to be independently optimized and maintained, reducing overall system complexity while enhancing data aggregation capability through specialized processing stages.
3Stability of the object's composition
If persistent member identifiers are generated and used, then data consistency across sources is improved, but implementation complexity and processing time increase
Solution Approach 1:
The system performs preliminary actions by pre-generating and storing persistent member identifiers alongside source data during data ingestion. This advance preparation eliminates the need for complex real-time identifier resolution during data aggregation, maintaining data consistency while significantly reducing processing time through cached identifier mappings.
Data Source
AI summary
Systems and methods for identifying related data for variant testing are disclosed. For example, data stored for records from disparate data sources may not include the same identifiers for all records such that it may not be readily identified as record for the same member. The presently-disclosed systems and methods generate data tagged as identifier information and determine the degree of similarity between the identifier information. Based at least in part on the degree of similarity meeting or exceeding a threshold amount of similarity, the data may be associated with a member identifier. By properly identifying user information corresponding to member identifiers, the members may be split in meaningful ways to perform variant test.


