Canonical Data Normalization for Heterogeneous Real-World Evidence
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies struggle to handle disparities in real-world data from disparate or heterogeneous sources, making it difficult to combine and analyze large datasets for reliable results in healthcare applications.
Innovation Solution
A computer-based real-world evidence (CRWE) solution that builds a relational database by converting raw real-world data into canonical versions, allowing for the reliable analysis and correlation of data from diverse sources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If data from disparate or heterogeneous sources is combined, then the quantity and diversity of data increases, but the reliability and consistency of results deteriorate due to data disparities
Solution Approach 1:
The system segments data processing into multiple stages: raw data collection from heterogeneous sources, canonicalization conversion to standardized formats, validation against predefined rules, and finally analysis. This segmentation allows each stage to handle specific aspects of data processing, ensuring that data quantity can be increased while maintaining reliability through systematic processing and validation at each segment.
Solution Approach 2:
The patent introduces a canonicalization layer as an intermediary between raw heterogeneous data and analysis processes. This intermediary converts diverse data formats into standardized canonical representations, enabling data from disparate sources to be combined without compromising reliability. The canonicalization process acts as a mediator that resolves format conflicts and ensures consistency.
2Reliability
If data validation and verification processes are implemented, then the reliability of results improves, but the processing time and complexity increase
Solution Approach 1:
The system performs validation and verification actions preliminarily during the canonicalization stage, before main analysis processes begin. By establishing data validity early through canonicalization and initial validation rules, the system avoids repeated validation cycles during analysis, thereby reducing overall processing time while maintaining high data verifiability.
Solution Approach 2:
The patent changes the state of data parameters during canonicalization, transforming raw data into standardized formats with embedded validation metadata. This parameter transformation includes converting diverse data types into canonical representations that inherently encode validation information, reducing the need for time-consuming validation checks during subsequent analysis.
3Stability of the object's composition
If canonicalization processes are applied to convert raw data, then data consistency improves, but the complexity of the processing system increases
Solution Approach 1:
The canonicalization process is designed as a universal, multi-functional component that handles various data types (structured, semi-structured, unstructured) through a single standardized interface. This universal approach consolidates multiple data processing functions into one system, achieving data consistency across heterogeneous sources without proportionally increasing overall system complexity.
Solution Approach 2:
The system applies homogeneity by converting all incoming diverse data into a unified canonical format with consistent structure, validation rules, and metadata standards. This homogeneous representation layer simplifies downstream processing by eliminating the need for multiple specialized handlers for different data types, thereby achieving data consistency while controlling system complexity.
Data Source
AI summary
A computer-based real-world evidence (CRWE) solution that can handle disparities in real-world data (RWD) that might otherwise not be combinable or that might originate from disparate or heterogeneous sources. The CRWE solution is designed such that it can build up RWD data from the ground up (for example, from the atomic level) into a canonical relational database by converting or linking the raw RWD data to its canonical versions such that reliable canonical answers can be mined from large datasets and consistently provided in response to queries to the solution for an answer. The CRWE solution is designed to gather and analyze large amounts of RWD data from heterogeneous, multi-national, and unverifiable data sources, and provide canonical results that are constituently reliable and can expose, for example, clinically-significant correlations between medical products or treatments and outcomes.


