Data Intelligence Platform Corpus Schema Validation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Organizations face challenges in understanding and managing vast amounts of data generated from various sources, including issues related to data completeness, accuracy, consistency, and validity, which impairs their ability to utilize this data effectively and efficiently.
Innovation Solution
A data intelligence platform utilizing machine learning techniques and data feature models to perform data discovery, analyze data quality, and automate processes for identifying and addressing data issues, enabling contextualized analysis across industries and organizations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If traditional data management methods are used to store and access data across multiple data sources, then data storage capacity is maintained, but data understanding and management efficiency deteriorate due to issues with data completeness, accuracy, consistency, and validity
Solution Approach 1:
The patent introduces a corpus as an intermediary layer between raw data sources and data analysis processes. The corpus standardizes data from multiple sources by defining schemas, data types, and validation rules, thereby improving data quality while maintaining management efficiency. The corpus acts as a mediator that transforms heterogeneous data into a unified, quality-assured format.
Solution Approach 2:
The patent performs preliminary data processing and validation by mapping data to a corpus schema before analysis. Data completeness, accuracy, and consistency are verified in advance through schema validation and data quality checks, preventing poor-quality data from propagating through the system and improving overall data reliability.
2Measurement precision
If manual data discovery and analysis processes are used, then data quality can be assessed, but the time and resources required for data discovery increase significantly
Solution Approach 1:
The patent replaces manual, mechanical data discovery processes with automated machine learning models and algorithms. These models automatically assess data quality metrics such as completeness, accuracy, and consistency by comparing data against the corpus schema, eliminating the need for time-consuming manual inspection while maintaining high measurement precision.
Solution Approach 2:
The system enables self-service data quality assessment where the corpus and associated validation rules automatically evaluate data quality without human intervention. The data quality metrics are computed and reported automatically, allowing users to obtain precise data quality assessments without investing significant time in manual analysis.
3Loss of information
If comprehensive data analysis is performed across multiple data sources, then data understanding improves, but computing resources and processing time increase
Solution Approach 1:
The patent extracts and standardizes essential data characteristics by mapping data to a predefined corpus schema. This extraction process identifies only the relevant data elements and their relationships, filtering out redundant information. By focusing analysis on the extracted, schema-compliant data rather than raw data from all sources, the system improves data understanding while reducing computing resource requirements.
Solution Approach 2:
The patent segments data analysis into distinct stages: data extraction, corpus mapping, validation, and analysis. Each stage processes only the necessary data subset, with validation checks filtering out problematic data early. This segmentation prevents unnecessary processing of invalid or irrelevant data, reducing overall computing resource consumption while maintaining comprehensive data understanding.
Data Source
AI summary
A device may receive data stored in one or more data sources associated with an organization based on utilizing one or more data discovery-related application programming interfaces (APIs) to access the data. The device may process, utilizing one or more data feature models via the one or more data discovery-related APIs, the data received from the one or more data sources to identify types of data included in the data based on a contextualization of the data. The one or more data feature models may identify a respective set of attributes expected to be included in the types of data. The device may perform multiple analyses of the data after identifying the types of data. The device may determine, based on a result of the multiple analyses, a score for the data. The device may perform one or more actions based on a respective result of the multiple analyses.


