Relational Database Probabilistic Graphical Model for Data Quality

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Relational databases often suffer from accuracy and quality issues due to human error, equipment limitations, and data availability constraints, with users being unaware of errors and inaccuracies present in the data.

Innovation Solution

A method is introduced to automatically create a probabilistic graphical model based on the relational database schema, using inference algorithms to detect errors, fill missing data, and suggest corrections, allowing for improved data understanding and quality without requiring users to have statistical or machine learning expertise.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If users manually verify data accuracy in relational databases, then data quality can be improved, but time consumption and operational complexity increase significantly

Engineering Contradiction:
Improvedata accuracyVSAvoidtime consumption
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system enables data to verify itself through automated probabilistic graphical models and inference algorithms. The database automatically detects errors, infers missing values, and validates data consistency without requiring manual user intervention, thus improving data accuracy while eliminating time-consuming manual verification processes

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical verification processes with automated computational systems. Probabilistic graphical models and inference algorithms automatically perform data validation, error detection, and quality assessment, substituting human-operated mechanical verification with automated electronic processing that is both faster and more accurate

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If comprehensive data validation is performed on all database entries, then data quality improves, but system complexity and processing overhead increase

Engineering Contradiction:
Improvedata qualityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system changes the parameters of data validation by using probabilistic models that continuously assess data quality based on learned patterns and relationships. Instead of rigid deterministic validation rules, the system uses probabilistic thresholds and inference confidence levels that adapt to data characteristics, maintaining high data quality while reducing system complexity through flexible, context-aware validation

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces probabilistic graphical models as intermediary structures that mediate between raw data and validation rules. These models capture complex relationships and dependencies in the data, allowing the system to perform comprehensive validation indirectly through the model's inference mechanisms rather than through direct complex rule checking, thus reducing apparent system complexity

Inventive Principle:
Principle #24Intermediary (Mediator)

3Loss of information

If users without statistical expertise use advanced data analysis tools, then data understanding improves, but ease of operation decreases

Engineering Contradiction:
Improvedata understandingVSAvoidease of use
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The system creates simplified copies or representations of complex statistical concepts through user-friendly interfaces. Instead of requiring users to directly interact with complex probabilistic models and statistical algorithms, the system provides intuitive visualizations, automated interpretations, and simplified data exploration tools that replicate the power of advanced analysis without the complexity, making deep data understanding accessible to non-experts

Inventive Principle:
Principle #26Copying

Data Source

PatentUS10685062B2Relational database management
Publication Date: 2020.06.16 MICROSOFT TECHNOLOGY LICENSING LLC
  • US10685062B2 patent drawing
  • US10685062B2 patent drawing
  • US10685062B2 patent drawing

AI summary

New methods of relational database management are described, for example, to enable completion and checking of data in relational databases, including completion of missing foreign key values, to facilitate understanding of data in relational databases, to highlight data that it would be useful to add to a relational database and for other applications. In various embodiments, the schema of a relational database is used to automatically create a probabilistic graphical model that has a structure related to the schema. For example, nodes representing individual rows are linked to rows of other tables according to the database schema. In examples, data in the relational database is used to carry out inference using inference algorithms derived from the probabilistic graphical model. In various examples, inference results, comprising probability distributions each for an individual table cell, are used to fill missing data, highlight errors, and for other purposes.