Semantic Inference for Disparate Data Integration

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current cloud services lack the ability to effectively provide information as a service across platforms, as disparate content providers do not coordinate their data publishing, leading to incompatible data sets that hinder data integration and utilization, requiring human intervention for semantic understanding and verification.

Innovation Solution

The system infers and updates semantic information about data sets through query analysis, maintaining mappings and adapting APIs to provide self-descriptive access, allowing for automatic joining and filtering of data sets without altering the underlying data, using weighted algorithms and probabilistic methods to determine column meanings and relationships.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If human intervention is used to verify and determine column relationships, then measurement precision of data semantics is improved, but device complexity and loss of time increase

Engineering Contradiction:
Improvesemantic understanding accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs self-verification of semantic mappings by automatically analyzing query patterns and data relationships without requiring manual human intervention. The automated semantic analyzer infers column relationships and verifies mappings by observing actual data queries and results, enabling the system to self-correct and improve its understanding over time.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system implements feedback mechanisms where query results and user interactions are analyzed to refine and verify semantic mappings. The automated analyzer continuously learns from actual usage patterns, adjusting its interpretations of column relationships to improve accuracy while reducing the need for manual verification.

Inventive Principle:
Principle #23Feedback

2Measurement precision

If human intervention is used to determine column relationships, then measurement precision of data semantics is improved, but loss of time increases

Engineering Contradiction:
Improvesemantic understanding accuracyVSAvoidtime for semantic verification
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system automatically performs semantic verification and column relationship determination without requiring manual human intervention. The automated semantic analyzer continuously processes data queries and infers relationships, eliminating the time-consuming manual verification process while maintaining high accuracy through algorithmic analysis.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs continuous automated semantic analysis alongside normal data operations, so that verification occurs in parallel rather than sequentially. The automated analyzer processes semantic mappings continuously as data is queried and updated, eliminating downtime and reducing total verification time while maintaining precision.

Inventive Principle:
Principle #20Continuity of useful action

3Adaptability or versatility

If data sets are published with proprietary schemas, then adaptability of data formats is improved, but ease of operation for data integration deteriorates

Engineering Contradiction:
Improvedata format flexibilityVSAvoiddata integration ease
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The system introduces a standardized semantic layer that acts as an intermediary between disparate data schemas and the user interface. This semantic abstraction layer translates proprietary data formats into unified conceptual models, allowing users to query and integrate data from multiple sources using consistent operations without needing to understand the underlying schema differences.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system creates a universal query interface that works across different data schemas and providers. The automated semantic analyzer enables the same query operations to function on diverse data formats, making the system multi-functional and schema-agnostic while preserving the adaptability of proprietary data structures.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Ease of operation

If automated semantic inference is implemented, then ease of operation for data integration is improved, but measurement precision of semantic understanding may worsen

Engineering Contradiction:
Improvedata integration easeVSAvoidsemantic understanding accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The system uses feedback from actual query patterns and user interactions to continuously refine automated semantic inferences. By observing how data is actually queried and used, the automated analyzer corrects and improves its interpretations, ensuring high precision while maintaining ease of operation through automation.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs preliminary automated semantic analysis and mapping before data integration operations are executed. This preliminary action establishes robust semantic understanding in advance, allowing subsequent operations to proceed easily without requiring manual verification, while the automated analyzer continues to refine its accuracy through ongoing learning.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS8635173B2Semantics update and adaptive interfaces in connection with information as a service
Publication Date: 2014.01.21 MICROSOFT TECHNOLOGY LICENSING LLC
  • US8635173B2 patent drawing
  • US8635173B2 patent drawing
  • US8635173B2 patent drawing

AI summary

Additional semantic information that describes data sets is inferred in response to a request for data from the data sets, e.g., in response to a query over the data sets, including analyzing a subset of results extracted based on the request for data to determine the additional semantic information. The additional semantic information can be verified by the publisher as correct, or satisfy correctness probabilistically. Mapping information based on the additional semantic information can be maintained and updated as the system learns additional semantic information (e.g., information about what a given column represents and data types represented), and the form of future data requests (e.g., URL based queries) can be updated to more closely correspond to the updated additional semantic information.