Cross-Provider Topic Conflation for Knowledge Graph Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Knowledge graphs often contain duplicate, conflicting, or erroneous data due to variations in data mining processes among content providers, leading to user confusion and misinformation when surfacing data from multiple sources.

Innovation Solution

Implement a system for cross-provider topic conflation that extracts document content from multiple providers, separates and clusters it into subparts, and merges unique information under a single topic, preserving source identification to prevent data duplication and ensure accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If data is mined from multiple content providers, then the quantity of data increases, but data accuracy and reliability deteriorate due to duplicates and conflicts

Engineering Contradiction:
Improvedata quantityVSAvoiddata accuracy
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent segments data from multiple content providers into discrete entities and topics, allowing individual processing and comparison. Each data element is broken down into可比 units that can be evaluated for duplication and conflict independently, enabling systematic resolution while maintaining overall data volume.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system changes parameters by introducing confidence scores, source identifiers, and entity types as new dimensions for data evaluation. These parameter transformations enable the system to weigh and prioritize data from different providers, resolving conflicts through quantitative comparison rather than simple duplication removal.

Inventive Principle:
Principle #35Parameter changes

2Loss of information

If data from multiple providers is aggregated, then information completeness improves, but data processing complexity increases

Engineering Contradiction:
Improveinformation completenessVSAvoidprocessing complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary processing layer that standardizes data from multiple providers before integration. This intermediary system applies consistent entity recognition, topic modeling, and conflict resolution rules, reducing the complexity burden on the overall system while maintaining comprehensive information aggregation.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system performs preliminary actions by pre-processing data from content providers through entity extraction, topic assignment, and initial deduplication before main aggregation. This preliminary filtering reduces the complexity of subsequent processing steps while ensuring no critical information is lost during integration.

Inventive Principle:
Principle #10Preliminary action

3Loss of substance

If duplicate data is removed through conflation, then data storage efficiency improves, but information loss may occur if unique details are merged

Engineering Contradiction:
Improvestorage efficiencyVSAvoidunique information
Core Design Contradiction:
Loss of substanceVSLoss of information

Solution Approach 1:

The patent applies local quality by treating different parts of conflated data differently based on their source and characteristics. Rather than uniform merging, the system preserves unique attributes from each provider while consolidating common information, ensuring that locally unique details are maintained even when global duplication is removed.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system implements feedback mechanisms that monitor the conflation process for potential information loss. By tracking source attribution and entity relationships, the system can identify and preserve unique information that would otherwise be lost during deduplication, while still achieving storage efficiency through intelligent merging.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12197421B2Cross-provider topic conflation
Publication Date: 2025.01.14 MICROSOFT TECHNOLOGY LICENSING LLC
  • US12197421B2 patent drawing
  • US12197421B2 patent drawing
  • US12197421B2 patent drawing

AI summary

Examples of the present disclosure describe systems and methods for cross-provider topic conflation. In aspects, a request relating to one or more topics may be received by a content surfacing platform. One or more data sources of multiple content providers may be searched for documents relating to the topic(s). Document content (e.g., document metadata and sentences, phrases, and other word content within the document) relating to the topic(s) may be extracted from the documents of the various content providers. The document content may be classified and/or separated into subparts. The subparts may be clustered and/or conflated by topic, thereby removing duplicated data while preserving the unique information in each subpart. The conflated topics may be stored in a single knowledge base, such as an enterprise knowledge graph, and/or presented in response to the request.