Topic-Specific Knowledge Graphs for Underrepresented Data Balancing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing systems face challenges in accurately identifying topics associated with data inputs due to underrepresented data sets, leading to decreased probability of topic identification and inefficient processing resources.

Innovation Solution

A method utilizing a domain knowledge graph to identify underrepresented topics, generating representation data through topic-specific knowledge graphs, and balancing data sets to enhance topic identification accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If data sets are used for topic identification, then topic classification can be performed, but underrepresented data sets lead to decreased probability of topic identification and inefficient processing

Engineering Contradiction:
Improvetopic identification accuracyVSAvoidprocessing efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system performs preliminary analysis to identify underrepresented topics before main processing, calculates representation scores, and generates synthetic data in advance to balance the data sets, thereby improving both accuracy and efficiency during actual topic identification

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system creates synthetic copies of data from represented topics to augment underrepresented topics, generating artificial data samples that mimic the structure and characteristics of existing data to balance data distribution across all topics

Inventive Principle:
Principle #26Copying

2Productivity

If traditional data processing methods are used, then all data is processed uniformly, but this wastes processor and memory resources on underrepresented topics with low identification probability

Engineering Contradiction:
Improveresource utilization efficiencyVSAvoidprocessor and memory resource wastage
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The system applies different processing strategies to different topics based on their representation scores - fully processing represented topics while generating only synthetic data for underrepresented topics, thereby optimizing resource allocation according to local data quality needs

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system changes the data generation parameter based on topic representation - using complete data processing for high-scoring topics and synthetic data generation for low-scoring topics, dynamically adjusting processing intensity to match topic importance

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS10915820B2Generating data associated with underrepresented data based on a received data input
Publication Date: 2021.02.09 ACCENTURE GLOBAL SOLUTIONS LTD
  • US10915820B2 patent drawing
  • US10915820B2 patent drawing
  • US10915820B2 patent drawing

AI summary

An example method described herein involves receiving a data input; identifying a plurality of topics in the data input; determining an underrepresented set of data for a first set of topics of the plurality of topics based on a plurality of knowledge graphs associated with the first set of topics; calculating a score for each topic of the first set of topics based on a representative learning technique; determining that the score for a first topic of the first set of topics satisfies a threshold score; selecting a topic specific knowledge graph based on the first topic; identifying representative objects that are similar to objects of the data input based on the topic specific knowledge graph; generating representation data that is similar to the data input based on the representative objects to balance the underrepresented set of data with a set of data associated with a second set of topics of the plurality of topics; and performing an action associated with the representation data.