Topic-Specific Knowledge Graphs for Underrepresented Data Balancing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing systems face challenges in accurately identifying topics associated with data inputs due to underrepresented data sets, leading to decreased probability of topic identification and inefficient processing resources.
Innovation Solution
A method utilizing a domain knowledge graph to identify underrepresented topics, generating representation data through topic-specific knowledge graphs, and balancing data sets to enhance topic identification accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If data sets are used for topic identification, then topic classification can be performed, but underrepresented data sets lead to decreased probability of topic identification and inefficient processing
Solution Approach 1:
The system performs preliminary analysis to identify underrepresented topics before main processing, calculates representation scores, and generates synthetic data in advance to balance the data sets, thereby improving both accuracy and efficiency during actual topic identification
Solution Approach 2:
The system creates synthetic copies of data from represented topics to augment underrepresented topics, generating artificial data samples that mimic the structure and characteristics of existing data to balance data distribution across all topics
2Productivity
If traditional data processing methods are used, then all data is processed uniformly, but this wastes processor and memory resources on underrepresented topics with low identification probability
Solution Approach 1:
The system applies different processing strategies to different topics based on their representation scores - fully processing represented topics while generating only synthetic data for underrepresented topics, thereby optimizing resource allocation according to local data quality needs
Solution Approach 2:
The system changes the data generation parameter based on topic representation - using complete data processing for high-scoring topics and synthetic data generation for low-scoring topics, dynamically adjusting processing intensity to match topic importance
Data Source
AI summary
An example method described herein involves receiving a data input; identifying a plurality of topics in the data input; determining an underrepresented set of data for a first set of topics of the plurality of topics based on a plurality of knowledge graphs associated with the first set of topics; calculating a score for each topic of the first set of topics based on a representative learning technique; determining that the score for a first topic of the first set of topics satisfies a threshold score; selecting a topic specific knowledge graph based on the first topic; identifying representative objects that are similar to objects of the data input based on the topic specific knowledge graph; generating representation data that is similar to the data input based on the representative objects to balance the underrepresented set of data with a set of data associated with a second set of topics of the plurality of topics; and performing an action associated with the representation data.


