Data Extraction System Minority Tag Identification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for extracting minority tags from large datasets struggle to efficiently identify useful minority tags and represent data groups accurately, often due to extreme bias in topic instance numbers and lack of guarantee on topic coherence.
Innovation Solution
A data extraction system that receives a data group with tag ID-appended data and time information, counts tag instances in each time slice, and extracts minority tags based on instance number and time slice ratio thresholds, while determining representative data using word appearance rates and peak timezones.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If keyword search is used to acquire desired text data, then the search can be executed by designating keywords representing features, but a huge amount of data is collected making it difficult to acquire desired text efficiently
Solution Approach 1:
The patent segments the large dataset into clusters based on dependency relationships between text documents. By converting dependency relationships into patterns and applying threshold values, the system divides the huge data volume into manageable clusters, allowing efficient extraction of desired text while maintaining search precision.
2Adaptability or versatility
If minority tags are used to represent minority text data, then diverse topics can be captured, but it is difficult to determine which minority tag to select and confirming all tags detracts from tagging advantages
Solution Approach 1:
The patent extracts representative text from each minority tag cluster based on specific criteria (number of instances and time slice distribution). By automatically selecting and extracting representative texts that best represent each minority tag, the system reduces the complexity of tag selection while maintaining comprehensive topic coverage.
Solution Approach 2:
The patent introduces representative text as an intermediary between minority tags and users. Instead of requiring users to directly evaluate and select from numerous minority tags, the system presents representative texts that encapsulate the essence of each tag cluster, making tag selection more manageable while preserving diverse topic representation.
3Productivity
If classification is performed based on words and grammar to extract minority clusters, then clusters can be obtained, but there is no guarantee that classified clusters represent the same topic
Solution Approach 1:
The patent employs feedback mechanisms by analyzing the distribution of text instances across time slices and using this information to refine cluster representation. By continuously evaluating whether texts within a cluster maintain consistent topic representation over time, the system improves topic coherence while maintaining classification efficiency.
Data Source
AI summary
An input of a data group is received in which tag ID-appended data and time information on time the tag ID-appended data was created are associated; the number of instances is counted of the tag ID-appended data, to which a tag identified by a tag ID has been appended, occurring in each time slice for each tag ID included in the tag ID-appended data and for each time slice obtained by dividing a timeline by a predetermined duration, and the tags are extracted as a few exuberant and useful minority tags in a case where the counted number of instances is greater than a predetermined instance number threshold value and where a ratio of the time slices in which the number of instances does not satisfy a predetermined criterion is greater than a predetermined ratio threshold value; and data in which a score satisfies a predetermined criterion is determined.


