Data Extraction System Minority Tag Identification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for extracting minority tags from large datasets struggle to efficiently identify useful minority tags and represent data groups accurately, often due to extreme bias in topic instance numbers and lack of guarantee on topic coherence.

Innovation Solution

A data extraction system that receives a data group with tag ID-appended data and time information, counts tag instances in each time slice, and extracts minority tags based on instance number and time slice ratio thresholds, while determining representative data using word appearance rates and peak timezones.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If keyword search is used to acquire desired text data, then the search can be executed by designating keywords representing features, but a huge amount of data is collected making it difficult to acquire desired text efficiently

Engineering Contradiction:
Improvesearch precisionVSAvoiddata volume
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent segments the large dataset into clusters based on dependency relationships between text documents. By converting dependency relationships into patterns and applying threshold values, the system divides the huge data volume into manageable clusters, allowing efficient extraction of desired text while maintaining search precision.

Inventive Principle:
Principle #1Segmentation

2Adaptability or versatility

If minority tags are used to represent minority text data, then diverse topics can be captured, but it is difficult to determine which minority tag to select and confirming all tags detracts from tagging advantages

Engineering Contradiction:
Improvetopic coverageVSAvoidtag selection difficulty
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The patent extracts representative text from each minority tag cluster based on specific criteria (number of instances and time slice distribution). By automatically selecting and extracting representative texts that best represent each minority tag, the system reduces the complexity of tag selection while maintaining comprehensive topic coverage.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces representative text as an intermediary between minority tags and users. Instead of requiring users to directly evaluate and select from numerous minority tags, the system presents representative texts that encapsulate the essence of each tag cluster, making tag selection more manageable while preserving diverse topic representation.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If classification is performed based on words and grammar to extract minority clusters, then clusters can be obtained, but there is no guarantee that classified clusters represent the same topic

Engineering Contradiction:
Improveclassification efficiencyVSAvoidtopic coherence
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent employs feedback mechanisms by analyzing the distribution of text instances across time slices and using this information to refine cluster representation. By continuously evaluating whether texts within a cluster maintain consistent topic representation over time, the system improves topic coherence while maintaining classification efficiency.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12210557B2Data extraction system and data extraction method
Publication Date: 2025.01.28 HITACHI LTD
  • US12210557B2 patent drawing
  • US12210557B2 patent drawing
  • US12210557B2 patent drawing

AI summary

An input of a data group is received in which tag ID-appended data and time information on time the tag ID-appended data was created are associated; the number of instances is counted of the tag ID-appended data, to which a tag identified by a tag ID has been appended, occurring in each time slice for each tag ID included in the tag ID-appended data and for each time slice obtained by dividing a timeline by a predetermined duration, and the tags are extracted as a few exuberant and useful minority tags in a case where the counted number of instances is greater than a predetermined instance number threshold value and where a ratio of the time slices in which the number of instances does not satisfy a predetermined criterion is greater than a predetermined ratio threshold value; and data in which a score satisfies a predetermined criterion is determined.