Meaning Extraction System Clustering Feature Vectors
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing meaning extraction systems face accuracy issues due to variations in text writing styles, leading to poor generalization capacity and decreased performance when extracting characteristic expressions from text data.
Innovation Solution
A meaning extraction system that includes clustering and cluster updating mechanisms to generate feature vectors, improve deviations in clusters, and create extraction rules with improved generalization capacity by clustering feature vectors based on similarity and updating them to reduce text writing style deviations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If extraction rules are generated using traditional machine learning on annotated text data, then the system can extract characteristic expressions from text, but the accuracy deteriorates when text writing styles vary due to poor generalization capacity
Solution Approach 1:
The patent segments the text data processing into distinct phases: creating multiple subsets of annotated text data divided by writing style, generating extraction rules for each subset separately, and then combining these rules. This segmentation allows the system to handle writing style variations more effectively by treating each style as a separate category that requires its own specialized extraction rules, thereby improving both accuracy and generalization capacity.
Solution Approach 2:
The patent changes the parameter of text data organization by dividing annotated text data into multiple subsets based on writing style characteristics. By organizing data according to writing style parameters and generating extraction rules specific to each subset, the system adapts to variations in writing styles while maintaining high extraction accuracy across different text types.
2Measurement precision
If the system processes text data with diverse writing styles using a single extraction rule set, then the device complexity remains low, but the extraction accuracy decreases due to deviation in text information
Solution Approach 1:
The patent divides the text data into multiple subsets based on writing style characteristics and generates separate extraction rules for each subset. This segmentation approach improves measurement precision by creating specialized rules tailored to each writing style, while the automated clustering and rule generation processes keep the increase in device complexity manageable.
Solution Approach 2:
The patent implements a feedback mechanism where extraction rules are generated for each text subset, applied to validate performance, and then refined based on results. This iterative feedback process improves extraction accuracy by continuously optimizing rules for each writing style subset while maintaining a systematic approach that prevents excessive complexity accumulation.
Data Source
AI summary
A meaning extraction device includes a clustering unit, an extraction rule generation unit and an extraction rule application unit. The clustering unit acquires feature vectors that transform numerical features representing the features of words having specific meanings and the surrounding words into elements, and clusters the acquired feature vectors into a plurality of clusters on the basis of the degree of similarity between feature vectors. The extraction rule generation unit performs machine learning based on the feature vectors within a cluster for each cluster, and generates extraction rules to extract words having specific meanings. The extraction rule application unit receives feature vectors generated from the words in documents which are subject to meaning extraction, specifies the optimum extraction rules for the feature vectors, and extracts the meanings of the words on the basis of which the feature vectors were generated by applying the specified extraction rules to the feature vectors.


