Semantic Unit Recognition via Coherence and Variation Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing text processing technologies fail to effectively identify and process multi-word text sequences as semantically meaningful units, which is crucial for applications like search engines, named entity learning, and document summarization.
Innovation Solution
A method and device that calculate the coherence and variation of terms in a sequence, using a coherence component, variation component, and decision component to determine if a sequence constitutes a semantic unit, by analyzing the sequence's occurrence in a document collection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If text processing systems treat each word as a separate unit, then processing simplicity is maintained, but semantic accuracy deteriorates
Solution Approach 1:
The patent segments text into semantic units by identifying multi-word sequences that function as cohesive entities. The system divides text not by individual words but by meaningful semantic boundaries, using coherence metrics to determine where sequences should be treated as single units versus separate words. This segmentation approach resolves the contradiction by maintaining processing simplicity at the semantic unit level while improving semantic accuracy.
Solution Approach 2:
The patent introduces coherence metrics and variation analysis as intermediary components between individual words and semantic units. These intermediaries evaluate the relationships between adjacent words and determine whether they form cohesive semantic units, thereby bridging the gap between simple word-level processing and complex semantic understanding without requiring full manual annotation.
2Reliability
If multi-word sequences are processed as single semantic units, then semantic meaning is preserved, but processing complexity increases
Solution Approach 1:
The patent applies partial action by focusing coherence analysis only on candidate multi-word sequences rather than analyzing every possible word combination in the text. The system identifies potential semantic units based on frequency and context, then applies coherence metrics only to these candidates, rather than exhaustively analyzing all possible sequences. This approach preserves semantic meaning while reducing processing complexity through selective analysis.
Solution Approach 2:
The patent changes the parameters used for text processing from individual word properties to sequence-level coherence parameters. Instead of analyzing words based on their individual characteristics, the system evaluates sequences based on coherence metrics that measure the strength of relationships between adjacent words. This parameter transformation enables reliable semantic unit identification while managing complexity through mathematical formulations.
3Measurement precision
If semantic units are identified using traditional methods, then processing speed is maintained, but identification accuracy deteriorates
Solution Approach 1:
The patent performs preliminary action by pre-identifying candidate semantic units based on their frequency and distribution patterns in the text corpus. Before applying complex coherence analysis, the system filters potential semantic units using simpler criteria such as occurrence frequency and positional patterns. This preliminary filtering reduces the number of sequences requiring detailed coherence evaluation, thereby improving identification accuracy while maintaining processing speed through staged analysis.
Solution Approach 2:
The patent substitutes mechanical or rule-based semantic unit identification methods with coherence metric-based analysis. Instead of using fixed linguistic rules or manual annotation mechanisms, the system employs computational coherence metrics that automatically evaluate the semantic relationships between words. This substitution improves identification accuracy through data-driven analysis while maintaining productivity through efficient algorithmic processing.
Data Source
AI summary
A semantic locator determines whether input sequences form semantically meaningful units. The semantic locator includes a coherence component that calculates a coherence of the terms in the sequence and a variation component that calculates the variation in terms that surround the sequence. A heuristics component may additionally refine results of the coherence component and the variation component. A decision component may make the determination of whether the sequence is a semantic unit based on the results of the coherence component, variation component, and heuristics component.


