Semantic Unit Recognition via Coherence and Variation Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing text processing technologies fail to effectively identify and process multi-word text sequences as semantically meaningful units, which is crucial for applications like search engines, named entity learning, and document summarization.

Innovation Solution

A method and device that calculate the coherence and variation of terms in a sequence, using a coherence component, variation component, and decision component to determine if a sequence constitutes a semantic unit, by analyzing the sequence's occurrence in a document collection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If text processing systems treat each word as a separate unit, then processing simplicity is maintained, but semantic accuracy deteriorates

Engineering Contradiction:
Improvesemantic accuracyVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments text into semantic units by identifying multi-word sequences that function as cohesive entities. The system divides text not by individual words but by meaningful semantic boundaries, using coherence metrics to determine where sequences should be treated as single units versus separate words. This segmentation approach resolves the contradiction by maintaining processing simplicity at the semantic unit level while improving semantic accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces coherence metrics and variation analysis as intermediary components between individual words and semantic units. These intermediaries evaluate the relationships between adjacent words and determine whether they form cohesive semantic units, thereby bridging the gap between simple word-level processing and complex semantic understanding without requiring full manual annotation.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If multi-word sequences are processed as single semantic units, then semantic meaning is preserved, but processing complexity increases

Engineering Contradiction:
Improvesemantic meaning preservationVSAvoidsequence identification complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies partial action by focusing coherence analysis only on candidate multi-word sequences rather than analyzing every possible word combination in the text. The system identifies potential semantic units based on frequency and context, then applies coherence metrics only to these candidates, rather than exhaustively analyzing all possible sequences. This approach preserves semantic meaning while reducing processing complexity through selective analysis.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent changes the parameters used for text processing from individual word properties to sequence-level coherence parameters. Instead of analyzing words based on their individual characteristics, the system evaluates sequences based on coherence metrics that measure the strength of relationships between adjacent words. This parameter transformation enables reliable semantic unit identification while managing complexity through mathematical formulations.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If semantic units are identified using traditional methods, then processing speed is maintained, but identification accuracy deteriorates

Engineering Contradiction:
Improvesemantic unit identification accuracyVSAvoidprocessing speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent performs preliminary action by pre-identifying candidate semantic units based on their frequency and distribution patterns in the text corpus. Before applying complex coherence analysis, the system filters potential semantic units using simpler criteria such as occurrence frequency and positional patterns. This preliminary filtering reduces the number of sequences requiring detailed coherence evaluation, thereby improving identification accuracy while maintaining processing speed through staged analysis.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent substitutes mechanical or rule-based semantic unit identification methods with coherence metric-based analysis. Instead of using fixed linguistic rules or manual annotation mechanisms, the system employs computational coherence metrics that automatically evaluate the semantic relationships between words. This substitution improves identification accuracy through data-driven analysis while maintaining productivity through efficient algorithmic processing.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS7580827B1Semantic unit recognition
Publication Date: 2009.08.25 GOOGLE LLC
  • US7580827B1 patent drawing
  • US7580827B1 patent drawing
  • US7580827B1 patent drawing

AI summary

A semantic locator determines whether input sequences form semantically meaningful units. The semantic locator includes a coherence component that calculates a coherence of the terms in the sequence and a variation component that calculates the variation in terms that surround the sequence. A heuristics component may additionally refine results of the coherence component and the variation component. A decision component may make the determination of whether the sequence is a semantic unit based on the results of the coherence component, variation component, and heuristics component.