Event Matching by Text Feature Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

It is time-consuming and challenging for readers to manually identify similar texts that refer to the same event, especially when dealing with diverse terminology and writing styles across multiple data sources.

Innovation Solution

A system and method for event matching by analyzing text characteristics, which involves acquiring a document collection, identifying document subsets based on structured metadata, extracting and normalizing salient text features, and generating an event similarity score to automatically identify documents describing the same event.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual reading and comparison of texts is used to identify documents describing the same event, then accuracy in identifying event similarities can be maintained, but the time required and manual effort increase significantly

Engineering Contradiction:
Improveaccuracy in identifying event similaritiesVSAvoidtime required for manual reading and comparison
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces the manual mechanical reading and comparison process with an automated computational system. The system uses text feature extraction, normalization, and similarity scoring algorithms to automatically identify documents describing the same event, eliminating the need for human readers to manually compare texts while maintaining high accuracy through structured text analysis.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent introduces intermediate text features as mediators between the raw documents and the final similarity assessment. By extracting and normalizing specific text features (such as event entities, temporal expressions, and spatial expressions) before comparison, the system creates a standardized intermediate representation that enables accurate automated matching without requiring direct manual reading of full documents.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If automated text comparison methods are used to identify similar documents, then processing speed and efficiency improve, but accuracy in identifying true event similarities may deteriorate due to diverse terminology and writing styles

Engineering Contradiction:
Improveprocessing speed and efficiencyVSAvoidaccuracy in identifying event similarities
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent transforms the text comparison problem by changing the parameters of analysis from raw text to normalized text features. By converting diverse terminology and writing styles into standardized feature representations (such as normalized event entities, temporal expressions, and spatial expressions), the system maintains high processing speed while improving accuracy through parameter standardization that eliminates variations in terminology and style.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent applies local quality by focusing analysis on specific critical text features rather than attempting to compare entire documents uniformly. By identifying and analyzing key local features such as event entities, temporal expressions, and spatial expressions that are most indicative of event similarity, the system achieves high accuracy in identifying true event matches while maintaining efficient processing through selective feature comparison.

Inventive Principle:
Principle #3Local quality

3Reliability

If comprehensive text analysis is performed on all documents in the collection, then complete event similarity assessment can be achieved, but the computational complexity and processing time increase

Engineering Contradiction:
Improvecompleteness of event similarity assessmentVSAvoidcomputational complexity and processing time
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent extracts only the essential text features that are most relevant to event similarity assessment, such as event entities, temporal expressions, and spatial expressions. By taking out and analyzing only these critical features rather than performing comprehensive analysis of all text content, the system achieves reliable event similarity assessment with reduced computational complexity and faster processing time.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent segments the text analysis process into distinct stages: document acquisition, text feature extraction, feature normalization, and similarity scoring. This segmentation allows the system to process documents in manageable steps, analyzing only the necessary features at each stage, thereby maintaining complete event similarity assessment while reducing overall computational complexity through structured modular processing.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10606869B2Event matching by analysis of text characteristics (E-MATCH)
Publication Date: 2020.03.31 THE BOEING CO
  • US10606869B2 patent drawing
  • US10606869B2 patent drawing
  • US10606869B2 patent drawing

AI summary

A system and method for event matching by analysis of text characteristics are presented. A document collection comprising documents is acquired. One or more document subsets of the document collection each comprising one or more documents potentially describing identical events are identified based on certain structured metadata fields of the documents. Salient text features are extracted from the documents in the document collection. An event similarity score for pairs of documents in the document collection is generated by comparing the text features extracted from the documents. A common event document list comprising sets of documents in the document collection whose event similarity scores with each other are above a similarity threshold is generated.