Automated Lexico-Syntactic Pattern Extraction for Named Entity Relations

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for identifying semantic relationships between named entities in text are inefficient due to the difficulty in establishing reliable lexico-syntactic patterns, leading to a large number of non-responsive search results.

Innovation Solution

A computer-implemented system and method that automatically generates lexico-syntactic patterns by retrieving text strings with named entities, extracting syntactic patterns, and generating generalized rules to identify candidate instances of semantic relations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual establishment of lexico-syntactic patterns is used, then pattern reliability can be improved, but time consumption and labor requirements increase significantly

Engineering Contradiction:
Improvepattern reliabilityVSAvoidtime consumption
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs self-service by automatically extracting and generating lexico-syntactic patterns from text data without requiring manual establishment. The pattern extraction module automatically analyzes text strings and generates patterns that reflect semantic relations between named entities, eliminating the need for manual pattern creation while maintaining reliability through systematic automated analysis

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual process of pattern establishment with an automated computational system. The pattern extraction module uses computational algorithms to automatically identify and generate lexico-syntactic patterns from text, substituting human manual analysis with automated mechanical processing that achieves both speed and reliability

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Loss of time

If automated pattern generation is used, then time consumption is reduced, but the complexity of the system increases

Engineering Contradiction:
Improvetime consumptionVSAvoidsystem complexity
Core Design Contradiction:
Loss of timeVSDevice complexity

Solution Approach 1:

The system segments the complex task of semantic relation extraction into distinct functional modules: a text retrieval module that obtains text strings, a pattern extraction module that identifies lexico-syntactic patterns, and a rule generation module that creates generalized rules. This segmentation divides the complex system into manageable, independent components that work together to achieve automated pattern generation

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The pattern extraction module serves multiple functions: it extracts patterns from text strings, identifies semantic relations between named entities, and generates generalized rules. This multi-functionality reduces the need for separate specialized components, thereby managing system complexity while achieving comprehensive automated processing

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Loss of information

If search for sentences with named entities is performed, then relevant information can be retrieved, but the number of non-responsive results increases

Engineering Contradiction:
Improveinformation retrievalVSAvoidnumber of results
Core Design Contradiction:
Loss of informationVSQuantity of substance

Solution Approach 1:

The system dynamically adapts the search process by first extracting actual lexico-syntactic patterns from the text data and then using these extracted patterns to filter and refine search results. This dynamic approach allows the system to adjust the search criteria based on the actual content and semantic relations found in the text, improving result relevance

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system implements feedback by extracting patterns from retrieved text strings and using these extracted patterns to refine subsequent search and filtering operations. The pattern extraction module analyzes the retrieved text and generates feedback information in the form of generalized rules that guide future retrieval operations, creating an iterative improvement cycle that reduces non-responsive results

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS8370128B2Semantically-driven extraction of relations between named entities
Publication Date: 2013.02.05 XEROX CORP
  • US8370128B2 patent drawing
  • US8370128B2 patent drawing
  • US8370128B2 patent drawing

AI summary

A system and method of developing rules for text processing enable retrieval of instances of named entities in a predetermined semantic relation (such as the DATE and PLACE of an EVENT) by extracting patterns from text strings in which attested examples of named entities satisfying the semantic relation occur. The patterns are generalized to form rules which can be added to the existing rules of a syntactic parser and subsequently applied to text to find candidate instances of other named entities in the predetermined semantic relation.