Bootstrapping Named Entity Canonicalizers via Machine Translation Alignment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing approaches to recognizing named entities in speech recognition technology are labor-intensive and inefficient, requiring hand-crafted grammars that are rigid and difficult to adapt to other languages, necessitating human expertise and manual translation of grammatical rules.

Innovation Solution

The method involves using a context-free grammar to generate sample expressions in one language, machine translating these expressions into another language, and associating them with canonical representations to create training data for a machine translator, allowing for automatic association of named entities without the need for hand-crafted grammars in the target language.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If hand-crafted grammars are used to recognize named entities, then recognition accuracy can be achieved, but the process becomes labor-intensive and inefficient

Engineering Contradiction:
Improvenamed entity recognition accuracyVSAvoiddevelopment efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent copies grammatical structures and named entity patterns from one language to another through machine translation, avoiding the need to manually craft grammars for each target language while maintaining recognition accuracy

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the manual mechanical process of hand-crafting grammars with an automated machine translation system that generates grammatical rules and named entity associations programmatically

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If hand-crafted grammars are used for named entity recognition, then recognition can be performed, but the approach is rigid and difficult to adapt to other languages

Engineering Contradiction:
Improvenamed entity recognition reliabilityVSAvoidlanguage adaptability
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal machine translation system that can handle multiple languages through a single framework, allowing the same named entity recognition approach to be applied across different languages without requiring language-specific manual grammar crafting

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent makes the grammar generation process dynamic and adaptable by using machine translation to automatically generate language-specific grammatical rules from a source language grammar, allowing easy adaptation to new languages

Inventive Principle:
Principle #15Dynamics

3Measurement precision

If hand-crafted grammars are used, then named entities can be recognized, but human expertise and manual translation of grammatical rules are required

Engineering Contradiction:
Improvenamed entity recognition precisionVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent enables the system to automatically generate its own language-specific grammars and named entity associations through machine translation, eliminating the need for human experts to manually translate grammatical rules for each target language

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS9146919B2Bootstrapping named entity canonicalizers from English using alignment models
Publication Date: 2015.09.29 GOOGLE LLC
  • US9146919B2 patent drawing
  • US9146919B2 patent drawing
  • US9146919B2 patent drawing

AI summary

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for training recognition canonical representations corresponding to named-entity phrases in a second natural language based on translating a set of allowable expressions with canonical representations from a first natural language, which may be generated by expanding a context-free grammar for the allowable expressions for the first natural language.