Semiotic Class Normalization via Universal Grammar and Lexical Maps

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech and language recognition systems require labor-intensive and resource-heavy processes for converting non-standard words into verbalizations, necessitating extensive manual grammar development and large training datasets, which is time-consuming and requires specialized linguistic expertise.

Innovation Solution

The implementation of a language universal covering grammar and language-specific lexical maps to generate and select verbalizations for input strings, reducing the need for extensive linguistic knowledge and parallel data, and enabling efficient text normalization across multiple languages.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If hand-built grammars are used for text normalization, then verbalization accuracy is improved, but development time and resource requirements increase significantly

Engineering Contradiction:
Improveverbalization accuracyVSAvoiddevelopment time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent uses existing linguistic resources and parallel data to automatically induce grammars, copying linguistic patterns from available data rather than manually constructing them. This allows the system to leverage existing linguistic knowledge embedded in parallel corpora, reducing the need for expert manual grammar development while maintaining accuracy.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the manual mechanical process of grammar writing with an automated computational process. Machine learning algorithms automatically induce grammars from parallel data, substituting the manual linguistic expertise requirement with computational methods that can process large amounts of data efficiently.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If extensive training data is used for selector training, then verbalization selection accuracy is improved, but computer resource requirements increase

Engineering Contradiction:
Improveverbalization selection accuracyVSAvoidcomputer resource requirement
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent uses a relatively small amount of parallel data compared to what would be needed for comprehensive training, achieving sufficient accuracy through selective use of linguistic patterns. The system doesn't require exhaustive training data by focusing on the most relevant linguistic variations for verbalization selection.

Inventive Principle:
Principle #16Partial or excessive action

3Measurement precision

If specialized linguistic expertise is required for grammar development, then verbalization quality is improved, but ease of operation deteriorates

Engineering Contradiction:
Improveverbalization qualityVSAvoidlinguistic expertise requirement
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The system performs self-service by automatically inducing grammars from parallel data without requiring external linguistic experts. The computational process autonomously discovers linguistic patterns and constructs grammars, making the system self-sufficient and eliminating the need for specialized human expertise in grammar development.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent creates a universal system that can handle multiple languages and semiotic classes through a single automated grammar induction process. The induced grammar framework is language-agnostic and can be applied across different languages and text types, making the system universally applicable without requiring language-specific manual grammar writing.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS10210153B1Semiotic class normalization
Publication Date: 2019.02.19 GOOGLE LLC
  • US10210153B1 patent drawing
  • US10210153B1 patent drawing
  • US10210153B1 patent drawing

AI summary

A language processing system for text normalization of an input string of a semiotic class. In an aspect, a method includes receiving an input string; accessing, for a semiotic class of non-standard words, a language universal covering grammar for a plurality of languages that generates, for each language of the plurality of languages, one or more sequences of word-level components for each instance of the semiotic class in the language; for each of the plurality of languages, accessing a lexical map specific to the language and that maps each sequence of word-level components for each instance of the semiotic class in the language verbalizations in the language; generating, from the language universal grammar and the lexical maps, a lattice of possible verbalizations of the input string; and selecting one of the possible verbalizations as a selected verbalization for the input string.