Location-Dependent Speech Recognition Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional speech recognition systems fail to account for location-specific context within buildings, leading to ambiguity in word usage and recognition accuracy, as they do not differentiate between areas like living rooms and kitchens based on unique vocabularies and acoustic signatures.

Innovation Solution

The development of location-dependent speech recognition models that estimate user location using wireless radio transponders or microphone arrays, generating location-specific acoustic and language models to accurately interpret utterances by correlating them with specific areas in a building, such as kitchens or living rooms.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional language models are used for speech recognition, then the system can recognize speech across all locations, but the recognition accuracy decreases due to location-specific word usage variations

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidlocation-specific context adaptation
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent implements location-specific language models that are tailored to each particular location within a building. Each location model captures the unique vocabulary and word usage patterns characteristic of that specific area (e.g., kitchen-specific terms vs. living room-specific terms), thereby improving recognition accuracy for location-dependent speech without requiring a single universal model to handle all variations.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent segments the building into multiple distinct locations, each with its own language model. This segmentation allows the system to divide the speech recognition task into location-specific sub-tasks, where each model focuses on the unique characteristics of its designated area, thereby resolving the contradiction between maintaining versatility across locations and achieving precision within each location.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If location-specific models are created for each area, then speech recognition accuracy improves, but the system complexity increases due to multiple models

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidnumber of speech recognition models
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent implements a dynamic model selection mechanism that automatically chooses the appropriate location-specific language model based on real-time location data from the mobile device. This dynamic approach allows the system to switch between different models as the user moves through the building, maintaining high recognition accuracy without requiring all models to be actively processed simultaneously, thereby managing system complexity more efficiently.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system uses the mobile device's existing location tracking capabilities (GPS, Wi-Fi positioning, or other location services) to automatically determine which location model to apply. This self-service approach eliminates the need for manual model selection or complex user input, allowing the system to autonomously manage the complexity of multiple models by leveraging data already collected by the device.

Inventive Principle:
Principle #25Self-service

3Device complexity

If a single universal language model is used, then the system structure remains simple, but word usage ambiguity increases in location-specific contexts

Engineering Contradiction:
Improvemodel structureVSAvoidlocation context information
Core Design Contradiction:
Device complexityVSLoss of information

Solution Approach 1:

The patent introduces location data from the mobile device as an intermediary that bridges the gap between the user's physical location and the appropriate language model. This intermediary provides the necessary context information to select the correct model, thereby preserving location-specific word usage patterns without requiring a complex manual configuration system. The location data acts as a key that connects the user's environment to the corresponding linguistic characteristics.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentEP2880844B1Speech recognition models based on location indicia
Publication Date: 2019.12.11 GOOGLE LLC
  • EP2880844B1 patent drawingFigure 1a
  • EP2880844B1 patent drawingFigure 1b
  • EP2880844B1 patent drawingFigure 2

AI summary

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for performing speech recognition using models that are based on where, within a building, a speaker makes an utterance are disclosed. The methods, systems, and apparatus include actions of receiving data corresponding to an utterance, and obtaining location indicia for an area within a building where the utterance was spoken. Further actions include selecting one or more models for speech recognition based on the location indicia, wherein each of the selected one or more models is associated with a weight based on the location indicia. Additionally, the actions include generating a composite model using the selected one or more models and the respective weights of the selected one or more models. And the actions also include generating a transcription of the utterance using the composite model.