Repeated Speech Recognition in Mobile Devices

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current mobile computing devices lack effective methods to recognize and analyze repeated vocal expressions spoken during conversations, which can lead to misinterpretation of intended actions or information exchange.

Innovation Solution

A processor-based method in mobile computing devices that detects and matches utterances between communicatively coupled devices, temporarily captures matching speech for processing, and activates functions based on the type of matched utterance, including displaying the match as text or hyperlinks, and providing alerts for mismatches.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If mobile computing devices perform basic voice communication, then communication functionality is provided, but the ability to recognize and process repeated utterances is lacking

Engineering Contradiction:
Improvevoice processing capabilityVSAvoidutterance recognition system
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The mobile computing device integrates multiple functions into a single system: voice detection via microphone, utterance matching through processor comparison, temporary storage in memory, and execution of subsequent actions. This allows the device to handle both basic communication and advanced repeated utterance recognition without requiring separate dedicated systems.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If the device captures and processes repeated utterances in real-time, then accurate recognition is achieved, but processing time and computational resources increase

Engineering Contradiction:
Improveutterance matching accuracyVSAvoidprocessing delay
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary capture of the first utterance and stores it temporarily in memory during the voice call. When a second utterance is detected, the processor immediately compares it against the pre-stored first utterance, enabling rapid matching without requiring complex real-time analysis of multiple utterances simultaneously.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If the system provides multiple processing functions based on utterance type, then user utility is enhanced, but system complexity increases

Engineering Contradiction:
Improveprocessing function varietyVSAvoidcontrol system
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system varies the subsequent processing function based on the type of matched utterance detected. Different utterance types trigger different actions (e.g., launching applications, searching information, dialing numbers), allowing the device to adapt its behavior to user intent without requiring a completely separate system for each function.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS9093075B2Recognizing repeated speech in a mobile computing device
Publication Date: 2015.07.28 GOOGLE TECHNOLOGY HOLDINGS LLC
  • US9093075B2 patent drawing
  • US9093075B2 patent drawing
  • US9093075B2 patent drawing

AI summary

A method is disclosed herein for recognizing a repeated utterance in a mobile computing device via a processor. A first utterance is detected being spoken into a first mobile computing device. Likewise, a second utterance is detected being spoken into a second mobile computing device within a predetermined time period. The second utterance substantially matches the first spoken utterance and the first and second mobile computing devices are communicatively coupled to each other. The processor enables capturing, at least temporarily, a matching utterance for performing a subsequent processing function. The performed subsequent processing function is based on a type of captured utterance.