Contextual Speech Recognition Using Context Deletion for Privacy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Speech recognition systems face challenges in accurately transcribing spoken input, especially on small-screen devices where users may find virtual keyboards cumbersome, and existing solutions do not effectively utilize context information to improve transcription accuracy while protecting user privacy.

Innovation Solution

A computer-implemented method that receives spoken input and associated context information, determines multiple hypotheses for transcription, selects likely intended transcriptions based on context, and deletes context information after use to ensure privacy, using context such as personal contacts, location, and recent user activity to enhance transcription accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If context information is used to improve speech transcription accuracy, then transcription accuracy is improved, but user privacy is compromised

Engineering Contradiction:
Improvetranscription accuracyVSAvoiduser privacy exposure
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent extracts only the essential context information needed for improving transcription accuracy while leaving out unnecessary personal data. The system selectively processes context to achieve the desired accuracy improvement without capturing excessive user information that would compromise privacy.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent implements a mechanism where context information is used temporarily for transcription processing and then discarded afterward. The system recovers the necessary transcription accuracy during the processing window but eliminates the stored context data afterward, preventing long-term privacy exposure while maintaining short-term accuracy benefits.

Inventive Principle:
Principle #34Discarding and recovering

2Ease of operation

If virtual keyboard is provided on small-screen devices, then text input capability is provided, but ease of operation deteriorates

Engineering Contradiction:
Improvetext input convenienceVSAvoidinterface complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent replaces the mechanical interaction of typing on a virtual keyboard with a voice-based acoustic input system. Users speak instead of manually selecting keys, substituting the mechanical typing process with speech recognition, thereby dramatically improving ease of operation on small-screen devices.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The speech recognition system automatically transcribes spoken input into text without requiring manual intervention. The system serves itself by processing the voice input and generating the corresponding text output, eliminating the need for users to manually navigate complex virtual keyboard interfaces.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS8862467B1Contextual speech recognition
Publication Date: 2014.10.14 GOOGLE LLC
  • US8862467B1 patent drawing
  • US8862467B1 patent drawing
  • US8862467B1 patent drawing

AI summary

A computer-implemented method can include receiving, by a computer system, a request to transcribe spoken input from a user of a computing device, the request including information that (i) characterizes a spoken input, and (ii) context information associated with the user or the computing device. The method can determine, based on the information that characterizes the spoken input, multiple hypotheses that each represent a possible textual transcription of the spoken input. The method can select, based on the context information, one or more of the multiple hypotheses for the spoken input as one or more likely intended hypotheses for the spoken input, and can send the one or more likely intended hypotheses for the spoken input to the computing device. In conjunction with sending the one or more likely intended hypotheses for the spoken input to the computing device, the method can delete the context information.