Real-Time Speech Recognition With Contextual Augmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In automated speech recognition systems, users often finish speaking before recognition results are displayed or acted upon, limiting the ability to provide real-time contextual suggestions during conversations.

Innovation Solution

A system processes conversational speech from multiple users to identify topics and key words/phrases by analyzing speech characteristics and visual cues, providing contextual information such as transcriptions, hyperlinks, and maps in real-time to enhance the conversation experience.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If speech recognition processes and displays results after user finishes speaking, then recognition accuracy is improved, but real-time interaction capability deteriorates

Engineering Contradiction:
Improverecognition accuracyVSAvoidreal-time interaction capability
Core Design Contradiction:
Measurement precisionVSSpeed

Solution Approach 1:

The system performs preliminary speech recognition processing during the user's speech, generating partial results before the user finishes speaking. This allows the system to prepare recognition outcomes in advance, enabling display or action to occur while the user is still speaking or immediately upon completion, thus maintaining both accuracy and real-time responsiveness.

Inventive Principle:
Principle #10Preliminary action

2Loss of information

If the system provides comprehensive recognition results after speech completion, then information completeness is improved, but user engagement during speech deteriorates

Engineering Contradiction:
Improveinformation completenessVSAvoiduser engagement during speech
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The system provides intermediate feedback by displaying partial recognition results, contextual suggestions, or relevant information during the user's speech. This continuous feedback loop keeps users engaged and informed about how their speech is being processed, while still delivering complete recognition results after speech completion.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs preliminary processing to identify and display relevant information, contextual suggestions, or key concepts during speech. This preliminary action provides value to users in real-time without compromising the completeness of final recognition results.

Inventive Principle:
Principle #10Preliminary action

3Loss of information

If the system processes speech to identify topics and keywords in real-time, then contextual suggestion quality is improved, but processing time increases

Engineering Contradiction:
Improvecontextual suggestion qualityVSAvoidprocessing time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The system performs partial speech processing during user speech to identify topics, keywords, or relevant concepts, generating contextual suggestions based on incomplete but progressively improving data. This partial action provides timely contextual feedback while avoiding the need to wait for complete speech processing, thus reducing perceived processing time.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system performs preliminary analysis of speech content during delivery, identifying topics and keywords as they are spoken. This preliminary action enables the generation of contextual suggestions without requiring complete speech processing to be finished first, thereby reducing processing delays.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20230315987A1Speech recognition and summarization
Publication Date: 2023.10.05 GOOGLE LLC
  • US20230315987A1 patent drawing
  • US20230315987A1 patent drawing
  • US20230315987A1 patent drawing

AI summary

The subject matter of this specification can be embodied in, among other things, a method that includes receiving two or more data sets each representing speech of a corresponding individual attending an internet-based social networking video conference session, decoding the received data sets to produce corresponding text for each individual attending the internet-based social networking video conference, and detecting characteristics of the session from a coalesced transcript produced from the decoded text of the attending individuals for providing context to the internet-based social networking video conference session.