Inline Response Generation for Voice and Video Messages

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for responding to video or voice messages are cumbersome, leading to user abandonment, especially when adding responses to long messages, as they require manual selection of response locations, which can elongate the message and deter senders from listening.

Innovation Solution

A method that automatically detects when a recipient is speaking, determines the appropriate location in the sender's media to insert the response, and generates combined media including both sender and recipient media, using machine learning to identify pauses or semantic breaks to summarize and contextualize the sender's media.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual technique is used to add responses to video or voice messages, then the recipient can respond to specific questions, but the process becomes arduous and may lead to user abandonment

Engineering Contradiction:
Improveresponse completion rateVSAvoidease of adding response
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The system automatically detects when the recipient is speaking and determines where to insert their response without requiring manual intervention. The recipient simply speaks their response, and the system handles the rest - detecting speech start points, identifying insertion locations in the sender's media, and generating the combined media file automatically.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The manual mechanical process of clicking and selecting response locations is replaced with an automated system that uses speech detection and machine learning models to automatically identify where responses should be inserted in the sender's video or voice message.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If responses are added to already long messages, then the recipient can provide feedback, but the message becomes even longer and the sender has no desire to listen

Engineering Contradiction:
Improvefeedback provisionVSAvoidmessage length
Core Design Contradiction:
ReliabilityVSLength of moving object

Solution Approach 1:

The system extracts only the relevant portions of the sender's media that contain the questions being asked, rather than including the entire message. By using machine learning models to identify questions and their context, the system creates a condensed version that includes only the necessary segments before inserting recipient responses.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The sender's media is segmented into relevant portions containing questions and their context. The system identifies and extracts specific segments that need responses, rather than treating the entire message as a single unit, thereby reducing overall length while maintaining responsiveness.

Inventive Principle:
Principle #1Segmentation

3Loss of information

If the entire sender media is included in combined media, then context is preserved, but the message length increases and reduces sender motivation to listen

Engineering Contradiction:
Improvecontext preservationVSAvoidcombined media length
Core Design Contradiction:
Loss of informationVSLength of moving object

Solution Approach 1:

The system applies different quality levels to different parts of the sender's media. Rather than uniformly including all segments, it identifies and preserves only the local portions that provide necessary context for understanding the recipient's responses - specifically the questions and their immediate context.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system performs partial inclusion of the sender's media by selecting only the portions that are excessive for understanding responses - specifically the questions and minimal surrounding context. This partial action is sufficient to maintain context while avoiding the excess of including the entire original message.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11425072B2Inline responses to video or voice messages
Publication Date: 2022.08.23 GOOGLE LLC
  • US11425072B2 patent drawing
  • US11425072B2 patent drawing
  • US11425072B2 patent drawing

AI summary

The method includes receiving sender media that was recorded by a sender device associated with a sender. The method further comprises playing, by a recipient device, the sender media for a recipient. The method further comprises detecting that the recipient is speaking. The method further comprises recording recipient media based on detecting that the recipient is speaking. The method further comprises determining a location in the sender media at which the recipient media is to be included. The method further comprises generating combined media that includes at least a portion of the sender media and the recipient media at the location.