Intent-Based Speech-to-Text Formatting and Text Deletion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users of automated assistants often need to manually manipulate textual content after dictation due to the lack of context-aware formatting and deletion capabilities, leading to inefficient use of computational resources and prolonged interaction.

Innovation Solution

An automated assistant processes spoken utterances to determine user intent and apply formatting operations automatically, using machine learning models and contextual data to arrange and delete textual content without explicit user commands.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If the application removes a standard length of text each time the user provides a clear command, then the deletion operation is simple to implement, but it does not accurately reflect user intent and requires additional manual manipulation

Engineering Contradiction:
Improvedeletion operation simplicityVSAvoiddeletion accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The system analyzes the spoken command context and provides feedback about which text will be deleted, allowing the user to confirm or adjust. The NLU engine processes the utterance to understand intent, and the system responds with the identified text segment for deletion, creating a feedback loop that improves precision while maintaining ease of use.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system autonomously determines the appropriate text segment to delete by analyzing the context of the spoken command, eliminating the need for manual text selection. The automated assistant performs the deletion operation based on its understanding of user intent, making the system self-serve the user's deletion needs without requiring additional manual manipulation.

Inventive Principle:
Principle #25Self-service

2Manufacturing precision

If the user provides explicit spoken commands for every formatting operation, then the formatting precision is high, but the interaction duration and computational resource usage increase

Engineering Contradiction:
Improveformatting precisionVSAvoidinteraction duration
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system performs preliminary analysis of the spoken utterance to predict the user's formatting intent before the user has to explicitly state it. The NLU engine processes the utterance contextually to anticipate desired formatting operations, applying them proactively based on the understood intent, thus reducing interaction duration while maintaining formatting precision.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The automated assistant autonomously determines and applies formatting operations based on its understanding of the user's spoken intent, eliminating the need for the user to provide explicit formatting commands. The system serves itself by interpreting context and automatically applying appropriate formatting, reducing both interaction time and computational overhead.

Inventive Principle:
Principle #25Self-service

3Adaptability or versatility

If the system requires users to switch between separate interfaces for speech input and text manipulation, then the functionality is clearly separated, but the ease of operation decreases and interaction becomes more complex

Engineering Contradiction:
Improveinterface functionality separationVSAvoidinterface usability
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The system merges the speech-to-text functionality with text manipulation capabilities into a single integrated interface. The automated assistant processes spoken commands and performs text operations within the same application context, eliminating the need for users to switch between separate interfaces while maintaining clear functional separation through contextual understanding.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The automated assistant serves multiple functions within a single interface, handling both speech-to-text conversion and text manipulation operations. The system is designed to be universal, accommodating various user needs through a unified interface that can interpret and execute diverse spoken commands without requiring users to switch applications or interfaces.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12431138B2Arranging and/or clearing speech-to-text content without a user providing express instructions
Publication Date: 2025.09.30 GOOGLE LLC
  • US12431138B2 patent drawing
  • US12431138B2 patent drawing
  • US12431138B2 patent drawing

AI summary

Implementations described herein relate to an application and/or automated assistant that can identify arrangement operations to perform for arranging text during speech-to-text operations—without a user having to expressly identify the arrangement operations. In some instances, a user that is dictating a document (e.g., an email, a text message, etc.) can provide a spoken utterance to an application in order to incorporate textual content. However, in some of these instances, certain corresponding arrangements are needed for the textual content in the document. The textual content that is derived from the spoken utterance can be arranged by the application based on an intent, vocalization features, and/or contextual features associated with the spoken utterance and/or a type of the application associated with the document, without the user expressly identifying the corresponding arrangements. In this way, the application can infer content arrangement operations from a spoken utterance that only specifies the textual content.