Automated Assistant Speech-to-Text Formatting

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users of automated assistants face inefficiencies in formatting and editing textual content generated through speech-to-text operations, as they often need to manually specify formatting and deletion commands, leading to prolonged interaction and resource wastage.

Innovation Solution

An automated assistant that processes spoken utterances to generate content arrangement data, determining user intent and applying formatting features based on context, application type, and prior interactions, allowing for automatic formatting and editing of textual content without explicit user commands.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If users manually specify formatting and deletion commands for speech-to-text content, then formatting precision is improved, but interaction duration and resource consumption increase

Engineering Contradiction:
Improveformatting precisionVSAvoidinteraction duration
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system automatically analyzes the transcribed text and applies appropriate formatting, punctuation, and structure without requiring user commands. The automated assistant detects context, identifies sentence boundaries, and formats content based on learned patterns from the application type and user preferences, allowing the system to serve itself rather than requiring continuous user intervention.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system pre-configures formatting rules and patterns based on the application type (email, text message, document) and user preferences before the user provides speech input. When speech is transcribed, the formatting is already prepared and automatically applied, eliminating the need for users to manually specify formatting commands after transcription.

Inventive Principle:
Principle #10Preliminary action

2Manufacturing precision

If users provide explicit formatting commands for each operation, then formatting accuracy is improved, but device complexity and operation difficulty increase

Engineering Contradiction:
Improveformatting accuracyVSAvoidoperation difficulty
Core Design Contradiction:
Manufacturing precisionVSEase of operation

Solution Approach 1:

The automated assistant autonomously determines formatting requirements by analyzing the transcribed text content, application context, and user preferences. Instead of requiring users to issue commands like 'add comma' or 'create new line,' the system independently identifies when formatting is needed and applies it automatically, making the system self-sufficient.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system implements a universal formatting engine that handles multiple formatting tasks (punctuation, line breaks, indentation, capitalization, structure) through a single automated process. This multi-functional approach replaces the need for multiple separate formatting commands, simplifying user interaction while maintaining comprehensive formatting capability.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Ease of operation

If the system deletes a standard length of text per clear command, then operation simplicity is improved, but information loss increases

Engineering Contradiction:
Improvecommand simplicityVSAvoidtext deletion precision
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

When the user issues a 'clear' or 'delete' command, the automated assistant autonomously analyzes the context to determine the appropriate scope of deletion. The system identifies sentence boundaries, paragraph structures, and semantic units to delete only the intended content rather than a fixed number of characters or words, preventing accidental loss of important information.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system provides contextual feedback when processing deletion commands by analyzing the surrounding text and presenting the user with options for what to delete (e.g., 'Delete this sentence?', 'Delete this paragraph?'). This feedback mechanism ensures that deletion actions match user intent while maintaining simple voice command interaction.

Inventive Principle:
Principle #23Feedback

4Adaptability or versatility

If users switch between audio and separate interfaces for content manipulation, then operation flexibility is improved, but interaction time increases

Engineering Contradiction:
Improveinterface flexibilityVSAvoidswitching time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system merges speech recognition with text manipulation capabilities into a unified interface. Users can perform formatting, editing, and content manipulation operations entirely through voice commands without switching to keyboard or mouse interfaces. The automated assistant integrates multiple functions (transcription, formatting, editing, punctuation) into a single voice-driven workflow.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The automated assistant serves as a universal interface that handles both speech-to-text conversion and subsequent text manipulation through the same voice command channel. This multi-functional approach eliminates the need to switch between audio input mode and separate text editing interfaces, maintaining flexibility while reducing interaction time.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12033637B2Arranging and/or clearing speech-to-text content without a user providing express instructions
Publication Date: 2024.07.09 GOOGLE LLC
  • US12033637B2 patent drawing
  • US12033637B2 patent drawing
  • US12033637B2 patent drawing

AI summary

Implementations described herein relate to an application and/or automated assistant that can identify arrangement operations to perform for arranging text during speech-to-text operations—without a user having to expressly identify the arrangement operations. In some instances, a user that is dictating a document (e.g., an email, a text message, etc.) can provide a spoken utterance to an application in order to incorporate textual content. However, in some of these instances, certain corresponding arrangements are needed for the textual content in the document. The textual content that is derived from the spoken utterance can be arranged by the application based on an intent, vocalization features, and/or contextual features associated with the spoken utterance and/or a type of the application associated with the document, without the user expressly identifying the corresponding arrangements. In this way, the application can infer content arrangement operations from a spoken utterance that only specifies the textual content.