Voice Instructed Document Authoring with Grammar Correction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users face challenges in composing electronic documents with correct grammar and sentence structure, especially when trying to convey intended messages through spoken speech, as existing systems require verbatim recitation and often require substantial editing time.

Innovation Solution

A computer-implemented system processes spoken speech to generate a draft electronic document by using machine-learning models for automatic speech recognition, normalization, comprehension, and enhancement, allowing users to dictate intent without precise language, and presenting enhancements for selection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If users recite specific text verbatim into the text processor using speech-to-text, then the text can be transformed into digital data, but users face difficulty in composing correct grammar and sentence structure and require substantial editing time

Engineering Contradiction:
Improvedocument composition speedVSAvoidediting time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent introduces an intermediary system between speech input and text output that includes natural language processing, grammar checking, and suggestion generation components. This intermediary automatically processes the raw speech-to-text output, identifies grammatical errors, and provides corrected suggestions, thereby reducing the manual editing time required while maintaining document composition speed

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system enables self-service by automatically detecting and correcting grammatical errors, suggesting improved phrasing, and providing real-time feedback without requiring user intervention. The text processor autonomously performs grammar checking and offers corrections, allowing users to compose documents quickly with minimal post-editing

Inventive Principle:
Principle #25Self-service

2Ease of operation

If users dictate speech expressing intent rather than specific language, then the system can generate draft text, but the transformation from speech to text requires sophisticated processing models

Engineering Contradiction:
Improvedictation convenienceVSAvoidprocessing model complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent segments the complex speech-to-text transformation process into distinct functional modules: speech recognition, natural language understanding, intent detection, text generation, and grammar checking. Each module handles a specific aspect of the transformation, making the overall complex system manageable and interpretable while maintaining ease of operation for users

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system employs a universal processing framework that handles multiple functions within a unified architecture: speech-to-text conversion, intent recognition, draft generation, and grammar correction all occur through an integrated multi-functional system that adapts to different user needs and contexts

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11941345B2Voice instructed machine authoring of electronic documents
Publication Date: 2024.03.26 SUPERHUMAN PLATFORM INC
  • US11941345B2 patent drawing
  • US11941345B2 patent drawing
  • US11941345B2 patent drawing

AI summary

A computer-implemented process is programmed to process a source input, determine text enhancements, and present the text enhancements to apply to the sentences dictated from the source input. A text processor may use machine-learning models to process an audio input to generate sentences in a presentable format. An audio input can be processed by an automatic speech recognition model to generate electronic text. The electronic text may be used to generate sentence structures using a normalization model. A comprehension model may be used to identify instructions associated with the sentence structures and generate sentences based on the instructions and the sentence structures. An enhancement model may be used to identify enhancements to apply to the sentences. The enhancements may be presented alongside sentences generated by the comprehension model to provide the user an option to select either the enhancements or the sentences.