Touchscreen Transcription Vetting via Gesture Correction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing communication technologies, such as video messaging and VoIP calls, face challenges in quickly and accurately correcting imperfect speech-to-text transcriptions, especially in fast-paced communication sessions, due to the limitations of traditional editing methods on user terminals like smartphones and tablets.
Innovation Solution
A user terminal equipped with a touchscreen interface and a communication client application that allows for fast transcription approval and editing through intuitive gestures, enabling users to vet and correct estimated transcriptions before sending, which can also include editing audio and video content based on transcription errors.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional text editing methods (mouse highlighting and keyboard retyping) are used to correct transcriptions, then correction accuracy can be achieved, but the correction time becomes too long for fast-paced communication sessions
Solution Approach 1:
The patent replaces the mechanical text editing system (mouse highlighting and keyboard retyping) with a voice-based correction system. Users speak the correct text, and the system automatically replaces the incorrect transcription portion, eliminating the need for manual selection and retyping while maintaining correction accuracy.
Solution Approach 2:
The patent introduces an intermediary voice recognition and text processing system between the user and the final transcription. This intermediary automatically processes the user's spoken correction and performs the text replacement, acting as a mediator that speeds up the correction process without requiring direct manual text manipulation.
2Productivity
If a stenographer's keyboard is used for fast transcription editing, then editing speed improves, but the device complexity and skill requirement increase significantly
Solution Approach 1:
The patent implements a self-service correction system where the user simply speaks the correction and the system automatically handles the text replacement. This eliminates the need for specialized stenographer keyboards or skilled operation, making fast correction accessible through a simple voice-based interface.
Solution Approach 2:
The patent substitutes complex mechanical stenographer keyboards with a voice-based input system combined with automatic text processing. This replacement maintains high editing speed while dramatically reducing device complexity and eliminating the need for specialized training.
3Ease of operation
If virtual QWERTYUIOP keyboard is used for transcription correction, then ease of operation is maintained, but the correction speed becomes too slow for fast-paced communication
Solution Approach 1:
The patent replaces the mechanical keyboard input system with a voice-based input system. Users speak corrections naturally without needing to physically type, maintaining ease of operation while dramatically increasing correction speed to match fast-paced communication sessions.
Solution Approach 2:
The patent introduces an intermediary voice processing system that converts spoken corrections into text replacements automatically. This intermediary layer maintains the natural ease of speaking while accelerating the correction process by eliminating manual typing and selection steps.
Data Source
AI summary
A portion of speech is captured when spoken by a near-end user. A near-end user terminal conducts a communication session, over a network, between the near-end user and one or more far-end users, the session including a message sent to the one or more far-end users. A vetting mechanism is provided via a touchscreen user interface of the near-end user terminal, to allow the near-end user to vet an estimated transcription of the portion of speech prior to being sent to the one or more far-end users in the message. According to the vetting mechanism: (i) a first gesture performed by the near-end user through the touchscreen user interface accepts the estimated transcription to be included in a predetermined role in the sent message, while (ii) one or more second gestures performed by the near-end user through the touchscreen user interface each reject the estimated transcription to be sent in the message.


