Vocal Interaction Segmentation for Speech-to-Text Accuracy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech recognition systems fail to provide accurate and complete transcription of vocal interactions, lacking in both recall and precision, and do not effectively capture the flow of interactions, which hinders analysis and visualization tools' effectiveness.
Innovation Solution
A method and apparatus that segment vocal interactions by combining textual and acoustic features to identify sections such as questions, answers, and non-verbal segments, using a trained model to improve text extraction, analysis, and visualization, and enhance speech-to-text quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If speech recognition systems are used to transcribe vocal interactions, then text output is obtained, but the recall and precision are not 100% accurate
Solution Approach 1:
The patent introduces discourse segments as an intermediary layer between raw audio and final text output. By first segmenting the interaction into structured discourse units (questions, answers, statements) and then performing text extraction within these segments, the system achieves more accurate transcription than direct speech recognition alone
Solution Approach 2:
The patent segments the vocal interaction into distinct discourse segments (questions, answers, statements, non-verbal segments) before text extraction. This segmentation allows the system to process and evaluate text quality segment-by-segment, improving overall transcription accuracy by addressing errors in specific segments rather than treating the entire interaction as one unit
2Ease of operation
If speech to text engines attempt to output syntactically correct sentences, then grammar is improved, but more correct words are lost
Solution Approach 1:
The patent inverts the traditional approach by not forcing syntactically correct output from speech recognition. Instead, it extracts text within discourse segments and uses the segment structure to guide interpretation, preserving more correct words even if the resulting text is less syntactically polished
Solution Approach 2:
The patent changes the evaluation parameter from syntactic correctness to discourse structure accuracy. By evaluating text extraction based on whether it correctly identifies discourse segments (questions, answers, statements) rather than grammatical correctness, the system preserves more accurate word transcription
3Quantity of substance
If full transcription is available, then complete text is obtained, but the interaction flow is not captured
Solution Approach 1:
The patent segments the interaction into structured discourse units (questions, answers, statements, non-verbal segments) to capture the interaction flow. This segmentation transforms complete transcription into organized discourse segments that reveal the conversational structure and flow between participants
Solution Approach 2:
The patent adds a structural dimension to complete transcription by organizing text into discourse segments with specific types (questions, answers, statements). This dimensional transformation converts flat transcription into structured interaction flow data that reveals conversational patterns
4Loss of information
If discourse segmentation is implemented, then interaction flow analysis is improved, but system complexity increases
Solution Approach 1:
The patent uses segmentation to divide the complex task of interaction analysis into manageable discourse segments. By processing interactions segment-by-segment rather than as a whole, the system reduces computational complexity while improving flow understanding
Solution Approach 2:
The patent performs preliminary discourse segmentation before detailed analysis. By pre-segmenting interactions into questions, answers, and statements, the system prepares data in advance for easier analysis, reducing the complexity of subsequent processing steps
Data Source
AI summary
A method and apparatus for analyzing and segmenting a vocal interaction captured in a test audio source, the test audio source captured within an environment. The method and apparatus first use text and acoustic features extracted from the interaction with tagging information, for constructing a model. Then, at production time, text and acoustic features are extracted from the interactions, and by applying the model, tagging information is retrieved for the interaction, enabling analysis, flow visualization or further processing of the interaction.


