Speech-to-text preprocessing with temporal keyword indexing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech-to-text conversion methods are inefficient in non-ideal conditions, such as multi-way telephone conversations or audio/video conferences with background music, resulting in poor accuracy for media indexing and content management systems.
Innovation Solution
A method and system for speech-to-text conversion preprocessing that captures and temporally associates speech audio input with supporting text sources to create an optimized keyword positional index metadata set, prioritizing keywords and using them as input for the speech recognition engine, enhancing the accuracy of text conversion.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional speech to text conversion technology is applied to non-ideal conditions such as multi-way telephone conversations or audio/video conferences with background music, then the system can process diverse audio inputs, but the accuracy of speech converted into text becomes poor and unsatisfactory
Solution Approach 1:
The patent segments the audio input into multiple channels or streams, separating speech signals from background noise and music. This segmentation allows the system to process each channel independently, improving accuracy by focusing on speech-specific characteristics while filtering out interfering elements from other channels.
Solution Approach 2:
The patent introduces intermediary processing components including audio preprocessing modules, speech activity detectors, and noise filters that act as mediators between the raw audio input and the speech-to-text conversion engine. These intermediaries prepare the audio signal by removing background music and isolating speech, thereby improving conversion accuracy in non-ideal conditions.
2Measurement precision
If the language vocabulary set is limited to optimized keywords with temporal locations, then the speech to text conversion accuracy is significantly improved, but the complexity of preprocessing and metadata creation increases
Solution Approach 1:
The patent performs preliminary actions by pre-processing the audio signal before speech-to-text conversion, including noise filtering, speech activity detection, and keyword extraction. It also creates metadata sets in advance that contain optimized keywords with temporal locations, so that when conversion occurs, the system already has prepared, context-aware vocabulary lists tailored to each audio segment, thereby improving accuracy without requiring complex real-time processing.
Solution Approach 2:
The patent dynamically changes parameters such as vocabulary size, keyword priority weights, and temporal window sizes based on the characteristics of the audio input. By adjusting these parameters according to the specific audio context (e.g., presence of background music, number of speakers), the system optimizes conversion accuracy for each situation without requiring a completely different processing architecture.
Data Source
AI summary
A computer-based system and method for speech to text conversion preprocessing of a presentation with a speech audio, useable in real time. The method captures a presentation speech audio input to be converted into text, temporally associates the speech audio input with at least one supporting text source from the same presentation containing common keywords and creates an optimized and prioritized keyword positional index metadata set for inputting into a speech to text conversion processor.


