Speech-to-text preprocessing with temporal keyword indexing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speech-to-text conversion methods are inefficient in non-ideal conditions, such as multi-way telephone conversations or audio/video conferences with background music, resulting in poor accuracy for media indexing and content management systems.

Innovation Solution

A method and system for speech-to-text conversion preprocessing that captures and temporally associates speech audio input with supporting text sources to create an optimized keyword positional index metadata set, prioritizing keywords and using them as input for the speech recognition engine, enhancing the accuracy of text conversion.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional speech to text conversion technology is applied to non-ideal conditions such as multi-way telephone conversations or audio/video conferences with background music, then the system can process diverse audio inputs, but the accuracy of speech converted into text becomes poor and unsatisfactory

Engineering Contradiction:
Improveability to process diverse audio inputsVSAvoidspeech to text conversion accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent segments the audio input into multiple channels or streams, separating speech signals from background noise and music. This segmentation allows the system to process each channel independently, improving accuracy by focusing on speech-specific characteristics while filtering out interfering elements from other channels.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces intermediary processing components including audio preprocessing modules, speech activity detectors, and noise filters that act as mediators between the raw audio input and the speech-to-text conversion engine. These intermediaries prepare the audio signal by removing background music and isolating speech, thereby improving conversion accuracy in non-ideal conditions.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If the language vocabulary set is limited to optimized keywords with temporal locations, then the speech to text conversion accuracy is significantly improved, but the complexity of preprocessing and metadata creation increases

Engineering Contradiction:
Improvespeech to text conversion accuracyVSAvoidpreprocessing and metadata creation complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent performs preliminary actions by pre-processing the audio signal before speech-to-text conversion, including noise filtering, speech activity detection, and keyword extraction. It also creates metadata sets in advance that contain optimized keywords with temporal locations, so that when conversion occurs, the system already has prepared, context-aware vocabulary lists tailored to each audio segment, thereby improving accuracy without requiring complex real-time processing.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent dynamically changes parameters such as vocabulary size, keyword priority weights, and temporal window sizes based on the characteristics of the audio input. By adjusting these parameters according to the specific audio context (e.g., presence of background music, number of speakers), the system optimizes conversion accuracy for each situation without requiring a completely different processing architecture.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS7908141B2Extracting and utilizing metadata to improve accuracy in speech to text conversions
Publication Date: 2011.03.15 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US7908141B2 patent drawing
  • US7908141B2 patent drawing
  • US7908141B2 patent drawing

AI summary

A computer-based system and method for speech to text conversion preprocessing of a presentation with a speech audio, useable in real time. The method captures a presentation speech audio input to be converted into text, temporally associates the speech audio input with at least one supporting text source from the same presentation containing common keywords and creates an optimized and prioritized keyword positional index metadata set for inputting into a speech to text conversion processor.