Multimodal Audio Editor Speech Indexing via ASR Timing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for indexing digitized speech in multimodal digital audio editors are cumbersome, especially in small handheld devices with limited interaction modalities, as they require building grammars and lexicons to restrict vocabulary for accurate speaker-independent voice recognition, limiting user interaction and requiring tedious analysis of audio data for word location.

Innovation Solution

A multimodal digital audio editor operating on a device with multiple interaction modes, coupled with an ASR engine, provides digitized speech for recognition, receives recognized words with timing information, and inserts this information into a speech recognition grammar to enable voice-enabled user interface commands, allowing for efficient indexing and visualization of recognized words within the audio data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If speaker-independent voice recognition is implemented with limited vocabulary, then voice recognition accuracy is improved, but user interaction capability deteriorates

Engineering Contradiction:
Improvevoice recognition accuracyVSAvoiduser interaction capability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent segments the speech recognition process into two distinct phases: an indexing phase where the complete audio stream is processed to identify all word boundaries and create a comprehensive word list with timing information, and a playback phase where only the identified words are presented to the speech recognizer. This segmentation allows the system to achieve both comprehensive vocabulary coverage during indexing while maintaining accurate recognition during playback, resolving the contradiction between vocabulary size and recognition accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies preliminary action by performing word detection and boundary identification on the complete audio stream before the actual speech recognition occurs. The system pre-processes the audio to identify all words, their positions, and durations, creating a ready-to-use word list that guides the subsequent speech recognition process. This preliminary indexing enables the system to handle unlimited vocabulary words while maintaining high recognition accuracy during user interaction.

Inventive Principle:
Principle #10Preliminary action

2Ease of operation

If multimodal access is combined in handheld devices, then ease of operation is improved, but device complexity increases

Engineering Contradiction:
Improveuser interaction easeVSAvoiddevice complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent implements universality by designing a unified speech recognition framework that handles multiple functions through a single integrated system. The same indexing mechanism serves both the speech recognizer and the playback system, while the word list generated during indexing is reused during playback without requiring separate processing. This multi-functional approach enables comprehensive voice interaction capabilities while minimizing redundant components and reducing overall device complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent applies self-service by having the speech recognition system utilize its own output (the word list and timing information generated during indexing) to guide subsequent recognition operations during playback. The system serves itself by using the indexing results to optimize the recognition process, eliminating the need for external intervention or separate processing systems. This self-service mechanism simplifies the overall system architecture while maintaining ease of operation through voice control.

Inventive Principle:
Principle #25Self-service

3Measurement precision

If manual analysis of audio data for word location is performed, then measurement precision is improved, but loss of time increases

Engineering Contradiction:
Improveword location accuracyVSAvoidanalysis time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces the mechanical manual analysis process with an automated computational indexing system. Instead of requiring human analysts to manually examine audio data to identify word locations, the system uses automated speech processing algorithms to detect words, determine their boundaries, and generate timing information. This substitution of manual mechanical analysis with automated computational processing achieves high measurement precision for word location while completely eliminating the time loss associated with manual analysis.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent implements continuity of useful action by performing the indexing operation once on the complete audio stream to establish all word boundaries and timing information, then reusing this pre-computed information throughout the entire playback process. Rather than performing word location analysis continuously or repeatedly, the system performs the indexing action once upfront and then continuously benefits from the results during playback. This approach maintains high measurement precision while minimizing the total time required for analysis.

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentUS9123337B2Indexing digitized speech with words represented in the digitized speech
Publication Date: 2015.09.01 MICROSOFT TECHNOLOGY LICENSING LLC
  • US9123337B2 patent drawing
  • US9123337B2 patent drawing
  • US9123337B2 patent drawing

AI summary

Indexing digitized speech with words represented in the digitized speech, with a multimodal digital audio editor operating on a multimodal device supporting modes of user interaction, the modes of user interaction including a voice mode and one or more non-voice modes, the multimodal digital audio editor operatively coupled to an ASR engine, including providing by the multimodal digital audio editor to the ASR engine digitized speech for recognition; receiving in the multimodal digital audio editor from the ASR engine recognized user speech including a recognized word, also including information indicating where, in the digitized speech, representation of the recognized word begins; and inserting by the multimodal digital audio editor the recognized word, in association with the information indicating where, in the digitized speech, representation of the recognized word begins, into a speech recognition grammar, the speech recognition grammar voice enabling user interface commands of the multimodal digital audio editor.