Speech Recognition Interface for Streaming Query Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users face difficulties in accessing relevant documents during conversations due to challenges in automated speech recognition accurately determining search terms, leading to inadequate real-time information retrieval solutions.

Innovation Solution

A system that analyzes audio data to produce word hypotheses and displays them on a graphical interface at varying speeds, allowing users to select relevant terms for information retrieval, combining automated speech recognition with manual user guidance to improve accuracy and relevance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If automated speech recognition is used to determine search terms from human speech, then information retrieval can be automated, but recognition errors occur and correct information may be overlooked

Engineering Contradiction:
Improveautomation of information retrievalVSAvoidaccuracy of search term recognition
Core Design Contradiction:
Extent of automationVSMeasurement precision

Solution Approach 1:

The system applies different processing qualities to different speech recognition results. High-confidence matches are processed automatically with high speed, while low-confidence matches are presented to users for manual selection. This local differentiation of processing quality resolves the contradiction by maintaining automation for reliable cases while applying human judgment only where needed.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system incorporates user feedback by allowing users to select or correct speech recognition results in real-time. This feedback loop improves the overall accuracy of search term recognition while maintaining automation for the majority of cases, resolving the contradiction between automation extent and recognition precision.

Inventive Principle:
Principle #23Feedback

2Measurement precision

If multiple word hypotheses are displayed simultaneously to allow user selection, then recognition accuracy improves, but interface clutter increases

Engineering Contradiction:
Improveaccuracy of term selectionVSAvoidvisual clutter of interface
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The interface displays multiple word hypotheses with differentiated visual prominence based on confidence levels. High-confidence matches are displayed prominently and automatically selected, while lower-confidence alternatives are displayed with reduced prominence for user review only if needed. This local quality differentiation reduces visual clutter while maintaining accurate term selection.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system extracts and displays only the most relevant word hypotheses based on confidence thresholds, rather than showing all possible interpretations. This selective extraction reduces interface clutter while preserving the accuracy benefits of multiple hypothesis evaluation by showing only the most plausible options.

Inventive Principle:
Principle #2Taking out (Extraction)

3Loss of time

If speech recognition results are processed in real-time during conversation, then information retrieval keeps pace with discussion, but processing speed requirements increase

Engineering Contradiction:
Improvedelay in information retrievalVSAvoidprocessing speed of speech recognition
Core Design Contradiction:
Loss of timeVSSpeed

Solution Approach 1:

The system processes speech recognition results partially in real-time, immediately handling high-confidence matches with automated extraction and search initiation. Lower-confidence results are processed with less urgency, allowing users to review them after the initial information retrieval. This partial real-time processing reduces the overall speed requirement while maintaining timely information delivery for critical terms.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The speech recognition results are segmented into different confidence levels and processed through different pipelines. High-confidence segments are processed rapidly through automated extraction, while lower-confidence segments are handled with extended timing for user review. This segmentation allows the system to meet real-time requirements for critical information while using more relaxed timing for less certain matches.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11461375B2User interface for streaming spoken query
Publication Date: 2022.10.04 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11461375B2 patent drawing
  • US11461375B2 patent drawing

AI summary

Methods and systems for information retrieval include analyzing audio data to produce word hypotheses. Displaying the word hypotheses in motion at different respective speeds at once across a graphical display. Information is retrieved in accordance with one or more selected terms from the displayed word hypotheses.