Selective Speech-to-Text Filtering for Hearing Impaired Users

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech-to-text systems lack selectivity, transcribing and displaying all detected speech regardless of relevance to the user, leading to distractions and decreased ability to focus on personally relevant conversations.

Innovation Solution

Implementing a voice transcription assistant that determines relevance based on criteria such as loudness thresholds and keyword identification, preventing irrelevant speech from being transcribed and displayed to the user.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If the speech-to-text system transcribes all detected speech, then the user receives complete information from the environment, but the user experiences distraction and difficulty focusing on relevant conversations

Engineering Contradiction:
Improvecompleteness of transcribed informationVSAvoiddistraction to the user
Core Design Contradiction:
Loss of informationVSObject-affected harmful factors

Solution Approach 1:

The patent extracts only the relevant speech content from the complete set of detected speech by introducing a relevance determination module. This module filters out irrelevant speech based on criteria such as speaker identification, contextual analysis, and user preferences, thereby eliminating distraction while preserving important information for the user.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies different processing qualities to different speech segments. Relevant speech receives full transcription and display treatment, while irrelevant speech is either filtered out or given reduced processing. This local differentiation allows the system to maintain information completeness where needed while reducing distraction in other areas.

Inventive Principle:
Principle #3Local quality

2Reliability

If the speech-to-text system transcribes all detected speech, then no relevant speech is missed, but the system becomes overwhelmed by overlapping speech from multiple sources

Engineering Contradiction:
Improveaccuracy of speech detectionVSAvoidcomplexity of speech processing
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent extracts only the speech segments that require processing by introducing a relevance determination module. This module filters out irrelevant speech based on criteria such as speaker identification, contextual analysis, and user preferences, thereby reducing the processing load while maintaining accuracy for relevant speech.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent segments the speech processing task into distinct stages: speech detection, relevance determination, and transcription. By dividing the processing pipeline, the system can handle overlapping speech more effectively, processing only relevant segments in detail while managing other speech at a higher level.

Inventive Principle:
Principle #1Segmentation

3Loss of information

If the speech-to-text system displays all transcribed speech, then the user has access to complete communication content, but relevant text becomes less obvious due to distraction from irrelevant content

Engineering Contradiction:
Improvecompleteness of displayed informationVSAvoidvisibility of relevant text
Core Design Contradiction:
Loss of informationVSMeasurement precision

Solution Approach 1:

The patent extracts only the relevant speech content from the complete set of detected speech by introducing a relevance determination module. This module filters out irrelevant speech based on criteria such as speaker identification, contextual analysis, and user preferences, thereby eliminating distraction while preserving important information for the user.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies different processing qualities to different speech segments. Relevant speech receives full transcription and display treatment, while irrelevant speech is either filtered out or given reduced processing. This local differentiation allows the system to maintain information completeness where needed while reducing distraction in other areas.

Inventive Principle:
Principle #3Local quality

4Ease of operation

If the speech-to-text system processes all speech equally, then the treatment is simple and consistent, but the system cannot adapt to different environments and user needs

Engineering Contradiction:
Improvesimplicity of system operationVSAvoidadaptability to different environments
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The patent introduces dynamic adaptability by allowing the relevance determination criteria to change based on the environment and user needs. The system can adjust speaker importance weights, contextual relevance thresholds, and filtering parameters dynamically, enabling it to adapt to different situations while maintaining a relatively simple operational interface for users.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS12334076B2Selective speech-to-text for deaf or severely hearing impaired
Publication Date: 2025.06.17 SONY GROUP CORP
  • US12334076B2 patent drawing
  • US12334076B2 patent drawing
  • US12334076B2 patent drawing

AI summary

A method for providing selectivity to a speech-to-text system operates a voice transcription assistant coupled to the speech-to-text system such that speech detected by the speech-to-text system is not transcribed into text for display to a user unless the voice transcription assistant determines that the detected speech is relevant to the user.Smart glasses, including at least one microphone, a speech-to-text processing unit, a display, and a voice transcription assistant, provide selective speech-to-text functionality. The voice transcription assistant has one or more processors; and logic encoded in one or more non-transitory media for execution by the one or more processors and when executed operable to prevent the speech-to-text processing unit providing text to the display in response to the reception of speech by the microphone unless the logic determines that the received speech is relevant to the user.