Speech to Text Conversion Using Spatial Audio Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for assisting hearing-impaired individuals, such as lip reading, third-party translation, and speech recognition technology, face limitations in noisy environments and require additional resources or training, and are impractical in social gatherings or crowded spaces.
Innovation Solution
A speech conversion system that uses a head-mounted display with a speech conversion program, including beamforming and face detection, to convert audio inputs from the environment into text, displayed in a mixed reality environment, allowing users to focus on specific speakers and receive accurate transcriptions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If speech recognition technology is used to receive and interpret speech, then speech can be visually presented to hearing impaired persons, but accuracy suffers when the speaker does not speak clearly into the microphone or when background noise is excessive
Solution Approach 1:
The system divides the audio environment into multiple directional segments using a microphone array, capturing sound from different spatial locations independently. This allows the system to isolate and process speech from specific directions while filtering out noise from other directions, thereby maintaining high recognition accuracy in noisy environments.
Solution Approach 2:
The system applies different processing qualities to different spatial locations. Speech sources are identified and enhanced with higher processing priority and noise tolerance, while background noise from other directions receives lower processing priority. This local differentiation allows accurate speech recognition even when overall background noise is excessive.
2Ease of operation
If a third party translates speech to sign language or transcribes it to written form, then hearing impaired persons can understand the content, but significant constraints are imposed by requiring a third party to be available
Solution Approach 1:
The system enables hearing impaired persons to independently capture, process, and interpret speech without requiring a third party. The wearable device with microphone array and speech recognition software allows the user to autonomously transcribe and understand spoken content, eliminating the need for external assistance while maintaining ease of operation.
3Reliability
If speech is converted to text in noisy and crowded environments, then hearing impaired persons can understand speech, but conventional speech recognition technology becomes impractical
Solution Approach 1:
The microphone array segments the crowded environment into multiple spatial zones, identifying and isolating speech signals from specific speakers while filtering out noise and speech from other directions. This spatial segmentation enables reliable speech understanding even in crowded environments with multiple noise sources.
Solution Approach 2:
The system dynamically changes processing parameters based on environmental conditions, adjusting noise thresholds, gain settings, and recognition sensitivity levels according to the measured noise floor and speech characteristics. This adaptive parameter adjustment maintains reliable speech understanding across varying environmental conditions from quiet rooms to crowded events.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Embodiments that relate to converting audio inputs from an environment into text are disclosed. For example, in one disclosed embodiment a speech conversion program receives audio inputs from a microphone array of a head-mounted display device. Image data is captured from the environment, and one or more possible faces are detected from image data. Eye-tracking data is used to determine a target face on which a user is focused. A beamforming technique is applied to at least a portion of the audio inputs to identify target audio inputs that are associated with the target face. The target audio inputs are converted into text that is displayed via a transparent display of the head-mounted display device.