Embedded Neural Codec for Low-Latency Hearable Speech Enhancement
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing hearable devices, such as hearing aids and earphones, rely heavily on Bluetooth connectivity and smartphone resources for enhanced capabilities, but this approach limits the scope of these features due to latency and computational burden, particularly in real-time speech enhancement and automatic speech recognition.
Innovation Solution
An embedded neural codec is integrated within hearable devices, combining speech enhancement and automatic speech recognition modules to encode and decode auditory signals efficiently, reducing latency and offloading processing to smartphones for enhanced functionality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If hearable devices rely on Bluetooth connectivity and smartphone resources for enhanced capabilities, then device complexity is reduced, but latency increases and real-time processing capability deteriorates
Solution Approach 1:
The system segments processing tasks by implementing a hybrid architecture where the hearable device performs critical real-time speech enhancement locally, while transcription and complex analysis are offloaded to the smartphone. This segmentation resolves the contradiction by maintaining low latency for time-sensitive audio processing while reducing overall device complexity through cloud/off-device assistance.
Solution Approach 2:
The patent introduces an intermediary processing architecture where the hearable device's neural codec acts as a mediator between the microphone and the smartphone. The neural codec performs preliminary speech enhancement and compression locally, then transmits processed data to the smartphone for further analysis, thereby reducing latency while sharing computational burden.
2Loss of time
If hearable devices perform real-time speech enhancement and transcription locally, then latency is reduced, but computational burden and energy consumption increase
Solution Approach 1:
The system applies partial action by implementing speech enhancement and compression locally on the hearable device, while deferring transcription and complex NLP tasks to the smartphone. This partial local processing reduces latency for critical audio path while distributing computational burden to avoid excessive energy consumption on the hearable device.
Solution Approach 2:
The computational tasks are segmented into two parts: time-critical speech enhancement and compression handled locally by the hearable device's neural codec, and non-time-critical transcription handled by the smartphone. This segmentation optimizes the trade-off between latency and computational energy consumption.
3Productivity
If an embedded neural codec is integrated within hearable devices, then real-time processing capability is improved, but device complexity increases
Solution Approach 1:
The neural codec is designed with multi-functionality, serving both as a compression encoder and a speech enhancement decoder within the hearable device. This universal approach improves real-time processing capability while minimizing the increase in device complexity by consolidating multiple functions into a single integrated component.
Solution Approach 2:
The patent merges the neural codec functionality directly into the hearable device's signal processing pipeline, combining compression and enhancement operations in an integrated architecture. This merging improves real-time processing capability while managing device complexity through unified hardware/software implementation.
Data Source
AI summary
In one embodiment, an example method for using an embedded neural codec for low latency speech enhancement and automatic speech recognition in hearables includes receiving, at a hearable device, an auditory signal and encoding, at the hearable device, the auditory signal into a compressed vector representation of the auditory signal. The method further includes decoding, at the hearable device using a speech enhancement decoder, the compressed vector representation of the auditory signal into denoised speech outputting, from the hearable device, the denoised speech, and transmitting, from the hearable device, an unprocessed vector representation of the auditory signal to a processing device to cause the processing device to decode, using a speech recognition decoder, the compressed vector representation of the auditory signal.


