Real-time machine learning assisted hearing aid

The selective audio amplification method uses a multi-modal data enhancement model to prioritize relevant audio signals in real-time by combining acoustic and visual data, addressing the limitations of traditional hearing aids in noisy environments with reduced latency and computational overhead.

GB2641109APending Publication Date: 2025-11-19THE COURT OF EDINBURGH NAPIER UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
GB2024006998
Authority / Receiving Office
GB · GB
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-17
Publication Date
2025-11-19

AI Technical Summary

Technical Problem

Traditional hearing aids struggle to differentiate and prioritize relevant sounds over unwanted noise, especially in complex environments, and require computationally heavy systems that are not suitable for real-time applications.

Method used

A selective audio amplification method using a real-time multi-modal data enhancement model that combines acoustic and visual data to determine priority audio based on context information, employing a single-headed transformer neural network or a lightweight convolutional recurrent neural network to achieve low latency processing.

Benefits of technology

Enables real-time enhancement of priority audio signals with reduced latency and computational requirements, effectively suppressing unwanted noise and preserving privacy by using low-dimensional visual features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A hearing aid selectively prioritises audio data (eg. one nearby speaker of several in a noisy environment) using visual data from a camera and contextual data (eg. lip movements, gaze, noise levels) input to a real time multi-modal data enhancement model comprising a single-headed transformer artificial neural network with self-attention followed by fully connected layers. The model may switch to a lightweight convolutional recurrent neural network.
Need to check novelty before this filing date? Find Prior Art

Description

Field of the Invention The present invention relates to a hearing aid and more specifically a machine learning assisted hearing aid. Background of the Invention A hearing aid is a device used to amplify audio for a user who may have impaired hearing, for example owing to a medical condition or environmental aspects (e.g., a noisy building site or an airport). A typical hearing aid may amplify audio within a predetermined range, for example audio frequency range, source direction or distance from a user. All audio within the predetermined range will be amplified accordingly and directed to the user’s ear (e.g., via an earpiece comprising a speaker). Although some degree of separation or filtering of unwanted noise is made available, traditional hearing aids are unable to differentiate and prioritize only those sounds which are important to the user e.g., a person speaking in a noisy environment, with relevant and irrelevant sounds being amplified to an undesirable degree. Accordingly, there is a desire to provide enhanced hearing aid devices and methods to enable such differentiation and prioritization of relevant audio. US2023 / 0122905 relates to audio-visual speech separation and discloses an audio-visual speech separation system and method. Said audio-visual speech separation method comprises obtaining, for a video stream, per-frame face embeddings for a face of each speaker present in the video stream to generate visual features of said faces, obtaining a spectrogram of an audio soundtrack for the video, generating a respective spectrogram mask for each of the speakers and determining a respective isolated speech spectrogram for each speaker. US2021 / 0134312A1 relates to audio-visual speech enhancement and discloses a speech enhancement system including a spatio-temporal residual network configured to receive video data comprising a target speaker and extract visual features from the video data, an autoencoder configured to receive input of an audio spectrogram and extract audio features. A mask is configured based on a fusion of said visual features and audio features in a squeeze-excitation fusion block. The mask may be applied to the audio spectrogram to generate an enhanced magnitude spectrogram. 2023 / 0267942A1 relates to an audio visual hearing aid and discloses a method comprising receiving a first indication of one or more first speakers visible by a current view of a camera of a user device, generating a respective isolated speech for the first speaker and sending the isolated speech for the first speaker to a listening device operatively coupled to the user device, and receiving a second indication of one or more second speakers visible by a current view of the camera of the user device, generating a respective isolated speech for the second speaker and sending the isolated speech for the second speaker to the listening device. The known methods and devices may use a combination of audio and visual data to isolate or enhance audio for a user. However, the known methods and devices focus on an audio source (e.g., a speaker) and do not consider information related to the environment of said audio source that is often disregarded as unwanted noise. It may be possible to use such environmental, or context, information, to enhance the identification and isolation of an audio source. Furthermore, state of the art methods may not generalize well to a full range of real-world noisy conversational environments that current hearing-aid users struggle with. In addition, they may be computationally heavy and require large complex computing systems to implement, and thus may not be suitable for real time or point of care applications. Summary of the Invention In accordance with a first aspect of the present disclosure, there is provided a selective audio amplification method, the method comprising: capturing acoustic data; capturing visual data; determining context information from said acoustic data and visual data; determining priority audio data based on the context information, visual data and acoustic data; and selectively amplifying said priority audio data. By determining context information, and determining priority audio based on said context information, said method may be less computationally heavy to execute and thus provide lower latency with respect to the prior art. The methods described herein provide for speech enhancement even when a speaker is not looking directly at a camera (i.e., the device used for capturing said visual data). The method may further comprise determining the priority audio data based on a real-time multi-modal data enhancement model. The real-time multi-modal data enhancement model may comprise a single-headed transformer artificial neural network with self-attention followed by fully connected layers. The real-time multi-modal data enhancement model may further comprise a lightweight convolutional recurrent neural network. The real-time multi-modal data enhancement model may switch between a single-headed transformer neural network and lightweight convolutional recurrent neural network based on the context information. In this way, the multi-modal data enhancement model may be tuned to different types of context which may lead to further reduced latency. For example, a singleheaded transformer network may be better suited to a wide range of environments, whereas a lightweight convolutional recurrent neural network may be a preferred lower latency option when the model is not required to be suited to a wide range of environments. The context information may comprise at least one of speech characteristics, lip movements, lip reading, eye gaze, facial expressions, environmental noise level or type of background noise. The step of determining context information may further comprise determining a priority audio data source embedding extraction model. The step of determining context information may further comprise determining an environmental context estimation model. The method may further comprise extracting an audio feature from the acoustic data, wherein the context information is based in part on the audio feature. The method may further comprise extracting the audio feature from the acoustic data using a convolutional neural network. The method may further comprise extracting a visual featurefrom the visual data, , wherein the context information is based in part on the visual feature. In some examples, the visual feature comprises a visual landmark, wherein the visual landmark is a lip outline of a target speaker. In some examples, the method comprises extracting the visual feature per-frame of the visual data. The method may further comprise extracting the visual feature from the per-frame visual data using a causal fast temporal convolution network Visual features and audio features may be considered low dimensional data, which may be less computationally heavy to process than, for example, raw image data. Furthermore, said audio and visual data may also preserve the privacy of the target speaker by comprising landmarks that may not be used to identify said target speaker (e.g., by using an outline of a target speakers lips rather than a full image of the target speaker). Thus, by utilizing visual features and audio features, lower latency execution may be achieved. Lower latency may be achieved by the use of low dimensional data, and may be further reduced by the use of context information. In this way, real-time audio enhancement may be realized. The step of selectively amplifying priority data may further comprise determining a priority data source based on the priority data source embedding extraction model; and amplifying priority data originating from said priority data source. The method may further comprise: determining an environmental context based on the environmental context estimation model; and suppressing data that does not originate from said priority data source based on said environmental context. The method may further comprise up-sampling said visual data and context information to match a sample rate of the acoustic data. The method may further comprise concatenating the visual data, context information and acoustic data to provide concatenated data, wherein the step of selectively amplifying priority data is based on said concatenated data. In a second aspect of the present disclosure there is provided an electronic selective data amplification system, the system comprising; an audio sensor operable to capture acoustic data; an optical sensor operable to capture visual data; and at least one processor, said processor operable to: determine context information from said acoustic data and visual data; determine priority data based on the context information, visual data and acoustic data; and selectively amplify said priority data. The system may comprise a wearable electronic system. The system may be configured to perform any of the methods described herein. In a third aspect, there is provided a computer program comprising instructions stored thereon which, when the program is executed by a computer, cause the computer to carry out any of the methods described herein. Brief Description of the Drawings Figure 1 depicts schematically a system architecture in accordance with some examples. Figure 2 depicts a method in accordance with some examples. Detailed Description To address the shortcomings of the prior art, the inventors have devised methods and associated systems that may contextually use audio and visual data to improve the quality and intelligibility of audio signals such as speech signals by removing or reducing unwanted audio such as background noise in real time. In contrast to the prior art, the presently proposed methods may be context aware and thus may dynamically determine what audio data to enhance and what audio data to suppress based on a context (e.g., an environmental context). Furthermore, the present solution may be employed independent of hardware constraints. That is, in contrast to models described in the prior art that may require complex and physically large hardware, the presently disclosed method may be implemented in e.g., a smart phone, smart watch, smart glasses, web-based communications technology, ear defenders, hearing aid or any other suitable device. The presently disclosed methods may be implemented in a standalone device that may be integrated into a pre-existing device having audio and optical data capturing abilities, for example a headset, ear defenders, smart glasses, smart watch, web-based communications device, or a hearing aid. Referring now to Figure 1 and Figure 2, there is provided a selective audio amplification method 200. The method may comprise at block 210 capturing acoustic data 120. The acoustic data 120 may comprise noisy audio having a mixture of target audio sources (e.g., a target speaker, or a target object such as a musical instrument) and background noise (e.g., conversation ambience, automotive traffic, air traffic, or any other noise unrelated to the target source that may overlap with the target audio source in time or frequency). The acoustic data may be captured by way of any suitable transducer such as a microphone, microphone array, or any other suitable transducer based on a frequency and / or amplitude of the acoustic data and digitized accordingly for processing. The acoustic data may be acquired using a sample rate based on any of a sample rate of a capturing device or a frequency of the acoustic data. In some examples, a spectrogram may be determined from the acoustic data in order to determine time-frequency information comprised in the acoustic data. Time-frequency information may describe how the frequency content (i.e., any of frequency, magnitude, or power) varies over time. In a specific example, Fourier transform (e.g., a short time Fourier transform) using a Hanning window may be determined from the acoustic data. The power spectral density of the acoustic data may be determined. In this way, the time-frequency information of the acoustic data is made available for processing. At block 230, audio features may be extracted from the acoustic data using an audio feature extraction network 140. In some examples the audio feature extraction network is a convolutional neural network (CNN) that may take the spectrogram of acoustic data and extract local patterns. The spectrogram may be a noisy spectrogram. That is, the spectrogram may comprise target audio (e.g., from a speaker) and background noise (e.g., environmental noise such as traffic, machinery or conversation). For example, the convolutional neural network may identify and extract audio features, for example acoustic data having a particular spectrogram (e.g., variation of frequency content over time). In some examples, raw audio data (e.g., a time series) may be used as input to audio feature extraction network 140. In a non-limiting example wherein the audio feature extraction network is a CNN, the CNN may comprise 2D convolutional layers followed by batch normalization and a parametric rectified linear unit (PReLU). At block 220, visual data 110 may be captured. The video data may comprise video frames comprising one or more target audio sources (e.g., a target speaker, or a target object such as a musical instrument) and background noise sources (e.g., non-target speakers, automotive traffic, air traffic, or any other potential noise source that is not the target source). The visual data may be captured by way of any suitable optical detector and digitized accordingly for processing. The video data may be acquired using a sample rate based on, for example, a sample rate of a capturing device, or any suitable sample rate to facilitate data processing. For example, the visual data may be sampled at, or re-sampled to match, a sample rate of the acoustic data. In some examples, capturing visual data may comprise extracting visual landmarks, for example a lip outline of a target speaker, using an enhanced deep neural network based landmark detection model.. In this way privacy of a target may be preserved, because only visual landmarks of an audio source target are extracted and not a full recognizable image (e.g., outlines of facial features rather than an image of a face). Furthermore, visual landmarks may be less computationally heavy to process than images because they have a relatively lower dimension and thus may lead to a reduced computational latency with respect to full image processing. At block 240, visual features may be extracted from the visual data using a visual feature extraction network 130. The visual feature extraction network may comprise a causal fast temporal convolutional network (TCN) that receives the visual data or extracted visual landmarks as landmark flow features and determines the visual features from these. In a nonlimiting example, the TCN may receive the lip outline of a target speaker as determined by the landmark detection model and may extract lip reading features based on the landmark flow features. Visual features may comprise contextually useful lip-reading features such as lip landmark flow, lip embeddings, thermal lip images or visemes. The feature extraction networks 130 140 described herein may be trained using synthetically generated video data with noisy audio generated from a dataset of video data comprising clean audio of a single speaker combined with real-world noise. Real world noise may comprise competing speech and non-speech noises such as cafeteria background noise, household appliance operational noise, restaurant background noise, public transport background noise and the like. Real world noise may be mixed with clean speech at different signal to noise levels to generate training data. In some examples, synthetically generated video data with noisy audio may be approximately 3 to 6 seconds long. The synthetically generated video data may be resampled to a set frequency (e.g. 30 fps for video and 48 kHz for audio) before further processing. The resampled data is fed to the extraction network 130 140 to generate an estimated output. The estimated output may be compared with clean features to calculate a loss function. The loss may be used to update the weights of the extraction network 130 140. At block 250 the method may comprise determining context information based on the audio features and / or visual features. Context information may be determined by a context extraction network 150 comprising a priority audio source embedding extraction model and / or an environmental context estimation model. The priority audio source embedding extraction model may comprise a CNN comprising dilated convolutional layers with Res2Net and squeeze-and-excitation blocks. The priority audio source embedding extraction model may take the visual features as input and determine priority audio source embeddings. The environmental context estimation model may receive the audio features (e.g., a spectrogram of said audio feature) as input and predict a segmental (i.e., per sample) signal-to-noise ratio (SNR) of the audio feature and the type of noise comprised in the audio feature. An environmental context may be determined based on the estimation of SNR and / or the type of noise. In some examples the environmental context model may comprise a convolutional neural network followed by a long short term memory network. Determining context information may further comprise up-sampling the priority audio source embeddings and the environmental context to match a sample rate of the audio features. The up-sampled priority audio source embeddings, environmental context and audio features may be concatenated across a time dimension to provide the context information. Said concatenating process may be considered a multi-modal fusion. Context information may comprise at least one of speech characteristics, lip movements, lip reading, eye gaze, facial expressions, environmental noise level or type of background noise. At block 260, the method may comprise receiving, by a multi-modal data enhancement model 160, the context information, the audio features and the visual features, and determining, by the multi-modal data enhancement model 160, priority audio data and / or a priority audio data source based on at least one of the context information, the audio feature and the visual feature. In some examples, the multi-modal data enhancement model 160 comprises a single headed transformer artificial neural network. In some examples, the single headed transformer artificial neural network comprises self-attention followed by fully connected layers. In other examples, the multi-modal data enhancement model 160 comprises a lightweight convolutional recurrent neural network. In yet another example, the multi-modal data enhancement model 160 comprises a lightweight convolutional recurrent neural network and a single headed transformer artificial neural network. The multi-modal data enhancement model 160 may switch between the single headed transformer network and the lightweight convolutional recurrent neural network based on the context information. For example, the multi-modal data enhancement model 160 may switch between the single headed transformer network and the lightweight convolutional recurrent neural network based on an environmental noise level or type of background noise. In this way, by dynamically selecting the lightweight convolutional recurrent neural network lower latency processing may be achieved. A lightweight convolutional recurrent neural network may be associated with lower latency processing compared to a single headed transformer network; however, a single headed transformer network may perform better than a lightweight convolutional recurrent neural network across a wider range of environments. Thus, for a narrower range of environments, the multi-modal data enhancement model 160 may select a lightweight convolutional recurrent neural network. In all examples, the multi-modal data enhancement model 160 may be a real-time multi-modal data enhancement model. That is, the multi-modal data-enhancement model may receive audio features, visual features and context information, and determine priority audio data in real time, such that a user may detect no delay between a priority audio data source (e.g., a speaker) and hearing the selectively amplified audio data. That is, the present architecture may exploit low-dimensional data in combination with context information to allow low latency, real time enhancement of priority audio signals. In some examples, priority audio data may be selectively amplified within 10 milliseconds. In some examples, latency may be hardware dependent. At block 270, the method may comprise selectively amplifying said priority audio data. Selectively amplifying priority data may comprise suppressing audio data based on said environmental context and / or the priority data source. That is, audio data that does not originate from said priority data source may be suppressed and / or priority audio data that originates from the priority data source may be amplified. There is further provided an electronic selective audio data amplification system, the system comprising; an audio sensor operable to capture acoustic data; an optical sensor operable to capture visual data; and at least one processor, said processor operable to: determine context information from said acoustic data and visual data; determine priority audio data based on the context information, visual data and acoustic data; and selectively amplify said priority data. The electronic selective audio data amplification system may for example be a laptop or desktop computer system comprising a microphone and a webcam or a wearable camera, optical sensor or the like. In other examples the electronic selective audio data amplification system is a wearable electronic system, for example smart ear defenders, web-based communications technology, or smart glasses comprising a microphone and a video camera. In some examples, the electronic selective audio data amplification system may comprise e.g., a smart watch, smart phone and smart glasses, wherein the audio sensor is comprised in the smart watch or smart phone, and wherein the optical sensor is comprised in the smart glasses. In some examples, the present methods for audio enhancement may be utilised by real-time web-based communications platforms such as videoconferencing, videotelephony platforms and the like. The description provided herein may be directed to specific implementations. It should be understood that the discussion provided herein is provided for the purpose of enabling a person with ordinary skill in the art to make and use any subject matter defined herein by the subject matter of the claims. It should be intended that the subject matter of the claims not be limited to the implementations and illustrations provided herein, but include modified forms of those implementations including portions of implementations and combinations of elements of different implementations in accordance with the claims. It should be appreciated that in the development of any such implementation, as in any engineering or design project, numerous implementation-specific decisions should be made to achieve a developers’ specific goals, such as compliance with system-related and business related constraints, which may vary from one implementation to another. Moreover, it should be appreciated that such a development effort may be complex and time consuming, but would nevertheless be a routine undertaking of design, fabrication, and manufacture for those of ordinary skill having benefit of this disclosure. Reference has been made in detail to various implementations, examples of which are illustrated in the accompanying drawings and figures. In the detailed description, numerous specific details are set forth to provide a thorough understanding of the disclosure provided herein. However, the disclosure provided herein may be practiced without these specific details. In some other instances, well-known methods, procedures, components, circuits, and networks have not been described in detail so as not to unnecessarily obscure details of the embodiments. It should also be understood that, although the terms first, second, etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element could be termed a second element, and, similarly, a second element could be termed a first element. The first element and the second element are both elements, respectively, but they are not to be considered the same element. The terminology used in the description of the disclosure provided herein is for the purpose of describing particular implementations and is not intended to limit the disclosure provided herein. As used in the description of the disclosure provided herein and appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. The term “and / or” as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. The terms “includes,” “including,” “comprises,” and / or “comprising,” when used in this specification, specify a presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. While the foregoing is directed to implementations of various techniques described herein, other and further implementations may be devised in accordance with the disclosure herein, which may be determined by the claims that follow. Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.

Claims

1. A selective audio amplification method, the method comprising:capturing acoustic data;capturing visual data;determining context information from said acoustic data and visual data;determining priority audio data based on the context information, visual data and acoustic data; andselectively amplifying said priority audio data.

2. The method of claim 1, further comprising determining the priority audio data based on a real-time multi-modal data enhancement model.

3. The method of claim 2, wherein the real-time multi-modal data enhancement model comprises a single-headed transformer artificial neural network with self-attention followed by fully connected layers.

4. The method of claim 2, wherein the real-time multi-modal data enhancement model further comprises a lightweight convolutional recurrent neural network.

5. The method of any preceding claim, wherein the real-time multi-modal data enhancement model switches between a single-headed transformer neural network and lightweight convolutional recurrent neural network based on the context information.

6. The method of any previous claim, wherein the context information comprises at least one of speech characteristics, lip movements, lip reading, eye gaze, facial expressions, environmental noise level or type of background noise.

7. The method of any previous claim, wherein the step of determining context information further comprises determining a priority audio data source embedding extraction model.

8. The method of any previous claim, wherein the step of determining context information further comprises determining an environmental context estimation model.

9. The method of any previous claim, further comprising extracting an audio feature from the acoustic data, wherein the context information is based in part on the audio feature.

10. The method of claim 9, further comprising extracting the audio feature from the acoustic data using a convolutional neural network.

11. The method of any previous claim, further comprising extracting a visual feature from the visual data, wherein the context information is based in part on the visual feature.

12. The method of claim 11, wherein the visual feature comprises a visual landmark, wherein the visual landmark is a lip outline of a target speaker.

13. The method of claim 11 or 12, further comprising extracting the visual feature from the visual data using a causal fast temporal convolution network.

14. The method of any of claims 11 to 13, further comprising extracting the visual feature per-frame of the visual data.

15. The method of claim 7, wherein the step of selectively amplifying priority data comprisesdetermining a priority data source based on the priority data source embedding extraction model; andamplifying priority data originating from said priority data source.

16. The method of claim 15, further comprising:determining an environmental context based on the environmental context estimation model; andsuppressing data that does not originate from said priority data source based on said environmental context.

17. The method of any previous claim, further comprising up-sampling said visual data and context information to match a sample rate of the acoustic data.

18. The method of claim 15 further comprising concatenating the visual data, context information and acoustic data to provide concatenated data, wherein the step of selectively amplifying priority data is based on said concatenated data.

19. An electronic selective data amplification system, the system comprising;an audio sensor operable to capture acoustic data;an optical sensor operable to capture visual data; andat least one processor, said processor operable to:determine context information from said acoustic data and visual data;determine priority data based on the context information, visual data and acoustic data; and selectively amplify said priority data.

20. The system of claim 19, wherein the electronic system is a wearable electronic system.

21. The system of any of claims 19 or 20, operable to perform the method of any of claims 1 to 18.

22. A computer program comprising instructions which, when the program is executed by a computer, cause the computer to carry out the method of any of claims 1 to 18.

Citation Information

Patent Citations

  • Audio-visual speech enhancement

    US20210134312A1

  • Audio-visual speech separation

    US20230122905A1

  • Audio-visual hearing aid

    US20230267942A1

  • Contextual audio ducking with situation aware devices

    US20140006026A1

  • Systems and methods for camera and microphone-based device

    US20200296521A1