An audio processing integration method and device for a hearing aid

Through multi-level network processing technology, including feature extraction, noise suppression, adaptive enhancement and feature fusion, the problem that hearing aids are difficult to effectively process voice signals in complex noise environments is solved, and higher quality audio processing effects are achieved, and users' auditory experience is improved.

CN119629561BActive Publication Date: 2025-05-27BOYIN HEARING TECH (SHANGHAI) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510150895.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-27
Estimated Expiration
2045-02-11

AI Technical Summary

Technical Problem

The existing hearing aid audio processing technology cannot effectively respond to the voice extraction and enhancement needs in complex noise environments, and it is difficult to adaptively adjust according to individual user differences, resulting in poor user auditory experience.

Method used

Multi-level network processing methods are adopted, including feature extraction network, noise suppression network, adaptive enhancement network, cross attention network and audio enhancement network, and the audio enhancement network are used to process the audio signals collected by the hearing aid microphone in real time. Through technical means such as multi-scale feature extraction, noise suppression, adaptive enhancement and feature fusion, the quality of the audio signal is optimized.

Benefits of technology

Effectively reduce the interference of environmental noise on voice signals, improve the clarity and intelligibility of voice signals, provide a clearer and more natural auditory experience, and significantly improve the adaptability and stability of hearing aids in different scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119629561B_ABST
    Figure CN119629561B_ABST
Patent Text Reader

Abstract

The present invention provides an audio processing integration method and device for a hearing aid. The method includes: acquiring in real time an audio signal collected by a hearing aid microphone; processing the audio signal through a feature extraction network to obtain first audio feature data; performing frequency domain conversion, noise distribution prediction and weighted filtering processing on the first audio feature data based on a noise suppression network to obtain first noise-reduced audio features; processing the first audio feature data by using an adaptive enhancement network, and obtaining second enhanced audio features by randomly selecting frequency reference points, collecting feature points and processing; fusing the first noise-reduced audio features and the second enhanced audio features through a cross-attention network to obtain fused audio features; and finally processing the fused audio features by using an audio enhancement network to obtain an enhanced audio signal and outputting it. The present invention can improve the speech clarity and sound quality performance of a hearing aid in a complex environment and improve the auditory experience of hearing-impaired users.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of hearing aid audio processing. More specifically, the present invention relates to an audio processing integration method and device for hearing aids. Background Art

[0002] In modern hearing aid technology, audio processing is a key link in enhancing the auditory experience of hearing-impaired users. Existing hearing aid audio processing technologies mainly rely on traditional noise reduction algorithms and simple audio enhancement methods. Although these methods can improve speech clarity to a certain extent, their effects are often limited in complex noise environments. Traditional noise reduction technologies usually rely on fixed filters or simple spectral processing, making it difficult to adapt to diverse noise scenarios and individual hearing needs. At the same time, most audio enhancement methods adopt a single amplification strategy, unable to effectively distinguish speech from noise, resulting in a decline in speech quality and a poor auditory experience for users.

[0003] In the process of implementing the embodiments of the present invention, the inventors found that there are at least the following problems or defects in the prior art: Existing audio processing technologies cannot effectively meet the requirements of speech extraction and enhancement in complex noise environments, and it is difficult to make adaptive adjustments according to individual differences of users, resulting in a significant reduction in the actual use effect of hearing aids and being unable to meet the needs of hearing-impaired users for a high-quality auditory experience. Summary of the Invention

[0004] The present invention provides an audio processing integration method and device for hearing aids.

[0005] In the first aspect of the present invention, an audio processing integration method for hearing aids is provided, including:

[0006] Real-time acquisition of audio signals collected by a microphone deployed in a hearing aid;

[0007] Processing the audio signals based on a feature extraction network to obtain first audio feature data;

[0008] Processing the first audio feature data based on a noise suppression network to obtain a first noise-reduced audio feature, including: presetting a noise intensity distribution threshold based on the installation position of the microphone; converting the first audio feature data to the frequency domain; predicting the noise distribution corresponding to the converted first audio feature data based on the noise intensity distribution threshold by using the noise suppression network; performing weighted, filtering, and minimum pooling aggregation processing on the converted first audio feature data and the noise distribution to obtain a first noise-reduced audio feature;

[0009] Process the first audio feature data based on an adaptive enhancement network to obtain a second enhanced audio feature, including: preset an audio frequency spectrum space, and convert each frequency point in the audio frequency spectrum space to the frequency domain; randomly select multiple frequency reference points above the converted audio frequency spectrum space, and project them to the time domain to obtain projection points; collect feature points around the projection points in the first audio feature data; process the feature points through the adaptive enhancement network to determine the eigenvalue corresponding to each projection point; determine the second enhanced audio feature based on the eigenvalue;

[0010] Based on a cross-attention network, fuse the first noise-reduced audio feature and the second enhanced audio feature to obtain a fused audio feature, including: perform linear transformations on the first noise-reduced audio feature and the second enhanced audio feature respectively to obtain corresponding queries, keys, and values, where the first noise-reduced audio feature corresponds to the first query, the first key, and the first value, and the second enhanced audio feature corresponds to the second query, the second key, and the second value; calculate a first attention matrix based on the first key and the second query; calculate a second attention matrix based on the first query and the second key; process the first value and the second value based on the first attention matrix and the second attention matrix to obtain the fused audio feature;

[0011] Process the fused audio feature using an audio enhancement network to obtain an enhanced audio signal. The audio enhancement network sequentially includes multiple filtering layers, inverse filtering layers, upsampling layers, and frame-by-frame classification layers.

[0012] Further, process the audio signal based on a feature extraction network to obtain the first audio feature data, including: process the audio signal using a WaveNet neural network to obtain multi-scale audio features, and the multi-scale audio features include audio feature data with different resolutions;

[0013] Fuse the multi-scale audio features in ascending order of resolution, including: perform 1×1 convolution and upsampling on the audio feature data with the first resolution, and fuse the processing result with the result of 1×1 convolution of the audio feature data with the second resolution;

[0014] The first resolution is the audio feature data with a lower resolution in the multi-scale audio features; the second resolution is the audio feature data with a higher resolution in the multi-scale audio features;

[0015] Use a convolutional neural network to aggregate the fusion result of the multi-scale audio features to obtain the first audio feature data.

[0016] Further, weighted, filtered, and minimum pooling aggregation processing is performed on the converted first audio feature data and the noise distribution to obtain the first noise-reduced audio feature, including: obtaining the adjusted audio feature by weighting the converted first audio feature data with the noise distribution; filtering the adjusted audio feature through a filter; dividing the filtered audio feature into multiple frequency bands, aggregating the audio features belonging to the same frequency band, and performing minimum pooling on the aggregation result to obtain the first noise-reduced audio feature.

[0017] Further, the feature points are processed by an adaptive enhancement network to determine the eigenvalue corresponding to each projection point, including: processing the feature points by an adaptive enhancement network to determine the corresponding weight value; performing weighted summation on the feature points corresponding to each projection point to obtain the corresponding eigenvalue, and determining the position of the eigenvalue in the audio spectrum space.

[0018] Further, the fused audio feature is processed by an audio enhancement network to obtain an enhanced audio signal, including: performing multi-layer filtering on the fused audio feature to obtain the first data;

[0019] performing an inverse filtering layer process on the first data to obtain the second data;

[0020] performing upsampling on the second data to obtain the third data with the same resolution as the fused audio feature;

[0021] performing frame-by-frame classification on the third data to obtain the enhanced audio signal.

[0022] Further, it also includes: constructing an audio processing model, where the audio processing model includes a feature extraction network, a noise suppression network, an adaptive enhancement network, a cross-attention network, and an audio enhancement network;

[0023] Collecting multiple audio data and annotating the collected audio data, including: starting from the starting point of the speech segment in the audio data and extending along the speech segment in the audio data until a complete speech annotation is formed;

[0024] Inputting the annotated audio data into the audio processing model for training to adjust the variable parameters in the audio processing model.

[0025] Further, the mean square error loss function is used to optimize the audio processing model.

[0026] Further, based on the first attention matrix and the second attention matrix, the first value and the second value are processed to obtain the fused audio feature, including:

[0027] Fused audio feature The calculation formula is:

[0028]

[0029] Among them, represents the second query, represents the first key, represents the first value, represents the first query and the second key, represents the second value, represents the normalization function.

[0030] Furthermore, the expression of the mean square error loss function is:

[0031]

[0032] Among them, represents the loss value, represents the number of samples in the audio data, represents the sample true value, represents the sample predicted value.

[0033] In the second aspect of the present invention, an audio processing integration device for a hearing aid is provided, including:

[0034] An acquisition module that acquires in real time the audio signal collected by a microphone deployed in the hearing aid;

[0035] A feature extraction module that processes the audio signal acquired by the acquisition module based on a feature extraction network to obtain first audio feature data;

[0036] A first processing module that processes the first audio feature data based on a noise suppression network to obtain a first noise-reduced audio feature;

[0037] A second processing module that processes the first audio feature data based on an adaptive enhancement network to obtain a second enhanced audio feature;

[0038] A feature fusion module that fuses the first noise-reduced audio feature output by the first processing module and the second enhanced audio feature output by the second processing module based on a cross-attention network to obtain a fused audio feature;

[0039] A third processing module that processes the fused audio feature output by the feature fusion module by using an audio enhancement network to obtain an enhanced audio signal, and the audio enhancement network sequentially includes a plurality of filtering layers, an inverse filtering layer, an upsampling layer, and a frame-by-frame classification layer;

[0040] An output module that outputs the enhanced audio signal to the speaker of the hearing aid.

[0041] The above embodiments of the present invention have at least the following beneficial effects: The hearing aid audio processing integration method and device of the present invention can effectively reduce the interference of environmental noise on speech signals and enhance the clarity and intelligibility of speech signals through multi-level network processing. The feature extraction network can obtain multi-scale audio features, providing a rich information basis for subsequent processing; the noise suppression network accurately predicts and suppresses noise based on a preset noise intensity distribution threshold to ensure the purity of speech signals; the adaptive enhancement network further improves the detailed performance of speech signals by randomly selecting frequency reference points and feature points for processing, thereby providing a clearer and more natural auditory experience for users in complex noise environments.

[0042] In addition, the cross-attention network fuses the noise-reduced and enhanced audio features, enabling more accurate feature extraction and fusion. The audio enhancement network further optimizes the quality of audio signals through multi-layer filtering, inverse filtering, upsampling, and frame-by-frame classification processing. This integration method can significantly improve the adaptability and stability of hearing aids in different scenarios, meet the needs of hearing-impaired users for high-quality audio processing, and provide new technical ideas and solutions for the development of hearing aid technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] By referring to the detailed description below with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present invention will become readily understood. In the drawings, several embodiments of the present invention are shown by way of illustration and not limitation, wherein:

[0044] Figure 1 is a flowchart showing the audio processing integration method of a hearing aid provided by an embodiment of the present invention;

[0045] Figure 2 is a schematic structural diagram of the audio processing integration device of a hearing aid provided by an embodiment of the present invention;

[0046] Figure 3 schematically shows a schematic structural diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] The principles and spirit of the present invention will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and implement the present invention, and do not limit the scope of the present invention in any way. On the contrary, these embodiments are provided to make the present invention more thorough and complete, and to be able to convey the scope of the present invention fully to those skilled in the art.

[0048] Those skilled in the art know that the embodiments of the present invention can be implemented as a system, a device, equipment, a method, or a computer program product. Therefore, the present invention can be specifically implemented in the following forms, namely: all hardware, all software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.

[0049] It should be noted that any number of elements in the drawings is for illustration rather than limitation, and any naming is only for distinction and does not have any limiting meaning.

[0050] The following refers to Figure 1 , Figure 1 which is a schematic flowchart of an audio processing integration method for a hearing aid provided in an embodiment of the present invention. As Figure 1 shown, an audio processing integration method 100 for a hearing aid includes:

[0051] Step 101, obtaining in real time an audio signal collected by a microphone deployed in the hearing aid;

[0052] Step 102, processing the audio signal based on a feature extraction network to obtain first audio feature data;

[0053] Step 103, processing the first audio feature data based on a noise suppression network to obtain a first noise-reduced audio feature, including: presetting a noise intensity distribution threshold based on the installation position of the microphone; converting the first audio feature data to the frequency domain; predicting the noise distribution corresponding to the converted first audio feature data based on the noise intensity distribution threshold by using the noise suppression network; performing weighted, filtering, and minimum pooling aggregation processing on the converted first audio feature data and the noise distribution to obtain the first noise-reduced audio feature;

[0054] Step 104, processing the first audio feature data based on an adaptive enhancement network to obtain a second enhanced audio feature, including: presetting an audio spectrum space and converting each frequency point in the audio spectrum space to the frequency domain; randomly selecting a plurality of frequency reference points above the converted audio spectrum space and projecting them to the time domain to obtain projection points; collecting feature points around the projection points in the first audio feature data; processing the feature points through the adaptive enhancement network to determine the eigenvalue corresponding to each projection point; determining the second enhanced audio feature based on the eigenvalue;

[0055] Step 105: Based on the cross-attention network, fuse the first noise-reduced audio feature and the second enhanced audio feature to obtain a fused audio feature, including: applying linear transformations to the first noise-reduced audio feature and the second enhanced audio feature respectively to obtain corresponding queries, keys, and values, where the first noise-reduced audio feature corresponds to the first query, the first key, and the first value, and the second enhanced audio feature corresponds to the second query, the second key, and the second value; calculating a first attention matrix based on the first key and the second query; calculating a second attention matrix based on the first query and the second key; processing the first value and the second value based on the first attention matrix and the second attention matrix to obtain the fused audio feature;

[0056] Step 106: Process the fused audio feature using an audio enhancement network to obtain an enhanced audio signal, where the audio enhancement network sequentially includes a plurality of filtering layers, an inverse filtering layer, an upsampling layer, and a frame-by-frame classification layer.

[0057] It should be noted that the present invention relates to an integrated method for audio processing of a hearing aid. The core lies in real-time processing of the audio signal collected by the hearing aid microphone through a series of network processing steps. First, the audio signal collected by the microphone deployed in the hearing aid is obtained in real time. Here, the audio signal refers to all sound information captured by the hearing aid microphone during use, including speech and environmental noise. Then, the audio signal is processed based on a feature extraction network to obtain the first audio feature data. The feature extraction network is a neural network structure based on deep learning, and its purpose is to extract useful feature information from the original audio signal for subsequent processing. After that, the first audio feature data is processed based on a noise suppression network to obtain the first noise-reduced audio feature. The noise suppression network realizes effective noise suppression by presetting a noise intensity distribution threshold and combining frequency domain conversion and noise distribution prediction. At the same time, the first audio feature data is processed by an adaptive enhancement network to obtain the second enhanced audio feature. This process randomly selects frequency reference points and projects them into the time domain, and weights the feature points to enhance the details of the audio signal. Finally, the first noise-reduced audio feature and the second enhanced audio feature are fused through a cross-attention network to obtain the fused audio feature, and the fused audio feature is processed using an audio enhancement network to obtain the enhanced audio signal. The audio enhancement network includes a plurality of filtering layers, an inverse filtering layer, an upsampling layer, and a frame-by-frame classification layer, which are used to further optimize the quality of the audio signal.

[0058] Specifically, the feature extraction network can adopt the WaveNet neural network structure, which can process audio signals and generate multi-scale audio features, including audio feature data with different resolutions. The fusion process of multi-scale audio features is achieved by performing 1×1 convolution and upsampling on the low-resolution audio feature data, and then fusing it with the high-resolution audio feature data. This fusion method can retain feature information at different scales and provide a more comprehensive feature basis for subsequent noise suppression and audio enhancement. In the noise suppression network, the preset noise intensity distribution threshold is pre-set according to the installation position and usage scenario of the microphone to distinguish between noise and speech signals. Frequency domain conversion is to convert the audio feature data from the time domain to the frequency domain for more effective noise distribution prediction and processing. In the adaptive enhancement network, the audio spectrum space refers to the distribution range of the audio signal in the frequency domain, and the randomly selected frequency reference points are used to determine the key frequency regions in the audio signal that need to be enhanced. By collecting feature points around these frequency points and performing weighted processing, the detailed performance of the audio signal can be enhanced. The cross-attention network calculates the attention matrix to dynamically adjust the weights of the fused audio features, thereby achieving more accurate feature fusion. In the audio enhancement network, the filtering layer, inverse filtering layer, upsampling layer, and frame-by-frame classification layer are respectively used to further optimize the spectral characteristics, resolution, and stability of the audio signal, and finally output a clear enhanced audio signal.

[0059] Preferably, in the feature extraction network, the fusion of multi-scale audio features can be achieved through multiple iterations of 1×1 convolution and upsampling operations to ensure smooth transition and information complementarity between features with different resolutions. In the noise suppression network, the noise intensity distribution threshold can be dynamically adjusted according to the actual usage scenario. For example, the threshold can be increased in a noisy environment to more effectively suppress noise. In the adaptive enhancement network, the selection of frequency reference points can be optimized based on the spectral characteristics of the audio signal. For example, the energy concentration regions of speech signals are preferentially selected. In addition, a normalization function can be introduced into the calculation of the attention matrix in the cross-attention network to enhance the stability and consistency of the fused audio features. In the audio enhancement network, the frame-by-frame classification layer can be optimized according to the temporal characteristics of the audio signal. For example, the long short-term memory network (LSTM) structure can be adopted to better process the dynamic changes of the audio signal.

[0060] In some embodiments, processing the audio signal based on the feature extraction network to obtain first audio feature data includes: processing the audio signal using the WaveNet neural network to obtain multi-scale audio features, where the multi-scale audio features include audio feature data with different resolutions;

[0061] Fuse the multi-scale audio features in ascending order of resolution, including: performing 1×1 convolution and upsampling on the audio feature data of the first resolution, and fusing the processing result with the result of 1×1 convolution of the audio feature data of the second resolution;

[0062] The first resolution is the audio feature data with a lower resolution in the multi-scale audio features; the second resolution is the audio feature data with a higher resolution in the multi-scale audio features;

[0063] Use a convolutional neural network to aggregate the multi-scale audio feature fusion result to obtain the first audio feature data.

[0064] It should be noted that the feature extraction network mentioned in the present invention is one of the core components of the audio processing method, and its function is to convert the audio signal collected by the microphone into data with multi-scale features. Here, the multi-scale audio features refer to the feature representations of the audio signal at different resolutions, and these features can capture the details and structural information in the audio signal. By using the Wave Net neural network to process the audio signal, multi-scale audio features can be obtained, and these features include audio feature data of different resolutions. Wave Net is a neural network architecture based on deep learning, which can generate high-quality audio features and is suitable for audio processing tasks. Subsequently, by fusing the multi-scale audio features in ascending order of resolution, a richer and more complete audio feature representation can be obtained, providing a basis for subsequent noise suppression and audio enhancement.

[0065] Specifically, the Wave Net neural network can process the audio signal frame by frame through its unique structure to generate audio feature data containing different resolutions. In the process of multi-scale audio feature fusion, first perform 1×1 convolution on the audio feature data of low resolution to reduce the dimension of the features and retain key information. Subsequently, adjust the size of the low-resolution features to be the same as that of the high-resolution features through upsampling, and then fuse them with the high-resolution features. This fusion method can effectively retain the detail information of the audio signal at different scales. For example, the low-resolution features can capture the global structure of the audio signal, while the high-resolution features can retain local details. In the fusion process, the parameters of the 1×1 convolution operation can be adjusted according to the characteristics of the audio signal to optimize the effect of feature fusion. In addition, the fused audio feature data can be further aggregated through a convolutional neural network to obtain a more compact and effective feature representation, providing high-quality input for subsequent audio processing steps.

[0066] Preferably, the number of layers and parameters of the Wave Net neural network can be adjusted according to the complexity of the audio signal. For example, in a high-noise environment, the number of network layers can be increased to improve the accuracy of feature extraction. During the multi-scale audio feature fusion process, an attention mechanism can be introduced to enable the network to automatically learn the weight relationship between features of different resolutions, thereby more effectively fusing features. In addition, in addition to 1×1 convolution and upsampling operations, more complex feature fusion strategies such as residual connections or dense connections can be adopted to enhance the information interaction between features. In the convolutional neural network aggregation stage, a deeper network structure can be used or the Batch Normalization technique can be introduced to improve the stability and efficiency of feature extraction.

[0067] In some embodiments, weighted, filtered, and minimum pooling aggregation processing is performed on the converted first audio feature data and the noise distribution to obtain a first noise-reduced audio feature, including: weighting the converted first audio feature data and the noise distribution to obtain an adjusted audio feature; filtering the adjusted audio feature through a filter; dividing the filtered audio feature into multiple frequency bands, aggregating the audio features belonging to the same frequency band, and performing minimum pooling on the aggregation result to obtain a first noise-reduced audio feature.

[0068] It should be noted that the steps of weighted, filtered, and minimum pooling aggregation processing of the converted first audio feature data and the noise distribution in the present invention are the key links to achieve the noise reduction effect. The weighted processing here refers to adjusting the weight of the audio feature according to the characteristics of the noise distribution and the audio feature data to reduce the influence of noise. The filtering process is to smooth the audio feature through a specific filter to remove high-frequency noise components. The minimum pooling aggregation processing is to divide the audio feature into multiple frequency bands, aggregate the features within each frequency band, and take the minimum value to further suppress noise. Through this series of operations, the noise-reduced audio feature can be effectively extracted, providing a cleaner signal basis for subsequent audio processing.

[0069] Specifically, the weighting process is to perform a weighting operation on the converted first audio feature data and the noise distribution to obtain the adjusted audio features. The weights here can be dynamically adjusted according to the noise intensity distribution threshold to ensure effective noise suppression in different noise environments. The filtering process is to process the adjusted audio features by designing a suitable filter, such as a low-pass filter or a band-pass filter, to remove the noise components. The minimum pooling aggregation process is to divide the filtered audio features by frequency bands. For example, the audio features can be divided into low-frequency, mid-frequency, and high-frequency bands, and then the features within each band are aggregated and the minimum value is taken. This processing method can effectively reduce the impact of noise on the audio signal while retaining the key information of the speech signal. In practical applications, the parameters of the filter can be optimized according to the characteristics of the audio signal, such as adjusting the cut-off frequency of the filter to adapt to different noise environments.

[0070] Preferably, an adaptive weight adjustment mechanism can be introduced in the weighting process to dynamically adjust the weights according to the real-time detected noise intensity, thereby improving the adaptability of the noise reduction effect. In the filtering process, a multi-stage filter structure can be adopted. For example, first, a low-pass filter is used to remove high-frequency noise, and then a band-pass filter is used to further optimize the spectral characteristics of the audio signal. In addition, a threshold judgment mechanism can be introduced in the minimum pooling aggregation process. When the minimum value within a frequency band is lower than a preset threshold, it is considered that the frequency band mainly contains noise and can be directly set to zero, thereby further improving the noise reduction effect. At the same time, other pooling methods, such as average pooling or max pooling, can also be considered, and a suitable pooling strategy can be selected according to different application scenarios to achieve the best noise reduction effect.

[0071] In some embodiments, the feature points are processed by an adaptive enhancement network to determine the eigenvalue corresponding to each projection point, including: processing the feature points by the adaptive enhancement network to determine the corresponding weight value; performing weighted summation on the feature points corresponding to each projection point to obtain the corresponding eigenvalue, and determining the position of the eigenvalue in the audio spectrum space.

[0072] It should be noted that the step of processing the feature points by the adaptive enhancement network to determine the eigenvalue corresponding to each projection point in the present invention is a key link in the audio enhancement process. The role of the adaptive enhancement network is to dynamically adjust the weights of the eigenvalues according to the distribution of the feature points in the audio spectrum space, thereby enhancing the detailed performance of the audio signal. Here, the feature points refer to the representative frequency points in the audio spectrum space, and the projection points are the positions after projecting these frequency points into the time domain. By processing the feature points by the adaptive enhancement network, the eigenvalue corresponding to each projection point can be determined, and its position in the audio spectrum space can be further determined, thereby realizing the enhancement of the audio signal.

[0073] Specifically, the processing of feature points by the adaptive enhancement network includes the determination of weight values and the weighted summation of feature values. The determination of weight values is achieved by the network learning the spectral characteristics of the audio signal and assigning a weight to each feature point to reflect its importance in the audio signal. The weighted summation of feature values is to sum the values of each feature point multiplied by its corresponding weight to obtain the final feature value of the projection point. This process can be realized through the training of a neural network. The network automatically adjusts the weight values by learning a large amount of audio data to achieve the best enhancement effect. In the audio spectrum space, the selection of frequency reference points can be based on the energy distribution of the audio signal or a specific frequency range. For example, speech signals are usually concentrated in the low-frequency and mid-frequency regions, so the frequency points in these regions can be preferentially selected as reference points. In this way, the adaptive enhancement network can effectively enhance the details of the audio signal while avoiding over-enhancement of noise.

[0074] Preferably, the weight values in the adaptive enhancement network can be further optimized by introducing normalization processing to ensure a more reasonable distribution of weight values and avoid audio distortion caused by overly large or small weights. In the selection of feature points, a dynamic selection mechanism can be adopted to dynamically adjust the position and number of frequency reference points according to the real-time characteristics of the audio signal to better meet the audio enhancement requirements in different scenarios. In addition, for the calculation of feature values, a non-linear activation function such as ReLU or Sigmoid can be introduced to enhance the non-linear fitting ability of the network and further improve the enhancement effect of the audio signal. At the same time, it can also be considered to combine the adaptive enhancement network with other audio processing technologies, such as working in cooperation with a noise suppression network to achieve more comprehensive audio optimization.

[0075] In some embodiments, processing the fused audio features through an audio enhancement network to obtain an enhanced audio signal includes: performing multi-layer filtering on the fused audio features to obtain first data;

[0076] Performing an inverse filtering layer process on the first data to obtain second data;

[0077] Upsampling the second data to obtain third data with the same resolution as the fused audio features;

[0078] Performing frame-by-frame classification on the third data to obtain the enhanced audio signal.

[0079] It should be noted that the step of processing the fused audio features through the audio enhancement network to obtain the enhanced audio signal in the present invention is the final output link of the entire audio processing flow. The role of the audio enhancement network is to further optimize the quality of the fused audio features through a series of processing operations to make it closer to the ideal audio signal. This process involves steps such as multi-layer filtering processing, inverse filtering layer processing, upsampling, and frame-by-frame classification, aiming to improve the clarity, stability, and naturalness of the audio signal, and finally output a high-quality enhanced audio signal to meet the hearing needs of hearing aid users in different environments.

[0080] Specifically, the multi-layer filtering processing in the audio enhancement network optimizes the fused audio features layer by layer by designing filters with different parameters. The types of filters can include low-pass filters, high-pass filters, or band-pass filters, and their parameters such as cut-off frequency and filter order can be adjusted according to the characteristics of the audio signal. The inverse filtering layer processing removes possible filtering distortions through inverse filtering operations to restore the original characteristics of the audio signal. The upsampling operation is used to increase the resolution of the audio signal to be the same as the input signal to ensure the integrity of the output signal. The frame-by-frame classification layer analyzes the processed audio signal frame by frame to determine whether each frame belongs to a speech signal, thereby further optimizing the clarity of the speech segment. In practical applications, the frame-by-frame classification layer can adopt deep learning models such as recurrent neural networks (RNNs) or convolutional neural networks (CNNs) to improve the accuracy of classification.

[0081] Preferably, the filtering layer in the audio enhancement network can adopt a multi-stage filtering structure. For example, first, a low-pass filter is used to remove high-frequency noise, and then a band-pass filter is used to optimize the spectral range of the speech signal. The parameters of the inverse filtering layer can be dynamically adjusted according to the actual effect of the filtering layer to achieve the best distortion compensation. In the upsampling process, interpolation algorithms such as linear interpolation or spline interpolation can be introduced to improve the quality of the upsampled audio signal. For the frame-by-frame classification layer, an attention mechanism can be introduced to make the network pay more attention to the key frames of the speech signal, thereby further enhancing the speech enhancement effect. In addition, the audio enhancement network can also combine a user feedback mechanism to dynamically adjust the network parameters according to the hearing needs of the user to achieve personalized audio enhancement effects.

[0082] In some embodiments, it further includes: constructing an audio processing model, where the audio processing model includes a feature extraction network, a noise suppression network, an adaptive enhancement network, a cross-attention network, and an audio enhancement network;

[0083] Collecting a plurality of audio data and annotating the collected audio data, including: starting from the starting point of the speech segment in the audio data and extending along the speech segment in the audio data until a complete speech annotation is formed;

[0084] Input the labeled audio data into the audio processing model for training to adjust the variable parameters in the audio processing model.

[0085] It should be noted that the process of constructing and training the audio processing model in the present invention is a fundamental link for achieving high-quality audio processing. The audio processing model is a complex system consisting of multiple sub-networks, including a feature extraction network, a noise suppression network, an adaptive enhancement network, a cross-attention network, and an audio enhancement network. These sub-networks work together to achieve noise reduction, enhancement, and optimization of audio signals. During the model training process, a large amount of audio data needs to be collected and labeled. The labeling process starts from the starting point of the speech segment in the audio data and extends along the speech segment until a complete speech annotation is formed. By inputting the labeled audio data into the model for training, the variable parameters in the model can be adjusted, thereby optimizing the audio processing effect.

[0086] Specifically, the construction of the audio processing model needs to comprehensively consider the functions and interrelationships of each sub-network. The feature extraction network is responsible for extracting multi-scale features from the original audio signal, the noise suppression network is used to reduce background noise, and the adaptive enhancement network focuses on enhancing the detail performance of the audio signal. The cross-attention network realizes more accurate feature extraction by fusing the output features of different networks, and the audio enhancement network is responsible for the final optimization of the audio signal. During the training process, the labeling of audio data is a key step, and the accuracy of labeling directly affects the performance of the model. When labeling, the starting point and ending point of the speech segment need to be accurately marked to ensure that the model can correctly distinguish speech and noise. In addition, the loss function used during the training process, such as the mean squared error loss function, is used to measure the difference between the model output and the true value and guide the adjustment of the model parameters. The calculation formula of the mean squared error loss function is:

[0087] Loss = N1∑i=1N(yi−y^i) 2

[0088] where N is the number of samples, yi is the true value of the sample, and y^i is the predicted value of the model.

[0089] Preferably, in the construction of the audio processing model, more complex network structures can be introduced. For example, a deeper Wave Net architecture can be used in the feature extraction network to extract richer audio features. In the noise suppression network, multiple noise intensity distribution thresholds can be set according to different noise types to adapt to diverse noise environments. During the model training process, data augmentation techniques such as adding different types of noise or adjusting the signal-to-noise ratio of the audio signal can be adopted to improve the robustness of the model. In addition, the mean squared error loss function can be combined with other loss functions such as spectral distortion loss or perceptual loss to further optimize the performance of the model. For labeled data, automated annotation tools can be introduced and combined with manual review to improve the annotation efficiency and accuracy.

[0090] In some embodiments, the mean squared error loss function is used to optimize the audio processing model.

[0091] It should be noted that the process of optimizing the audio processing model using the mean squared error loss function in the present invention is a key link to ensure the model performance. The mean squared error loss function is a commonly used optimization metric for measuring the difference between the model output and the true value. By minimizing this loss function, the variable parameters in the audio processing model can be adjusted, thereby optimizing the model performance so that it can more accurately restore the speech signal and suppress noise when processing audio signals. The calculation formula of the mean squared error loss function is:

[0092] Loss = 1 / N ∑_{i = 1}^{N} (y_i - \hat{y}_i)^2 2

[0093] where N represents the number of samples in the audio data, y_i represents the true value of the sample, and \hat{y}_i represents the prediction value of the model for this sample. By optimizing this loss function, it can be ensured that the audio processing model gradually approaches the ideal audio processing effect during the training process.

[0094] Specifically, the core of the mean squared error loss function is to calculate the squared difference between the model output and the true value and take the average of all samples. During the training process of the audio processing model, the model will adjust the weights and other variable parameters in the network according to the value of this loss function. For example, when the difference between the model output and the true value is large, the loss value will also increase accordingly, and the model will adjust the parameters through the backpropagation algorithm to reduce this difference. In practical applications, the parameter setting of the mean squared error loss function is relatively simple, mainly involving the determination of the sample number N, which is usually consistent with the size of the training dataset. In addition, the mean squared error loss function is applicable to various audio processing tasks because it can directly reflect the error of the audio signal in the time domain or frequency domain and is a general and effective optimization objective.

[0095] Preferably, in order to further improve the performance of the audio processing model, other auxiliary loss functions can be introduced based on the mean squared error loss function. For example, the spectral distortion loss function can be used to optimize the spectral characteristics of the audio signal, making it closer to the spectral distribution of real speech; the perceptual loss function can optimize the audio quality from the perspective of auditory perception, making the audio signal output by the model more natural in subjective listening. In addition, during the optimization process, a dynamic learning rate adjustment strategy can be adopted, such as learning rate decay or an adaptive learning rate algorithm (such as the Adam optimizer), to accelerate the convergence speed of the model and improve the optimization effect. At the same time, in order to improve the generalization ability of the model, regularization techniques such as weight decay or Dropout can be introduced during training to prevent the model from overfitting.

[0096] In some embodiments, processing the first value and the second value based on the first attention matrix and the second attention matrix to obtain a fused audio feature, including:

[0097] Fused audio feature The calculation formula is:

[0098]

[0099] Where, represents the second query, represents the first key, represents the first value, represents the first query, represents the second key, represents the second value, represents the normalization function.

[0100] It should be noted that the calculation formula of the fused audio feature mentioned in the present invention is the key part to implement the function of the cross-attention network. This formula processes the query (Query), key (Key), and value (Value) through the normalization function, thereby realizing the dynamic fusion of features. Among them, the query and the key are used to calculate the attention matrix, while the value is used to store the feature information. In this way, the fused audio feature can better retain the important information in the audio signal while suppressing noise interference. The role of the normalization function is to adjust the feature values to a suitable range to improve the stability and generalization ability of the model.

[0101] Specifically, the calculation formula of the fused audio feature is:

[0102] FusionFeature = Normalize(Q ⋅ K^T) ⋅ V

[0103] Among them, Q represents the query, K represents the key, and V represents the value. The normalization function (Normalize) usually adopts the Softmax function or other similar normalization methods to normalize the attention scores into a probability distribution. In the cross-attention network, the first denoised audio feature and the second enhanced audio feature correspond to different queries, keys, and values respectively. By calculating the attention matrix between the first key and the second query, and the attention matrix between the first query and the second key, effective fusion of the two features can be achieved. In practical applications, the calculation of the attention matrix can be realized through matrix multiplication, and the parameters of the normalization function can be adjusted according to the characteristics of the audio signal to optimize the fusion effect.

[0104] Preferably, the normalization function can adopt a more complex multi-head attention mechanism (Multi-Head Attention) to further improve the effect of feature fusion. The multi-head attention mechanism divides the query, key, and value into multiple heads, calculates the attention matrix respectively, and then concatenates the results together, so as to capture the features at different frequencies and time scales in the audio signal. In addition, the output of the normalization function can be combined with a residual connection to avoid the problem of gradient vanishing or explosion in the multi-layer network. When calculating the attention matrix, a scaling factor (such as dk1) can be introduced, where dk is the dimension of the key, to stabilize the propagation of the gradient. At the same time, in order to further optimize the fusion effect, an adaptive weight adjustment mechanism can be introduced to dynamically adjust the weights of the query, key, and value according to the real-time characteristics of the audio signal.

[0105] In some embodiments, the expression of the mean square error loss function is:

[0106]

[0107] Among them, represents the loss value, represents the number of samples in the audio data, represents the sample 's true value, represents the sample 's predicted value.

[0108] It should be noted that the mean squared error loss function mentioned in the present invention is a key tool for optimizing the audio processing model. This loss function quantifies the performance of the model by calculating the squared difference between the model output and the true value and taking the average over all samples. During the training process of the audio processing model, the mean squared error loss function can effectively measure the accuracy of audio signal processing, help the model adjust parameters to minimize the error, and thus achieve more accurate audio enhancement and noise reduction effects. By optimizing this loss function, the model can better restore the speech signal in a complex noise environment and improve the user's auditory experience.

[0109] Specifically, the expression of the mean squared error loss function is:

[0110] Loss = \(\frac{1}{N}\sum_{i = 1}^{N}(y_{i}-\hat{y}_{i})^{2}\) 2

[0111] Among them, \(N\) represents the number of samples in the audio data, \(y_{i}\) is the true value of the sample, and \(\hat{y}_{i}\) is the predicted value of the model for this sample. In the training of the audio processing model, the true value \(y_{i}\) is usually a high-quality audio signal that has been manually labeled or preprocessed, while the predicted value \(\hat{y}_{i}\) is the audio signal processed by the model. By calculating the squared error of each sample and taking the average, the model can obtain a global error evaluation. In practical applications, the smaller the value of the loss function, the better the performance of the model. To optimize the model, the weights and other parameters in the network are usually adjusted through the backpropagation algorithm to gradually reduce the loss value.

[0112] Preferably, to further improve the optimization effect of the model, other auxiliary loss functions can be introduced on the basis of the mean squared error loss function. For example, the spectral distortion loss function can be used to optimize the spectral characteristics of the audio signal to make it closer to the spectral distribution of real speech; the perceptual loss function can optimize the audio quality from the perspective of auditory perception to make the audio signal output by the model more natural in subjective listening. In addition, the mean squared error loss function can be weighted, and different weights can be assigned according to the importance of different frequency bands or time periods of the audio signal, so as to more accurately optimize the model performance. During the training process, a dynamic learning rate adjustment strategy can also be combined, such as learning rate decay or an adaptive learning rate algorithm (such as the Adam optimizer), to accelerate the convergence speed of the model and improve the optimization effect.

[0113] The above embodiments of the present invention have the following beneficial effects: The hearing aid audio processing integration method and device of the present invention can effectively improve the speech processing performance of hearing aids in complex noise environments through multi-level network processing. The feature extraction network can obtain multi-scale audio features, providing a rich information basis for subsequent processing; the noise suppression network can accurately predict and suppress noise based on a preset noise intensity distribution threshold, ensuring the purity of the speech signal. The adaptive enhancement network can further improve the detail performance of the speech signal through randomly selecting frequency reference points and feature points for processing, thereby providing a clearer and more natural auditory experience for users in complex noise environments.

[0114] In addition, the cross-attention network can efficiently fuse the denoised and enhanced audio features, and achieve dynamic adjustment and optimization of the features by calculating the attention matrix, further improving the quality of the audio signal. The multi-layer filtering, inverse filtering, upsampling, and frame-by-frame classification processing of the audio enhancement network can further optimize the resolution and stability of the audio signal, ensuring that the output enhanced audio signal is more natural and clear. Overall, the method and device can provide a more personalized and accurate audio processing solution for hearing aid users, and at the same time provide strong support for the intelligent development of hearing aid technology.

[0115] As Figure 2 shown, an audio processing integration device 200 of a hearing aid according to some embodiments, the device 200 includes:

[0116] An acquisition module 201, which acquires in real time the audio signal collected by a microphone deployed in a hearing aid;

[0117] A feature extraction module 202, which processes the audio signal acquired by the acquisition module based on a feature extraction network to obtain first audio feature data;

[0118] A first processing module 203, which processes the first audio feature data based on a noise suppression network to obtain a first denoised audio feature;

[0119] A second processing module 204, which processes the first audio feature data based on an adaptive enhancement network to obtain a second enhanced audio feature;

[0120] A feature fusion module 205, which fuses the first denoised audio feature output by the first processing module and the second enhanced audio feature output by the second processing module based on a cross-attention network to obtain a fused audio feature;

[0121] A third processing module 206, which processes the fused audio feature output by the feature fusion module by using an audio enhancement network to obtain an enhanced audio signal, and the audio enhancement network sequentially includes a plurality of filtering layers, inverse filtering layers, upsampling layers, and frame-by-frame classification layers;

[0122] The output module 207 outputs the enhanced audio signal to the speaker of the hearing aid.

[0123] It can be understood that the various modules described in the audio processing integration device 200 of the hearing aid correspond to the respective steps in the audio processing integration method of the hearing aid described in the reference. Figure 1 Therefore, the operations, features, and beneficial effects described above for the audio processing integration method of the hearing aid also apply to the audio processing integration device 200 of the hearing aid and the modules included therein, and will not be elaborated herein.

[0124] Next, referring to Figure 3 , which shows a schematic structural diagram of an electronic device 300 suitable for implementing some embodiments of the present invention. The electronic device in some embodiments of the present invention may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 3 The terminal device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present invention.

[0125] As Figure 3 shown, the electronic device 300 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 302 or the program loaded from the storage device 308 into the random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The processing device 301, the ROM 302, and the RAM 303 are connected to each other through a bus 304. The input / output (I / O) interface 305 is also connected to the bus 304.

[0126] Generally, the following devices may be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 309. The communication device 309 may allow the electronic device 300 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 3 the electronic device 300 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had. Figure 3Each box shown may represent a device or, as required, multiple devices.

[0127] Furthermore, the storage medium of the embodiments of the present application stores program instructions capable of implementing all of the above methods. Among them, the program instructions may be stored in the above storage medium in the form of a software product, including several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes, or terminal devices such as computers, servers, mobile phones, and tablets.

[0128] The above description is only some preferred embodiments of the present invention and an explanation of the technical principles applied. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present invention is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features with similar functions disclosed in the embodiments of the present invention.

Claims

1. An audio processing integration method for a hearing aid, characterized in that: include: Acquire the audio signal collected by the microphone deployed in the hearing aid in real time; Processing the audio signal based on a feature extraction network to obtain first audio feature data; The first audio feature data is processed based on a noise suppression network to obtain a first noise reduction audio feature, including: presetting a noise intensity distribution threshold based on the installation position of the microphone; converting the first audio feature data to the frequency domain; based on the noise intensity distribution threshold, using the noise suppression network to predict the noise distribution corresponding to the converted first audio feature data; weighting, filtering and minimum pooling aggregation processing are performed on the converted first audio feature data and the noise distribution to obtain the first noise reduction audio feature; The first audio feature data is processed based on an adaptive enhancement network to obtain a second enhanced audio feature, including: presetting an audio spectrum space, and converting each frequency point in the audio spectrum space to a frequency domain; randomly selecting a plurality of frequency reference points above the converted audio spectrum space, and projecting them to a time domain to obtain projection points; collecting feature points around the projection points in the first audio feature data; processing the feature points through an adaptive enhancement network to determine a feature value corresponding to each projection point; and determining a second enhanced audio feature based on the feature value; Based on the cross attention network, the first noise reduction audio feature and the second enhanced audio feature are fused to obtain the fused audio feature, including: applying linear transformation to the first noise reduction audio feature and the second enhanced audio feature respectively to obtain corresponding queries, keys and values, wherein the first noise reduction audio feature corresponds to the first query, the first key and the first value, and the second enhanced audio feature corresponds to the second query, the second key and the second value; calculating the first attention matrix based on the first key and the second query; calculating the second attention matrix based on the first query and the second key; processing the first value and the second value based on the first attention matrix and the second attention matrix to obtain the fused audio feature; The fused audio features are processed using an audio enhancement network to obtain an enhanced audio signal, wherein the audio enhancement network sequentially includes a plurality of filtering layers, an inverse filtering layer, an upsampling layer, and a frame-by-frame classification layer.

2. The audio processing integration method of a hearing aid according to claim 1, characterized in that: Processing the audio signal based on a feature extraction network to obtain first audio feature data, including: using a WaveNet neural network to process the audio signal to obtain a multi-scale audio feature, wherein the multi-scale audio feature includes a plurality of audio feature data with different resolutions; The multi-scale audio features are sequentially fused according to the resolution from low to high, including: performing 1×1 convolution and upsampling processing on the audio feature data of the first resolution, and fusing the processing result with the audio feature data of the second resolution through the 1×1 convolution result; The first resolution is audio feature data with a lower resolution among the multi-scale audio features; the second resolution is audio feature data with a higher resolution among the multi-scale audio features; A convolutional neural network is used to aggregate the multi-scale audio feature fusion results to obtain first audio feature data.

3. The audio processing integration method of a hearing aid according to claim 1, characterized in that: The converted first audio feature data and noise distribution are weighted, filtered and minimum pooled to obtain a first noise reduction audio feature, including: weighting the converted first audio feature data and the noise distribution to obtain an adjusted audio feature; filtering the adjusted audio feature through a filter; dividing the filtered audio feature into multiple frequency bands, aggregating audio features belonging to the same frequency band, and performing minimum pooling on the aggregation results to obtain the first noise reduction audio feature.

4. The audio processing integration method of a hearing aid according to claim 1, characterized in that: The feature points are processed by an adaptive enhancement network to determine the feature value corresponding to each projection point, including: processing the feature points by an adaptive enhancement network to determine the corresponding weight value; performing weighted summation on the feature points corresponding to each projection point to obtain the corresponding feature value, and determining the position of the feature value in the audio spectrum space.

5. The audio processing integration method of a hearing aid according to claim 1, characterized in that: Performing audio enhancement network processing on the fused audio features to obtain an enhanced audio signal, including: performing multi-layer filtering processing on the fused audio features to obtain first data; Performing inverse filter layer processing on the first data to obtain second data; Upsampling the second data to obtain third data having the same resolution as the fused audio feature; The third data is classified frame by frame to obtain an enhanced audio signal.

6. The audio processing integration method of a hearing aid according to claim 1, characterized in that: Also includes: Constructing an audio processing model, wherein the audio processing model includes a feature extraction network, a noise suppression network, an adaptive enhancement network, a cross-attention network, and an audio enhancement network; Collecting multiple audio data and annotating the collected audio data, including: starting from a starting point of a speech segment in the audio data, extending along the speech segment in the audio data until a complete speech annotation is formed; The annotated audio data is input into the audio processing model for training to adjust the variable parameters in the audio processing model.

7. The audio processing integration method of a hearing aid according to claim 6, characterized in that: The audio processing model is optimized using the mean square error loss function.

8. The audio processing integration method of a hearing aid according to claim 3, characterized in that: The first value and the second value are processed based on the first attention matrix and the second attention matrix to obtain a fused audio feature, including: Fusion of audio features The calculation formula is: in, represents the second query, Indicates the first key, represents the first value, represents the first query, represents the second key, represents the second value, Represents the normalization function.

9. The audio processing integration method of a hearing aid according to claim 7, characterized in that: The expression of the mean square error loss function is: in, represents the loss value, Represents the number of samples in the audio data, Representation sample The true value of Representation sample The predicted value of .

10. A hearing aid device, characterized in that: A hearing aid audio processing integration method according to any one of claims 1 to 9 is implemented, wherein the hearing aid device comprises: an acquisition module for acquiring in real time an audio signal collected by a microphone deployed in the hearing aid; A feature extraction module processes the audio signal acquired by the acquisition module based on a feature extraction network to obtain first audio feature data; A first processing module processes the first audio feature data based on a noise suppression network to obtain a first noise reduction audio feature; A second processing module processes the first audio feature data based on an adaptive enhancement network to obtain a second enhanced audio feature; A feature fusion module, based on a cross attention network, fuses the first noise reduction audio feature output by the first processing module and the second enhanced audio feature output by the second processing module to obtain a fused audio feature; A third processing module processes the fused audio features output by the feature fusion module using an audio enhancement network to obtain an enhanced audio signal, wherein the audio enhancement network sequentially includes a plurality of filtering layers, an inverse filtering layer, an upsampling layer, and a frame-by-frame classification layer; The output module outputs the enhanced audio signal to the speaker of the hearing aid.

Citation Information

Patent Citations

  • Speech recognition method, electronic equipment, vehicle-mounted speech recognition system and automobile

    CN115019778A

  • Multi-mode speech enhancement system based on audio and video

    CN119380742A