Ambient sound processing method applied to hearing aid and related product

By using front and rear microphones in the hearing aid to acquire signals, performing frequency response compensation and neural network processing, and outputting sound location and scene category, the problem of traditional hearing aids failing to utilize sound location information in scene classification is solved, thus improving the user's auditory experience and comfort.

CN121940703APending Publication Date: 2026-04-28AUSTAR HEARING SCI & TECH XIAMEN CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
AUSTAR HEARING SCI & TECH XIAMEN CO LTD
Filing Date
2024-10-25
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Traditional hearing aids fail to effectively utilize sound location information in scene classification, resulting in poor hearing performance and an inability to provide optimal parameter adjustment and processing strategies in different scenarios.

Method used

By acquiring ambient sound signals, using front and rear microphones to obtain directional and omnidirectional signals, and performing frequency response compensation to generate multidimensional features, these features are input into a preset neural network model for splitting processing, and the output sound location information, type, and scene category are used to adjust algorithm parameters.

Benefits of technology

It improves the user's auditory experience and comfort in different scenarios by optimizing the parameter adjustment and processing strategies of hearing aids through more detailed scenario division and location information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121940703A_ABST
    Figure CN121940703A_ABST
Patent Text Reader

Abstract

The invention provides an environmental sound processing method applied to a hearing aid and a related product. The method comprises the following steps: acquiring a signal corresponding to environment sound; processing the signal to obtain a multi-dimensional feature corresponding to the environment sound; and the multi-dimensional features are input into a preset neural network model for branching processing, one path outputs the azimuth information and the sound type of the environmental sound, and the other path outputs the scene type of the environmental sound. Through the method, information in different directions can be provided while background information classification is provided, the same scene is divided more meticulously, a basis is provided for parameter adjustment and strategies of other algorithms such as WDRC, noise reduction, directivity and feedback, and finally the intelligibility and comfort of a user are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of hearing aids, and in particular to an environmental sound processing method and related products for use in hearing aids. Background Technology

[0002] In the field of hearing aids, scene classification technology can significantly improve the user's listening experience. By analyzing the characteristics of ambient sound, the current environment is determined, and then the hearing aid parameters are adjusted or the program settings are switched to optimize the user's auditory experience in different scenarios. Traditional processing methods extract audio features or audio features from a specific location for classification. However, these methods only obtain environmental information at the current time or from a specific location, failing to utilize spatial information to improve classification accuracy. Furthermore, the classified scene does not reflect the sound's location, limiting subsequent parameter adjustments and processing strategies for the hearing aid. The processing methods for sounds from different locations within the same scene vary greatly. For example, speech preservation and noise suppression differ between multi-person and single-person conversations in an office setting, and directional noise is handled in street scenes. Only by determining both the scene type and the current sound's location can a better user experience be provided. Summary of the Invention

[0003] In view of the above problems, the present invention proposes an environmental sound processing method and related products for hearing aids that overcomes or at least partially solves the above problems.

[0004] One objective of this invention is to determine the location information, sound type, and scene category of the current sound.

[0005] A further objective of this invention is to determine the type of algorithm for acquiring ambient sound based on different actual scenarios.

[0006] A further objective of this invention is to enhance the user experience in different scenarios.

[0007] Specifically, the present invention provides an environmental sound processing method for hearing aids, comprising:

[0008] Acquire signals corresponding to ambient sounds;

[0009] The signal is processed to obtain the multidimensional features corresponding to the ambient sound;

[0010] The multidimensional features are input into a preset neural network model for split processing. One path outputs the location information and sound type of the ambient sound, while the other path outputs the scene category of the ambient sound.

[0011] Optionally, the step of inputting multidimensional features into a preset neural network model for demultiplexing processing includes:

[0012] Obtain multidimensional features;

[0013] Multidimensional features are convolved in two dimensions to obtain low-dimensional feature representations;

[0014] The feature representation is input into the first layer block for processing to obtain the first processing result;

[0015] The first processing result is input into the second and third layer blocks respectively for processing to obtain the second and third processing results.

[0016] The second processing result is processed to output location information and sound type;

[0017] The third processing result is processed to output the scene category.

[0018] Optionally, the first layer block includes: a first convolutional layer and a second convolutional layer;

[0019] The first convolutional layer uses depthwise separable convolution;

[0020] The second convolutional layer uses pointwise convolution;

[0021] Both the second and third layers include a third convolutional layer and a fourth convolutional layer. The second and third layers introduce a layer attention mechanism and deploy the layer attention mechanism between the third and fourth convolutional layers. The third convolutional layer uses depthwise separable convolution, and the fourth convolutional layer uses pointwise convolution.

[0022] Optionally, the hearing aid includes a front microphone and a rear microphone;

[0023] Signals include directional signals and omnidirectional signals. Directional signals include forward signals and backward signals.

[0024] The steps for acquiring the signal corresponding to ambient sound include:

[0025] Acquire forward ambient sound signals through the front microphone;

[0026] Acquire ambient sound signals from behind using the rear microphone;

[0027] The forward and backward ambient sound signals are processed to obtain an omnidirectional signal;

[0028] The forward and backward ambient sound signals are delayed according to a preset delay to obtain forward and backward signals. The delay is determined based on the distance between the front and rear microphones.

[0029] Optionally, the steps of processing the signal to obtain the multidimensional features corresponding to the ambient sound include:

[0030] Frequency response compensation is performed on both forward and backward signals;

[0031] The forward, backward, and omnidirectional signals after frequency response compensation are processed to obtain a three-dimensional feature map as a multi-dimensional feature.

[0032] Optionally, the steps for frequency response compensation of the forward and backward signals include:

[0033] The frequency responses of the forward and backward signals are determined as the frequency responses to be compensated;

[0034] Design a compensation frequency response with the opposite frequency response to the frequency response to be compensated;

[0035] The frequency response to be compensated is normalized by compensating the frequency response to be compensated, thereby normalizing the forward signal, backward signal and omnidirectional signal to the same frequency response.

[0036] Optionally, after the step of inputting multidimensional features into a preset neural network model for demultiplexing processing, the method further includes:

[0037] The algorithm for acquiring ambient sound is adjusted based on location information, sound type, and scene category.

[0038] According to another aspect of the present invention, a computer-readable storage medium is also provided, on which a computer program is stored, wherein the computer program, when executed by a processor, implements the steps of any of the above-described environmental sound processing methods applied to a hearing aid.

[0039] According to another aspect of the present invention, a computer program product is also provided, comprising a computer program that, when executed by a processor, implements the steps of any of the above-described environmental sound processing methods applied to a hearing aid.

[0040] According to another aspect of the present invention, a hearing aid is also provided, including a memory, a processor, and a machine-executable program stored in the memory and running on the processor, wherein the processor executes the machine-executable program to implement the steps of any of the above-described environmental sound processing methods applied to the hearing aid.

[0041] The environmental sound processing method for hearing aids of the present invention first acquires the signal corresponding to the environmental sound; then, the signal is processed to obtain multi-dimensional features corresponding to the environmental sound; next, the multi-dimensional features are input into a preset neural network model for split processing, with one path outputting the location information and sound type of the environmental sound, and the other path outputting the scene category of the environmental sound. This method can provide information on different locations while providing background information classification, allowing for more detailed segmentation of the same scene. This provides a basis for parameter adjustment and strategies for other algorithms such as WDRC, noise reduction, directivity, and feedback, ultimately improving user intelligibility and comfort.

[0042] Furthermore, the environmental sound processing method for hearing aids of the present invention includes the following steps for inputting multi-dimensional features into a preset neural network model for multi-path processing: acquiring multi-dimensional features; performing two-dimensional convolution on the multi-dimensional features to obtain a low-dimensional feature representation; inputting the feature representation into a first layer block for processing to obtain a first processing result; inputting the first processing result into a second layer block and a third layer block for processing to obtain a second processing result and a third processing result; processing the second processing result to output location information and sound type; and processing the third processing result to output scene category. The second and third layer blocks introduce a layer attention mechanism, which is deployed between the third and fourth convolutional layers. By introducing the layer attention mechanism, the second and third layer blocks can focus on their respective classification tasks while sharing the first layer block.

[0043] Furthermore, the environmental sound processing method for hearing aids of the present invention, after the step of inputting multi-dimensional features into a preset neural network model for de-processing, further includes: adjusting the algorithm for acquiring environmental sounds based on location information, sound type, and scene category. This method can combine sound location information, sound type, and scene type to determine how to adjust the algorithm's effect, thereby enhancing the user experience in different scenarios and improving hearing aid performance.

[0044] The above and other objects, advantages and features of the present invention will become more apparent to those skilled in the art from the following detailed description of specific embodiments of the invention in conjunction with the accompanying drawings. Attached Figure Description

[0045] The following sections will describe some specific embodiments of the invention in detail by way of example and not limitation, with reference to the accompanying drawings. The same reference numerals in the drawings denote the same or similar parts or portions. Those skilled in the art should understand that these drawings are not necessarily drawn to scale. In the drawings:

[0046] Figure 1 This is a flowchart illustrating an environmental sound processing method for hearing aids according to an embodiment of the present invention.

[0047] Figure 2 This is a schematic diagram of the architecture of an environmental sound processing method for hearing aids according to an embodiment of the present invention;

[0048] Figure 3 This is a schematic diagram of the processing flow of a neural network model for an environmental sound processing method applied to a hearing aid according to an embodiment of the present invention;

[0049] Figure 4 This is a schematic diagram of the processing flow of a neural network model for an environmental sound processing method applied to a hearing aid according to another embodiment of the present invention;

[0050] Figure 5 This is a schematic diagram of the architecture of the first layer block of an environmental sound processing method for hearing aids according to an embodiment of the present invention;

[0051] Figure 6 This is a schematic diagram of the architecture of the second layer block of an environmental sound processing method for hearing aids according to an embodiment of the present invention;

[0052] Figure 7 This is a schematic diagram of the frequency response of different signals in an environmental sound processing method for a hearing aid according to an embodiment of the present invention;

[0053] Figure 8 This is a schematic diagram of a computer program product according to an embodiment of the present invention;

[0054] Figure 9 This is a schematic diagram of a computer-readable storage medium according to an embodiment of the present invention; and

[0055] Figure 10 This is a schematic diagram of a hearing aid according to an embodiment of the present invention. Detailed Implementation

[0056] Those skilled in the art should understand that the embodiments described below are merely a part of the embodiments of the present invention, and not all of the embodiments of the present invention. These partial embodiments are intended to explain the technical principles of the present invention and are not intended to limit the scope of protection of the present invention. Based on the embodiments provided by the present invention, all other embodiments obtained by those skilled in the art without creative effort should still fall within the scope of protection of the present invention.

[0057] It should be noted that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus or device (such as a computer-based system, a processor-included system or other system that can fetch and execute instructions from, an instruction execution system, apparatus or device).

[0058] This invention provides an environmental sound processing method for hearing aids. Figure 1 This is a flowchart illustrating an environmental sound processing method for hearing aids according to an embodiment of the present invention, as shown below. Figure 1 As shown, the environmental sound processing method applied to hearing aids includes at least the following steps S101 to S103.

[0059] Step S101: Acquire the signal corresponding to the ambient sound. In some optional embodiments of the present invention, the hearing aid generally includes a front microphone and a rear microphone. The signal generally includes a directional signal and an omnidirectional signal, and the directional signal generally includes a forward signal and a backward signal. The step of acquiring the signal corresponding to the ambient sound generally includes: acquiring a forward ambient sound signal through the front microphone; acquiring a backward ambient sound signal through the rear microphone; processing the forward and backward ambient sound signals to obtain an omnidirectional signal; delaying the forward and backward ambient sound signals according to a preset delay to obtain the forward signal and the backward signal, the delay being determined according to the distance between the front and rear microphones. Acquiring the forward ambient sound signal through the front microphone means using the microphone located in front to capture all sounds coming from in front of it, and the sound is converted into a corresponding signal, i.e., the forward ambient sound signal, after passing through the front microphone. Similarly, the microphone located at the rear is responsible for collecting sound information from behind it.

[0060] Optionally, the omnidirectional signal can be obtained by processing the forward and backward ambient sound signals, which can be achieved by adding the forward and backward ambient sound signals and dividing by 2. The omnidirectional signal represents information about the entire surroundings of the hearing aid. Those skilled in the art can choose the appropriate processing method to obtain the omnidirectional signal based on the actual situation.

[0061] Step S102: Process the signal to obtain the multidimensional features corresponding to the ambient sound.

[0062] Due to the differences in directional frequency response characteristics, signals from different directions have different frequency responses. Furthermore, the canceling effect of signal overlap results in different frequency responses at different frequencies. Therefore, to eliminate the influence of frequency response, in some optional embodiments, the step of processing the signal to obtain multidimensional features corresponding to the ambient sound generally includes: performing frequency response compensation on the forward and backward signals; and processing the frequency-response-compensated forward, backward, and omnidirectional signals to obtain a three-dimensional feature map as multidimensional features.

[0063] Optionally, the steps for frequency response compensation of the forward and backward signals may generally include: determining the frequency responses of the forward and backward signals as the frequencies to be compensated; designing compensation frequencies with opposite frequencies based on the frequencies to be compensated; and normalizing the frequencies to be compensated through the compensation frequencies, thereby normalizing the forward, backward, and omnidirectional signals to the same frequency response.

[0064] One possible example of frequency response compensation is as follows: Figure 7 As shown, Figure 7This is a schematic diagram illustrating the frequency response of different signals in an environmental sound processing method for hearing aids according to an embodiment of the present invention. 710 represents an example of the frequency response of a forward or backward signal. 720 represents an example of the frequency response of an omnidirectional signal, typically exhibiting a flat frequency response. Due to the destructive effect of signal overlap, forward or backward signals exhibit different frequency responses at different frequencies. Therefore, frequency response normalization is required before feature fusion. This is achieved by designing an IIR filter with the opposite frequency response to 710, for example... Figure 7 The reshape IIR has a frequency response of 730. This filter normalizes the three signals to the same frequency response.

[0065] After frequency response compensation is completed, the forward, backward, and omnidirectional signals after frequency response compensation are processed to obtain a three-dimensional feature map as a multi-dimensional feature. One possible processing method is to analyze the short-time Fourier spectrum representation or Mel spectrum features of the three signals using WOLA (Weighted Overlap-Add) technology, and then recombine the processed frames into a continuous output signal through the Combine operation, and finally fuse them to form the above-mentioned three-dimensional feature map.

[0066] Step S103: Input the multidimensional features into the preset neural network model for split processing. One path outputs the location information and sound type of the ambient sound, and the other path outputs the scene category of the ambient sound.

[0067] In some optional embodiments, the step of inputting multidimensional features into a preset neural network model for demultiplexing processing generally includes: acquiring multidimensional features; performing two-dimensional convolution on the multidimensional features to obtain a low-dimensional feature representation; inputting the feature representation into a first layer block for processing to obtain a first processing result; inputting the first processing result into a second layer block and a third layer block for processing to obtain a second processing result and a third processing result; processing the second processing result to output location information and sound type; and processing the third processing result to output scene category. It should be noted that the first layer block, second layer block, and third layer block can be composed of multiple corresponding cascaded layers, and those skilled in the art can determine their number according to the actual situation.

[0068] The first layer typically includes a first convolutional layer and a second convolutional layer. The first convolutional layer can use depthwise separable convolution, and the second convolutional layer can use pointwise convolution. Both the second and third layers include a third convolutional layer and a fourth convolutional layer. Layer attention mechanisms are introduced in the second and third layers and deployed between the third and fourth convolutional layers. This layer attention mechanism allows the second and third layers to focus on their respective classification tasks while sharing the initial processing results. The third convolutional layer can use depthwise separable convolution, and the fourth convolutional layer can use pointwise convolution.

[0069] The process of processing the second processing result to output location information and sound type, and processing the third processing result to output scene category, can generally be achieved through point-wise convolution, pooling layers, fully connected layers, and softmax. The final output signal's location information and sound type typically include forward, backward, left / right, and omnidirectional directions; the sound type can generally include speech and noise. This location information and sound type can be matched and combined into 8 labels, and the probabilities corresponding to these 8 labels are output. Another output shows the probabilities of different scene categories. Scene categories can generally include: 1. Quiet scenes (e.g., home, library); 2. Social scenes (e.g., multiple people talking in an office or noisy social occasions); 3. Noisy indoor public places (e.g., shopping malls, airports); 4. Outdoor scenes (e.g., streets, parks, squares); 5. Transportation scenes (e.g., buses, subways); 6. Music scenes, etc.

[0070] This method can provide information from different directions while providing background information classification, enabling more detailed division of the same scene. It provides a basis for parameter adjustment and strategies for other algorithms such as WDRC, noise reduction, directionality, and feedback, ultimately improving user understanding and comfort.

[0071] In some optional embodiments, after the step of inputting multidimensional features into a preset neural network model for demultiplexing processing, the method may further include: adjusting the algorithm for acquiring ambient sound based on location information, sound type, and scene category. To clearly illustrate the adjustment process, some specific embodiments are provided, such as the application of this invention to a street scene:

[0072] Assume the output of the second layer block is label1 (label 1), and the output of the third layer block is label2 (label 2).

[0073] Scenario A: The user is walking on the street and only hears background street noise. In this scenario, label1 will identify it as omnidirectional noise, and label2 will identify it as an outdoor scene. In this case, the directional algorithm will be adjusted to omnidirectional mode, and the noise reduction algorithm will only suppress the background noise to a certain extent.

[0074] Scenario B: A user hears a car horn behind them while walking on the street. In this scenario, label1 identifies it as noise from behind, while label2 identifies it as an outdoor scene. The directional algorithm is then adjusted to a backward orientation to pick up the noise from behind (because the sound from behind on the street is a warning sound) and to suppress background noise to a certain extent. This allows the user to immediately perceive the approaching vehicle, improving their sense of security.

[0075] Scenario C: A user is having a face-to-face conversation with someone on the street. In this scenario, label1 will identify the speech as forward, while label2 will identify it as an outdoor scene. In this case, the directional algorithm is adjusted to be forward-facing to pick up the speech signal from the front, and the noise reduction algorithm suppresses a certain degree of background noise, thereby improving the intelligibility of the speech while ensuring safety.

[0076] When the background is outdoor, the noise reduction level is limited, and a certain amount of background noise is provided to ensure user safety. Based on this, combined with the sound's location information and whether the sound type is speech, the directional algorithm effect can be adjusted to enhance the user experience in street scenes. Similarly, in other scenarios, scene classification combined with spatial characteristics can provide richer information to improve the hearing aid effect.

[0077] Figure 2 This is a schematic diagram of the architecture of an environmental sound processing method for hearing aids according to an embodiment of the present invention, as shown below. Figure 2 As shown, it generally includes a front microphone 211 and a rear microphone 212. The front microphone 211 is used to acquire the forward ambient sound signal, and the rear microphone 212 is used to acquire the backward ambient sound signal. The forward and backward ambient sound signals are added together and divided by 2 to obtain an omnidirectional signal 233 (using "two mic omni" to represent the omnidirectional signal). The forward ambient sound signal is delayed at 221 and combined with the backward ambient sound signal to obtain a rear-facing signal 232; the backward ambient sound signal is delayed at 222 and combined with the forward ambient sound signal to obtain a front-facing signal 231. Subsequently, the forward signal passes through filter 241 (reshape IIR filter), and the backward signal passes through filter 242 for frequency response compensation to obtain three signals with consistent frequency response (forward signal, backward signal, and omnidirectional signal). Next, the three signals are analyzed by WOLA at 250 to obtain the short-time Fourier spectrum representation or Mel spectrum features of the three signals. The processed frames are then recombined into a continuous output signal through the Combine operation, and finally fused to form a three-dimensional feature map. The three-dimensional feature map is input into the neural network 260 for data processing, and the signal information and background information of the front, back, left and right directions can be obtained, providing a basis for subsequent signal processing and strategies.

[0078] Figure 3 This is a schematic diagram of the processing flow of a neural network model for an environmental sound processing method applied to a hearing aid according to an embodiment of the present invention, as shown below. Figure 3 As shown, the environmental sound processing method applied to hearing aids includes at least the following steps S301 to S305.

[0079] Step S301: Obtain multidimensional features.

[0080] Step S302: Perform two-dimensional convolution on the multidimensional features to obtain a low-dimensional feature representation.

[0081] Step S303: Input the feature representation into the first layer block for processing to obtain the first processing result.

[0082] Step S304: The first processing result is input into the second and third layer blocks for processing to obtain the second and third processing results. The first, second, and third layer blocks can be formed by cascading multiple corresponding layer blocks; those skilled in the art can determine their number according to the actual situation.

[0083] Step S305: Process the second processing result to output location information and sound type, and process the third processing result to output scene category. In the final output signal, the location information generally includes forward, backward, left / right, and omnidirectional; the sound type generally includes speech and noise. Thus, the location information and sound type can be matched and combined into 8 tags, and the probabilities corresponding to the 8 tags are output. The other output outputs the probabilities of different scene categories.

[0084] This method can provide information from different directions while providing background information classification, enabling more detailed division of the same scene. It provides a basis for parameter adjustment and strategies for other algorithms such as WDRC, noise reduction, directionality, and feedback, ultimately improving user understanding and comfort.

[0085] about Figure 3 An alternative implementation of the processing procedure, for example Figure 4 As shown, Figure 4 This is a schematic diagram of the processing flow of a neural network model for an environmental sound processing method applied to a hearing aid according to another embodiment of the present invention, as shown below. Figure 4As shown, the environmental sound processing method applied to hearing aids includes at least the following steps: First, at 410, the stacked features corresponding to the processed three signals are used as the input to the neural network model; then, conv2d (two-dimensional convolution) is performed at 420, and a low-dimensional feature representation can be obtained after standard two-dimensional convolution; the processing result is input to 430 and processed by layer1 blocks (i.e., the first layer block, which is composed of multiple layer1 blocks cascaded) to obtain the first processing result; the first processing result is input to layer2 blocks A (i.e., the second layer block, which is composed of multiple layer2 blocks cascaded) at 440 and layer2 blocks B (i.e., the third layer block, which is composed of multiple layer2 blocks cascaded) at 470 for processing.

[0086] The processing result at position 440 is called the second processing result. The second processing result is then subjected to pointwise convolution (pointwise convolution is a special type of convolution operation where the kernel size is 1x1) at position 450. The specific convolution process generally includes: conv2d(1*1) + bn (Batch Normalization) + relu (Rectified Linear Unit), pooling layer (AvgPool), fully connected layer (FC), and softmax (an activation function). Finally, the processing result output at position 450 is the probability of 8 labels composed of the signal's orientation information (front, back, left, right, omnidirectional) and sound type (speech, noise).

[0087] The processing result at position 470 is called the third processing result. This third processing result is then subjected to point-wise convolution (a special type of convolution operation where the kernel size is 1x1) at position 480. The specific convolution process generally includes: conv2d(1*1) + bn (Batch Normalization) + relu (Rectified Linear Unit), pooling layer (AvgPool), fully connected layer (FC), and softmax (an activation function). The final output at position 480 represents the probabilities of different scene categories. Optional scene categories include: 1. Quiet scenes (e.g., home, library); 2. Social scenes (e.g., multiple people talking in an office or noisy social occasions); 3. Noisy indoor public places (e.g., shopping malls, airports); 4. Outdoor scenes (e.g., streets, parks, squares); 5. Transportation scenes (e.g., buses, subways); 6. Music scenes, etc.

[0088] Figure 5 This is a schematic diagram of the architecture of the first layer block of an environmental sound processing method for a hearing aid according to an embodiment of the present invention. The first layer block is composed of multiple cascaded layer 1 blocks, and each layer 1 block is a block composed of multiple layer 1 blocks. The specific structure of layer 1 is as follows: Figure 5 As shown, the typical structure includes convolutional layers + batch normalization + activation layers. To achieve lightweight design, depthwise and pointwise convolutions are used, and skip connections are introduced to ensure information transfer within the model. Specifically, at step 431, a depthwise convolution + batch normalization + ReLU architecture is chosen. The data processed at step 431 is then input to step 432 for further processing. Step 432 uses a pointwise convolution + batch normalization architecture. After processing at step 432, the result is input to the activation layer at step 433 for further processing. Step 433 uses a ReLU architecture.

[0089] Figure 6 This is a schematic diagram of the architecture of the second layer block in an environmental sound processing method for hearing aids according to an embodiment of the present invention. The second layer block is composed of multiple cascaded layer 2 blocks, and each layer 2 block is a block composed of multiple layers 2. The specific structure of the layer 2 is as follows: Figure 6As shown, a layer attention mechanism is added on top of layer 1. The specific structure of layer 2 is as follows: at 441, the architecture is depthwise convolution (Depthwise Convolution) + batch normalization (BN) + relu (Rectified Linear Unit). The data processed at 441 is input into 442 for further processing. At 442, the newly introduced layer attention mechanism helps the second layer block focus on its classification task when using the results of the first processing. The architecture at 443 is pointwise convolution (Pointwise Convolution) + batch normalization (BN). After processing at 443, the result is input into the activation layer at 444 for further processing. The architecture at 444 is relu (Rectified Linear Unit).

[0090] Similarly, the architecture of the third layer block is the same as that of the second layer block. Both introduce a layer attention mechanism and deploy the layer attention mechanism between the third and fourth convolutional layers. This allows the third layer block and the second layer block to focus on their respective classification tasks while sharing layer1 blocks.

[0091] The flowchart provided in this embodiment is not intended to indicate that the operations of the method will be performed in any particular order, or that all operations of the method are included in every case. Furthermore, the method may include additional operations. Within the scope of the technical concept provided by the method in this embodiment, additional variations can be made to the above method.

[0092] It should be understood that in some embodiments, the components may be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods may be implemented using software or firmware stored in memory and executed by a suitable instruction execution system.

[0093] This embodiment also provides a computer program product 10, a computer-readable storage medium 20, and a hearing aid 30. Figure 8 This is a schematic diagram of a computer program product 10 according to an embodiment of the present invention. Figure 9 This is a schematic diagram of a computer-readable storage medium 20 according to an embodiment of the present invention. Figure 10This is a schematic diagram of a hearing aid 30 according to an embodiment of the present invention. The computer program product 10 includes a computer program 11, which, when executed by the processor 32, implements the steps of any of the above-described environmental sound processing methods applied to a hearing aid. A computer-readable storage medium 20 stores the computer program 11 thereon, which, when executed by the processor 32, implements the steps of any of the above-described environmental sound processing methods applied to a hearing aid. The hearing aid may include a memory 31, a processor 32, and the computer program 11 stored in the memory 31 and running on the processor 32.

[0094] The computer program 11 used to perform the operations of this invention may be assembly instructions, Instruction Set Architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, integrated circuit configuration data, or source code or object code written in any combination of one or more programming languages ​​and procedural programming languages. The computer program 11 may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer via any type of network, including a Local Area Network (LAN) or Wide Area Network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, to perform aspects of this invention, electronic circuits, including, for example, programmable logic circuits, Field-Programmable Gate Arrays (FPGAs), or Programmable Logic Arrays (PLAs), may execute computer-readable program instructions using status information from computer-readable program instructions to personalize the electronic circuits.

[0095] For the purposes of this embodiment, computer program product 10 is a related product containing computer program 11. For the purposes of this embodiment, computer-readable storage medium 20 is a tangible device capable of holding and storing computer program 11, and can be any device capable of containing, storing, communicating, propagating, or transmitting program 11 for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable storage medium 20 include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable optical disc read-only memory (CD-ROM), digital versatile disc (DVD), memory stick, floppy disk, mechanical encoding device, and any suitable combination thereof.

[0096] The hearing aid may include a processor 32 adapted to execute stored instructions and a memory 31 that provides temporary storage for the operation of said instructions during operation. The processor 32 may be a single-core processor, a multi-core processor, a computing cluster, or any other configuration. The memory 31 may include random access memory (RAM), read-only memory, flash memory, or any other suitable storage system.

[0097] Therefore, those skilled in the art should recognize that although numerous exemplary embodiments of the present invention have been shown and described in detail herein, many other variations or modifications conforming to the principles of the present invention can be directly determined or derived from the disclosure of the present invention without departing from the spirit and scope of the invention. Thus, the scope of the present invention should be understood and construed as covering all such other variations or modifications.

Claims

1. An environmental sound processing method for use in hearing aids, comprising: Acquire the signal corresponding to the ambient sound; The signal is processed to obtain the multidimensional features corresponding to the ambient sound; The multidimensional features are input into a preset neural network model for split processing. One path outputs the location information and sound type of the environmental sound, while the other path outputs the scene category of the environmental sound.

2. The environmental sound processing method for hearing aids according to claim 1, wherein, The step of inputting the multidimensional features into a preset neural network model for demultiplexing processing includes: Obtain the multidimensional features; The multidimensional features are then subjected to two-dimensional convolution to obtain a low-dimensional feature representation; The feature representation is input into the first layer block for processing to obtain the first processing result; The first processing result is input into the second layer block and the third layer block respectively for processing to obtain the second processing result and the third processing result; The second processing result is processed to output the location information and the sound type; The third processing result is processed to output the scene category.

3. The environmental sound processing method for hearing aids according to claim 2, wherein, The first layer block includes: a first convolutional layer and a second convolutional layer; The first convolutional layer uses depthwise separable convolution technology; The second convolutional layer uses pointwise convolution. Both the second layer block and the third layer block include a third convolutional layer and a fourth convolutional layer. The second layer block and the third layer block introduce a layer attention mechanism and deploy the layer attention mechanism between the third convolutional layer and the fourth convolutional layer. The third convolutional layer uses depthwise separable convolution technology, and the fourth convolutional layer uses pointwise convolution technology.

4. The environmental sound processing method for hearing aids according to claim 1, wherein, The hearing aid includes a front microphone and a rear microphone; The signal includes a directional signal and an omnidirectional signal, and the directional signal includes a forward signal and a backward signal; The step of acquiring the signal corresponding to the ambient sound includes: Acquire forward ambient sound signals through the front microphone; The rear ambient sound signal is acquired through the rear microphone; The omnidirectional signal is obtained by processing the forward ambient sound signal and the backward ambient sound signal. The forward ambient sound signal and the backward ambient sound signal are delayed according to a preset delay to obtain the forward signal and the backward signal. The delay is determined based on the distance between the front microphone and the rear microphone.

5. The environmental sound processing method for hearing aids according to claim 4, wherein, The step of processing the signal to obtain the multidimensional features corresponding to the ambient sound includes: Frequency response compensation is performed on the forward signal and the backward signal; The forward signal, the backward signal, and the omnidirectional signal after frequency response compensation are processed to obtain a three-dimensional feature map as the multidimensional feature.

6. The environmental sound processing method for hearing aids according to claim 5, wherein, The step of performing frequency response compensation on the forward signal and the backward signal includes: The frequency responses of the forward signal and the backward signal are determined as the frequency responses to be compensated; Design a compensation frequency response with the opposite frequency response based on the frequency response to be compensated; The frequency response to be compensated is normalized by the compensation frequency response, so that the forward signal, the backward signal and the omnidirectional signal are normalized to the same frequency response.

7. The environmental sound processing method for hearing aids according to claim 1, wherein, After the step of inputting the multidimensional features into a preset neural network model for demultiplexing processing, the method further includes: The algorithm for acquiring the environmental sound is adjusted based on the location information, the sound type, and the scene category.

8. A computer-readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the steps of the environmental sound processing method for a hearing aid as described in any one of claims 1 to 7.

9. A computer program product comprising a computer program that, when executed by a processor, implements the steps of the environmental sound processing method for a hearing aid as described in any one of claims 1 to 7.

10. A hearing aid, comprising a memory, a processor, and a machine-executable program stored in the memory and running on the processor, wherein the processor, when executing the machine-executable program, implements the steps of the environmental sound processing method for a hearing aid according to any one of claims 1 to 7.