Audio enhancement method and apparatus, computer storage medium

By introducing the adaptive adjustment of the attenuation function μ(t) of the WANC filter with time-varying weight coefficient matrix, the problem of over-suppression of sound in a specific direction in existing beamforming algorithms is solved, and flexible adjustment of sound from different directions and frequencies is realized.

CN114550734BActive Publication Date: 2026-01-06ORKA HEALTH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210199889.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-02
Publication Date
2026-01-06
Estimated Expiration
2042-03-02

AI Technical Summary

Technical Problem

Existing beamforming algorithms can only retain sound from a specific direction, while completely reducing sound from other directions. This fails to meet the requirements of simulating the sound reception effect of the human ear in applications such as hearing aids.

Method used

By introducing an adaptive filtering matrix WANC into the beamforming algorithm, the weight coefficient matrix of the attenuation function μ(t) changes with time to control the attenuation degree of sound in different directions and frequencies. Combined with the delay unit 305 to generate the attenuation function μ(t), customized adjustment of sound in different directions and frequencies can be achieved.

Benefits of technology

Under low power consumption, beamforming algorithms can better simulate the auditory experience of the human ear, adjust the frequency response in different directions, simulate the auditory effect of the human ear, solve the problems that existing technologies have not been able to effectively solve, and achieve the effect of targeted adjustment of sound from different directions and frequencies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114550734B_ABST
    Figure CN114550734B_ABST
Patent Text Reader

Abstract

This application discloses an audio enhancement method and apparatus, and a computer storage medium. The method includes: generating a set of audio acquisition signals from a microphone array; performing delay summation processing on the set of audio acquisition signals to generate a delay summation signal; performing blocking matrix processing on the set of audio acquisition signals to generate a blocking matrix signal; filtering the blocking matrix signal using an adaptive filtering matrix, and removing the filtered blocking matrix signal from the delay summation signal to obtain an enhanced audio output signal. The adaptive filtering matrix is ​​based on at least one attenuation function, and each of the at least one attenuation function is updated at a corresponding predetermined update interval T.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to a beamforming technique, and more specifically, to an audio enhancement method and apparatus, and a computer storage medium. Background Technology

[0002] Beamforming algorithms are commonly used in audio devices such as headphones, hearing aids, and speakers. Their basic principle is to use two or more microphones to pick up sound and calculate the time it takes for the same sound to reach different microphones, thus determining the source of the sound. In subsequent processes, algorithms can be used to retain or eliminate sounds coming from a particular direction. For example, Bluetooth wireless headphones with ambient noise cancellation can configure two microphones to be placed vertically, so that the wearer's mouth is roughly in a straight line connecting the two microphones. Picking up the wearer's voice in this way helps to eliminate ambient noise, thereby improving the sound quality during calls. Currently, hearing aids on the market generally have two microphones, which can be placed front-to-back. Beamforming algorithms can then be used to extract sounds from the front (relative to the wearer's orientation) and eliminate sounds from behind, allowing the wearer to better focus on sounds coming from the front during conversations.

[0003] However, typical beamforming algorithms can only preserve sound from a specific direction, while completely attenuating sound from other directions. This is unsuitable for applications such as hearing aids where two or more microphones are used to simulate the sound pickup effect of the human ear. Therefore, it is necessary to provide an improved beamforming algorithm. Summary of the Invention

[0004] One objective of this application is to provide an audio enhancement method and apparatus, and a computer storage medium, to solve the problem of over-suppression of sound in non-target directions by beamforming algorithms.

[0005] In one aspect of this application, an audio enhancement method is provided, the method comprising: generating a set of audio acquisition signals from a microphone array, wherein each audio acquisition signal in the set of audio acquisition signals is generated by one microphone in the microphone array, and each microphone in the microphone array is spaced apart from each other; and performing a delay summation process on the set of audio acquisition signals to generate a delay summation signal Y. DSB (k,l), where k represents the frequency window and l represents the frame index; the audio acquisition signals are processed by a blocking matrix to generate a blocking matrix signal Y. BM (k,l); using the adaptive filtering matrix W ANC For the blocking matrix signal Y BM (k,l) is filtered, and the filtered blocking matrix signal is then extracted from the delayed summation signal Y.DSB Remove from (k,l) to obtain the enhanced audio output signal Y. OUT (k,l); where the adaptive filtering matrix W ANC It is based on at least one attenuation function μ(t), which varies with the audio output signal Y. OUT (k,l) and the blocking matrix signal Y BM The weight coefficient matrix varies with (k,l), and each of the at least one decay function μ(t) is updated with a corresponding predetermined update interval T.

[0006] In some embodiments, the microphone array may optionally include at least two microphones located on the same audio processing device.

[0007] In some embodiments, the audio processing device may optionally be worn inside the ear.

[0008] In some embodiments, optionally, one of the at least two microphones is oriented toward the auricle, while the other of the at least two microphones is oriented away from the auricle.

[0009] In some embodiments, the audio output signal is optionally determined by the following equation: Furthermore, the adaptive filtering matrix W ANC Determined by the following equation: Among them, P est (k,l) is determined by the following equation: Where α is the forgetting factor and M is the number of microphones in the microphone array.

[0010] In some embodiments, optionally, the at least one attenuation function includes a first attenuation function and a second attenuation function, wherein the first attenuation function is updated at a first predetermined update interval, and the second attenuation function is updated at a second predetermined update interval; wherein the first attenuation function corresponds to a high-frequency signal greater than or equal to a predetermined frequency threshold; and the second attenuation function corresponds to a low-frequency signal less than the predetermined frequency threshold, and the first predetermined update interval is shorter than the second predetermined update interval.

[0011] In some embodiments, optionally, each of the decay functions μ(t) is updated in the current update interval based on its value in the first update interval.

[0012] In some embodiments, optionally, each point in the decay function μ(t) within the current update interval is updated based on a change weight between 0 and 1 assigned to the value of its corresponding point within the first update interval.

[0013] In some embodiments, optionally, the weight is a linear function of time within the current update interval.

[0014] In some embodiments, optionally, the weight is a linearly increasing function of time within the current update interval.

[0015] In some embodiments, optionally, the weight is a non - linear function of time within the current update interval.

[0016] In some embodiments, optionally, each of the attenuation functions μ(t) is also updated within the current update interval based on its value at the end of the previous update interval.

[0017] In some embodiments, optionally, each of the attenuation functions μ(t) satisfies the following equation within the current update interval (NT, (N + 1)T]: N*T < t ≤ (N + 1)*T; where N is a positive integer.

[0018] In another aspect of the present application, there is also provided an audio enhancement device, which includes a non - transient computer storage medium, on which one or more executable instructions are stored, and after being executed by a processor, the one or more executable instructions perform any of the audio enhancement methods described above.

[0019] In some embodiments, optionally, the audio enhancement device can be a hearing aid device.

[0020] In yet another aspect of the present application, there is also provided a non - transient computer storage medium, on which one or more executable instructions are stored, and after being executed by a processor, the one or more executable instructions perform any of the audio enhancement methods described above.

[0021] The above is an overview of the present application. There may be simplifications, generalizations, and omissions of details. Therefore, those skilled in the art should recognize that this part is only illustrative and not intended to limit the scope of the present application in any way. This overview section is neither intended to identify the key features or essential features of the claimed subject matter nor intended to be used as an aid in determining the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Through the following description of the specification and the appended claims in combination with the drawings, the above and other features of the present application will be more fully and clearly understood. It can be understood that these drawings only depict several embodiments of the present application, and thus should not be considered as limiting the scope of the present application. By using the drawings, the present application will be described more clearly and in detail.

[0023] Figure 1A schematic diagram of a beamforming algorithm based on an example is shown;

[0024] Figure 2 A schematic diagram of a beamforming algorithm based on an example is shown;

[0025] Figure 3 A schematic diagram of a beamforming algorithm according to an embodiment of this application is shown;

[0026] Figure 4 An audio enhancement method according to an embodiment of this application is shown;

[0027] Figure 5 A schematic diagram of a beamforming algorithm according to an embodiment of this application is shown;

[0028] Figure 6 A schematic diagram of a beamforming algorithm according to an embodiment of this application is shown;

[0029] Figure 7 A schematic diagram illustrating the effect of a beamforming algorithm according to an embodiment of this application is shown;

[0030] Figure 8 A schematic diagram illustrating the effect of a beamforming algorithm according to an embodiment of this application is shown;

[0031] Figure 9 A schematic diagram illustrating the effect of a beamforming algorithm according to an embodiment of this application is shown.

[0032] Before explaining any embodiments of the invention in detail, it should be understood that the application of the invention is not limited to the details of the construction and the arrangement of components set forth in the following description or shown in the following drawings. The invention can have other embodiments and can be practiced or implemented in various ways. Moreover, it should be understood that the wording and terminology used herein are for descriptive purposes and should not be considered limiting. Detailed Implementation

[0033] In the following detailed description, reference is made to the accompanying drawings, which form a part thereof. In the drawings, similar symbols generally denote similar components unless the context otherwise requires. The illustrative embodiments described in the detailed description, drawings, and claims are not intended to be limiting. Other embodiments and variations may be employed without departing from the spirit or scope of the subject matter of this application. It will be understood that various different configurations, substitutions, combinations, and designs can be made to the various aspects of the general description and illustrated in the drawings of this application, all of which explicitly form part of the subject matter of this application.

[0034] Figure 1 and Figure 2Beamforming algorithms based on some examples are shown. Figure 1 As shown, the sound emitted by sound source 101 can be picked up by microphones 102-1 and 102-2, such as those of a hearing aid. Microphones 102-1 and 102-2 can be positioned on the left and right sides of the hearing aid wearer 103 (e.g., inside the bilateral auricles), with a distance d between them. For example, the distance d can depend on the interauricular distance of the wearer 103. The wearer 103 faces the hearing aid at an angle of 0° as shown in the figure. Figure 1 Above (i.e., in front of the wearer). Sound source 101 is located to the left front of wearer 103, forming an angle θ with the midline of wearer 103's field of vision. Since the distance between sound source 101 and wearer 103 (and their two ears) is much greater than the distance between their ears, it can be assumed that sound source 101 is approximately at the angle θ shown in the figure with respect to both microphones 102-1 and 102-2. From geometric relationships, assuming the speed of sound in air is v, and the signal received by microphone 102-1 is y1(t), then the signal received by microphone 102-2 is y2(t) = y1(t-τ), where τ = (d*sin(θ)) / v.

[0035] Short-time Fourier transforms are performed on the sound signals received by microphones 102-1 and 102-2, respectively. Assume the transform result of y1(t) is Y1(k,l) and the transform result of y2(t) is Y2(k,l), where k represents the frequency window and l represents the frame index. Then Y1(k,l) and Y2(k,l) satisfy the following relationship: Y2(k,l) = Y1(k,l) * e -jωτ .

[0036] Transfer to Figure 2 The delayed beamformer 201 and the blocking matrix 202 receive and process signals from microphones 102-1 and 102-2, respectively. In some schemes, the signal Y obtained after processing by the delayed beamformer 201... DSB For example, it can satisfy The signal Y obtained after processing by blocking matrix 202 BM For example, it can satisfy Y BM =Y1(k,l)-Y2(k,l)e jωτ The parameter-adjustable minimum mean square adaptive filter (LMS filter) 203 further processes the YBM and sends the processed result to the summing unit 204. The signal Y output from the summing unit 204... GSC (k,l) satisfies Among them W ANC(k,l) are the iteration coefficients of the LMS filter 203, and * indicates conjugate.

[0037] Furthermore, W ANC (k,l) satisfy the following relationship:

[0038]

[0039] P est (k,l)=αP est (k,l-1)+(1-α)(|Y BM (k,l)| 2 +|Y GSC (k,l)| 2 (2)

[0040] If the hearing aid includes M microphones for collecting sound signals, then equation (2) can be expressed as:

[0041]

[0042] In equations (2) and (2') above, α is the forgetting factor. As understood, the introduction of the forgetting factor α can emphasize the amount of information provided by new data and gradually reduce the influence of earlier data, preventing data saturation.

[0043] However, as mentioned above, the beamforming algorithm described can only retain sound from a pre-set direction, while completely reducing sound from other directions. For example, returning... Figure 1 If the retention direction is set to 90°, this algorithm will retain almost all sound from the 90° direction, but will almost completely eliminate the signal from the 0° direction. Furthermore, the sound from the 0° to 90° direction will attenuate depending on the angle. For applications such as hearing aids that use two or more microphones to simulate the sound pickup effect of the human ear, this directional retention signal processing method may be undesirable. In real life, the structure of the human ear's auricle has an auxiliary sound pickup effect, making it easier to pick up sound from the front than from the back, and it has different effects on different frequencies. Therefore, to achieve the effect of simulating the human ear's auricle in hearing aids, a beamforming method that can customize the adjustment of sound from different directions is needed. Furthermore, it would be desirable for this method to also be able to specifically adjust for different frequencies of sound.

[0044] This application proposes an algorithm that can control the attenuation level and / or the attenuation level of signals at different frequencies with low power consumption, so that applications based on the algorithm are more in line with the auditory perception of the human ear.

[0045] Figure 3A schematic diagram of a beamforming algorithm according to an embodiment of this application is shown. (Distinguished from the above regarding...) Figure 1 and Figure 2 The described scheme will vary depending on the configuration of the iteration coefficients of the LMS filter 303 in the beamforming algorithm of some examples of this application: the coefficient μ is set to a constant value in the above equation (1), while the coefficient μ is set to a function μ(t) that can change with time in the beamforming algorithm of some examples of this application, and in some examples, different functions μ1(t), μ2(t), ... can be set for different frequencies (or frequency bands). The setting of this coefficient will be described in detail below.

[0046] like Figure 3 As shown, compared to Figure 2 The proposed scheme Figure 3 A delay unit 305 is added. The delay unit 305 can delay a series of coefficients U for a period of time (referred to as the update interval, denoted as T, in the context of this application) and then use them to calculate the attenuation function μ(t) for the LMS filter 303, thereby updating the parameters of the LMS filter 303. As described below, the coefficients U can be the values ​​of the attenuation function μ(t) within the first update interval, and the delay unit 305 can delay and output this portion of coefficients U multiple times. This portion of coefficients U is also referred to as the reduction coefficients U in the context of this application.

[0047] According to some examples of this application, after each update interval, the beamforming attenuation coefficient U is iterated again to form a time-varying attenuation function μ(t). In this way, the intensity of sound signal attenuation can be controlled, thereby preventing excessive suppression of sound in non-target directions. Figure 5 A schematic diagram of a beamforming algorithm according to an embodiment of this application is shown. Figure 5 As shown, curves A, B, and C represent the reduction coefficient U updated in time periods #1, #2, and #3, respectively. Figure 5 Curves A, B, and C shown have the same shape, indicating that the reduction factor U is the same across time periods #1, #2, and #3. Specifically, the reduction factor U represented by curve A is the initial part of the decay function μ(t), and can be expressed through parameters such as... Figure 3 The delay unit 305 shown continuously updates and replicates curve A at an update interval T, resulting in curves B and C, as well as subsequent curves (not shown). This update and replication process is equivalent to delaying curve A multiple times and then outputting the result.

[0048] On the other hand, to maintain the continuity of the audio attenuation function μ(t), the updated reduction coefficient U is not applied immediately; it is gradually applied to the attenuation function μ(t) after a delay of an update interval T. For example... Figure 5 As shown, the decay coefficient U of the previous update replication will be applied to the next update interval. Specifically, the updated curves A, B, and C generated in time periods #1, #2, and #3 will be applied to time periods #2, #3, and #4, respectively, to form the corresponding curves A', B', and C'. Curves A', B', and C' will serve as the corresponding parts of the decay function μ(t).

[0049] The values ​​of the decay function μ(t) at each point within the current update interval can be updated based on the value of the corresponding point in the decay coefficient U. For example, a weight between 0 and 1 can be assigned to the value of the corresponding point in the decay coefficient U. In this way, the updated values ​​at each point within the current update interval will be limited to a controllable range. It should be noted that, in the context of this application, each point within the current update interval and its corresponding point in the decay coefficient U are specified in a one-to-one correspondence over time. In some examples, the assigned weights within the current update interval can be linear functions of time. In other examples, the assigned weights within the current update interval can be non-linear functions of time.

[0050] As mentioned above, in some examples, the weights assigned to the decay function μ(t) can be either linear or nonlinear functions of time. For example, when the weights are linear functions of time (linearly increasing functions), the decay function μ(t) can be expressed by equation (3):

[0051]

[0052] Where N represents the number of updates most recent to the current time point. For example, within time period #3 (2T to 3T), the decay function μ(t) can be expressed by equation (4):

[0053]

[0054] As can be seen from equations (3) and (4) above, setting the weights as a linearly increasing function with respect to time can, to some extent, offset the "over-convergence" characteristic of μ(tN*T), thus providing a compensation mechanism.

[0055] In some examples, the weights assigned to the decay function μ(t) can be nonlinear functions with respect to time. For example, the time-dependent decay function μ(t) can be expressed as:

[0056]

[0057] Where N represents the number of the most recent update from the current time point.

[0058] The mathematical description of the decay function μ(t) above will help in understanding the generation mechanism of μ(t), but the generation method of the decay function μ(t) in the real world can still be aided by... Figure 3 The delay unit 305 is shown in the figure. From equation (4) above, it can be seen that the value of μ(t) in the range (2T, 3T] is related to the value of μ(t) in (0, T] and the value of μ(t) at the end of the previous update interval, μ(2T). Therefore, the value of μ(t) in the range (2T, 3T] (or, the shape of curve B') is related to the value of μ(t) in (0, T] (or, the shape of curve A). Since... Figure 5 Curves A, B, and C are updated within time periods #1, #2, and #3, respectively. Therefore, the shape of curve B is consistent with that of curve A; in other words, the shape of curve B' is related to that of curve B. Curve B is an update copy of curve A within time period 2#, thus, within time periods 2T to 3T, the updated coefficients can be used to adjust the LMS filter 303. The continuous copying and updating of the curve within the update interval T will cause the attenuation function μ(t) to be generated and updated according to the update interval T, thereby avoiding excessive suppression of sound in non-target directions due to over-convergence of the filter. On the other hand, since the value of μ(t) in the range (2T, 3T] is related to the value of μ(t) at the end of the previous update interval, μ(2T), μ(t) will not experience drastic changes around time 2T. The smoothness of μ(t) can prevent users of hearing aids from being bothered by unexpected fluctuations in volume.

[0059] As mentioned above, curves B and C are copies of curve A, therefore the decay coefficient can have the same value (the starting value of curves B and C) at the beginning of each predetermined update interval. In other examples, curves B and C can also be fine-tuned relative to curve A, in which case the decay coefficient can have different values ​​(the starting value of curves B and C) at the beginning of each predetermined update interval.

[0060] Furthermore, due to factors such as the human ear's auricle, the human ear responds differently to sounds of different frequencies in different directions. Therefore, it is also expected that beamforming algorithms can provide different responses to sounds of different frequencies. In some examples of this application, the aforementioned response adjustment can be achieved by setting different update intervals for sound signals of different frequencies. For example, by setting separate update intervals for low-frequency and high-frequency sounds, the attenuation of low-frequency and high-frequency sounds can be controlled separately, thereby simulating the frequency response of the human ear's auricle.

[0061] Figure 6 A schematic diagram of a beamforming algorithm according to an embodiment of this application is shown. Figure 6 As shown, an update interval T1 = 5T0 can be configured for low-frequency sounds (e.g., frequencies less than 4000Hz), while an update interval T2 = T0 can be configured for high-frequency sounds (e.g., frequencies greater than or equal to 4000Hz). The update interval T1 for low-frequency sounds is greater than that for high-frequency sounds, so that the attenuation function μ(t) exhibits stronger suppression of low-frequency sounds. This is because low-frequency sounds have better diffraction capabilities than high-frequency sounds, and low-frequency sounds emitted from sound sources outside the target direction are more likely to propagate to the microphone than high-frequency sounds. Furthermore, this configuration also better suppresses low-frequency noise from non-target directions.

[0062] In other examples, the threshold for distinguishing low-frequency and high-frequency sounds can be any frequency other than 4000Hz, or it can be customized based on, for example, different hearing aid wearers to better suit their physiological characteristics. These customized thresholds can be determined, for example, through actual testing or statistical data. In other examples, other methods can be used to distinguish low-frequency and high-frequency sounds, and the methods are not limited to dividing the audible frequency into two intervals. Correspondingly, the number of attenuation functions is not limited to two. For example, audio can be divided into three intervals—low-frequency sounds (e.g., frequencies less than 2000Hz), mid-frequency sounds (e.g., between 2000Hz and 6000Hz), and high-frequency sounds (e.g., frequencies greater than or equal to 6000Hz)—using thresholds of 2000Hz and 6000Hz. Different update intervals can be configured for each interval. For example, an update interval T3 = 5T0 can be configured for low-frequency sounds, an update interval T4 = 3T0 for mid-frequency sounds, and an update interval T5 = T0 for high-frequency sounds.

[0063] In some examples of this application, the hearing aid device is adapted to be worn inside the ear, for example, one microphone in the hearing aid may be oriented toward the ear and the other microphone may be oriented away from the ear.

[0064] Figure 4 An audio enhancement method 40 according to an embodiment of this application is illustrated. The audio enhancement method 40 includes the illustrated steps S402, S404, S406, and S408. It should be noted that, although... Figure 4 The diagram illustrates one possible sequence, but the execution of steps S402, S404, S406, and S408 is not limited to this; steps S402, S404, S406, and S408 can also be executed in other possible sequences. The following will focus on... Figure 4The working principle of steps S402, S404, S406 and S408 of the mid-frequency enhancement method 40 is also cited in the above text along with the corresponding examples in other figures, and will not be repeated here due to space limitations.

[0065] like Figure 4 As shown, in step S402, the audio enhancement method 40 generates an audio acquisition signal. In some examples, as described above, the sound emitted by a sound source 101 can be picked up by microphones 102-1 and 102-2, such as those of a hearing aid. Microphones 102-1 and 102-2 can be positioned on the left and right sides of the hearing aid wearer 103, with a distance d between them. For example, the distance d can depend on the interaural distance of the wearer 103. The wearer 103 faces the device at an angle of 0° as illustrated. Figure 1 Above the wearer 103. Sound source 101 is located to the left front of the wearer 103, forming an angle θ with the center line of the wearer 103's field of vision. Since the distance between sound source 101 and the wearer 103 (and both ears) is much greater than the distance between the ears, it can be assumed that sound source 101 forms an angle θ with respect to both microphones 102-1 and 102-2 as shown in the figure. From geometric relationships, assuming the speed of sound in air is v, and the signal received by microphone 102-1 is y1(t), then the signal received by microphone 102-2 is y2(t) = y1(t-τ), where τ = (d*sin(θ)) / v.

[0066] Short-time Fourier transforms are performed on the signals received by microphones 102-1 and 102-2, respectively. Let the transform result of y1(t) be Y1(k,l) and the transform result of y2(t) be Y2(k,l), where k represents the frequency bin and l represents the frame index. The generated audio acquisition signals Y1(k,l) and Y2(k,l) will satisfy the following relationship:

[0067] Y2(k,l)=Y1(k,l)*e -jωτ .

[0068] In step S404 of audio enhancement method 40, the audio acquisition signal is subjected to delay and summation processing. (Proceed to...) Figure 3 As described above, the delayed beamformer 201 can receive and process signals from microphones 102-1 and 102-2. In some embodiments, the signal Y obtained after processing by the delayed beamformer 201... DSB For example, it can satisfy

[0069] In step S406 of the audio enhancement method 40, the audio acquisition signal undergoes blocking matrix processing. (Continue to refer to...) Figure 3As described above, the blocking matrix 202 can receive and process signals from microphones 102-1 and 102-2. In some schemes, the signal Y obtained after processing by the blocking matrix 202... BM For example, it can satisfy Y BM =Y1(k,l)-Y2(k,l)e jωτ .

[0070] In step S408 of the audio enhancement method 40, the blocking matrix signal Y is... BM (k,τ) is filtered. See also... Figure 3 As described above, the parameter-adjustable LMS filter 303 will affect Y. BM The signal Y is further processed and the processed result is sent to the summing unit 204. OUT (k,l) satisfies Among them W ANC (k,l) are the iteration coefficients of the LMS filter 303, and * indicates conjugate.

[0071] Furthermore, W ANC (k,l) satisfies the following relationship defined by equations (5) and (6):

[0072]

[0073] P est (k,l)=αP est (k,l-1)+(1-α)(|Y BM (k,l)| 2 +|Y OUT (k,l)| 2 (6)

[0074] The decay function μ(t) satisfies the relationship defined in equation (3). As described above, the delay unit 305 enables μ(t) to be updated at a predetermined update interval T, which will not be elaborated further here.

[0075] Figure 7 , Figure 8 and Figure 9 They are shown respectively in Figure 1The beamforming algorithm according to some examples of this application was tested in three directions: 90°, 0°, and -90°. As shown in the figure, the beamforming algorithm according to some examples of this application can obtain the frequency response curve of the beamforming shown in the figure based on the frequency response curves of microphones 1 and 2 in the microphone array, and the obtained frequency response curve matches the frequency response curve of the real human ear quite well. The simulation results show that the frequency response curve obtained by the beamforming algorithm does not over-suppress any specific direction; therefore, the beamforming algorithm according to some examples of this application has good adaptability to applications that need to simulate the response characteristics of the human ear. The beamforming algorithm according to some examples of this application, while effectively suppressing noise, also takes into account the response characteristics of the human ear, making it particularly suitable for applications such as hearing aids that require a faithful reflection of the physical world.

[0076] Another aspect of this application proposes an audio enhancement device comprising a non-transitory computer storage medium storing one or more executable instructions, which, when executed by a processor, perform any of the audio enhancement methods described above. In some examples, this audio enhancement device may be a hearing aid device.

[0077] Another aspect of this application proposes a non-transitory computer storage medium storing one or more executable instructions, which, when executed by a processor, perform any of the audio enhancement methods described above.

[0078] Embodiments of the present invention can be implemented in hardware, software, or a combination of both. The hardware portion can be implemented using dedicated logic; the software portion can be stored in memory and executed by a suitable instruction execution system, such as a microprocessor or dedicated-design hardware. Those skilled in the art will understand that the above-described devices and methods can be implemented using computer-executable instructions and / or included in processor control code, for example, on a carrier medium such as a disk, CD, or DVD-ROM, a programmable memory such as read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of the present invention can be implemented using hardware circuitry such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, or programmable hardware devices such as field-programmable gate arrays, programmable logic devices, etc., or using software executed by various types of processors, or using a combination of the above-described hardware circuitry and software, such as firmware.

[0079] It should be noted that although several steps or modules of the audio enhancement method, apparatus, and storage medium have been mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, according to embodiments of this application, the features and functions of two or more modules described above can be embodied in one module. Conversely, the features and functions of one module described above can be further divided and embodied by multiple modules.

[0080] Those skilled in the art can understand and implement other modifications to the disclosed embodiments by studying the specification, the disclosure, the drawings, and the appended claims. In the claims, the word "comprising" does not exclude other elements and steps, and the words "a" or "an" do not exclude a plurality. In practical applications of this application, a single part may perform the function of multiple technical features referenced in the claims. Any reference numerals in the claims should not be construed as limiting the scope.

Claims

1. An audio enhancement method, characterized by, The method comprises: generating a set of audio acquisition signals by a microphone array, wherein each audio acquisition signal in the set of audio acquisition signals is generated by one microphone in the microphone array, and each microphone in the microphone array is spaced apart from each other; delayed sum processing is performed on the set of audio acquisition signals to generate a delayed sum signal Y DSB (k, l), where k denotes a frequency window and l denotes a frame index; The set of audio acquisition signals is subjected to a blocking matrix processing to generate a blocking matrix signal Y BM (k,l); Utilizing an adaptive filter matrix W ANC filtering the blocking matrix signal Y BM (k,l) and removing the filtered blocking matrix signal from the delayed sum signal Y DSB (k,l) to obtain an enhanced audio output signal Y OUT (k,l); wherein the adaptive filter matrix W ANC is a weight coefficient matrix that varies based on at least one decay function μ(t), the audio output signal Y OUT (k,l) and the blocking matrix signal Y BM (k,l), and each of the at least one decay function μ(t) is updated at a corresponding predetermined update interval T.

2. The method of claim 1, wherein, the microphone array comprises at least two microphones located on the same audio processing device.

3. The method of claim 2, wherein, The audio processing device is adapted to be worn in the pinna of a human ear.

4. The method of claim 3, wherein, One of the at least two microphones is oriented towards the pinna, while the other of the at least two microphones is oriented away from the pinna.

5. The method of claim 1, wherein, The audio output signal is determined by the following equation: And, the adaptive filter matrix W ANC is determined by the equation: Among them, P est (k,l) is determined by the following equation: wherein a is a forgetting factor, and M is the number of microphones in the microphone array.

6. The method of claim 1, wherein, The at least one attenuation function comprises a first attenuation function and a second attenuation function, the first attenuation function is updated at a first predetermined update interval, and the second attenuation function is updated at a second predetermined update interval; wherein the first attenuation function corresponds to high frequency signals greater than or equal to a predetermined frequency threshold, and the second attenuation function corresponds to low frequency signals less than the predetermined frequency threshold, and the first predetermined update interval is shorter than the second predetermined update interval.

7. The method of claim 1, wherein, Each of the attenuation functions μ(t) is updated in the current update interval based on its value in the first update interval.

8. The method of claim 7, wherein, Each of the attenuation functions μ(t) is updated in each point in the current update interval based on the value of the corresponding point in the first update interval, and is assigned a change weight between 0 and 1.

9. The method of claim 8, wherein, The weight is a linear function of time in the current update interval.

10. The method of claim 9, wherein, The weight is a linearly increasing function of time in the current update interval.

11. The method of claim 8, wherein, The weight is a non-linear function of time in the current update interval.

12. The method of claim 9 or 10, wherein, Each of the attenuation functions μ(t) is also updated in the current update interval based on its value at the end of the last update interval.

13. The method of claim 12, wherein, Each of the attenuation functions μ(t) satisfies the following equation in the current update interval (NT, (N+1)T]: where N takes a positive integer.

14. An audio enhancement device, characterized by The device comprises a non-transitory computer storage medium having stored thereon one or more executable instructions which, when executed by a processor, perform the following steps: generating a set of audio acquisition signals by a microphone array, wherein each audio acquisition signal in the set of audio acquisition signals is generated by one microphone in the microphone array, and each microphone in the microphone array is spaced apart from each other; delayed sum processing is performed on the set of audio acquisition signals to generate a delayed sum signal Y DSB (k, l), where k denotes a frequency window and l denotes a frame index; The set of audio acquisition signals is subjected to a blocking matrix processing to generate a blocking matrix signal Y BM (k,l); Utilizing an adaptive filter matrix W ANC filtering the blocking matrix signal Y BM (k,l) and removing the filtered blocking matrix signal from the delayed sum signal Y DSB (k,l) to obtain an enhanced audio output signal Y OUT (k,l); wherein the adaptive filter matrix W ANC is a weight coefficient matrix that varies based on at least one decay function μ(t) with the audio output signal Y OUT (k,l) and the blocking matrix signal Y BM (k,l), and each of the at least one decay function μ(t) is updated at a corresponding predetermined update interval T.

15. The apparatus of claim 14, wherein, The device is a hearing aid.

16. A non-transitory computer storage medium having stored thereon one or more executable instructions which, when executed by a processor, perform an audio enhancement method, the method comprising the following steps: generating a set of audio acquisition signals by a microphone array, wherein each audio acquisition signal in the set of audio acquisition signals is generated by one microphone in the microphone array, and each microphone in the microphone array is spaced apart from each other; delayed sum processing is performed on the set of audio acquisition signals to generate a delayed sum signal Y DSB (k, l), where k denotes a frequency window and l denotes a frame index; The set of audio acquisition signals is subjected to a blocking matrix processing to generate a blocking matrix signal Y BM (k,l); Utilizing an adaptive filter matrix W ANC filtering the blocking matrix signal Y BM (k,l) and removing the filtered blocking matrix signal from the delayed sum signal Y DSB (k,l) to obtain an enhanced audio output signal Y OUT (k,l); wherein, The adaptive filter matrix W ANC is based on at least one decay function μ(t), which varies with the audio output signal Y OUT (k,l) and the blocking matrix signal Y BM (k,l), and each of the at least one decay function μ(t) is updated at a corresponding predetermined update interval T.

Citation Information

Patent Citations

  • Adaptive beam forming method for reducing voice distortion

    CN106653043A

  • Signal enhancement method and device, computer readable storage medium and electronic equipment

    CN110689900A