A digital hearing aid automatic gain control method and system
By acquiring and analyzing Mel speech spectrograms in hearing aids and combining them with adaptive step size adjustment using the natural gradient algorithm, effective separation of target sound source signals and noise signals is achieved, improving speech clarity and comfort of digital hearing aids in noisy environments.
Patent Information
- Application Number
- CN202511451378.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-11
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2045-10-11
AI Technical Summary
Existing digital hearing aids have difficulty effectively separating the target sound source signal from the ambient noise signal when the wearer communicates with others, resulting in a decrease in the accuracy of speech signal transmission and affecting the user experience of hearing-impaired patients.
By acquiring sound signals through the built-in microphone of the hearing aid, obtaining the Mel speech spectrogram, analyzing the energy concentration and noise salience, and combining the iterative step size adaptive adjustment of the natural gradient algorithm, blind source separation and automatic gain control are achieved.
It improves the clarity of the target signal and the applicability of hearing aids, enhances the auditory effect in noisy environments, and improves the auditory experience of hearing-impaired patients.
Smart Images

Figure CN120980427B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of hearing aids, in particular to a digital hearing aid automatic gain control method and system. BACKGROUND
[0002] A hearing aid is an effective device for improving the hearing level of a hearing-impaired patient, helping the patient to hear clear and complete sounds from the outside world. With the development of digital signal processing technology, a digital hearing aid has excellent sound signal processing capability, and an automatic gain control method built-in the digital hearing aid can process the collected sound signals such as noise reduction, compression, compensation, and direction positioning, so as to reduce environmental noise and make the wearer accurately identify and enhance the required sound signals, thereby improving the hearing ability and comfort of the hearing-impaired patient.
[0003] The automatic gain control method in the hearing aid can use linear amplification for small and medium sounds, and use compression amplification for sounds above 65 dB (medium) and with a large volume. However, in actual application scenarios, the wearer will usually include the influence of ambient noise signals in the collected speech signals of the other party during communication with the other party, which affects the accuracy of speech signal transmission. The effect of the automatic gain control of the hearing aid depends on the noise reduction effect of the sound signal. If the target sound source signal and the noise signal can be separated, the automatic gain control effect of the hearing aid can be effectively improved. If the wearer of the hearing aid is interested in the speech of a certain person, he or she will mainly hear the sound of the person and ignore other sounds around him or her. Such a process of extracting an unknown source signal from a large number of mixed signals is called blind source separation (BSS), which provides technical support for speech signal enhancement and improves the sound quality of the hearing aid.
[0004] The natural gradient algorithm is a commonly used method in blind source separation, which has a fast convergence speed and good separation performance. However, the step size of the natural gradient algorithm is fixed. If the step size is too large, the steady-state error is also large. If the step size is too small, the convergence speed is too slow. In the early stage of sound signal separation, the correlation is generally large, and a large step size can be used. As the separation degree becomes larger, the correlation between the sound signals is smaller, and a smaller step size should be used. Therefore, if the natural gradient algorithm with a fixed step size is used, the separated sound source signal and the environmental noise signal will have poor effects, and cannot meet the real-time requirements of the hearing aid, thereby affecting the experience of the hearing-impaired person using the hearing aid. SUMMARY
[0005] To solve the above technical problems, the purpose of the present application is to provide a digital hearing aid automatic gain control method and system, and the technical scheme adopted is as follows:
[0006] In a first aspect, the embodiments of the present application provide a digital hearing aid automatic gain control method, which comprises the following steps:
[0007] acquiring a mel-spectrogram of the sound signal;
[0008] dividing each frame of the mel-spectrogram into mel-frequency bands;
[0009] analyzing differences between high-frequency energy and low-frequency energy in each frame of the mel-spectrogram, and energy distribution in each frame of the mel-spectrogram, and determining a noise salience of each frame of the mel-spectrogram based on the energy concentration degree;
[0010] performing blind source separation on the acquired sound signal using a natural gradient algorithm, determining a separation difference degree of each iteration based on similarity between a target signal and a noise signal obtained after each iteration, similarity between the acquired sound signal and the separated signal, and the noise salience of all frames of the mel-spectrogram, to determine a step size of the next iteration;
[0011] performing automatic gain control on the target signal obtained by the blind source separation.
[0012] In one embodiment, the determining of the energy concentration degree of each mel-frequency band includes:
[0013] calculating a product of an energy value of each coordinate point in each mel-frequency band and a local density thereof, determining a fusion result of the product of all coordinate points in each mel-frequency band, denoted as a first fusion value, and determining a fusion result of a metric distance between all adjacent coordinate points in each mel-frequency band, denoted as a second fusion value;
[0014] determining the energy concentration degree based on the first fusion value and the second fusion value.
[0015] In one embodiment, the energy concentration degree is positively correlated with the first fusion value and negatively correlated with the second fusion value.
[0016] In one embodiment, the determining of the noise salience of each frame of the mel-spectrogram includes:
[0017] dividing all mel-frequency bands of each frame of the mel-spectrogram into low-frequency and high-frequency parts, and calculating a ratio between an energy sum value of all coordinate points of all mel-frequency bands of the high-frequency and an energy sum value of all coordinate points of all mel-frequency bands of the low-frequency;
[0018] determining an energy mean value of all coordinate points in each frame of the mel-spectrogram;
[0019] Determine a fusion result of the energy concentration degrees of all the mel frequency bands in each frame of mel spectrogram, and mark it as a third fusion value;
[0020] The noise prominence is positively correlated with the ratio, the energy mean value, and the third fusion value.
[0021] In one embodiment, the noise prominence is a normalized result of the product of the ratio, the energy mean value, and the third fusion value.
[0022] In one embodiment, the obtaining of the separation difference degree of each iteration includes:
[0023] Determine a proportion of the iteration number corresponding to each iteration of the natural gradient algorithm in the total preset iteration number, and calculate an average value of the noise prominences of all the frames of mel spectrogram;
[0024] Mark a similarity degree of the target signal and the noise signal obtained after the separation of each iteration as a first similarity degree, and mark a fusion result of the similarity degrees of the collected sound signal and all the separated signals as a fourth fusion value;
[0025] Determine the separation difference degree by the proportion, the average value, and the first similarity degree and the fourth fusion value.
[0026] In one embodiment, obtain an inverse proportional mapping result of the first similarity degree and the fourth fusion value, and positively fuse the proportion and the average value to obtain the separation difference degree.
[0027] In one embodiment, the determining of the step length of the next iteration includes:
[0028] Multiply a normalized result of the separation difference degree of each iteration by a preset step length adjustment sensitivity, combine the step length of each iteration, and obtain the step length of the next iteration.
[0029] In one embodiment, calculate a difference value between a natural number 1 and the multiplied result, and the step length of the next iteration is a product of the step length of each iteration and the difference value.
[0030] In a second aspect, the embodiments of the present application further provide a digital hearing aid automatic gain control system, including a memory, a processor, and a computer program stored in the memory and running on the processor, and the processor implements the steps of the method in any one of the above aspects when executing the computer program.
[0031] The present application has at least the following beneficial effects:
[0032] The application acquires a sound signal by a built-in microphone of a hearing aid; acquires a mel-spectrogram of the sound signal; divides each frame of the mel-spectrogram into each mel-frequency band, determines the energy concentration degree of each mel-frequency band, accurately quantifies the energy distribution in each mel-frequency band, helps to analyze the noise content in each frame of the mel-spectrogram, and improves the reliability of subsequent noise salience calculation; further, determines the noise salience of each frame of the mel-spectrogram, reflects the noise content in each frame of the mel-spectrogram, and improves the accuracy of noise extraction in each frame of the mel-spectrogram; the similarity degree of the target signal and the noise signal obtained after each iteration of separation by the natural gradient algorithm, and the similarity degree of the collected sound signal and the separated signal, combined with the noise salience of all frames of the mel-spectrogram, obtain the separation difference degree of each iteration to determine the step size of the next iteration, so that the step size of each iteration integrates the signal separation effect of the last iteration, improves the step size determination accuracy and suitability in the natural gradient algorithm, thereby further improving the accuracy of blind source separation, avoiding signal distortion, optimizing the automatic gain control effect of the hearing aid, improving the clarity of the target signal, improving the applicability and comfort of the hearing aid, helping to provide a more natural hearing effect for the user, effectively improving the performance of the hearing aid in a noisy environment, and solving the problem that the user cannot clearly hear the conversation or surrounding sound in a complex environment. BRIEF DESCRIPTION OF DRAWINGS
[0033] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present application or prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without any creative effort.
[0034] Figure 1 A step flow chart of a digital hearing aid automatic gain control method provided by an embodiment of the present application;
[0035] Figure 2 A flow chart for determining the step size of the natural gradient algorithm iteration. DETAILED DESCRIPTION
[0036] In order to further illustrate the technical means and effects adopted by the present application to achieve the predetermined invention purpose, the following will combine the drawings and preferred embodiments to specifically describe the specific implementation, structure, features and effects of the digital hearing aid automatic gain control method and system according to the present application. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.
[0037] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs.
[0038] The specific scheme of the digital hearing aid automatic gain control method and system provided by the application will be specifically described below in combination with the drawings.
[0039] Please refer to Figure 1 , which shows the step flow chart of a digital hearing aid automatic gain control method provided by an embodiment of the application, which comprises the following steps:
[0040] S1, collect the sound signal through the built-in microphone of the hearing aid; obtain the mel spectrogram of the sound signal.
[0041] The hearing aid device mainly includes a microphone, a digital-to-analog conversion, a digital processing chip, and a speaker structure. When the hearing aid wearer talks with a person, the hearing aid first collects the received sound signal through the built-in two microphones, and the sound signal includes the sound signal of the other party and the noise signal of the surrounding environment, so the sound signal collected by the microphone is recorded as a mixed signal; the sampling frequency is 32 kHz, which can be set by the implementer according to the actual situation. Then, the collected mixed signal is converted into a digital signal through AD analog-to-digital conversion. Since the actual conversation environment is complex and diverse, there are low-frequency noise (such as wind noise) and high-frequency noise (such as electronic interference), and the digital signal uses the embedded Butterworth filter during the digital processing chip to preliminarily remove the noise influence, and outputs the digital signal containing the sound source signal and the environmental noise signal, which is recorded as a mixed digital signal. The mixed digital signal will be separated in the subsequent process to separate the target signal, i.e. the sound source signal, which the wearer wants to hear, and mainly to automatically control the gain of the target signal to improve the sound quality. Among them, AD analog-to-digital conversion and Butterworth filter are both prior art known technologies, and the specific process will not be described in detail.
[0042] When the hearing aid wearer talks with a person outdoors, there are often some noises around, such as car horn sound, wind noise, etc., which cause the sound signal collected by the hearing aid to contain a large amount of noise signal, which can be simply recorded as , wherein is the observation signal, i.e. the mixed digital signal collected by the microphone, A is the mixing matrix, which represents the mixing of the signal, is the source signal, i.e. the actual multi-source sound signal in the mixed digital signal, in this embodiment is the actual target signal and noise signal. The basis for blind source separation is only to obtain the separation matrix W according to the observation signal , W is the inverse matrix of A, so that the transformed output is the source signal An estimate.
[0043] Since the target signal and noise signal are independent of each other, the energy of the sound signal generated by the same person speaking will usually be concentrated in the same frequency range on the vertical axis of the Mel speech spectrogram, and this frequency range changes smoothly over time. The energy distribution of adjacent time frames has a strong correlation, and the color changes smoothly. However, the noise signal that suddenly appears in the surrounding area will have higher energy, and the energy distribution on the vertical axis of the Mel speech spectrogram will be wider, significantly exceeding the high-frequency area covered by the normal speaking sound signal. When the noise signal appears at a certain moment, the Mel speech spectrogram will show a bright area at that moment, that is, a high energy, which usually lasts for a certain period of time on the horizontal axis, forming a short-term energy distribution segment. Blind source separation is required for the mixed digital signal during this period.
[0044] Based on the above analysis, the Mel spectrogram of the mixed digital signal is obtained. The mixed digital signal is processed by frame-by-frame windowing with a frame length of 10ms, a step size of 5ms, and a Hamming window function. Short-time Fourier transform is used for spectrum calculation, and RGB values are used for amplitude quantization. The Mel-language spectrogram is converted to a known technique; specific implementation details are not elaborated here. In the Mel-language spectrogram, the horizontal axis represents time (frame number), and the vertical axis represents frequency. The coordinates of the i-th row and j-th column of the Mel-language spectrogram are... The amplitude corresponding to this coordinate in the Melan spectrum is denoted as the energy value. , representing the signal energy at frequency j in the mixed digital signal of the i-th frame, i.e., the sound energy value, is represented by the Mel spectrogram. The i-th frame is denoted as In obtaining the Mel language spectrogram The implementer can set the frame length and step size according to the actual situation, and this embodiment does not impose any restrictions on this.
[0045] S2, divide each frame of Mel spectrogram into Mel frequency bands; determine the energy concentration of each Mel frequency band by the energy value and local density of each coordinate point in each Mel frequency band, combined with the metric distance between adjacent coordinate points in each Mel frequency band.
[0046] When two people are talking normally, the fluctuations in their voices are relatively small, and the corresponding sound signal frequency and energy will also fluctuate within a certain range. However, sudden sharp noise usually has a large sound energy value, and the sound signal shows obvious peak fluctuations. It has more high-frequency energy in the Mel spectrogram, and the energy density at each coordinate point is greater. Based on this, the distribution area of sharp noise in the Mel spectrogram can be determined.
[0047] Based on the above analysis, this embodiment uses the density peak clustering algorithm to calculate the coordinates of each coordinate point in the Melanographic spectrogram of each frame. The local density of the corresponding energy value is denoted as a local energy density, wherein the difference of the energy values between each coordinate point is taken as the distance between the coordinate points in the density peak clustering algorithm, which is a known technology and will not be described in detail. Then, the corresponding mel spectrum of each frame is divided into N mel frequency bands The mel frequency bands are numbered from low frequency to high frequency, and each frequency band is denoted as , In this embodiment, N = 20, and the implementer can set it according to the actual situation, which is not limited in this embodiment.
[0048] Based on this, the energy concentration of each mel frequency band in each frame of mel spectrum is calculated, which is used to represent the distribution state of the energy in the mel frequency band, and the specific process is as follows:
[0049] The product of the energy value of each coordinate point in each mel frequency band and the local density thereof is calculated, the fusion result of the products of all coordinate points in each mel frequency band is determined, and the first fusion value is denoted, and the fusion result of the metric distance between all adjacent coordinate points in each mel frequency band is determined, and the second fusion value is denoted.
[0050] Based on the first fusion value and the second fusion value, the energy concentration is determined, wherein the energy concentration is positively correlated with the first fusion value and negatively correlated with the second fusion value.
[0051] It should be noted that fusion means combining multiple variables, which can be calculated by addition, multiplication, addition-multiplication mixing, and mean value calculation.
[0052] In this embodiment, the expression of the energy concentration of each mel frequency band in each frame of mel spectrum is as follows:
[0053] ; in the formula, denotes the energy concentration of the nth mel frequency band of the ith frame of mel spectrum, denotes the number of coordinate points contained in each mel frequency band, denotes the kth coordinate point, denotes the local energy density of the kth coordinate point in the nth mel frequency band of the ith frame of mel spectrum, denotes the energy value corresponding to the kth coordinate point in the nth mel frequency band of the ith frame of mel spectrum, , denote the kth and k+1th coordinate points in the nth mel frequency band of the ith frame of mel spectrum, respectively, represents the calculation of the metric distance, and the metric distance in the embodiment is calculated by using the Euclidean distance. The implementer can select other feasible metric distance calculation methods, and the embodiment does not limit this. denoted as a first fusion value, denoted as a second fusion value.
[0054] S3, the difference between the high-frequency energy and the low-frequency energy in each frame of the mel spectrogram is analyzed, and the energy distribution in each frame of the mel spectrogram is analyzed, and the noise saliency of each frame of the mel spectrogram is determined in combination with the energy concentration degree.
[0055] Further, the noise saliency of each frame of the mel spectrogram is calculated, which is used to represent the possibility of containing sharp noise in each frame of the mel spectrogram, and specifically:
[0056] All mel frequency bands of each frame of the mel spectrogram are divided into low-frequency and high-frequency two parts. In the embodiment, each frame of the mel spectrogram is divided into N mel frequency bands, and the N mel frequency bands are arranged in order from low frequency to high frequency. Therefore, the first N / 2 mel frequency bands are marked as low-frequency frequency bands, and the last N / 2 mel frequency bands are marked as high-frequency frequency bands.
[0057] In the embodiment, the expression of the noise saliency of each frame of the mel spectrogram is:
[0058] ; in the formula, denotes the noise saliency of the i-th frame of the mel spectrogram, denotes a normalization function, denotes the ratio of the energy sum value of all coordinate points of all high-frequency frequency bands in the i-th frame of the mel spectrogram to the energy sum value of all coordinate points of all low-frequency frequency bands, is the average energy of all coordinate points in the i-th frame of the mel spectrogram, denotes the energy concentration degree of the n-th mel frequency band of the i-th frame of the mel spectrogram, and N is the number of all mel frequency bands in each frame of the mel spectrogram. In the embodiment, N=20. The energy concentration degree of each mel frequency band in the i-th frame of the mel spectrogram is denoted as: denoted as a third fusion value.
[0059] The greater the energy concentration degree of each mel frequency band in the i-th frame of the mel spectrogram, the more high-frequency energy, and the greater the noise saliency, which indicates that the i-th frame of the mel spectrogram is more likely to contain sharp noise.
[0060] S4, the natural gradient algorithm is used to perform blind source separation on the collected sound signal, the similarity degree between the target signal and the noise signal obtained after each iteration, and the similarity degree between the collected sound signal and the separated signal are obtained, and the noise saliency of all frames of the mel spectrogram is combined to obtain the separation difference degree of each iteration, so as to determine the step length of the next iteration.
[0061] In the natural gradient algorithm for blind source separation of digital mixed signals, the greater the noise salience of the overall digital mixed signal, the greater the separation degree of the digital mixed signal with the iteration of the natural gradient algorithm, the greater the difference between the separated target signal and the noise signal after each iteration, and the iteration step should be adaptively updated and reduced to improve the separation accuracy and avoid loss of target signal and distortion of the separated target signal due to excessive step size.
[0062] Based on the above analysis, the separation difference degree of the target signal and the noise signal separated after each iteration of the natural gradient algorithm is calculated, which is used to represent the separation degree and difference degree of the two signals separated based on the digital mixed signal, specifically:
[0063] The proportion of the iteration number corresponding to each iteration of the natural gradient algorithm in the total iteration number is determined, and the average value of the noise salience of all frame mel spectrograms is calculated;
[0064] The similarity between the target signal and the noise signal obtained after each iteration is denoted as the first similarity, and the fusion result of the similarity between the collected sound signal and all separated signals is denoted as the fourth fusion value;
[0065] The inverse proportional mapping result of the first similarity and the fourth fusion value is obtained, and is forward fused with the proportion and the average value to obtain the separation difference degree.
[0066] It should be noted that the similarity between signals can be calculated by Pearson correlation coefficient, cosine similarity, etc.; the inverse proportional mapping represents a mapping mode between dependent variable and independent variable, i.e. the dependent variable will decrease with the increase of the independent variable, and will increase with the decrease of the independent variable, such as reciprocal relationship, negative exponential relationship, etc.
[0067] In this embodiment, the expression of the separation difference degree is:
[0068] In the formula, represents the separation difference degree of the target signal and the noise signal separated after the dth iteration of the natural gradient algorithm, represents the dth iteration, represents the total iteration number, D=100 in this embodiment, and the implementer can set it according to the actual situation; represents the total number of mel spectrogram frames, , represents the target signal and the noise signal separated after the dth iteration of the natural gradient algorithm, respectively, Q represents the total number of signal types separated by the natural gradient algorithm, and in this embodiment, Q=2, indicating that two signals, i.e., the target signal and the noise signal, are separated, represents the mixed digital signal, represents the calculation of the Pearson correlation coefficient, represents the qth signal separated after the dth iteration of the natural gradient algorithm, represents the noise saliency of the ith frame of the mel spectrogram. The noise saliency is calculated as denoted as the first similarity, denoted as the second similarity, denoted as the fourth fusion value.
[0069] It should be understood that, the noise saliency is positively correlated, the more significant the noise in the mel spectrogram, the greater the difference between the target signal and the noise signal separated by the natural gradient algorithm when separating the mixed digital signal, i.e., the greater the separation difference; and as the iteration proceeds, the value of increases, the correlation coefficient between the target signal and the noise signal, and the correlation coefficient between the mixed digital signal and the separated signal gradually decreases, so that the value of increases more and more, indicating that the signal separation degree is getting larger and larger.
[0070] Further, based on the separation difference between the target signal and the noise signal separated after each iteration, the step size of the next iteration of the natural gradient algorithm is adaptively corrected, specifically:
[0071] In the formula, represents the step size of the dth iteration of the natural gradient algorithm, represents the step size of the (d-1)th iteration of the natural gradient algorithm, is a preset step size adjustment sensitivity, and the value range is 0-1, the greater the step size adjustment sensitivity, the greater the change in the step size of the dth iteration relative to the step size of the (d-1)th iteration, and in this embodiment, the implementer can set it according to the actual situation, and this embodiment does not limit it, is the separation difference between the target signal and the noise signal separated after the (d-1)th iteration of the natural gradient algorithm, and Norm() is a normalization function. In this embodiment, the initial step size of the natural gradient algorithm iteration is 0.05, and the implementer can set it according to the actual situation. The step size determination flowchart of the natural gradient algorithm iteration is shown in Figure 2 .
[0072] It should be understood that the greater the separation difference between the target signal and the noise signal separated by the natural gradient algorithm after each iteration, the smaller the iteration step should be in the next iteration to improve the separation accuracy and avoid the loss of the target signal and the distortion of the separated target signal caused by the too large step.
[0073] The natural gradient algorithm with adaptive update step is used to perform blind source separation on the mixed digital signal, and a voice activity detection (VAD) technique is used to identify the target signal in the mixed digital signal based on the result of the blind source separation. The natural gradient algorithm for blind source separation and the voice activity detection (VAD) are both known techniques, and the specific process is not described in detail.
[0074] S5. Automatically controlling the gain of the target signal obtained by the blind source separation.
[0075] The target signal obtained by the blind source separation is first restored to the original sound pressure level range through amplitude correction to obtain an original speech signal, and then gain is applied to the original speech signal according to the automatic gain control requirement of the hearing aid, such as linear amplification for weak sound frequency bands with a gain limit of 20 decibels, and nonlinear compression for strong sound frequency bands with a compression ratio of 1:2, to ensure that the output sound of the hearing aid is clear and comfortable in the ear of the hearing-impaired person. The amplitude correction and the automatic gain control are both known techniques, and the specific implementation details are not described in detail.
[0076] Based on the same inventive concept as the above method, the embodiments of the present application also provide a digital hearing aid automatic gain control system, which includes a memory, a processor, and a computer program stored in the memory and running on the processor, and the processor implements the steps of any one of the above digital hearing aid automatic gain control methods when executing the computer program.
[0077] It should be noted that the above-mentioned embodiments of the present application are only for description, and do not represent the advantages and disadvantages of the embodiments. The above describes specific embodiments of the present application. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multi-task processing and parallel processing are also possible or may be advantageous.
[0078] Each embodiment in the present specification is described in a progressive manner, and the same or similar parts between each embodiment can be referred to each other. Each embodiment focuses on the differences from other embodiments.
[0079] The above only describes the preferred embodiments of the present application, and does not limit the present application. Any modification, equivalent replacement, improvement, etc. made within the principles of the present application shall be included in the protection scope of the present application.
Claims
1. A method for automatic gain control in a digital hearing aid, characterized in that, The method includes the following steps: The sound signal is collected through the built-in microphone of the hearing aid; the Melanographic spectrogram of the sound signal is obtained. Each frame of the Mel spectrogram is divided into Mel frequency bands; the energy concentration of each Mel frequency band is determined by the energy value and local density of each coordinate point in each Mel frequency band, combined with the metric distance between adjacent coordinate points in each Mel frequency band. The differences between high-frequency and low-frequency energies in each frame of the Mel speech spectrogram, as well as the energy distribution in each frame of the Mel speech spectrogram, are analyzed. Combined with the energy concentration, the noise significance of each frame of the Mel speech spectrogram is determined. The natural gradient algorithm is used to perform blind source separation on the acquired sound signal. The separation difference degree of each iteration is obtained by combining the similarity between the target signal and the noise signal obtained after each iteration, the similarity between the acquired sound signal and the separated signal, and the noise saliency of all frames of Mel spectrograms, so as to determine the step size of the next iteration. Automatic gain control is applied to the target signal obtained from blind source separation.
2. The automatic gain control method for a digital hearing aid as described in claim 1, characterized in that, Determining the energy concentration of each Mel frequency band includes: Calculate the product of the energy value and its local density at each coordinate point in each Mel frequency band, determine the fusion result of the product of all coordinate points in each Mel frequency band, and record it as the first fusion value; determine the fusion result of the metric distance between all adjacent coordinate points in each Mel frequency band, and record it as the second fusion value. The energy concentration is determined based on the first fusion value and the second fusion value.
3. The automatic gain control method for a digital hearing aid as described in claim 2, characterized in that, The energy concentration is positively correlated with the first fusion value and negatively correlated with the second fusion value.
4. The automatic gain control method for a digital hearing aid as described in claim 1, characterized in that, Determining the noise saliency of each frame of Mel spectrogram includes: Divide all Mel frequency bands in each frame of Mel spectrogram into low frequency and high frequency parts, and calculate the ratio of the sum of energy and value of all coordinate points in all Mel frequency bands in the high frequency range to the sum of energy and value of all coordinate points in all Mel frequency bands in the low frequency range. Determine the average energy of all coordinate points in each frame of the Mel language spectrogram; The fusion result of the energy concentration of all Mel frequency bands in each frame of Mel spectrogram is determined and denoted as the third fusion value; The noise significance is positively correlated with the ratio, the mean energy, and the third fusion value.
5. The automatic gain control method for a digital hearing aid as described in claim 4, characterized in that, The noise significance is the normalized result of the product of the ratio, the mean energy, and the third fusion value.
6. The automatic gain control method for a digital hearing aid as described in claim 1, characterized in that, The process of obtaining the separation difference degree for each iteration includes: Determine the proportion of the number of iterations corresponding to each iteration of the natural gradient algorithm in the preset total number of iterations, and calculate the average noise significance of all frames of Mel spectrograms; The similarity between the target signal and the noise signal obtained after each iteration is denoted as the first similarity. The fusion result of the similarity between the collected sound signal and all the separated signals is denoted as the fourth fusion value. The separation difference is determined by the proportion, the average value, and the first similarity and the fourth fusion value.
7. The automatic gain control method for a digital hearing aid as described in claim 6, characterized in that, Obtain the inverse proportional mapping result between the first similarity and the fourth fusion value, and perform positive fusion with the proportion and the average value to obtain the separation difference.
8. The automatic gain control method for a digital hearing aid as described in claim 1, characterized in that, The step size for determining the next iteration includes: The normalized result of the separation difference in each iteration is multiplied by the preset step size adjustment sensitivity, and combined with the step size of each iteration, the step size of the next iteration is obtained.
9. The automatic gain control method for a digital hearing aid as described in claim 8, characterized in that, Calculate the difference between the natural number 1 and the result of multiplying them. The step size of the next iteration is the product of the step size of each iteration and the difference.
10. A digital hearing aid automatic gain control system, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1-9.
Citation Information
Patent Citations
Speech enhancement method, device and system based on cross-domain feature fusion
CN119360874A
Voice processing method and device, storage medium and computer equipment
CN119517068A