Hearing aid enhancement strength adaptive adjustment method and system based on voice distortion perception index
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-20
- Publication Date
- 2026-08-11
AI Technical Summary
[0003]然而,单纯依赖上述客观声学指标的调节方式,存在明显的技术缺陷:客观声学指标无法准确反映用户主观听感上的失真风险,导致增强强度调节缺乏对听感的有效约束
本实施例适用于双耳协同助听器系统,核心解决双耳独立调节导致的听感不一致问题,具体实现方式为:
Smart Images

Figure CN122554767A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of digital hearing aids, hearing enhancement devices and their speech processing and control technology, specifically to a method and system for adaptive adjustment of hearing aid enhancement intensity based on speech distortion perception index. This method and system can be deployed on low-power hearing aid platforms that employ frequency domain WOLA / FFT processing chains, traditional noise reduction algorithms, neural network speech enhancement algorithms or combinations thereof. Background Technology
[0002] As hearing enhancement devices, the core need of hearing aids is to improve speech clarity and reduce background noise interference in complex acoustic environments such as restaurants, shopping malls, streets, or multi-person conversations by enhancing speech and suppressing noise. Most existing hearing aid enhancement intensity adjustment schemes determine the enhancement level based on environmental classification, input signal-to-noise ratio, noise energy, or speech activity detection results. Examples include increasing noise reduction gain, increasing neural network masking depth, decreasing minimum spectral gain, or increasing post-spectral filtering intensity.
[0003] However, relying solely on the aforementioned objective acoustic indicators for adjustment has significant technical limitations: objective acoustic indicators cannot accurately reflect the risk of distortion in the user's subjective hearing, resulting in a lack of effective constraints on hearing perception when adjusting the enhancement intensity. When the enhancement intensity is too high, the hearing aid system is prone to producing auditory artifacts such as musical noise, metallic sounds, slurring, envelope breathing, or thin timbre, severely reducing wearing comfort; when the enhancement intensity is too low, it is difficult to effectively improve speech intelligibility, failing to meet the user's listening needs in complex noisy environments.
[0004] Ultimately, existing hearing aids generally lack a negative feedback control loop with "auditory distortion" as the core constraint. They cannot perceive and suppress subjective auditory degradation in real time during speech enhancement. Therefore, in complex environments, it is difficult to simultaneously ensure speech clarity, natural sound quality, and long-term wearing comfort. The stability and robustness of the algorithm also need to be improved. Summary of the Invention
[0005] To address the aforementioned deficiencies in existing technologies, this invention proposes a method and system for adaptive adjustment of hearing aid enhancement intensity based on speech distortion perception indices. The core objective is to assess the distortion risk of enhanced speech in real time or near real time without relying on clean reference speech, and dynamically adjust the enhancement intensity parameters accordingly. This enables the speech enhancement system to effectively suppress the subjective auditory degradation caused by excessive enhancement while improving speech clarity, achieving a controllable balance between clarity and naturalness.
[0006] Specifically, the objectives of this invention include: Reduce the generation of artifacts such as music noise, metallic sounds, sucking sensations, and envelope anomalies; Establish a quantifiable and adjustable balance mechanism between improving speech clarity and natural sound quality; Improve the stability and robustness of speech enhancement algorithms in complex acoustic scenarios, while balancing clarity and naturalness; This enables enhanced control logic to be adapted to the real-time implementation requirements of low-power DSPs, MCUs, or hearing aid-specific SoCs. It is compatible with various speech enhancement architectures such as traditional frequency domain noise reduction and neural network masking enhancement, and also supports various hearing aid system forms such as monoauricular and binauricular collaboration.
[0007] To achieve the aforementioned objectives, the core technical concept of this invention is as follows: abandoning the traditional approach of simply maximizing the signal-to-noise ratio, this invention introduces a referenceless distortion perception index into the speech enhancement processing chain, constructing a negative feedback closed-loop control system of "enhanced output - distortion assessment - parameter update - subsequent enhancement". By extracting perceptual features that characterize subjective distortion risk from the enhanced speech, a comprehensive distortion risk score is calculated. Combined with the dynamic adjustment of the enhancement intensity parameter based on the ambient noise intensity, this prevents algorithm overshoot and achieves adaptive and intelligent adjustment of the enhancement intensity.
[0008] 1. Overall System Structure The hearing aid enhancement intensity adaptive adjustment system based on speech distortion perception index provided by this invention includes at least a microphone acquisition and preprocessing module, a speech enhancement module, a distortion perception index calculation module, an enhancement intensity control module, and a parameter update module. The inputs, outputs, and key implementation points of each module are shown in Table 1 below: Table 1 Core parameters of each module in the system The speech enhancement module can select a traditional frequency domain noise reduction module, a neural network masking enhancement module, a time domain enhancement module, or a combination thereof, according to the needs of the hearing aid platform; the continuous or hierarchical control quantity u(t) output by the enhancement intensity control module is used to adjust core parameters such as noise reduction depth, masking index, minimum gain, spectral smoothing factor, or post-filtering intensity.
[0009] 2. Distortion Perception Index Design To ensure the stability and engineering feasibility of this invention, the distortion perception index is defined as a combination of at least one or more no-reference indices. Preferably, three complementary no-reference distortion perception sub-indices are used to characterize the subjective auditory distortion risk from three dimensions: spectrum, envelope, and harmonic structure. The specific definitions and functions of the three sub-indices are as follows: Scope of application and symbol conventions This section is used to supplement the explanation of the specific calculation process of the distortion perception sub-indices D1, D2, and D3 in the "Adaptive Adjustment Method and System for Hearing Aid Enhancement Intensity Based on Speech Distortion Perception Index", as well as the normalization, dimension unification, weight acquisition, and time smoothing methods of each sub-indice, so as to enhance the sufficiency of disclosure in the specification.
[0010] The following symbols have the following meanings: : Input time-domain speech signal acquired by the microphone; : No. Frame number Input frame spectrum at each frequency point; : No. Frame number Enhanced spectrum at each frequency point; : Enhanced time-domain output signal; The enhanced spectral amplitude is defined as follows: ; : Input smoothed reference amplitude spectrum obtained by time smoothing or frequency band smoothing of the input frame spectrum; Total number of frequency points per frame; Total number of sub-bands; : No. The frequency value corresponding to each frequency point; Modulation frequency; : No. The temporal sample set corresponding to the frame; : To prevent extremely small positive numbers with a denominator of zero; , , Three original distortion perception sub-indices; , , Three distortion perception sub-indices after normalization and dimensional unification; Comprehensive distortion risk score; : Overall distortion risk score after time smoothing; , , : Weighting coefficients for the overall distortion risk score; , , , , , , , , : Local weighting coefficients when constructing each sub-index.
[0011] D1: Spectral peak / spectral flux index: used to reflect musical noise, metallic sounds and sudden spectral roughness anomalies, constructed by the ratio of the maximum value to the average value of the enhanced spectrum, the degree of abrupt changes between adjacent frequency points, or the spectral flux of adjacent frames; D2: Envelope Change Rate / Modulation Anomaly Index: Used to reflect the unnatural oscillations of the speech energy envelope, such as spitting, wheezing, or enhanced speech energy. It is constructed by short-time energy change rate, subband envelope fluctuation rate, or the proportion of abnormal modulation energy in the 2Hz to 20Hz modulation range. D3: Spectral flatness / harmonic retention index: used to reflect the naturalness of speech, timbre fidelity and harmonic structure damage, constructed by enhancing the spectral flatness deviation, bandwidth drift or harmonic peak-valley structure difference between the output and the input smoothing reference.
[0012] 3.D1: Calculation process of spectral peak / spectral flux quantum index It is used to characterize abnormal peaks, spectral roughness, and abrupt changes between adjacent frames in the enhanced spectrum, in order to reflect the risk of subjective auditory artifacts such as music noise and metallic sounds.
[0013] 3.1 Calculation of Intermediate Quantities First, calculate the enhanced spectral amplitude: Based on the amplitude spectrum, construct one or more of the following intermediate quantities: (1) Peak-to-average ratio in, Used to characterize the degree of abnormal spikes within a single frame.
[0014] (2) Frequency mutation term in, Used to characterize unnatural and drastic changes between adjacent frequency points.
[0015] (3) Spectral flux term in, Used to characterize the rate of spectral change between adjacent frames.
[0016] 3.2 Normalization Method because , , The units and ranges of values are different, in the construction Previously, it was preferred to perform normalization separately. In the linear normalization implementation, we have: in, , , These are the lower bound reference values for the corresponding intermediate quantities. , , These are the upper bound reference values for the corresponding intermediate quantities. This represents a truncation function used to restrict the result to a specific range. Interval.
[0017] In another implementation, mean-variance standardization can be used, followed by mapping to a sigmoid compression function. An interval, for example: and The same method can also be used, in which, , To calculate the mean and standard deviation, This is the compression slope coefficient.
[0018] 3.3 Sub-index Construction After normalization, construct the original values of the spectral peak / spectral flux quantum index: in, , , The coefficients are non-negative and preferably satisfy the following conditions: If hardware resources are limited, you can also select only... , , One or two of the structures .
[0019] 4. D2: Calculation process of envelope change rate / modulation anomaly sub-index It is used to characterize the unnatural fluctuations in the short-term envelope of enhanced speech, reflecting the sensation of spitting, wheezing, and envelope modulation abnormalities caused by excessive enhancement.
[0020] 4.1 Calculation of Intermediate Quantities (1) Short-time energy change rate term First, calculate the short-time energy of the enhanced temporal output frame by frame: Redefining the short-time rate of energy change: (2) Sub-band envelope volatility term Let the first The enhanced envelope of the individual bands is Then we have: (3) Modulation anomalies Let the modulation spectrum of the enhanced output envelope be... The modulation frequency weighting function is Then we can define: in, Used for characterization to The percentage of abnormal modulation energy within the modulation range.
[0021] 4.2 Normalization Method right , and When performing linear normalization, we have: in, , , This serves as the lower bound reference value for the corresponding intermediate quantity. , , This is the upper bound reference value for the corresponding intermediate quantity.
[0022] In another implementation, normalization can also be performed by combining mean-variance standardization with sigmoid compression, which is the same as in Section 2.2 and will not be repeated here.
[0023] 4.3 Sub-index Construction After normalization, construct the original values of the envelope change rate / modulation anomaly sub-index: in, , , The coefficients are non-negative and preferably satisfy the following conditions: If hardware resources are limited, only one method can be used. or As a construction The basic quantity.
[0024] 5. D3: Calculation process of spectral flatness / harmonic retention sub-index It is used to characterize the naturalness of the enhanced speech, the fidelity of the timbre, and the preservation of the harmonic structure, so as to reflect the risk of distortion such as thinning of the timbre and destruction of harmonics.
[0025] 5.1 Calculation of Intermediate Quantities (1) Spectral flatness deviation term The spectral flatness of the enhanced output is defined as: The spectral flatness of the input smooth reference spectrum is defined as: Based on this, the spectral flatness deviation term is constructed: (2) Bandwidth drift term The enhanced spectral centroid is defined as: The centroid of the input smooth reference spectrum is defined as: Therefore, the bandwidth drift term is constructed: (3) Harmonic peak-valley structure difference item In a preferred embodiment, local peaks and valleys of the enhanced spectrum and the input smoothed reference spectrum can be extracted within several dominant frequency bands, and a harmonic peak-valley structure difference term can be constructed based on peak position offset, peak-valley ratio deviation, or peak-valley spacing difference, denoted as... .for The specific implementation can be achieved by using the local peak matching method, the in-band peak-valley comparison method, or the harmonic template difference method, depending on the hardware resources and product positioning.
[0026] 5.2 Normalization Method right , and When performing linear normalization, we have: in, , , This serves as the lower bound reference value for the corresponding intermediate quantity. , , This serves as the upper bound reference value for the corresponding intermediate quantity. If linear normalization is not used, a combination of mean-variance standardization and sigmoid compression can also be employed.
[0027] 5.3 Sub-index Construction After normalization, the original values of the spectral flatness / harmonic preservation sub-indices are constructed as follows: in, , , The coefficients are non-negative and preferably satisfy the following conditions: If hardware resources are limited, only one method can be used. or and Combination construction .
[0028] 6. Unification and normalization of the three sub-indicators and unification of dimensions In a preferred embodiment, to make , , Since they are comparable to each other and facilitate subsequent weighted fusion and closed-loop control, the three original sub-indices are further unified and normalized to obtain: in, For the first The lower bound reference value of each original sub-indicator. For the first The upper bound reference value of each original sub-indicator.
[0029] In another implementation, mean-variance standardization can also be used: And mapped to via the S-type compression function Interval: in, and The first The statistical mean and standard deviation of each original sub-indicator. This is the compression slope coefficient.
[0030] The term "dimensional unification" refers to the process of normalizing or standardizing the original distorted representations with different physical meanings, dimensions, and numerical ranges, and then unifying them into a single entity. The dimensionless risk quantity within the interval is used for subsequent fusion and control.
[0031] 7. Weighting method and overall distortion risk score Sub-indices after normalization and dimensional unification , , Then, a comprehensive distortion risk score is constructed using a weighted method: in, , , The weighting coefficients are non-negative and preferably satisfy the following conditions: The weighting coefficients can be obtained in any of the following ways: Experience-based setting method: Pre-set according to the degree of impact of different distortion types on subjective listening experience; Offline calibration method: determined based on subjective evaluation data through least squares fitting, correlation analysis, or regression analysis; Data-driven approach: trained on a dataset with subjective labels; Scene adaptive mode: dynamically adjusts based on environment category, noise type, or user configuration.
[0032] To avoid the overall distortion risk score exceeding the controller's processing range, additional measures can be taken. Amplitude limiting is applied: Unless otherwise specified, the following text refers to... It can also represent the overall distortion risk score after amplitude limiting.
[0033] 8. Time smoothing processing Because speech signals and enhanced outputs exhibit significant short-term fluctuations, directly applying the instantaneous composite distortion risk score to enhancement intensity control can easily lead to control jitter. Therefore, a first-order recursive smoothing method is preferred. in, This is the time smoothing coefficient, and its preferred value range is [value range missing]. to ; The smoothing overall distortion risk score for the previous frame; This is the overall distortion risk score for the current frame after amplitude limiting.
[0034] In a preferred embodiment .when When the value is large, the control process is more stable; when When the size is smaller, the control response is faster.
[0035] 9. Explanation In this invention, the specific calculation forms of D1, D2, and D3 are not limited to the above embodiments. Any technical solution that is based on the enhanced output or the enhanced output and input smoothing reference construction, capable of representing spectral roughness anomalies, envelope anomalies, and harmonic / timbre anomalies respectively, and used to construct a comprehensive distortion risk score after normalization, dimension unification, weighted fusion, and time smoothing, falls within the protection scope of this invention.
[0036] The three types of distortion sub-indicators mentioned above are combined in a weighted manner to obtain a comprehensive distortion risk score. The calculation formula is as follows: in, , , This is a weighting coefficient that can be optimized based on the product positioning and clinical adjustment needs of the hearing aid.
[0037] To avoid abrupt changes in the overall distortion risk score due to instantaneous signal fluctuations, D(t) is time-smoothed to obtain the smoothed overall distortion risk score. The calculation formula is: Wherein, β is the time smoothing coefficient, and the preferred value range is 0.70 to 0.95. The larger the β is, the smoother the control process is, and the more effectively the frequent jitter of the enhanced intensity is avoided.
[0038] 10. Definition of reinforcement strength parameters In this invention, "enhancement strength" is not a single parameter, but refers to one or more control parameters that have a substantial impact on the strength of speech enhancement processing, covering the core adjustment dimensions of various architectures such as traditional noise reduction and neural network enhancement, specifically including: Neural Network Masking Index For example, the enhancement gain is calculated as G(k,t)=M(k,t)^α(t), and the larger α is, the stronger the noise suppression. Minimum spectral gain Gmin(t): The smaller the minimum gain, the greater the allowable noise reduction depth; Spectral smoothing or time-domain smoothing coefficient γ(t): The larger the coefficient, the more beneficial it is to suppress musical noise, but it may reduce transient fidelity; Post-filtering gain, noise estimation update rate, sub-band gain upper limit or compression mapping intensity.
[0039] To facilitate the implementation of closed-loop control and the generalization of the patent protection scope, the controller outputs a continuous control quantity u(t) and converts it into specific algorithm parameters through a linear mapping relationship. A typical mapping method is: Neural network masking index: , / The preferred value range is 0.6 to 2.0; Spectral smoothing coefficient: , / The preferred value range is 0.2 to 0.8.
[0040] 11. Closed-loop control strategy One of the core innovations of this invention is to construct an enhanced intensity adaptive negative feedback closed-loop, with the smoothed comprehensive distortion risk score as the core control basis, combined with the environmental noise intensity N(t), and adopt a piecewise control law with hysteresis to achieve dynamic adjustment of the enhanced intensity. The core control logic is: when the distortion risk increases, the controller automatically reduces the enhanced intensity or increases the smoothness; when the distortion risk is low and the noise is strong, moderately increase the enhanced intensity; when the distortion risk is in the middle range, maintain the current enhanced intensity.
[0041] The specific piecewise control rules are: When > Dhi (preset upper threshold): It means that the enhancement process has generated an obvious subjective distortion risk. Execute a rapid fallback of the enhanced intensity, quickly reduce the noise suppression strength, and suppress auditory artifacts; When < Dlo (preset lower threshold) and N(t) > Nth (noise threshold): It means that the current distortion risk is low and the environmental noise is strong. Execute a slow enhancement of the enhanced intensity, gradually increase the noise suppression strength, and improve speech clarity; In other states: Keep the current enhanced intensity unchanged, or only make minor adjustments to ensure the stability of the auditory sense.
[0042] To further improve the robustness of the control, the preset control rules also include the following constraint mechanisms: Rate of change limit: The maximum rate of change per frame , The preferred value range is 0.05 to 0.10 to prevent jumps in the enhanced intensity parameters; Peak Protection Mode: Special processing is performed on transient strong distortion signals to avoid over-adjustment caused by a single peak; Scene switching freeze strategy: When a rapid change in acoustic scene is detected, the enhancement intensity parameter update is frozen to avoid erroneous adjustment caused by scene switching.
[0043] The preferred control parameter range of this invention is: Dlo=0.45, Dhi=0.65, enhancement step size Δup=0.02, and backoff step size Δdown=0.08. The backoff step size is greater than the enhancement step size to ensure a rapid response when the distortion risk increases and a slow increase when the distortion risk decreases, thus avoiding overshoot.
[0044] 12. Typical Implementation Process The adaptive adjustment method of this invention can operate at the frame level, subframe level, or scene level, with frame-level implementation being the preferred method, as it can balance real-time performance and control accuracy. A typical frame-level implementation process is as follows: The microphone collects the time-domain speech signal x[n], and the preprocessing module performs pre-emphasis, framing, windowing and STFT processing on it to obtain the input frame spectrum X(k,t). The frame length is preferably 2ms, 5ms or 10ms. The speech enhancement module uses a preset speech enhancement algorithm to process X(k,t) to obtain the enhanced spectrum Ŝ(k,t) or the enhanced time-domain output signal ŷ[n]. The distortion perception index calculation module calculates three types of distortion perception sub-indices, D1, D2, and D3, based on Ŝ(k,t) or ŷ[n], and obtains the comprehensive distortion risk score D(t) according to the weighted fusion formula. Time smoothing is applied to D(t) to obtain the smoothed comprehensive distortion risk score. ; The enhanced intensity control module collects the current ambient noise intensity N(t) and, according to... Given N(t) and the piecewise control law with hysteresis, calculate the enhancement intensity control quantity u(t+1) for the next frame; The parameter update module converts u(t+1) into specific enhancement intensity parameters (such as α(t), γ(t), Gmin(t), etc.) through mapping relationships and updates them to the speech enhancement module; Repeat steps 1)-6) above for subsequent voice frames to form a continuous adaptive negative feedback closed-loop control.
[0045] 13. Preferred Embodiment To further illustrate the technical solution of the present invention, the following three preferred embodiments are provided in combination with different hearing aid speech enhancement architectures and system forms. Each embodiment is based on the above core technical solution, and the parameters and modules are adapted only according to the specific application scenario.
[0046] 13.1 Frequency Domain Neural Network Masking Enhancement Example This embodiment is applicable to low-power hearing aids employing the WOLA / FFT frequency domain processing architecture. The speech enhancement module uses a lightweight neural network masking enhancement algorithm, specifically implemented as follows: 1. The input time-domain speech signal is divided into 5ms frames, processed by WOLA to obtain frequency domain features, input into a lightweight neural network, and outputs a masking function M(k,t); 2. Enhance gain by pressing The calculation is performed, where α(t) is the neural network masking index, which serves as the core enhancement strength parameter and is adaptively updated by the controller of this invention. 3. When the distortion perception module detects an increase in the overall distortion risk score, the controller decreases α(t) and increases the spectral smoothing coefficient γ(t) to suppress music noise and metallic sounds; 4. When the overall distortion risk score is low and the environmental noise is strong, the controller slowly increases α(t) to enhance noise suppression and improve speech clarity.
[0047] This embodiment requires minimal modification to the existing neural network enhancement main link, achieving adaptive adjustment only by updating the masking index and spectral smoothing coefficient, making it easy to upgrade and implement on existing hearing aid platforms.
[0048] 13.2 Traditional Frequency Domain Noise Reduction Examples This embodiment is applicable to hearing aids that employ traditional frequency domain noise reduction algorithms. The speech enhancement module uses Wiener filtering or spectral subtraction, and the specific implementation method is as follows: The preprocessing module obtains the input frame spectrum through FFT, and the speech enhancement module uses the Wiener filtering algorithm to calculate the noise reduction gain and performs noise reduction processing on the input spectrum to obtain the enhanced spectrum; The controller uses the minimum spectral gain Gmin(t) and the post-filter strength as the core enhancement strength parameters and adjusts them directly. When the distortion sensing module detects a spectral peak (D1 increases) or an envelope anomaly (D2 increases), the controller appropriately increases Gmin(t) and enhances gain smoothing to reduce sucking sensation and music noise. When the risk of distortion is low and the noise is strong, the controller appropriately reduces Gmin(t) to increase the noise reduction depth and improve the clarity.
[0049] This embodiment is adapted to the algorithm architecture of traditional hearing aids, with simple control logic, low computational load, and meets the requirements for low-power real-time implementation.
[0050] 13.3 Binaural Synergy Example This embodiment applies to binaural hearing aid systems, and its core solution addresses the inconsistency in hearing perception caused by independent binaural adjustment. The specific implementation method is as follows: 1. Both the left and right ear hearing aids are equipped with independent preprocessing modules, speech enhancement modules, and distortion perception index calculation modules, which respectively collect local speech signals and calculate a comprehensive distortion risk score; 2. A master-slave collaborative approach is adopted, with one side designated as the master and the other as the slave. The master's enhancement intensity control module integrates the distortion risk scores of both ears and the ambient noise intensity to calculate a unified enhancement intensity control value. 3. The host broadcasts the enhancement intensity control value to the slave, and the parameter update modules of both ears synchronously map the control value to the same or adapted enhancement intensity parameters and update their respective voice enhancement modules; 4. Supports binaural synchronous control or asymmetric collaborative control. Asymmetric collaborative control can make minor adjustments to the enhancement intensity parameters according to the sound field differences between the left and right ears, taking into account both auditory consistency and sound field adaptability.
[0051] This embodiment can be considered as a dependent claim, extending the scope of protection of the present invention in binaural hearing aid systems and improving the comfort and consistency of binaural listening.
[0052] 14. Recommended parameter range The core control parameters of this invention are all set with reasonable value ranges, which not only ensures the stability and robustness of the algorithm, but also reserves space for product optimization of hearing aids. The example ranges, functional descriptions and remarks of each core parameter are shown in Table 2 below: Table 2. Range of core control parameters The hearing aid enhancement intensity adaptive adjustment method and system based on speech distortion perception index of the present invention has the following significant technical effects compared with the prior art: Reference-free distortion assessment, adaptable to real-time implementation on the device: It can realize the distortion risk assessment of enhanced speech without the need for clean reference speech, avoiding the problem that clean reference speech cannot be obtained in the actual environment. It has low computational requirements and is adapted to the real-time implementation requirements of low-power DSP, MCU or hearing aid-specific SoC. Negative feedback closed-loop control precisely suppresses auditory artifacts: By constructing a negative feedback closed loop of "enhanced output - distortion assessment - parameter update - subsequent enhancement", subjective auditory distortion is used as the core constraint. When a high risk of distortion is detected, the enhancement intensity is automatically reduced, effectively suppressing enhancement artifacts such as metallic sounds, musical noise, sucking sensation and envelope abnormalities. Balancing clarity and naturalness to meet subjective listening needs: Abandoning the traditional approach of simply pursuing maximum signal-to-noise ratio, it dynamically adjusts the enhancement intensity based on the intensity of ambient noise. When the risk of distortion is low, the enhancement intensity is increased to ensure clarity, and when the risk of distortion is high, the enhancement intensity is reduced to ensure naturalness. A controllable balance is established between the two, which is closer to the user's subjective listening needs. The algorithm is highly interpretable and easy to debug in products: it establishes an explicit quantitative mapping relationship between enhancement intensity, smoothness and distortion risk, and the physical meaning of each control parameter is clear, which facilitates clinical debugging, product optimization and subsequent writing of medical device registration documents. Highly compatible and widely applicable: It is suitable for various hearing aid systems such as single-microphone, dual-microphone, and binaural collaboration. It is compatible with various speech enhancement architectures such as traditional frequency domain noise reduction, neural network masking enhancement, and time domain enhancement. It also supports multi-algorithm hybrid architecture, and has strong engineering compatibility and practicality. High robustness and adaptability to complex acoustic scenarios: By introducing multiple constraint mechanisms such as time smoothing, rate of change limit, spike protection, and scene switching freeze, the algorithm effectively avoids frequent parameter jitter and misadjustment, and improves the stability and robustness of the algorithm in complex acoustic scenarios such as restaurants, shopping malls, and streets.
[0053] 15. Alternative implementation methods and expansion points To broaden the scope of this invention and adapt to the development and productization needs of hearing aid technology, the following alternative embodiments and extensions may be retained in the formal invention patent application documents, all of which do not depart from the core technical concept of this invention: Alternative forms of distortion perception indicators: Any one or more of D1, D2, and D3 can be used as distortion perception indicators, or a comprehensive distortion risk score can be directly predicted through a lightweight neural network to replace the calculation method of rule-based algorithms. Alternative options for enhancement intensity parameters: Depending on the hearing aid's algorithm architecture, one or more of the following can be selected as the core adjustment parameters: masking index, minimum gain, post-filter intensity, compression mapping parameter, or subband gain upper limit; Alternatives to the controller strategy: The controller can adopt a segmented threshold strategy, a proportional control strategy, a lookup table strategy, or a reinforcement learning strategy to replace the segmented control law with hysteresis preferred in this invention. Multi-dimensional information overlay and fusion: User subjective feedback, environmental recognition credibility, binaural synchronization strategy or long-term usage behavior trends can be overlaid in the control logic to achieve more personalized and intelligent enhancement intensity adjustment; Alternative options for processing granularity: The method of the present invention can operate at the frame level, subframe level, or scene level, and different processing granularities can be selected according to the computing power and real-time requirements of the hearing aid. Attached Figure Description
[0054] Figure 1 The overall flowchart of the adaptive adjustment of voice enhancement shows the complete process of the method of the present invention from microphone acquisition to parameter update; Figure 2A schematic diagram of the speech distortion perception index shows the functions of the three sub-indicators D1, D2, and D3, as well as the fusion method of the comprehensive distortion risk score. Figure 3 The diagram illustrates the enhanced strength adaptive negative feedback closed loop, showing the signal flow and negative feedback control relationship between the modules. Figure 4 The diagram illustrates the relationship between enhanced intensity control and the overall distortion risk score, showing the segmented control relationship between enhanced intensity control and the overall distortion risk score, including different control strategies in the low distortion zone, the maintain / gradual adjustment zone, and the high distortion zone.
[0055] Explanation of the attached diagram: x[n] is the microphone time-domain signal, X(k,t) is the frame spectrum, Ŝ(k,t) is the enhanced spectrum, ŷ[n] is the enhanced time-domain output signal, D1 / D2 / D3 are the three types of distortion perception sub-indicators, and D(t) is the comprehensive distortion risk score. The smoothed comprehensive distortion risk score is given by u(t), where u(t) is the enhancement intensity control value, α is the neural network masking index, γ is the spectral smoothing coefficient, and N(t) is the environmental noise intensity. Detailed Implementation
[0056] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments, but the scope of protection of the present invention is not limited to the following embodiments.
[0057] The core of the hearing aid enhancement intensity adaptive adjustment method and system based on speech distortion perception index of the present invention is to construct a negative feedback closed-loop control system with no-reference distortion perception index as the core. All embodiments are based on this core concept, and the parameters and modules are adapted only according to the hardware platform, algorithm architecture and system form of the hearing aid.
[0058] In actual product implementation, the algorithm for calculating the distortion perception index can be lightweighted and optimized to reduce computational load, based on the low power consumption requirements of hearing aids; the weighting coefficients can be adjusted according to clinical test results. , , The algorithm is optimized by adjusting the control thresholds Dlo and Dhi to better suit the hearing needs of different users. For binaural hearing aid systems, the master-slave collaborative communication protocol can be optimized to ensure the synchronization and real-time performance of the enhancement intensity control information.
[0059] Those skilled in the art can make partial modifications or equivalent substitutions to the methods and systems of this invention based on the core technical concept of this invention. As long as these modifications and substitutions do not depart from the spirit and scope of the technical solution of this invention, they shall all fall within the protection scope of this invention.
[0060] Industrial applicability The adaptive adjustment method and system for hearing aid enhancement intensity based on speech distortion perception index of the present invention can be directly deployed on existing digital hearing aid platforms with minimal changes to the algorithm architecture of existing hearing aids, making it easy to upgrade and implement; the computational load of each module is low, adapting to the real-time implementation requirements of low-power DSP, MCU or hearing aid-specific SoC, and meeting the industrial productization needs of hearing aids.
[0061] This invention effectively solves the technical problem of difficulty in balancing clarity and naturalness in the enhancement intensity adjustment of existing hearing aids, significantly improves the listening effect and wearing comfort of hearing aids in complex acoustic environments, has high clinical application value and market promotion value, can be widely used in various digital hearing aid products such as monoauricular and binauricular collaborative hearing aids, and has good industrial applicability.
Claims
1. A method of adaptive adjustment of the intensity of hearing aid amplification based on a speech distortion perceptual index, characterized in that, Includes the following steps: 1) The time-domain speech signal acquired by the microphone is preprocessed by pre-emphasis, framing, windowing and short-time Fourier transform to obtain the frame spectrum. The preprocessing is implemented by WOLA, FFT or filter bank. 2) The frame spectrum is processed using a speech enhancement algorithm to obtain an enhanced spectrum or an enhanced time-domain output signal. The speech enhancement algorithm is any one or a combination of traditional frequency domain noise reduction algorithm, neural network masking enhancement algorithm, and time-domain enhancement algorithm. 3) Based on the enhanced spectrum or enhanced time-domain output signal, calculate a non-reference distortion perception sub-indicator. The distortion perception sub-indicator includes at least one or more of the following: spectral peak / spectral flux index, envelope change rate / modulation anomaly index, and spectral flatness / harmonic preservation index. The distortion perception sub-indicators are weighted and fused to obtain a comprehensive distortion risk score, and the comprehensive distortion risk score is subjected to time smoothing. 4) Collect the ambient noise intensity, and calculate the enhancement intensity control amount based on the smoothed comprehensive distortion risk score, the ambient noise intensity, and the pre-set segmented control rules with hysteresis; 5) Map the enhancement intensity control quantity to specific enhancement intensity parameters and update them to the speech enhancement algorithm. Repeat steps 1)-5) for the subsequently acquired speech signals to form a negative feedback closed-loop control for speech enhancement.
2. The hearing aid gain enhancement adaptive adjustment method based on voice distortion perception index according to claim 1, characterized in that, The frame length of the frame segmentation in step 1) is 2ms, 5ms or 10ms; the traditional frequency domain noise reduction algorithm includes Wiener filtering algorithm and spectral subtraction, and the neural network masking enhancement algorithm is a lightweight neural network masking enhancement algorithm.
3. The hearing aid gain enhancement adaptive adjustment method based on voice distortion perception index according to claim 1, characterized in that, The spectral peak / spectral flux index is constructed by the ratio of the maximum to the average value of the enhanced spectrum, the degree of abrupt changes in adjacent frequency points, or the spectral flux of adjacent frames; the envelope change rate / modulation anomaly index is constructed by the short-time energy change rate, the sub-band envelope fluctuation rate, or the proportion of abnormal modulation energy in the modulation range of 2Hz to 20Hz; the spectral flatness / harmonic retention index is constructed by the spectral flatness deviation between the enhanced output and the input smoothing reference, the bandwidth drift, or the difference in harmonic peak-valley structure.
4. The hearing aid gain enhancement adaptive adjustment method based on voice distortion perception index according to claim 1, characterized in that, The comprehensive distortion risk score is calculated in the following manner: wherein is the comprehensive distortion risk score, , , is a weight coefficient, and D1, D2, and D3 are, in sequence, a spectral peak / spectral flux, an envelope change rate / modulation abnormality, and a spectral flatness / harmonic preservation index. The time smoothing manner is as follows: wherein is the smoothed comprehensive distortion risk score, and β is a time smoothing coefficient, with a value range of 0.70-0.
95.
5. The hearing aid gain enhancement adaptive adjustment method based on voice distortion perception index according to claim 1, characterized in that, The enhancement strength parameter includes any one or more of the following: neural network masking exponent, minimum spectral gain, spectral smoothing coefficient, temporal smoothing coefficient, post-filter gain, noise estimation update rate, subband gain upper limit, and compression mapping strength; when the enhancement strength parameter is the neural network masking exponent... When, the mapping relationship is: , , The minimum and maximum values of the masking index are taken from 0.6 to 2.0; when the enhancement intensity parameter is the spectral smoothing coefficient... When, the mapping relationship is: , , These are the minimum and maximum values of the spectral smoothing coefficient, ranging from 0.2 to 0.
8. To enhance the control of intensity.
6. The hearing aid enhancement intensity adaptive adjustment method based on speech distortion perception index according to claim 1, characterized in that, The hysteresis-based segmented preset control rule is specifically as follows: when the smoothed comprehensive distortion risk score... When the value exceeds the preset upper threshold Dhi, a rapid rollback of the enhancement intensity is performed; when the smoothed overall distortion risk score... When the ambient noise intensity N(t) is less than the preset lower threshold Dlo and the ambient noise intensity N(t) is greater than the noise threshold Nth, the enhancement intensity is slowly increased; in other states, the current enhancement intensity remains unchanged. The preferred values for Dlo and Dhi are 0.45 and 0.65, respectively, and the preferred values for the enhancement step size Δup and the backoff step size Δdown are 0.02 and 0.08, respectively.
7. The hearing aid enhancement intensity adaptive adjustment method based on speech distortion perception index according to claim 6, characterized in that, The preset control rule further comprises a change rate limit of the enhanced intensity control quantity, and a maximum change rate of a single frame , The value range of the value of the enhanced intensity control quantity is 0.05-0.10; a spike protection mode and a scene switching freezing strategy are configured simultaneously to avoid frequent shaking of the enhanced intensity parameter.
8. The hearing aid gain boost adaptive adjustment method based on a voice distortion perceptual index according to any one of claims 1-7, characterized in that, This method is applicable to single-microphone, dual-microphone, or binaural collaborative hearing aid systems. When applied to binaural collaborative hearing aid systems, distortion risk scores are calculated separately for each ear, or enhancement intensity control information is shared through master-slave collaboration to achieve synchronous binaural control or asymmetric collaborative control.
9. A hearing aid gain enhancement adaptive adjustment system based on a speech distortion perceptual index implementing the method according to any one of claims 1-8, characterized in that, The system includes at least: a preprocessing module for preprocessing the time-domain speech signal acquired by the microphone and outputting a frame-by-frame spectrum; a speech enhancement module connected to the preprocessing module for processing the frame-by-frame spectrum using a speech enhancement algorithm and outputting an enhanced spectrum or an enhanced time-domain output signal; a distortion perception index calculation module connected to the speech enhancement module for calculating a referenceless distortion perception sub-index based on the enhanced spectrum or enhanced time-domain output signal, weighted fusion to obtain a comprehensive distortion risk score, and time smoothing processing; an enhancement intensity control module connected to the distortion perception index calculation module and an external noise detection unit, for calculating an enhancement intensity control quantity based on the smoothed comprehensive distortion risk score, environmental noise intensity, and preset control rules, wherein the enhancement intensity control module supports control strategies such as dual threshold, hysteresis, amplitude limiting, and spike protection; and a parameter update module connected to the enhancement intensity control module and the speech enhancement module, for mapping the enhancement intensity control quantity to specific enhancement intensity parameters and updating them to the speech enhancement module; the system is deployed on a low-power DSP, MCU, or hearing aid-specific SoC platform.
10. A digital hearing aid, characterized by The device includes a microphone, a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the hearing aid enhancement intensity adaptive adjustment method based on speech distortion perception index as described in any one of claims 1-8. The digital hearing aid is a monoaural hearing aid or a binaural hearing aid, wherein the left and right ears of the binaural hearing aid share enhancement intensity control information through a master-slave cooperative mode.