A smart hearing aid mode switching system and its switching method
By combining panoramic sound field perception and proprioceptive motion intent capture in a joint decision, and integrating the sound field uncertainty index and gaze shift intensity, a parameter freeze signal and beam lead compensation angle are generated. This solves the problems of mode switching error and lag in hearing aids in high dynamic sound fields, achieves synchronization of hearing and vision, and improves the user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-22
- Publication Date
- 2026-04-03
AI Technical Summary
In highly dynamic and complex sound fields, hearing aids may experience problems such as incorrect mode switching and auditory beam tracking lag due to user head rotation, which can affect audiovisual synchronization and user experience.
The system employs a combined decision-making process involving a panoramic sound field perception unit and a body motion intent capture unit. By judging the sound field uncertainty index and the intensity of gaze shift, it generates a parameter freezing signal to lock the scene classification results. Furthermore, it utilizes the beam advance compensation angle to perform beam deflection, constructing an auditory perception feedback optimization closed loop to achieve synchronization between auditory focus and visual focus.
It effectively suppresses erroneous mode switching, ensures the continuity and steady state of auditory perception, reduces the system's sensitivity to invalid motion data, achieves spatial synchronization between hearing and vision, and enhances the robustness and adaptability of system decision-making.
Smart Images

Figure CN121357481B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent hearing enhancement and digital signal processing technology, specifically to an intelligent hearing aid mode switching system and its switching method. Background Technology
[0002] With the increasing diversification of hearing aid applications, the demand for speech enhancement in highly dynamic and complex sound field environments has increased significantly. This complexity stems not only from the disordered fluctuations of the acoustic environment itself, but also from the user's own body movement behavior.
[0003] Currently, scene classification and mode switching in hearing aids typically rely on a single analysis of acoustic signal characteristics. However, in non-steady-state noise or highly uncertain sound fields, the reliability of environmental assessment based solely on acoustic characteristics is low. In particular, when users turn their heads to search for sound sources, signal fluctuations caused by airflow noise or relative displacement of the sound source can easily lead to misjudgments by the hearing aid, resulting in frequent mode switching errors or gain parameter jumps. Furthermore, the inherent system lag between electronic signal processing and mechanical motion causes the auditory beam to fail to follow the user's visual gaze in real time, resulting in spatial asynchrony of audiovisual perception. Therefore, how to effectively suppress mode switching errors caused by body motion coupling and eliminate the lag effect of beam tracking to achieve audiovisual synchronization in highly dynamic and complex sound fields has become an urgent problem to be solved in this field. Summary of the Invention
[0004] The purpose of this invention is to provide an intelligent hearing aid mode switching system and method, which can effectively suppress erroneous mode switching caused by the user's movement in highly dynamic and complex sound fields, and can eliminate system lag to achieve spatial synchronization between auditory focus and visual focus. Specifically, the technical solution of this invention is as follows:
[0005] A smart hearing aid mode switching system, comprising:
[0006] A panoramic sound field sensing unit is used to acquire microphone array signals in real time and extract acoustic features; the panoramic sound field sensing unit is used to calculate the sound field uncertainty index based on the variance of short-time energy and spectral entropy value;
[0007] The proprioception intent capture unit is used to collect head motion data through an inertial measurement unit; the proprioception intent capture unit is used to calculate the gaze shift intensity and real-time motion vector based on the head motion data;
[0008] A heterogeneous feature decoupling and fusion unit is connected to both the panoramic sound field perception unit and the body motion intent capture unit. The heterogeneous feature decoupling and fusion unit synchronously receives the sound field uncertainty index, the gaze point shift intensity, and the real-time motion vector. When the sound field uncertainty index is greater than a preset sound field uncertainty threshold and the gaze point shift intensity is greater than a preset gaze point shift intensity threshold, a parameter freezing signal is generated. This parameter freezing signal is used to lock the scene classification result and gain parameters. The heterogeneous feature decoupling and fusion unit calculates the beam advance compensation angle based on the real-time motion vector. The heterogeneous feature decoupling and fusion unit superimposes the beam advance compensation angle into the beamforming parameters to drive the beam pointing to perform a unidirectional advance deflection.
[0009] An auditory perception feedback optimization unit is used to monitor the clarity index of the target speech within a preset time window after beam deflection; the auditory perception feedback optimization unit is used to generate a correction feedback signal when the clarity index decreases;
[0010] The proprioception intent capture unit is used to increase the trigger threshold of the gaze shift intensity in response to the correction feedback signal.
[0011] Preferably, the panoramic sound field perception unit is specifically used for:
[0012] Continuously monitor the microphone array signal;
[0013] Calculate the short-time energy variance and spectral entropy of the microphone array signal;
[0014] The sound field uncertainty index is obtained by weighting the short-time energy variance and the spectral entropy value.
[0015] The sound field uncertainty index is transmitted as a weighted reference signal to the heterogeneous feature decoupling and fusion unit;
[0016] The sound field uncertainty index is used to characterize the complexity of the current acoustic environment.
[0017] Preferably, the body motion intent capture unit is specifically used for:
[0018] Collect raw data from the gyroscope and accelerometer;
[0019] The original data is then denoised.
[0020] Determine whether the angular velocity of head rotation exceeds a preset physiological threshold;
[0021] Determine whether the direction of head rotation remains singular and continuous;
[0022] When the angular velocity exceeds the preset physiological threshold and the direction remains single and continuous, the output value of the fixation point shift intensity is increased;
[0023] The gaze shift intensity is used to characterize whether the user has a clear auditory search intent.
[0024] Preferably, the parameter freezing signal generated by the heterogeneous feature decoupling and fusion unit is specifically used for:
[0025] The acoustic classifier is prohibited from updating the scene mode based on fluctuations in the microphone array signal;
[0026] The current scene classification result and gain parameters remain unchanged until the sound field uncertainty index or the gaze shift intensity is lower than the corresponding threshold.
[0027] Preferably, the heterogeneous feature decoupling and fusion unit is used to calculate the beam lead compensation angle specifically for:
[0028] Obtain a preset dynamic compensation coefficient, which is positively correlated with the system processing delay time;
[0029] The product of the preset dynamic compensation coefficient and the angular velocity in the real-time motion vector is calculated to obtain the beam lead compensation angle;
[0030] The beam advance compensation angle is used to offset the lag time between electronic signal processing and mechanical rotation.
[0031] Preferably, the auditory perception feedback optimization unit is specifically used for:
[0032] The envelope completeness is obtained as the sharpness indicator;
[0033] When a decrease in envelope integrity is detected, it is determined that a beam deflection has been misjudged;
[0034] The correction feedback signal is generated and sent to the body motion intent capture unit to reduce the system's sensitivity to the head motion data.
[0035] A method for switching modes in a smart hearing aid includes the following steps:
[0036] The ambient sound signal is collected by the panoramic sound field sensing unit, and the variance of short-time energy and spectral entropy value are calculated to generate the sound field uncertainty index.
[0037] The proprioception capture unit collects data from the inertial measurement unit, calculates the angular velocity and direction of head rotation, and outputs the gaze shift intensity and real-time motion vector.
[0038] The sound field uncertainty index and the gaze point shift intensity are received by the heterogeneous feature decoupling and fusion unit;
[0039] When the sound field uncertainty index is greater than a preset sound field uncertainty threshold and the gaze point shift intensity is greater than a preset gaze point shift intensity threshold, a parameter freeze signal is generated to lock the scene parameters.
[0040] The heterogeneous feature decoupling and fusion unit calculates the beam advance compensation angle based on the real-time motion vector, and performs a unidirectional advance deflection of the beam direction.
[0041] The auditory perception feedback optimization unit monitors the speech intelligibility index after beam deflection, and generates a correction feedback signal when the speech intelligibility index decreases.
[0042] The proprioception capture unit adjusts the trigger threshold of the gaze point shift intensity according to the correction feedback signal.
[0043] Compared with the prior art, the present invention has the following beneficial effects:
[0044] 1. This invention uses a combined judgment of panoramic sound field perception and body motion intention to trigger a parameter freezing mechanism when both high uncertainty in the sound field and gaze point shift intensity exceed the limit. This mechanism forcibly maintains the current scene classification and gain parameters, effectively shielding parameter jumps caused by airflow noise or relative displacement of the sound source due to rapid head rotation. It solves the problem of frequent false switching in highly dynamic and complex environments, ensuring the continuity and steady state of the user's auditory perception during the search for sound sources.
[0045] 2. This invention utilizes a heterogeneous feature decoupling and fusion unit to calculate the beam advance compensation angle based on real-time motion vectors and superimposes it into the beamforming parameters; through a feedforward control strategy, the pickup beam is deflected in the same direction relative to the physical array, accurately offsetting the inherent system lag between electronic signal processing and mechanical rotation; this mechanism achieves spatial synchronization between auditory focus and visual focus, ensuring that the beam can reach the target location in real time or slightly ahead when the user turns their head quickly;
[0046] 3. This invention introduces the sound field uncertainty index as the core indicator for environmental assessment, which comprehensively quantifies the complexity of the sound field environment by combining short-time energy variance and spectral entropy. This index does not rely on a single feature, but reflects the degree of disorder in the environment, providing the system with a confidence reference beyond simple acoustic classification. Combined with motion intention recognition, it effectively avoids blindly classifying scenes in environments with non-steady-state noise or drastic fluctuations in signal-to-noise ratio, significantly improving the robustness and accuracy of system decision-making.
[0047] 4. This invention constructs an auditory perception feedback optimization closed loop, which evaluates the execution effect by monitoring the integrity of the speech envelope after beam deflection; when a decrease in clarity is detected, it is determined to be a misjudgment, and a correction signal is generated to adaptively increase the trigger threshold of gaze shift intensity; this dynamic adjustment mechanism reduces the system's sensitivity to invalid motion data, suppresses erroneous responses in high misjudgment environments, and endows the system with adaptive learning capabilities for different usage habits and environments. Attached Figure Description
[0048] The present invention will be further explained below with reference to the accompanying drawings and embodiments:
[0049] Figure 1 This is a structural diagram of the system of the present invention;
[0050] Figure 2 This is a flowchart of the method of the present invention. Detailed Implementation
[0051] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments.
[0052] Example 1:
[0053] Please see Figure 1 A smart hearing aid mode switching system, comprising:
[0054] The panoramic sound field sensing unit is used to acquire microphone array signals in real time and extract acoustic features; the panoramic sound field sensing unit is used to calculate the sound field uncertainty index based on the variance of short-time energy and spectral entropy value.
[0055] The proprioception intent capture unit is used to collect head motion data through an inertial measurement unit; the proprioception intent capture unit is used to calculate the gaze shift intensity and real-time motion vector based on the head motion data;
[0056] The heterogeneous feature decoupling and fusion unit is connected to the panoramic sound field perception unit and the body motion intent capture unit, respectively. The heterogeneous feature decoupling and fusion unit is used to synchronously receive the sound field uncertainty index, gaze point shift intensity, and real-time motion vector. The heterogeneous feature decoupling and fusion unit is used to generate a parameter freezing signal when the sound field uncertainty index is greater than a preset sound field uncertainty threshold and the gaze point shift intensity is greater than a preset gaze point shift intensity threshold. The parameter freezing signal is used to lock the scene classification result and gain parameters. The heterogeneous feature decoupling and fusion unit is used to calculate the beam advance compensation angle based on the real-time motion vector. The heterogeneous feature decoupling and fusion unit is used to superimpose the beam advance compensation angle into the beamforming parameters to drive the beam pointing to perform a unidirectional advance deflection.
[0057] The auditory perception feedback optimization unit is used to monitor the clarity index of the target speech within a preset time window after beam deflection; the auditory perception feedback optimization unit is used to generate a correction feedback signal when the clarity index decreases; the proprioceptive motion intent capture unit is used to increase the trigger threshold of gaze shift intensity in response to the correction feedback signal.
[0058] A hearing aid control logic based on multi-source heterogeneous data fusion is proposed to solve the technical challenge of mode switching errors caused by the coupling of non-steady-state noise and body motion in highly dynamic and complex sound fields. The core of the system operation lies in the parallel processing mechanism of the panoramic sound field perception unit and the body motion intention capture unit. The panoramic sound field perception unit, as an environmental assessment module, performs quantitative analysis of the complexity of the acoustic environment. It not only extracts conventional acoustic features, but more importantly, it calculates and outputs the sound field uncertainty index. This index is a comprehensive quantitative indicator reflecting the disorder and energy fluctuation of the current acoustic environment. Its physical meaning is to characterize the confidence risk of scene classification based solely on acoustic signals. The higher the value, the more chaotic the environment and the lower the reliability of acoustic classification. At the same time, the body motion intention capture unit uses a microelectromechanical system inertial measurement unit to track head posture in real time and calculates the gaze shift intensity and real-time motion vector through a built-in algorithm.
[0059] Fixation shift intensity is defined as a continuous quantitative indicator characterizing whether a user has a clear auditory search intent; specifically, this intensity... It's about the head rotation angular velocity. With duration The function, whose value range is normalized to the [0,1] interval, is used to quantify the strength of auditory search intent, rather than a simple yes / no logical state; it is based on the physiological characteristics of rapid eye movement and head coordination in biokinetics to distinguish between unconscious head shaking and purposeful sound source search actions.
[0060] The heterogeneous feature decoupling and fusion unit, as the core decision-making center of the system, performs deep coupling and arbitration of acoustic features and kinematic features. Under specific conditions, when the sound field uncertainty index exceeds the preset sound field uncertainty threshold and the gaze point shift intensity also exceeds the preset gaze point shift intensity threshold, the unit triggers a protection mechanism and generates a parameter freeze signal. This signal directly acts on the digital signal processor, forcibly maintaining the current scene classification mode and gain parameters, blocking parameter jumps caused by airflow noise or relative displacement of the sound source, thereby constructing a steady-state time window in the auditory dimension.
[0061] Within this steady-state window, the unit further performs feedforward control based on real-time motion vectors, calculates the beam advance compensation angle and superimposes it into the beamforming algorithm, driving the pickup beam to deflect in the same direction relative to the physical array, thus offsetting the system lag between electronic processing and mechanical motion. The auditory perception feedback optimization unit constructs a closed-loop verification circuit to monitor the clarity index of the speech signal within a specific time window after beam deflection. If the index decay is detected, it indicates a mismatch between the feedforward prediction and the actual sound source location, and a correction feedback signal is generated and fed back to the front end to drive the body motion intention capture unit to adaptively increase the trigger threshold of gaze point shift intensity, thereby reducing the system's sensitivity to motion data.
[0062] This technical solution introduces non-acoustic variables as arbitration criteria, maintaining high sensitivity to sudden environmental danger signals while eliminating false triggers caused by user actions through a parameter freezing mechanism, thus achieving a balance between high sensitivity and high auditory steady-state performance; and through beam lead compensation technology, it achieves spatial synchronization between auditory focus and visual focus.
[0063] Example 2:
[0064] The panoramic sound field perception unit is specifically used for:
[0065] Continuously monitor the microphone array signal;
[0066] Calculate the short-time energy variance and spectral entropy of the microphone array signal;
[0067] The sound field uncertainty index is obtained by weighting the short-time energy variance and spectral entropy.
[0068] The sound field uncertainty index is transmitted as a weighted reference signal to the heterogeneous feature decoupling and fusion unit;
[0069] The acoustic field uncertainty index is used to characterize the complexity of the current acoustic environment.
[0070] The panoramic sound field perception unit executes a rigorous signal quality assessment algorithm. This unit continuously performs analog-to-digital conversion on the analog signals acquired by the microphone array and calculates the short-time energy variance and spectral entropy within a preset time frame. The short-time energy variance quantifies the energy fluctuation amplitude of the sound signal in the time domain; a larger variance indicates more drastic transient changes in environmental noise. The spectral entropy quantifies the disorder of the frequency domain energy distribution based on the principle of information entropy; a higher entropy indicates a flatter sound spectrum, i.e., close to white noise or irregular background noise. To synthesize a unified evaluation standard, this unit uses a weighted summation algorithm, expressed by the formula: The system maintains a length of ,For example A historical sliding window of frames records the original energy variance in real time. Compared with the original spectral entropy The extreme values; the normalization calculation uses the following formula:
[0071]
[0072]
[0073] in, and These are the minimum and maximum values within the sliding window, respectively. To prevent extremely small values where the denominator is zero, such as 1e-6, the above normalization should be performed before substituting into the formula. Calculation; where This represents the uncertainty index of the sound field. The normalized energy variance The normalized spectral entropy, and For example, take the preset weighting coefficients. The value of this index is determined based on the sensitivity requirements of the application scenario to energy fluctuations or spectral complexity. The calculated sound field uncertainty index does not directly trigger mode switching, but is transmitted to the backend as a weighted reference signal to dynamically adjust the system's trust in the acoustic classification results. The introduction of this index enables the system to effectively identify high-dynamic instantaneous noise environments, avoid blindly classifying scenes when the signal-to-noise ratio fluctuates drastically, and ensure the robustness of subsequent fusion logic.
[0074] Example 3:
[0075] The body motion intent capture unit is specifically used for:
[0076] Collect raw data from the gyroscope and accelerometer;
[0077] Denoise the raw data;
[0078] Determine whether the angular velocity of head rotation exceeds a preset physiological threshold;
[0079] Determine whether the direction of head rotation remains singular and continuous;
[0080] When the angular velocity exceeds the preset physiological threshold and the direction remains unidirectional and continuous, increase the output value of the fixation point shift intensity;
[0081] Fixation shift intensity is used to characterize whether a user has a clear auditory search intent.
[0082] The specific intensity calculation follows the cumulative integral model:
[0083]
[0084] in, The linear decay constant for a single frame has a value of [value missing]. ; The intensity of the fixation shift at the current moment. For real-time angular velocity, To preset physiological thresholds, This is the normalized constant for the head rotation limit velocity. It is a cumulative coefficient; when a reversal or stop in the direction of motion is detected, Reset in a linear decreasing manner;
[0085] The body motion intent capture unit implements multi-level logic thresholds to filter out invalid motion interference; the acquisition end obtains high-frequency raw data from the gyroscope and accelerometer, and removes high-frequency jitter noise through a low-pass filter; the logic decision stage includes two core dimensions: the first dimension is angular velocity threshold discrimination, which compares the real-time angular velocity with a preset physiological threshold; this threshold is not set arbitrarily, but is a critical value determined based on ergonomic statistical data. Motion below this value is usually judged as unconscious body swaying, while motion above this value is regarded as rotation with potential intent;
[0086] Dimension two is the determination of directional continuity. The system detects whether the sign of the angular velocity direction flips within a continuous time window to eliminate repetitive oscillations. Only when the angular velocity exceeds the physiological threshold and the direction of movement remains unidirectional and continuous, will this unit significantly increase the output value of the gaze shift intensity. This processing logic converts biokinematic characteristics into digital control signals, accurately extracting the auditory search intent from complex head movements, ensuring that subsequent beam deflection operations are triggered only when the user actually performs the search action, effectively preventing the ineffective occupation of system resources.
[0087] Example 4:
[0088] The parameter freezing signal generated by the heterogeneous feature decoupling and fusion unit is specifically used for:
[0089] The acoustic classifier is prohibited from updating the scene pattern based on fluctuations in the microphone array signal;
[0090] The current scene classification results and gain parameters remain unchanged until the sound field uncertainty index or gaze shift intensity is lower than the corresponding threshold.
[0091] The parameter freezing signal generated by the heterogeneous feature decoupling and fusion unit performs a mandatory state locking operation. When the trigger condition is met, this signal is sent to the digital signal processing layer as the highest priority control command, cutting off the acoustic scene classifier's update response path to the real-time microphone signal. During this period, regardless of how the environmental signal picked up by the microphone changes due to wind noise or signal-to-noise ratio fluctuations, the system forcibly maintains the scene classification label and corresponding frequency gain curve unchanged at the moment of entering the high-dynamic scene. This locking state is maintained until the monitoring data indicates that the system has left the high-dynamic condition, that is, the sound field uncertainty index falls below the safety threshold, or the gaze point shift intensity indicates that the head rotation has stopped. This logical design fundamentally solves the scene misjudgment caused by the drastic changes in acoustic features induced by the head turning action itself in the existing technology, ensuring the continuity and stability of the user's auditory perception during the search for sound sources.
[0092] Example 5:
[0093] The heterogeneous feature decoupling and fusion unit is specifically used to calculate the beam lead compensation angle for:
[0094] Obtain the preset dynamic compensation coefficient, which is positively correlated with the system processing delay time;
[0095] The beam lead compensation angle is obtained by calculating the product of the preset dynamic compensation coefficient and the angular velocity in the real-time motion vector.
[0096] The beam lead compensation angle is used to compensate for the lag time between electronic signal processing and mechanical rotation.
[0097] The heterogeneous feature decoupling and fusion unit predicts the relative azimuth of the target sound source through a mathematical model; the core of the calculation lies in determining the beam lead compensation angle. Its calculation logic follows the formula ,in The real-time detected head rotation angular velocity, This is a preset dynamic compensation coefficient; The physical meaning corresponds to the total processing delay time of the system, and this coefficient can be dynamically calibrated by the system according to the current computing load; the system monitors the CPU utilization rate of the digital signal processor (DSP) in real time. Dynamic compensation coefficient The calculation follows a linear mapping model; to ensure prediction accuracy, the system uses a task scheduling strategy to... The response is controlled within the linear response range of [0, 85%]. Within this range, the delay coefficient is approximately linearly related to the load, and the calculation follows the linear mapping model below:
[0098]
[0099] in, This is due to the inherent hardware latency when the system is idle. For reference load, for example, take the CPU utilization rate of the system in idle state as 10%. The load delay conversion factor, for example, is 0.5ms / %, representing the additional queuing delay introduced by each 1% increase in load. By multiplying the angular velocity by the delay time, the system accurately calculates the head rotation angle during the signal processing lag. The system superimposes this compensation angle into the azimuth parameter of the beamforming filter, so that the generated pickup beam has a forward deflection in the same direction relative to the current head orientation in physical space. This technique effectively aligns the auditory focus and visual focus on the time axis, canceling the inherent lag effect of the electronic system. This allows the beam to reach the target position synchronously or even slightly ahead of the line of sight when the user quickly turns their head to find the sound source, achieving a zero-delay spatial auditory experience.
[0100] Example 6:
[0101] The auditory perception feedback optimization unit is specifically used for:
[0102] Obtain envelope integrity as a sharpness indicator;
[0103] When a decrease in envelope integrity is detected, it is determined that a beam deflection has been misjudged;
[0104] A correction feedback signal is generated and sent to the body motion intent capture unit to reduce the system's sensitivity to head motion data.
[0105] The auditory perception feedback optimization unit performs the system's self-diagnosis and parameter calibration functions. Within the monitoring window after beam deflection, this unit analyzes the output audio stream and extracts envelope integrity as a key technical indicator for measuring speech intelligibility. The normalized variance stability of the speech signal envelope is obtained by calculating it using the Hilbert transform. The specific formula is as follows:
[0106]
[0107] in The Hilbert envelope amplitude of the audio frame after beam deflection. The mean; To calculate the total number of sampling points within the window, The sampling point number; A higher value indicates that the envelope fluctuations are more consistent with the characteristics of natural speech, and are not truncated or smoothed. Envelope integrity reflects the clarity of the temporal envelope of the speech signal and is positively correlated with speech intelligibility. If the monitoring data shows a significant decrease in envelope integrity after beam deflection, the system logic judges that the motion-based feedforward deflection was not aligned with the effective sound source or introduced additional noise, which is considered a misjudgment. Based on this result, the unit generates a correction feedback signal and sends it back to the front end. The body motion intent capture unit responds to this signal and executes an adaptive adjustment strategy, gradually increasing the trigger threshold of gaze shift intensity. The threshold adjustment follows the following step formula:
[0108]
[0109] in, The preset penalty step size, for example, 0.05. This is a preset lower threshold for sensitivity; this operation aims to suppress frequent switching in environments with high false positive rates by raising the trigger threshold; making the system more conservative in judging head movements in similar subsequent working conditions; this mechanism gives the system the ability to learn adaptively for different users, ensuring the accuracy of judgment and auditory comfort during long-term use.
[0110] In addition, to prevent the system from remaining in a low-sensitivity state for extended periods, the auditory perception feedback optimization unit also maintains a recovery timer; when envelope integrity is detected... In consecutive time windows For example, continuously exceeding the preset excellent threshold for 5 seconds. When this happens, the system executes a threshold callback operation:
[0111]
[0112] in The preset recovery step size ensures that the system can gradually recover its sensitive response to head movements after the environment stabilizes.
[0113] Example 7:
[0114] Please see Figure 2 A method for switching modes in a smart hearing aid includes the following steps:
[0115] The ambient sound signal is collected by the panoramic sound field sensing unit, and the variance of short-time energy and spectral entropy value are calculated to generate the sound field uncertainty index.
[0116] The proprioception capture unit collects data from the inertial measurement unit, calculates the angular velocity and direction of head rotation, and outputs the gaze shift intensity and real-time motion vector.
[0117] The sound field uncertainty index and gaze shift intensity are received by a heterogeneous feature decoupling and fusion unit.
[0118] When the sound field uncertainty index is greater than the preset sound field uncertainty threshold and the gaze point shift intensity is greater than the preset gaze point shift intensity threshold, a parameter freeze signal is generated to lock the scene parameters.
[0119] The beam advance compensation angle is calculated based on real-time motion vector by the heterogeneous feature decoupling and fusion unit, and the beam direction is deflected in the same direction.
[0120] The auditory perception feedback optimization unit monitors the speech intelligibility index after beam deflection, and generates a correction feedback signal when the speech intelligibility index decreases.
[0121] The proprioception intent capture unit adjusts the trigger threshold for gaze shift intensity based on the correction feedback signal.
[0122] This embodiment details the execution flow of a smart hearing aid mode switching method, which strictly follows the control closed loop of perception, decision-making, execution and feedback;
[0123] The system's multidimensional perception foundation was established. At this stage, the system processes acoustic and kinematic data in parallel: the panoramic sound field perception unit quantifies the complexity of the environment by calculating energy variance and spectral entropy, and outputs the sound field uncertainty index; the ontological motion intent capture unit analyzes IMU data and outputs the gaze shift intensity and specific motion vectors that characterize the user's intent; the heterogeneous feature decoupling and fusion unit serves as a data convergence point, receiving and aligning these two heterogeneous data streams in real time.
[0124] The execution of the core control logic is explained; the system execution condition judgment: the parameter freeze is triggered only when the dual conditions of complex environment and clear user intention are met, and the scene parameters are locked to shield dynamic noise interference; the system calls the feedforward control algorithm, uses real-time motion vector and system delay coefficient to calculate compensation angle, drives beam to actively deflect ahead, and realizes predictive following of auditory focus;
[0125] An adaptive feedback mechanism for the system was established; after beam deflection, the auditory perception feedback optimization unit immediately evaluates the speech clarity; once the clarity index is found to deteriorate, indicating that the deflection strategy has failed, the system immediately generates a correction feedback signal to dynamically increase the trigger threshold of the front end.
[0126] This method, through the coordinated operation of the above steps, achieves accurate capture and response to the user's auditory intent in complex sound field environments, effectively solving the problems of parameter jumps and tracking lag in traditional technologies.
[0127] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A smart hearing aid mode switching system, characterized in that, include: A panoramic sound field sensing unit is used to acquire microphone array signals in real time and extract acoustic features; The panoramic sound field sensing unit is used to calculate the sound field uncertainty index based on the variance of short-time energy and the spectral entropy value. The proprioception intent capture unit is used to collect head motion data through an inertial measurement unit; the proprioception intent capture unit is used to calculate the gaze shift intensity and real-time motion vector based on the head motion data; A heterogeneous feature decoupling and fusion unit is connected to both the panoramic sound field perception unit and the body motion intent capture unit. The heterogeneous feature decoupling and fusion unit synchronously receives the sound field uncertainty index, the gaze point shift intensity, and the real-time motion vector. When the sound field uncertainty index is greater than a preset sound field uncertainty threshold and the gaze point shift intensity is greater than a preset gaze point shift intensity threshold, a parameter freezing signal is generated. This parameter freezing signal is used to lock the scene classification result and gain parameters. The heterogeneous feature decoupling and fusion unit calculates the beam advance compensation angle based on the real-time motion vector. The heterogeneous feature decoupling and fusion unit superimposes the beam advance compensation angle into the beamforming parameters to drive the beam pointing to perform a unidirectional advance deflection. An auditory perception feedback optimization unit is used to monitor the clarity index of the target speech within a preset time window after beam deflection; the auditory perception feedback optimization unit is used to generate a correction feedback signal when the clarity index decreases; The proprioception intent capture unit is used to increase the trigger threshold of the gaze shift intensity in response to the correction feedback signal.
2. The intelligent hearing aid mode switching system as described in claim 1, characterized in that, The panoramic sound field sensing unit is specifically used for: Continuously monitor the microphone array signal; Calculate the short-time energy variance and spectral entropy of the microphone array signal; The sound field uncertainty index is obtained by weighting the short-time energy variance and the spectral entropy value. The sound field uncertainty index is transmitted as a weighted reference signal to the heterogeneous feature decoupling and fusion unit; The sound field uncertainty index is used to characterize the complexity of the current acoustic environment.
3. The intelligent hearing aid mode switching system as described in claim 1, characterized in that, The body motion intent capture unit is specifically used for: Collect raw data from the gyroscope and accelerometer; The original data is then denoised. Determine whether the angular velocity of head rotation exceeds a preset physiological threshold; Determine whether the direction of head rotation remains singular and continuous; When the angular velocity exceeds the preset physiological threshold and the direction remains single and continuous, the output value of the fixation point shift intensity is increased; The gaze shift intensity is used to characterize whether the user has a clear auditory search intent.
4. The intelligent hearing aid mode switching system as described in claim 1, characterized in that, The parameter freezing signal generated by the heterogeneous feature decoupling and fusion unit is specifically used for: The acoustic classifier is prohibited from updating the scene mode based on fluctuations in the microphone array signal; The current scene classification result and gain parameters remain unchanged until the sound field uncertainty index or the gaze shift intensity is lower than the corresponding threshold.
5. The intelligent hearing aid mode switching system as described in claim 1, characterized in that, The heterogeneous feature decoupling and fusion unit is specifically used to calculate the beam lead compensation angle for: Obtain a preset dynamic compensation coefficient, which is positively correlated with the system processing delay time; The product of the preset dynamic compensation coefficient and the angular velocity in the real-time motion vector is calculated to obtain the beam lead compensation angle; The beam advance compensation angle is used to offset the lag time between electronic signal processing and mechanical rotation.
6. The intelligent hearing aid mode switching system as described in claim 1, characterized in that, The auditory perception feedback optimization unit is specifically used for: The envelope completeness is obtained as the sharpness indicator; When a decrease in envelope integrity is detected, it is determined that a beam deflection has been misjudged; The correction feedback signal is generated and sent to the body motion intent capture unit to reduce the system's sensitivity to the head motion data.
7. A method for switching modes in an intelligent hearing aid, applied to an intelligent hearing aid mode switching system as described in any one of claims 1 to 6, characterized in that, Includes the following steps: The ambient sound signal is collected by the panoramic sound field sensing unit, and the variance of short-time energy and spectral entropy value are calculated to generate the sound field uncertainty index. The proprioception capture unit collects data from the inertial measurement unit, calculates the angular velocity and direction of head rotation, and outputs the gaze shift intensity and real-time motion vector. The sound field uncertainty index and the gaze point shift intensity are received by the heterogeneous feature decoupling and fusion unit; When the sound field uncertainty index is greater than a preset sound field uncertainty threshold and the gaze point shift intensity is greater than a preset gaze point shift intensity threshold, a parameter freeze signal is generated to lock the scene parameters. The heterogeneous feature decoupling and fusion unit calculates the beam advance compensation angle based on the real-time motion vector, and performs a unidirectional advance deflection of the beam direction. The auditory perception feedback optimization unit monitors the speech intelligibility index after beam deflection, and generates a correction feedback signal when the speech intelligibility index decreases. The proprioception capture unit adjusts the trigger threshold of the gaze point shift intensity according to the correction feedback signal.
Citation Information
Patent Citations
Phase synchronization compensation algorithm of earphone bone conduction-air conduction mixed sound field
CN119922457A
Spatial sound effect testing method and system of sound system
CN120881494A