Multi-modal sensor fused earphone hybrid noise reduction control system
By using a multimodal sensor fusion system that combines information from multiple sensors and deep learning algorithms, the noise reduction mode is dynamically adjusted, solving the problems of insufficient noise detection accuracy and environmental adaptability in existing headphone noise reduction technologies, and realizing personalized headphone noise reduction control.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ANHUI XINYUCHI SEMICONDUCTOR TECHNOLOGY CO LTD
- Filing Date
- 2026-01-20
- Publication Date
- 2026-05-05
AI Technical Summary
Existing headphone noise cancellation technology fails to fully integrate information from multiple sensors, making it unable to effectively cope with dynamic changes in environmental noise, differences in headphone wearing status, and the impact of user movement, resulting in poor noise detection accuracy and environmental adaptability.
A multimodal sensor fusion system is adopted, including a microphone array, bone conduction vibration sensor, triaxial accelerometer, capacitive proximity sensor and pressure sensor. A correlation model between ambient sound intensity and user hearing safety threshold is established through deep learning algorithm. The weights of active noise cancellation, passive noise cancellation and directional noise cancellation are dynamically adjusted by combining wearing status and motion status characteristics.
It improves the accuracy and environmental adaptability of noise detection, realizes personalized noise reduction control, balances noise reduction effect with health and safety, and solves the problem that a single sensor cannot cope with complex scenarios.
Smart Images

Figure CN121985246A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of headphone noise reduction technology, specifically to a headphone hybrid noise reduction control system based on multimodal sensor fusion. Background Technology
[0002] As people's demands for audio experience continue to increase, headphone noise cancellation technology is becoming increasingly important. Traditional noise-canceling headphones mainly use either active noise cancellation (ANC) or passive noise cancellation (PNC) technology.
[0003] The reference patent is titled: "A Noise Reduction Method and System for Headphones Based on Multiple Noise Sources" (Patent Publication No.: CN119893362A, Patent Publication Date: 2025-04-25). It includes a multi-array microphone system for collecting effective user sound signals and signals from different noise sources; a noise reduction mode control system that enables switching between sports mode, video viewing mode, call mode, and audio loading mode, as well as standard noise reduction mode, comfortable noise reduction mode, and enhanced noise reduction mode, all through a user terminal interface; a wireless headphone noise reduction system, a user terminal headphone noise reduction system, and a data computing cloud service center employing a three-tiered "cloud-edge-device" architecture, using a built-in intelligent voice noise reduction method to eliminate noise signals in mixed voice signals, improving user experience; a cloud transmission architecture design combining wired and wireless composite information transmission systems to achieve long-distance, high-capacity, and efficient information transmission; and a voice enhancement system that uses a built-in intelligent voice enhancement method to perform final processing on the noise-reduced sound signal, eliminating mixing and echo interference, and outputting a high-fidelity sound signal.
[0004] Based on the description in the above documents, existing headphones process noise using a single sensor without combining information from other sensors for comprehensive judgment. Furthermore, they fail to fully consider the dynamic changes in environmental noise, differences in headphone wearing status, and the impact of user movement on noise reduction performance, resulting in poor noise detection accuracy and environmental adaptability. Therefore, this invention provides a multimodal sensor fusion-based hybrid noise reduction control system for headphones. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention provides a multimodal sensor fusion-based hybrid noise reduction control system for headphones. This system solves the problems of existing headphones processing noise through a single sensor without combining information from other sensors for comprehensive judgment, and failing to fully consider the dynamic changes in environmental noise, differences in headphone wearing status, and the impact of user movement on noise reduction performance, resulting in poor noise detection accuracy and environmental adaptability.
[0006] To achieve the above objectives, the present invention provides the following technical solution: a multimodal sensor fusion-based headphone hybrid noise reduction control system, comprising:
[0007] The multimodal information acquisition module uses multiple types of sensors to collect environmental noise and real-time user parameter information;
[0008] The data processing module preprocesses the collected data, establishes a correlation model between ambient sound intensity and the user's hearing safety threshold, and assesses the user's hearing status.
[0009] The scene recognition and noise processing module prioritizes determining the user's wearing status of the headphones and determines the motion state after wearing them based on posture data. It fuses the features of the wearing status parameters and motion state parameters and performs simultaneous active noise reduction, passive noise reduction and directional noise reduction based on the fused features.
[0010] The control feedback module generates noise processing instructions and controls the output audio signal based on these instructions. It also dynamically adjusts parameters based on user noise reduction feedback logs and electromyographic signals to achieve personalized noise reduction.
[0011] Preferably, the multi-type sensors of the multimodal information acquisition module include:
[0012] Microphone array: Each earphone integrates four microphones and adopts a four-microphone array. The feedforward microphone is located on the outside of the earphone to collect ambient noise in real time, and the feedback microphone is located on the inside of the ear canal to monitor the actual noise reduction effect. The signal is converted into a digital signal by an ADC for subsequent processing.
[0013] Bone conduction vibration sensors are used to capture speech signals transmitted through the jawbone and the fundamental frequency of vocal cord vibration.
[0014] A combination of a three-axis accelerometer and a three-axis gyroscope is used to collect the motion parameters of the user's head;
[0015] The system uses a combination of capacitive proximity sensors and pressure sensors to automatically identify the wearing status of the headphones.
[0016] Preferably, the data processing module assesses the user's hearing status as follows:
[0017] The signals collected by each sensor are filtered, denoised, and standardized. The Kalman filter algorithm is used to eliminate random interference in the noise acquisition unit signal, and the moving average filter is used to eliminate high-frequency jitter in the motion sensing unit signal. The standardization process normalizes each signal to the [0,1] interval to ensure data consistency.
[0018] Then, a correlation model between ambient sound intensity and user hearing safety threshold is established using deep learning algorithms. Through transfer learning on a large amount of ambient sound data and user hearing data, the model determines the current ambient noise and outputs a dynamic safe sound pressure threshold curve.
[0019] Preferably, the operation for forming the safe sound pressure threshold curve is as follows:
[0020] Extract environmental sound parameters from historical data, and extract the user's hearing sound pressure level over sequential time based on the same environmental sound parameters;
[0021] That is, after extracting the hearing sound pressure value according to the sequential time, if it is determined that the hearing sound pressure value has not changed within the set time interval, then the hearing sound pressure value is extracted, and then the average of multiple hearing sound pressure values is obtained to obtain the safe sound pressure threshold under the current environmental sound parameters.
[0022] Using the changes in environmental sound parameters from small to large as the parameters of the horizontal axis, and the corresponding safe sound pressure threshold under the environmental sound parameters as the parameters of the vertical axis, the parameter points located on the horizontal and vertical axes are connected by curves to obtain the safe sound pressure threshold curve.
[0023] Preferably, the scene recognition and noise processing module determines the user's wearing status of the headphones by:
[0024] The wearing status is determined by extracting the parameter data collected by the capacitive proximity sensor and the pressure sensor and the factory-set threshold. The capacitive proximity sensor is used to detect the contact distance between the headphones and the user's ear as 'a', and the pressure sensor is set on the edge of the earcup to detect the contact pressure between the headphones and the ear as 'b'.
[0025] Furthermore, the factory-set distance threshold is A, and the ear contact pressure threshold is B. The specific results are as follows:
[0026] Result 1: When a < A and b ≥ B, the current wearing state is stable.
[0027] Result 2: When a≥A and b<B, the current wearing state is loose.
[0028] Result 3: When a has no data and b=0, the current state is not wearing the device.
[0029] Preferably, the operation of determining the motion state after wearing the device based on posture data in the scene recognition and noise processing module is as follows:
[0030] The built-in three-axis gyroscope tracks the three-dimensional motion of the user's head in real time, and determines the direction and speed of the user's head based on the data obtained from the three-axis accelerometer.
[0031] Taking the initial position of the sensor when the headphones are first worn as the initial point, and extracting the triaxial accelerometer data in the subsequent time period, we can obtain the angular velocity c and acceleration d of the headphone sensor at the initial point.
[0032] Based on historical data, the angular velocity range [m, n] and acceleration range [p, q] for different motion states were derived.
[0033] If c < m, d < p, and the duration is t, then the current user is in a static state. , If c > n and d > q, and the duration is t, then the current user is in a walking state. If c > n and d > q, and the duration is t, then the current user is in a running state.
[0034] Preferably, the feature fusion operation of wearing state parameters and motion state parameters in the scene recognition and noise processing module is as follows:
[0035] That is, by extracting real-time wearing status parameters and motion status parameters and comparing them with the set contact distance threshold and ear contact pressure threshold, it is determined whether there are changes under different motion status parameters.
[0036] It also generates adjustment instructions and provides feedback in a timely manner based on changes in the situation.
[0037] Preferably, the simultaneous operation of active noise reduction, passive noise reduction, and directional noise reduction based on fused features in the scene recognition and noise processing module is as follows:
[0038] By extracting the feature parameters of the noise signal after feature fusion, and combining the wearing state parameters and motion state parameters, the weights of active noise reduction and passive noise reduction are dynamically adjusted.
[0039] Simultaneously, the direction of the sound wave from the sound source is located based on the motion state parameters, and a reverse sound wave is emitted in a specific direction to achieve noise suppression.
[0040] Preferably, the operation of adjusting the weighting values of active noise reduction and passive noise reduction is as follows:
[0041] If the wearing state is a stable wearing state, the amplitude of the noise signal is extracted and compared with the set signal amplitude;
[0042] If the amplitude of the extracted noise signal is less than the set signal amplitude, then the active noise reduction weight is set to u1, and the passive noise reduction weight is set to u1. ,and If the amplitude of the extracted noise signal is greater than the set signal amplitude, then the active noise reduction weight is set to u2, and the passive noise reduction weight is set to u2. ,and ;
[0043] If the wearing condition is loose, then set the active noise cancellation weight to [value]. Passive noise reduction accounts for a significant portion of the weighting. ,and ;
[0044] If the device is not worn, both active and passive noise cancellation functions will be automatically turned off to reduce power consumption.
[0045] Preferably, the operation of locating the sound wave direction of the sound source based on motion state parameters is as follows:
[0046] By determining the arrival timestamps of the same noise signal received by each microphone in the four-microphone array, the direction of the sound wave of the sound source can be located according to the order of the timestamps.
[0047] It also updates the change in the direction of the sound source by combining the change in motion state, so as to generate a noise reduction signal with the opposite direction of the sound wave of the sound source for cancellation.
[0048] This invention provides a multimodal sensor fusion-based hybrid noise cancellation control system for headphones. Compared with existing technologies, it has the following advantages:
[0049] 1. This multimodal sensor fusion-based headphone hybrid noise reduction control system, by employing multimodal sensor fusion and combining information from a dual-microphone array, wear detection sensor, and motion sensing sensor, achieves comprehensive perception of environmental noise, wearing status, and user movement status. This improves the accuracy of noise detection and environmental adaptability, solves the problem that a single sensor cannot cope with complex scenarios, and realizes dynamic weight adjustment of active and passive noise reduction. It also introduces directional noise reduction to improve the quality of headphone hybrid noise reduction.
[0050] 2. This multimodal sensor-integrated headphone hybrid noise reduction control system establishes a correlation model between ambient sound intensity and user hearing safety threshold through deep learning algorithms. It generates a dynamic safe sound pressure threshold curve through transfer learning, breaking through the limitations of traditional noise reduction that only focuses on noise cancellation and ignores hearing protection. It can dynamically adjust the sound pressure range of the output audio based on the user's historical hearing data and real-time ambient noise intensity. At the same time, the scene recognition and noise processing modules fuse features of wearing status and movement status to balance noise reduction effect and health and safety.
[0051] 3. The headphone hybrid noise reduction control system, which integrates multimodal sensors, dynamically adjusts the weight ratio of active and passive noise reduction by fusing features of wearing and motion states through scene recognition and noise processing modules. Combined with directional noise reduction technology, it achieves precise noise reduction. The multi-mode collaborative strategy solves the problems of fixed hybrid noise reduction modes and noise reduction effect that are greatly affected by wearing and motion states in existing systems. Attached Figure Description
[0052] Figure 1 This is a schematic diagram of the noise reduction control system of the present invention. Detailed Implementation
[0053] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0054] Please see Figure 1 This invention provides a technical solution: a multimodal sensor fusion-based headphone hybrid noise reduction control system, comprising:
[0055] The multimodal information acquisition module uses multiple types of sensors to collect environmental noise and real-time user parameter information;
[0056] The data processing module preprocesses the collected data, establishes a correlation model between ambient sound intensity and the user's hearing safety threshold, and assesses the user's hearing status.
[0057] The scene recognition and noise processing module prioritizes determining the user's wearing status of the headphones and determines the motion state after wearing them based on posture data. It fuses the features of the wearing status parameters and motion state parameters and performs simultaneous active noise reduction, passive noise reduction and directional noise reduction based on the fused features.
[0058] The control feedback module generates noise processing instructions and controls the output audio signal based on these instructions. It also dynamically adjusts parameters based on user noise reduction feedback logs and electromyographic signals to achieve personalized noise reduction.
[0059] Among them, noise processing instruction generation: Based on the scene recognition results and the safe sound pressure threshold curve, the main control unit generates control instructions including noise reduction weight, pointing angle, and upper limit of output sound pressure. According to the control instructions, the audio signal is amplified to the safe sound pressure threshold range while suppressing signal distortion.
[0060] Furthermore, the accompanying app records user actions (such as manually adjusting noise cancellation intensity and switching modes), stores them in the format of "time + scenario + adjustment parameters", and generates a user preference report every month. An EMG electromyography sensor is integrated inside the earcup to collect electromyography signals from the muscles behind the ear. When the user frowns or clenches their teeth due to poor noise cancellation, the amplitude of the electromyography signal exceeds the threshold, triggering parameter adjustment.
[0061] Every 30 seconds, user feedback logs and electromyography (EMG) signal data are integrated. If the user increases the noise reduction intensity multiple times in the walking scenario, the active noise reduction weight in that scenario will be automatically increased by 5%. If the EMG signal is frequently triggered and the safe sound pressure threshold has not reached the upper limit, the output sound pressure will be appropriately increased to achieve personalized adaptation.
[0062] By employing multimodal sensor fusion, combining information from a dual-microphone array, wear detection sensors, and motion sensing sensors, comprehensive perception of environmental noise, wearing status, and user movement status is achieved. This improves the accuracy of noise detection and environmental adaptability, solves the problem that a single sensor cannot cope with complex scenarios, and enables dynamic weight adjustment of active and passive noise cancellation. The introduction of directional noise cancellation also improves the quality of hybrid noise cancellation in headphones.
[0063] In this embodiment of the invention, the multi-type sensors of the multimodal information acquisition module include:
[0064] Microphone array: Each earphone integrates four microphones and adopts a four-microphone array. The feedforward microphone is located on the outside of the earphone to collect ambient noise in real time, and the feedback microphone is located on the inside of the ear canal to monitor the actual noise reduction effect. The signal is converted into a digital signal by an ADC for subsequent processing.
[0065] Bone conduction vibration sensors are used to capture speech signals transmitted through the jawbone and the fundamental frequency of vocal cord vibration.
[0066] A combination of a three-axis accelerometer and a three-axis gyroscope is used to collect the motion parameters of the user's head;
[0067] The system uses a combination of capacitive proximity sensors and pressure sensors to automatically identify the wearing status of the headphones.
[0068] The bone conduction vibration sensor uses the SBT001 bone conduction sensor, which is attached to the inside of the earmuff near the temporal bone to capture the voice signal conducted by the jawbone. The communication module integrates a Bluetooth 5.2 module (nRF52840), which supports BLE audio transmission with a latency of ≤20ms, and is used to synchronize the noise reduction feedback commands of the user terminal.
[0069] In this embodiment of the invention, the data processing module assesses the user's hearing status as follows:
[0070] The signals collected by each sensor are filtered, denoised, and standardized. The Kalman filter algorithm is used to eliminate random interference in the noise acquisition unit signal, and the moving average filter is used to eliminate high-frequency jitter in the motion sensing unit signal. The standardization process normalizes each signal to the [0,1] interval to ensure data consistency.
[0071] Then, a correlation model between ambient sound intensity and user hearing safety threshold is established using deep learning algorithms. Through transfer learning on a large amount of ambient sound data and user hearing data, the model determines the current ambient noise and outputs a dynamic safe sound pressure threshold curve.
[0072] Among them, filtering processing: Kalman filtering algorithm is used for microphone array signals to eliminate random interference such as airflow and electromagnetic interference; moving average filtering with a window size of 5 is used for motion sensing sensor signals to eliminate high-frequency jitter.
[0073] The association model training adopts a CNN-LSTM hybrid deep learning model. The input layer is the ambient sound intensity (after normalization), and the output layer is the hearing safety threshold. The training dataset contains tens of thousands of sets of ambient sound data and tens of thousands of sets of user hearing data (different ages and hearing conditions). The model parameters are fine-tuned through transfer learning. The safe sound pressure threshold is calculated by extracting historical ambient sound parameters from the past 3 months. For the same ambient sound parameter (e.g., 60dB), the hearing sound pressure value that has not changed within 5 consecutive seconds is extracted (at least 20 sets are collected), and the mean is calculated as the safe sound pressure threshold for that environment.
[0074] In this embodiment of the invention, the operation for forming the safe sound pressure threshold curve is as follows:
[0075] Extract environmental sound parameters from historical data, and extract the user's hearing sound pressure level over sequential time based on the same environmental sound parameters;
[0076] That is, after extracting the hearing sound pressure value according to the sequential time, if it is determined that the hearing sound pressure value has not changed within the set time interval, then the hearing sound pressure value is extracted, and then the average of multiple hearing sound pressure values is obtained to obtain the safe sound pressure threshold under the current environmental sound parameters.
[0077] Using the changes in environmental sound parameters from small to large as the parameters of the horizontal axis, and the corresponding safe sound pressure threshold under the environmental sound parameters as the parameters of the vertical axis, the parameter points located on the horizontal and vertical axes are connected by curves to obtain the safe sound pressure threshold curve.
[0078] By establishing a correlation model between ambient sound intensity and user hearing safety threshold through deep learning algorithms, and generating a dynamic safe sound pressure threshold curve through transfer learning, this approach breaks through the limitations of traditional noise reduction that only focuses on noise cancellation and ignores hearing protection. It can dynamically adjust the sound pressure range of the output audio based on the user's historical hearing data and real-time ambient noise intensity. At the same time, the scene recognition and noise processing modules fuse features of wearing status and movement status to balance noise reduction effect and health and safety.
[0079] In this embodiment of the invention, the scene recognition and noise processing module determines the user's wearing status of the headphones as follows:
[0080] The wearing status is determined by extracting the parameter data collected by the capacitive proximity sensor and the pressure sensor and the factory-set threshold. The capacitive proximity sensor is used to detect the contact distance between the headphones and the user's ear as 'a', and the pressure sensor is set on the edge of the earcup to detect the contact pressure between the headphones and the ear as 'b'.
[0081] Furthermore, the factory-set distance threshold is A, and the ear contact pressure threshold is B. The specific results are as follows:
[0082] Result 1: When a < A and b ≥ B, the current wearing state is stable.
[0083] Result 2: When a≥A and b<B, the current wearing state is loose.
[0084] Result 3: When a has no data and b=0, the current state is not wearing the device.
[0085] In this embodiment of the invention, the operation of determining the motion state after wearing the device based on posture data in the scene recognition and noise processing module is as follows:
[0086] The built-in three-axis gyroscope tracks the three-dimensional motion of the user's head in real time, and determines the direction and speed of the user's head based on the data obtained from the three-axis accelerometer.
[0087] Taking the initial position of the sensor when the headphones are first worn as the initial point, and extracting the triaxial accelerometer data in the subsequent time period, we can obtain the angular velocity c and acceleration d of the headphone sensor at the initial point.
[0088] Based on historical data, the angular velocity range [m, n] and acceleration range [p, q] for different motion states were derived.
[0089] If c < m, d < p, and the duration is t, then the current user is in a static state. If c > n and d > q, and the duration is t, then the current user is in a walking state. If c > n and d > q, and the duration is t, then the current user is in a running state.
[0090] For example, setting a threshold range: Based on a large amount of user test data, setting an angular velocity range. acceleration range The duration is t = 2 seconds;
[0091] State determination logic: Taking the initial wearing position as the origin, calculate the average angular velocity c and average acceleration d every 100ms thereafter; if and If it lasts for 2 seconds, it is considered a stationary state (such as working in an office); if and If it lasts for 2 seconds, it is determined to be a walking state (such as outdoor commuting); if and The signal lasts for 2 seconds and is considered to be in a running state (such as a morning run).
[0092] State Correction: When the motion state changes abruptly (such as changing from walking to running), the judgment duration is shortened to 1 second to improve response speed.
[0093] In this embodiment of the invention, the feature fusion operation of wearing state parameters and motion state parameters in the scene recognition and noise processing module is as follows:
[0094] That is, by extracting real-time wearing status parameters and motion status parameters and comparing them with the set contact distance threshold and ear contact pressure threshold, it is determined whether there are changes under different motion status parameters.
[0095] It also generates adjustment instructions and provides feedback in a timely manner based on changes in the situation.
[0096] In this embodiment of the invention, the simultaneous operation of active noise reduction, passive noise reduction, and directional noise reduction based on fusion features in the scene recognition and noise processing module is as follows:
[0097] By extracting the feature parameters of the noise signal after feature fusion, and combining the wearing state parameters and motion state parameters, the weights of active noise reduction and passive noise reduction are dynamically adjusted.
[0098] Simultaneously, the direction of the sound wave from the sound source is located based on the motion state parameters, and a reverse sound wave is emitted in a specific direction to achieve noise suppression.
[0099] By integrating features from wearing status and motion state through scene recognition and noise processing modules, the weight ratio of active noise cancellation and passive noise cancellation is dynamically adjusted. Combined with directional noise cancellation technology, precise noise cancellation is achieved. The multi-mode collaborative strategy solves the problems of fixed hybrid noise cancellation modes and noise cancellation effect that are greatly affected by wearing status and motion state in existing systems.
[0100] In this embodiment of the invention, the operation of adjusting the weighting values of active noise reduction and passive noise reduction is as follows:
[0101] If the wearing state is a stable wearing state, the amplitude of the noise signal is extracted and compared with the set signal amplitude;
[0102] If the amplitude of the extracted noise signal is less than the set signal amplitude, then the active noise reduction weight is set to [value missing]. Passive noise reduction accounts for a significant portion of the weighting. ,and , If the amplitude of the extracted noise signal is greater than the set signal amplitude, then the active noise reduction weight is set to [value missing]. Passive noise reduction accounts for a significant portion of the weighting. ,and , ;
[0103] If the wearing condition is loose, then set the active noise cancellation weight to [value]. Passive noise reduction accounts for a significant portion of the weighting. ,and , ;
[0104] If the device is not worn, both active and passive noise cancellation functions will be automatically turned off to reduce power consumption.
[0105] For example, noise reduction weight adjustment:
[0106] Stable wearing condition: If the noise signal amplitude (after normalization) is <0.5 (corresponding to a low-noise environment, such as a library), set the active noise reduction weight. Passive noise reduction weights If the noise signal amplitude is ≥0.5 (corresponding to a high-noise environment, such as a subway station), set... ;
[0107] Loose fit: Due to decreased passive noise cancellation effect, the settings... The problem of insufficient sealing is compensated by enhancing active noise reduction;
[0108] When not worn: Immediately shuts down the active noise cancellation module and passive noise cancellation drive circuit to reduce power consumption (standby power consumption ≤ 5mA).
[0109] In this embodiment of the invention, the operation of locating the sound wave direction of the sound source based on motion state parameters is as follows:
[0110] By determining the arrival timestamps of the same noise signal received by each microphone in the four-microphone array, the direction of the sound wave of the sound source can be located according to the order of the timestamps.
[0111] It also updates the change in the direction of the sound source by combining the change in motion state, so as to generate a noise reduction signal with the opposite direction of the sound wave of the sound source for cancellation.
[0112] The directional noise reduction is achieved by setting the spacing between the four microphone arrays to 2cm, and by detecting the time difference (accuracy) of the same noise signal arriving at the four microphones. The sound source direction is calculated using the TDOA (Time Difference of Arrival) algorithm (angle error ≤ 3°). Combined with motion state data (such as head swing angle during running), the sound source direction is updated every 50ms to generate a noise reduction signal with the opposite direction, equal amplitude, and opposite phase to the sound source direction. This signal is then output to the speaker through the audio processing unit to achieve directional noise cancellation.
[0113] Furthermore, any content not described in detail in this specification is existing technology known to those skilled in the art.
[0114] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0115] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A multimodal sensor fusion-based hybrid noise reduction control system for headphones, characterized in that: include: The multimodal information acquisition module uses multiple types of sensors to collect environmental noise and real-time user parameter information; The data processing module preprocesses the collected data, establishes a correlation model between ambient sound intensity and the user's hearing safety threshold, and assesses the user's hearing status. The scene recognition and noise processing module prioritizes determining the user's wearing status of the headphones and determines the motion state after wearing them based on posture data. It fuses the features of the wearing status parameters and motion state parameters and performs simultaneous active noise reduction, passive noise reduction and directional noise reduction based on the fused features. The control feedback module generates noise processing instructions and controls the output audio signal based on these instructions. It also dynamically adjusts parameters based on user noise reduction feedback logs and electromyographic signals to achieve personalized noise reduction.
2. The headphone hybrid noise reduction control system based on multimodal sensor fusion according to claim 1, characterized in that: The multi-modal information acquisition module includes the following types of sensors: Microphone array: Each earphone integrates four microphones and adopts a four-microphone array. The feedforward microphone is located on the outside of the earphone to collect ambient noise in real time, and the feedback microphone is located on the inside of the ear canal to monitor the actual noise reduction effect. The signal is converted into a digital signal by an ADC for subsequent processing. Bone conduction vibration sensors are used to capture speech signals transmitted through the jawbone and the fundamental frequency of vocal cord vibration. A combination of a three-axis accelerometer and a three-axis gyroscope is used to collect the motion parameters of the user's head; The system uses a combination of capacitive proximity sensors and pressure sensors to automatically identify the wearing status of the headphones.
3. The headphone hybrid noise reduction control system based on multimodal sensor fusion according to claim 1, characterized in that: The data processing module assesses the user's hearing status as follows: The signals collected by each sensor are filtered, denoised, and standardized. The Kalman filter algorithm is used to eliminate random interference in the noise acquisition unit signal, and the moving average filter is used to eliminate high-frequency jitter in the motion sensing unit signal. The standardization process normalizes each signal to the [0,1] interval to ensure data consistency. Then, a correlation model between ambient sound intensity and user hearing safety threshold is established using deep learning algorithms. Through transfer learning on a large amount of ambient sound data and user hearing data, the model determines the current ambient noise and outputs a dynamic safe sound pressure threshold curve.
4. The headphone hybrid noise reduction control system based on multimodal sensor fusion according to claim 3, characterized in that: The operation for forming the safe sound pressure threshold curve is as follows: Extract environmental sound parameters from historical data, and extract the user's hearing sound pressure level over sequential time based on the same environmental sound parameters; That is, after extracting the hearing sound pressure value according to the sequential time, if it is determined that the hearing sound pressure value has not changed within the set time interval, then the hearing sound pressure value is extracted, and then the average of multiple hearing sound pressure values is obtained to obtain the safe sound pressure threshold under the current environmental sound parameters. Using the changes in environmental sound parameters from small to large as the parameters of the horizontal axis, and the corresponding safe sound pressure threshold under the environmental sound parameters as the parameters of the vertical axis, the parameter points located on the horizontal and vertical axes are connected by curves to obtain the safe sound pressure threshold curve.
5. The headphone hybrid noise reduction control system based on multimodal sensor fusion according to claim 2, characterized in that: The scene recognition and noise processing module determines the user's operation regarding the wearing status of the headphones as follows: The wearing status is determined by extracting the parameter data collected by the capacitive proximity sensor and the pressure sensor and the factory-set threshold. The capacitive proximity sensor is used to detect the contact distance between the headphones and the user's ear as 'a', and the pressure sensor is set on the edge of the earcup to detect the contact pressure between the headphones and the ear as 'b'. Furthermore, the factory-set distance threshold is A, and the ear contact pressure threshold is B. The specific results are as follows: Result 1: When a < A and b ≥ B, the current wearing state is stable. Result 2: When a≥A and b<B, the current wearing state is loose. Result 3: When a has no data and b=0, the current state is not wearing the device.
6. The headphone hybrid noise reduction control system based on multimodal sensor fusion according to claim 2, characterized in that: The scene recognition and noise processing module performs the following operation to determine the motion state after wearing the device based on posture data: The built-in three-axis gyroscope tracks the three-dimensional motion of the user's head in real time, and determines the direction and speed of the user's head based on the data obtained from the three-axis accelerometer. Taking the initial position of the sensor when the headphones are first worn as the initial point, and extracting the triaxial accelerometer data in the subsequent time period, we can obtain the angular velocity c and acceleration d of the headphone sensor at the initial point. Based on historical data, the angular velocity range [m, n] and acceleration range [p, q] for different motion states were derived. If c < m, d < p, and the duration is t, then the current user is in a static state. , If c > n and d > q, and the duration is t, then the current user is in a walking state. If c > n and d > q, and the duration is t, then the current user is in a running state.
7. The headphone hybrid noise reduction control system based on multimodal sensor fusion according to claim 1, characterized in that: The feature fusion operation of wearing state parameters and motion state parameters in the scene recognition and noise processing module is as follows: That is, by extracting real-time wearing status parameters and motion status parameters and comparing them with the set contact distance threshold and ear contact pressure threshold, it is determined whether there are changes under different motion status parameters. It also generates adjustment instructions and provides feedback in a timely manner based on changes in the situation.
8. The headphone hybrid noise reduction control system based on multimodal sensor fusion according to claim 2, characterized in that: The simultaneous operation of active noise reduction, passive noise reduction, and directional noise reduction based on fused features in the scene recognition and noise processing module is as follows: By extracting the feature parameters of the noise signal after feature fusion, and combining the wearing state parameters and motion state parameters, the weights of active noise reduction and passive noise reduction are dynamically adjusted. Simultaneously, the direction of the sound wave from the sound source is located based on the motion state parameters, and a reverse sound wave is emitted in a specific direction to achieve noise suppression.
9. A multimodal sensor fusion-based headphone hybrid noise reduction control system according to claim 7, characterized in that: The operation of adjusting the weighting values of active noise reduction and passive noise reduction is as follows: If the wearing state is a stable wearing state, the amplitude of the noise signal is extracted and compared with the set signal amplitude; If the amplitude of the extracted noise signal is less than the set signal amplitude, then the active noise reduction weight is set to u1, and the passive noise reduction weight is set to u1. ,and If the amplitude of the extracted noise signal is greater than the set signal amplitude, then the active noise reduction weight is set to [value missing]. Passive noise reduction accounts for a significant portion of the weighting. ,and ; If the wearing condition is loose, then set the active noise cancellation weight to [value]. Passive noise reduction accounts for a significant portion of the weighting. ,and ; If the device is not worn, both active and passive noise cancellation functions will be automatically turned off to reduce power consumption.
10. A multimodal sensor fusion-based headphone hybrid noise reduction control system according to claim 8, characterized in that: The operation of locating the sound source direction based on motion state parameters is as follows: By determining the arrival timestamps of the same noise signal received by each microphone in the four-microphone array, the direction of the sound wave of the sound source can be located according to the order of the timestamps. It also updates the change in the direction of the sound source by combining the change in motion state, so as to generate a noise reduction signal with the opposite direction of the sound wave of the sound source for cancellation.
Citation Information
Patent Citations
Earphone noise reduction method and system based on multivariate noise
CN119893362A