Noise control method and close-to-ear open audio device

By setting up a microphone in a near-ear open-back audio device to identify noise patterns and directions, and selecting a matching noise reduction controller for active noise control, the problem of poor noise control in existing technologies is solved, and more precise noise suppression is achieved.

CN122313939BActive Publication Date: 2026-08-25GOERTEK INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610771532.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-01
Publication Date
2026-08-25
Estimated Expiration
2046-06-01

AI Technical Summary

Technical Problem

Existing active noise control technologies are ineffective in near-ear open-back audio devices and may introduce additional noise, making it difficult to effectively suppress ambient noise.

Method used

By setting at least two microphones in an open-back audio device near the ear, the ambient noise pattern and noise direction are identified, and an appropriate noise reduction controller is selected for active noise control. The ambient noise is then precisely canceled out using anti-phase noise cancellation waves.

Benefits of technology

It achieves more precise noise suppression in complex acoustic environments, improving the noise control performance of near-ear open-back audio devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122313939B_ABST
    Figure CN122313939B_ABST
Patent Text Reader

Abstract

The application discloses a noise control method and an open ear near-field audio device, and relates to the technical field of acoustic processing. The noise control method is applied to the open ear near-field audio device, at least two first microphones are arranged in the open ear near-field audio device, and the noise control method comprises the following steps: acquiring environmental noise signals collected by the first microphones; performing noise mode identification according to the environmental noise signals to obtain a target noise mode; performing noise positioning according to the environmental noise signals to obtain a target noise direction; selecting a target noise reduction controller corresponding to the target noise mode and the target noise direction from a plurality of preset noise reduction controllers; and performing active noise control based on the target noise reduction controller. The application provides a scheme for applying active noise control in the open ear near-field audio device and improving the noise control effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of acoustic processing technology, and in particular to a noise control method and a near-ear open-back audio device. Background Technology

[0002] Near-ear open-back audio devices, as a typical representative of wearable devices, have seen rapid development in the field of open-back acoustic electronics in recent years. When worn, these devices keep the user's ear canal open or semi-open, with a certain distance between the speaker and the ear canal entrance, allowing ambient sound to pass through naturally. Users can perceive the audio content while maintaining auditory awareness of their surroundings. This open-back acoustic structure provides a user experience drastically different from traditional in-ear or closed-back headphones, but it also presents new technical challenges for the application of active noise control technology.

[0003] Existing active noise control technologies are mostly applied to in-ear or closed-back headphones, whose acoustic environment is relatively closed and controllable. However, for near-ear open-back audio devices, due to their open acoustic structure, the coupling between the speaker and the ear canal is low, and the acoustic propagation path is more complex. When traditional noise reduction solutions are directly transplanted and applied, they often face problems such as poor noise reduction effect or even the introduction of additional noise.

[0004] Therefore, providing a noise control method that can effectively suppress environmental noise, taking into account the open acoustic structure characteristics of near-ear open audio devices, remains a technical problem that urgently needs to be solved in this field.

[0005] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention

[0006] The main objective of this application is to provide a noise control method and a near-ear open-back audio device, aiming to provide a solution for applying active noise control in a near-ear open-back audio device and improving the noise control effect.

[0007] To achieve the above objectives, this application proposes a noise control method applied to a near-ear open-back audio device, wherein at least two first microphones are provided in the near-ear open-back audio device, and the noise control method includes: Acquire the ambient noise signals collected by each first microphone; The target noise pattern is obtained by performing noise pattern recognition based on the environmental noise signal. Noise localization is performed based on environmental noise signals to determine the direction of the target noise. Select the target noise reduction controller that corresponds to the target noise mode and the target noise direction from a set of preset noise reduction controllers; Active noise control based on target noise reduction controller.

[0008] Optionally, the step of obtaining the target noise pattern by performing noise pattern recognition based on the environmental noise signal includes: A bandpass filter is used to filter the target environmental noise signal to obtain a filtered signal. The target environmental noise signal is any one of the environmental noise signals or the one with the largest signal energy. The center frequency of the bandpass filter is designed according to the energy spectrum of the preset type of environmental noise. The first signal feature value is obtained by weighting the signal energy of multiple frames of filtered signals, wherein the weighting weight of the signal energy of the later filtered signal is greater than the weighting weight of the signal energy of the earlier filtered signal. The target noise pattern is determined based on the comparison between the first signal feature value and the first preset threshold.

[0009] Optionally, there are two bandpass filters. The center frequencies of the two bandpass filters are designed based on the two frequency bands with the largest energy contrast in the energy spectrum of various preset types of environmental noise. The first preset threshold includes a first threshold, a second threshold, a third threshold, and a fourth threshold, with the third threshold being less than the first threshold. The steps for determining the target noise pattern based on the comparison between the first signal feature value and the preset threshold include: If the difference between the first feature value and the second feature value is greater than the first threshold, then return to the step of calculating the signal feature value of the filtered signal, wherein the first feature value is the first signal feature value of the filtered signal obtained by filtering through a bandpass filter with a larger center frequency, and the second feature value is the first signal feature value of the filtered signal obtained by filtering through a bandpass filter with a smaller center frequency. If the sum of the first feature value and the second feature value is greater than the second threshold, and the difference between the first feature value and the second feature value is greater than the third threshold, then the preset first noise mode is determined to be the target noise mode. If the sum of the first feature value and the second feature value is greater than the second threshold, and the difference between the first feature value and the second feature value is less than or equal to the third threshold, then the preset second noise mode is determined to be the target noise mode. If the sum of the first feature value and the second feature value is less than or equal to the second threshold and greater than the fourth threshold, then the preset third noise mode is determined as the target noise mode. If the sum of the first feature value and the second feature value is less than or equal to the fourth threshold, then the preset fourth noise mode is determined as the target noise mode.

[0010] Optionally, the steps of active noise control based on the target noise reduction controller include: The signal weight of each microphone in the noise-canceling microphone group is determined according to the target noise direction, wherein the signal weight of the microphone closer to the target noise direction is greater than the signal weight of the microphone farther away from the target noise direction. The noise-canceling microphone group includes a reference microphone and / or an error microphone set in a near-ear open audio device, and the first microphone includes all or part of the reference microphone. Active noise control is performed based on the signals collected by each microphone in the noise-canceling microphone group and the target noise-canceling controller. During the active noise control process, the signals collected by the corresponding microphones are weighted based on the signal weights.

[0011] Optionally, the near-ear open-back audio device includes a reference microphone and an error microphone, the first microphone comprising all or part of the reference microphone, and after the step of active noise control based on the target noise reduction controller, it further includes: The signal acquired by one of the reference microphones is filtered by a bandpass filter with a preset frequency band to obtain a first signal, and the signal acquired by one of the error microphones is filtered by a bandpass filter to obtain a second signal. The preset frequency band does not overlap with the noise reduction frequency band of the target noise reduction controller. The gain of the target noise reduction controller is adjusted based on the comparison of the signal energy of the first and second signals.

[0012] Optionally, the step of adjusting the gain of the target noise reduction controller based on the signal energy comparison result of the first signal and the second signal includes: Calculate the absolute value of the difference between the signal energies of the first and second signals; If the absolute value of the difference is greater than or equal to the second preset threshold, the gain of the target noise reduction controller will be attenuated at the first preset rate until the attenuation reaches the first target value. If the state where the absolute value of the difference is less than the second preset threshold continues for a first preset duration, the gain of the target noise reduction controller will be increased at a second preset rate until it is restored to the initial gain of the target noise reduction controller or the absolute value of the difference is greater than or equal to the second preset threshold. If the absolute value of the difference continues to increase to the third preset threshold after the initial attenuation at the first preset speed, the gain of the target noise reduction controller will be attenuated at the third preset speed until the absolute value of the difference is less than the second preset threshold. Then, the gain of the target noise reduction controller will be increased at the fourth preset speed until it returns to the second target value. The second target value is the initial gain minus the first target value, the third preset threshold is greater than the second preset threshold, and the third preset speed is greater than the first preset speed. If, after restoring the gain of the target noise reduction controller to the second target value, the absolute value of the difference is less than the second preset threshold for a second preset duration, then the gain of the target noise reduction controller will be increased at a fifth preset speed until it is restored to the initial gain or the absolute value of the difference is greater than or equal to the second preset threshold.

[0013] Optionally, the step of acquiring the ambient noise signal collected by each first microphone includes: Acquire the target signal, wherein the target signal is the original signal collected by the target microphone or the signal after high-pass filtering of the original signal in this embodiment at a preset frequency. In this embodiment, the target microphone is one of the microphones set in the near-ear open audio device of this embodiment. Calculate the average signal energy of the target signal in the most recent first preset frame to obtain the second signal characteristic value; If the second signal feature value is greater than or equal to the fourth preset threshold, then wearer voice activity detection is performed based on the signal energy of the target signal of this embodiment at the most recent second preset frame number; otherwise, return to the step of obtaining the target signal in this embodiment, wherein the second preset frame number in this embodiment is greater than the first preset frame number in this embodiment. If the wearer's voice activity is detected, the process returns to the step of obtaining the target signal in this embodiment; If no voice activity of the wearer is detected, the original signal collected by the first microphone in each embodiment is subjected to echo cancellation to obtain the environmental noise signal of this embodiment.

[0014] Optionally, the step of detecting wearer voice activity based on the signal energy of the target signal at the most recent second preset frame number includes: For the target signal of the most recent second preset frame number, calculate the upper and lower boundaries of the box line according to the preset box line parameters and the signal energy of the target signal of each frame; The proportion of outliers in the signal energy of the target signal in each frame is determined based on the upper and lower boundaries. If the proportion of outliers of a preset number of consecutively calculated outliers is greater than the fifth preset threshold, then it is determined that the wearer's voice activity has been detected.

[0015] In addition, to achieve the above objectives, this application also proposes a near-ear open-back audio device, which is provided with at least two first microphones for collecting ambient noise signals. The near-ear open-back audio device further includes a memory, a processor, and a computer program stored in the memory and executable on the processor. The computer program is configured to implement the steps of the noise control method described above.

[0016] Optionally, the near-ear open-back audio device is smart glasses, wherein at least one temple of the smart glasses has at least one speaker and at least one error microphone on its ear hook portion; at least one first microphone is located on the temple side further away from the temple root than on the ear hook portion; and at least one first microphone is located on the temple side closer to the temple root than on the ear hook portion, or, at least one first microphone is located on the frame of the smart glasses. In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the noise control method described above.

[0017] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the noise control method described above.

[0018] One or more technical solutions proposed in this application have at least the following technical effects: This application acquires environmental noise signals collected by each first microphone, identifies the target noise pattern from the environmental noise signals, and determines the main characteristics of the current environmental noise, providing a basis for selecting a noise reduction strategy. Furthermore, by determining the direction of the target noise from the environmental noise signals, it obtains the spatial orientation information of the main noise source relative to the audio device, supplementing the selection of the noise reduction strategy with spatial dimension reference. After obtaining information on both the noise pattern and noise direction, this application selects a noise reduction controller that matches the current acoustic scene. This allows subsequent active noise control to use control parameters pre-optimized for this specific scene, avoiding the performance degradation caused by using a fixed set of parameters to cope with ever-changing acoustic environments. Finally, the selected noise reduction controller actually performs active noise control, precisely applying the anti-phase noise cancellation wave to the near-ear region to cancel out the environmental noise. Ultimately, this achieves more precise noise suppression in the complex acoustic environment of near-ear open-back audio devices, thereby improving the noise control effect of near-ear open-back audio devices. Attached Figure Description

[0019] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0020] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a flowchart illustrating the first embodiment of the noise control method of this application; Figure 2 This is a flowchart illustrating the second embodiment of the noise control method of this application. Figure 3 A waveform diagram of real speech provided for the fourth embodiment of the noise control method of this application; Figure 4 The fourth embodiment of the noise control method provided in this application will Figure 3 Waveforms of real speech mixed with different types of noise; Figure 5 This is an interference detection result diagram provided in the fourth embodiment of the noise control method of this application; Figure 6 An exemplary spectrogram of the control effect provided by the third embodiment of the noise control method of this application, wherein (a) is an exemplary spectrogram of the control effect without adopting the noise control scheme of this application, and (b) is an exemplary spectrogram of the control effect with adopting the noise control scheme of this application; Figure 7 This is a schematic diagram of a near-ear open-back audio device provided in the fifth embodiment of the noise control method of this application; Figure 8 This is a schematic diagram of a smart glasses device provided in the fifth embodiment of the noise control method of this application.

[0022] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0023] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0024] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0025] Near-ear open-back audio devices, as a typical representative of wearable devices, have seen rapid development in the field of open-back acoustic electronics in recent years. When worn, these devices leave the user's ear canal open or semi-open, allowing ambient sound to pass through naturally. However, precisely because of the open-back design, external ambient noise can easily enter the ear canal, interfering with the clarity of the audio content being played. Especially in noisy environments, users often need to increase the volume to suppress ambient noise, which not only affects auditory comfort but may also pose potential risks to hearing health. Existing active noise control technologies are mostly applied to closed-back or in-ear headphones, where the acoustic environment is relatively controllable. However, for near-ear open-back audio devices, due to the complex acoustic leakage paths and low coupling between the speaker and the ear canal, directly applying traditional noise reduction solutions often results in poor noise reduction, decreased sound field fidelity, or even the introduction of additional noise. Therefore, how to provide an effective noise control method to suppress ambient noise, tailored to the unique acoustic structure of near-ear open-back audio devices, remains a pressing issue.

[0026] This application provides a solution that identifies noise patterns and locates the direction of ambient noise, selects a corresponding noise reduction controller based on the combination of noise pattern and direction, and performs active noise control to achieve noise suppression that matches the current acoustic environment in near-ear open-back audio devices. It should be noted that the execution subject of the noise control method in each embodiment of this application is a near-ear open-back audio device (hereinafter also referred to as "audio device"). It is understood that there are many ways to implement the hardware architecture and software system of audio devices, and this application does not limit the specific hardware architecture and software system implementation of the audio device to which the noise control method is applied in each embodiment.

[0027] The following presents a first embodiment of the noise control method of this application. (Refer to...) Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the noise control method of this application. In this embodiment, the audio device is equipped with at least two first microphones. These first microphones are used to collect ambient noise signals, and their specific placement and number can be flexibly adjusted according to the device configuration; no limitation is imposed in this embodiment. In this embodiment, the noise control method includes steps S10 to S30: Step S10: Acquire the ambient noise signal collected by each first microphone.

[0028] It should be noted that in practical applications of audio devices, the audio device may have only one channel or two channels. In the case of two channels, the two channels may correspond to the user's left ear channel and right ear channel, respectively. That is, in the case of two channels, the audio device may have at least two first microphones for the user's left ear and right ear, respectively. In addition, the audio device may also have a noise-canceling microphone group for the user's left ear and right ear, respectively. The specific composition of the noise-canceling microphone group is related to the architecture adopted by the active noise control. The first microphones and the microphones in the noise-canceling microphone group may overlap or be independent of each other, which is not limited in this embodiment. In this binaural wearing scenario, the noise control method described in this application embodiment can be executed separately based on the hardware configuration corresponding to the left ear channel or separately based on the hardware configuration corresponding to the right ear channel.

[0029] To simplify the description, the technical solutions of the various embodiments of this application will be described below using only the hardware configuration corresponding to a single channel as an example. However, those skilled in the art should understand that the description below can be applied equally and unambiguously to the case where the other channel or both channels are used simultaneously.

[0030] It should be noted that, in this embodiment, the first microphone refers to a microphone used to collect ambient noise signals, and the collected ambient noise signals are used for noise pattern recognition and noise localization. It should be understood that the first microphone is not equivalent to the reference microphone or error microphone that may be involved in subsequent active noise control. Specifically, the first microphone may include all reference microphones, or only some reference microphones, or it may include microphones used for other functions, as long as the ambient noise signals collected by the first microphone are used for noise pattern recognition and noise localization.

[0031] In one feasible embodiment, each first microphone is disposed on the main body of the audio device and is located further away from the ear canal entrance than the speaker's sound outlet. Specifically, when the audio device is worn, the acoustic environment of the first microphone is less directly affected by the sound played by the speaker and is mainly exposed to the ambient noise field of the external space. For example, the first microphone can be disposed on the outer surface, front edge, rear edge, or top / bottom of the main body of the audio device, as long as the ambient noise component in the sound signal it picks up is significantly higher than the speaker feedback component.

[0032] It should be noted that active noise control typically includes three basic architectures: feedforward, feedback, and hybrid. The embodiments of this application do not limit the specific active noise control architecture used in the audio device. Under different architectures, the microphones in the noise-canceling microphone group play different functional roles. For example, in a feedforward architecture, the reference microphone is used to collect ambient noise as a feedforward reference signal, while in a feedback architecture, the error microphone is positioned near the ear to collect the residual noise signal after noise reduction processing as a feedback signal. The first microphone and the microphones in the noise-canceling microphone group are functionally independent, but physically they can be multiplexed or separated, and can be flexibly configured according to actual product requirements.

[0033] Furthermore, it should be noted that the environmental noise signal collected and provided by the first microphone can be the raw audio signal directly output by the first microphone, or it can be a signal obtained after the raw audio signal has undergone at least one preprocessing operation. Preprocessing operations include, for example, echo cancellation, gain adjustment, equalization filtering, dynamic range compression, or high-pass filtering. For instance, in a specific embodiment, considering that near-ear open-back audio devices may be affected by wind noise interference or circuit background noise during use, a high-pass filter can be performed on the raw audio signal directly output by the first microphone before active noise control to filter out low-frequency wind noise components or DC offset components, thereby improving the robustness of subsequent noise reduction algorithms.

[0034] Step S20: Based on the environmental noise signal, noise pattern recognition is performed to obtain the target noise pattern.

[0035] After acquiring the ambient noise signal, noise pattern recognition is performed to determine the target noise pattern corresponding to the current acoustic environment. It should be noted that a noise pattern refers to a type identifier obtained by classifying the acoustic characteristics of the ambient noise signal.

[0036] In one feasible implementation, multiple candidate noise modes can be predefined, each corresponding to different acoustic characteristics. Specifically, the characteristic changes of environmental noise signals within a time window can be statistically analyzed. Based on the comparison between the statistical results and preset judgment conditions, it can be determined which candidate noise mode the current environmental noise belongs to.

[0037] In specific implementations, the number of candidate noise patterns can be flexibly set according to actual needs, such as two, three, four, or more. Each candidate noise pattern corresponds to a set of preset judgment conditions. When the statistical characteristics of the environmental noise signal meet a certain set of judgment conditions, the candidate noise pattern is determined to be the target noise pattern. It is understood that in actual implementation, more or fewer noise patterns can be set according to the computing power of the device or the application scenario requirements. For example, if it is necessary to reduce the algorithm complexity, only two or three candidate noise patterns can be distinguished, and the number of judgment conditions can be reduced accordingly. This embodiment does not limit the specific number of candidate noise patterns and their corresponding judgment conditions.

[0038] Step S30: Based on the environmental noise signal, noise localization is performed to obtain the target noise direction.

[0039] After acquiring the ambient noise signal, noise localization is performed on the ambient noise signal to determine the directional information of the main noise sources in the current acoustic environment relative to the audio equipment.

[0040] In one feasible implementation, due to differences in the distance and relative angle between each first microphone and the noise source, the sound emitted by the same noise source will exhibit spatially related differences in signal amplitude, phase, or arrival time when it reaches different first microphones. Based on these differences, cross-correlation analysis or phase comparison can be performed on the environmental noise signals collected by each first microphone to estimate the directional information of the main noise source relative to the audio device, i.e., the target noise direction.

[0041] Furthermore, in one feasible embodiment, ambient noise signals collected by two first microphones on an audio device can be used for noise localization. For ease of description, the two first microphones are referred to as first microphone A and first microphone B, respectively. Specifically, the signals collected by first microphone A and first microphone B within the same time period are processed in frames, and the duration of each frame can be set according to actual needs. A temporal correlation operation is performed on the two signals within a frame to calculate the sound path difference between them. To improve the stability of the estimation, the above correlation operation can be repeated multiple times using an overlapping sliding frame window. Specifically, the frame window is slid backward sequentially with a preset overlap rate, and after each slide, a temporal correlation operation is performed again on the two signals in the current frame to obtain a new sound path difference. This is repeated multiple times to obtain a set of sound path difference data. Then, statistical processing is performed on this set of sound path difference data, such as calculating the average value of the set of sound path differences, and the incident direction of the main noise source is determined based on the statistical results. For example, if the average sound path difference is close to zero, it indicates that the noise source is roughly located in the direction of the perpendicular plane of the line connecting the two first microphones; if the average value is positive and the absolute value is large, it indicates that the noise source is biased towards the side where the first microphone A is located; if the average value is negative and the absolute value is large, it indicates that the noise source is biased towards the side where the first microphone B is located. The sign of the sound path difference can be pre-defined according to the relative position of the two microphones.

[0042] It should be noted that the above method of noise localization using two first microphones is merely an illustrative example. In actual implementation, an appropriate localization algorithm can be selected based on the specific number and layout of the first microphones on the audio device. This embodiment does not limit the specific algorithm used for noise localization.

[0043] Step S40: Select a target noise reduction controller that corresponds to the target noise mode and the target noise direction from a plurality of preset noise reduction controllers.

[0044] It should be noted that after determining the current target noise mode and target noise direction, a matching noise reduction controller needs to be selected for targeted active noise control. Ambient noise varies under different noise modes, and the sound path and attenuation characteristics of ambient noise reaching each microphone also differ under different noise directions. Therefore, pre-setting corresponding noise reduction controllers for different combinations of noise modes and noise directions helps to more accurately match the current acoustic environment during active noise control, improving the noise reduction effect. It should be noted that in this embodiment, the noise reduction controller refers to the noise control algorithm used to perform active noise control. The difference between different noise reduction controllers can be reflected in different algorithm types or different parameters under the same algorithm type. For example, for a certain combination of noise mode and noise direction, the corresponding noise reduction controller can use a first type of algorithm and configure a first set of parameters; for another combination, the corresponding noise reduction controller can use the same algorithm type but configure a second set of parameters, or it can use a different second type of algorithm. In actual implementation, different noise reduction controllers only need to switch different algorithm programs or parameter configurations when called. Furthermore, this embodiment does not limit which noise control algorithm the noise reduction controller can specifically use; it can be understood that any noise control algorithm capable of achieving active noise control is acceptable.

[0045] In one feasible implementation, M candidate noise patterns and N candidate noise directions can be predefined. The number M of candidate noise patterns can be set according to actual needs. The number N of candidate noise directions can also be set according to actual needs, for example, dividing the space into N directional intervals based on the placement of the first microphone on the audio device. The M noise patterns and N noise directions can form M×N combinations. For each combination, a noise reduction controller can be pre-set, for a total of M×N noise reduction controllers. Further, in one feasible implementation, each noise reduction controller can include a set of preset filter coefficients. Once the target noise pattern and target noise direction are determined, the set of filter coefficients corresponding to the combination of the target noise pattern and target noise direction can be selected from the M×N sets of filter coefficients to generate an anti-phase noise cancellation wave.

[0046] In one feasible implementation, the parameters of the noise reduction controller can be obtained through pre-tuning. For each combination of candidate noise mode and candidate noise direction, ambient noise of that candidate noise mode can be played through a speaker in another device in that candidate noise direction. The audio device performs active noise control through the configured noise reduction controller to be tuned, and determines a set of parameters for the noise reduction controller using a preset optimization algorithm or through manual tuning, so that the noise reduction performance in that candidate noise direction and candidate noise mode meets preset requirements, such as the residual noise energy of the error signal acquisition being lower than a preset threshold. The noise reduction controller containing this set of parameters is then associated with the combination of candidate noise direction and candidate noise mode in the audio device.

[0047] Step S50: Perform active noise control based on the target noise reduction controller.

[0048] After selecting a noise reduction controller that corresponds to the target noise mode and the target noise direction, the noise reduction controller is used to perform active noise control processing on the current ambient noise.

[0049] It should be noted that active noise control, also known as active noise cancellation, works by generating an anti-phase noise cancellation wave that is out of phase with the ambient noise. This wave cancels out the original ambient noise within the spatial superposition region, thereby reducing the noise level perceived by the human ear. In engineering implementation, active noise control typically requires using a reference microphone to collect ambient noise as a reference signal and / or using an error microphone to collect near-ear residual noise as a feedback signal.

[0050] In one feasible implementation, an ambient noise signal collected by a microphone used as a reference signal in a first microphone is taken as input, and a corresponding anti-phase noise cancellation signal is generated based on a target noise reduction controller. This anti-phase noise cancellation signal is fed to the speaker of an audio device for playback. The anti-phase noise cancellation signal played by the speaker acoustically superimposes with the ambient noise in the near-ear region, thereby reducing the ambient noise level near the user's ear canal.

[0051] In one feasible implementation, it should be noted that during practical applications, the acoustic environment may change, causing a change in the target noise pattern or direction. In such cases, when switching from the currently used noise reduction controller to a newly selected one, a smooth transition can be employed. Specifically, during the switching process, the transition of the noise reduction control function is smoothed by adjusting the gain parameters of the current and new noise reduction controllers. In one feasible implementation, the smooth transition can be achieved by gradually increasing and decreasing the gain of the filters within the noise reduction controller. That is, the gain of the current noise reduction controller decreases to 0 at a preset rate, while the gain of the new noise reduction controller increases to a preset value at a preset rate, thus avoiding instantaneous anomalies in the output signal caused by sudden changes in the noise reduction controller. In one feasible implementation, the switching duration (which can be converted to the switching speed) can be set according to actual needs. For example, the switching duration can be set to no more than 30 milliseconds to achieve a balance between controlling the smoothness of the switching and the response speed.

[0052] This embodiment acquires environmental noise signals collected by each first microphone, identifies the target noise pattern from the environmental noise signals, and determines the main characteristics of the current environmental noise, providing a basis for selecting a noise reduction strategy. Simultaneously, it determines the direction of the target noise from the environmental noise signals, obtaining the spatial orientation information of the main noise source relative to the audio device, supplementing the selection of the noise reduction strategy with spatial dimension reference. After obtaining information on both the noise pattern and noise direction, this embodiment selects a noise reduction controller that matches the current acoustic scene, enabling subsequent active noise control to use pre-optimized control parameters for this specific scene, avoiding the performance degradation caused by using a fixed set of parameters to cope with ever-changing acoustic environments. Finally, the selected noise reduction controller actually performs active noise control, precisely applying the anti-phase noise cancellation wave to the near-ear region to cancel out the environmental noise. Ultimately, this achieves more precise noise suppression in the complex acoustic environment of near-ear open-back audio devices, thereby improving the noise control effect of near-ear open-back audio devices.

[0053] In one feasible implementation, during the active noise control process, the signals collected by each microphone in the noise-canceling microphone group can be weighted based on the target noise direction determined in step S30. Microphones at different locations have different pickup sensitivities to environmental noise from different directions. Microphones closer to the target noise direction collect signals containing stronger noise source components and more accurate spatial information. By assigning higher signal weights to microphones closer to the noise direction, the reference signal and / or feedback signal input to the noise-canceling controller can more accurately reflect the characteristics of the main noise source, thereby further improving the targeting and noise reduction effect of the active noise control. Based on the above principle, step S50 can further include the following steps S501~S502: Step S501: Determine the signal weight of each microphone in the noise-canceling microphone group according to the target noise direction, wherein the signal weight of the microphone closer to the target noise direction is greater than the signal weight of the microphone farther away from the target noise direction. The noise-canceling microphone group includes a reference microphone and / or an error microphone set in a near-ear open audio device, and the first microphone includes all or part of the reference microphone.

[0054] Step S502: Active noise control is performed based on the signals collected by each microphone in the noise reduction microphone group and the target noise reduction controller. During the active noise control process, the signals collected by the corresponding microphones are weighted based on the signal weights.

[0055] The signal weights of each microphone in the noise-canceling microphone group are determined based on the direction of the target noise. It should be noted that the noise-canceling microphone group includes a reference microphone and / or an error microphone positioned in a near-ear open-back audio device, wherein the first microphone includes all or part of the reference microphone. Microphones closer to the direction of the target noise are assigned larger signal weights, while microphones farther from the direction of the target noise are assigned smaller signal weights.

[0056] In one alternative implementation, the determination of whether a direction is close can be made by pre-establishing a spatial coordinate system, mapping the noise direction and the positions of each microphone in the noise-canceling microphone group to this coordinate system, and representing each direction with angles. By calculating the difference between the target noise direction angle and the angles of each microphone, the smaller the absolute value of the difference, the closer the microphone is to the target noise direction, and correspondingly, a higher signal weight is assigned to it.

[0057] Furthermore, as a feasible implementation, an example of using three microphones on an audio device will be provided. Assume the first microphone includes a first microphone A positioned on the side of the device closest to the wearer's facing direction, a second microphone B positioned on the side of the device closest to the wearer's back facing direction, and a third microphone C positioned between the two. For ease of description, the orientation of the user when normally wearing an open-back audio device is used as a reference, defining the direction the user's face faces as forward and the opposite direction as backward. When the target noise direction is determined to be close to the front, the signal weight of the first microphone A is amplified; when the target noise direction is determined to be close to the rear, the signal weight of the second microphone B is amplified; when the target noise direction is determined to be approximately perpendicular to the line connecting the first microphone A and the second microphone B, the signal weight of the third microphone C is amplified.

[0058] After determining the signal weights of each microphone, directional weighted active noise control can be performed based on these weights and the selected target noise reduction controller. The signals collected by each microphone in the noise-reducing microphone group are weighted according to their corresponding signal weights. The weighted signal is then input to the selected target noise reduction controller to generate an inverse-phase noise cancellation signal, which is then fed to the speaker for playback. Through weighting, sound components from the target noise direction are given greater importance during the control process, thereby improving the suppression depth of noise in that direction.

[0059] It is understood that the above methods of determining directional proximity based on angle differences and illustrating weight allocation using three microphones are merely illustrative. In actual implementation, these methods can be flexibly adjusted according to specific needs. This embodiment does not limit the specific values ​​of each parameter or the implementation details.

[0060] Based on the first embodiment described above, a second embodiment of the noise control method of this application is proposed. In this embodiment, content that is the same as or similar to that in the first embodiment can be referred to the above description, and will not be repeated hereafter. In this embodiment, refer to... Figure 2 Step S20 includes S201~S203: Step S201: The target environmental noise signal is filtered by a bandpass filter to obtain a filtered signal. The target environmental noise signal is any one of the environmental noise signals or the one with the largest signal energy. The center frequency of the bandpass filter is designed according to the energy spectrum of the preset type of environmental noise.

[0061] It should be noted that before bandpass filtering the environmental noise signal, a target environmental noise signal needs to be identified as the filter input. This target environmental noise signal can be any one of the various environmental noise signals or the one with the highest signal energy. After identifying the target environmental noise signal, bandpass filtering is performed on it. The center frequency of the bandpass filter can be designed based on the energy spectrum of the preset type of environmental noise.

[0062] In one feasible implementation, taking an audio device equipped with two first microphones as an example, the two first microphones are denoted as first microphone A and first microphone B, respectively. Ambient noise signals collected by first microphone A and first microphone B within the same time period are acquired, and the root mean square (RMS) values ​​of the two signals within a preset time window are calculated. The RMS values ​​of the two signals are compared, and the signal with the larger RMS value is determined as the target ambient noise signal. It is understood that in actual implementation, the settings of relevant parameters can be adjusted according to the actual requirements of the device, and this embodiment does not impose any restrictions on this.

[0063] In the specific design of bandpass filters, one or more bandpass filters can be set according to actual needs. When using a single bandpass filter, its center frequency can be determined based on the peak position of the energy spectrum of a preset type of ambient noise. The peak position refers to the frequency point with the highest energy value in the energy spectrum of that preset noise type. When using two or more bandpass filters, the passband of each filter can be designed separately based on the energy spectra of two or more preset types of ambient noise.

[0064] In one feasible implementation, if two bandpass filters are used, the energy spectrum of all preset types of environmental noise to be distinguished can be analyzed to find one or more frequency bands where the energy difference between different preset types of environmental noise is most significant. That is, the center frequencies of the two bandpass filters are designed based on the two frequency bands with the greatest energy contrast in the energy spectrum of various preset types of environmental noise.

[0065] As an example, let's assume two ambient noise types: a first type and a second type. Energy spectrum analysis reveals that the energy peak of the first type is located around 200Hz, while the energy peak of the second type is around 80Hz. However, at 300Hz, the energy difference between the first and second types is significantly greater than the energy difference at their respective peaks. Simultaneously, at another frequency of 20Hz, the energy difference between the first and second types is less significant than at 300Hz. Therefore, if two bandpass filters are used, the center frequency of one bandpass filter can be set at 300Hz, and the center frequency of the other at 20Hz. It is understood that the above frequency values ​​and relationships are merely examples; in practical applications, the center frequency should be determined based on the measured energy spectrum of the specific ambient noise type.

[0066] Step S202: The signal energy of the multi-frame filtered signals is weighted and calculated to obtain the first signal feature value, wherein the weighting weight of the signal energy of the later filtered signal is greater than the weighting weight of the signal energy of the earlier filtered signal.

[0067] After obtaining the filtered signal, it is processed into frames. The signal energy of multiple frames of filtered signals is weighted and calculated to obtain the first signal feature value. In the weighting calculation, the weight corresponding to the signal energy of each frame is related to the time order of the frames; the later the frame is in time, the greater the weight, and the earlier the frame is in time, the smaller the weight. This setting allows the first signal feature value to more sensitively reflect the recent trend of environmental noise changes. It should be noted that signal energy refers to a parameter reflecting the strength of the signal within a preset time window. The preset time window refers to the signal duration used to define a single signal energy calculation. Each frame refers to a segment of signal data extracted from the filtered signal according to the preset time window; that is, each frame corresponds to the signal duration of a preset time window. By calculating the signal energy of each frame separately, the signal energy corresponding to each frame can be obtained. In a specific implementation, the signal energy can be obtained by calculating the sum of squares and / or the mean square of each sample value of the filtered signal within the preset time window, or other calculation methods that can characterize signal strength can be used; this embodiment does not limit this.

[0068] In one feasible implementation, the filtered signal is divided into frames with a preset frame length, and the frame window is updated with a preset overlap rate. The root mean square value of the filtered signal within each frame is calculated as the signal energy corresponding to that frame. It should be understood that the specific values ​​of the frame length and overlap rate can be adjusted according to actual needs, and other statistical quantities can also be used to calculate the signal energy; this embodiment does not impose any restrictions on this.

[0069] Further, in a feasible implementation, the signal energy calculated for the current frame can be recorded in a first variable. For ease of description, the first variable is denoted as MB1. After obtaining MB1, the first cumulative variable used to characterize the signal energy value of the current frame is updated. This first cumulative variable can be denoted as B1. Each time the signal energy value MB1 of the current frame is calculated, B1 is updated. In a specific implementation, the update method can be expressed as B1 = (B1 + MB1) / 2. The purpose of this is to give a larger weight to the energy in subsequent weightings, resulting in a faster response to noise changes. Since the instantaneous signal energy value MB1 of the current frame occupies half the weight in each update, and the signal energy from earlier moments contained in the historical cumulative value B1 is gradually diluted with multiple updates, the signal energy of the current frame has a larger effective weight in B1. It should be noted that the above update formula is only an example. In actual implementation, other weighted update methods can be used, as long as the weight of the signal energy in later moments is greater than the weight of the signal energy in earlier moments. When the above calculation and update operations are performed a preset number of times, it is considered that noise feature information for a sufficient duration has been accumulated.

[0070] Furthermore, in one feasible implementation, a counter TB can be defined. Each time B1 is updated, the counter TB is incremented by one. When the counter TB reaches a preset threshold TB1, it is considered that sufficient noise feature information has been accumulated, and B1 becomes the first signal feature value used for subsequent judgment. The preset threshold TB1 can be set according to actual needs, for example, as an integer between 1 and 3. This embodiment does not limit the specific value of the preset threshold. Before the counter TB reaches the preset threshold TB1, the above operation can continue until the count value meets the condition.

[0071] It is understood that the definitions, values, and formulas of the above parameters are merely illustrative examples. In actual implementation, they can all be adjusted according to equipment requirements, and this embodiment does not impose any restrictions on them.

[0072] Furthermore, in one feasible embodiment, if multiple bandpass filters are used to process the target environmental noise signal, for example, bandpass filter A and bandpass filter B are used simultaneously, two filtered signals can be obtained. For each filtered signal, its corresponding first signal feature value can be obtained according to the above operation method. In subsequent steps, threshold comparisons can be performed based on the first signal feature values ​​corresponding to each filtered signal, and the target noise mode can be determined by combining the comparison results. It should be noted that the number of bandpass filters can be flexibly selected according to actual needs, and this embodiment does not limit this.

[0073] Step S203: Determine the target noise mode based on the comparison result between the first signal feature value and the first preset threshold.

[0074] It should be noted that the first preset threshold is a pre-defined numerical boundary used to divide the range of values ​​for the first signal feature value. Different numerical ranges correspond to different candidate noise patterns. By comparing the first signal feature value with the first preset threshold, the numerical range in which the first signal feature value falls can be determined, thereby identifying the candidate noise pattern corresponding to that numerical range as the target noise pattern. There can be one or more first preset thresholds. Setting multiple first preset thresholds allows for the division of more numerical ranges, corresponding to more refined noise pattern classifications. After obtaining the first signal feature value, the first signal feature value is compared with the preset first preset threshold. Based on the comparison result, the target noise pattern corresponding to the current acoustic environment is determined.

[0075] In one feasible implementation, if a bandpass filter is used, i.e., if there is only one filtered signal, a first signal characteristic value is generated accordingly, and only a first preset threshold can be set. The first signal characteristic value is compared with the first preset threshold. If the first signal characteristic value is greater than the threshold, the target noise mode is a first candidate noise mode; if the first signal characteristic value is not greater than the threshold, the target noise mode is a second candidate noise mode. It is understood that the specific value, number, and correspondence between the first preset threshold and the candidate noise modes can be set according to actual needs, and this embodiment does not limit this.

[0076] In one feasible embodiment, two bandpass filters can be used. The center frequencies of the two bandpass filters can be designed based on the two frequency bands with the largest energy contrast in the energy spectrum of various preset types of environmental noise. The first preset threshold can include a first threshold, a second threshold, a third threshold, and a fourth threshold, wherein the third threshold is smaller than the first threshold, in order to distinguish four different noise modes. Accordingly, step S203 includes the following steps S2031 to S2035: Step S2031: If the difference between the first feature value and the second feature value is greater than the first threshold, then return to the step of calculating the signal feature value of the filtered signal, wherein the first feature value is the first signal feature value of the filtered signal obtained by filtering through a bandpass filter with a larger center frequency, and the second feature value is the first signal feature value of the filtered signal obtained by filtering through a bandpass filter with a smaller center frequency.

[0077] It should be noted that, for ease of distinction, the first signal characteristic value of the filtered signal obtained by filtering through a bandpass filter with a higher center frequency is called the first characteristic value, and the first signal characteristic value of the filtered signal obtained by filtering through a bandpass filter with a lower center frequency is called the second characteristic value. Calculating the difference between the first and second characteristic values ​​specifically involves subtracting the second characteristic value from the first characteristic value.

[0078] If the difference is greater than the first threshold, it indicates that the energy of the low-frequency components in the current environmental noise is abnormally prominent, and there may be low-frequency interference such as walking vibration or vehicle bumps. At this time, return to step S202, that is, initialize the accumulated first signal characteristic value related variables, and re-perform the statistical and weighted calculation of signal energy to eliminate the influence of low-frequency interference.

[0079] If the difference is less than or equal to the first threshold, it indicates that the low-frequency interference is not significant. Further, the target noise pattern is determined by comparing the sum of the first and second feature values ​​with the preset second and fourth thresholds, and the difference with the preset third threshold.

[0080] Step S2032: If the sum of the first feature value and the second feature value is greater than the second threshold, and the difference between the first feature value and the second feature value is greater than the third threshold, then the preset first noise mode is determined to be the target noise mode.

[0081] If the sum is greater than the second threshold and the difference is greater than the third threshold, it indicates that the overall energy is high and the higher frequency band energy is relatively prominent. In this case, the preset first noise mode is determined as the target noise mode.

[0082] Step S2033: If the sum of the first feature value and the second feature value is greater than the second threshold, and the difference between the first feature value and the second feature value is less than or equal to the third threshold, then the preset second noise mode is determined to be the target noise mode.

[0083] If the sum is greater than the second threshold and the difference is less than or equal to the third threshold, it indicates that the overall energy is high but the energy contrast between frequency bands is not prominent. In this case, the preset second noise mode is determined as the target noise mode.

[0084] Step S2034: If the sum of the first feature value and the second feature value is less than or equal to the second threshold and greater than the fourth threshold, then the preset third noise mode is determined as the target noise mode.

[0085] If the sum is less than or equal to the second threshold and greater than the fourth threshold, it indicates that the overall energy of the current environmental noise is at a medium level, and the preset third noise mode is determined as the target noise mode.

[0086] Step S2035: If the sum of the first feature value and the second feature value is less than or equal to the fourth threshold, then the preset fourth noise mode is determined as the target noise mode.

[0087] If the sum is less than or equal to the second threshold and the sum is less than or equal to the fourth threshold, it indicates that the overall energy of the current environmental noise is low, and the preset fourth noise mode is determined as the target noise mode.

[0088] It is understood that the above threshold numbers, specific values, and correspondence with noise patterns are all illustrative examples and can be adjusted according to specific application scenarios and noise classification requirements in actual implementation.

[0089] In this embodiment, by using the combination of the sum and difference of the first feature value and the second feature value, the noise pattern can be accurately identified from two dimensions: energy contrast and overall energy.

[0090] Based on the first and / or second embodiments described above, a third embodiment of the noise control method of this application is proposed. In this embodiment, content that is the same as or similar to the first and second embodiments described above can be referred to the above description and will not be repeated hereafter. In this embodiment, the noise control method further includes steps A10 to A20: Step A10: The signal acquired by one of the reference microphones is filtered using a bandpass filter with a preset frequency band to obtain a first signal, and the signal acquired by one of the error microphones is filtered using a bandpass filter to obtain a second signal, wherein the preset frequency band does not overlap with the noise reduction frequency band of the target noise reduction controller.

[0091] It should be noted that the audio device can be equipped with a reference microphone and an error microphone. The first microphone includes all or part of the reference microphone. When performing this step, only one reference microphone and one error microphone can be used.

[0092] In one feasible implementation, the preset frequency band can be selected as a mid-to-high frequency band outside the noise reduction band. When the noise reduction controller is working normally, noise within the noise reduction band is suppressed, while noise outside the noise reduction band should neither be reduced nor increased. If controller instability is caused by mismatched controller parameters or sudden changes in the acoustic environment, abnormal increases in noise energy often occur outside the noise reduction band, which are usually more pronounced in the mid-to-high frequency band. Therefore, by monitoring signal energy changes within a specific frequency band outside the noise reduction band, signs of controller instability can be detected in a timely manner.

[0093] Furthermore, as a feasible implementation, the signal collected by a reference microphone closer to the target noise direction can be selected as the source of the first signal. Optionally, the orientation of the user when normally wearing an open-back audio device is used as a reference, with the direction the user's face faces defined as forward and the opposite direction defined as backward. Multiple reference microphones can be provided on the audio device. For ease of distinction, the reference microphone located closer to the wearer's face is designated as the first reference microphone, and the reference microphone located closer to the wearer's back is designated as the second reference microphone. When the target noise direction is determined to be closer to the front, the signal collected by the first reference microphone can be selected as the source of the first signal; when the target noise direction is determined to be closer to the rear, the signal collected by the second reference microphone can be selected as the source of the first signal. The signal collected by the error microphone is selected as the source of the second signal. Both signals are filtered using the same bandpass filter, and the passband range of the bandpass filter is the aforementioned preset frequency band.

[0094] Step A20: Adjust the gain of the target noise reduction controller based on the signal energy comparison results of the first signal and the second signal.

[0095] Calculate the signal energy of the first signal within a preset time period, and the signal energy of the second signal within the same preset time period. Compare the energies of the two signals, and adjust the gain of the target noise reduction controller accordingly based on the comparison results.

[0096] In this embodiment, the stability of the controller is indirectly monitored by comparing the signal energy outside the noise reduction band, and the controller gain is adjusted at different speeds according to the degree of abnormality. This effectively avoids the noise rise problem caused by controller mismatch or sudden changes in the acoustic environment without sacrificing the normal noise reduction effect, thereby improving the robustness of the overall active noise control.

[0097] In a specific implementation, step A20 can adopt a multi-level threshold differential gain speed control mechanism, and the number of threshold levels and parameters can be expanded in practical applications.

[0098] For example, in one feasible implementation, a two-level threshold control mechanism can be employed. Specifically, step A20 includes the following steps A201 to A205: Step A201: Calculate the absolute value of the difference between the signal energy of the first signal and the second signal.

[0099] After acquiring the first and second signals, the signal energy values ​​of the two signals are calculated for the same preset time period. The absolute value of the difference is obtained by subtracting the energy values ​​of the two signals. This absolute value of the difference is used to measure the consistency of the energy of the two signals outside the noise reduction frequency band, and thus serves as a basis for judging whether there is an instability risk in the noise reduction system.

[0100] Step A202: If the absolute value of the difference is greater than or equal to the second preset threshold, the gain of the target noise reduction controller is attenuated at a first preset rate until the attenuation reaches the first target value.

[0101] It should be noted that the second preset threshold is a critical value used to determine the instability risk of the noise reduction system. When the absolute value of the difference is greater than or equal to the second preset threshold, it indicates that there is a significant noise increase outside the noise reduction frequency band, and the noise reduction controller has an instability risk. At this time, the gain is rapidly attenuated at a first preset rate. The first preset rate refers to the rate of gain attenuation, which can be configured to a relatively fast rate to quickly reduce the controller's intensity when abnormal symptoms appear. The attenuation process continues until the cumulative attenuation of the gain reaches the first target value and remains unchanged.

[0102] Step A203: If the state where the absolute value of the difference is less than the second preset threshold continues for a first preset duration, the gain of the target noise reduction controller will be increased at a second preset rate until it is restored to the initial gain of the target noise reduction controller or the absolute value of the difference is greater than or equal to the second preset threshold.

[0103] After gain decay, if the absolute value of the difference falls below the second preset threshold and this state continues for the first preset duration, it indicates that the unstable factors have been eliminated or weakened. The first preset duration is an observation period set to avoid triggering recovery due to instantaneous fluctuations. At this point, the gain is gradually restored at a second preset rate. The second preset rate refers to the rate of gain recovery, which is less than the first preset rate to prevent excessively rapid gain recovery from causing repeated instability. The gain increase process continues until the initial gain is restored, or until the absolute value of the difference is greater than or equal to the second preset threshold again during this period, at which point the increase stops.

[0104] Step A204: If the absolute value of the difference continues to increase to the third preset threshold after the initial attenuation at the first preset speed, the gain of the target noise reduction controller is attenuated at the third preset speed until the absolute value of the difference is less than the second preset threshold. Then, the gain of the target noise reduction controller is increased at the fourth preset speed until it is restored to the second target value. The second target value is the initial gain minus the first target value, the third preset threshold is greater than the second preset threshold, and the third preset speed is greater than the first preset speed.

[0105] If, during the attenuation process at the first preset speed, the absolute value of the difference not only fails to decrease but continues to increase and becomes greater than or equal to the third preset threshold, it indicates that the current anomaly level has exceeded the response range of the first-level response. The third preset threshold is greater than the second preset threshold and is used to define a more severe anomaly level. At this point, the attenuation is switched to the third preset speed for faster attenuation. The third preset speed is greater than the first preset speed, achieving faster gain suppression. Attenuation continues until the absolute value of the difference falls below the second preset threshold, and then the gain is increased at the fourth preset speed to restore it to the second target value. The second target value is the initial gain minus the first target value, which allows recovery to the level after only deducting the first-level attenuation amount. The fourth preset speed is less than the first preset speed but greater than the second preset speed, ensuring that the recovery speed after a severe anomaly is between rapid attenuation and slow recovery, guaranteeing a certain recovery efficiency while avoiding secondary instability caused by excessively rapid recovery.

[0106] Step A205: If, after restoring the gain of the target noise reduction controller to the second target value, the absolute value of the difference is less than the second preset threshold for a second preset duration, then the gain of the target noise reduction controller is increased at a fifth preset speed until it is restored to the initial gain or the absolute value of the difference is greater than or equal to the second preset threshold.

[0107] After the gain recovers to the second target value, the absolute value of the difference continues to be monitored. If it remains below the second preset threshold for a second preset duration, it indicates that the system has further stabilized. It should be noted that the second preset duration can be the same as or different from the first preset duration. The fifth preset speed is the same as the second preset speed, meaning a slower recovery speed is used to complete the final gain recovery phase. The gain increase process continues until the initial gain is fully recovered, or until the absolute value of the difference again exceeds or equals the second preset threshold during this process, at which point the increase stops.

[0108] In a specific implementation, during the normal operation of the noise reduction controller, the absolute value of the signal energy difference between the first signal and the second signal is monitored in real time. A second preset threshold is denoted as DT1, and a third preset threshold is denoted as DT2, where DT2 is greater than DT1. Optionally, DT1 can be 3dB, and DT2 can be 6dB.

[0109] When the absolute value of the difference is greater than or equal to DT1, it indicates a significant noise increase outside the noise reduction band. At this point, the gain is rapidly attenuated at a first preset rate, for example, 20 dB per second, until the total attenuation reaches the first target value GT1 and remains constant. The first target value GT1 can be set according to the actual needs of the equipment, for example, 3 dB. If, after this, the absolute value of the difference falls below DT1 and this state is maintained for a first preset duration HT1, it indicates that the unstable factor has been eliminated or weakened. HT1 can be set, for example, to 1 second, to avoid triggering recovery due to instantaneous fluctuations. At this point, the gain of the target noise reduction controller is gradually increased at a second preset rate until it recovers to the initial gain, or the recovery stops when the absolute value of the difference again exceeds or equals DT1 during this process. The second preset rate can be set to a slower rate, less than the first preset rate, for example, 3 dB per second, to prevent excessively rapid gain recovery from causing repeated instability. If, after the gain has begun to attenuate at the first preset rate, the absolute value of the difference not only fails to decrease but continues to increase and exceeds or equals DT2, it indicates that the current anomaly level has exceeded the response range of the first-level response. At this point, the gain of the target noise reduction controller is switched to the third preset speed for faster attenuation, for example, 40dB per second, until the absolute value of the difference falls below DT1. Then, the gain is increased at a fourth preset speed, which is between the first and second preset speeds, for example, 10dB per second, until it recovers to the second target value and then pauses. This setting ensures that during the recovery phase after a severe anomaly, the gain recovery speed is between a fast attenuation speed and a slow recovery speed, guaranteeing a certain recovery efficiency while avoiding secondary instability caused by excessively rapid recovery. The energy difference between the two signals continues to be monitored. If the absolute value of the difference is less than DT1 and can be maintained for a second preset duration HT2, it indicates that the system has further stabilized. At this point, the gain of the target noise reduction controller continues to increase at a fifth preset speed until it is fully restored to the initial gain. If the absolute value of the difference is again greater than or equal to the second preset threshold during the increase, the increase is stopped.

[0110] It is understandable that, in actual implementation, the values ​​of each parameter can be flexibly adjusted according to actual needs. This embodiment does not impose restrictions on the specific values ​​of each parameter.

[0111] To more intuitively demonstrate the effect of this embodiment, refer to... Figure 6 , Figure 6 This is an exemplary spectrogram illustrating the control effect in this embodiment. Figure 6 The two spectrograms (a) and (b) in the figure record the sound signals recorded near the speaker. The horizontal axis represents time, and the vertical axis represents frequency. At approximately 2 seconds, an obstacle was placed close to the speaker and an error microphone positioned near the ear. Figure 6Figure (a) shows the case where the control scheme of this embodiment is not adopted: when the obstacle approaches, the acoustic transfer function is severely disturbed, and the change is close to or exceeds the system design margin, resulting in the noise reduction controller mismatch. The spectrogram shows obvious howling energy concentration, and the system howls. Figure 6 Figure (b) illustrates the situation using the control scheme of this embodiment: after the same obstacle interference is introduced, when the abnormal howling just begins to occur, the stability control mechanism quickly detects the abnormality and executes gain control, effectively suppressing the howling. No continuous strong energy howling pattern appears in the spectrogram, avoiding continuous auditory disturbance to the user.

[0112] This embodiment provides a multi-level threshold differential gain speed control mechanism. When the anomaly is mild, it decays at a faster speed and recovers at a slower speed. When the anomaly worsens, it decays at a faster speed and recovers in steps at a moderate speed between the fast and slow speeds. This graded and speed-based control method can effectively suppress abnormal noise increases caused by controller mismatch or sudden changes in the acoustic environment while ensuring noise reduction effect.

[0113] Based on the first, second, and / or third embodiments described above, a fourth embodiment of the noise control method of this application is proposed. In this embodiment, content that is the same as or similar to the first, second, and third embodiments described above can be referred to the above description and will not be repeated hereafter. In this embodiment, step S10: acquiring the environmental noise signals collected by each first microphone can be further refined into steps S101 to S104: Step S101: Acquire target signal, wherein the target signal is the original signal collected by the target microphone or the signal after high-pass filtering of the original signal at a preset frequency, and the target microphone is one of the microphones set in the near-ear open audio device.

[0114] In this embodiment, to make a preliminary judgment on the current acoustic environment before entering complex noise pattern recognition and noise localization, a microphone needs to be selected as the signal source. This microphone is referred to as the target microphone in this embodiment. By analyzing the signal collected by the target microphone, the environmental noise level can be efficiently judged and the wearer's voice activity can be detected, thereby terminating the process in advance when appropriate and saving subsequent processing resources. The target microphone can be any one of the microphones set in the near-ear open-back audio device. The microphones here include both the aforementioned first microphone and any microphone in the noise-canceling microphone group. For example, if the noise-canceling microphone group includes an error microphone, the error microphone can also be used as the target microphone when the speaker is not playing sound. When the speaker is playing sound, the error microphone may be directly affected by the content played by the speaker. In this case, other microphones can be selected as the target microphone first. This embodiment does not limit the specific method of selecting the target microphone. The target signal refers to the signal collected by the target microphone. The target signal can be the original signal output by the target microphone directly, or it can be the signal after high-pass filtering of the original signal at a preset frequency. The purpose of high-pass filtering is to filter out interference components such as low-frequency wind noise or circuit noise in order to more accurately reflect the actual level of environmental noise.

[0115] In one feasible implementation, the signal acquired by the target microphone can be sampled at a lower sampling rate, for example, a sampling rate of 4kHz. The sampled signal is then processed into frames at certain time intervals, for example, one frame is acquired every 1 second. The signal data within each frame can first undergo high-pass filtering to remove low-frequency interference; the cutoff frequency of the high-pass filter can be set, for example, to 100Hz. The filtered signal is the target signal. It is understood that the sampling rate, time interval, and filter cutoff frequency involved in the above steps are illustrative. In actual implementation, each parameter can be flexibly adjusted according to device requirements. This embodiment does not limit the specific values ​​of each parameter.

[0116] Step S102: Calculate the average signal energy of the target signal in the most recent first preset frame number to obtain the second signal feature value.

[0117] It should be noted that the second signal feature value refers to a feature parameter obtained by statistically processing the signal energy of the target signal across multiple time frames. This parameter characterizes the energy level of the target signal over a period of time. The energy level can be used as a basis for subsequent detection of the wearer's voice activity. The second signal feature value can be obtained by averaging the signal energy of multiple frames of the target signal; the number of frames involved in the averaging can be preset according to actual needs.

[0118] In a specific implementation, the target signal is segmented into multiple time frames, and the signal energy of each frame is calculated. The signal energy of each frame can be obtained by calculating the sum of squares and / or root mean square value of the sampled values ​​within the frame, which can characterize the signal strength. The calculated signal energy of each frame is temporarily stored. As time progresses, the signal energy of each frame is continuously calculated and recorded. When the accumulated number of frames reaches a first preset number of frames, the signal energy values ​​within the most recent first preset number of frames are averaged to obtain the average value, which is the second signal characteristic value. The first preset number of frames can be determined based on the time window length and sampling interval.

[0119] Step S103: If the second signal feature value is greater than or equal to the fourth preset threshold, then the wearer's voice activity is detected based on the signal energy of the target signal of the most recent second preset frame number; otherwise, return to step S101, wherein the second preset frame number is greater than the first preset frame number.

[0120] The obtained second signal feature value is compared with a preset fourth threshold. If the second signal feature value is greater than or equal to the fourth threshold, it indicates that the current environmental noise level has reached a point where noise reduction processing is required. Further, based on the signal energy of the target signal in the most recent second preset frame, speech activity detection is performed to determine if the wearer is speaking. If the wearer's speech activity is detected, it indicates that the current target signal contains strong interference signals generated by the wearer's own speech. The energy of the wearer's speech is usually much greater than the environmental noise; if processing continues, it will severely interfere with the accuracy of subsequent noise pattern recognition and noise localization. Therefore, the process returns directly to the step of acquiring the target signal, re-acquires the signal, and waits for the wearer's speech to end before proceeding with subsequent processing. If no wearer's speech activity is detected, the subsequent step of echo cancellation on the original signals acquired by each first microphone continues. The second preset frame number is greater than the first preset frame number, meaning the number of frames used for speech activity detection is greater than the number of frames used for environmental noise level judgment, in order to obtain more comprehensive speech feature information.

[0121] If the second signal feature value is less than the fourth preset threshold, it indicates that the current environment is relatively quiet and no further processing is required. The process can return to step S101 to continue collecting the target signal and monitoring changes in environmental noise.

[0122] Step S104: If the wearer's voice activity is detected, return to step S101.

[0123] It should be noted that when the wearer is speaking, their voice is picked up by the first microphone through air conduction or device structure conduction, creating an interference signal with an intensity much greater than the ambient noise. If this signal mixed with the wearer's voice is directly used as the ambient noise signal for subsequent processing, the wearer's own voice components will interfere with the accuracy of noise pattern recognition and noise localization, thus affecting the noise reduction effect. Therefore, when the wearer's voice activity is detected, the currently acquired signal is no longer suitable for ambient noise analysis, and the process returns directly to step S101 to reacquire the target signal, waiting for the wearer's voice activity to end before proceeding with subsequent processing.

[0124] In step S105, if no voice activity of the wearer is detected, echo cancellation is performed on the original signals collected by each first microphone to obtain an environmental noise signal.

[0125] It should be noted that when the wearer is not engaged in any voice activity, the signal captured by the first microphone may still contain distant speech content played by the speaker. The sound from the speaker can be picked up again by the first microphone through spatial coupling or device structure conduction, forming an acoustic echo. If this echo component is not eliminated, it will interfere with the accuracy of subsequent noise pattern recognition and noise localization. Therefore, after confirming that the wearer is not engaged in any voice activity, echo cancellation processing is performed on the signal captured by the first microphone to filter out the distant speech echo component, resulting in a cleaner environmental noise signal.

[0126] In one feasible implementation, an adaptive filtering method can be used for echo cancellation. Specifically, the echo cancellation process can begin by using white noise played by an audio device's speaker as a reference signal, and then using a designated microphone to collect the echo signal fed back through the acoustic path. Based on the white noise signal and the collected echo signal, a preset adaptive algorithm is used to iteratively update the filter coefficients until the filter converges. The set of filter coefficients obtained after convergence characterizes the echo transmission path characteristics in the current acoustic environment. Commonly used adaptive algorithms can be, for example, least mean square algorithms. During actual product operation, the pre-converged filter coefficients can be loaded into a finite impulse response (FIR) filter. When the speaker plays normal audio content, the audio content signal is input into the FIR filter loaded with coefficients, and the output of the filter is the estimated echo signal. Subsequently, the estimated echo signal is subtracted from the original signal collected by the microphone to obtain the echo-cancelled ambient noise signal.

[0127] Furthermore, in one feasible embodiment, to reduce the time delay introduced by the finite impulse response (FIR) filter during signal processing, the converged filter coefficients can be transformed from the time domain to the complex frequency domain, and a set of infinite impulse response (IR) filters can be designed accordingly, such that the frequency response of the IIR filters is the same as or approximately the same as that of the FIR filters. In actual operation, the IIR filters are used instead of the FIR filters to filter the audio signal, thereby reducing processing delay while achieving echo cancellation. It is understood that the specific implementation of the echo cancellation process described above is only an illustrative example, and adjustments can be made according to the actual needs of the device in actual implementation. This embodiment does not impose any limitations on this.

[0128] In another alternative implementation, if the near-ear open-back audio device integrates a physical detection module such as a voice pickup unit or an accelerometer, the wearer's voice activity detection involved in steps S103 to S105 can directly utilize the detection results of the physical detection module.

[0129] Specifically, the voice pickup unit can directly detect the physical vibrations generated when the wearer speaks through bone conduction or vibration sensing, thereby accurately determining whether the wearer is engaged in voice activity. An accelerometer can assist in determining the presence or absence of voice activity by detecting the micro-vibration characteristics of the wearer's head or facial bones. When the physical detection module indicates that the wearer is not engaged in voice activity, the echo cancellation step can be skipped, or other preset methods can be used to process the environmental noise signal. It is understood that the above-described method of voice activity detection using a voice pickup unit or accelerometer is applicable to near-ear open-back audio devices that have integrated the corresponding hardware modules. For devices that do not integrate such modules, the aforementioned method of voice activity detection based on signal energy can still be used. This embodiment does not limit the specific implementation means of wearer voice activity detection.

[0130] In this embodiment, during the acquisition of environmental noise signals, the environmental noise level is first monitored to determine whether to proceed to subsequent processing. When necessary, the wearer's voice activity is further detected to determine whether to enable echo cancellation, thereby effectively controlling the overall power consumption of the device while ensuring noise reduction effect.

[0131] In one feasible implementation, step S103 includes the following steps A131 to A133: Step A131: For the target signal of the most recent second preset frame number, calculate the upper and lower boundaries of the box line according to the preset box line parameters and the signal energy of the target signal of each frame.

[0132] It should be noted that during the operation of the audio device, when the wearer is not engaged in speech activity, the signal energy of the target signal typically fluctuates within a relatively stable range. When the wearer begins to speak, the signal energy of the target signal will show a significant increase, exceeding the normal fluctuation range. Therefore, by pre-determining the upper and lower boundaries of the normal energy fluctuation range, a quantitative basis can be provided for subsequent judgment of whether wearer speech activity exists. The second preset frame number refers to the number of frames of the target signal participating in this statistical processing; its specific value can be set according to actual needs. The box-line parameter defines a set of rules or conditions used to define the normal data distribution range. Obtain the signal energy of each frame of the target signal within the most recent second preset frame number. Perform statistical processing on this set of signal energy according to the preset box-line parameter to obtain the upper and lower boundaries of the box-line. The interval between the upper and lower boundaries represents the normal energy fluctuation range under the preset statistical meaning; data points exceeding this interval can be considered outliers.

[0133] In one feasible implementation, a first quantile level and a second quantile level are preset, with the first quantile level being lower than the second quantile level. From the signal energy of the target signal in the most recent second preset frame number, the signal energy corresponding to the first quantile level and the signal energy corresponding to the second quantile level are determined. The difference between the two signal energies is calculated, reflecting the degree of dispersion in the statistical distribution of this group of signal energies. Based on the difference, a lower boundary is obtained by shifting the signal energy corresponding to the first quantile level downwards by a preset multiple; an upper boundary is obtained by shifting the signal energy corresponding to the second quantile level upwards by the same multiple. The interval between the lower and upper boundaries is the normal energy fluctuation range; signal energy exceeding this range can be considered outliers.

[0134] Furthermore, in a specific implementation, the box-line parameters may include a first quantile level, a second quantile level, and a preset multiple. The box-line parameters can be set as follows: A first quantile level q1 is pre-configured. The specific value of q1 can be set according to actual needs during the product development and debugging phase. Different settings will affect the position of the box-line boundary. The second quantile level q2 can be determined based on q1, such as q2 = 100% - q1. In this group of signal energies, the signal energy at position q1 is denoted as Q1, and the signal energy at position q2 is denoted as Q2. The difference between Q2 and Q1 is calculated and denoted as the quantile difference. Based on the quantile difference and the preset boundary multiple, Q1 is subtracted from the product of the quantile difference and the boundary multiple to obtain the lower boundary of the box-line. Q2 is added to the product of the quantile difference and the boundary multiple to obtain the upper boundary of the box-line. That is, the lower boundary of the box-line = Q1 - boundary multiple × quantile difference, and the upper boundary of the box-line = Q2 + boundary multiple × quantile difference. Optionally, a quartile approach is used as an example for explanation. At this point, the first quantile level q1 = 25%, and the second quantile level q2 = 100% - q1 = 100% - 25% = 75%. Under this configuration, the signal energy at the first quantile level is denoted as Q1, and the signal energy at the second quantile level is denoted as Q3. The difference between the two is calculated, i.e., the interquartile range IQR = Q3 - Q1. For example, if the boundary multiple is 1.5, then the lower boundary of the box line = Q1 - 1.5 * IQR, and the upper boundary = Q3 + 1.5 * IQR. The interval between these upper and lower boundaries is the normal energy fluctuation range. When the signal energy of a frame exceeds the upper boundary, that frame can be considered an outlier, indicating that there may be abnormal increases in wearer voice activity or environmental noise. It is understood that the specific definitions and values ​​of the box line parameters involved in the above steps are illustrative. In actual implementation, each parameter can be flexibly adjusted according to actual needs. This embodiment does not impose any limitations on this.

[0135] Step A132: Determine the outlier ratio of the signal energy of the target signal in each frame based on the upper and lower boundaries.

[0136] After obtaining the upper and lower boundaries, it is determined whether the signal energy of each frame within the most recent second preset frame number exceeds the upper or lower boundary. If the signal energy of a frame is greater than the upper boundary or less than the lower boundary, that frame is marked as an outlier. The number of outliers is counted, and the proportion of outliers to the second preset frame number is calculated to obtain the outlier ratio. It can be understood that the outlier ratio reflects the degree of abnormal fluctuation in signal energy within that time period.

[0137] Step A133: If the proportion of outliers of the preset number of consecutively calculated outliers is greater than the fifth preset threshold, then it is determined that the wearer's voice activity has been detected.

[0138] It should be noted that the fifth preset threshold is used to distinguish between normal energy fluctuations and continuous energy anomalies caused by the wearer's speech. The calculated outlier ratio is compared with the preset fifth threshold. To avoid misjudgments caused by transient interference, the outlier ratios calculated multiple times consecutively are monitored. When the outlier ratios of a preset number of consecutive calculations are all greater than the fifth preset threshold, it indicates that there is a continuous and significant abnormal fluctuation in signal energy, consistent with the energy change characteristics of the wearer's speech activity. At this point, the wearer's speech activity is detected. If the outlier ratio calculated in any instance does not exceed the fifth preset threshold, the continuous counting can be reset, and monitoring can continue.

[0139] To facilitate understanding of the above judgment logic, refer to Figure 3 , Figure 4 as well as Figure 5 The three figures provide examples of outlier analysis for a real speech segment under different noise backgrounds. The horizontal axis of each figure represents time.

[0140] Figure 3 The waveform of a real speech segment is shown, with the vertical axis representing the signal amplitude. This speech segment is the wearer's speech sample used in the subsequent mixed noise test; its duration is relatively short, and the energy distribution is concentrated within the speech activity period. Figure 4 The diagram shows the situation where the speech is mixed with different types of noise. The vertical axis also represents the signal amplitude. Each type of noise lasts for about 15 seconds. As can be seen from the waveform, the original speech is submerged by background noise of different intensities. The noise floor is significantly different in different time periods, perfectly simulating the real environment of complex non-stationary noise and sudden interference in actual scenarios.

[0141] Figure 5 The image shows the detection results using the algorithm in this embodiment. The red horizontal line represents the fifth preset threshold, and the black dotted, solid, and dashed lines represent the outlier detection curves for different values ​​of the box-circuit parameter. Different box-circuit parameter values ​​affect the detection sensitivity and response speed, specifically manifested in the difference in the rise and trend of the outlier curve. Figure 5 It can be seen that during periods without voice activity, the outlier detection curves for different box-line parameter values ​​all remained below the red threshold, indicating that the system can maintain a low false trigger rate in the absence of voice activity. When the wearer engages in voice activity, the outlier curves for all three parameter configurations show a significant increase, successively exceeding the red threshold. Specifically, some parameter configurations exhibit a faster increase in the outlier curve response, triggering detection earlier in the voice activity; others show higher peak values, reflecting a stronger sensitivity to changes in voice activity energy. After the voice activity ends, the outlier curves gradually fall back below the threshold, returning to normal monitoring status.

[0142] The detection results show that even when the target speech has an extremely low energy proportion in strong background noise, it can still be identified as an abnormal interference signal through quantile statistics. By adjusting the values ​​of the box-line parameters, the detection sensitivity of the algorithm can be flexibly adjusted to adapt to different application scenarios.

[0143] This embodiment provides a method for detecting wearer speech activity based on box plot outlier analysis. Unlike conventional speech activity detection methods based on fixed energy thresholds or single feature comparisons, this embodiment utilizes the statistical properties of box plots to dynamically determine the normal range of the current signal energy distribution and determines speech activity by continuously monitoring the persistence of the outlier ratio. This method does not rely on a preset absolute energy threshold and can better adapt to changes in energy baseline under different environmental noise backgrounds, thereby improving the robustness and accuracy of speech activity detection. It should be noted that applying box plot outlier analysis to wearer speech activity detection in near-ear open-back audio devices is not a common application in this field; the solution provided in this embodiment fills this gap.

[0144] This application also provides a near-ear open-back audio device, which can be used to execute the noise control method described in any of the foregoing embodiments. The device includes at least two first microphones for acquiring ambient noise signals, providing raw acoustic input for noise pattern recognition, noise localization, and active noise control in the foregoing embodiments. The specific placement and number of the first microphones on the device body can be flexibly configured according to the device form. The near-ear open-back audio device also includes a memory and a processor. The memory stores a computer program, which the processor can execute. When the computer program is executed by the processor, the near-ear open-back audio device performs the steps of the noise control method in any of the foregoing embodiments, including acquiring ambient noise signals, identifying target noise patterns, determining the direction of target noise, selecting a noise reduction controller, and performing active noise control. Furthermore, the near-ear open-back audio device may also include an error microphone and a speaker positioned near the ear.

[0145] In one feasible implementation, refer to Figure 7 Near-ear open-back audio devices can be OWS (Open Wearable Stereo) headphones. OWS headphones include a housing 1, which is positioned near the user's ear without obstructing the ear canal when worn. At least two primary microphones are located within the housing 1, for example... Figure 7 The image shows two first microphones 2, and Figure 7 An error microphone 3 and a speaker 4 are also shown positioned near the ear within the housing 1. It should be noted that... Figure 7This is only used to illustrate the placement of the microphone and speaker in an OWS headset; therefore, other components, such as support structures, may also be included in the headset. Figure 7 It is not shown in the text.

[0146] In another feasible implementation, the near-ear open-back audio device can be smart glasses. (See reference) Figure 8 The smart glasses 5 include a frame and temples connected to both sides of the frame. At least one temple has at least one speaker 4 and at least one error microphone 3 mounted on its ear loop. The ear loop is the area where the temple rests above the user's ear when worn. The speaker is located on or near the ear loop facing the ear canal and is used to play audio content and generate anti-phase noise cancellation waves. The error microphone 3 is located near the ear on the ear loop facing the ear canal or near the speaker's sound outlet, and is used to collect near-ear acoustic signals after noise reduction processing. Regarding the placement of the first microphone, in the form of smart glasses, at least one of the following layouts can be adopted: at least one first microphone is located on the temple on a side further away from the temple root than on the ear loop; at least one first microphone is located on the temple on a side closer to the temple root than on the ear loop; at least one first microphone is located on the frame of the smart glasses.

[0147] Optionally, refer to Figure 8 Taking the setup of two first microphones as an example, one first microphone 2 can be located in the front area of ​​the temple, that is, closer to the connection between the temple and the frame compared to the ear loops; the other first microphone 2 can be located in the rear area of ​​the temple, that is, further away from the base of the temple compared to the ear loops. This arrangement allows for the acquisition of environmental noise reference signals with different acoustic characteristics by utilizing the spatial difference between the two first microphones, thereby providing spatial information for direction estimation in noise characteristic analysis.

[0148] Further, refer to Figure 8 In a specific implementation, the distance between the two first microphones 2 and the error microphone 3 can be set to be the same. This arrangement ensures that the acoustic path lengths of environmental noise from different directions, after reaching their respective first microphones, tend to be consistent before propagating to the error microphone. In subsequent active noise control, this symmetrical layout helps simplify the computational complexity of noise localization and makes the acoustic transmission relationship between the reference signal and the error signal more symmetrical, facilitating the noise reduction controller to generate anti-phase noise cancellation waves more stably, thereby improving the overall noise reduction effect.

[0149] In some alternative implementations, the number of first microphones may exceed two. For example, see reference... Figure 8An additional first microphone 2 can be installed on the frame to collect ambient noise from above or the side, further enriching the spatial dimension information. The specific installation position and pointing angle of each first microphone can be optimized according to the industrial design, acoustic structure, and target noise reduction scenario of the smart glasses.

[0150] It is understood that the above hardware structure description using smart glasses as an example is only used to explain the implementation of the technical solution of this application in a specific product form, and does not constitute a limitation on the scope of protection of this application. In other possible implementations, the near-ear open-back audio device may also be an ear-hook audio player, an augmented reality headset, or other wearable devices with a near-ear open-back acoustic structure. As long as the device has at least two first microphones, a memory, and a processor, and is able to execute the steps of the noise control method in any embodiment of this application, it falls within the scope of protection of this application.

[0151] Furthermore, the number of first microphones, their specific installation locations, the specific layout of the speakers and error microphones, as well as hardware parameters such as processor type and memory capacity, can all be adjusted according to actual product requirements. This application does not impose any restrictions on the specific values ​​of the aforementioned hardware parameters.

[0152] Compared with the prior art, the beneficial effects of the near-ear open-back audio device provided in this application embodiment are the same as the beneficial effects of the noise control method provided in the above embodiments, and other technical features in the near-ear open-back audio device correspond to the features disclosed in the foregoing embodiments, which will not be repeated here.

[0153] The above description is merely an embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0154] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, which are used to execute the noise control method in the above embodiments.

[0155] The computer-readable storage medium provided in this application embodiment may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0156] The aforementioned computer-readable storage medium may be included in a near-ear open-back audio device; or it may exist independently and not assembled into a near-ear open-back audio device.

[0157] The aforementioned computer-readable storage medium carries one or more programs that, when executed by a near-ear open-back audio device, cause the near-ear open-back audio device to perform the functions defined in the methods of the embodiments disclosed in this application.

[0158] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0159] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0160] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0161] The readable storage medium provided in this application embodiment is a computer-readable storage medium, which stores computer-readable program instructions (i.e., computer programs) for performing the above-described noise control method. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application embodiment are the same as the beneficial effects of the noise control method provided in the above-described embodiments, and will not be repeated here.

[0162] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the noise control method described above.

[0163] Compared with the prior art, the beneficial effects of the computer program product provided in this application embodiment are the same as the beneficial effects of the noise control method provided in the above embodiments, and will not be repeated here.

[0164] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A noise control method, characterized in that, The noise control method is applied to a near-ear open-back audio device, wherein at least two first microphones are provided in the near-ear open-back audio device, and the noise control method includes: Acquire the ambient noise signals collected by each of the first microphones; Based on the environmental noise signal, noise pattern recognition is performed to obtain the target noise pattern; Based on the environmental noise signal, noise localization is performed to obtain the target noise direction; Select a target noise reduction controller from a set of preset noise reduction controllers that corresponds to the target noise mode and the target noise direction; Active noise control is performed based on the target noise reduction controller; The step of acquiring the environmental noise signals collected by each of the first microphones includes: Acquire a target signal, wherein the target signal is the original signal collected by the target microphone or the signal after the original signal has been high-pass filtered at a preset frequency, and the target microphone is one of the microphones provided in the near-ear open audio device; Calculate the average signal energy of the target signal over the most recent first preset frame number to obtain the second signal characteristic value; If the second signal feature value is greater than or equal to the fourth preset threshold, then wearer voice activity detection is performed based on the signal energy of the target signal at the most recent second preset frame number; otherwise, return to the step of obtaining the target signal, wherein the second preset frame number is greater than the first preset frame number. If the wearer's voice activity is detected, return to the step of acquiring the target signal; If no voice activity of the wearer is detected, the original signals collected by each of the first microphones are subjected to echo cancellation to obtain the environmental noise signal.

2. The noise control method as described in claim 1, characterized in that, The step of performing noise pattern recognition based on the environmental noise signal to obtain the target noise pattern includes: A bandpass filter is used to filter the target environmental noise signal to obtain a filtered signal. The target environmental noise signal is any one of the environmental noise signals or the one with the largest signal energy. The center frequency of the bandpass filter is designed according to the energy spectrum of the preset type of environmental noise. The first signal feature value is obtained by weighting the signal energy of multiple frames of the filtered signal, wherein the weighting weight of the signal energy of the later filtered signal is greater than the weighting weight of the signal energy of the earlier filtered signal; The target noise pattern is determined based on the comparison result between the first signal feature value and the first preset threshold.

3. The noise control method as described in claim 2, characterized in that, The number of bandpass filters is two, and the center frequencies of the two bandpass filters are designed based on the two frequency bands with the largest energy contrast in the energy spectrum of various preset types of environmental noise. The first preset threshold includes a first threshold, a second threshold, a third threshold, and a fourth threshold, and the third threshold is less than the first threshold. The step of determining the target noise pattern based on the comparison result of the first signal feature value and the preset threshold includes: If the difference between the first feature value and the second feature value is greater than the first threshold, then return to the step of weighting the signal energy of the filtered signals of multiple frames to obtain the first signal feature value, wherein the first feature value is the first signal feature value of the filtered signal obtained by filtering with the bandpass filter with a larger center frequency, and the second feature value is the first signal feature value of the filtered signal obtained by filtering with the bandpass filter with a smaller center frequency; If the sum of the first feature value and the second feature value is greater than the second threshold, and the difference between the first feature value and the second feature value is greater than the third threshold, then the preset first noise mode is determined to be the target noise mode. If the sum of the first feature value and the second feature value is greater than the second threshold, and the difference between the first feature value and the second feature value is less than or equal to the third threshold, then the preset second noise mode is determined to be the target noise mode. If the sum of the first feature value and the second feature value is less than or equal to the second threshold and greater than the fourth threshold, then the preset third noise mode is determined as the target noise mode. If the sum of the first feature value and the second feature value is less than or equal to the fourth threshold, then the preset fourth noise mode is determined as the target noise mode.

4. The noise control method as described in claim 1, characterized in that, The steps of performing active noise control based on the target noise reduction controller include: The signal weight of each microphone in the noise-canceling microphone group is determined according to the target noise direction, wherein the signal weight of the microphone closer to the target noise direction is greater than the signal weight of the microphone farther away from the target noise direction, and the noise-canceling microphone group includes a reference microphone and / or an error microphone disposed in the near-ear open audio device, and the first microphone includes all or part of the reference microphone. Active noise control is performed based on the signals collected by each microphone in the noise-canceling microphone group and the target noise-canceling controller, and the signals collected by the corresponding microphones are weighted based on the signal weights during the active noise control process.

5. The noise control method as described in claim 1, characterized in that, The near-ear open-back audio device includes a reference microphone and an error microphone, the first microphone comprising all or part of the reference microphone, and after the step of performing active noise control based on the target noise reduction controller, it further includes: A first signal is obtained by filtering the signal acquired by one of the reference microphones using a bandpass filter with a preset frequency band, and a second signal is obtained by filtering the signal acquired by one of the error microphones using the same bandpass filter, wherein the preset frequency band does not overlap with the noise reduction frequency band of the target noise reduction controller; The gain of the target noise reduction controller is adjusted based on the signal energy comparison results of the first signal and the second signal.

6. The noise control method as described in claim 5, characterized in that, The step of adjusting the gain of the target noise reduction controller based on the signal energy comparison result of the first signal and the second signal includes: Calculate the absolute value of the difference between the signal energies of the first signal and the second signal; If the absolute value of the difference is greater than or equal to the second preset threshold, the gain of the target noise reduction controller is attenuated at a first preset rate until the attenuation reaches the first target value. If the state where the absolute value of the difference is less than the second preset threshold continues for a first preset duration, the gain of the target noise reduction controller will be increased at a second preset rate until it is restored to the initial gain of the target noise reduction controller or the absolute value of the difference is greater than or equal to the second preset threshold. If the absolute value of the difference continues to increase to a third preset threshold after the initial attenuation at the first preset speed, the gain of the target noise reduction controller is attenuated at a third preset speed until the absolute value of the difference is less than the second preset threshold. Then, the gain of the target noise reduction controller is increased at a fourth preset speed until it returns to the second target value. The second target value is the initial gain minus the first target value. The third preset threshold is greater than the second preset threshold, and the third preset speed is greater than the first preset speed. If, after restoring the gain of the target noise reduction controller to the second target value, the absolute value of the difference is less than the second preset threshold for a second preset duration, then the gain of the target noise reduction controller will be increased at a fifth preset rate until it is restored to the initial gain or the absolute value of the difference is greater than or equal to the second preset threshold.

7. The noise control method as described in claim 1, characterized in that, The step of detecting wearer voice activity based on the signal energy of the target signal at the most recent second preset frame number includes: For the target signal of the most recent second preset frame number, the upper and lower boundaries of the box line are calculated according to the preset box line parameters and the signal energy of the target signal in each frame; The proportion of outliers in the signal energy of the target signal in each frame is determined based on the upper and lower boundaries. If the proportion of outliers of a preset number of consecutively calculated outliers is greater than a fifth preset threshold, then it is determined that the wearer's voice activity has been detected.

8. A near-ear open-back audio device, characterized in that, The near-ear open-back audio device is provided with at least two first microphones, the first microphones being used to collect ambient noise signals. The near-ear open-back audio device further includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the noise control method as described in any one of claims 1 to 7.

9. The near-ear open-back audio device as described in claim 8, characterized in that, The near-ear open-back audio device is a smart glasses, wherein at least one temple of the smart glasses has at least one speaker and at least one error microphone on the ear hook portion; at least one first microphone is located on the temple on a side further away from the root of the temple than on the ear hook portion; and at least one first microphone is located on the temple on a side closer to the root of the temple than on the ear hook portion, or, at least one first microphone is located on the frame of the smart glasses.

Citation Information

Patent Citations

  • Noise reduction processing method and device, electronic equipment, earphone and storage medium

    CN113949955A

  • Open type wearable acoustic equipment and active noise reduction method

    CN119096556A