A noise reduction regulation method for earphones and a noise reduction earphone
By constructing a human voice feature template library in the headphones and using Fourier transform to distinguish human voice from noise, and employing an active noise reduction algorithm to generate canceling sound waves and amplify human voice signals, the problem of headphones being unable to transmit human voice in high-noise environments is solved, achieving a combination of clear communication and safe production.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUANGZHOU SOUNDBOX ACOUSTIC TECH
- Filing Date
- 2025-08-12
- Publication Date
- 2026-04-21
AI Technical Summary
Existing headphone noise cancellation technology struggles to accurately identify and transmit specific human voices in high-noise environments, resulting in users being unable to clearly receive instructions or information during voice communication, posing safety risks and providing a poor user experience.
By collecting ambient sound signals at preset time intervals, constructing a human voice feature template library using an attention mechanism, extracting spectral features using Fourier transform, accurately distinguishing human voice from noise, and using an active noise reduction algorithm to generate canceled sound waves and amplified human voice signal output.
It effectively filters noise in high-noise environments, clearly transmits specific human voices, and allows for smooth communication without removing headphones, thus improving user experience and production safety.
Smart Images

Figure CN120833774B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of headphone noise reduction technology, and more specifically, to a noise reduction control method for headphones and noise-canceling headphones. Background Technology
[0002] Currently, mainstream headphone noise cancellation technologies are mainly achieved through a combination of active noise cancellation (ANC) and passive noise cancellation. Active noise cancellation technology uses microphones to collect ambient noise and then uses noise cancellation circuitry to generate sound waves with the opposite phase to cancel out the noise, thereby reducing ambient noise. Passive noise cancellation, on the other hand, relies on the physical structure of the headphones, such as sealed earplugs and earcups, to block the entry of external sounds.
[0003] However, existing noise reduction technologies have significant limitations in practical applications. When users are in scenarios requiring voice interaction, such as receiving instructions, alarms, or engaging in voice communication in high-noise environments like construction sites, manufacturing workshops, mines, and hydroelectric power stations, traditional noise reduction modes filter out most of the ambient sound, including useful sounds like human voices. This makes it difficult for users to receive instructions, alarms, or engage in voice communication normally, requiring them to remove their headphones to hear the voice. This is not only cumbersome but may also block speech, easily leading to safety accidents and potential safety hazards, thus creating a conflict between occupational health and safety.
[0004] While some headphones are equipped with ambient sound pass-through functionality, attempting to pick up external sounds through microphones and transmit them into the ear canal to address the aforementioned issues, this pass-through function often transmits all ambient sounds indiscriminately. This includes both the desired human voice and various unwanted noises, such as daily work noises (e.g., drilling, chainsaws, pile driving), equipment operating noises, and unwanted human voices. This means that while users can hear human voices after enabling pass-through, they are also subjected to significant sound interference, failing to achieve a clear and comfortable communication experience and thus failing to meet their needs in complex environments.
[0005] Therefore, the headphone noise reduction control method that can accurately identify and transmit specific human voices while efficiently filtering other environmental noises is an important problem that urgently needs to be solved in the current headphone noise reduction technology field. It is of great significance for improving user experience and ensuring safe production. Summary of the Invention
[0006] In response, the present invention provides a noise reduction control method for headphones, noise-canceling headphones, electronic devices, computer storage media, and computer program products to solve at least one of the above-mentioned technical problems.
[0007] A first aspect of the present invention provides a noise reduction control method for headphones, comprising the following method steps:
[0008] Before receiving a noise reduction control command, the system collects a first ambient sound signal of the surrounding environment at a preset time interval using at least one microphone on the headphones. An attention mechanism is then used to extract human voice features from the first ambient sound signal to construct a human voice feature template library. The preset time interval is determined based on the ambient sound features representing the mobility of people extracted from the first ambient sound signal.
[0009] In response to the noise reduction control command, the second ambient sound signal of the environmental area is collected, the high-frequency interference and low-frequency background noise in the second ambient sound signal are removed to obtain the time domain signal, the time domain signal is converted into the frequency domain signal through Fourier transform, and the spectral features are extracted from the frequency domain signal.
[0010] The spectral features are matched with the human voice feature template library to determine the signals in the environmental sound signals that conform to the human voice frequency range and mark them as human voice signals; signals in other frequency ranges are marked as noise signals.
[0011] An active noise cancellation algorithm is used to generate a canceling sound wave that is out of phase with the noise signal, and the human voice signal is amplified. The canceling sound wave and the amplified human voice signal are then output through the speaker of the headphones.
[0012] A second aspect of the present invention provides noise-canceling headphones, including a control unit, a storage unit, a speaker, and at least one microphone; the control unit performs the following steps by calling computer program code stored in the storage unit:
[0013] Before receiving a noise reduction control command, the system collects a first ambient sound signal of the surrounding environment at a preset time interval using at least one microphone on the headphones. An attention mechanism is then used to extract human voice features from the first ambient sound signal to construct a human voice feature template library. The preset time interval is determined based on the ambient sound features representing the mobility of people extracted from the first ambient sound signal.
[0014] In response to the noise reduction control command, the second ambient sound signal of the environmental area is collected, the high-frequency interference and low-frequency background noise in the second ambient sound signal are removed to obtain the time domain signal, the time domain signal is converted into the frequency domain signal through Fourier transform, and the spectral features are extracted from the frequency domain signal.
[0015] The spectral features are matched with the human voice feature template library to determine the signals in the environmental sound signals that conform to the human voice frequency range and mark them as human voice signals; signals in other frequency ranges are marked as noise signals.
[0016] An active noise cancellation algorithm is used to generate a canceling sound wave that is out of phase with the noise signal, and the human voice signal is amplified. The canceling sound wave and the amplified human voice signal are then output through the speaker of the headphones.
[0017] A third aspect of the present invention provides an electronic device comprising: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, wherein the computer program, when executed by the processor, implements the method as described in any of the preceding claims.
[0018] A fourth aspect of the present invention provides a computer storage medium storing a computer program that can be executed by a processor to implement the method as described in any of the preceding claims.
[0019] A fifth aspect of the present invention provides a computer program product comprising a computer program executable by a processor to implement the method as described in any of the preceding claims.
[0020] This solution collects ambient sound at preset time intervals and constructs a human voice template library using an attention mechanism. It then combines Fourier transform to extract spectral features for matching, accurately distinguishing human voices from noise. Active noise reduction generates canceling sound waves while simultaneously amplifying and transmitting the specific human voice corresponding to the human voice template library. This invention achieves effective noise filtering and clear transmission of specific human voices, enabling smooth communication without removing headphones, thereby improving the user experience in complex environments and ensuring production safety. Attached Figure Description
[0021] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a flowchart illustrating a noise reduction control method for headphones disclosed in an embodiment of the present invention;
[0023] Figure 2 This is a schematic diagram of constructing a human voice feature template library based on the remaining human voice signal, as disclosed in an embodiment of the present invention;
[0024] Figure 3 This is a schematic diagram of the electrical structure of a noise-canceling headphone disclosed in an embodiment of the present invention. Detailed Implementation
[0025] The following specific embodiments illustrate the implementation of this application. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0026] Furthermore, the technical features involved in the different embodiments of this application described below can be combined with each other as long as they do not conflict with each other.
[0027] like Figure 1 As shown in the figure, an embodiment of the present invention discloses a noise reduction control method for headphones, including the following method steps:
[0028] S10, before receiving the noise reduction control command, the system collects the first ambient sound signal of the surrounding environment area through at least one microphone set on the headphones at a preset time interval, and extracts human voice features from the first ambient sound signal using an attention mechanism to construct a human voice feature template library; wherein, the preset time interval is determined based on the ambient sound features representing the mobility of people extracted from the first ambient sound signal.
[0029] like Figure 3 As shown, the noise-canceling headphones of this invention, in addition to being equipped with conventional built-in speakers and their electronic control and power supply components, also feature microphones for collecting ambient sound signals from the surrounding environment. These microphones are typically located on the headphone housing, but this invention does not specifically limit their installation location or number. Furthermore, users can activate or deactivate the noise-canceling pass-through function for specific human voices using a switch on the player or the headphones themselves.
[0030] When a user wears the aforementioned noise-canceling headphones and enters or remains in a specific environmental area, the headphones collect the first ambient sound signal of the surrounding area through at least one microphone. Simultaneously, an attention mechanism is employed to extract human voice features, thereby constructing a human voice feature template library. This attention mechanism focuses on key information related to the human voice, reducing interference from other noises and thus improving the accuracy of human voice feature extraction.
[0031] Furthermore, the acquisition of the first ambient sound signal before receiving the noise reduction control command is performed at preset time intervals. This setting allows for continuous updates to the human voice feature template library to adapt to changes in the flow of people within the environment. Specifically, the aforementioned preset time interval is determined based on the ambient sound features representing human mobility extracted from the first ambient sound signal. This setting makes the time interval setting more targeted, conforming to the sound change characteristics under different human movement conditions, thereby constructing a human voice feature template library that is more suitable for the current environment.
[0032] S20, in response to the noise reduction control command, acquire the second ambient sound signal of the environmental area, remove high-frequency interference and low-frequency background noise from the second ambient sound signal to obtain a time-domain signal, convert the time-domain signal into a frequency-domain signal through Fourier transform, and extract the spectral features from the frequency-domain signal.
[0033] When a user needs to enable a specific voice pass-through function, they can turn on the specific voice pass-through function of the noise-canceling headphones by using the switch button on the player or noise-canceling headphones, which will then automatically trigger the noise-canceling control command.
[0034] In response to the noise reduction control command, the sound signal of the environmental area is re-acquired, resulting in a second environmental sound signal, which is the sound signal requiring noise reduction control. First, high-frequency interference and low-frequency background noise are removed to obtain a cleaner time-domain signal, avoiding irrelevant noise from affecting subsequent analysis. Then, a Fourier transform is used to convert the previously processed time-domain signal into a frequency-domain signal, transforming the sound from a time-varying form into a frequency-distributed form, facilitating the analysis of sound characteristics from a frequency perspective. Next, spectral features are extracted from the obtained frequency-domain signal to distinguish human voices from other noise. Spectral features include, but are not limited to, the fundamental frequency, the formants to be matched, the spectral envelope to be matched, and the rate of change of the intensity to be matched.
[0035] S30, the spectral features are matched with the human voice feature template library to determine the signals in the environmental sound signals that conform to the human voice frequency range and are marked as human voice signals, while signals in other frequency ranges are marked as noise signals.
[0036] The spectral features extracted in step S20 are matched with the human voice feature template library constructed in step S10. Using the accurate human voice feature information already present in the template library, it is possible to precisely determine which parts of the second environmental sound signal belong to the human voice frequency range corresponding to these human voice feature templates. Signals that conform to the human voice frequency range are identified through matching and marked as human voice signals; the rest are marked as noise signals.
[0037] S40, an active noise reduction algorithm is used to generate a canceling sound wave that is out of phase with the noise signal, and the human voice signal is amplified. The canceling sound wave and the amplified human voice signal are then output through the speaker of the headphones.
[0038] For the identified noise signals, an Active Noise Cancellation (ANC) algorithm is used to generate canceling sound waves with opposite phases. These are then output through a speaker, canceling out ambient noise and reducing overall noise levels. Simultaneously, the human voice signal is amplified before being output through the speaker, ensuring the user can clearly hear specific human voices. Finally, the canceling sound waves and the amplified human voice signal are output simultaneously, achieving the goal of effectively filtering out other noise while ensuring clear transmission of specific human voices corresponding to the human voice feature template. This satisfies the user's need to enjoy noise reduction while maintaining smooth voice communication in complex environments.
[0039] This solution collects ambient sound at preset time intervals and constructs a human voice template library using an attention mechanism. It then combines Fourier transform to extract spectral features for matching, accurately distinguishing human voices from noise. Active noise reduction generates canceling sound waves while simultaneously amplifying and transmitting the specific human voice corresponding to the human voice template library. This invention achieves effective noise filtering and clear transmission of specific human voices, enabling smooth communication without removing headphones, thereby improving the user experience in complex environments and ensuring production safety.
[0040] It should be noted that the solution in this invention can be packaged in software as a voice pass-through function for noise-canceling headphones. This voice pass-through function is only used for the pass-through of the voices of other people matched with the voice feature template library, which is updated periodically according to the aforementioned time intervals. Alternatively, a conventional voice pass-through function that does not rely on the voice feature template library can also be designed for the user, that is, outputting all voice signals to the user. This function can also be turned on or off by the user via a switch button on a player (such as a mobile phone, tablet, etc.) or noise-canceling headphones; details will not be elaborated further.
[0041] As an example, the step of extracting human voice features from the first ambient sound signal using an attention mechanism to construct a human voice feature template library includes:
[0042] The system uses an attention mechanism to process multiple human voice signals in the first environmental sound signal, and identifies virtual human voice signals played by electronic devices; and performs distance detection and tracking on the remaining human voice signals, obtaining tracking results showing that the human voice signals are gradually moving away from the environmental area.
[0043] Virtual human voice signals played by electronic devices and human voice signals that are gradually moving away from the environment are removed, and a human voice feature template library is constructed based on the remaining human voice signals.
[0044] First, the attention mechanism focuses on analyzing the spectral characteristics, rhythm, and other aspects of the sound signal. For example, virtual voices played by electronic devices often differ from real human voices in terms of timbre consistency and background noise characteristics. The attention mechanism focuses on these differences to accurately distinguish real human voices from virtual voices played by electronic devices such as mobile phones, televisions, and radios. These virtual voice signals are then eliminated to prevent them from entering the voice feature template library, thus preventing irrelevant voice features from being mixed into the template library and affecting the accuracy of subsequent recognition of real human voices.
[0045] Then, distance detection and tracking are performed on the remaining human voice signals that do not belong to the virtual human voice signal. The tracking results show that the human voice signals are gradually moving away from the environmental area. Distance detection and tracking combines parameters such as changes in sound signal intensity and propagation delay to determine the distance relationship between the source of the human voice and the user. When a human voice signal is detected to be continuously weakening in intensity and the trend of change conforms to the characteristic of gradually moving away, it is marked as a human voice signal that is gradually moving away (such as the sound of vendors hawking outside the window, or the singing voice of a passerby). These types of human voice signals are those human voices in the environment that are highly mobile and have a low probability of subsequent interaction with the user. Therefore, they are not used to build the human voice feature template library.
[0046] Therefore, the human voice feature template library constructed through this embodiment has purer and more targeted features, which can significantly improve the accuracy of subsequent identification and matching of target human voices in the second ambient sound signal, thereby optimizing the effect of the entire headphone noise reduction control method and allowing users to hear the human voices they need to communicate with more clearly.
[0047] As an example, the preset time interval is determined based on ambient sound features extracted from the first ambient sound signal, including:
[0048] The ambient sound features representing the mobility of people are extracted from the first ambient sound signal. The ambient sound features include the switching frequency of human voice signals, the occurrence rate of new human voice features, and the frequency of sudden changes in the intensity of human voice signals.
[0049] A quantitative index for personnel mobility is established based on the environmental sound characteristics, and a time interval correction value is derived based on the quantitative index for personnel mobility and a preset correlation; wherein, the quantitative index for personnel mobility is positively correlated with the time interval correction value;
[0050] The preset time interval is obtained by subtracting the base time interval from the time interval correction value.
[0051] The switching frequency of human voice signals refers to the number of times different human voice signals alternate within a unit of time, reflecting the speed at which conversation partners change in the environment; the occurrence rate of new human voice features is the number of human voice features appearing for the first time within a unit of time, reflecting the situation of newly entered personnel; the frequency of sudden changes in human voice signal intensity is the number of times the intensity of human voice signal changes significantly within a unit of time, indirectly reflecting the movement status of personnel. Based on the above environmental sound characteristics, the activity level of personnel movement in the environment can be described in multiple dimensions.
[0052] By weighting the frequency of voice switching, the rate of new voices, and the frequency of intensity abrupt changes, these disparate features are integrated into a comprehensive quantitative indicator of population mobility, achieving accurate quantification of the degree of population movement. The calculation formula is, for example: ,in, As a quantitative indicator of personnel mobility, It is the standardized value of the switching frequency of human voice signals (i.e., the ratio of the actual switching frequency to the maximum frequency threshold). This is the standardized value of the occurrence rate of newly added human voice features (i.e., the ratio of the actual rate to the maximum rate threshold). It is the standardized value of the frequency of abrupt changes in the intensity of the human voice signal (i.e., the ratio of the actual frequency to the maximum frequency threshold). , , These are the weighting coefficients, and The default values are 0.4, 0.3, and 0.3 (which can be dynamically adjusted according to the actual scenario).
[0053] The preset correlation is a pre-defined rule that maps quantitative indicators of personnel mobility to correction values. For example, for each level increase in the quantitative indicator, the correction value increases accordingly, ensuring that the correction value increases as personnel mobility increases. An example is shown in the table below:
[0054]
[0055] The base time interval is a preset initial acquisition interval, such as 5 minutes. When there is high personnel mobility (large quantifiable index), the time interval correction value is large, and the preset time interval obtained by subtracting the two is short. This allows for more frequent acquisition of environmental sound signals, enabling rapid updates to the human voice feature template library to adapt to environments with frequent personnel changes. When there is low personnel mobility (small quantifiable index), the correction value is small, and the preset time interval is long. This ensures the effectiveness of the template library while reducing unnecessary acquisition and calculation, lowering headphone power consumption, and balancing the timeliness of the template library with resource consumption.
[0056] As an example, such as Figure 2 As shown, the construction of the human voice feature template library based on the remaining human voice signals includes:
[0057] The core feature parameter set is extracted from the remaining human voice signals, including fundamental frequency, formant, spectral envelope and sound intensity change rate, and the first set of human voice feature templates is constructed based on the core feature parameter set;
[0058] The feature variation range of each core feature parameter in the core feature parameter group is simulated under a preset scenario to obtain a variation feature parameter group. A second set of human voice feature templates is constructed based on the variation feature parameter group. The preset scenario refers to a scenario that causes short-term changes in voice features.
[0059] A human voice feature template library is constructed based on the first set of human voice feature templates and the second set of human voice feature templates.
[0060] The fundamental frequency in the core feature parameter set reflects the pitch of the sound, the formants characterize the resonance characteristics of the vocal tract, the spectral envelope characterizes the frequency energy distribution profile of the sound, and the rate of change of sound intensity characterizes the dynamic changes in sound strength. These core feature parameters characterize the basic acoustic features of the human voice from different dimensions. Based on the above core feature parameter set for the remaining human voice signals, corresponding human voice feature templates are constructed, resulting in the first set of human voice feature templates. These human voice feature templates are equivalent to the standard sound image of each voice. The human voice feature templates contain the typical value range and probability distribution of each core feature parameter, such as the normal fluctuation range of the fundamental frequency and the main distribution frequency of the formants. It can be understood that the human voice feature templates in the subsequent second set of human voice feature templates also contain the typical value range and probability distribution of the variant feature parameters corresponding to each core feature parameter, which will not be elaborated further.
[0061] Meanwhile, this invention also sets up several preset scenarios, focusing on situations that cause short-term changes in sound. For example, changes in throat condition after drinking water (or certain special beverages) may lead to a slight shift in the fundamental frequency; the resonant peak of the sound may temporarily change after coughing; and the rate of change in sound intensity may significantly increase during emotional excitement. By simulating the possible fluctuation range of each core parameter in these scenarios (such as a change in the fundamental frequency within ±5%, a shift in a certain resonant peak frequency of 100-200Hz, etc.), a set of variable characteristic parameters is generated. The specific processing procedure is roughly as follows:
[0062] (1) Classify and define the preset scenarios, and identify typical scenarios that cause short-term changes in voice characteristics, including changes in throat state after drinking water, vocal cord adjustment after a short cough, changes in vocal state caused by emotional excitement, and changes in oral resonance after eating. For each type of scenario, establish a correlation model between the scenario and the feature changes in advance through sample collection (preferably based on CNN) - collect human voice signals of different people in the above scenarios, extract core parameters such as fundamental frequency, formants, spectral envelope and sound intensity change rate, compare the parameter differences before and after the scenario, and statistically analyze the distribution of the change amplitude of feature parameters under various scenarios.
[0063] (2) Set scenario-based variation thresholds for each core feature parameter. For example, for the drinking water scenario: the fundamental frequency is allowed to fluctuate within ±8% of the original value, the first formant frequency offset is controlled within 50-150Hz, the energy peak attenuation rate of the spectral envelope does not change by more than 20%, and the fluctuation range of the sound intensity change rate is limited to ±15% of the original value. For the cough scenario: the short-term fluctuation range of the fundamental frequency is expanded to ±12%, the second formant frequency offset can reach 80-200Hz, and the energy proportion of the spectral envelope in the mid-to-high frequency band (2-4kHz) is allowed to change by ±25%. For the emotional fluctuation scenario: the variation range of the sound intensity change rate is widened to ±30%, and the dynamic change rate (change per second) of the fundamental frequency can be increased to 1.5 times the original value.
[0064] (3) Generate variant feature parameter sets corresponding to each preset scenario based on the original core feature parameter set and the above-mentioned variant threshold. The original core feature parameter set is adjusted within the set threshold using a random perturbation algorithm, generating multiple sets of samples that conform to the variation law of the scenario for each core parameter. For example, for the fundamental frequency parameter, five variant values with different offsets are generated within ±8% range; for the formant parameter, 3-4 sets of variant frequency values are generated according to the specific frequency offset range of the scenario. Several variant feature parameter sets are then obtained.
[0065] Based on the remaining sets of variation parameters corresponding to each human voice signal, a second set of human voice feature templates is constructed. This is equivalent to adding a dynamically changing sound image to each voice, making up for the limitation of the first set of templates which can only cope with stable states, and covering the characteristic fluctuations of sound caused by short-term physiological or state changes.
[0066] Finally, by integrating the first set of human voice feature templates and the second set of human voice feature templates, the final human voice feature template library is constructed.
[0067] As an example, the spectral features include the fundamental frequency to be matched, the formants to be matched, the spectral envelope to be matched, and the rate of change of sound intensity to be matched; then, the spectral features are matched and calculated with the human voice feature template library to determine the signals in the environmental sound signal that conform to the human voice frequency range, including:
[0068] Calculate the feature similarity between the spectral features and each human voice feature template in the human voice feature template library. The feature similarity is calculated by weighted Euclidean distance, wherein the weight coefficients of the fundamental frequency to be matched and the formant to be matched are higher than those of the spectral envelope to be matched and the rate of change of sound intensity to be matched.
[0069] When the feature similarity is greater than or equal to the similarity threshold, the corresponding signal is determined to conform to the human voice frequency range; otherwise, the corresponding signal is determined to not conform to the human voice frequency range. The similarity threshold is dynamically adjusted according to the intensity of environmental noise interference.
[0070] The extracted spectral features include the fundamental frequency to be matched, the formants to be matched, the spectral envelope to be matched, and the rate of change of sound intensity to be matched. These features correspond one-to-one with the core feature parameter set (fundamental frequency, formants, spectral envelope, and rate of change of sound intensity) extracted when constructing the human voice feature template library.
[0071] The similarity between the spectral feature and each voice feature template in the human voice feature template library is calculated using weighted Euclidean distance. The fundamental frequency determines the pitch of the sound, and the formants reflect the resonance characteristics of the vocal tract. These two are core identifiers that distinguish different human voices and have a greater impact on the accuracy of voice recognition. Therefore, the weight coefficients of the fundamental frequency and formants to be matched are set higher than those of the spectral envelope and the rate of change of sound intensity to be matched; that is, the accuracy of matching is improved by assigning higher weights.
[0072] When the feature similarity is greater than or equal to the similarity threshold, it indicates that the spectral features of the signal have a high degree of matching with the features in the human voice feature template library, and it is determined to be within the human voice frequency range; otherwise, it is determined to be non-matching.
[0073] Furthermore, the similarity threshold is dynamically adjusted based on the intensity of environmental noise interference. Specifically, in environments with high noise levels, human voice signals may be masked by noise, leading to a decrease in feature similarity. In this case, appropriately lowering the similarity threshold can prevent noise-affected human voices from being misidentified as non-human voices. Conversely, in environments with low noise levels, human voice features are clearer, allowing for a more appropriate increase in the threshold to reduce the likelihood of noise being misidentified as human voices. This dynamic adjustment mechanism enables the matching judgment to adapt to different environmental noise conditions, further enhancing the robustness of voice recognition.
[0074] like Figure 3 As shown, this embodiment of the invention also provides a noise-canceling headset, including a control unit 1001, a storage unit 1002, a speaker 1003, and at least one microphone 1004; the control unit 1001 implements the following steps by calling computer program code stored in the storage unit 1002:
[0075] Before receiving a noise reduction control command, the system collects a first ambient sound signal of the surrounding environment at a preset time interval using at least one microphone on the headphones. An attention mechanism is then used to extract human voice features from the first ambient sound signal to construct a human voice feature template library. The preset time interval is determined based on the ambient sound features representing the mobility of people extracted from the first ambient sound signal.
[0076] In response to the noise reduction control command, the second ambient sound signal of the environmental area is collected, the high-frequency interference and low-frequency background noise in the second ambient sound signal are removed to obtain the time domain signal, the time domain signal is converted into the frequency domain signal through Fourier transform, and the spectral features are extracted from the frequency domain signal.
[0077] The spectral features are matched with the human voice feature template library to determine the signals in the environmental sound signals that conform to the human voice frequency range and mark them as human voice signals; signals in other frequency ranges are marked as noise signals.
[0078] An active noise cancellation algorithm is used to generate a canceling sound wave that is out of phase with the noise signal, and the human voice signal is amplified. The canceling sound wave and the amplified human voice signal are then output through the speaker of the headphones.
[0079] It is understandable that noise-canceling headphones usually also include power supply components, shell structure, and other components, which will not be elaborated here.
[0080] As an example, an attention mechanism is used to extract human voice features from the first ambient sound signal to construct a human voice feature template library, including:
[0081] The system uses an attention mechanism to process multiple human voice signals in the first environmental sound signal, and identifies virtual human voice signals played by electronic devices; and performs distance detection and tracking on the remaining human voice signals, obtaining tracking results showing that the human voice signals are gradually moving away from the environmental area.
[0082] Virtual human voice signals played by electronic devices and human voice signals that are gradually moving away from the environment are removed, and a human voice feature template library is constructed based on the remaining human voice signals.
[0083] As an example, the preset time interval is determined based on ambient sound features extracted from the first ambient sound signal, including:
[0084] The ambient sound features representing the mobility of people are extracted from the first ambient sound signal. The ambient sound features include the switching frequency of human voice signals, the occurrence rate of new human voice features, and the frequency of sudden changes in the intensity of human voice signals.
[0085] A quantitative index for personnel mobility is established based on the environmental sound characteristics, and a time interval correction value is derived based on the quantitative index for personnel mobility and a preset correlation; wherein, the quantitative index for personnel mobility is positively correlated with the time interval correction value;
[0086] The preset time interval is obtained by subtracting the base time interval from the time interval correction value.
[0087] As an example, a vocal feature template library is constructed based on the remaining vocal signal, including:
[0088] The core feature parameter set is extracted from the remaining human voice signals, including fundamental frequency, formant, spectral envelope and sound intensity change rate, and the first set of human voice feature templates is constructed based on the core feature parameter set;
[0089] The feature variation range of each core feature parameter in the core feature parameter group is simulated under a preset scenario to obtain a variation feature parameter group. A second set of human voice feature templates is constructed based on the variation feature parameter group. The preset scenario refers to a scenario that causes short-term changes in voice features.
[0090] A human voice feature template library is constructed based on the first set of human voice feature templates and the second set of human voice feature templates.
[0091] As an example, the spectral features include the fundamental frequency to be matched, the formants to be matched, the spectral envelope to be matched, and the rate of change of sound intensity to be matched; then, the spectral features are matched and calculated with the human voice feature template library to determine the signals in the environmental sound signal that conform to the human voice frequency range, including:
[0092] Calculate the feature similarity between the spectral features and each human voice feature template in the human voice feature template library. The feature similarity is calculated by weighted Euclidean distance, wherein the weight coefficients of the fundamental frequency to be matched and the formant to be matched are higher than those of the spectral envelope to be matched and the rate of change of sound intensity to be matched.
[0093] When the feature similarity is greater than or equal to the similarity threshold, the corresponding signal is determined to conform to the human voice frequency range; otherwise, the corresponding signal is determined to not conform to the human voice frequency range. The similarity threshold is dynamically adjusted according to the intensity of environmental noise interference.
[0094] This invention also provides an electronic device comprising: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, wherein the computer program, when executed by the processor, implements the method as described in any of the foregoing embodiments.
[0095] This invention also provides a computer storage medium storing a computer program that can be executed by a processor to implement the methods described in any of the foregoing claims.
[0096] This invention also provides a computer program product comprising a computer program that can be executed by a processor to implement the method as described in any of the foregoing embodiments.
[0097] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0098] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A noise reduction control method for headphones, characterized in that, Includes the following steps: Before receiving a noise reduction control command, the system collects a first ambient sound signal of the surrounding environment at a preset time interval using at least one microphone on the headphones. An attention mechanism is then used to extract human voice features from the first ambient sound signal to construct a human voice feature template library. The preset time interval is determined based on the ambient sound features representing the mobility of people extracted from the first ambient sound signal. In response to the noise reduction control command, the second ambient sound signal of the environmental area is collected, the high-frequency interference and low-frequency background noise in the second ambient sound signal are removed to obtain the time domain signal, the time domain signal is converted into the frequency domain signal through Fourier transform, and the spectral features are extracted from the frequency domain signal. The spectral features are matched with the human voice feature template library to determine the signals in the environmental sound signals that conform to the human voice frequency range and mark them as human voice signals; signals in other frequency ranges are marked as noise signals. An active noise cancellation algorithm is used to generate a canceling sound wave that is out of phase with the noise signal, and the human voice signal is amplified. The canceling sound wave and the amplified human voice signal are then output through the speaker of the headphones. An attention mechanism is used to extract human voice features from the first ambient sound signal to construct a human voice feature template library, including: The system uses an attention mechanism to process multiple human voice signals in the first environmental sound signal, and identifies virtual human voice signals played by electronic devices; and performs distance detection and tracking on the remaining human voice signals, obtaining tracking results showing that the human voice signals are gradually moving away from the environmental area. Virtual human voice signals played by electronic devices and human voice signals that are gradually moving away from the environment are removed, and a human voice feature template library is constructed based on the remaining human voice signals. The preset time interval is determined based on the ambient sound features extracted from the first ambient sound signal, including: The environmental sound features characterizing personnel mobility are extracted from the first environmental sound signal. These environmental sound features include the switching frequency of human voice signals, the occurrence rate of newly added human voice features, and the frequency of sudden changes in human voice signal intensity. The switching frequency of human voice signals refers to the number of times different human voice signals alternate within a unit of time, reflecting the speed of change in the people conversing in the environment. The occurrence rate of newly added human voice features is the number of human voice features appearing for the first time within a unit of time, reflecting the situation of newly entering personnel. The frequency of sudden changes in human voice signal intensity is the number of times the intensity of the human voice signal changes significantly within a unit of time, indirectly reflecting the movement status of personnel. A quantitative index for personnel mobility is established based on the environmental sound characteristics, and a time interval correction value is derived based on the quantitative index for personnel mobility and a preset correlation; wherein, the quantitative index for personnel mobility is positively correlated with the time interval correction value; The preset time interval is obtained by subtracting the base time interval from the time interval correction value.
2. The noise reduction control method for headphones according to claim 1, characterized in that: A human voice feature template library is constructed based on the remaining human voice signals, including: The core feature parameter set is extracted from the remaining human voice signals, including fundamental frequency, formant, spectral envelope and sound intensity change rate, and the first set of human voice feature templates is constructed based on the core feature parameter set; The feature variation range of each core feature parameter in the core feature parameter group is simulated under a preset scenario to obtain a variation feature parameter group. A second set of human voice feature templates is constructed based on the variation feature parameter group. The preset scenario refers to a scenario that causes short-term changes in voice features. A human voice feature template library is constructed based on the first set of human voice feature templates and the second set of human voice feature templates.
3. The noise reduction control method for headphones according to claim 2, characterized in that: The spectral features include the fundamental frequency to be matched, the formant to be matched, the spectral envelope to be matched, and the rate of change of sound intensity to be matched; then, the spectral features are matched with the human voice feature template library to determine the signals in the environmental sound signal that conform to the human voice frequency range, including: Calculate the feature similarity between the spectral features and each human voice feature template in the human voice feature template library. The feature similarity is calculated by weighted Euclidean distance, wherein the weight coefficients of the fundamental frequency to be matched and the formant to be matched are higher than those of the spectral envelope to be matched and the rate of change of sound intensity to be matched. When the feature similarity is greater than or equal to the similarity threshold, the corresponding signal is determined to conform to the human voice frequency range; otherwise, the corresponding signal is determined to not conform to the human voice frequency range. The similarity threshold is dynamically adjusted according to the intensity of environmental noise interference.
4. A noise-canceling headphone, characterized in that, It includes a control unit, a storage unit, a speaker, and at least one microphone; the control unit performs the following steps by calling computer program code stored in the storage unit: Before receiving a noise reduction control command, the system collects a first ambient sound signal of the surrounding environment at a preset time interval using at least one microphone on the headphones. An attention mechanism is then used to extract human voice features from the first ambient sound signal to construct a human voice feature template library. The preset time interval is determined based on the ambient sound features representing the mobility of people extracted from the first ambient sound signal. In response to the noise reduction control command, the second ambient sound signal of the environmental area is collected, the high-frequency interference and low-frequency background noise in the second ambient sound signal are removed to obtain the time domain signal, the time domain signal is converted into the frequency domain signal through Fourier transform, and the spectral features are extracted from the frequency domain signal. The spectral features are matched with the human voice feature template library to determine the signals in the environmental sound signals that conform to the human voice frequency range and mark them as human voice signals; signals in other frequency ranges are marked as noise signals. An active noise cancellation algorithm is used to generate a canceling sound wave that is out of phase with the noise signal, and the human voice signal is amplified. The canceling sound wave and the amplified human voice signal are then output through the speaker of the headphones. An attention mechanism is used to extract human voice features from the first ambient sound signal to construct a human voice feature template library, including: The system uses an attention mechanism to process multiple human voice signals in the first environmental sound signal, and identifies virtual human voice signals played by electronic devices; and performs distance detection and tracking on the remaining human voice signals, obtaining tracking results showing that the human voice signals are gradually moving away from the environmental area. Virtual human voice signals played by electronic devices and human voice signals that are gradually moving away from the environment are removed, and a human voice feature template library is constructed based on the remaining human voice signals. The preset time interval is determined based on the ambient sound features extracted from the first ambient sound signal, including: The environmental sound features characterizing personnel mobility are extracted from the first environmental sound signal. These environmental sound features include the switching frequency of human voice signals, the occurrence rate of newly added human voice features, and the frequency of sudden changes in human voice signal intensity. The switching frequency of human voice signals refers to the number of times different human voice signals alternate within a unit of time, reflecting the speed of change in the people conversing in the environment. The occurrence rate of newly added human voice features is the number of human voice features appearing for the first time within a unit of time, reflecting the situation of newly entering personnel. The frequency of sudden changes in human voice signal intensity is the number of times the intensity of the human voice signal changes significantly within a unit of time, indirectly reflecting the movement status of personnel. A quantitative index for personnel mobility is established based on the environmental sound characteristics, and a time interval correction value is derived based on the quantitative index for personnel mobility and a preset correlation; wherein, the quantitative index for personnel mobility is positively correlated with the time interval correction value; The preset time interval is obtained by subtracting the base time interval from the time interval correction value.
5. An electronic device comprising: At least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, characterized in that: when the computer program is executed by the processor, it implements the method as described in any one of claims 1-3.
6. A computer storage medium, characterized in that: The computer storage medium stores a computer program that can be executed by a processor to implement the method as described in any one of claims 1-3.
7. A computer program product, characterized in that: The computer program product includes a computer program that can be executed by a processor to implement the method as described in any one of claims 1-3.
Citation Information
Patent Citations
Bluetooth earphone call noise reduction method and device based on AI human voice extraction
CN118890575A