Self-adaptive earphone active noise reduction method and device
By using feedback and feedforward microphones in headphones, combined with temporal saliency maps and one-dimensional convolutional neural networks, the noise reduction algorithm is adaptively adjusted, solving the problem of insufficient noise suppression in existing headphones in changing environments, and achieving more efficient noise cancellation and a better listening experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN LIGHKEEP CO LTD
- Filing Date
- 2026-02-03
- Publication Date
- 2026-05-01
AI Technical Summary
Existing active noise-canceling headphone algorithms cannot effectively cope with the variability and diversity of environmental noise, resulting in their inability to effectively suppress noise in certain situations.
The system receives in-ear noise signals from the headset terminal via a mobile terminal, collects signals using feedback and feedforward microphones, generates a time-domain saliency map, identifies specific acoustic events, adjusts the initial antiphase sound wave to generate the target antiphase sound wave, and performs precise noise reduction processing by combining a one-dimensional convolutional neural network and filters.
It achieves real-time adaptive noise cancellation for headphones in complex or changing noise environments, improves noise cancellation performance, optimizes noise cancellation effect, and enhances the user's listening experience.
Smart Images

Figure CN121967954A_ABST
Abstract
Description
An adaptive headphone active noise cancellation method and device Technical Field
[0001] This invention belongs to the technical field of audio processing, and particularly relates to an adaptive headphone active noise cancellation method and device. Background Technology
[0002] With the continuous advancement of technology and the improvement of people's living standards, headphones, as an important audio device, have gradually become integrated into people's daily lives. Especially in noisy environments, active noise cancellation (ANC) technology in headphones has received increasing attention. Active noise cancellation technology cancels out noise by generating sound waves that are out of phase with external noise, thereby improving the user's listening experience.
[0003] Existing active noise-canceling headphones typically rely on fixed noise-canceling algorithms, which are often based on preset noise characteristics. However, the variability and diversity of environmental noise mean that these fixed algorithms cannot effectively suppress noise in certain situations. Summary of the Invention
[0004] In view of this, embodiments of the present invention provide an adaptive headphone active noise cancellation method and apparatus to solve the technical problem that the variability and diversity of environmental noise makes it impossible for a fixed algorithm to effectively suppress noise in certain situations.
[0005] A first aspect of this invention provides an adaptive active noise cancellation method for headphones, the method comprising: a mobile terminal receiving an in-ear noise signal sent by an earphone terminal; wherein the in-ear noise signal is a noise sound wave collected by a feedback microphone after noise reduction processing based on an initial anti-phase sound wave; determining the sound wave regularity characteristics corresponding to the in-ear noise signal according to a time-domain saliency map corresponding to the in-ear noise signal; wherein the time-domain saliency map is used to identify specific acoustic events; and adjusting the initial anti-phase sound wave according to the sound wave regularity characteristics to obtain a target anti-phase sound wave, and sending it to the earphone terminal for noise reduction processing.
[0006] Furthermore, before the step of the mobile terminal receiving the in-ear noise signal sent by the earphone terminal, the method further includes: collecting ambient sound waves based on a feedforward microphone and generating an initial anti-phase sound wave; performing noise reduction processing on the initial anti-phase sound wave; collecting in-ear sound waves based on a feedback microphone during the noise reduction processing; removing audio data from the in-ear sound waves to obtain the in-ear noise signal; and sending the in-ear noise signal to the mobile terminal.
[0007] Further, the step of determining the acoustic wave regularity characteristics corresponding to the intra-ear noise signal based on the time-domain saliency map of the intra-ear noise signal includes: Step A1: Calculating the instantaneous amplitude corresponding to the discrete intra-ear noise signal based on the Hilbert transform; Step A2: Calculating the time-domain saliency map based on the instantaneous amplitude; the time-domain saliency map is used to identify specific acoustic events; Step A3: Extracting the highest peak point in the time-domain saliency map; Step A4: Constructing an adaptive resonant core based on the highest peak point; Step A5: Projecting the intra-ear noise signal onto the constructed optimal resonant core to obtain the target mode components corresponding to each discrete intra-ear noise signal; Step A6: Using the target mode components corresponding to each discrete intra-ear noise signal as the current noise feature.
[0008] Further, step A5 includes: step A51: projecting the intra-ear noise signal onto the constructed optimal resonant core to obtain the initial mode components corresponding to each discrete intra-ear noise signal; step A52: subtracting the initial mode components corresponding to each discrete intra-ear noise signal from the discrete intra-ear noise signal to obtain the current residual signal; step A53: using the current residual signal as the discrete intra-ear noise signal, iteratively executing steps A1 to A53 until a termination condition is met to obtain the target mode components corresponding to each discrete intra-ear noise signal; wherein, the termination condition includes the total energy of the residual signal being lower than a threshold.
[0009] Further, step A2 includes: substituting the instantaneous amplitude into the time-domain saliency function to obtain the time-domain saliency map; the time-domain saliency function is:
[0010] in, This represents the time-domain saliency map, where n represents the time point corresponding to the discrete intraocular noise signal. This represents a time window centered at n. Represents variance. This represents the instantaneous amplitude of the nth discrete signal in the discrete intra-auricular noise signal. This represents the instantaneous amplitude of the m-th discrete signal in the discrete intraocular noise signal, where m represents a time point within the time window.
[0011] Further, step A4 includes: substituting the highest peak point into a preset function to obtain the adaptive resonant core; wherein the preset function is:
[0012] in, This represents an adaptive resonant core. Indicates the highest peak point. Indicates the width of the Gaussian envelope. , Indicates a point in time The corresponding signal value at that location, Indicates the resonant frequency. Indicates the initial phase.
[0013] Further, the step of adjusting the initial antiphase sound wave according to the sound wave characteristics to obtain the target antiphase sound wave and sending it to the headphone terminal for noise reduction processing includes: inputting multiple target modal components into a one-dimensional convolutional neural network to obtain the sound category and confidence level corresponding to each of the multiple target modal components; wherein, the sound category includes human voice, horn sound, wind sound, water sound, knocking sound, and keyboard sound; obtaining a preset cancellation weight mask corresponding to each sound category; multiplying the preset cancellation weight mask, confidence level, and in-ear noise signal to obtain a cancellation signal; processing the cancellation signal through a filter to obtain the target antiphase sound wave; and sending the corrected antiphase sound wave to the headphone terminal for noise reduction processing.
[0014] A second aspect of this invention provides an adaptive headphone active noise cancellation device, comprising: an acquisition unit, configured to acquire research and development data text to be identified, and to segment the research and development data text into sentences to obtain multiple sentences to be identified; an extraction unit, configured to extract the number of entity words and the number of logical connectors in the sentences to be identified; wherein, entity words include at least one of personal names, place names, organizations, dates, and professional terms; a calculation unit, configured to calculate the topic coherence index, positional influence index, and format influence index of the sentences to be identified; wherein, the topic coherence index refers to the degree of correlation between the sentences and the topic; and an annotation unit, configured to extract key sentences based on the number of entity words, the number of logical connectors, the topic coherence index, the positional influence index, and the format influence index, and to annotate the key sentences.
[0015] A third aspect of the present invention provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps in the adaptive headphone active noise cancellation method described in the first aspect.
[0016] A fourth aspect of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the adaptive headphone active noise cancellation method described in the first aspect.
[0017] The beneficial effects of this invention compared to existing technologies are as follows: It receives in-ear noise signals sent by the earphone terminal via a mobile terminal and generates a time-domain saliency map based on these signals to identify specific acoustic events. This design allows the earphone to analyze changes in ambient noise in real time, ensuring that noise reduction processing can dynamically adapt to different noise environments. Compared to traditional fixed noise reduction algorithms, this real-time adaptability significantly improves noise reduction performance, especially in complex or changing noise environments. By determining the sound wave characteristics corresponding to the in-ear noise signals, this method can more accurately understand the nature of the noise and adjust the initial anti-phase sound wave accordingly to generate the target anti-phase sound wave. This process optimizes the noise reduction effect, effectively eliminating noise within a specific frequency range and improving the user's listening experience in various usage scenarios. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 shows a schematic flowchart of an adaptive headphone active noise cancellation method provided by the present invention; Figure 2 shows a schematic diagram of an adaptive headphone active noise cancellation device provided by an embodiment of the present invention; Figure 3 shows a schematic diagram of a terminal device provided by an embodiment of the present invention. Detailed Implementation
[0020] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of the invention. However, those skilled in the art will understand that the invention can be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods are omitted so as not to obscure the description of the invention with unnecessary detail.
[0021] This invention provides an adaptive headphone active noise cancellation method and apparatus to address the technical problem that the variability and diversity of environmental noise makes it difficult for fixed algorithms to effectively suppress noise in certain situations.
[0022] First, this invention provides an adaptive active noise cancellation method for headphones. Please refer to Figure 1, which shows a schematic flowchart of an adaptive active noise cancellation method for headphones provided by this invention. As shown in Figure 1, the adaptive active noise cancellation method for headphones may include the following steps: Step 101: The mobile terminal receives an in-ear noise signal sent by the headphone terminal; wherein, the in-ear noise signal is a noise sound wave collected by a feedback microphone after noise reduction processing based on an initial anti-phase sound wave; the in-ear noise signal is not the original external noise, but rather the noise that still leaks into the user's ear after the first round of noise reduction. This can be understood as the residual error of noise reduction. This in-ear noise signal is the data basis for all subsequent analysis and decisions.
[0023] The initial out-of-phase sound wave refers to the out-of-phase sound wave currently being used by the headphones, which has not been adjusted in this round. It is the result of the last adjustment.
[0024] The feedback microphone, a key piece of hardware for active noise cancellation, is located inside the headphones, close to the user's eardrum. Its function is to collect the sound waves that actually reach the eardrum after initial noise cancellation. Therefore, the signal it collects perfectly reflects the true state of the current noise cancellation effect.
[0025] As an optional embodiment of the present invention, before step 101, steps B1 to B5 are included: Step B1: Based on the ambient sound waves collected by the feedforward microphone, an initial anti-phase sound wave is generated; the feedforward microphone is located outside the earphone and is used to collect external ambient noise that has not yet entered the earphone and has not been processed. The initial anti-phase sound wave is a mirror sound wave calculated in real time based on the feedforward microphone signal. The principle is that sound waves cancel each other out, generating a sound wave with the same amplitude but opposite phase (180 degrees out of phase) to the ambient noise, and the two can cancel each other out after being superimposed.
[0026] Step B2: Noise reduction is performed using the initial anti-phase sound wave; inside the ear canal, ambient noise from the outside encounters and interferes with the initial anti-phase sound wave played by the headphones, thereby physically reducing or eliminating the original noise. The user has already experienced a certain degree of noise reduction effect at this point.
[0027] Step B3: During the noise reduction process, the sound waves inside the ear are collected based on the feedback microphone; the feedback microphone is located inside the earphone and is used to monitor the final sound effect that actually reaches the eardrum after the first round of noise reduction.
[0028] The sound waves inside the ear are a mixed signal, including: residual ambient noise (parts that are not completely canceled out by the initial out-of-phase sound waves), the audio data that the user is listening to (such as music, podcasts), and the initial out-of-phase sound waves played by the headphones themselves.
[0029] Collect raw data that best reflects the actual effect of current noise reduction. This step is the cornerstone of achieving high-precision adaptive noise reduction because it reflects the inadequacies of noise reduction in the real world.
[0030] Step B4: Remove the audio data from the intra-ear sound waves to obtain the intra-ear noise signal; ensure that the signal sent to the mobile terminal does not contain the music / voice the user wants to hear, avoiding interference from these contents in the identification and analysis of noise features. This allows subsequent time-domain saliency mapping analysis to accurately focus on the noise itself.
[0031] Step B5: Send the in-ear noise signal to the mobile terminal.
[0032] In the embodiments corresponding to steps B1 to B5, a combination of feedforward and feedback is used in the actual noise reduction process to achieve precise capture and processing of in-ear noise. This method can not only perform preliminary noise reduction based on the external environment, but also achieve dynamic adjustment by monitoring in-ear sound waves in real time, thereby improving the effectiveness of noise reduction and user experience.
[0033] Step 102: Determine the acoustic wave regularity characteristics corresponding to the intra-ear noise signal based on the time-domain saliency map corresponding to the intra-ear noise signal; wherein, the time-domain saliency map is used to identify specific acoustic events; the time-domain saliency map is used to identify which parts on the time axis are prominent, abnormal or important.
[0034] Specific acoustic events are the targets sought in time-domain saliency maps, referring to non-stationary, irregular, and sudden noises, such as subway announcements, keyboard typing, and car horns. In contrast, stationary noises, such as air conditioner noise and engine roar, can usually be handled well by traditional FFT (Fast Fourier Transform) spectrum analysis.
[0035] Acoustic wave regularity characteristics refer to the quantitative parameters extracted from saliency maps that can be used to guide the generation of antiphase acoustic waves.
[0036] Specifically, step 102 includes steps A1 to A6: Step A1: Calculate the instantaneous amplitude of the discrete intra-ear noise signal based on the Hilbert transform; the original intra-ear noise signal is a waveform that varies with time. The Hilbert transform is a mathematical tool that can construct an analytic signal for a real signal, thereby enabling accurate calculation of the instantaneous amplitude and instantaneous frequency of the signal at any point in time.
[0037] This step converts the signal from a waveform perspective to an energy envelope perspective. The curve formed by the instantaneous amplitude clearly depicts the fluctuations in signal energy, where high amplitude points often correspond to important acoustic events (such as knocking sounds or whistling sounds).
[0038] Assume the discrete-time signal acquired from the feedforward microphone is Where n = 1, 2, ..., N. The Hilbert transform calculation process is as follows: .in This represents the Hilbert transform.
[0039] Step A2: Based on the instantaneous amplitude, calculate the temporal saliency map; the temporal saliency map is used to identify specific acoustic events; based on the instantaneous amplitude sequence obtained in the previous step, through further processing (e.g., calculating its envelope, or comparing it with the local mean / variance), a temporal saliency map is generated. This map can be understood as a curve, the position of its peak on the time axis marking the moment when the specific acoustic event occurred.
[0040] The goal is to visualize / quantify the temporal characteristics of the signal, generating a heat map that directly indicates at which time points noteworthy acoustic events occurred.
[0041] Specifically, step A2 includes: substituting the instantaneous amplitude into the time-domain significance function to obtain the time-domain significance map; the time-domain significance function is:
[0042] in, This represents the time-domain saliency map, where n represents the time point corresponding to the discrete intraocular noise signal. This represents a time window centered at n. Represents variance. This represents the instantaneous amplitude of the nth discrete signal in the discrete intra-auricular noise signal. This represents the instantaneous amplitude of the m-th discrete signal in the discrete intraocular noise signal, where m represents a time point within the time window.
[0043] Part One: (Instantaneous amplitude) This represents the absolute intensity or loudness of the signal at time point n. For a sound event to be considered significant, it must first be loud enough. This is the most basic criterion. A knocking sound is certainly more significant than a faint breathing sound. This part ensures that high-energy points are given priority.
[0044] Part Two The absolute value of the instantaneous amplitude derivative is crucial for identifying abrupt changes or transients. It measures the drastic change in the signal envelope (instantaneous amplitude). It is specifically designed to capture the leading edge (attack) of a signal. A sudden noise (such as a snap or ding) is characterized by its amplitude rapidly climbing from zero to a peak value in a very short time. The steeper this climb (i.e., the larger the derivative), the higher the value of |dA[n] / dn|. For stationary noise (such as the sound of an air conditioner), the amplitude changes slowly, and the derivative is close to zero, so this part suppresses its significance.
[0045] Part Three: (The reciprocal of the local variance) Here, a local time window is introduced. And calculate the variance of the instantaneous amplitude within that window. If a time window The background itself is full of irregular fluctuations (i.e., the background noise is chaotic), so even if a new event occurs, its prominence will be weakened by the chaos of the background. For example, in a bustling market, a sudden shout might not be very noticeable. If the background is very calm and quiet (like in a library), then even a small sound (like turning a page) will stand out.
[0046] Therefore, variance This serves as a measure of background clutter. A cluttered background (higher variance) results in a larger denominator, thus suppressing the overall significance score. A more stable background (lower variance) results in a smaller denominator, thus amplifying the significance of the event. The +1 is added to prevent the denominator from being zero, ensuring computational stability.
[0047] Consider the three parts together: significance = (Absolute signal strength) × (Signal change intensity) × (1 / Background clutter). It will not misjudge a simple, loud but slowly changing sound (such as a distant, continuous truck roar) as the most significant event. This is because it requires that the loud sound must be a sudden burst (Part 2), and that the background environment in which it occurs cannot be too cluttered (Part 3).
[0048] This function can precisely highlight the sudden, specific acoustic events that are of real interest on the timeline, while suppressing the effects of stationary noise and highly non-stationary background noise.
[0049] Step A3: Extract the highest peak point from the time-domain saliency map; on the generated saliency map, find the peak point corresponding to the most prominent and energetic event. The time coordinate and amplitude of this point represent the primary problem that needs to be solved most in the current residual noise.
[0050] This achieves a reduction in complexity and prioritizes key events. Instead of processing all minor events simultaneously, the system prioritizes identifying and resolving the noise that has the greatest impact on the user experience.
[0051] Peak point The calculation method is as follows: . Represents the time-domain saliency map. Find the current saliency map. The highest peak point in This marks the starting anchor point of the Kth major acoustic event.
[0052] Step A4: Based on the highest peak point, construct an adaptive resonant core; the system analyzes a small segment of the signal centered at the highest peak point. The shape of this resonant core (e.g., a decaying oscillating waveform) will match the inherent vibrational mode of that particular noise event. For example, if it is a clanging sound, the resonant core might be a waveform with a specific frequency and decay rate near that time point.
[0053] Specifically, step A4 includes: substituting the highest peak point into a preset function to obtain the adaptive resonant core; wherein the preset function is:
[0054] in, This represents an adaptive resonant core. Indicates the highest peak point. Indicates the width of the Gaussian envelope. , Indicates a point in time The corresponding signal value at that location, Indicates the resonant frequency. Indicates the initial phase.
[0055] Part 1: Gaussian Envelope It is a bell-shaped curve (Gaussian function), with the highest peak point... Centered on the resonant core, its function is to give the resonant core a start and an end, making it a local rather than an infinitely extending filter. This represents the current time point n and the peak point. The distance.
[0056] The most crucial adaptive parameter in the entire function is the width of the Gaussian envelope. It directly determines the range of influence of this probe on the time axis. Here, the amplitude envelope is used in... The sharpness of an event is estimated by the second derivative (curvature) at a given point. A very sharp pulse (like a click) will have a large curvature, resulting in a very narrow... This ensures that only the event itself is captured without any trailing.
[0057] The instantaneous amplitude A[n] at the peak point The second derivative at the peak point. The second derivative measures the curvature or bending of a function. At this point, the first derivative is zero, while the second derivative describes whether the peak is sharp or flat. A sharp peak (like a snap) has a large curvature at its apex (i.e., a large absolute value of the second derivative). According to the formula, this leads to... It becomes smaller. A flat peak (like the top of a whistle) with very small curvature (small absolute value of the second derivative), thus leading to It gets bigger.
[0058] therefore, The logic is: for sudden, transient events (sharp peaks). It will automatically shrink, generating a temporally concentrated and narrow resonant core to precisely match the transient characteristics of the event. This is for continuous, steady-state events (flat peaks). It will automatically increase in size, generating a resonant core with a wider time span to match its longer duration. It is an adjustable scaling factor used to globally control the width of the resonant core.
[0059] Part Two: Cosine Carrier It is an oscillating component that determines the frequency characteristics of the noise events that the resonant core must match.
[0060] resonant frequency From peak point The local spectrum of nearby signals (through short-time Fourier transform, etc.) is used to lock onto the dominant frequency of the event. Resonant frequency. and initial phase This is obtained through a local parsing match: Within a short time window nearby, for signal s[n] and kernel function Projection is performed, and the resonant frequency is fine-tuned using gradient descent. and initial phase , making Maximize the correlation with local signals. .
[0061] initial phase Used for fine-tuning to ensure that the waveform of the resonant core is optimally aligned with the waveform of the noise event.
[0062] Combining the two parts, the Gaussian envelope determines the temporal locality of the resonant core (when it occurs and how long it lasts). The cosine carrier determines the frequency response of the resonant core (what tone it is). The product of the two... This creates a probe that is highly matched to the target event in both time and frequency.
[0063] This preset function can dynamically generate the most effective extraction tool based on the unique physical properties (sharpness, dominant frequency) of the noise event represented by each peak point.
[0064] Step A5: Project the intra-ear noise signal onto the constructed optimal resonant core to obtain the target mode components corresponding to the discrete intra-ear noise signals. Projection is a mathematical operation, essentially measuring how similar the original signal is to the custom resonant core. The portion similar in shape to the resonant core (i.e., the specific noise event) is separated from the original signal. This separated portion is the target mode component.
[0065] Specifically, step A5 includes steps A51 to A53: Step A51: Project the intra-ear noise signal onto the constructed optimal resonant core to obtain the initial modal components corresponding to the discrete intra-ear noise signals; this is the first step of each loop. Using the resonant core constructed based on the highest peak point in the current loop, extract the component that best matches it from the current signal (the original signal in the first loop, and the residual signal in subsequent loops). This extracted component is the initial modal component, which represents the acoustic event with the highest energy and most significant value in the current signal.
[0066] The k-th initial mode component is extracted by projecting the signal onto the constructed optimal resonant core. . The calculation process is as follows:
[0067]
[0068] Step A52: Subtract the discrete intra-ear noise signal from its corresponding initial mode component to obtain the current residual signal; subtract the extracted initial mode component from the original signal (or the residual signal from the previous round). This yields the signal remaining after removing the identified events. This can be understood as: Residual signal = Original signal - Extracted most significant event. This residual signal contains all remaining noise with lower energy or less distinct characteristics.
[0069] Step A53: Treat the current residual signal as a discrete intra-ear noise signal, and iteratively execute steps A1 to A53 until the termination condition is met to obtain the target mode components corresponding to each discrete intra-ear noise signal; wherein, the termination condition includes the total energy of the residual signal being lower than a threshold.
[0070] The current residual signal obtained from the calculation in this round is used as the input signal for the next round of the cycle (i.e., the new in-ear noise signal).
[0071] For this new, cleaner residual signal, repeat the entire process (① to ⑤): ① Calculate its instantaneous amplitude. ② Generate its own time-domain saliency spectrum (at this point, the highest peak in the spectrum is the second most significant event in the original signal). ③ Extract the highest peak from its spectrum. ④ Construct a new, tailor-made resonant kernel for this second most significant event. ⑤ Extract the second initial mode component from this residual signal. Repeat the above process continuously, extracting the most significant mode component from the remaining residual signal each time.
[0072] The loop will not continue indefinitely. When the total energy of the residual signal falls below a preset threshold, the loop stops. This means that the remaining signal is very weak and no longer contains acoustic events with actual noise reduction value (it may just be background noise or random noise).
[0073] After N iterations, what is obtained is no longer a single target modal component, but a set: {IMF1, IMF2, IMF3, ..., IMFn, final residual}.
[0074] Each IMF (Intrinsic Mode Function) is a target mode component, which represents all the specific acoustic events in the original intra-ear noise signal, from the most important to the least important.
[0075] For example, the final decomposition result might be: IMF1: a brief, high-energy keyboard click.
[0076] IMF2: A periodic, low-energy mouse click sound.
[0077] IMF3: The undulating envelope of distant colleagues' conversations.
[0078] Residual: The low-frequency humming sound of the air conditioner (energy is already very low).
[0079] In the embodiments corresponding to steps A51 to A53, these three sub-steps constitute an iterative processing mechanism that gradually acquires the target modal components of the in-ear noise signal by continuously projecting, subtracting, and updating the residual signal. This process not only improves the accuracy of noise feature extraction but also ensures dynamic adaptation to changing noise environments during noise reduction. The ultimate goal is to enable headphones to effectively reduce interference under various noise conditions and provide a clearer sound experience.
[0080] Step A6: Use the target mode components corresponding to each of the discrete intra-ear noise signals as the current noise features.
[0081] The target modal component separated in the previous step, or its characteristic parameters (such as energy, dominant frequency, and duration) after simple calculation, are formally defined as the acoustic wave regularity characteristics (i.e., the current noise characteristics).
[0082] In the embodiments corresponding to steps A1 to A6, these steps constitute a systematic process, from acquiring and processing the in-ear noise signal to feature extraction, ultimately generating noise features that can be used for adaptive noise reduction. This method, through the combination of temporal saliency maps and resonant kernels, helps improve the noise reduction performance of headphones in complex noise environments and enhances the user experience.
[0083] Step 103: Based on the characteristics of the sound wave, adjust the initial antiphase sound wave to obtain the target antiphase sound wave, and send it to the headphone terminal for noise reduction processing.
[0084] Based on the extracted sound wave patterns, a new target reverse sound wave is generated and sent back to the headphones.
[0085] Specifically, step 103 includes steps 1031 to 1035: Step 1031: Input multiple target modal components into a one-dimensional convolutional neural network to obtain the sound category and its confidence level corresponding to each target modal component; wherein, the sound category includes human voice, horn sound, wind sound, water sound, knocking sound, and keyboard sound; a one-dimensional convolutional neural network (1D-CNN) is a deep learning model specifically designed for processing sequential data (such as audio signals, time series). It can automatically learn and extract features from data, making it very suitable for sound classification. The one-dimensional convolutional neural network is a traditional technique and will not be elaborated upon here.
[0086] A one-dimensional convolutional neural network outputs a specific label (such as a human voice, a horn sound, a keyboard sound, etc.) and a confidence score (a value between 0 and 1, representing the model's confidence in its classification result).
[0087] The system not only knows that there is a sudden high-energy event, but also knows what the event is. This is the first step in achieving context-aware noise reduction.
[0088] Step 1032: Obtain the preset cancellation weight mask corresponding to each sound category; this is a predefined strategy library. The system has pre-set a cancellation weight mask for each known sound category (such as human voice, horn sound).
[0089] The cancellation weight mask can be understood as a function or vector that specifies at which frequencies and with what intensity antiphase sound waves need to be generated in order to cancel out this type of sound.
[0090] Example: A mask for human voices might primarily focus on the mid-frequency range and have a low overall weight (because users might want to hear some of the human voice); while a mask for keyboard sounds might cover the high-frequency range and have a high weight (aimed at completely eliminating it).
[0091] For important warning sounds (such as horns) or sounds that users might want to hear (such as human voices), the system will adopt a more conservative strategy; for purely interfering noises (such as wind noise or keyboard noise), it will adopt an aggressive strategy.
[0092] For example, the preset offset weight mask is shown in Table 1 below: Table 1:
[0093] Step 1033: Multiply the preset cancellation weight mask, confidence level, and in-ear noise signal to obtain the cancellation signal; Step 1034: Process the cancellation signal through a filter to obtain the target reverse sound wave; The cancellation signal obtained in the previous step is not perfect and needs to be adjusted by a filter. The filter is a phase correction filter to ensure that the reverse sound wave is accurately aligned with the original noise in the ear. This ensures that the generated reverse sound wave is physically effective and can achieve the best destructive interference effect.
[0094] Step 1035: Send the corrected antiphase sound wave to the headphone terminal to perform noise reduction processing.
[0095] Once the headphones receive new and improved noise cancellation instructions, they play them through the speakers, thereby achieving precise and intelligent cancellation of specific acoustic events in the current environment.
[0096] In the embodiments corresponding to steps 1031 to 1035, through the above steps, the system can flexibly and accurately adjust the reverse sound wave based on the regular characteristics of sound waves to achieve a highly efficient active noise reduction effect. Each step uses advanced algorithms and technologies to ensure the real-time performance and effectiveness of the noise reduction process.
[0097] In the embodiments corresponding to steps 101 to 103, the mobile terminal receives the in-ear noise signal sent by the headphone terminal, and generates a time-domain saliency map based on these signals to identify specific acoustic events. This design enables the headphones to analyze changes in ambient noise in real time, ensuring that noise reduction processing can dynamically adapt to different noise environments. Compared to traditional fixed noise reduction algorithms, this real-time adaptability significantly improves noise reduction performance, especially in complex or changing noise environments. By determining the sound wave characteristics corresponding to the in-ear noise signal, this method can more accurately understand the nature of the noise and adjust the initial anti-phase sound wave accordingly to generate the target anti-phase sound wave. This process optimizes the noise reduction effect, effectively eliminating noise within a specific frequency range and improving the user's listening experience in various usage scenarios.
[0098] Figure 2 shows a schematic diagram of an adaptive active noise cancellation device for headphones provided by the present invention. The device includes: a receiving unit 21, used by a mobile terminal to receive an in-ear noise signal sent by a headphone terminal; wherein the in-ear noise signal is a noise sound wave collected by a feedback microphone after noise reduction processing based on an initial anti-phase sound wave; a determining unit 22, used to determine the sound wave regularity characteristics corresponding to the in-ear noise signal according to the time-domain saliency spectrum corresponding to the in-ear noise signal; wherein the time-domain saliency spectrum is used to identify specific acoustic events; and an adjusting unit 23, used to adjust the initial anti-phase sound wave according to the sound wave regularity characteristics to obtain a target anti-phase sound wave, and send it to the headphone terminal for noise reduction processing.
[0099] This invention provides an adaptive active noise cancellation device for headphones. It receives in-ear noise signals from the headphone terminal via a mobile terminal and generates a time-domain saliency map based on these signals to identify specific acoustic events. This design allows the headphones to analyze changes in ambient noise in real time, ensuring that noise cancellation dynamically adapts to different noise environments. Compared to traditional fixed noise cancellation algorithms, this real-time adaptability significantly improves noise cancellation performance, especially in complex or changing noise environments. By determining the sound wave characteristics corresponding to the in-ear noise signals, this method can more accurately understand the nature of the noise and adjust the initial anti-phase sound wave accordingly to generate the target anti-phase sound wave. This process optimizes the noise cancellation effect, effectively eliminating noise within a specific frequency range and improving the user's listening experience in various usage scenarios.
[0100] Figure 3 is a schematic diagram of a terminal device according to an embodiment of the present invention. As shown in Figure 3, the terminal device 3 of this embodiment includes: a processor 30, a memory 31, and a computer program 32 stored in the memory 31 and executable on the processor 30, such as an adaptive headphone active noise cancellation program. When the processor 30 executes the computer program 32, it implements the steps in the various embodiments of the adaptive headphone active noise cancellation method described above, such as steps 101 to 103 shown in Figure 1. Alternatively, when the processor 30 executes the computer program 32, it implements the functions of each unit in the various device embodiments described above, such as the functions of the unit shown in Figure 2.
[0101] For example, the computer program 32 can be divided into one or more units, which are stored in the memory 31 and executed by the processor 30 to complete the present invention. The one or more units can be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program 32 in the terminal device 3. For example, the specific functions of each unit of the computer program 32 are as follows: a receiving unit, used by the mobile terminal to receive an in-ear noise signal sent by the headphone terminal; wherein the in-ear noise signal is a noise sound wave collected by a feedback microphone after noise reduction processing based on an initial anti-phase sound wave; a determining unit, used to determine the sound wave regularity characteristics corresponding to the in-ear noise signal according to the time-domain saliency spectrum corresponding to the in-ear noise signal; wherein the time-domain saliency spectrum is used to identify specific acoustic events; and an adjusting unit, used to adjust the initial anti-phase sound wave according to the sound wave regularity characteristics to obtain a target anti-phase sound wave, and send it to the headphone terminal for noise reduction processing.
[0102] The terminal device includes, but is not limited to, a processor 30 and a memory 31. Those skilled in the art will understand that Figure 3 is merely an example of a terminal device 3 and does not constitute a limitation on a single terminal device 3. It may include more or fewer components than illustrated, or combine certain components, or use different components. For example, the terminal device may also include input / output devices, network access devices, buses, etc.
[0103] The processor 30 can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0104] The memory 31 can be an internal storage unit of the terminal device 3, such as a hard disk or memory of the terminal device 3. The memory 31 can also be an external storage device of the terminal device 3, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the terminal device 3. Furthermore, the memory 31 can include both internal and external storage units of the terminal device 3. The memory 31 is used to store the computer program and other programs and data required by the roaming control device. The memory 31 can also be used to temporarily store data that has been output or will be output.
[0105] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0106] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of the present invention. For details on their specific functions and technical effects, please refer to the method embodiments section, which will not be repeated here.
[0107] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this invention. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0108] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps described in the various method embodiments above.
[0109] This invention provides a computer program product that, when run on a mobile terminal, enables the mobile terminal to implement the steps described in the above-described method embodiments.
[0110] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying the computer program code to a photographing device / terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks.
[0111] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0112] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0113] In the embodiments provided by this invention, it should be understood that the disclosed apparatus / network devices and methods can be implemented in other ways. For example, the apparatus / network device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0114] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; they may be located in one place or distributed across multiple network units.
[0115] It should be understood that, when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0116] It should also be understood that the term “and / or” as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0117] As used in this specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if [the described condition or event] is detected" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once [the described condition or event] is detected," or "in response to detection of [the described condition or event]."
[0118] Furthermore, in the description of this invention and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0119] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of the invention include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0120] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. An adaptive headphone active noise cancellation method, characterized in that, The adaptive headphone active noise cancellation method includes: a mobile terminal receiving an in-ear noise signal sent by a headphone terminal; wherein the in-ear noise signal is a noise sound wave collected by a feedback microphone after noise reduction processing based on an initial anti-phase sound wave; determining the sound wave regularity characteristics corresponding to the in-ear noise signal according to the time-domain saliency map corresponding to the in-ear noise signal; wherein the time-domain saliency map is used to identify specific acoustic events; and adjusting the initial anti-phase sound wave according to the sound wave regularity characteristics to obtain a target anti-phase sound wave, and sending it to the headphone terminal for noise reduction processing.
2. The adaptive headphone active noise cancellation method as described in claim 1, characterized in that, Before the step of the mobile terminal receiving the in-ear noise signal sent by the earphone terminal, the method further includes: collecting ambient sound waves based on a feedforward microphone and generating an initial anti-phase sound wave; performing noise reduction processing on the initial anti-phase sound wave; collecting in-ear sound waves based on a feedback microphone during the noise reduction processing; removing audio data from the in-ear sound waves to obtain the in-ear noise signal; and sending the in-ear noise signal to the mobile terminal.
3. The adaptive headphone active noise cancellation method as described in claim 1, characterized in that, The step of determining the acoustic wave regularity characteristics corresponding to the intra-ear noise signal based on the time-domain saliency map of the intra-ear noise signal includes: Step A1: Calculating the instantaneous amplitude corresponding to the discrete intra-ear noise signal based on the Hilbert transform; Step A2: Calculating the time-domain saliency map based on the instantaneous amplitude; the time-domain saliency map is used to identify specific acoustic events; Step A3: Extracting the highest peak point in the time-domain saliency map; Step A4: Constructing an adaptive resonant core based on the highest peak point; Step A5: Projecting the intra-ear noise signal onto the constructed optimal resonant core to obtain the target mode components corresponding to each discrete intra-ear noise signal; Step A6: Using the target mode components corresponding to each discrete intra-ear noise signal as the current noise feature.
4. The adaptive headphone active noise cancellation method as described in claim 3, characterized in that, Step A5 includes: Step A51: Projecting the intra-ear noise signal onto the constructed optimal resonant core to obtain the initial mode components corresponding to each discrete intra-ear noise signal; Step A52: Subtracting the initial mode components corresponding to each discrete intra-ear noise signal from the discrete intra-ear noise signal to obtain the current residual signal; Step A53: Using the current residual signal as the discrete intra-ear noise signal, iteratively executing steps A1 to A53 until a termination condition is met to obtain the target mode components corresponding to each discrete intra-ear noise signal; wherein, the termination condition includes the total energy of the residual signal being lower than a threshold.
5. The adaptive headphone active noise cancellation method as described in claim 3, characterized in that, Step A2 includes: substituting the instantaneous amplitude into the time-domain significance function to obtain the time-domain significance map; the time-domain significance function is: ;in, This represents the time-domain saliency map, where n represents the time point corresponding to the discrete intraocular noise signal. This represents a time window centered at n. Represents variance. This represents the instantaneous amplitude of the nth discrete signal in the discrete intra-auricular noise signal. This represents the instantaneous amplitude of the m-th discrete signal in the discrete intraocular noise signal, where m represents a time point within the time window.
6. The adaptive headphone active noise cancellation method as described in claim 3, characterized in that, Step A4 includes: substituting the highest peak point into a preset function to obtain the adaptive resonant core; wherein the preset function is: ;in, This represents an adaptive resonant core. Indicates the highest peak point. Indicates the width of the Gaussian envelope. , Indicates a point in time The corresponding signal value at that location, Indicates the resonant frequency. Indicates the initial phase.
7. The adaptive headphone active noise cancellation method as described in claim 1, characterized in that, The step of adjusting the initial antiphase sound wave according to the sound wave characteristics to obtain the target antiphase sound wave and sending it to the headphone terminal for noise reduction processing includes: inputting multiple target modal components into a one-dimensional convolutional neural network to obtain the sound category and confidence level corresponding to each of the multiple target modal components; wherein, the sound category includes human voice, horn sound, wind sound, water sound, knocking sound and keyboard sound; obtaining a preset cancellation weight mask corresponding to each sound category; multiplying the preset cancellation weight mask, confidence level and in-ear noise signal to obtain a cancellation signal; processing the cancellation signal through a filter to obtain the target antiphase sound wave; and sending the corrected antiphase sound wave to the headphone terminal for noise reduction processing.
8. An adaptive headphone active noise cancellation device, characterized in that, The adaptive headphone active noise cancellation device includes: a receiving unit for receiving an in-ear noise signal sent by a headphone terminal from a mobile terminal; wherein the in-ear noise signal is a noise sound wave collected by a feedback microphone after noise reduction processing based on an initial anti-phase sound wave; a determining unit for determining the sound wave regularity characteristics corresponding to the in-ear noise signal according to the time-domain saliency spectrum corresponding to the in-ear noise signal; wherein the time-domain saliency spectrum is used to identify specific acoustic events; and an adjusting unit for adjusting the initial anti-phase sound wave according to the sound wave regularity characteristics to obtain a target anti-phase sound wave, and sending it to the headphone terminal for noise reduction processing.
9. A terminal device, characterized in that, The terminal device includes: a memory, a processor, and an adaptive headphone active noise cancellation program stored in the memory and executable on the processor, the adaptive headphone active noise cancellation program being configured to implement the steps in the adaptive headphone active noise cancellation method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps in the adaptive headphone active noise cancellation method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Waveform music component separation method based on deep learning
CN117253502A
Bluetooth headset noise reduction processing method and device, equipment and storage medium
CN118368561A
Earphone adjusting system with active noise reduction function
CN121300414A
Noise reducing headphone
JP1994006886A
Ear-wearable device with active noise cancellation system that uses internal and external microphones
US20230300516A1