Method and system for optimizing noise reduction effect of earphone

By collecting and analyzing ambient sound in the headset, and dynamically switching the noise reduction mode using sliding time windows and machine learning models, the problem of insufficient response in a fast acoustic environment is solved, achieving a more efficient noise reduction and comfort experience.

CN120343456APending Publication Date: 2025-07-18SHENZHEN KINGSTAR IND CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510526608.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

When facing a rapidly changing acoustic environment, existing active noise reduction systems have insufficient response capabilities for dynamic phase changes, resulting in phase deviation between reverse sound waves and noise signals, and cannot achieve effective sound wave cancellation, which may even cause noise perception enhancement and ear pressure discomfort, affecting comfort and safety.

Method used

Ambient sound acquisition is performed through the headphones' built-in microphone, the audio stream is divided into sub-sections using the sliding time window, short-time Fourier transform and spectrum energy distribution analysis are performed, feature indicators are extracted, intelligent prediction is combined with machine learning models, dynamically switch noise reduction mode and adjust frequency to adapt to fast acoustic disturbances.

Benefits of technology

It realizes high-precision identification of rapidly changing acoustic environments and intelligent switching of dynamic noise reduction structures, avoids phase lag and structural mismatch, improves noise reduction performance and user auditory comfort, and optimizes the user experience in variable sound fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343456A_ABST
    Figure CN120343456A_ABST
Patent Text Reader

Abstract

The invention discloses a method and system for optimizing the noise reduction effect of an earphone, and relates to the technical field of earphone noise reduction, and the method comprises the following steps: carrying out the environment sound collection at a fixed sampling rate through employing a built-in microphone of the earphone, dividing an audio stream into a plurality of sub-segments through employing a sliding time window, and forming a continuous acoustic frame sequence; performing short-time Fourier transform processing on each sub-segment obtained by division, and extracting spectrum energy distribution in the sub-segment; and for each sub-fragment, extracting a characteristic index representing the rapid change of the audio from the spectrum energy distribution diagram, and performing deep analysis on the extracted characteristic index. Through feature extraction and machine learning prediction, noise reduction control from passive response to active prediction is realized, when rapid disturbance is detected, a noise reduction structure is dynamically switched, the switching frequency is adaptively adjusted, phase lag and structure mismatch are effectively avoided, and noise reduction performance and user comfort are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of headphone noise reduction, and particularly to a method and system for optimizing the noise reduction effect of headphones. Background Art

[0002] Optimizing the noise reduction effect of headphones means that during the process of wearing headphones, through a series of technical means, the ability to suppress ambient noise is enhanced, so that when users listen to music, make calls or are in a quiet environment, they can more effectively isolate external noise interference. The prior art usually adopts two types of methods to achieve noise reduction optimization: one is passive noise reduction, which blocks the physical propagation of medium and high-frequency external noise by designing an earcup structure that fits tightly against the auricle and using high-density sound insulation materials; the other is active noise reduction, which collects external noise signals in real time through a built-in microphone, and uses a noise reduction chip to generate "anti-phase sound waves" with opposite phases for sound wave cancellation to weaken the perception of low-frequency noise. In addition, in recent years, advanced methods such as hybrid noise reduction structures (combining feedforward and feedback microphones), adaptive filtering algorithms, personalized ear canal modeling, and AI-based noise scene recognition and adjustment technologies have also emerged to further improve the intelligence and adaptability of noise reduction, thereby enhancing the auditory experience of users in complex acoustic environments.

[0003] The working principle of active noise reduction is based on the principle of sound wave phase interference. Its core lies in that the microphone integrated inside the headphone will pick up the noise signals in the external environment in real time, and these noise signals are transmitted to the noise reduction processing chip inside the headphone. After high-speed operation by this chip, "anti-phase sound waves" with the same amplitude but opposite phases as the noise signals are generated. Since sound waves are superimposable, when the original noise wave meets its anti-phase wave in the ear canal, phase cancellation (also known as "destructive interference") will occur, thereby greatly reducing the noise energy felt by the user in the ear. This technology is particularly good at eliminating low-frequency continuous noise, such as aircraft engine noise, subway rumbling or air conditioner operation noise. The whole process relies on the precise tracking and real-time response of the high-speed digital signal processor (DSP) to noise changes, so as to achieve the closed-loop control effect of dynamic noise reduction.

[0004] The prior art has the following deficiencies:

[0005] In the prior art, when an active noise cancellation system faces a rapidly changing acoustic environment, there is generally a problem of insufficient response ability to the dynamic phase change of the noise signal. Due to the lack of a high-frequency dynamic tracking and real-time adjustment mechanism in the system, the generated reverse sound wave lags in response during the phase synchronization process and cannot achieve an accurate anti-correlation control relationship with the original noise signal. When there is a phase deviation between the reverse sound wave and the noise signal, it is not only difficult to achieve an effective sound wave cancellation effect, but may also cause an abnormal increase in the sound pressure in the target frequency band due to phase superposition, resulting in a phenomenon of enhanced noise perception. This type of phase mismatch problem is particularly concentrated in the low-frequency region during actual use, easily causing a "head-banging feeling" in the user's subjective perception, and causing ear pressure discomfort, auditory fatigue during continuous wearing, and even potentially interfering with the user's hearing health, seriously affecting the comfort and safety of the active noise cancellation device during use.

[0006] The above information disclosed in the background art section is only used to enhance the understanding of the background of the present disclosure, and thus it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0007] The object of the present invention is to provide a method and system for optimizing the noise cancellation effect of headphones. By continuously sampling the audio signal, performing short-time spectrum analysis and dynamic feature extraction, and combining a machine learning model to intelligently predict the trend of acoustic disturbances, the noise cancellation decision is changed from "passive response" to "active prediction". When a rapidly changing acoustic disturbance is detected, the system can timely switch to a more suitable noise cancellation structure (feedforward or feedback mode) and adaptively adjust the switching frequency based on the audio disturbance rate, effectively avoiding "noise cancellation failure" or "head-banging feeling" caused by phase lag and structure mismatch, not only improving the overall noise cancellation performance, but also optimizing the user's auditory comfort and usage experience in a changing sound field to solve the problems in the above background art.

[0008] To achieve the above object, the present invention provides the following technical solution: A method for optimizing the noise cancellation effect of headphones, comprising the following steps:

[0009] Use the built-in microphone of the headphones to collect ambient sound at a fixed sampling rate, and use a sliding time window to divide the audio stream into multiple sub-fragments to form a continuous sequence of acoustic frames;

[0010] Perform short-time Fourier transform processing on each divided sub-fragment to extract the spectral energy distribution within the sub-fragment;

[0011] For each sub - segment, extract the characteristic indicators representing the rapid changes in the audio from the spectral energy distribution diagram. After deeply analyzing the extracted characteristic indicators, use the analyzed characteristic indicators as feature vectors and input them into a pre - trained machine learning model to intelligently predict the audio change situation in the external environment of the earphone;

[0012] When it is recognized that the external environment audio is in a rapid change state, start the noise - reduction model switching engine, dynamically switch the working modes of feedback ANC and feed - forward ANC, and adaptively adjust the switching frequency based on the acoustic environment change rate.

[0013] Preferably, use a sliding time window to divide the audio stream into multiple sub - segments to form a continuous sequence of acoustic frames, which specifically includes the following steps:

[0014] S1. Set a time window with a fixed length as the time span of each sub - segment;

[0015] S2. Starting from the audio acquisition starting point, intercept the first sub - segment from the continuous audio stream in units of this time window;

[0016] S3. Slide the time window backward according to the set sliding step size, and intercept the next sub - segment starting from the current position;

[0017] S4. Repeat S3 to successively intercept multiple time - overlapping sub - segments, and finally construct a continuous acoustic frame sequence covering the entire time period.

[0018] Preferably, perform short - time Fourier transform processing on each divided sub - segment to extract the spectral energy distribution within the sub - segment. The specific steps are as follows:

[0019] Take the time - domain audio signal of each sub - segment as the input, apply a window function to it to reduce the interference of edge effects on the spectrum;

[0020] Use the fast Fourier transform to perform frequency - domain transformation on the windowed signal to convert the time - domain signal into a spectrum containing complex coefficients;

[0021] Calculate the square of the magnitude of each frequency point in the spectrum to obtain the energy value of the corresponding frequency component, and form the energy spectrum diagram of the sub - segment within a certain time interval;

[0022] Arrange the spectral results of multiple sub - segments in chronological order to construct a two - dimensional spectral energy distribution matrix reflecting the audio change over time.

[0023] Preferably, for each sub - segment, characteristic indicators representing the rapid change of the audio are extracted from the spectral energy distribution diagram. Among them, the extracted characteristic indicators include the degree of change in frequency - energy distribution between adjacent spectral frames and the proportional change of energy jumping from the mid - low frequency band to the high - frequency band. After in - depth analysis of the extracted characteristic indicators, a spectral transition reference value and a high - frequency escape reference value are generated respectively. The overall fluctuation intensity of the audio signal in the frequency structure is characterized by the spectral transition reference value, and the density and trend of the sudden high - frequency energy surge phenomenon in the continuous time domain are characterized by the high - frequency escape reference value.

[0024] Preferably, the analyzed spectral transition reference value and high - frequency escape reference value are used as feature vectors and input into a pre - trained machine - learning model. The audio change coefficient is output by the machine - learning model, and the audio change situation of the external environment of the earphone is intelligently predicted based on the audio change coefficient.

[0025] Preferably, the audio change coefficient generated when the machine - learning model intelligently predicts the audio change situation of the external environment of the earphone is compared and analyzed with a pre - set audio change coefficient reference threshold to classify the audio change of this sub - segment. The specific classification steps are as follows:

[0026] If the audio change coefficient is greater than the audio change coefficient reference threshold, the external - environment audio change of this sub - segment is classified as a rapid change;

[0027] If the audio change coefficient is less than or equal to the audio change coefficient reference threshold, the external - environment audio change of this sub - segment is classified as a regular change.

[0028] Preferably, when it is recognized that the external - environment audio is in a rapid - change state, a noise - reduction model switching engine is started to dynamically switch the working modes of feedback ANC and feed - forward ANC, and the switching frequency is adaptively adjusted based on the acoustic - environment change rate. The specific steps are as follows:

[0029] Perform trend modeling on the change states of audio sub - segments within a continuous time period. Specifically: after obtaining the audio change coefficients corresponding to each sub - segment, construct an audio perturbation trend function to measure the perturbation acceleration in the acoustic environment. The expression of the audio perturbation trend function is as follows:

[0030]

[0031] Where: represents the audio change coefficient corresponding to the (k - 1) - th frame sub - segment in the current sliding window, represents the audio change coefficient of the k - th frame in the sliding window, which is the central value of the reference frame within this local window; It represents the audio change coefficient of the (k + 1)-th frame in the sliding window; n is the number of frames in the sliding time window, and Ω t is the target disturbance trend intensity index;

[0032] After obtaining the target disturbance trend intensity index, adaptively and dynamically adjust the switching frequency between the feedback ANC and the feedforward ANC. The expression for the adaptive dynamic adjustment is:

[0033]

[0034] , where: f switch (t) represents the model switching frequency at the current moment; f min is the minimum base switching frequency set by the system in a stable environment; α p is the maximum frequency adjustment amplitude coefficient, representing the maximum frequency increase value that can be achieved in a rapidly changing environment; β r is the disturbance trend sensitivity adjustment coefficient.

[0035] Preferably, for each sub - segment, the specific steps to generate the spectral transition reference value after deeply analyzing the change degree of the frequency energy distribution between adjacent spectral frames are as follows:

[0036] Express the short - time Fourier transform result in this sub - segment as a two - dimensional spectral energy matrix F, F = {F i,j}, where i represents the i - th frame, j represents the j - th frequency point, and F i,j represents the energy component of the i - th frame at the j - th frequency. Then calculate the spectral difference degree between each pair of adjacent frames. The calculation expression is:

[0037]

[0038] , where: D i represents the spectral difference degree between the i - th frame and the (i + 1)-th frame; N is the spectral dimension; λ is the spectral difference normalization factor, which suppresses the suppression of the difference by the high - amplitude main frequency;

[0039] After obtaining the inter - frame spectral difference sequence, calculate the spectral transition reference value based on the ratio relationship between the maximum spectral perturbation amplitude and the weighted overall perturbation. The specific calculation formula is as follows:

[0040]

[0041] , where: S tt is the spectral transition reference value, which is used to characterize the frequency structure fluctuation intensity within this sub - segment; max(D i ) represents the spectral mutation amplitude between the strongest pair of frames; M is the total number of frames in this sub - segment; ω iis the modulation weight, emphasizing the perturbation sensitivity in the middle of the i-th frame sequence; ∈ is the minimum stable constant.

[0042] Preferably, for each sub - segment, the specific steps of generating the high - frequency escape reference value after deeply analyzing the ratio change of energy jumping from the middle - low frequency band to the high - frequency band are as follows:

[0043] In the spectral energy distribution diagram corresponding to each sub - segment, set two frequency - band sets: the middle - low - frequency energy region A and the high - frequency energy region B. By calculating the total energy belonging to the high - frequency region and the total energy belonging to the middle - low - frequency region in this sub - segment, construct the high - frequency transition energy ratio factor. The constructed expression is:

[0044]

[0045] , where: E B is the cumulative energy of the high - frequency region, reflecting the intensity of burst enhancement; E A is the cumulative energy of the middle - low - frequency region, used to measure the background energy baseline; E B 2 is used to enhance the sensitivity of sudden high - frequency escape behavior; Λ is the high - frequency transition energy ratio factor of the current segment;

[0046] After obtaining the high - frequency transition energy ratio factor, further perform a structural non - linear mapping on it to construct the high - frequency escape reference value, which is used to enhance the expression of the directional surge trend of the transition phenomenon. The construction expression of the high - frequency escape reference value is:

[0047] γ dix =|tanh(Λ - Λ 00 )| δ

[0048] , where: Λ0 is the system - set transition stability threshold, indicating the reasonable distribution boundary of high - frequency and middle - low - frequency energies in the normal acoustic structure; tanh(*) is used to limit extreme values and enhance non - linear response; δ is the trend enhancement factor, used to regulate the sensitivity of the system to the surge trend; Υ dix is the high - frequency escape reference value.

[0049] A system for optimizing the noise reduction effect of headphones includes an audio acquisition and window framing module, a spectrum extraction and time - frequency transformation module, a feature analysis and intelligent prediction module, and a model switching and frequency adaptive control module;

[0050] The audio acquisition and window framing module uses the built - in microphone of the headphones to collect environmental sounds at a fixed sampling rate, and uses a sliding time window to divide the audio stream into multiple sub - segments, forming a continuous acoustic frame sequence;

[0051] Spectrum extraction and time-frequency transformation module, which performs short-time Fourier transform processing on each divided sub-segment to extract the spectrum energy distribution within the sub-segment;

[0052] Feature analysis and intelligent prediction module. For each sub-segment, it extracts feature indicators characterizing the rapid changes in the audio from the spectrum energy distribution diagram. After deeply analyzing the extracted feature indicators, it inputs the analyzed feature indicators as feature vectors into a pre-trained machine learning model, and uses the machine learning model to intelligently predict the audio change situation in the external environment of the earphone;

[0053] Model switching and frequency adaptive control module. When it recognizes that the external environment audio is in a rapidly changing state, it starts the noise reduction model switching engine, dynamically switches the working modes of feedback ANC and feedforward ANC, and adaptively adjusts the switching frequency based on the change rate of the acoustic environment.

[0054] In the above technical solution, the technical effects and advantages provided by the present invention:

[0055] The present invention performs continuous sampling, short-time spectrum analysis and dynamic feature extraction on the audio signal, combines a machine learning model to intelligently predict the trend of acoustic disturbances, so that the noise reduction decision changes from "passive response" to "active prediction". When detecting rapidly changing acoustic disturbances, the system can timely switch to a more suitable noise reduction structure and adaptively adjust the switching frequency based on the audio disturbance rate, effectively avoiding "noise reduction failure" or "boom feeling" caused by phase lag and structural mismatch, not only improving the overall noise reduction performance, but also optimizing the user's auditory comfort and usage experience in a changing sound field. Description of the Drawings

[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.

[0057] Figure 1 It is a flowchart of a method for optimizing the noise reduction effect of the earphone of the present invention.

[0058] Figure 2 It is a schematic diagram of the system module for optimizing the noise reduction effect of the earphone of the present invention. Detailed Embodiments

[0059] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these example embodiments are provided so that this disclosure will be more complete and thorough, and will fully convey the concept of the example embodiments to those skilled in the art.

[0060] The present invention provides a method for optimizing the noise reduction effect of headphones as shown in Figure 1 and includes the following steps:

[0061] Using the built-in microphone of the headphones to collect ambient sound at a fixed sampling rate (such as 16 kHz or higher), and using a sliding time window (such as every 500 ms) to divide the audio stream into multiple sub-fragments, forming a continuous sequence of acoustic frames;

[0062] This step provides a continuous and stable basic data stream for subsequent spectral analysis and change trend modeling, ensuring that the fine-grained changes of ambient sound can be continuously captured in various dynamic acoustic environments, and is the front-end perception basis for the whole system to perceive the external environment.

[0063] Using a sliding time window to divide the audio stream into multiple sub-fragments and form a continuous sequence of acoustic frames specifically includes the following steps:

[0064] First, set a time window with a fixed length (such as 500 ms) as the time span of each sub-fragment;

[0065] Then, starting from the audio acquisition start point, intercept the first sub-fragment from the continuous audio stream in units of this time window;

[0066] Next, slide the time window backward according to the set sliding step (such as 250 ms, representing a 50% overlap rate), and intercept the next sub-fragment starting from the current position;

[0067] Repeat this process in sequence to intercept multiple time-partially overlapping sub-fragments, and finally construct a continuous and full-time coverage sequence of acoustic frames.

[0068] Intercepting multiple time-partially overlapping sub-fragments in sequence means that when using a sliding time window to divide the audio stream, there is a part of the time interval that overlaps between two adjacent sub-fragments, that is, they share a part of the same audio data, rather than intercepting completely non-overlapping. This partially overlapping division is a characteristic of the sliding window mechanism, and the purpose is to improve the time resolution and avoid the signal change being "cut off" at the window boundary.

[0069] For example, assume that the sliding time window length is set to 500 milliseconds and the sliding step is 250 milliseconds. Then the first sub - segment covers the audio from 0 milliseconds to 500 milliseconds, and the second sub - segment starts at 250 milliseconds and ends at 750 milliseconds. It can be seen that the 250 - millisecond time interval (i.e., from 250 ms to 500 ms) between these two sub - segments is repeated. Such a division method will continue. For example, the third sub - segment is from 500 ms to 1000 ms, and so on. In this way, the constructed sub - segments have partial overlap, which can ensure that in a rapidly changing sound environment, any sudden audio changes will not be truncated by the window boundaries, facilitating the improvement of the time continuity and change perception accuracy of subsequent spectral analysis or feature extraction.

[0070] This method can improve the capture accuracy of rapidly changing audio features on the basis of ensuring time continuity, and is a fundamental link for subsequent short - time spectral analysis and dynamic noise recognition.

[0071] Perform short - time Fourier transform processing on each divided sub - segment to extract the spectral energy distribution within the sub - segment;

[0072] This step is the key to realizing time - frequency domain mapping, which can convert the sound fluctuation features that are difficult to observe in the time domain into a clear and computable frequency response structure, facilitating the capture of short - time change features such as high - frequency jumps and low - frequency oscillations, and is the core data source for subsequent change detection and machine learning prediction.

[0073] Perform short - time Fourier transform (STFT) processing on each divided sub - segment to extract the spectral energy distribution within the sub - segment. The specific steps are as follows: First, take the time - domain audio signal of each sub - segment as the input and apply a window function (such as a Hamming window or a Hanning window) to reduce the interference of edge effects on the spectrum. Then, use the fast Fourier transform (FFT) to perform a frequency - domain transformation on the windowed signal to convert the time - domain signal into a spectrum containing complex coefficients. Next, calculate the square of the magnitude of each frequency point in the spectrum to obtain the energy value of the corresponding frequency component, forming the energy spectrum diagram of the sub - segment within a certain time interval. Finally, arrange the spectral results of multiple sub - segments in chronological order to construct a two - dimensional spectral energy distribution matrix (i.e., a time - frequency diagram) that reflects the audio change over time, providing basic data for subsequent audio feature extraction and change trend analysis.

[0074] For each sub - segment, extract the feature indicators characterizing the rapid audio changes from the spectral energy distribution diagram. After deeply analyzing the extracted feature indicators, use the analyzed feature indicators as feature vectors and input them into a pre - trained machine learning model to intelligently predict the audio change situation in the external environment of the earphone through the machine learning model;

[0075] For each sub - segment, characteristic index indicators that characterize the rapid change of audio are extracted from the spectral energy distribution diagram. Among them, the extracted characteristic index indicators include the degree of change in frequency - energy distribution between adjacent spectral frames and the proportional change of energy leaping from the mid - low frequency band to the high - frequency band. After in - depth analysis of the extracted characteristic index indicators, a spectral transition reference value and a high - frequency escape reference value are generated respectively. The overall fluctuation intensity of the audio signal in the frequency structure is characterized by the spectral transition reference value, and the density and trend of the sudden high - frequency energy surge phenomenon in the continuous time domain are characterized by the high - frequency escape reference value, providing an accurate quantitative basis for subsequent judgment of whether the external environment is in a rapidly changing state;

[0076] If the degree of change in frequency - energy distribution between adjacent spectral frames in the spectral energy distribution diagram corresponding to the sub - segment is relatively high, it usually indicates that the audio signal of this sub - segment has changed relatively quickly. The reason is as follows: After mapping the audio signal from the time domain to the frequency domain by the short - time Fourier transform, the spectral frames reflect the energy distribution of each frequency component at different time points; when the energy distribution difference between adjacent spectral frames is significant, it indicates that the signal has experienced a drastic frequency - structure adjustment or energy transition in a very short time. This phenomenon usually appears in scenarios such as rapid sound mutation, short - time strong noise occurrence, and dynamic movement of the sound source. Therefore, the high degree of change in frequency - energy distribution is an important feature of rapid perturbation of audio in the time - frequency joint space, which can effectively characterize the suddenness and non - stationarity of the sound within this sub - segment.

[0077] For each sub - segment, the specific steps to generate the spectral transition reference value after in - depth analysis of the degree of change in frequency - energy distribution between adjacent spectral frames are as follows:

[0078] Express the short - time Fourier transform result of this sub - segment as a two - dimensional spectral energy matrix F, F = {F i,j}, where i represents the i - th frame, j represents the j - th frequency point, and F i,j represents the energy component of the i - th frame at the j - th frequency. Then calculate the spectral difference degree (i.e., the total relative gradient of frequency - energy) between each pair of adjacent frames. The calculation formula is:

[0079]

[0080] , where: D i represents the spectral difference degree between the i - th frame and the (i + 1) - th frame; N is the spectral dimension (the number of frequency points); λ is the spectral - difference normalization factor to suppress the suppression of the difference by high - amplitude main frequencies, and the recommended value range is 0.01 - 0.1; the denominator introduces 1+λ·|F i+1,j +F i,j | to avoid the over - dominance of high - energy frequency bands and enhance the sensitivity to medium - low amplitude fluctuations;

[0081] This step constructs a non-linear, amplitude-adaptive sequence of inter-frame spectral change metrics to capture the perturbation of the energy distribution of the spectrum in the time direction and avoid masking the change trend due to the strength of the main frequency energy.

[0082] After obtaining the inter-frame spectral difference sequence, calculate the spectral transition reference value based on the ratio relationship between the maximum spectral perturbation amplitude and the weighted overall perturbation. The specific calculation formula is as follows:

[0083]

[0084] , where: S tt is the spectral transition reference value, used to characterize the intensity of frequency structure fluctuations within this sub-segment; max(D i ) represents the spectral mutation amplitude between the strongest pair of frames; M is the total number of frames in this sub-segment; ω i is the modulation weight, emphasizing the perturbation sensitivity in the middle of the i-th frame sequence (can be deformed into a Bessel distribution) to enhance the response ability to changes in the core area of the spectral structure; ∈ is a minimum stable constant to prevent the denominator from being zero, usually taking the order of 10 -6 ;

[0085] This step captures the concentrated characteristics of spectral changes in the form of "maximum perturbation / overall perturbation", combined with non-linear modulation weights, making the spectral transition reference value not only highly responsive to sudden transitions but also weakening the influence of edge perturbations, achieving highly sensitive and structured quantification of frequency structure fluctuations, suitable for transition identification in a fast dynamic sound field.

[0086] From the spectral transition reference value, it can be seen that for each sub-segment, the larger the value of the spectral transition reference value generated after in-depth analysis of the degree of change in the frequency energy distribution between adjacent spectral frames, the faster the audio change in this sub-segment, and vice versa, the slower the audio change in this sub-segment. The reason is that: the spectral transition reference value characterizes the structural fluctuation intensity of the audio signal by measuring the non-linear difference degree of the frequency energy distribution between adjacent spectral frames. When the audio is in a fast-changing state, significant energy migration or frequency peak drift phenomena will occur between consecutive frames of the spectrum, resulting in an increase in the difference degree between adjacent frames in the spectrogram, thus causing a significant increase in the numerator term representing the mutation amplitude in the spectral transition reference value; at the same time, although the overall perturbation intensity also increases with weighting, the ratio of the maximum mutation to the weighted total perturbation still tends to increase, ultimately resulting in an increase in the spectral transition reference value. On the contrary, if the audio signal remains stable or changes slowly, the difference in the frequency energy distribution between adjacent frames is small, and the spectral transition reference value also decreases accordingly. Therefore, the spectral transition reference value can sensitively reflect the degree of drastic changes in the frequency structure within a short period of time and is an important quantitative indicator for identifying rapid audio change behaviors.

[0087] If the proportion change of the energy jumping from the mid - low frequency band to the high - frequency band in the corresponding spectrum energy distribution map of a sub - segment is relatively high, it usually indicates that the audio of this sub - segment changes rapidly. The reason is that in natural environments or structured sound sources, most background noises and stable sound sources (such as wind sounds, machine running sounds, voice fundamentals) are mainly concentrated in the mid - low frequency region, while high - frequency components are often closely related to instantaneous changes, impulsive events or non - stationary sound sources, such as metal collisions, sharp voices, impulse noises, etc. When the energy jumps from the mid - low frequency concentration to the high - frequency region, it indicates that sudden and short - time non - continuous events have occurred in the audio signal. Such events show a sudden increase in high - frequency energy in the time - frequency domain, which is a typical fast - change feature. Therefore, the high - frequency escape behavior not only reflects the dynamic mutation of the spectrum structure but also reflects the non - linear perturbation phenomenon in the physical propagation process of the audio, and it is a sensitive and effective indicator for judging the audio change rate.

[0088] For each sub - segment, the specific steps to generate the high - frequency escape reference value after in - depth analysis of the proportion change of the energy jumping from the mid - low frequency band to the high - frequency band are as follows:

[0089] In the spectrum energy distribution map corresponding to each sub - segment, set two frequency - band sets: the mid - low frequency energy region A and the high - frequency energy region B. By calculating the total energy belonging to the high - frequency region and the total energy belonging to the mid - low frequency region in this sub - segment, construct the high - frequency transition energy ratio factor. The constructed expression is:

[0090]

[0091] , where: E B is the cumulative energy of the high - frequency region, reflecting the intensity of sudden excitation; E A is the cumulative energy of the mid - low frequency region, used to measure the background energy baseline; adding 1 is used to suppress the extreme case where the denominator is zero and maintain numerical stability; E B 2 is used to enhance the sensitivity of sudden high - frequency escape behavior; Λ is the high - frequency transition energy ratio factor of the current segment;

[0092] This step extracts the enhancement degree of high - frequency energy relative to mid - low frequency energy in the spectrum and amplifies the influence of "structural jumps" through non - linear processing, which is used to accurately identify whether high - frequency transition events have occurred.

[0093] After obtaining the high - frequency transition energy ratio factor, further perform a structural non - linear mapping on it to construct the high - frequency escape reference value, which is used to enhance the expression of the "directional surge trend" of the transition phenomenon. The construction expression of the high - frequency escape reference value is:

[0094] Υ dix =|tanh(Λ - Λ0)| δ

[0095] , where: Λ0 is the transition stability threshold set by the system (such as set by experience or obtained through model training), representing the reasonable distribution boundary of high-frequency and mid-low-frequency energies in a normal acoustic structure; tanh(*) is used to limit extreme values and enhance the non-linear response, enhancing the response to large jumps and compressing the response to weak perturbations; δ is the trend enhancement factor (recommended to be 1.5-3), used to regulate the sensitivity of the system to the surging trend; Υ dix is the high-frequency escape reference value;

[0096] This step performs normalization and amplification on the high-frequency surging trend through the non-linear response mechanism of the hyperbolic tangent function, thereby generating a structural and numerically stable high-frequency escape reference value Υ dix , which can be directly used to identify whether there is a non-linear excitation behavior of the high-frequency band energy structure in the current sub-fragment and has the ability to perform highly sensitive modeling on sudden perturbations.

[0097] From the high-frequency escape reference value, it can be seen that for each sub-fragment, the larger the performance value of the high-frequency escape reference value generated after in-depth analysis of the proportion change of the energy jumping from the mid-low frequency band to the high frequency band, the faster the audio change of the sub-fragment, and vice versa, the slower the audio change of the sub-fragment. The reason is that the high-frequency escape reference value essentially reflects the degree of discontinuous transition of the spectral energy in the structure, especially focusing on whether the energy jumps from the low-frequency band to the high-frequency band in a very short time. In physical acoustic characteristics, high-frequency components usually correspond to fast-changing behaviors such as short-time impacts, sharp transients, and structural perturbations. Therefore, when the mid-low frequency energy decreases rapidly and the high-frequency energy increases significantly, it represents that a severe time-frequency mutation has occurred in the signal in this fragment. Through the non-linear amplification mechanism introduced in the high-frequency escape reference value (such as the square enhancement of the transition ratio and the tanh mapping), this change trend will be further amplified, thus forming a highly positive correlation between the size of the high-frequency escape reference value and the audio change speed. Therefore, the high-frequency escape reference value can be regarded as a sensitive quantitative index for judging the speed of audio change within a single fragment.

[0098] Input the analyzed spectral transition reference value and high-frequency escape reference value as feature vectors into a pre-trained machine learning model, and output the audio change coefficient through the machine learning model, and perform intelligent prediction on the audio change situation of the external environment of the headset based on the audio change coefficient.

[0099] A pre-trained machine learning model refers to a model that, before system implementation, is trained on a large amount of historical data and preset scenarios using machine learning algorithms (such as supervised learning, deep learning, etc.) so that the model can learn to extract effective information from input features and make predictions. Specifically, a "pre-trained model" goes through a process in which the model continuously adjusts its parameters to learn the patterns and rules of audio signal changes from input features (such as spectral transition reference values and high-frequency escape reference values), thereby accurately predicting audio changes in the external environment.

[0100] The construction process of a pre-trained machine learning model is divided into several stages: data collection and preparation, feature selection and extraction, model selection and training, validation and optimization. In the first stage, the system first collects a large number of audio data samples, which usually come from different audio environments, such as static environments, traffic noise, industrial noise, music scenes, etc. During the data collection process, each sample also needs to be labeled, and the labeling content may include the category of the audio signal (such as a stable environment or a rapidly changing environment) or the intensity of a certain change. Through these labels, the model can learn how to make corresponding predictions from the given input data based on the features.

[0101] Next, in the feature selection and extraction stage, the system extracts some key audio features from the audio data. These features include spectral transition reference values and high-frequency escape reference values, which can effectively describe the dynamic changes in the spectral structure of the audio signal. The spectral transition reference value mainly focuses on the degree of mutation in the spectral energy distribution, while the high-frequency escape reference value reflects whether the energy jumps from the mid-low frequency to the high-frequency band. These features are extracted through mathematical models and algorithms (such as Fourier transform, time-frequency analysis, etc.) to generate data vectors suitable for input into the machine learning model.

[0102] In the model selection and training stage, the selected machine learning algorithms (such as support vector machine (SVM), decision tree, neural network, etc.) are trained on these feature vectors. The model is trained and optimized multiple times on the relationship between the input data and the labeled data so that it can find the mapping relationship between audio changes and these features. This training process is usually adjusted and optimized through techniques such as gradient descent, error backpropagation, and cross-validation to ensure that the model can complete the prediction task with the minimum error. Through this training process, the model learns how to judge whether the audio changes rapidly or predict the intensity of the audio changes based on the input features.

[0103] Once the training is completed, the machine learning model becomes a tool that can make intelligent predictions in practical applications. When new audio data (such as environmental audio data collected in real time by headphones) is input, the system uses these input features (spectral transition reference value and high-frequency escape reference value) to generate feature vectors and passes them to the pre-trained machine learning model for prediction. The output of the model is the audio change coefficient, which reflects the intensity and trend of the external environmental audio change.

[0104] The machine learning model is not specifically limited here, as long as it can realize the comprehensive analysis of the spectral transition reference value S tt and the high-frequency escape reference value Υ dix to generate the audio change coefficient Audio change is acceptable. To implement the technical solution of the present invention, the present invention provides a specific implementation; the expression for generating the audio change coefficient Audio change is: Audio change = p m / S tt + p n *Υ dix , where p m , p n are the preset proportionality coefficients of the spectral transition reference value S tt and the high-frequency escape reference value Υ dix respectively, and p m , p n are both greater than 0.

[0105] The preset proportionality coefficient refers to the weighting factor applied to different feature indicators when generating the audio change coefficient Audio change , that is, the constant coefficients p tt and p dix used to adjust the weights of the spectral transition reference value S m and the high-frequency escape reference value Υ n in the calculation of the final audio change coefficient. Specifically, these preset proportionality coefficients play a role in controlling and adjusting the influence degree of different features in the comprehensive analysis result. Since the spectral transition reference value S tt and the high-frequency escape reference value Υ dix respectively reflect the transition amplitude of the audio in the spectral structure and the high-frequency energy perturbation characteristics, and their contributions to the change trend may be different in different acoustic scenarios, so it is necessary to set p m and p n to apply differential weighting in the calculation process.

[0106] For example, in a scenario where low-frequency perturbations are dominant in the environment, p mThe value of n makes the influence of spectral transition on the audio change coefficient more significant; while in an environment with frequent high-frequency mutations, the proportion of p

[0107] can be increased to make high-frequency escape dominate the system's judgment. Such preset proportional coefficients can be set by experience, adjusted by expert knowledge, or automatically fitted through machine learning model training. Their core role is to enhance the adaptability and sensitivity of the model to different change patterns. It is mentioned in the text that the mean of these two proportional coefficients is greater than 0 to ensure the positive contribution and perceptibility of the two features in the audio change coefficient, avoiding being ignored by the system or generating ineffective activations.

[0108] Compare and analyze the audio change coefficient generated by the machine learning model for intelligent prediction of the audio change situation in the external environment of the headset with the preset audio change coefficient reference threshold to divide the audio change of this sub-segment. The specific division steps are as follows:

[0109] If the audio change coefficient is greater than the audio change coefficient reference threshold, the audio change in the external environment of this sub-segment is divided into rapid change;

[0110] If the audio change coefficient is less than or equal to the audio change coefficient reference threshold, the audio change in the external environment of this sub-segment is divided into normal change.

[0111] When it is recognized that the external environment audio is in a rapid change state, start the noise reduction model switching engine, dynamically switch the working modes of feedback ANC and feedforward ANC, and adaptively adjust the switching frequency based on the acoustic environment change rate to achieve rapid response and precise noise reduction for complex acoustic disturbances;

[0112] When it is recognized that the external environment audio is in a rapid change state, start the noise reduction model switching engine, dynamically switch the working modes of feedback ANC and feedforward ANC, and adaptively adjust the switching frequency based on the acoustic environment change rate. The specific steps are as follows:

[0113] Trend modeling is performed on the change status of audio sub - segments within consecutive time periods. Specifically: After obtaining the audio change coefficients corresponding to each sub - segment (generated by a machine - learning model combining the spectral transition reference value and the high - frequency escape reference value), an audio perturbation trend function is constructed to measure the perturbation acceleration in the acoustic environment. The expression of the audio perturbation trend function is as follows:

[0114]

[0115] Where: represents the audio change coefficient corresponding to the (k - 1)-th frame sub - segment in the current sliding window (one frame behind the reference frame); represents the audio change coefficient of the k - th frame in the sliding window, which is the central value of the reference frame within this local window; represents the audio change coefficient of the (k + 1)-th frame in the sliding window (one frame in front of the reference frame); n is the number of frames in the sliding time window, and Ω t is the target perturbation trend intensity index. Through the construction of this function, the second - order change of the audio perturbation trend in the current environment can be quantitatively measured, reflecting whether the current sound field is in a trend of increasing perturbation;

[0116] The role of this step is to construct a perturbation trend index with high forward - looking and sensitivity based on the historical multi - frame audio change trend, providing a decision - making basis for the subsequent frequency control model.

[0117] After obtaining the target perturbation trend intensity index, the switching frequency between the feedback - type ANC and the feed - forward - type ANC is adaptively and dynamically adjusted. To achieve a balance between fast response and system stability, the switching frequency is adaptively and dynamically adjusted through the following switching - frequency regulation function:

[0118]

[0119] , where: f switch (t) represents the model switching frequency at the current moment, with the unit of times per second; f min is the minimum basic switching frequency set by the system in a stable environment; α p is the maximum frequency adjustment amplitude coefficient, representing the maximum frequency increase value that can be achieved in a rapidly changing environment; β r is the perturbation trend sensitivity adjustment coefficient, which determines the activation speed of the frequency response change by the degree of perturbation;

[0120] Through the above-mentioned frequency control function, when facing a higher disturbance trend, the system can quickly increase the switching frequency through an exponential response function, thereby enhancing its ability to respond to external complex acoustic disturbances; and when the disturbance trend slows down or stabilizes, the function automatically enters the saturation range, effectively suppressing frequency overshoot and ensuring the continuity of system operation and the stability of audio output.

[0121] The purpose of this step is to realize the noise reduction structural response mechanism based on adaptive adjustment of disturbance trend, so that the system can achieve the optimal model switching rhythm and response performance under conditions of different acoustic disturbance intensities.

[0122] The core function of this step is to achieve rapid adaptation and fine control of the active noise reduction system in a complex acoustic disturbance environment. When it is recognized that the audio signal of the external environment of the headset is in a rapidly changing state, it means that there may be sudden noise, drastic frequency fluctuations or energy transitions in the current sound field. At this time, the traditional fixed structure ANC (such as single feedback or feedforward) is difficult to strike a balance between timeliness and adaptability. By starting the "noise reduction model switching engine", the system can intelligently switch between feedback ANC (suitable for stable noise) and feedforward ANC (suitable for dynamic noise) to ensure that the noise reduction structure matches the environmental state. At the same time, the system dynamically adjusts the model switching frequency according to the rate of change of the acoustic environment, improves the response sensitivity of the system in high-disturbance scenarios, and avoids stability problems caused by frequent switching in low-disturbance scenarios, thereby optimizing the user's wearing experience while ensuring the noise reduction effect, and achieving a dual improvement in noise reduction performance and system energy efficiency.

[0123] Feedback Active Noise Cancellation (FAN) and feedforward Active Noise Cancellation (FAN) are two basic working modes in active noise reduction systems. The core difference between them lies in the position of the microphone and the timing of noise reduction control. Feedforward ANC collects noise signals in the external environment through a microphone set outside the headset. Before the noise enters the ear canal, the system uses the noise reduction processor to generate reverse sound waves with opposite phases to cancel it in advance. It is suitable for active suppression of sudden and highly structured external noise (such as traffic noise, wind noise, etc.). Feedback ANC collects the residual noise actually heard by the user through a microphone set inside the earmuff or near the ear canal, and feeds it back to the control system in real time for reverse sound wave adjustment to continuously correct the noise reduction effect. It is more suitable for processing low- and medium-frequency continuous noise (such as air conditioning sound, engine sound, etc.). These two modes have their own advantages. The feedforward mode has a fast response and a wide range, while the feedback mode has high accuracy and fits the actual listening experience. Reasonable switching in a dynamic noise environment can achieve a more stable and intelligent noise reduction control strategy.

[0124] Through the above method for optimizing the noise reduction effect of headphones, it is possible to achieve high-precision recognition of a rapidly changing acoustic environment and intelligent switching control of a dynamic noise reduction structure, significantly improving the response speed and noise reduction stability of the active noise reduction system in complex noise scenarios. Specifically, the system performs continuous sampling, short-time spectrum analysis, and dynamic feature extraction on the audio signal, and combines a machine learning model to intelligently predict the trend of acoustic disturbances, enabling the noise reduction decision to change from "passive response" to "active prediction". When a rapidly changing acoustic disturbance is detected, the system can promptly switch to a more suitable noise reduction structure (feedforward or feedback mode) and adaptively adjust the switching frequency based on the audio disturbance rate, effectively avoiding "noise reduction failure" or "boom feeling" caused by phase lag and structural mismatch. This solution not only improves the overall noise reduction performance but also optimizes the user's auditory comfort and usage experience in a changing sound field, with strong practicality and technical promotion value.

[0125] The present invention provides a system for optimizing the noise reduction effect of headphones as shown in Figure 2 which includes an audio acquisition and window framing module, a spectrum extraction and time-frequency transformation module, a feature analysis and intelligent prediction module, and a model switching and frequency adaptive control module;

[0126] The audio acquisition and window framing module uses the built-in microphone of the headphones to collect ambient sound at a fixed sampling rate, and uses a sliding time window to divide the audio stream into multiple sub-fragments to form a continuous sequence of acoustic frames;

[0127] The spectrum extraction and time-frequency transformation module performs short-time Fourier transform processing on each divided sub-fragment to extract the spectrum energy distribution within the sub-fragment;

[0128] The feature analysis and intelligent prediction module extracts characteristic indicators representing rapid audio changes from the spectrum energy distribution map for each sub-fragment. After in-depth analysis of the extracted characteristic indicators, the analyzed characteristic indicators are used as feature vectors and input into a pre-trained machine learning model, and the machine learning model is used to intelligently predict the audio change situation in the external environment of the headphones;

[0129] The model switching and frequency adaptive control module, when it recognizes that the external environment audio is in a rapidly changing state, starts the noise reduction model switching engine, dynamically switches the working modes of feedback-type ANC and feedforward-type ANC, and adaptively adjusts the switching frequency based on the change rate of the acoustic environment;

[0130] The method for optimizing the noise reduction effect of headphones provided by the embodiments of the present invention is implemented through the above system for optimizing the noise reduction effect of headphones. The specific methods and processes of the system for optimizing the noise reduction effect of headphones are detailed in the embodiments of the above method for optimizing the noise reduction effect of headphones, and will not be elaborated here.

[0131] Only some exemplary embodiments of the present invention have been described above by way of illustration. Without doubt, for those of ordinary skill in the art, the described embodiments can be modified in various different ways without departing from the spirit and scope of the present invention. Therefore, the above drawings and description are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.

Claims

1. A method for optimizing the noise reduction effect of headphones, characterized in that, It includes the following steps: Use the built-in microphone of the headset to collect environmental sounds at a fixed sampling rate, and use a sliding time window to divide the audio stream into multiple sub-fragments to form a continuous sequence of acoustic frames; Perform short-time Fourier transform processing on each divided sub-fragment to extract the spectral energy distribution within the sub-fragment; For each sub-fragment, extract the characteristic indicators representing the rapid changes in the audio from the spectral energy distribution diagram. After deeply analyzing the extracted characteristic indicators, use the analyzed characteristic indicators as feature vectors and input them into a pre-trained machine learning model. Through the machine learning model, make an intelligent prediction of the audio change situation in the external environment of the headset; When it is recognized that the external environment audio is in a rapidly changing state, start the noise reduction model switching engine, dynamically switch the working modes of feedback ANC and feedforward ANC, and adaptively adjust the switching frequency based on the acoustic environment change rate.

2. The method for optimizing the noise reduction effect of the earphone according to claim 1, wherein Using a sliding time window to divide the audio stream into multiple sub-fragments and form a continuous sequence of acoustic frames specifically includes the following steps: S1. Set a time window with a fixed length as the time span of each sub-fragment; S2. Starting from the audio acquisition starting point, intercept the first sub-fragment from the continuous audio stream in units of this time window; S3. Slide the time window backward according to the set sliding step length, and intercept the next sub-fragment starting from the current position; S4. Repeat S3 in sequence to intercept multiple time-overlapping sub-fragments, and finally construct a continuous acoustic frame sequence covering the entire time period.

3. A method for optimizing the noise reduction effect of a headset according to claim 1, characterized in that, Perform short-time Fourier transform processing on each divided sub-fragment to extract the spectral energy distribution within the sub-fragment. The specific steps are as follows: Use the time-domain audio signal of each sub-fragment as the input, apply a window function to it to reduce the interference of edge effects on the spectrum; Use the fast Fourier transform to perform a frequency-domain transformation on the windowed signal to convert the time-domain signal into a spectrum containing complex coefficients; Calculate the square of the amplitude of each frequency point in the spectrum to obtain the energy value of the corresponding frequency component, and form an energy spectrum diagram of the sub-fragment within a certain time interval; Arrange the spectral results of multiple sub-fragments in chronological order to construct a two-dimensional spectral energy distribution matrix reflecting the audio change over time.

4. A method for optimizing the noise reduction effect of an earphone according to claim 1, characterized in that For each sub-fragment, extract the characteristic indicators representing the rapid changes in the audio from the spectral energy distribution diagram. Among them, the extracted characteristic indicators include the degree of change in the frequency energy distribution between adjacent spectral frames and the ratio change of the energy jumping from the mid-low frequency band to the high frequency band. After deeply analyzing the extracted characteristic indicators, generate a spectral transition reference value and a high-frequency escape reference value respectively. The overall fluctuation intensity of the audio signal in the frequency structure is characterized by the spectral transition reference value, and the density and trend of the sudden high-frequency energy surge phenomenon in the continuous time domain are characterized by the high-frequency escape reference value.

5. A method for optimizing the noise reduction effect of a headset, characterized in that, Use the analyzed spectral transition reference value and high-frequency escape reference value as feature vectors and input them into a pre-trained machine learning model. Through the machine learning model, output the audio change coefficient, and make an intelligent prediction of the audio change situation in the external environment of the headset based on the audio change coefficient.

6. A method for optimizing the noise reduction effect of an earphone according to claim 5, characterized in that, When the audio change coefficient generated by the machine learning model for the intelligent prediction of the audio change of the external environment of the earphone is compared with the preset audio change coefficient reference threshold, the audio change of this sub - segment is divided. The specific division steps are as follows: If the audio change coefficient is greater than the audio change coefficient reference threshold, the external environment audio change of this sub - segment is divided into rapid change; If the audio change coefficient is less than or equal to the audio change coefficient reference threshold, the external environment audio change of this sub - segment is divided into normal change.

7. A method for optimizing the noise reduction effect of an earphone according to claim 6, wherein, When it is recognized that the external environment audio is in a rapid change state, start the noise reduction model switching engine, dynamically switch the working modes of feedback ANC and feed - forward ANC, and adaptively adjust the switching frequency based on the acoustic environment change rate. The specific steps are as follows: Conduct trend modeling on the change state of audio sub - segments within a continuous time period. Specifically, after obtaining the audio change coefficients corresponding to each sub - segment, construct an audio perturbation trend function to measure the perturbation acceleration in the acoustic environment. The expression of the audio perturbation trend function is as follows: , Wherein: represents the audio change coefficient corresponding to the (k - 1)-th frame sub-segment in the current sliding window; represents the audio change coefficient of the k-th frame in the sliding window, which is the reference frame center value within this local window; represents the audio change coefficient of the (k + 1)-th frame in the sliding window; n is the number of frames in the sliding time window, and Ω t is the target perturbation trend intensity index; After obtaining the target perturbation trend intensity index, adaptively and dynamically adjust the switching frequency between feedback ANC and feed - forward ANC. The expression of the adaptive dynamic adjustment is: , Where: f switch (t) represents the model switching frequency at the current moment; f min is the minimum basic switching frequency set by the system in a stable environment; α p is the maximum frequency adjustment amplitude coefficient, indicating the maximum frequency increase value that can be achieved in a rapidly changing environment; β r is the disturbance trend sensitive adjustment coefficient.

8. A method for optimizing the noise reduction effect of an earphone according to claim 4, characterized in that, For each sub - segment, the specific steps to generate a spectrum transition reference value after deeply analyzing the change degree of the frequency energy distribution between adjacent spectrum frames are as follows: Express the short-time Fourier transform result in this sub-segment as a two-dimensional spectral energy matrix F, F = {F i,j}, where i represents the i-th frame and j represents the j-th frequency point, and F i,j represents the energy component of the i-th frame at the j-th frequency. Then, calculate the spectral difference degree between each pair of adjacent frames. The calculation expression is as follows: , Where: D i represents the spectral difference degree between the i-th frame and the (i + 1)-th frame; N is the spectral dimension; λ is the spectral difference normalization factor, which suppresses the suppression of the difference by the high-amplitude main frequency; After obtaining the inter - frame spectrum difference sequence, calculate the spectrum transition reference value based on the ratio relationship between the maximum amplitude of spectral perturbation and the weighted overall perturbation. The specific calculation formula is as follows: , Where: S tt is the spectral transition reference value, used to characterize the frequency structure fluctuation intensity within this sub - segment; max(D i ) represents the spectral mutation amplitude between the strongest pair of frames; M is the total number of frames in this sub - segment; ω i is the modulation weight, emphasizing the perturbation sensitivity in the middle of the i - th frame sequence; ∈ is the minimum stable constant.

9. A method for optimizing the noise reduction effect of a headset, characterized in that, For each sub - segment, the specific steps to generate a high - frequency escape reference value after deeply analyzing the proportion change of the energy jumping from the mid - low frequency band to the high - frequency band are as follows: In the spectrum energy distribution diagram corresponding to each sub - segment, set two frequency band sets: the mid - low frequency energy region A and the high - frequency energy region B. By calculating the total energy belonging to the high - frequency region and the total energy belonging to the mid - low frequency region in this sub - segment, construct a high - frequency transition energy ratio factor. The constructed expression is: , Where: E B is the cumulative energy of the high-frequency region, reflecting the intensity of the burst enhancement; E A is the cumulative energy of the medium-low frequency region, used to measure the background energy baseline; E B 2 is used to enhance the sensitivity of the sudden high-frequency escape behavior; Λ is the high-frequency transition energy ratio factor of the current segment; After obtaining the high - frequency transition energy ratio factor, further perform a structural non - linear mapping on it to construct a high - frequency escape reference value for enhancing the expression of the directional surge trend of the transition phenomenon. The construction expression of the high - frequency escape reference value is: Υ dix = |tanh(Λ - Λ0)| δ , Among them: Λ0 is the transition stability threshold set by the system, representing the reasonable distribution boundary of high-frequency and medium-low-frequency energies in a normal acoustic structure; tanh(*) is used to limit extreme values and enhance the nonlinear response; δ is the trend enhancement factor, used to regulate the sensitivity of the system to the surge trend; Υ dix is the high-frequency escape reference value.

10. A system for optimizing the noise reduction effect of headphones, which is used to implement the method for optimizing the noise reduction effect of headphones described in any one of the above claims 1-9, characterized in that, Including an audio acquisition and window framing module, a spectrum extraction and time - frequency transformation module, a feature analysis and intelligent prediction module, and a model switching and frequency adaptive control module; The audio acquisition and window framing module uses the built - in microphone of the earphone to collect environmental sounds at a fixed sampling rate, and uses a sliding time window to divide the audio stream into multiple sub - segments to form a continuous acoustic frame sequence; The spectrum extraction and time - frequency transformation module performs short - time Fourier transform processing on each divided sub - segment to extract the spectrum energy distribution within this sub - segment; Feature analysis and intelligent prediction module. For each sub - segment, extract feature indicators representing the rapid changes of audio from the spectral energy distribution map. After deeply analyzing the extracted feature indicators, use the analyzed feature indicators as feature vectors and input them into a pre - trained machine learning model, and through the machine learning model, intelligently predict the audio change situation of the external environment of the earphone; Model switching and frequency adaptive control module. When it is recognized that the external environment audio is in a rapid change state, start the noise reduction model switching engine, dynamically switch the working modes of feedback ANC and feed - forward ANC, and adaptively adjust the switching frequency based on the change rate of the acoustic environment.