AI noise reduction speech enhancement method for multi-mode wireless earphone

By extracting acoustic features and analyzing neural networks from wireless headphones, a collaborative control vector is generated, which solves the problem of inconsistent auditory experience in complex environments and optimizes auditory comfort and sound quality consistency in multiple modes.

CN122269185APending Publication Date: 2026-06-23SHENZHEN BOLUKE ELECTRONIC TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN BOLUKE ELECTRONIC TECH CO LTD
Filing Date
2026-05-06
Publication Date
2026-06-23

AI Technical Summary

Technical Problem

Existing wireless headphones struggle to achieve smooth coordination and intelligent adaptation between different operating modes in complex and ever-changing acoustic environments and dynamic usage scenarios, resulting in inconsistent auditory experiences and limited system adaptability.

Method used

By synchronously acquiring feedforward microphone signals, feedback microphone signals, and reference audio signals, acoustic features are extracted and analyzed to generate wind noise scene description vectors, sound quality deviation vectors, and multi-objective optimization weight vectors. A pre-trained neural network model is then used to generate collaborative control vectors for adaptive noise reduction and speech enhancement processing.

Benefits of technology

It achieves the maintenance of auditory comfort, mode target fidelity, and timbre consistency in multiple working modes and complex environments, thereby improving user experience, system decision rationality, and auditory experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122269185A_ABST
    Figure CN122269185A_ABST
Patent Text Reader

Abstract

The application discloses an AI noise reduction voice enhancement method of a multi-mode wireless earphone, and relates to the field of audio signal processing, which comprises the following steps: synchronously collecting feedforward, feedback microphone signals and reference audio signals, and obtaining a working mode instruction, and obtaining an acoustic feature vector after processing; performing wind noise analysis based on the vector to generate a wind noise scene description vector, simultaneously generating a sound quality deviation degree vector based on the feedback and reference signals, and analyzing a multi-objective optimization weight vector based on the working mode; calculating a mode fidelity and comfort degree conflict factor based on the wind noise scene description vector and the multi-objective optimization weight vector; splicing all the vectors and the conflict factor into an enhancement system state vector, inputting the enhancement system state vector into a second neural network model which is pre-trained, and outputting a cooperative control vector which is used for coordinating downstream processing; and performing adaptive noise reduction, voice enhancement or environmental sound mixing processing and dynamic equalization compensation based on the cooperative control vector, and synthesizing a final output signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio signal processing, specifically to an AI noise reduction and voice enhancement method for multimode wireless headphones. Background Technology

[0002] With the continuous advancement of wireless communication technology and digital audio processing algorithms, wireless headphones have become the mainstream personal audio device. To adapt to diverse usage scenarios, modern wireless headphones generally support multiple operating modes, such as active noise cancellation, ambient sound transparency, and voice call, aiming to provide users with an immersive listening experience, clear environmental awareness, and high-quality voice communication, respectively. Currently, the implementation of these functions mainly relies on a series of mature audio signal processing technologies. Active noise cancellation typically employs a combined feedforward and feedback architecture, utilizing adaptive filtering algorithms to generate anti-noise signals with the opposite phase to the noise, thereby achieving acoustic cancellation. Ambient sound transparency is achieved by mixing the external sound collected by the feedforward microphone after gain and phase processing into the audio path. In call mode, beamforming technology is often used to spatially filter the signal received by the microphone array, thereby enhancing the voice from the target direction and suppressing interference.

[0003] Despite the widespread application of these technologies, there remains room for further exploration in achieving smooth coordination and intelligent adaptation between different operating modes in complex and ever-changing real-world acoustic environments and dynamic usage scenarios. In existing solutions, the processing links corresponding to each mode are typically designed and tuned relatively independently or simply in series. When faced with strong interference such as wind noise or frequent mode switching, the system still faces challenges in achieving real-time, dynamic, globally optimal coordination across multiple dimensions, including maintaining timbre consistency across all modes, achieving acoustic goals for specific modes, and ensuring auditory comfort. Therefore, researching how to construct an intelligent system that can more comprehensively perceive the acoustic environment, more accurately understand user intentions, and thereby coordinate and optimize audio processing parameters across the entire link, thus providing multi-mode wireless headphones with a consistent, comfortable, and high-quality listening experience in various scenarios, is currently an important development direction in this field. Summary of the Invention

[0004] Based on the shortcomings of the prior art described above, the purpose of this invention is to provide an AI noise reduction and voice enhancement method for multi-mode wireless headphones to solve the aforementioned technical problems.

[0005] To achieve the above objectives, the present invention provides the following technical solution: an AI noise reduction and voice enhancement method for multi-mode wireless headphones, comprising: S1: Synchronously acquire feedforward microphone signals, feedback microphone signals, and reference audio signals, and simultaneously obtain working mode commands; preprocess and extract features from the feedforward microphone signals, feedback microphone signals, and reference audio signals to obtain acoustic feature vectors; S2: Analyze wind noise characteristics based on acoustic feature vectors to generate a wind noise scene description vector; analyze output frequency response and distortion based on feedback microphone signals and reference audio signals to generate a sound quality deviation vector; parse target mode intent based on working mode commands to generate a multi-objective optimization weight vector. S3: Calculate the conflict factor between mode fidelity and comfort based on the wind noise scene description vector and the multi-objective optimization weight vector; S4: Concatenate the wind noise scene description vector, sound quality deviation vector, multi-objective optimization weight vector, and mode fidelity and comfort conflict factor to generate an enhanced system state vector; input the enhanced system state vector into the pre-trained second neural network model to output a cooperative control vector; S5: Based on the cooperative control vector, adaptive noise reduction, speech enhancement or ambient sound mixing and dynamic equalization compensation are performed on the feedforward microphone signal and the feedback microphone signal, and the processed signals are synthesized into the final output signal.

[0006] The present invention is further configured such that S1 includes: At a preset sampling rate and resolution, the system simultaneously acquires feedforward microphone signals, feedback microphone signals, and reference audio signals, and simultaneously obtains working mode commands. High-pass filtering, frame-segmentation and windowing, and time-frequency transformation are performed on the feedforward microphone signal, feedback microphone signal, and reference audio signal respectively to convert each signal into its corresponding frequency domain complex spectrum. Subband energy, spectral centroid, spectral flatness, and zero-crossing rate are extracted from the complex frequency domain spectra of the feedforward microphone signal, feedback microphone signal, and reference audio signal, respectively, as underlying acoustic description features; based on the complex frequency domain spectra of the feedforward microphone signal and feedback microphone signal, the complex spectral coherence coefficients of the two signals in the preset frequency band are calculated. The underlying acoustic descriptive features are combined with the complex spectral coherence coefficients to form an acoustic feature vector.

[0007] The present invention is further configured such that S2 includes: Based on acoustic feature vectors, a pre-trained first neural network model is used to identify and quantify wind noise characteristics, generating a wind noise scene description vector; the wind noise scene description vector includes a wind noise existence probability scalar, wind noise intensity level, wind noise spectrum energy distribution vector, and wind noise impact index. At each effective frequency point, the complex spectrum of the feedback microphone signal in the frequency domain is divided by the complex spectrum of the reference audio signal in the frequency domain to obtain the real-time frequency response curve characterizing the frequency response characteristics of the current system; the average amplitude deviation between the real-time frequency response curve and the target frequency response curve preset for the current working mode in the preset evaluation frequency band is calculated, and a sound quality deviation vector is generated based on each average amplitude deviation. Based on the working mode command, a multi-objective optimization weight vector is obtained by parsing according to the preset mapping rules; the multi-objective optimization weight vector includes noise suppression weight, voice communication weight, ambient sound transparency weight, and timbre consistency weight.

[0008] The present invention is further configured such that the effective frequency point is determined by judging whether the energy of the complex spectrum of the frequency domain of the reference audio signal is higher than a preset energy threshold.

[0009] The present invention is further configured such that S3 includes: The current working mode is determined based on the multi-objective optimization weight vector, and the intention frequency band range and the auditory discomfort frequency band range corresponding to the current working mode are determined according to the preset acoustic knowledge base. Based on the wind noise spectral energy distribution vector in the wind noise scene description vector, the feedback microphone signal and the reference audio signal are respectively multiplied by the wind noise energy value of each frequency point within the intended frequency band by the preset corresponding frequency point importance weight, and all product results are summed to obtain the first conflict energy value; the wind noise energy value of each frequency point within the hearing discomfort frequency band is multiplied by the preset corresponding frequency point sensitivity weight, and all product results are summed to obtain the second conflict energy value; Logarithmic operations are performed on the first and second conflict energy values ​​respectively, and the results of the two logarithmic operations are weighted by the first and second preset coefficients related to the current working mode. The two weighted results are then added to the value obtained by multiplying the wind noise impact index in the wind noise scene description vector by the third preset coefficient to generate the mode fidelity and comfort conflict factor.

[0010] The present invention is further configured such that S4 includes: The wind noise scene description vector, sound quality deviation vector, multi-objective optimization weight vector, and mode fidelity and comfort conflict factor are concatenated to generate an enhanced system state vector. The enhanced system state vector is input into a pre-trained second neural network model, which performs forward propagation calculations to output a cooperative control vector; wherein, the cooperative control vector includes a noise reduction control sub-vector, a beamforming control sub-vector, an equalization control sub-vector, and a hybrid control sub-vector.

[0011] The present invention is further configured such that: the noise reduction control sub-vector is used to configure the type of the noise reduction filter and its noise reduction depth in each frequency band; the beamforming control sub-vector is used to configure the beam pointing and null depth when beamforming the feedforward microphone signal; the equalization control sub-vector is used to configure the gain compensation value of the dynamic equalizer in multiple adjustable frequency bands; and the mixing control sub-vector is used to configure the mixing ratio and sidetone gain of the ambient sound pass-through.

[0012] The present invention is further configured such that S5 includes: Based on the noise reduction control sub-vector in the cooperative control vector, adaptive noise reduction processing is performed on the feedforward microphone signal and the feedback microphone signal. Based on the current working mode, the feedforward microphone signal is subjected to speech enhancement processing or ambient sound mixing processing based on the beamforming control sub-vector or hybrid control sub-vector in the cooperative control vector. Based on the equalization control sub-vector in the cooperative control vector, dynamic equalization compensation is performed on the signal after adaptive noise reduction processing, speech enhancement processing, or ambient sound mixing processing. The signal after adaptive noise reduction, speech enhancement or ambient sound mixing and dynamic equalization compensation is synthesized to generate the final output signal.

[0013] The present invention is further configured such that the sampling rate is not less than 16 kHz.

[0014] The present invention is further configured such that the resolution is not less than 16 bits.

[0015] This invention provides an AI noise reduction and voice enhancement method for multi-mode wireless headphones. The method comprises: S1: Simultaneously acquiring feedforward microphone signals, feedback microphone signals, and reference audio signals, and simultaneously obtaining working mode commands; preprocessing and feature extraction of the feedforward microphone signals, feedback microphone signals, and reference audio signals to obtain acoustic feature vectors; S2: Analyzing wind noise characteristics based on the acoustic feature vectors to generate a wind noise scene description vector; analyzing output frequency response and distortion based on the feedback microphone signals and reference audio signals to generate a sound quality deviation vector; and parsing the target mode intent based on the working mode commands to generate a multi-objective optimization weight vector; S3: Based on the wind noise scene description vector and the multi-objective optimization weight vector, the mode fidelity and comfort conflict factor is calculated; S4: The wind noise scene description vector, sound quality deviation vector, multi-objective optimization weight vector, and mode fidelity and comfort conflict factor are concatenated to generate the enhanced system state vector; The enhanced system state vector is input into the pre-trained second neural network model to output the cooperative control vector; S5: Based on the cooperative control vector, adaptive noise reduction, speech enhancement, or ambient sound mixing processing and dynamic equalization compensation are performed on the feedforward microphone signal and the feedback microphone signal, and the processed signals are synthesized into the final output signal. The beneficial effects include: 1. By constructing a complete processing flow from synchronous acquisition and feature extraction, environmental state and user intent perception, conflict prediction, intelligent collaborative decision-making to parameterized dynamic execution, the system solves the problems of inconsistent experience when switching between different working modes of multi-mode wireless headphones and limited system adaptability in complex acoustic environments such as wind noise. By introducing a conflict factor between mode fidelity and comfort, the system achieves a quantitative assessment of the difficulties and auditory costs faced in executing the mode target in the current scenario. This allows the system to make strategic adjustments based on the assessment, thereby achieving an optimized balance between auditory comfort, mode target fidelity, and timbre consistency in various working modes and complex environments including strong wind noise, thus improving the user experience. 2. By determining the intended frequency band range and the auditory discomfort frequency band range corresponding to the current working mode, the weighted conflict energy in the above frequency bands is calculated based on the wind noise spectrum energy distribution. Combined with the wind noise impact index to synthesize the mode fidelity and comfort conflict factor, a quantitative assessment of the degree of contradiction between achieving the mode acoustic goal and maintaining auditory comfort is achieved. This mode fidelity and comfort conflict factor provides prior information for subsequent neural network decision-making, enabling the system to adjust the strategy emphasis and adjustment degree of multi-objective optimization according to its value, thereby improving the rationality of decision-making and the adaptive ability of auditory experience in complex scenarios. 3. By concatenating the wind noise scene description vector, sound quality deviation vector, multi-objective optimization weight vector, and mode fidelity and comfort conflict factor to form an enhanced system state vector, a structured collaborative control vector is generated through forward propagation using a pre-trained second neural network model. This collaborative control vector includes noise reduction control, beamforming control, equalization control, and hybrid control sub-vectors, which can perform synchronous and coordinated parameterized configuration of downstream audio processing units, realizing centralized collaborative decision-making. The clear division of sub-vector functions makes the control commands clear and the module interfaces clear, facilitating system implementation, debugging, and optimization.

[0016] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings: Figure 1The flowchart illustrates an AI noise reduction and voice enhancement method for multimode wireless headphones, as an exemplary embodiment of the present invention. Detailed Implementation

[0018] The embodiments of the present invention will be described below with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are only for illustrating the present invention and not for limiting the scope of protection of the present invention.

[0019] It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Therefore, the drawings only show the components related to the present invention and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.

[0020] In the following description, numerous details are explored to provide a more thorough explanation of embodiments of the invention. However, it will be apparent to those skilled in the art that embodiments of the invention may be practiced without these specific details. In other embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring embodiments of the invention.

[0021] A method for AI noise reduction and voice enhancement in multi-mode wireless headphones, such as Figure 1 As shown, it includes: S1: Synchronously acquire feedforward microphone signals, feedback microphone signals, and reference audio signals, and simultaneously obtain working mode commands; preprocess and extract features from the feedforward microphone signals, feedback microphone signals, and reference audio signals to obtain acoustic feature vectors; S2: Analyze wind noise characteristics based on acoustic feature vectors to generate a wind noise scene description vector; analyze output frequency response and distortion based on feedback microphone signals and reference audio signals to generate a sound quality deviation vector; parse target mode intent based on working mode commands to generate a multi-objective optimization weight vector. S3: Calculate the conflict factor between mode fidelity and comfort based on the wind noise scene description vector and the multi-objective optimization weight vector; S4: Concatenate the wind noise scene description vector, sound quality deviation vector, multi-objective optimization weight vector, and mode fidelity and comfort conflict factor to generate an enhanced system state vector; input the enhanced system state vector into the pre-trained second neural network model to output a cooperative control vector; S5: Based on the cooperative control vector, adaptive noise reduction, speech enhancement or ambient sound mixing and dynamic equalization compensation are performed on the feedforward microphone signal and the feedback microphone signal, and the processed signals are synthesized into the final output signal.

[0022] The present invention is further configured such that S1 includes: The present invention synchronously acquires feedforward microphone signals, feedback microphone signals, and reference audio signals at a preset sampling rate and resolution, and synchronously obtains working mode instructions; the present invention is further configured such that the sampling rate is not less than 16kHz; the present invention is further configured such that the resolution is not less than 16bit. High-pass filtering, frame-segmentation and windowing, and time-frequency transformation are performed on the feedforward microphone signal, feedback microphone signal, and reference audio signal respectively to convert each signal into its corresponding frequency domain complex spectrum. Subband energy, spectral centroid, spectral flatness, and zero-crossing rate are extracted from the complex frequency domain spectra of the feedforward microphone signal, feedback microphone signal, and reference audio signal, respectively, as underlying acoustic description features; based on the complex frequency domain spectra of the feedforward microphone signal and feedback microphone signal, the complex spectral coherence coefficients of the two signals in the preset frequency band are calculated. The underlying acoustic description features are combined with complex spectral coherence coefficients to form an acoustic feature vector. Specifically, time-domain signals from the feedforward physical microphone channel, the feedback physical microphone channel, and the reference audio electrical signal channel are simultaneously acquired at a sampling rate of at least 16 kHz and a resolution of at least 16 bits. The feedforward microphone signal is generated by a sensor used to pick up ambient sound and user voice from the headphones. The feedback microphone signal is generated by a sensor located in the ear canal to acquire the synthesized sound field after speaker playback and ear canal acoustic modulation. The reference audio signal is directly taken from the raw audio data stream to be fed to the speaker for electroacoustic conversion. Simultaneously, the signal obtained in real time from user-interactive operations is acquired by accessing designated registers or software interfaces of the main control chip. The system generates digital operating mode instruction codes. These codes are synchronized with the time-domain signal acquisition, and their values ​​uniquely correspond to the specific audio processing state selected by the user. The defined states include at least a noise reduction mode (focused on noise reduction), a transparency mode (focused on transmitting ambient sound), and a call mode (focused on optimizing voice communication quality). This synchronization mechanism ensures that at any given processing moment, multiple audio samples are strictly aligned with the user's intent, providing a unified time reference for subsequent algorithms that rely on multi-source information fusion. The three acquired raw time-domain audio signals then enter a standardized preprocessing flow, which sequentially performs high-pass filtering, frame windowing, and time-frequency transformation. First, in the high-pass filtering step, each channel... First, a high-pass digital filter with a preset cutoff frequency is applied to the signal to filter out ultra-low frequency noise components that do not carry effective auditory information and DC bias components. Second, in the frame-segmentation and windowing process, the filtered continuous time-domain signal is segmented according to a preset time length, with each segment called an analysis frame. Each time-domain signal sample point within each analysis frame is multiplied by a pre-calculated sequence of window function coefficients, thus suppressing spectral energy leakage caused by truncating the continuous signal. Finally, in the time-frequency transformation step, a fast Fourier transform algorithm is executed on each windowed time-domain signal frame to transform the signal from a time-domain representation to a frequency-domain representation, obtaining the complex frequency spectrum corresponding to that signal frame. This complex spectrum is represented in array form. The formula fully describes the amplitude and phase information of the signal at each frequency point. Based on the frequency domain complex spectrum obtained after preprocessing, for each signal in the feedforward microphone signal, feedback microphone signal and reference audio signal, four basic characteristic parameters are independently calculated for each frame of its frequency domain complex spectrum: subband energy, spectral centroid, spectral flatness and zero crossing rate. These constitute the underlying acoustic description features that describe the basic acoustic properties of the signal. The calculation process of subband energy is to divide the main frequency range of human hearing perception into multiple continuous frequency bands uniformly or non-uniformly. For each subband, the square values ​​of the amplitudes corresponding to all frequency points in the frequency domain complex spectrum of the signal that fall within the subband are accumulated and summed. The summation result is represented as the energy value of the subband.The calculation of the spectral centroid involves obtaining the weighted average frequency of the entire complex spectrum in the analysis frequency domain, where the weights are the amplitude energy values ​​corresponding to each frequency point. This parameter reflects the concentration trend of spectral energy along the frequency axis. The calculation of spectral flatness involves obtaining a ratio by comparing the geometric mean and arithmetic mean of the signal's complex spectrum in the frequency domain. This ratio is used to quantify the harmonic or noise characteristics of the spectrum. The calculation of the zero-crossing rate involves tracing back to the original time-domain waveform corresponding to the signal frame and obtaining the rate by counting the number of times the time-domain signal waveform crosses the zero-level axis. This parameter can be used to roughly characterize the dominance frequency of the signal. Further... The method calculates a joint feature characterizing the spatial acoustic relationship between the feedforward microphone signal and the feedback microphone signal, namely the complex spectral coherence coefficient. This calculation is based on the complex frequency domain spectra of the feedforward microphone signal and the feedback microphone signal, and is performed separately on several preset frequency bands that are crucial for active noise cancellation performance and environmental condition assessment. For each preset frequency band, the statistical correlation within that band is quantified by calculating the ratio of the covariance between the complex spectral values ​​of the two signals at the same frequency point to the square root of their respective variance products, thereby generating a coefficient between zero and... The real numerical results between 1 and 0; when the complex spectral coherence coefficient value approaches 1, it indicates that the complex spectra of the two signals are highly correlated in the preset frequency band, indicating that the propagation path of sound from the external environment to the ear canal is highly consistent; when the complex spectral coherence coefficient value approaches zero, it indicates that the correlation between the two signals in this frequency band is significantly reduced. This phenomenon is often caused by the loss of connection between the feedforward and feedback signal paths due to special interference such as wind noise; after completing the extraction of all single-channel features and dual-channel joint features, the feature combination step is performed, that is, according to a predefined and fixed splicing order, all the low-level features extracted from the feedforward microphone signal are combined. The acoustic descriptive features, all low-level acoustic descriptive features extracted from the feedback microphone signal, all low-level acoustic descriptive features extracted from the reference audio signal, and the calculated complex spectral coherence coefficients of the feedforward and feedback microphone signals are all sequentially concatenated to form a high-dimensional, comprehensive one-dimensional digital array. This array integrates information from different physical sensors, covers different analytical dimensions in the time and frequency domains, and includes information on the relationships between the two channels. This array is the acoustic feature vector. This acoustic feature vector provides a comprehensive and well-organized data foundation for subsequent refined environmental analysis, real-time sound quality assessment, and user intent interpretation.

[0023] The present invention is further configured such that S2 includes: Based on acoustic feature vectors, a pre-trained first neural network model is used to identify and quantify wind noise characteristics, generating a wind noise scene description vector; the wind noise scene description vector includes a wind noise existence probability scalar, wind noise intensity level, wind noise spectrum energy distribution vector, and wind noise impact index. At each effective frequency point, the complex frequency spectrum of the feedback microphone signal is divided by the complex frequency spectrum of the reference audio signal to obtain a real-time frequency response curve characterizing the current system frequency response characteristics; the average amplitude deviation between the real-time frequency response curve and the target frequency response curve preset for the current working mode in the preset evaluation frequency band is calculated, and a sound quality deviation vector is generated based on each average amplitude deviation; the present invention is further configured such that the effective frequency point is determined by judging whether the energy of the complex frequency spectrum of the reference audio signal is higher than a preset energy threshold; Based on the working mode instructions, a multi-objective optimization weight vector is obtained by parsing according to preset mapping rules. This multi-objective optimization weight vector includes noise suppression weights, voice communication weights, ambient sound transparency weights, and timbre consistency weights. Specifically, the wind noise characteristic analysis process uses the acoustic feature vector generated in step S1 as input. This acoustic feature vector contains low-level acoustic descriptive features extracted from feedforward and feedback microphone signals, such as sub-band energy distribution, spectral flatness measurement, and complex spectral coherence coefficient. The core of this process is a pre-trained offline first neural network model, whose function is to automatically identify complex acoustic patterns of wind noise through learning. This first neural network model is constructed as a feedforward deep neural network structure. The number of neurons in its input layer precisely corresponds to the dimension of the acoustic feature vector to ensure complete reception of all input feature data. The model contains at least one hidden layer, each composed of multiple neurons. Each node integrates a non-linear activation function to perform layer-by-layer, non-linear mathematical transformations and higher-order feature abstractions on the input feature data. The number of neurons in its output layer is specifically configured to perfectly match the total dimension of the desired structured wind noise scene description vector. During operation, this first neural network model maps the input high-dimensional hybrid acoustic features to a wind noise scene description vector through a forward propagation calculation process defined by its internal connection weights and activation functions. This wind noise scene description vector contains… The four quantitative components are: the first component is a wind noise presence probability scalar, whose value is a continuous value between zero and one, used to objectively quantify the possibility of wind noise interference in the current environment; the second component is the wind noise intensity level, whose output is a discrete integer value determined according to a preset intensity threshold, used to identify different intensity levels such as no wind, light wind, moderate wind, and strong wind; the third component is the wind noise spectrum energy distribution vector, whose dimension is consistent with the sub-band division of the acoustic feature vector. Each element in the wind noise spectrum energy distribution vector is a normalized wind noise energy value, used to describe the relative distribution of wind noise energy on different frequency sub-bands, where the sub-band corresponding to the maximum value indicates the main concentrated frequency band of wind noise energy; the fourth component is the wind noise impact index, which is the wind noise impact... The frequency response index is calculated based on underlying acoustic characteristics that reflect the time-varying transient properties of a signal, such as the degree of fluctuation in the zero-crossing rate within a short time window. It is used to characterize the temporal stability or sudden impact of wind noise, and a higher value indicates a stronger impact characteristic. The output frequency response and distortion analysis process takes the complex frequency spectrum of the feedback microphone signal and the complex frequency spectrum of the reference audio signal as input. The process first performs real-time frequency response estimation. Specifically, it traverses each frequency point in the frequency domain and determines whether it is a valid frequency point according to a preset energy threshold. The determination criterion is whether the energy of the complex frequency spectrum of the reference audio signal at that frequency point is higher than the energy threshold, so as to ensure that subsequent calculations are only performed in the valid frequency band with sufficient signal energy.For all valid frequency points, a point-by-point complex division operation is performed between the complex frequency spectrum of the feedback microphone signal and the complex frequency spectrum of the reference audio signal to obtain a real-time frequency response curve characterizing the frequency response characteristics of the complete physical path from the reference signal input to the acoustic pickup of the feedback microphone. This real-time frequency response curve is compared with a preset target frequency response curve, which is retrieved from a pre-stored database based on the current system's operating mode instruction code and represents the desired frequency response shape in the corresponding operating mode. The comparison operation is performed separately in multiple preset key evaluation frequency bands. Within each evaluation frequency band, The amplitude differences between the real-time frequency response curves and the target frequency response curve at all effective frequency points covered by the frequency band are calculated, and the arithmetic mean of these amplitude differences is obtained to obtain the average amplitude deviation characterizing the deviation degree of the evaluation frequency band. The average amplitude deviations calculated for all evaluation frequency bands are arranged and combined in ascending order of frequency band to form a sound quality deviation vector. At the same time, the target frequency response curve data obtained in this process is also passed as output to the subsequent signal processing stage as the benchmark target for frequency response compensation by the dynamic equalizer. The target mode intent parsing process receives the working mode instruction code obtained in step S1 as input. Based on its internally maintained preset mapping rule table, the process parses the input instruction code into a four-dimensional multi-objective optimization weight vector through a query operation. The four dimensions of this multi-objective optimization weight vector correspond to noise suppression weight, voice communication weight, ambient sound transparency weight, and timbre consistency weight, respectively. Specifically, the noise suppression weight is assigned a higher value when set to noise reduction mode and a lower value when set to transparency mode; the voice communication weight is assigned a higher value only when set to call mode and a lower value in both noise reduction and transparency modes; and the ambient sound transparency weight is assigned a higher value when set to transparency mode and a lower value in noise reduction mode. The timbre consistency weight is assigned a relatively high and stable value across all supported operating modes. This is to ensure that the perceived timbre base remains consistent when switching between different operating modes, thus maintaining the continuity of the overall listening experience. Ultimately, this parsed multi-objective optimization weight vector serves as a key prior control parameter, fed into the subsequent intelligent decision-making process. It guides the entire audio processing system in real-time dynamic trade-offs and global optimization among multiple competing optimization objectives such as noise suppression, voice communication, ambient sound transparency, and timbre consistency.

[0024] The present invention is further configured such that S3 includes: The current working mode is determined based on the multi-objective optimization weight vector, and the intention frequency band range and the auditory discomfort frequency band range corresponding to the current working mode are determined according to the preset acoustic knowledge base. Based on the wind noise spectral energy distribution vector in the wind noise scene description vector, the feedback microphone signal and the reference audio signal are respectively multiplied by the wind noise energy value of each frequency point within the intended frequency band by the preset corresponding frequency point importance weight, and all product results are summed to obtain the first conflict energy value; the wind noise energy value of each frequency point within the hearing discomfort frequency band is multiplied by the preset corresponding frequency point sensitivity weight, and all product results are summed to obtain the second conflict energy value; Logarithmic operations are performed on the first and second conflict energy values, and the results are weighted by a first and a second preset coefficient related to the current working mode. The two weighted results are then added to the wind noise impact index in the wind noise scene description vector multiplied by a third preset coefficient to generate the mode fidelity and comfort conflict factor. Specifically, the current dominant working mode is determined by analyzing the input multi-objective optimization weight vector. The process involves comparing the values ​​of the four components within this multi-objective optimization weight vector: noise suppression weight, voice communication weight, ambient sound transparency weight, and timbre consistency weight. The system prioritizes specific audio processing objectives based on the weighted component with the highest value, identifying it as the core operating mode that needs to be prioritized and implemented. For example, if the ambient sound transparency weight is higher than all other weights, the current operating state is determined to be a transparency mode with ambient sound transmission as the core objective. If the voice communication weight is higher than all other weights, the current operating state is determined to be a call mode with optimized voice pickup and communication quality as the core objective. If the noise suppression weight is higher than all other weights, the current operating state is determined to be a noise reduction mode with broadband ambient noise suppression as the core objective. The system identifies the current dominant operating mode and accesses a pre-defined acoustic knowledge base. This knowledge base stores two types of key acoustic parameter definitions associated with each operating mode in a structured format: the intended frequency band and the auditory discomfort frequency band. The intended frequency band is the critical frequency region essential for achieving the core acoustic functions of the corresponding operating mode and ensuring its acoustic fidelity. The auditory discomfort frequency band, on the other hand, is the sensitive frequency region where, under this operating mode, excessive interference sound energy at certain frequencies can easily cause auditory discomfort, auditory fatigue, or adverse physiological reactions in the user. Based on the identified current operating mode, the system retrieves data from this acoustic knowledge base. And extract the uniquely corresponding intention frequency band range data set and auditory discomfort frequency band range data set; for example, corresponding to the transparency mode with environmental sound transmission as the core, its intention frequency band range usually covers the mid-frequency region where the human ear is most sensitive to speech clarity and environmental sound naturalness, while its auditory discomfort frequency band range usually includes the low-frequency region that is prone to low-frequency booming due to wind noise and the high-frequency region that is prone to producing harsh sounds or electroacoustic feedback howling; based on the wind noise spectral energy distribution vector in the wind noise scene description vector, perform two parallel conflict energy calculations, which describe the relative energy of wind noise in each frequency sub-band in a normalized form;The first step is the calculation of conflict energy in the intended frequency band. This involves extracting the wind noise energy value from the wind noise spectrum energy distribution vector for each frequency point within the intended frequency band extracted from the acoustic knowledge base. This wind noise energy value is then multiplied by a preset importance weighting coefficient reflecting the frequency point's importance in achieving the current working mode objective, yielding a weighted conflict contribution value for that frequency point. Subsequently, these weighted conflict contribution values ​​for all frequency points within the intended frequency band are summed to obtain the first conflict energy value. The second step is the calculation of conflict energy in the unsuitable frequency band. This involves extracting the auditory unsuitable frequency band from the acoustic knowledge base... Within the range of auditory discomfort, for each frequency point within the frequency band, the wind noise energy value of the corresponding frequency point is obtained from the wind noise spectrum energy distribution vector. This wind noise energy value is multiplied by a preset sensitivity weighting coefficient that reflects the sensitivity of the frequency point to causing auditory discomfort, resulting in another weighted conflict contribution value for that frequency point. Subsequently, these weighted conflict contribution values ​​of all frequency points within the auditory discomfort frequency band are summed to obtain the second conflict energy value. The synthesis calculation of mode fidelity and comfort conflict factors begins with applying a base-10 logarithmic operation to the first and second conflict energy values, respectively. The core purpose of this logarithmic operation is to compress the dynamics of the data. The range is expanded and the stability of subsequent numerical processing is enhanced. Two preset weighting coefficients, namely the first preset coefficient and the second preset coefficient, are invoked according to the current working mode. The magnitudes of these two coefficients are preset to reflect the differentiated emphasis of different working modes on intended frequency band conflicts and auditory discomfort frequency band conflicts. For example, in the transparent mode that pursues high-fidelity ambient sound reproduction, the first preset coefficient assigned to intended frequency band conflicts has a relatively higher value, while in the noise reduction mode that prioritizes auditory comfort, the second preset coefficient assigned to uncomfortable frequency band conflicts is configured to be relatively higher. The first preset coefficient is used to perform a scalar multiplication on the first conflict energy value after logarithmic processing to achieve... The weighting is then applied by multiplying the logarithmically processed second conflict energy value by a second preset coefficient. The results of the two weighting processes are then algebraically summed to obtain an intermediate sum. Simultaneously, a wind noise impact index, representing the temporal impact characteristics of wind noise, is extracted from the wind noise scene description vector. This wind noise impact index is multiplied by a third preset coefficient to obtain an impact correction term used to quantify the impact of sudden wind noise on the overall conflict. The intermediate sum obtained above is algebraically summed with this impact correction term to generate a single scalar value, which is defined as the mode fidelity and comfort conflict factor.The absolute value of the fidelity-compatibility conflict factor in this mode comprehensively quantifies the intensity of the potential contradiction between fully achieving the acoustic design goals of the original operating mode and maintaining user auditory comfort under current wind noise interference conditions. A higher value indicates a more acute conflict, thus providing a clear signal for subsequent intelligent decision-making steps to drive the co-controller to adopt more adaptive or compromise-oriented strategies among multiple competing audio processing goals.

[0025] The present invention is further configured such that S4 includes: The wind noise scene description vector, sound quality deviation vector, multi-objective optimization weight vector, and mode fidelity and comfort conflict factor are concatenated to generate an enhanced system state vector. The enhanced system state vector is input to a pre-trained second neural network model, which performs forward propagation calculations to output a cooperative control vector. This cooperative control vector includes a noise reduction control sub-vector, a beamforming control sub-vector, an equalization control sub-vector, and a hybrid control sub-vector. Further, the noise reduction control sub-vector configures the type of the noise reduction filter and its noise reduction depth in each frequency band. The beamforming control sub-vector configures the beam pointing and null depth when beamforming the feedforward microphone signal. The equalization control sub-vector configures the gain compensation values ​​of the dynamic equalizer in multiple adjustable frequency bands. The hybrid control sub-vector configures the mixing ratio of ambient sound pass-through. Example and sidetone gain; specifically, the construction of the enhanced system state vector is performed, which receives the wind noise scene description vector, sound quality deviation vector, and multi-objective optimization weight vector from step S2, and the mode fidelity and comfort conflict factor from step S3; according to a preset and unchangeable arrangement order, the above four data entities are sequentially concatenated and combined, specifically in the following order: first, the wind noise presence probability scalar, wind noise intensity level, wind noise spectrum energy distribution vector, and wind noise impact index contained in the wind noise scene description vector are arranged in order; then, all dimensional elements of the sound quality deviation vector are sequentially concatenated thereafter; finally, the noise suppression weight, voice communication weight, ambient sound transparency weight, and sound quality deviation weight in the multi-objective optimization weight vector are concatenated and combined. Color consistency weights are arranged sequentially; mode fidelity and comfort conflict factors are appended to the end; through this operation, a full-dimensional enhanced system state vector is generated, integrating acoustic scene features, system performance feedback, user intent, and potential conflict levels; after the enhanced system state vector is constructed, it is immediately input into a pre-trained offline second neural network model. This second neural network model is architecturally a feedforward multilayer perceptron, and the number of neurons in its input layer is strictly equal to the total dimension of the enhanced system state vector to ensure that information in each dimension is received independently; this model contains at least two hidden layers, each hidden layer consisting of multiple neurons, and each neuron integrates a type of rectified linear unit. Nonlinear activation functions enable the network to perform hierarchical nonlinear transformations and high-order feature abstractions on the high-dimensional input state information, thereby learning a deep mapping relationship between state information and optimal control parameters. The output layer of this model is specially designed so that the total number of its neurons is exactly equal to the sum of the parameters required by all downstream audio processing modules to be controlled. When the enhanced system state vector is input into this model, the data undergoes linear transformation of the weight matrices of each layer and nonlinear processing of the activation function, i.e., forward propagation calculation, and finally generates a numerical sequence of a specific length in the output layer. This sequence is the cooperative control vector. The cooperative control vector is a structured one-dimensional numerical array, which is logically divided into four functional sub-vector segments according to its control purpose.The first sub-vector segment is the noise reduction control sub-vector, whose internal parameters are used to configure the structure type of the adaptive active noise cancellation filter and set differentiated noise reduction depth target values ​​for different frequency bands to achieve adjustable suppression of broadband or specific frequency band noise. The second sub-vector segment is the beamforming control sub-vector, whose internal parameters are used to guide the adaptive beamforming algorithm for the feedforward microphone signal, including setting the pointing angle of the main beam and the null depth for the interference direction, to enhance the speech in the target direction and suppress background noise in voice communication. The third sub-vector segment is the equalization control sub-vector, whose values ​​correspond to multiple independently controllable parameters in the dynamic equalizer. The gain compensation amount required for adjusting the frequency band is used to adjust the overall frequency response characteristics of the system in real time; the fourth sub-vector segment is the mixing control sub-vector, whose parameters are used to manage the audio mixing ratio, including controlling the mixing ratio of ambient sound in transparency mode and the gain of sidetone signal in call mode, to balance the needs of auditory privacy and environmental perception; the ability of this second neural network model to generate optimal collaborative control parameters comes from its offline supervised training phase, which uses a large amount of sample data covering various acoustic scenarios, system states and user operations. This data is constructed through high-fidelity acoustic simulation and field acquisition; each training sample contains an enhancement The system state vector serves as input, and a cooperative control vector, calculated by an expert system or multi-objective optimization algorithm under that state to optimize overall system performance, serves as the desired output label. Training adjusts the network's internal connection weights through backpropagation to minimize a carefully designed comprehensive objective function. This objective function comprehensively considers multiple key performance indicators, including the total noise suppression, speech intelligibility evaluated by an objective algorithm, and the overall consistency error between the system's actual output frequency response and the target frequency response curve. A mode fidelity and comfort conflict factor is integrated into this objective function as a dynamic weight: when the mode fidelity and comfort conflict factor is high, the objective function is more tolerant of timbre consistency errors caused by pursuing comfort; when the mode fidelity and comfort conflict factor is low, the objective function more strictly penalizes deviations from the target frequency response curve. Through large-scale training guided by such conflict-aware objective functions, the second neural network model ultimately learns the ability to generate a set of cooperative control parameters that achieves a dynamic optimal balance among multiple competing performance objectives such as noise suppression, speech quality, timbre consistency, and ambient sound transparency, based on the real-time, full-dimensional system state.

[0026] The present invention is further configured such that S5 includes: Based on the noise reduction control sub-vector in the cooperative control vector, adaptive noise reduction processing is performed on the feedforward microphone signal and the feedback microphone signal. Based on the current working mode, the feedforward microphone signal is subjected to speech enhancement processing or ambient sound mixing processing based on the beamforming control sub-vector or hybrid control sub-vector in the cooperative control vector. Based on the equalization control sub-vector in the cooperative control vector, dynamic equalization compensation is performed on the signal after adaptive noise reduction processing, speech enhancement processing, or ambient sound mixing processing. The signal after adaptive noise reduction, speech enhancement, or ambient sound mixing and dynamic equalization compensation is synthesized to generate the final output signal. Specifically, adaptive noise reduction is first performed. This process receives the feedforward microphone signal and feedback microphone signal from step S1, and simultaneously receives the noise reduction control sub-vector from the cooperative control vector. Based on the filter structure type identifier contained in the noise reduction control sub-vector, the processing logic selects a specific structure from a set of pre-stored filter implementation schemes for instantiation, such as a minimum mean square adaptive filter structure based on the filter reference signal, or a specific form of feedback infinite impulse response filter structure. After selecting the filter structure, the processing logic... The logic, based on the noise reduction depth target values ​​set for different frequency sub-bands in the noise reduction control sub-vector, drives the adaptive algorithm corresponding to the selected filter structure. It iteratively updates the required coefficients within the selected filter and completes the coefficient refresh. The filters with configured parameters act on the input feedforward and feedback microphone signals respectively, generating corresponding anti-noise signals. These anti-noise signals are designed to maintain opposite phase and amplitude correlation with the original noise signals picked up by their respective paths, achieving noise cancellation in acoustic superposition. Next, according to the currently determined operating mode, corresponding speech enhancement processing or ambient sound mixing processing is performed. This processing receives the feedforward microphone signal and coordinates... The control vector includes beamforming control sub-vectors and hybrid control sub-vectors. When a call is detected, voice enhancement processing is activated. The processing logic, based on the desired pointing angle parameter of the main beam and the null depth parameter for interference directions defined in the beamforming control sub-vector, dynamically calculates the complex weighting coefficients applied to the signals of each physical unit of the feedforward microphone array by solving a spatially constrained optimization problem (such as the MVDR algorithm). Through this spatial weighting, a beam pattern with a specific directionality is formed. The main lobe direction of this beam is designed to align with the typical position of the user's mouth to enhance voice pickup sensitivity, while simultaneously forming the deepest possible gain suppression in the direction of existing interference sources. The system separates and enhances the voice signal from interference noise in the spatial domain. When the system is determined to be in transparency mode, ambient sound mixing is enabled. The processing logic performs necessary gain adjustment and phase compensation calculations on the feedforward microphone signal based on the preset ambient sound mixing ratio parameters in the mixing control sub-vector. Then, the processed ambient sound signal is linearly mixed into the main audio signal path according to the set ratio to achieve natural transmission of external ambient sound. In addition, the side tone gain parameter included in the mixing control sub-vector is used in call mode to adjust the signal amplitude of the user's own voice signal after it is fed back through the internal path and mixed into the earpiece path, aiming to maintain the naturalness of the auditory experience during the call.Next, dynamic adaptive equalization compensation processing is performed. This processing receives the intermediate audio signal after the aforementioned noise reduction processing and speech enhancement or ambient sound mixing processing, and receives two key control inputs: one is the target frequency response curve data from step S2, which strictly corresponds to the current operating mode; the other is the equalization control sub-vector from the cooperative control vector. This processing contains two logically connected compensation stages. The first stage is basic equalization compensation, which applies a fixed, global frequency response correction to the input intermediate audio signal based on the desired system frequency response characteristics described by the target frequency response curve data, aiming to initially adjust the overall frequency response of the system to a reference shape close to the target curve. The second stage is dynamic fine-tuning equalization compensation, which performs real-time, fine-tuning of the signal's frequency response based on the gain compensation values ​​set for multiple pre-divided, independently adjustable frequency bands in the equalization control sub-vector. This dynamic fine-tuning compensation is specifically used to offset and correct the additional frequency response deviation from the basic target frequency response curve introduced by the strategic adjustments made by the pre-processing to cope with complex acoustic scenarios. This step ensures the tonal characteristics of the final output sound of the headphones in different environments. The core execution mechanism maintains high consistency and meets user expectations in listening to interference and operating mode switching. Finally, it performs the synthesis and output of all signals. The synthesis processing logic is responsible for converging and mixing all processed signal components. At the synthesis point, based on a predefined signal flow graph, the following three signals are linearly superimposed: the first is the main audio signal after dynamic adaptive equalization compensation; the second is the signal selected according to the current mode after speech enhancement or ambient sound mixing; the third is the anti-noise signal generated by adaptive noise reduction. After linearly superimposing these three signals in the digital domain, a composite, full-bandwidth time-domain digital audio sequence is generated. This time-domain digital audio sequence is the final output signal, which is then sent to a digital-to-analog converter to be converted into a continuous analog electrical signal. This analog electrical signal is then amplified by a power amplifier to obtain sufficient power to drive the speaker, and finally delivered to the speaker unit. The speaker converts this into sound waves that radiate into the user's ear canal, thus completing a complete closed-loop audio processing flow from multi-source signal synchronous acquisition, environmental state perception, intelligent collaborative decision-making to final acoustic signal output.

[0027] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for AI noise reduction and voice enhancement in multi-mode wireless headphones, characterized in that, include: S1: Synchronously acquire feedforward microphone signals, feedback microphone signals, and reference audio signals, and simultaneously obtain working mode commands; Preprocessing and feature extraction are performed on the feedforward microphone signal, feedback microphone signal, and reference audio signal to obtain acoustic feature vectors; S2: Analyze wind noise characteristics based on acoustic feature vectors and generate wind noise scene description vectors; Output frequency response and distortion are analyzed based on feedback microphone signals and reference audio signals to generate a sound quality deviation vector; target mode intent is parsed based on working mode commands to generate a multi-objective optimization weight vector. S3: Calculate the conflict factor between mode fidelity and comfort based on the wind noise scene description vector and the multi-objective optimization weight vector; S4: Concatenate the wind noise scene description vector, sound quality deviation vector, multi-objective optimization weight vector, and mode fidelity and comfort conflict factor to generate an enhanced system state vector; input the enhanced system state vector into the pre-trained second neural network model to output a cooperative control vector; S5: Based on the cooperative control vector, adaptive noise reduction, speech enhancement or ambient sound mixing and dynamic equalization compensation are performed on the feedforward microphone signal and the feedback microphone signal, and the processed signals are synthesized into the final output signal.

2. The AI ​​noise reduction and voice enhancement method for multi-mode wireless headphones according to claim 1, characterized in that, S1 includes: At a preset sampling rate and resolution, the system simultaneously acquires feedforward microphone signals, feedback microphone signals, and reference audio signals, and simultaneously obtains working mode commands. High-pass filtering, frame-segmentation and windowing, and time-frequency transformation are performed on the feedforward microphone signal, feedback microphone signal, and reference audio signal respectively to convert each signal into its corresponding frequency domain complex spectrum. Subband energy, spectral centroid, spectral flatness, and zero-crossing rate are extracted from the complex frequency domain spectra of the feedforward microphone signal, feedback microphone signal, and reference audio signal, respectively, as underlying acoustic description features; based on the complex frequency domain spectra of the feedforward microphone signal and feedback microphone signal, the complex spectral coherence coefficients of the two signals in the preset frequency band are calculated. The underlying acoustic descriptive features are combined with the complex spectral coherence coefficients to form an acoustic feature vector.

3. The AI ​​noise reduction and voice enhancement method for multi-mode wireless headphones according to claim 1, characterized in that, S2 includes: Based on acoustic feature vectors, a pre-trained first neural network model is used to identify and quantify wind noise characteristics, generating a wind noise scene description vector; the wind noise scene description vector includes a wind noise existence probability scalar, wind noise intensity level, wind noise spectrum energy distribution vector, and wind noise impact index. At each effective frequency point, the complex spectrum of the feedback microphone signal in the frequency domain is divided by the complex spectrum of the reference audio signal in the frequency domain to obtain the real-time frequency response curve characterizing the frequency response characteristics of the current system; the average amplitude deviation between the real-time frequency response curve and the target frequency response curve preset for the current working mode in the preset evaluation frequency band is calculated, and a sound quality deviation vector is generated based on each average amplitude deviation. Based on the working mode command, a multi-objective optimization weight vector is obtained by parsing according to the preset mapping rules; the multi-objective optimization weight vector includes noise suppression weight, voice communication weight, ambient sound transparency weight, and timbre consistency weight.

4. The AI ​​noise reduction and voice enhancement method for multi-mode wireless headphones according to claim 3, characterized in that, The effective frequency point is determined by judging whether the energy of the complex spectrum of the frequency domain of the reference audio signal is higher than a preset energy threshold.

5. The AI ​​noise reduction and voice enhancement method for multi-mode wireless headphones according to claim 1, characterized in that, S3 includes: The current working mode is determined based on the multi-objective optimization weight vector, and the intention frequency band range and the auditory discomfort frequency band range corresponding to the current working mode are determined according to the preset acoustic knowledge base. Based on the wind noise spectral energy distribution vector in the wind noise scene description vector, the feedback microphone signal and the reference audio signal are respectively multiplied by the wind noise energy value of each frequency point within the intended frequency band by the preset corresponding frequency point importance weight, and all product results are summed to obtain the first conflict energy value; the wind noise energy value of each frequency point within the hearing discomfort frequency band is multiplied by the preset corresponding frequency point sensitivity weight, and all product results are summed to obtain the second conflict energy value; Logarithmic operations are performed on the first and second conflict energy values ​​respectively, and the results of the two logarithmic operations are weighted by the first and second preset coefficients related to the current working mode. The two weighted results are then added to the value obtained by multiplying the wind noise impact index in the wind noise scene description vector by the third preset coefficient to generate the mode fidelity and comfort conflict factor.

6. The AI ​​noise reduction and voice enhancement method for multi-mode wireless headphones according to claim 1, characterized in that, S4 includes: The wind noise scene description vector, sound quality deviation vector, multi-objective optimization weight vector, and mode fidelity and comfort conflict factor are concatenated to generate an enhanced system state vector. The enhanced system state vector is input into a pre-trained second neural network model, which performs forward propagation calculations to output a cooperative control vector; wherein, the cooperative control vector includes a noise reduction control sub-vector, a beamforming control sub-vector, an equalization control sub-vector, and a hybrid control sub-vector.

7. The AI ​​noise reduction and voice enhancement method for multi-mode wireless headphones according to claim 6, characterized in that, The noise reduction control subvector is used to configure the type of noise reduction filter and its noise reduction depth in each frequency band; the beamforming control subvector is used to configure the beam pointing and null depth when beamforming the feedforward microphone signal; the equalization control subvector is used to configure the gain compensation value of the dynamic equalizer in multiple adjustable frequency bands; and the mixing control subvector is used to configure the mixing ratio and sidetone gain of the ambient sound pass-through.

8. The AI ​​noise reduction and voice enhancement method for multi-mode wireless headphones according to claim 1, characterized in that, S5 includes: Based on the noise reduction control sub-vector in the cooperative control vector, adaptive noise reduction processing is performed on the feedforward microphone signal and the feedback microphone signal. Based on the current working mode, the feedforward microphone signal is subjected to speech enhancement processing or ambient sound mixing processing based on the beamforming control sub-vector or hybrid control sub-vector in the cooperative control vector. Based on the equalization control sub-vector in the cooperative control vector, dynamic equalization compensation is performed on the signal after adaptive noise reduction processing, speech enhancement processing, or ambient sound mixing processing. The signal after adaptive noise reduction, speech enhancement or ambient sound mixing and dynamic equalization compensation is synthesized to generate the final output signal.

9. The AI ​​noise reduction and voice enhancement method for multi-mode wireless headphones according to claim 2, characterized in that, The sampling rate is not less than 16 kHz.

10. The AI ​​noise reduction and voice enhancement method for multi-mode wireless headphones according to claim 2, characterized in that, The resolution is no less than 16 bits.