An earphone control method based on voiceprint recognition and reverse wave cancellation wind noise reduction

By employing a headphone control method based on voiceprint recognition and reverse wave cancellation, the problem of poor noise cancellation performance in traditional headphones in noisy environments is solved. This achieves personalized adaptive noise cancellation and efficient voiceprint recognition, thereby improving the user experience of headphones in noisy environments.

CN119052696BActive Publication Date: 2025-11-04MINAMI ACOUSTICS LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411164094.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-23
Publication Date
2025-11-04
Estimated Expiration
2044-08-23

AI Technical Summary

Technical Problem

Traditional headphones have poor noise cancellation in noisy environments, especially for high-frequency noise and wind noise. Furthermore, existing voiceprint recognition technology is easily affected by environmental noise, has low recognition accuracy, excessive power consumption, and short battery life.

Method used

A headphone control method based on voiceprint recognition and inverse wave cancellation is adopted. Data enhancement is performed by acquiring user voiceprint data, speech and noise are separated by long short-term memory network, noise source localization is performed by combining microphone array technology, and real-time noise tracking and inverse wave adjustment are performed by recursive least mean square algorithm. Power consumption is dynamically adjusted to optimize inverse wave data, and finally clear voiceprint audio is generated.

Benefits of technology

It improves the noise cancellation effect and voiceprint recognition accuracy of headphones in noisy environments, reduces power consumption, extends battery life, and achieves personalized voiceprint processing and adaptive noise cancellation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119052696B_ABST
    Figure CN119052696B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of earphone control, in particular to an earphone control method based on voiceprint recognition and reverse wave wind noise reduction. The method comprises the following steps: obtaining user voiceprint data; performing data enhancement on the user voiceprint data to obtain voiceprint enhanced data; performing voice and noise separation on the voiceprint enhanced data by using a long short-term memory network, so that voice signal data is obtained; performing voice spectrum feature and phase feature fusion on the voice signal data, so that user voiceprint feature data is generated; generating a user voiceprint model according to the user voiceprint feature data; positioning a noise source of the user voiceprint data by using a microphone array technology, and determining optimal waveforms and phases of reverse waves according to the noise source and earphone loudspeaker layout, so that optimal reverse wave data is obtained. The application can more effectively reduce complex noises such as wind noise by using the reverse wave technology and the adaptive algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of earphone control, and particularly relates to an earphone control method based on voiceprint recognition and reverse wave cancellation for wind noise reduction. BACKGROUND

[0002] With the rapid development of wireless communication technology and audio processing technology, earphones have become one of the indispensable electronic devices in people's daily life. However, the use experience of traditional earphones in noisy environments is often unsatisfactory, especially under the interference of wind noise and other environmental noise, users are difficult to obtain clear audio effect. In addition, with the increasing demand for personalization, users hope that earphones can be adaptively adjusted according to their own voiceprint characteristics, so as to provide better sound quality and noise reduction effect.

[0003] Traditional active noise reduction earphones mainly rely on microphones to capture environmental noise and generate reverse waveforms with opposite phases to cancel noise. This method has good noise reduction effect on low-frequency noise, but has limited effect on high-frequency noise and wind noise in complex environments, and it is difficult to realize real-time noise change tracking. At the same time, existing voiceprint recognition technology is also often applied to wind noise reduction in earphones.

[0004] However, the traditional noise reduction method using voiceprint recognition and reverse wave cancellation often has the following problems: high-performance noise reduction and voice recognition often lead to high power consumption of earphones and short battery life; conventional voice control is easily disturbed by environmental noise and has low recognition accuracy. SUMMARY

[0005] Therefore, it is necessary to provide an earphone control method based on voiceprint recognition and reverse wave cancellation for wind noise reduction to solve at least one of the above technical problems.

[0006] To achieve the above purpose, an earphone control method based on voiceprint recognition and reverse wave cancellation for wind noise reduction comprises the following steps:

[0007] Step S1: obtaining user voiceprint data; performing data enhancement on the user voiceprint data to obtain voiceprint enhanced data; using a long short-term memory network to separate the voiceprint enhanced data into speech and noise, thereby obtaining speech signal data;

[0008] Step S2: performing speech spectrum feature and phase feature fusion on the speech signal data to generate user voiceprint feature data; establishing a user voiceprint model according to the user voiceprint feature data;

[0009] Step S3: positioning the noise source of the user voiceprint data through microphone array technology, and determining the optimal waveform and phase of the reverse wave according to the noise source and the layout of the earphone loudspeaker, thereby obtaining the optimal reverse wave data;

[0010] Step S4: Real-time tracking of noise changes using a recursive least squares algorithm based on user voiceprint data, and adjusting the optimal data of the reverse wave to obtain adaptive reverse wave data;

[0011] Step S5: Dynamic adjustment of the computing power and power consumption of the adaptive reverse wave data based on a digital signal processor to obtain power consumption control parameter data; noise cancellation and voiceprint extraction of user voiceprint data based on power consumption control parameter data, adaptive reverse wave data and user voiceprint model to obtain voiceprint clear audio data.

[0012] The application of data enhancement technology to user voiceprint data can increase the diversity of data samples and improve the robustness and generalization ability of the voiceprint model. This helps the voiceprint model to have good recognition performance for voiceprint data under different speech environments and noise conditions. Through technologies such as long short-term memory network, voiceprint enhancement data can be separated from speech and noise, which can remove the interference of background noise on voiceprint features, which helps to improve the accuracy and stability of the voiceprint recognition system. By performing spectral analysis on the speech signal data and fusing the speech spectrum features and phase features, the information of both can be utilized comprehensively to improve the representation ability of voiceprint features, which helps to enhance the understanding and differentiation ability of voiceprint model for sound features, and improves the accuracy and robustness of voiceprint recognition. By training a voiceprint model based on user voiceprint feature data, a personalized voiceprint model for the user can be established, which helps to identify and distinguish the user's voiceprint in the future, and realizes personalized voiceprint processing and service. Through the microphone array technology, the position of the noise source in the user voiceprint data can be determined, which helps to accurately calculate the optimal waveform and phase of the reverse wave to minimize the interference of noise on the voiceprint data. According to the noise source and the layout of the earphone speaker, the optimal waveform and phase of the reverse wave are determined to realize the cancellation and reduction of noise, which helps to improve the clarity and distinguishability of the voiceprint signal. Through technologies such as recursive least squares algorithm, the changes of noise in user voiceprint data are tracked in real time, which helps to dynamically adjust the waveform of the reverse wave to adapt to different step changes in the noise environment, further reducing the influence of noise on voiceprint data. By adjusting the waveform of the optimal reverse wave data, the noise can be counteracted according to the real-time noise change, which helps to improve the noise cancellation effect and make the voiceprint data clearer and more distinguishable. Through the dynamic adjustment of the computing power and power consumption of the adaptive reverse wave data by the digital signal processor, the calculation performance and power consumption can be balanced according to the actual demand, which helps to improve the efficiency of the system and save energy. According to the power consumption control parameter data, the adaptive reverse wave data and the user voiceprint model, the noise cancellation and voiceprint extraction of the user voiceprint data are carried out, which helps to eliminate the interference of noise on the voiceprint and extract clear voiceprint audio data, thereby improving the accuracy and reliability of voiceprint recognition.

[0013] Preferably, step S1 comprises the following steps:

[0014] Step S11: Collecting user voice through the built-in microphone of the earphone to obtain original voiceprint data;

[0015] Step S12: Downsampling and filtering the original voiceprint data, and performing mute detection and endpoint detection to extract valid voice segments, thereby obtaining user voiceprint data;

[0016] Step S13: Data enhancement is performed on the user voiceprint data to obtain voiceprint enhanced data;

[0017] Step S14: Constructing a long short-term memory network model according to the voiceprint enhanced data, wherein the long short-term memory network model comprises an input layer for receiving the voiceprint enhanced data, a multi-layer LSTM layer for capturing time sequence dependence, a full connection layer for feature fusion, and an output layer for generating a time-frequency mask of voice and noise;

[0018] Step S15: Processing input features according to the long short-term memory network model to obtain voice-noise separation mask data;

[0019] Step S16: Reconstructing voice from the user voiceprint data according to the voice-noise separation mask data to obtain voice signal data.

[0020] The application can conveniently collect user's voice data through the built-in microphone of the earphone, provide input for subsequent voiceprint processing, help establish personal voiceprint model, and perform voiceprint recognition and verification. By reducing the sampling rate and applying filter, the number of sampling points of original voiceprint data can be reduced, the calculation and storage overhead can be reduced, and the key frequency components can be retained, which helps improve the efficiency and processing speed of the system. Through the mute detection and endpoint detection algorithm, the silent segment and non-speech segment can be identified and removed, and only the effective speech segment is retained, which helps reduce the interference of noise and non-speech components on voiceprint feature extraction, and improve the accuracy and reliability of voiceprint recognition. By applying data enhancement techniques such as sound reverberation, speed change, volume adjustment, etc. to user voiceprint data, the diversity of data samples can be increased, and the robustness and generalization ability of the voiceprint model can be improved, which helps the voiceprint model to have good recognition performance under different voice environments and noise conditions. By constructing an LSTM model, the time sequence dependence relationship in the voiceprint enhanced data can be captured, and the feature representation of speech and noise can be extracted. This helps improve the understanding and differentiation ability of the voiceprint recognition system for voiceprint data. Through the full connection layer for feature fusion, the features extracted by the LSTM model are integrated to generate the time-frequency mask of speech and noise, which helps separate the speech and noise components and improve the clarity and noise suppression effect of the speech signal. By applying the long short-term memory network model to process the input features, a speech-noise separation mask data can be generated to indicate the relative energy distribution of speech and noise, which helps accurately separate the speech and noise components and improve the clarity and audibility of the sound. By applying the speech-noise separation mask to the user voiceprint data, the original speech signal can be restored and the noise component can be removed, which helps improve the quality and clarity of the speech signal and makes it easier to understand and analyze.

[0021] Preferably, step S13 comprises the following steps:

[0022] Step S131: Adding different types and intensities of environmental noise to the user voiceprint data, and changing the speech rate, thereby generating audio data of different speech rate versions;

[0023] Step S132: Perform pitch shift simulation on the audio data to obtain a voiceprint enhanced data set;

[0024] Step S133: Perform short-time Fourier transform on the voiceprint enhanced data set to convert the time domain signal to time-frequency representation, thereby obtaining spectrogram data;

[0025] Step S134: Extract Mel-frequency cepstral coefficient features from the spectrogram data and calculate the pitch frequency contour, thereby obtaining pitch feature data;

[0026] Step S135: performing spectral centroid and spectral flux acoustic feature extraction on the pitch feature data to obtain acoustic feature data;

[0027] Step S136: performing acoustic feature intensity-based screening on the voiceprint enhancement data set according to the acoustic feature data to obtain voiceprint enhancement data.

[0028] The present application simulates different speech environments and noise conditions by adding different types and intensities of environmental noise, making the voiceprint model robust to voiceprint data in different environments. At the same time, by generating audio data of different speech rates by changing the speech rate, the diversity of data samples is increased, and the generalization ability of the voiceprint model is improved. By simulating pitch shift on audio data, the pitch of the sound can be changed, the diversity of data samples is increased, and the recognition performance of the voiceprint model on voiceprint data under different pitch conditions is improved, which helps to make the voiceprint model robust to changes in the pitch of different people's voices. By applying short-time Fourier transform, the time-domain signals in the voiceprint enhancement data set are converted into time-frequency representation, i.e. spectral graph data, which helps to capture the energy distribution and spectral features of the sound at different frequencies, providing a basis for voiceprint feature extraction and analysis. By extracting Mel-frequency cepstral coefficient (MFCC) features from spectral graph data, the spectral envelope and resonance characteristics of the sound can be captured for voiceprint feature representation. In addition, pitch frequency contour calculation can estimate the fundamental frequency information of the sound, providing pitch-related feature data. By calculating the spectral centroid and spectral flux of the pitch feature data, the acoustic features of the sound can be extracted. The spectral centroid reflects the concentration and distribution of the spectrum, and the spectral flux represents the rate of change of the spectrum, which has strong discriminative ability to distinguish different sounds and noise components. By screening based on acoustic feature intensity, voiceprint enhancement data with more obvious and prominent acoustic features can be selected, reducing data samples with more noise interference or less obvious features, which helps to improve the training effect and recognition accuracy of the voiceprint model.

[0029] Preferably, step S16 comprises the following steps:

[0030] Step S161: performing short-time Fourier transform on the user voiceprint data to obtain original spectral graph data;

[0031] Step S162: applying speech-noise separation mask data to the original spectral graph data to obtain clean speech spectral graph data;

[0032] Step S163: performing phase reconstruction on the clean speech spectral graph data using the Griffin-Lim algorithm to obtain speech spectral graph optimization data;

[0033] Step S164: performing inverse short-time Fourier transform on the speech spectral graph optimization data to obtain preliminary reconstructed speech data;

[0034] Step S165: smoothing the preliminary reconstructed speech data using a Wiener filter to obtain smoothed speech signal data;

[0035] Step S166: performing dynamic range compression based on audio loudness and intelligibility on the smoothed speech signal data to obtain dynamically compressed speech signal data.

[0036] The Short-Time Fourier Transform (STFT) decomposes the sound signal into different frequency components and provides information on time and frequency. By applying STFT to the user's voiceprint data, the energy distribution of the sound signal at different time periods and frequencies can be obtained. The speech-noise separation mask is used to separate the noise components from the speech components in the original spectrogram. By applying the speech-noise separation mask data, the pure speech spectrogram data can be extracted, thereby removing the interference of background noise on the speech signal. The Griffin-Lim algorithm is used to estimate the phase information of the speech signal. By combining the pure speech spectrogram data with the estimated phase information, the optimized data of the speech spectrogram is obtained. This step helps to restore the time-domain characteristics of the speech signal, making the reconstructed speech more natural and understandable. The Inverse Short-Time Fourier Transform (ISTFT) converts the optimized speech spectrogram data back to the time-domain signal. By applying the Inverse Short-Time Fourier Transform, the preliminary reconstructed speech data can be recovered from the frequency domain. The Wiener filter is a commonly used signal processing filter that reduces noise interference and enhances the intelligibility of the speech signal. By applying the Wiener filter, the preliminary reconstructed speech data can be smoothed, reducing the impact of noise and improving the quality and intelligibility of the speech. Dynamic range compression is an audio processing technique that adjusts the volume dynamic range of the speech signal, making the audio listening experience more balanced and consistent at different volume levels. By applying dynamic range compression based on audio loudness and intelligibility, the volume of the speech signal can be more stable, improving the audibility and comfort of the speech.

[0037] Preferably, step S2 comprises the following steps:

[0038] Step S21: performing linear prediction coefficient calculation on the speech signal data to obtain LPC feature data;

[0039] Step S22: performing phase information extraction on the speech signal data and performing group delay feature evaluation based on the phase information to obtain group delay feature data;

[0040] Step S23: performing feature concatenation on the LPC feature data and the group delay feature data to obtain preliminary fused feature data;

[0041] Step S24: performing z-score normalization on the preliminary fused feature data to obtain user voiceprint feature data;

[0042] Step S25: modeling the user voiceprint feature data using a Gaussian mixture model to obtain a GMM voiceprint model, and performing feature extraction based on an i-vector framework on the GMM voiceprint model to obtain an i-vector voiceprint representation;

[0043] Step S26: constructing a deep neural network (DNN) voiceprint model based on the user voiceprint feature data, and performing d-vector feature extraction from a bottleneck layer of the DNN voiceprint model to obtain a d-vector voiceprint representation;

[0044] Step S27: integrating the GMM voiceprint model, the i-vector voiceprint representation, and the d-vector voiceprint representation to obtain a comprehensive user voiceprint model.

[0045] The linear prediction coefficient (LPC) in the present application is a method for analyzing speech signals. By calculating the linear prediction coefficient of the speech signal, the formant frequency and bandwidth of the speech signal can be extracted, which has high discrimination ability for voiceprint recognition. The group delay feature is a feature that describes the phase information of the speech signal. By extracting the phase information of the speech signal and evaluating the group delay feature, the phase structure of the speech signal in the frequency domain can be obtained, which is very important for speech periodicity modeling and voiceprint feature extraction in voiceprint recognition. Feature concatenation is a process of combining and fusing different types of feature data. By concatenating LPC feature data and group delay feature data, more rich and diverse voiceprint features can be obtained, improving the performance and robustness of voiceprint recognition. z-score normalization is a common feature standardization method used to eliminate the scale difference between features. By performing z-score normalization on the preliminary fused feature data, the feature data can be mapped to a standard normal distribution with a mean of 0 and a standard deviation of 1, making the weights of different features more balanced and improving the robustness and reliability of the voiceprint model. Gaussian mixture model (GMM) is a common voiceprint modeling method used for voiceprint feature modeling and recognition. By applying GMM modeling to user voiceprint feature data, a probability model can be obtained to describe the distribution of voiceprint features. At the same time, by extracting i-vector features based on the i-vector framework, a more compact and discriminative i-vector voiceprint representation can be obtained. DNN has strong non-linear modeling ability and can learn more complex and abstract voiceprint feature representations, which helps to distinguish the voice features of different people and improve the discrimination performance of voiceprint recognition. DNN voiceprint model can learn robust feature representation to noise and changes through the use of large-scale training data, which makes the voiceprint model have certain anti-interference ability to environmental noise, speech quality change and other factors. DNN voiceprint model can be flexibly adjusted and expanded by increasing the number of network layers and adjusting the number of neurons, which makes the voiceprint model adapt to different voiceprint recognition tasks and application scenarios. The bottleneck layer of the DNN voiceprint model generally has a lower dimension, which is usually much smaller than the dimension of the original speech feature. This dimension reduction can reduce the redundancy of the feature and improve the efficiency and accuracy of subsequent processing. d-vector feature is a high-discriminative voiceprint representation learned by DNN voiceprint model. Compared with the original speech feature, d-vector feature can better distinguish the voice features of different people, thereby improving the accuracy and robustness of voiceprint recognition.

[0046] Preferably, step S3 comprises the following steps:

[0047] Step S31: audio channel division is performed on the user voiceprint data by using a microphone array technology, so as to obtain spatial audio data;

[0048] Step S32: time delay calculation is performed on the spatial audio data based on a generalized cross-correlation algorithm, so as to obtain sound source direction data;

[0049] Step S33: noise source positioning processing is performed on the sound source direction data by using a multiple signal classification algorithm, so as to obtain noise source position data;

[0050] Step S34: an acoustic transfer function model is constructed according to the noise source position data and the geometric layout of the earphone loudspeaker, and optimal reverse wave waveform calculation is performed on the acoustic transfer function model according to a least mean square error criterion, so as to obtain initial reverse waveform data;

[0051] Step S35: phase correction is performed on the initial reverse waveform data, so as to obtain reverse wave optimal data.

[0052] The microphone array of the present application can receive sound signals in different positions and directions, and by dividing the audio channel of the voiceprint data, the source direction and position information of the sound can be determined, which helps subsequent sound source positioning and reverse wave calculation. The microphone array can suppress environmental noise and noise through spatial difference reception, and by selecting appropriate microphone combinations or applying array signal processing algorithms, the noise interference on the voiceprint data can be effectively reduced, and the accuracy and robustness of voiceprint recognition can be improved. By calculating the time delay between different microphones, the direction and position of the sound source can be determined, and the generalized cross-correlation algorithm can effectively estimate the time difference of the sound signal between different microphones, thereby realizing the positioning of the sound source. The generalized cross-correlation algorithm can utilize more abundant sound feature information when calculating the time delay, providing higher positioning accuracy than traditional cross-correlation algorithms, which helps to accurately obtain the sound source direction data and provide accurate input for subsequent noise source positioning and reverse wave calculation. The multiple signal classification algorithm can classify the sound signal according to the sound source direction data, and divide the sound source into target sound source and noise source. By identifying and positioning the noise source, the noise situation in the environment can be better understood, providing important information for subsequent reverse wave calculation. Noise source positioning processing can help effectively separate the target sound source from the background noise, and by separating the target sound source and noise source, the reverse wave form can be calculated more accurately, improving the noise suppression effect and the performance of voiceprint recognition. The acoustic transfer function model describes the propagation process of sound between the earphone speaker and the microphone, and by constructing the acoustic transfer function model and calculating the optimal reverse wave form, the interference generated by the noise source on the earphone speaker can be canceled out, thereby realizing noise suppression and elimination. The optimal reverse wave form calculation is based on the acoustic transfer function model and the minimum mean square error criterion, and by calculating the optimal reverse wave form, a wave form opposite to the noise can be generated to cancel out the influence of the noise on the sound signal, which helps to improve the clarity and audibility of the sound. Phase correction can ensure that the reverse wave form is consistent with the phase of the original sound signal, and phase consistency is very important for sound reconstruction and restoration, which can reduce sound distortion and aberration, and improve sound quality and accuracy. Through phase correction, the original sound signal can be better restored, and phase correction can correct the phase shift of the reverse wave form to make it completely match the original sound signal, thereby realizing more accurate sound restoration effect. BRIEF DESCRIPTION OF DRAWINGS

[0053] Other features, objects, and advantages of the present application will become more apparent from the following detailed description of non-limiting embodiments made with reference to the following drawings:

[0054] Figure 1 The step flowchart of the earphone control method based on voiceprint recognition and reverse wave cancellation wind noise of the present application;

[0055] Figure 2 isFigure 1 Detailed step flow diagram in step S1 is shown in the following figure;

[0056] Figure 3 For Figure 1 Detailed step flow diagram in step S2 is shown in the following figure. DETAILED DESCRIPTION

[0057] The technical method of the present application will be described clearly and completely in combination with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all the other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0058] In addition, the accompanying drawings are only schematic illustrations of the present application, and are not necessarily drawn to scale. The same reference numerals in the drawings represent the same or similar parts, and thus repeated description thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities, which do not necessarily have to correspond to physically or logically independent entities. The functional entities can be implemented in the form of software, or in one or more hardware modules or integrated circuits, or in different network and / or processor methods and / or microcontroller methods.

[0059] It should be understood that although the terms "first", "second" and the like can be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of the exemplary embodiments, a first element can be called a second element, and similarly a second element can be called a first element. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0060] To achieve the above-mentioned purpose, please refer to Figures 1 to 3 The present application provides an earphone control method based on voiceprint recognition and reverse wave cancellation wind noise reduction, which comprises the following steps:

[0061] Step S1: obtaining user voiceprint data; performing data enhancement on the user voiceprint data to obtain voiceprint enhanced data; using a long short-term memory network to separate the voiceprint enhanced data into speech and noise, thereby obtaining speech signal data;

[0062] Step S2: performing speech spectrum feature and phase feature fusion on the speech signal data to generate user voiceprint feature data; generating a user voiceprint model according to the user voiceprint feature data;

[0063] Step S3: The noise source of the user voiceprint data is located through the microphone array technology, and the optimal waveform and phase of the reverse wave are determined according to the noise source and the layout of the earphone speaker, so as to obtain the optimal data of the reverse wave;

[0064] Step S4: The noise change is tracked in real time according to the user voiceprint data by using the recursive least square algorithm, and the waveform of the reverse wave is adjusted according to the optimal data of the reverse wave, so as to obtain the adaptive reverse wave data;

[0065] Step S5: The adaptive reverse wave data is dynamically adjusted in terms of computing power and power consumption based on the digital signal processor, so as to obtain the power consumption control parameter data; the user voiceprint data is noise-canceled and voiceprint extracted according to the power consumption control parameter data, the adaptive reverse wave data and the user voiceprint model, so as to obtain the voiceprint clear audio data.

[0066] In the embodiment of the application, reference Figure 1 The application is a step flow diagram of an earphone control method based on voiceprint recognition and reverse wave cancellation wind noise reduction, and in the example, the earphone control method based on voiceprint recognition and reverse wave cancellation wind noise reduction comprises the following steps:

[0067] Step S1: Obtain user voiceprint data; perform data enhancement on the user voiceprint data to obtain voiceprint enhanced data; and use a long short-term memory network to separate speech and noise from the voiceprint enhanced data, so as to obtain speech signal data;

[0068] The embodiment of the application can record the user's voice sample through a microphone or other audio equipment, and save it as a digital audio file format (such as WAV). For the obtained user voiceprint data, a series of data enhancement techniques such as sound speed change, time domain transformation, frequency domain transformation, etc. can be applied to increase the diversity and richness of the data. The long short-term memory network (LSTM) in deep learning or other neural network models suitable for speech signal processing are used to train and infer the voiceprint enhanced data, so as to separate the speech and noise, and obtain pure speech signal data.

[0069] Step S2: Perform speech spectrum feature and phase feature fusion on the speech signal data, so as to generate user voiceprint feature data; and generate a user voiceprint model according to the user voiceprint feature data;

[0070] Embodiments of the present application use signal processing techniques, such as Short-Time Fourier Transform (STFT), to convert the speech signal into a time-frequency domain representation. Speech spectral features, such as Mel-frequency Cepstral Coefficients (MFCC), are extracted from the time-frequency domain representation. Phase information is extracted from the time-frequency domain representation of the speech signal. Commonly used methods include extracting the phase spectrum or phase-based feature representation. The speech spectral features and phase features are combined or fused to generate user voiceprint feature data. Using the generated voiceprint feature data, a machine learning method (such as a support vector machine, a deep neural network, etc.) or a statistical modeling method (such as a Gaussian mixture model) can be used to establish a user's voiceprint model.

[0071] Step S3: Noise source localization is performed on the user voiceprint data by using a microphone array technique, and the optimal waveform and phase of the counter wave are determined according to the noise source and the layout of the earphone speaker, thereby obtaining optimal counter wave data.

[0072] Embodiments of the present application use an array with multiple microphones, which is deployed in a suitable position to receive user voiceprint data and noise in the environment. By using signal processing and sound source positioning algorithms, the position of the noise source is determined by analyzing the sound received by the microphone array. According to the position of the noise source and the layout of the earphone speaker, the optimal counter wave form and phase are determined using the counter wave form synthesis technique and the phase adjustment method to cancel the noise. The determined counter wave form and phase are applied to the user voiceprint data to obtain the optimal counter wave data after the noise is canceled.

[0073] Step S4: Real-time tracking of noise changes is performed on the user voiceprint data using a recursive least squares algorithm, and the counter wave optimal data is adjusted, thereby obtaining adaptive counter wave data.

[0074] Embodiments of the present application use a recursive least squares algorithm to track the user voiceprint data in real time to estimate the statistical properties of the noise and update the adjustment parameters of the counter wave form in real time. According to the estimation results of the recursive least squares algorithm, the noise changes are tracked in real time, and the counter wave form is adjusted accordingly. According to the real-time tracking of the noise changes, the previously determined counter wave optimal data is adjusted. An adaptive filter or filtering algorithm, such as a recursive least squares (RLS) algorithm, can be used to dynamically adjust the counter wave form to adapt to the changes in the noise. By adjusting the counter wave optimal data in real time, adaptive counter wave data is obtained. These data have been optimized according to the real-time noise changes to maximize the cancellation of noise and provide clear sound signals.

[0075] Step S5: dynamically adjusting the computing power and power consumption of the adaptive backwave data based on the digital signal processor, thereby obtaining power consumption control parameter data; performing noise cancellation and voiceprint extraction on the user voiceprint data according to the power consumption control parameter data, the adaptive backwave data and the user voiceprint model, thereby obtaining voiceprint clear audio data.

[0076] The voice input signal of the user is acquired in the embodiment of the application, which can be collected through a microphone or other audio equipment. The adaptive backwave data is superimposed on the user voice input signal, and the signal superposition or mixing technology can be used to superimpose the adaptive backwave data and the user voice input signal to achieve noise cancellation. Through the user voice input signal superimposed with the adaptive backwave data, the clear user voice output is finally obtained. In this way, the influence of noise can be effectively reduced, and the voice quality and clarity are improved.

[0077] The present application can increase the diversity of data samples, improve the robustness and generalization ability of voiceprint model by applying data enhancement technology to user voiceprint data; this helps the voiceprint model to have good recognition performance under different voice environments and noise conditions. By separating the voiceprint enhanced data through long short-term memory network and other technologies, the interference of background noise on voiceprint features can be removed, which helps to improve the accuracy and stability of the voiceprint recognition system. By performing spectral analysis on the voice signal data and fusing the voice spectrum features and phase features, the information of both can be utilized comprehensively to improve the representation ability of voiceprint features, which helps to enhance the understanding and differentiation ability of voiceprint model to sound features, and improves the accuracy and robustness of voiceprint recognition. By training a voiceprint model according to user voiceprint feature data, a personalized voiceprint model for the user can be established, which helps to identify and distinguish the user's voiceprint in the future, and realizes personalized voiceprint processing and service. Through the microphone array technology, the position of the noise source in the user voiceprint data can be determined, which helps to accurately calculate the optimal waveform and phase of the reverse wave to reduce the interference of noise on voiceprint data to the greatest extent. According to the noise source and earphone speaker layout, the optimal waveform and phase of the reverse wave are determined, which can realize the cancellation and reduction of noise, which helps to improve the clarity and distinguishability of voiceprint signal. Through recursive least square algorithm and other technologies, the change of noise in user voiceprint data is tracked in real time, which helps to dynamically adjust the waveform of the reverse wave to adapt to different step changes in the noise environment, further reducing the influence of noise on voiceprint data. By adjusting the waveform of the optimal reverse wave data, the noise can be counteracted according to the real-time noise change, which helps to improve the noise cancellation effect and make the voiceprint data clearer and more distinguishable. Through the dynamic adjustment of the computing power and power consumption of the adaptive reverse wave data by the digital signal processor, the calculation performance and power consumption can be balanced according to the actual demand, which helps to improve the efficiency of the system and save energy. According to the power consumption control parameter data, the adaptive reverse wave data and the user voiceprint model, the noise cancellation and voiceprint extraction are performed on the user voiceprint data, which helps to eliminate the interference of noise on voiceprint and extract clear voiceprint audio data, thereby improving the accuracy and reliability of voiceprint recognition.

[0078] Preferably, step S1 comprises the following steps:

[0079] Step S11: acquiring user voice through the earphone built-in microphone to obtain original voiceprint data;

[0080] Step S12: downsampling and filtering the original voiceprint data, and performing mute detection and endpoint detection to extract valid voice segments, thereby obtaining user voiceprint data;

[0081] Step S13: data enhancement is performed on the user voiceprint data to obtain voiceprint enhanced data;

[0082] Step S14: Constructing a long short-term memory network model according to the voiceprint enhancement data, wherein the long short-term memory network model comprises an input layer for receiving the voiceprint enhancement data, a multi-layer LSTM layer for capturing time-dependent, a full connection layer for feature fusion, and an output layer for generating a time-frequency mask of speech and noise;

[0083] Step S15: Processing the input features according to the long short-term memory network model to obtain speech-noise separation mask data;

[0084] Step S16: Reconstructing the speech from the user voiceprint data according to the speech-noise separation mask data to obtain speech signal data.

[0085] As an embodiment of the present application, referring to Figure 2 Fig. 1 shows a detailed step flowchart of step S1 in the present application, wherein step S1 in the embodiment of the present application comprises the following steps: Figure 1

[0086] Step S11: Collecting user speech through the built-in microphone of the earphone to obtain original voiceprint data;

[0087] The embodiment of the present application uses the built-in microphone of the earphone to collect speech, which can be realized by accessing the audio input interface of the device, such as using the Web Audio API of the Web browser or the audio collection function of the mobile application. The speech signal collected by the microphone is used to obtain the original voiceprint data; this can be realized by using an audio processing library or software to obtain an audio data stream and storing it as original voiceprint data.

[0088] Step S12: Downsampling and filtering the original voiceprint data, and performing silence detection and endpoint detection to extract valid speech segments to obtain user voiceprint data;

[0089] The embodiment of the present application downsamples the original voiceprint data to reduce the data volume and adapt to the requirements of subsequent processing. Common downsampling methods include average extraction and maximum value extraction. The voiceprint data after downsampling is filtered to remove high-frequency noise and irrelevant signal components; a digital filter such as a low-pass filter or a band-pass filter can be used. Through a silence detection algorithm, the silence segments and non-silence segments in the voiceprint data are identified. The silence segments usually do not contain valid speech information, so they can be excluded from the voiceprint data. Through an endpoint detection algorithm, the starting point and ending point of the voiceprint data are determined to extract valid speech segments. Endpoint detection can be based on energy threshold, short-time energy or other features.

[0090] Step S13: Data enhancement is performed on the user voiceprint data to obtain voiceprint enhancement data;

[0091] ​The embodiment of the application applies data enhancement technology to expand the user voiceprint data set to improve the robustness and generalization ability of the model. Common data enhancement technologies include time shift, frequency shift, speed disturbance, noise addition, etc. The data enhancement technology is applied to the user voiceprint data, for example, time shift in the time domain, frequency shift in the frequency domain, changing the speech speed, etc. These operations can be realized through an audio processing library or software.

[0092] Step S14: constructing a long short-term memory network model according to the voiceprint enhanced data, wherein the long short-term memory network model comprises an input layer for receiving the voiceprint enhanced data, a multi-layer LSTM layer for capturing time sequence dependence, a full connection layer for feature fusion, and an output layer for generating a time-frequency mask of speech and noise;

[0093] In the embodiment of the application, the voiceprint enhanced data is divided into a training set and a validation set, input features and target labels are prepared, and an input layer is designed to receive the processed voiceprint enhanced data. A multi-layer LSTM structure is constructed, each layer captures the time sequence dependence of the speech data, and the number of layers and neurons is adjusted according to actual requirements. The output of the LSTM layer is fused through a full connection layer. An output layer is designed to generate a time-frequency mask of speech and noise for separating speech and noise. The LSTM model is trained using training data, and parameters and hyperparameters are adjusted to improve the performance of the model.

[0094] Step S15: processing the input features according to the long short-term memory network model to obtain speech-noise separation mask data;

[0095] In the embodiment of the application, the enhanced voiceprint data is input into the trained LSTM model, and a time-frequency mask data of speech and noise is generated through LSTM model prediction. The model output is processed to obtain speech-noise separation mask data, which represents the distribution of speech and noise in the time-frequency domain.

[0096] Step S16: reconstructing the speech from the user voiceprint data according to the speech-noise separation mask data to obtain speech signal data.

[0097] In the embodiment of the application, the noise component in the user voiceprint data is suppressed according to the speech-noise separation mask data, and the pure speech component is retained. A speech enhancement algorithm such as spectral subtraction or Wiener filtering is used to reconstruct the speech signal, further improving the speech quality. The processed speech signal data is output to ensure that its intelligibility and quality meet the user's requirements.

[0098] The application can conveniently collect user's voice data through the built-in microphone of the earphone, provide input for subsequent voiceprint processing, help establish personal voiceprint model, and perform voiceprint recognition and verification. By reducing the sampling rate and applying filter, the number of sampling points of original voiceprint data can be reduced, the calculation and storage overhead can be reduced, and the key frequency components can be retained, which helps improve the efficiency and processing speed of the system. Through the mute detection and endpoint detection algorithm, the silent segment and non-speech segment can be identified and removed, and only the effective speech segment is retained, which helps reduce the interference of noise and non-speech components on voiceprint feature extraction, and improve the accuracy and reliability of voiceprint recognition. By applying data enhancement techniques such as sound reverberation, speed variation, volume adjustment, etc. to user voiceprint data, the diversity of data samples can be increased, and the robustness and generalization ability of the voiceprint model can be improved, which helps the voiceprint model to have good recognition performance under different voice environments and noise conditions. By constructing an LSTM model, the time sequence dependence relationship in the voiceprint enhanced data can be captured, and the feature representation of speech and noise can be extracted. This helps improve the understanding and differentiation ability of the voiceprint recognition system for voiceprint data. Through the full connection layer for feature fusion, the features extracted by the LSTM model are integrated to generate the time-frequency mask of speech and noise, which helps separate the speech and noise components and improve the clarity of the speech signal and the noise suppression effect. By applying the long short-term memory network model to process the input features, a speech-noise separation mask data can be generated to indicate the relative energy distribution of speech and noise, which helps accurately separate the speech and noise components and improve the clarity and audibility of the sound. By applying the speech-noise separation mask to the user voiceprint data, the original speech signal can be restored and the noise component can be removed, which helps improve the quality and clarity of the speech signal and makes it easier to understand and analyze.

[0099] Preferably, step S13 comprises the following steps:

[0100] Step S131: adding different types and intensities of environmental noise to the user voiceprint data, and changing the speech speed, thereby generating audio data of different speech speed versions;

[0101] The embodiment of the application randomly selects different types and intensities of environmental noise and mixes them into the user voiceprint data; for example, by linear mixing, the noise signal and the speech signal are added with different weights to generate audio data of multiple noise versions. The playback speed of the speech signal is changed using an audio processing tool. By adjusting the sampling rate of the audio or using a time stretching algorithm (such as PSOLA), audio data of different speech speeds (such as 0.8x, 1.2x) is generated.

[0102] Step S132: performing pitch shift simulation on the audio data to obtain a voiceprint enhanced data set;

[0103] The embodiment of the present application uses a pitch shift algorithm to process audio data, generates different pitch version audio data by changing the fundamental frequency of the audio signal without changing the speech rate, and forms a voiceprint enhancement dataset containing multiple noise, speech rate and pitch changes by using the processed audio data set.

[0104] Step S133: Short-time Fourier transform is performed on the voiceprint enhancement dataset to convert the time domain signal into a time-frequency representation, thereby obtaining the spectrogram data.

[0105] In the embodiment of the present application, short-time Fourier transform is applied to each audio segment in the voiceprint enhancement dataset. First, the audio signal is segmented into multiple short time frames (such as 25 ms), and then Fourier transform is applied to each frame to obtain the spectrum of each frame. The spectrum information of all frames is spliced to generate a spectrogram of the entire audio segment, which is represented as a time-frequency two-dimensional matrix.

[0106] Step S134: Mel frequency cepstral coefficient features are extracted from the spectrogram data, and a fundamental frequency contour is calculated, thereby obtaining pitch feature data.

[0107] In the embodiment of the present application, a Mel frequency filter bank is applied to the spectrum data of each frame to calculate the Mel frequency cepstral coefficient (MFCC) of each frame, which reflects the short-time power spectrum characteristics of the speech signal. The YIN algorithm or ACF method is used to calculate the fundamental frequency of each frame to generate a fundamental frequency contour, which reflects the pitch change of the speech.

[0108] Step S135: Spectral centroid and spectral flux acoustic feature extraction is performed on the pitch feature data, thereby obtaining acoustic feature data.

[0109] In the embodiment of the present application, the spectral centroid, i.e., the center of gravity of the spectrum, is calculated for each frame of spectrum, which reflects the center frequency of the spectrum. The spectral change rate of adjacent frames is calculated, which reflects the dynamic change of the spectrum. The difference between the spectrum of each frame and the spectrum of the previous frame is summed to obtain the spectral flux.

[0110] Step S136: The voiceprint enhancement dataset is screened based on the acoustic feature intensity according to the acoustic feature data, thereby obtaining the voiceprint enhancement data.

[0111] In the embodiment of the present application, all extracted acoustic features, including MFCC, fundamental frequency contour, spectral centroid and spectral flux, are calculated for each audio segment in the voiceprint enhancement dataset. According to the set feature intensity threshold, audio segments meeting the requirements are screened out. The K-means clustering method can be used to divide the audio segments into several categories, and representative and high-quality acoustic feature data is selected. After screening, audio segments with high feature intensity and good quality are retained to form the final voiceprint enhancement data.

[0112] The voiceprint model has robustness to voiceprint data in different environments by adding different types and intensities of environmental noise to simulate different speech environments and noise conditions. Meanwhile, the diversity of data samples is increased and the generalization ability of the voiceprint model is improved by generating audio data of different speech rates by changing the speech rate. The recognition performance of the voiceprint model to voiceprint data under different pitch conditions is improved by simulating pitch shift on the audio data to change the pitch of the sound and increase the diversity of data samples, which helps to make the voiceprint model robust to changes in the pitch of different people's voices. The time-frequency representation, i.e., the spectrogram data, is converted from the time-domain signal in the voiceprint enhancement dataset by applying the short-time Fourier transform, which helps to capture the energy distribution and spectral characteristics of the sound at different frequencies, providing a basis for voiceprint feature extraction and analysis. The mel-frequency cepstral coefficient (MFCC) feature is extracted from the spectrogram data to capture the spectral envelope and resonance characteristics of the sound for voiceprint feature representation. In addition, the pitch frequency contour calculation can estimate the fundamental frequency information of the sound to provide pitch-related feature data. The acoustic features of the sound can be extracted by calculating the spectral centroid and spectral flux of the pitch feature data. The spectral centroid reflects the concentration and distribution of the spectrum, and the spectral flux represents the rate of change of the spectrum, and these features have strong discriminability for distinguishing different sounds and noise components. By screening based on the strength of the acoustic features, voiceprint enhancement data with more obvious and prominent acoustic features can be selected, and data samples with more noise interference or less obvious features can be reduced, which helps to improve the training effect and recognition accuracy of the voiceprint model.

[0113] Preferably, step S16 comprises the following steps:

[0114] Step S161: performing short-time Fourier transform on the user voiceprint data to obtain original spectrogram data;

[0115] In the embodiment of the present application, the user voiceprint data is segmented into multiple short time frames (such as 25ms per frame), and the frames overlap (such as 50%). A Hamming or Hanning window function is applied to each frame to reduce spectral leakage. Fourier transform is applied to each frame to convert the time-domain signal to the frequency domain signal to obtain the spectrum of each frame. The spectrum information of all frames is spliced to generate a spectrogram of the entire audio segment, represented as a time-frequency two-dimensional matrix.

[0116] Step S162: applying the speech-noise separation mask data to the original spectrogram data to obtain pure speech spectrogram data;

[0117] In the embodiment of the present application, a pre-trained LSTM or CNN model is used to input the original spectrogram data to generate a speech-noise separation mask. The separation mask is multiplied point by point with the original spectrogram data to filter out the noise spectrum and retain the speech spectrum to obtain pure speech spectrogram data.

[0118] Step S163: phase reconstruction is performed on the pure speech spectrogram data by using a Griffin-Lim algorithm, so as to obtain speech spectrogram optimization data;

[0119] In the embodiment of the application, the amplitude spectrum of the pure speech spectrogram is used to randomly initialize the phase spectrum; the amplitude spectrum and the initial phase spectrum are combined to perform inverse Fourier transform to obtain a time domain signal; the time domain signal is subjected to short-time Fourier transform to update the phase spectrum; the new phase spectrum is combined with the amplitude spectrum of the pure speech spectrogram to perform inverse Fourier transform again; the above steps are repeated multiple times to gradually optimize the phase spectrum until convergence, and the optimized speech spectrogram data is obtained.

[0120] Step S164: inverse short-time Fourier transform is performed on the speech spectrogram optimization data, so as to obtain preliminary reconstructed speech data;

[0121] In the embodiment of the application, the inverse short-time Fourier transform is performed on the optimized speech spectrogram data to convert the data into a time domain signal, the overlapping area between frames is processed to ensure smooth connection of the audio signal, and all time domain signal frames are combined to obtain complete preliminary reconstructed speech data.

[0122] Step S165: smoothing processing is performed on the preliminary reconstructed speech data by using a Wiener filter, so as to obtain smoothed speech signal data;

[0123] In the embodiment of the application, the noise power spectrum in the preliminary reconstructed speech data is estimated, and a Wiener filter is designed according to the speech signal and the noise power spectrum; the preliminary reconstructed speech data is subjected to frequency domain smoothing processing by the Wiener filter to reduce residual noise, and the smoothed speech signal data is output to reduce the influence of noise.

[0124] Step S166: dynamic range compression is performed on the smoothed speech signal data based on audio loudness and intelligibility, so as to obtain dynamically compressed speech signal data.

[0125] In the embodiment of the application, the loudness analysis tool is used to measure the overall loudness of the smoothed speech signal, and the threshold, ratio, attack time and release time and other parameters of the dynamic range compression are set according to the loudness and intelligibility requirements; the smoothed speech signal is subjected to dynamic range compression to adjust the dynamic range of the audio signal and maintain the consistency of the intelligibility and loudness; the dynamically compressed speech signal data is obtained to ensure the consistency of the auditory effect of the audio under different loudness conditions.

[0126] The Short-Time Fourier Transform (STFT) of the present application decomposes the sound signal into different frequency components and provides information on time and frequency. By applying STFT to the user's voiceprint data, the energy distribution of the sound signal at different time periods and frequencies can be obtained. The speech-noise separation mask is used to separate the noise components from the speech components in the original spectrogram. By applying speech-noise separation mask data, pure speech spectrogram data can be extracted, thereby removing the interference of background noise on the speech signal. The Griffin-Lim algorithm is used to estimate the phase information of the speech signal. By combining the pure speech spectrogram data with the estimated phase information, the optimization data of the speech spectrogram is obtained. This step helps to restore the time-domain characteristics of the speech signal, making the reconstructed speech more natural and understandable. The Inverse Short-Time Fourier Transform (ISTFT) converts the optimized speech spectrogram data back to the time-domain signal. By applying the Inverse Short-Time Fourier Transform, the preliminary reconstructed speech data can be recovered from the frequency domain. The Wiener filter is a commonly used signal processing filter used to reduce noise interference and enhance the clarity of the speech signal. By applying the Wiener filter, the preliminary reconstructed speech data can be smoothed, reducing the impact of noise and improving the quality and intelligibility of the speech. Dynamic range compression is an audio processing technique used to adjust the volume dynamic range of the speech signal, making the audio listening experience more balanced and consistent at different volume levels. By applying dynamic range compression based on audio loudness and clarity, the volume of the speech signal can be more stable, improving the audibility and comfort of the speech.

[0127] Preferably, step S2 comprises the following steps:

[0128] Step S21: performing linear prediction coefficient calculation on the speech signal data to obtain LPC feature data;

[0129] Step S22: performing phase information extraction on the speech signal data and performing group delay feature evaluation based on the phase information to obtain group delay feature data;

[0130] Step S23: performing feature concatenation on the LPC feature data and the group delay feature data to obtain preliminary fusion feature data;

[0131] Step S24: performing z-score normalization on the preliminary fusion feature data to obtain user voiceprint feature data;

[0132] Step S25: modeling the user voiceprint feature data by using a Gaussian mixture model, thereby obtaining a GMM voiceprint model; performing feature extraction based on an i-vector framework on the GMM voiceprint model, thereby obtaining an i-vector voiceprint representation;

[0133] Step S26: constructing a deep neural network-based DNN voiceprint model according to the user voiceprint feature data, and performing d-vector feature extraction from a bottleneck layer of the DNN voiceprint model, thereby obtaining a d-vector voiceprint representation;

[0134] Step S27: integrating the GMM voiceprint model, the i-vector voiceprint representation, and the d-vector voiceprint representation, thereby obtaining a comprehensive user voiceprint model.

[0135] As an embodiment of the present application, refer to Figure 3 Fig. 1 shows a detailed step flowchart of step S2 in the present application, and step S2 in the embodiment of the present application includes the following steps: Figure 1

[0136] Step S21: performing linear prediction coefficient calculation on the speech signal data, thereby obtaining LPC feature data;

[0137] In the embodiment of the present application, the continuous speech signal is cut into short time period frames, usually with a duration of 20-30 milliseconds, and pre-emphasis is a kind of high-pass filter, which is used to emphasize the high frequency part and reduce the energy attenuation of the low frequency part. This can be achieved by performing first-order difference operation on the samples of each speech frame. Autocorrelation is used to estimate the linear correlation of the speech signal, and the autocorrelation coefficient is obtained by calculating the correlation between each frame of speech signal and its delayed version. The Levinson-Durbin algorithm is a recursive algorithm used to calculate the linear prediction coefficient from the autocorrelation coefficient, which recursively updates the prediction error and the linear prediction coefficient until the desired prediction order is obtained. The linear prediction coefficient extracted from the output of the Levinson-Durbin algorithm is used to describe the spectral characteristics of the speech signal, i.e., the LPC feature data.

[0138] Step S22: performing phase information extraction on the speech signal data, and performing group delay feature evaluation based on the phase information, thereby obtaining group delay feature data;

[0139] ​The embodiment of the present application converts each speech frame from time domain to frequency domain to obtain the spectral representation of the speech frame; the phase information is extracted from the spectrum after Fourier transform, which is usually obtained by calculating the complex amplitude angle of each spectral point. The group delay feature is used to describe the short-time dynamic characteristics of the spectrum, which can be evaluated by calculating the difference of the spectral phase. A common method is to calculate the phase difference between adjacent spectral frames and statistically analyze these differences, such as calculating the mean, standard deviation, etc. These statistical features constitute the group delay feature data.

[0140] Step S23: concatenating the LPC feature data and the group delay feature data at the feature level to obtain preliminary fusion feature data;

[0141] The embodiment of the present application ensures that the number of frames of the LPC feature data and the group delay feature data is the same, which can be realized by interpolation or truncation of the frames. The LPC feature data and the group delay feature data in each frame are connected together in a certain order to form a larger feature vector. For example, the LPC feature data can be used as the first few feature dimensions, and then the group delay feature data can be used as the subsequent feature dimensions.

[0142] Step S24: z-score normalization of the preliminary fusion feature data to obtain user voiceprint feature data;

[0143] The embodiment of the present application calculates the mean and standard deviation of each feature dimension over the entire data set, subtracts the mean from the value of each feature dimension, and then divides by the standard deviation to obtain the normalized feature value.

[0144] Step S25: modeling the user voiceprint feature data using a Gaussian mixture model to obtain a GMM voiceprint model; and performing feature extraction based on an i-vector framework on the GMM voiceprint model to obtain an i-vector voiceprint representation;

[0145] The embodiment of the present application trains a GMM model using the user voiceprint feature data, which is composed of multiple Gaussian distributions, each of which represents the distribution of voiceprint features in different voiceprint categories; i-vector is a voiceprint representation method used to extract a low-dimensional representation of voiceprint features; it maps each voiceprint feature to a low-dimensional latent space, and then extracts the mapped vector as an i-vector voiceprint representation; this mapping process can be calculated by the parameters and statistical characteristics of the GMM voiceprint model.

[0146] Step S26: constructing a deep neural network-based DNN voiceprint model according to the user voiceprint feature data, and performing d-vector feature extraction from the bottleneck layer of the DNN voiceprint model to obtain a d-vector voiceprint representation;

[0147] The embodiment of the present application designs a deep neural network structure, including multiple hidden layers and activation functions, for modeling user voiceprint feature data. A large amount of labeled voiceprint data is used to train the DNN voiceprint model, and the model parameters are optimized by minimizing the difference between voiceprint categories. The bottleneck layer is a layer in the middle of the DNN voiceprint model, which can be regarded as a more abstract representation of the voiceprint features. By inputting the user voiceprint feature data into the DNN voiceprint model, the output of the bottleneck layer is obtained as the d-vector voiceprint representation. The output of this bottleneck layer is a low-dimensional vector that captures the high-level abstract representation of the voiceprint features.

[0148] Step S27: integrating the GMM voiceprint model, the i-vector voiceprint representation and the d-vector voiceprint representation to obtain a comprehensive user voiceprint model.

[0149] The embodiment of the present application calculates the mean and standard deviation of all components (i.e. each dimension of the vector) of the d-vector voiceprint representation, subtracts the mean from each component of the d-vector voiceprint representation, and then divides by the standard deviation to obtain the normalized voiceprint features.

[0150] The linear prediction coefficient (LPC) in the present application is a method for analyzing speech signals. By calculating the linear prediction coefficient of the speech signal, the formant frequency and bandwidth of the speech signal can be extracted, which has high discrimination ability for voiceprint recognition. The group delay feature is a feature that describes the phase information of the speech signal. By extracting the phase information of the speech signal and evaluating the group delay feature, the phase structure of the speech signal in the frequency domain can be obtained, which is very important for speech periodicity modeling and voiceprint feature extraction in voiceprint recognition. Feature concatenation is a process of combining and fusing different types of feature data. By concatenating LPC feature data and group delay feature data, more rich and diverse voiceprint features can be obtained, improving the performance and robustness of voiceprint recognition. z-score normalization is a common feature standardization method used to eliminate the scale difference between features. By performing z-score normalization on the preliminary fused feature data, the feature data can be mapped to a standard normal distribution with a mean of 0 and a standard deviation of 1, making the weights of different features more balanced and improving the robustness and reliability of the voiceprint model. Gaussian mixture model (GMM) is a common voiceprint modeling method used for voiceprint feature modeling and recognition. By applying GMM modeling to user voiceprint feature data, a probability model can be obtained to describe the distribution of voiceprint features. At the same time, by extracting i-vector features based on the i-vector framework, a more compact and discriminative i-vector voiceprint representation can be obtained. DNN has strong non-linear modeling ability and can learn more complex and abstract voiceprint feature representations, which helps to distinguish the voice features of different people and improve the discrimination performance of voiceprint recognition. DNN voiceprint model can learn robust feature representation to noise and changes through the use of large-scale training data, which makes the voiceprint model have certain anti-interference ability to environmental noise, speech quality change and other factors. DNN voiceprint model can be flexibly adjusted and expanded by increasing the number of network layers and adjusting the number of neurons, which makes the voiceprint model adapt to different voiceprint recognition tasks and application scenarios. The bottleneck layer of DNN voiceprint model generally has a lower dimension, which is usually much smaller than the dimension of the original speech feature. This dimension reduction can reduce the redundancy of the feature and improve the efficiency and accuracy of subsequent processing. d-vector feature is a high-discriminative voiceprint representation learned by DNN voiceprint model. Compared with the original speech feature, d-vector feature can better distinguish the voice features of different people, thereby improving the accuracy and robustness of voiceprint recognition.

[0151] Preferably, step S3 comprises the following steps:

[0152] Step S31: audio channel division of user voiceprint data is performed through a microphone array technology, so as to obtain spatial audio data;

[0153] In the embodiment of the application, a proper number of microphones are selected and arranged according to a certain geometric layout, for example, linear array, circular array or spherical array, etc. The user voiceprint data is input into the microphone array, and the sound signal is collected by the microphone, and a proper distance between the microphone array and the user is ensured to obtain a clear sound signal. The collected sound signal is decomposed into different audio channels by using the geometric layout of the microphone array and the signal processing algorithm, and each channel corresponds to a microphone in the array. In this way, the spatial audio data can be obtained, wherein each channel represents a sound signal in a different direction.

[0154] Step S32: time delay calculation based on a generalized cross-correlation algorithm is performed on the spatial audio data, so as to obtain sound source direction data;

[0155] In the embodiment of the application, the spatial audio data is preprocessed, including denoising, gain adjustment and the like, to improve the accuracy of calculation; the generalized cross-correlation algorithm is used to perform cross-correlation calculation on the signals between the audio channels to determine the relative time delay. The peak position of the cross-correlation function is usually used to estimate the time delay. According to the calculated time delay, the geometric layout of the microphone array and the sound source positioning algorithm are combined to estimate the direction of the sound source, and the commonly used methods include the maximum cross-correlation method, the minimum variance method and the like.

[0156] Step S33: noise source positioning processing is performed on the sound source direction data by using a multiple signal classification algorithm, so as to obtain noise source position data;

[0157] In the embodiment of the application, some features are extracted from the sound source direction data, for example, the angle of the sound source direction, energy and the like; a multiple signal classifier is trained by using known noise source data and target sound source data, for example, support vector machine (SVM), random forest (Random Forest) and the like. The sound source direction data is input into the trained classifier, and the classifier is used for prediction, so that the sound source is classified as a noise source or a target sound source, and the position data of the noise source is obtained according to the classification result.

[0158] Step S34: an acoustic transfer function model is constructed according to the noise source position data and the geometric layout of the earphone loudspeaker, and optimal reverse wave form calculation is performed on the acoustic transfer function model according to the least mean square error criterion, so as to obtain initial reverse waveform data;

[0159] According to the position data of the noise source and the geometric layout of the earphone speaker, an acoustic transfer function model is established, which describes the sound transmission process from the noise source to each earphone speaker, including the attenuation, reflection, etc. of the sound. By using the least mean square error criterion, the optimal reverse wave waveform is calculated by an optimization algorithm (such as the least square method), so that the sound played through the earphone speaker after the reverse wave waveform minimizes the superposition of the noise generated by the noise source at the user's ear position. According to the acoustic transfer function model and the optimal reverse wave waveform, the initial reverse wave waveform data is calculated.

[0160] Step S35: Phase correction is performed on the initial reverse wave waveform data to obtain optimal reverse wave data.

[0161] The embodiments of the present application use the initial reverse wave waveform data and the noise source waveform data to estimate the phase difference therebetween, and common methods include cross-correlation method, least mean square method, etc. According to the estimated phase difference, phase correction is performed on the initial reverse wave waveform data, which can be achieved by phase rotation or using a filter, etc. After phase correction, optimal reverse wave data is obtained, which can be used to generate a reverse sound wave to offset the noise source.

[0162] The microphone array of the present application can receive sound signals in different positions and directions, and by dividing the audio channel of the voiceprint data, the source direction and position information of the sound can be determined, which helps subsequent sound source positioning and reverse wave calculation. The microphone array can suppress environmental noise and noise through spatial difference reception, and by selecting appropriate microphone combinations or applying array signal processing algorithms, the noise interference on the voiceprint data can be effectively reduced, and the accuracy and robustness of voiceprint recognition can be improved. By calculating the time delay between different microphones, the direction and position of the sound source can be determined, and the generalized cross-correlation algorithm can effectively estimate the time difference of the sound signal between different microphones, thereby realizing the positioning of the sound source. The generalized cross-correlation algorithm can utilize more abundant sound feature information when calculating the time delay, providing higher positioning accuracy than traditional cross-correlation algorithms, which helps to accurately obtain sound source direction data and provide accurate input for subsequent noise source positioning and reverse wave calculation. The multiple signal classification algorithm can classify the sound signal according to the sound source direction data, and divide the sound source into target sound source and noise source. By identifying and positioning the noise source, the noise situation in the environment can be better understood, providing important information for subsequent reverse wave calculation. Noise source positioning processing can help effectively separate the target sound source from the background noise, and by separating the target sound source and the noise source, the reverse wave form can be more accurately calculated, improving the noise suppression effect and the performance of voiceprint recognition. The acoustic transfer function model describes the propagation process of sound between the earphone speaker and the microphone, and by constructing the acoustic transfer function model and calculating the optimal reverse wave form, the interference generated by the noise source on the earphone speaker can be canceled out, thereby realizing noise suppression and elimination. The optimal reverse wave form calculation is based on the acoustic transfer function model and the minimum mean square error criterion, and by calculating the optimal reverse wave form, a wave form opposite to the noise can be generated to cancel out the influence of the noise on the sound signal, which helps to improve the clarity and audibility of the sound. Phase correction can ensure that the reverse wave form is consistent with the phase of the original sound signal, and phase consistency is very important for sound reconstruction and restoration, which can reduce sound distortion and aberration, and improve sound quality and accuracy. Through phase correction, the original sound signal can be better restored, and phase correction can correct the phase shift of the reverse wave form to match the original sound signal perfectly, thereby realizing more accurate sound restoration effect.

[0163] Preferably, step S4 comprises the following steps:

[0164] Step S41: power spectral density and spectral entropy are performed on the original spectrum graph data to obtain noise feature data;

[0165] The embodiment of the present application converts the original audio signal into spectrum graph data, and can convert the time domain signal into the frequency domain signal by using a fast Fourier transform (FFT) method or the like. The amplitude square operation is performed on the spectrum graph data to obtain the power spectrum density of each frequency point, and the power spectrum density represents the signal strength at different frequencies. The probability density estimation is performed on the spectrum graph data, and then the estimated probability density is used to calculate the spectrum entropy, and the spectrum entropy is used to describe the complexity and randomness of the spectrum graph data.

[0166] Step S42: classifying the noise feature data by using the pre-trained convolutional neural network, and performing noise type identification, so as to obtain noise type identification data;

[0167] The embodiment of the present application trains a convolutional neural network model with good classification ability for noise types by using a large amount of labeled noise data set. The noise feature data extracted in step S41 is input into the pre-trained convolutional neural network model as input. According to the output of the convolutional neural network model, it is determined which pre-defined noise type the input noise belongs to, so as to obtain noise type identification data.

[0168] Step S43: performing adaptive filter initialization according to the noise type identification data, and setting the forgetting factor and the regularization parameter of the recursive least square algorithm, so as to obtain RLS configuration parameter data;

[0169] The embodiment of the present application selects an adaptive filter model suitable for the current noise type according to the noise type identification data, and performs initialization. For example, a recursive least square (RLS) filter can be used. According to the noise type identification data, the forgetting factor and the regularization parameter in the recursive least square algorithm are set, and these parameters are used to control the convergence speed and stability of the adaptive filter. The set recursive least square algorithm parameters are used as RLS configuration parameter data, which is used for subsequent noise estimation and reverse waveform adjustment.

[0170] Step S44: performing real-time processing on the user voiceprint data according to the RLS configuration parameter data, and performing time-varying characteristic estimation of the noise signal, so as to obtain noise estimation data;

[0171] The voiceprint data of the user is input into the adaptive filter for processing, and the adaptive filter usually uses the recursive least squares (RLS) algorithm for adaptive filtering. In step S43, the initialization of the RLS configuration parameter data is performed, and these parameters include the forgetting factor and the regularization parameter of the recursive least squares algorithm. The forgetting factor determines the degree of forgetting of historical data by the filter, and the regularization parameter is used to control the stability and convergence speed of the filter. According to the processed voiceprint data of the user in step S44 and the output of the adaptive filter, the time-varying characteristics of the noise signal can be estimated, which can be achieved by comparing the difference between the filter output and the voiceprint data of the user. According to the noise signal time-varying characteristic estimation in step S44, noise estimation data can be obtained, which reflects the noise situation in the current environment.

[0172] Step S45: adjusting the reverse wave waveform of the reverse wave optimal data according to the noise estimation data to obtain adaptive reverse wave data.

[0173] The reverse wave optimal data of the embodiment of the present application refers to the optimal reverse wave data obtained by an optimization algorithm under the condition of known noise estimation data. This optimization algorithm can be selected according to specific requirements and application scenarios, such as the mean square error (MSE) optimization algorithm. According to the noise estimation data, the waveform of the reverse wave optimal data is adjusted, which can be achieved by adding the reverse waveform of the noise estimation data to the reverse wave optimal data. The purpose of waveform adjustment is to offset the influence of noise on the voiceprint data of the user, so as to improve the accuracy and reliability of voiceprint recognition. After waveform adjustment, the obtained data is the adaptive reverse wave data, which can be used for noise suppression or noise reduction applications to improve the quality and clarity of the voiceprint signal.

[0174] The power spectral density and spectral entropy are important statistical features of the spectrogram data, the power spectral density represents the energy of each frequency component in the spectrum, and the spectral entropy reflects the complexity of the spectrum; by calculating these features, the relevant information of the noise signal can be extracted, which provides a useful reference for subsequent noise classification and processing. Noise feature data can be used to analyze and describe noise signals, power spectral density can reveal the energy distribution of noise signals at different frequencies, and spectral entropy can reflect the spectral complexity of noise signals; through the analysis of noise feature data, the characteristics and features of noise can be better understood, which provides guidance for subsequent noise processing algorithms. The convolutional neural network can learn the feature representation of different noise types during training. By inputting the noise feature data into the pre-trained convolutional neural network, the noise can be classified into different noise types, such as white noise, noise, traffic noise, etc. This helps better understand and identify the characteristics of different noise types, and provides targeted algorithms and parameter settings for subsequent noise processing. Noise type identification data represents the noise type information classified by the convolutional neural network. By obtaining noise type identification data, the main noise type contained in the input signal can be known, which provides an important basis for subsequent adaptive filter initialization and parameter setting. According to the noise type identification data, the appropriate adaptive filter structure and initialization parameters can be selected. The adaptive filter can be adjusted in real time according to the characteristics of the input signal and the noise type to better suppress the noise component. By initializing according to the noise type identification data, the performance and adaptability of the adaptive filter can be improved. The recursive least squares algorithm is a commonly used adaptive filter algorithm, which is used to adjust the filter weights in real time to minimize the error between the output signal and the expected signal; according to the noise type identification data, appropriate forgetting factor and regularization parameter can be set to balance the convergence speed and stability of the filter. The forgetting factor controls the degree of forgetting of historical input data by the filter, and the regularization parameter is used to control the smoothing degree of the filter weights; by reasonably setting these parameters, the adaptive filter can obtain better performance and stability in different noise environments. The RLS configuration parameter data provides the parameter setting of the adaptive filter, which can be applied to real-time processing of user voiceprint data; by inputting the user voiceprint data into the adaptive filter, the noise component can be suppressed in real time, and the quality and distinguishability of the voiceprint signal can be improved. According to the RLS configuration parameter data, the adaptive filter can estimate the time-varying characteristics of the noise signal, and by analyzing the difference between the input signal and the filter output signal, the estimation information about the noise signal can be obtained, which helps to understand the statistical characteristics and time-varying properties of the noise signal, and provides a basis for subsequent noise processing and suppression.According to the noise estimation data, the reverse wave optimal data can be adjusted to suppress the noise component. By appropriately weighting and adjusting the reverse wave according to the noise estimation data, the interference of noise on the signal can be reduced, and the quality and clarity of the signal can be improved. The reverse wave waveform adjustment according to the noise estimation data can realize adaptive noise suppression. The characteristics and intensity of the noise signal can change over time. By adjusting the reverse wave waveform according to the real-time noise estimation data, different noise environments can be adapted to and better noise suppression effects can be provided.

[0175] Preferably, step S5 comprises the following steps:

[0176] Step S51: complexity analysis is performed on the adaptive reverse wave data to obtain calculation complexity data, and the complexity of the current noise environment is estimated to obtain environment complexity data;

[0177] The embodiment of the present application performs complexity analysis on the adaptive reverse wave data, which can consider factors such as data length, data dimension, and calculation complexity of the algorithm. By analyzing these factors, calculation complexity data of the adaptive reverse wave data can be obtained. According to the characteristics and statistical information of the current noise environment, the complexity of the noise environment is estimated. The complexity of the noise environment can consider factors such as the intensity, spectral distribution, and time-varying nature of the noise.

[0178] Step S52: current battery power data is obtained, and the current working frequency and load of the digital signal processor are monitored to obtain processor state data;

[0179] The embodiment of the present application obtains the current power data of the battery through a system interface or a sensor. The current working frequency and load information of the digital signal processor (DSP) are obtained through a system interface or a performance monitoring tool.

[0180] Step S53: a power consumption prediction model is constructed according to the calculation complexity data, the environment complexity data, and the processor state data. Power consumption estimation is performed under different configurations according to the power consumption prediction model, so as to obtain power consumption prediction data;

[0181] The embodiment of the present application constructs a power consumption prediction model according to the calculation complexity data, the environment complexity data, and the processor state data. The model can be modeled using a machine learning algorithm (such as a regression model) or a rule-based method. The constructed power consumption prediction model is used to estimate power consumption under different configurations, which can include different processor working frequencies, load distribution, and noise reduction algorithm parameters.

[0182] Step S54: a multi-level performance-power balance strategy is formulated according to the power consumption prediction data and the battery power data, and the clock frequency of the DSP is dynamically adjusted to obtain power consumption control parameter data.

[0183] The embodiment of the present application formulates a multi-level performance-power balance strategy according to the power consumption prediction data and the battery power data, which can dynamically adjust the clock frequency of the DSP to realize the control of power consumption according to the actual demand and constraint conditions. According to the formulated multi-level performance-power balance strategy, the dynamic adjustment of the power consumption of the DSP is realized by controlling the clock frequency of the DSP, and the adjustment of the clock frequency can be realized by the system interface or the power management setting of the adjusting processor.

[0184] Step S55: According to the power consumption control parameter data, the parameter configuration of the digital signal processor is performed, the adaptive reverse wave data is input into the configured digital signal processor for waveform cancellation processing, so as to obtain the noise reduction audio data;

[0185] The embodiment of the present application formulates a multi-level performance-power balance strategy according to the power consumption prediction data and the battery power data, which can dynamically adjust the clock frequency of the DSP to realize the control of power consumption according to the actual demand and constraint conditions. According to the formulated multi-level performance-power balance strategy, the dynamic adjustment of the power consumption of the DSP is realized by controlling the clock frequency of the DSP, and the adjustment of the clock frequency can be realized by the system interface or the power management setting of the adjusting processor.

[0186] Step S56: The noise reduction audio data is input into the user voiceprint model for voiceprint feature extraction, so as to obtain the voiceprint clear audio data.

[0187] The embodiment of the present application formulates a multi-level performance-power balance strategy according to the power consumption prediction data and the battery power data, which can dynamically adjust the clock frequency of the DSP to realize the control of power consumption according to the actual demand and constraint conditions. According to the formulated multi-level performance-power balance strategy, the dynamic adjustment of the power consumption of the DSP is realized by controlling the clock frequency of the DSP, and the adjustment of the clock frequency can be realized by the system interface or the power management setting of the adjusting processor.

[0188] The present application can evaluate the running efficiency of algorithms on different hardware platforms by calculating the complexity data, thereby selecting the most suitable platform to perform processing tasks. Through the environmental complexity data, the parameters and algorithms in the subsequent steps can be adjusted according to the complexity of the noise environment, to improve the noise reduction effect and performance. Obtaining battery power data and monitoring the state of the processor can provide information about system resources, and the battery power data can help determine the energy supply of the current system, while the working frequency and load of the processor can reflect the performance and load of the current processor. By obtaining battery power data, corresponding power consumption control strategies and adjustments can be made according to the energy situation of the system to avoid energy depletion leading to system interruption or performance degradation. By monitoring the processor state data, the working condition of the processor can be understood in real time, and corresponding optimization can be made according to the load condition to improve the power consumption and performance balance of the system. Power consumption prediction data can help evaluate the power consumption of the system under different configurations, which helps to select the best configuration to minimize power consumption while meeting performance requirements. Through power consumption prediction data, power consumption optimization can be performed during the design and development stage to identify and solve possible power consumption problems in advance. Through the multi-level performance-power balance strategy, the clock frequency of the processor can be dynamically adjusted according to the power consumption demand and battery power of the system to ensure performance while controlling power consumption. The power consumption control parameter data provides guidance for power consumption optimization according to system requirements, which can help the system achieve the best performance and power consumption balance in different scenarios. By configuring the parameters of the digital signal processor, the power consumption optimization of the system can be realized according to the power consumption control parameter data, so that the processor can meet the performance requirements while minimizing power consumption. The generation of noise reduction audio data can provide clearer sound and reduce the interference of noise on voiceprint feature extraction, thereby improving the accuracy and performance of subsequent voiceprint recognition. Voiceprint clear audio data can provide better voiceprint features, which can help improve the accuracy and robustness of the voiceprint recognition system. Through noise reduction processing and voiceprint feature extraction, better user experience and voiceprint recognition performance can be provided, thereby supporting security authentication, identity verification and other functions in the voiceprint recognition application field.

[0189] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting, the scope of the present application being defined by the appended claims and not by the above description, therefore all variations falling within the meaning and scope of the equivalent elements of the application file are intended to be included in the present application.

[0190] The foregoing is considered as illustrative only of the principles of the application. Numerous modifications and changes will readily occur to those skilled in the art, and it is intended to embrace all such modifications and changes that fall within the scope of the application. Accordingly, the application is not to be restricted in scope to the specific embodiments disclosed herein but is to be accorded the full scope that the principles and novel features request appropriately granted.

Claims

1. A headphone control method based on voiceprint recognition and reverse wave cancellation for wind noise reduction, characterized in that, Includes the following steps: Step S1: Obtain user voiceprint data; The user's voiceprint data is augmented to obtain voiceprint augmented data; the voiceprint augmented data is separated from noise using a long short-term memory network to obtain the voice signal data. Step S2: Fusion of speech spectrum features and phase features on the speech signal data to generate user voiceprint feature data; A user voiceprint model is generated and established based on user voiceprint feature data; Step S3: The noise source is located in the user's voiceprint data using microphone array technology, and the optimal waveform and phase of the reverse wave are determined based on the noise source and the headphone speaker layout, thereby obtaining the optimal reverse wave data. Step S4: Use the recursive least mean square algorithm to track noise changes in real time based on user voiceprint data, and adjust the reverse wave waveform of the optimal reverse wave data to obtain adaptive reverse wave data. Step S5: Dynamically adjust the computing power and power consumption of the adaptive backwave data based on the digital signal processor to obtain power consumption control parameter data; perform noise cancellation and voiceprint extraction on the user voiceprint data according to the power consumption control parameter data, the adaptive backwave data and the user voiceprint model to obtain clear voiceprint audio data.

2. The headphone control method based on voiceprint recognition and reverse wave cancellation for wind noise reduction according to claim 1, characterized in that, Step S1 includes the following steps: Step S11: Collect the user's voice through the built-in microphone of the earphone to obtain the raw voiceprint data; Step S12: Downsample and filter the original voiceprint data, perform silence detection and endpoint detection, extract the effective speech segments, and thus obtain the user voiceprint data; Step S13: Perform data augmentation on the user's voiceprint data to obtain voiceprint augmentation data; Step S14: Construct a long short-term memory network model based on the voiceprint enhancement data. The long short-term memory network model includes an input layer for receiving voiceprint enhancement data, multiple LSTM layers for capturing temporal dependencies, a fully connected layer for feature fusion, and an output layer for generating time-frequency masks for speech and noise. Step S15: Process the input features according to the Long Short-Term Memory Network Model to obtain speech-noise separation mask data; Step S16: Reconstruct the user's voiceprint data based on the speech-noise separation mask data to obtain the speech signal data.

3. The headphone control method based on voiceprint recognition and reverse wave cancellation for wind noise reduction according to claim 2, characterized in that, Step S13 includes the following steps: Step S131: Add environmental noise of different types and intensities to the user's voiceprint data and change the speech rate to generate audio data with different speech rates; Step S132: Perform pitch shift simulation on the audio data to obtain the voiceprint enhancement dataset; Step S133: Perform a short-time Fourier transform on the voiceprint enhancement dataset to convert the time-domain signal into a time-frequency representation, thereby obtaining the spectrogram data; Step S134: Extract the Mel frequency cepstral coefficient features from the spectrogram data and perform fundamental frequency profile calculation to obtain pitch feature data; Step S135: Extract acoustic features of the spectral centroid and spectral flux from the pitch feature data to obtain acoustic feature data; Step S136: Based on the acoustic feature data, the voiceprint enhancement dataset is filtered according to the acoustic feature intensity to obtain voiceprint enhancement data.

4. The headphone control method based on voiceprint recognition and reverse wave cancellation for wind noise reduction according to claim 3, characterized in that, Step S16 includes the following steps: Step S161: Perform a short-time Fourier transform on the user's voiceprint data to obtain the original spectrogram data; Step S162: Apply the speech-noise separation mask data to the original spectrogram data to obtain clean speech spectrogram data; Step S163: Use the Griffin-Lim algorithm to reconstruct the phase of the clean speech spectrogram data to obtain optimized speech spectrogram data; Step S164: Perform inverse short-time Fourier transform on the optimized speech spectrogram data to obtain preliminary reconstructed speech data; Step S165: Use a Wiener filter to smooth the initially reconstructed speech data to obtain smoothed speech signal data; Step S166: Perform dynamic range compression on the smooth speech signal data based on audio loudness and intelligibility to obtain dynamically compressed speech signal data.

5. The headphone control method based on voiceprint recognition and reverse wave cancellation for wind noise reduction according to claim 4, characterized in that, Step S2 includes the following steps: Step S21: Calculate the linear prediction coefficients for the speech signal data to obtain LPC feature data; Step S22: Extract phase information from the speech signal data and evaluate the group delay features based on the phase information to obtain group delay feature data; Step S23: Perform feature concatenation on LPC feature data and group delay feature data to obtain preliminary fused feature data; Step S24: Perform z-score normalization on the preliminary fused feature data to obtain user voiceprint feature data; Step S25: Model the user's voiceprint feature data using Gaussian Mixture Model to obtain the GMM voiceprint model; perform feature extraction on the GMM voiceprint model based on the i-vector framework to obtain the i-vector voiceprint representation. Step S26: Construct a DNN voiceprint model based on deep neural network according to the user's voiceprint feature data, and extract d-vector features from the bottleneck layer of the DNN voiceprint model to obtain the d-vector voiceprint representation. Step S27: Integrate the GMM voiceprint model, i-vector voiceprint representation, and d-vector voiceprint representation to obtain a comprehensive user voiceprint model.

6. The headphone control method based on voiceprint recognition and reverse wave cancellation for wind noise reduction according to claim 5, characterized in that, Step S3 includes the following steps: Step S31: Divide the user's voiceprint data into audio channels using microphone array technology to obtain spatial audio data; Step S32: Calculate the time delay of the spatial audio data based on the generalized cross-correlation algorithm to obtain the sound source direction data; Step S33: Use a multi-signal classification algorithm to perform noise source localization processing on the sound source direction data to obtain noise source location data; Step S34: Construct an acoustic transfer function model based on the noise source location data and the geometric layout of the headphone speaker, and calculate the optimal reverse wave waveform based on the acoustic transfer function model using the minimum mean square error criterion to obtain the initial reverse wave waveform data. Step S35: Perform phase correction on the initial reverse waveform data to obtain the optimal reverse wave data.

7. The headphone control method based on voiceprint recognition and reverse wave cancellation for wind noise reduction according to claim 6, characterized in that, Step S4 includes the following steps: Step S41: Calculate the power spectral density and spectral entropy of the original spectrogram data to obtain noise characteristic data; Step S42: Use a pre-trained convolutional neural network to classify the noise feature data and identify the noise type to obtain noise type identification data; Step S43: Initialize the adaptive filter based on the noise type identification data, and set the forgetting factor and regularization parameters of the recursive least mean square algorithm to obtain the RLS configuration parameter data; Step S44: Process the user's voiceprint data in real time according to the RLS configuration parameter data, and estimate the time-varying characteristics of the noise signal to obtain noise estimation data; Step S45: Adjust the reverse wave waveform of the optimal reverse wave data based on the noise estimation data to obtain adaptive reverse wave data.

8. The headphone control method based on voiceprint recognition and reverse wave cancellation for wind noise reduction according to claim 7, characterized in that, Step S5 includes the following steps: Step S51: Perform complexity analysis on the adaptive backwave data to obtain computational complexity data; and estimate the complexity of the current noise environment to obtain environmental complexity data. Step S52: Obtain the current battery power data and monitor the current operating frequency and load of the digital signal processor to obtain processor status data; Step S53: Construct a power consumption prediction model based on computational complexity data, environmental complexity data, and processor state data; Power consumption is estimated under different configurations based on the power consumption prediction model, thereby obtaining power consumption prediction data; Step S54: Formulate a multi-level performance-power balance strategy based on power consumption prediction data and battery power data, and dynamically adjust the DSP clock frequency to obtain power consumption control parameter data. Step S55: Configure the parameters of the digital signal processor according to the power consumption control parameter data, and input the adaptive reverse wave data into the configured digital signal processor for waveform cancellation processing to obtain noise reduction frequency data; Step S56: Input the noise-reduced audio data into the user's voiceprint model to extract voiceprint features, thereby obtaining clear audio data with a clear voiceprint.

Citation Information

Patent Citations

  • Intelligent noise elimination system with directional noise reduction function

    CN106504761A

  • Call noise reduction method based on voiceprint recognition, call noise reduction device and earphone

    CN114724565A