Earphone-based noise processing method, earphone and storage medium
The complex spectrum matrix of noisy audio is processed by a sliding converter network and enhanced by ratio masking, which solves the speech distortion problem caused by energy threshold suppression in existing headphone technology and improves speech clarity and recognition accuracy in open environments.
Patent Information
- Application Number
- CN202511022615.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-24
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-07-24
AI Technical Summary
Existing headphone technology that uses energy threshold-based noise suppression can easily cause speech distortion, resulting in low speech clarity, especially in open environments with severe wind noise interference.
A sliding converter network is used to process the complex spectrum matrix of the noisy audio collected by a microphone. The ratio mask is calculated by predicting the complex spectrum consisting of the real and imaginary parts, and the complex spectrum matrix is enhanced. The inverse short-time Fourier transform is then performed to obtain the noise-suppressed audio.
It effectively avoids noise cutting directly based on audio energy threshold, improves the audio clarity after noise reduction, and enhances speech recognition accuracy.
Smart Images

Figure CN120568245B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of wireless earphones, and particularly relates to an earphone-based noise processing method, an earphone and a storage medium. BACKGROUND
[0002] TWS (True Wireless Stereo, true wireless surround sound) earphones are prone to wind noise interference in open environments such as streets, cycling, subway stations, etc., leading to sound pickup distortion and seriously affecting voice clarity and recognition accuracy.
[0003] Current noise reduction techniques are usually based on the energy difference between noise and target signals, and a set threshold to achieve noise suppression. However, this energy threshold-based noise suppression method usually cuts off all noise, which can easily cause voice distortion and reduce voice clarity.
[0004] The above content is only used to assist in understanding the technical solutions of the present application and does not represent the acknowledgement of the above content as prior art. SUMMARY
[0005] The main purpose of the present application is to provide an earphone-based noise processing method, an earphone and a storage medium, aiming to solve the technical problem of low voice clarity caused by energy threshold-based noise suppression.
[0006] To achieve the above purpose, the present application provides an earphone-based noise processing method, which comprises:
[0007] inputting the real part and the imaginary part of the complex spectrum matrix of the noise audio collected by the microphone into a sliding converter network to obtain a predicted real part and a predicted imaginary part;
[0008] based on the complex spectrum composed of the predicted real part and the predicted imaginary part, calculating a ratio mask between the complex spectrum and the complex spectrum matrix;
[0009] enhancing the complex spectrum matrix based on the ratio mask to obtain a target complex spectrum;
[0010] performing inverse short-time Fourier transform on the target complex spectrum to obtain a noise-reduced audio.
[0011] In an embodiment, the output result of the sliding converter network further includes an amplitude spectrum, and before the step of inputting the real part and the imaginary part of the complex spectrum matrix of the noise audio collected by the microphone into the sliding converter network to obtain the predicted real part and the predicted imaginary part, the earphone-based noise processing method further comprises:
[0012] obtaining a noise spectrum corresponding to a pre-trained noise audio and a real spectrum corresponding to a reference audio;
[0013] inputting the noise spectrum into the sliding transformer network to be trained to obtain a real part prediction, an imaginary part prediction, and an amplitude spectrum, wherein the real part prediction and the imaginary part prediction constitute a predicted spectrum;
[0014] calculating a ratio mask between the predicted spectrum and the noise spectrum, and enhancing the noise spectrum according to the ratio mask to obtain a fusion spectrum;
[0015] updating generator parameters and discriminator parameters of the sliding transformer network according to the fusion spectrum, the predicted spectrum, the real spectrum, and the amplitude spectrum based on a loss function to obtain the trained sliding transformer network.
[0016] In an embodiment, the step of updating the generator parameters and the discriminator parameters of the sliding transformer network according to the fusion spectrum, the predicted spectrum, the real spectrum, and the amplitude spectrum based on a loss function to obtain the trained sliding transformer network comprises:
[0017] performing inverse short-time Fourier transform on the fusion spectrum to obtain a denoised signal, and calculating a first loss value of the denoised signal on a time domain reconstruction error based on a first loss function;
[0018] calculating a second loss value of the denoised signal on a signal reconstruction quality based on a second loss function;
[0019] calculating a third loss value of the amplitude spectrum based on a third loss function;
[0020] calculating a fourth loss value between the predicted spectrum and the real spectrum based on a discriminator;
[0021] fusing the first loss value, the second loss value, the third loss value, and the fourth loss value to obtain a target loss value;
[0022] updating the generator parameters and the discriminator parameters based on the target loss value to obtain the trained sliding transformer network.
[0023] In an embodiment, the step of fusing the first loss value, the second loss value, the third loss value, and the fourth loss value to obtain a target loss value comprises:
[0024] determining a norm of the first loss value, a complement of the second loss value and a first weight, a second weight of the third loss value, and a third weight of the fourth loss value;
[0025] accumulating a product of the norm, the complement, and the first weight, a product of the third loss value and the second weight, and a product of the fourth loss value and the third weight to obtain the target loss value.
[0026] In one embodiment, the step of enhancing the complex spectrum matrix based on the ratio mask to obtain a target complex spectrum includes:
[0027] The product of the ratio mask and the complex spectrum matrix is calculated, and the product is set as the target complex spectrum.
[0028] In one embodiment, the step of inputting the real part and the imaginary part of the complex spectrum matrix of the noise audio collected by the microphone into the sliding converter network to obtain the predicted real part and the predicted imaginary part includes:
[0029] Performing windowing and short-time Fourier transform on the noise audio to obtain the complex spectrum matrix;
[0030] Inputting the real and imaginary parts of the complex spectrum matrix of the noisy audio into a sliding converter network, and extracting and correlating local structural features of the complex spectrum matrix based on the complex convolutional layer of the multi-scale encoder of the sliding converter network;
[0031] The cross-layer transfer features are obtained based on the decoder, and after upsampling processing is performed on the cross-layer transfer features, the predicted real part and the predicted imaginary part are output based on the output layer.
[0032] In one embodiment, the microphone is a dual-channel microphone. Before the step of inputting the real and imaginary parts of the complex spectrum matrix of the noise audio collected by the microphone into a sliding converter network to obtain the predicted real and imaginary parts, the headphone-based noise processing method further includes:
[0033] Acquire dual-channel audio captured by the dual-channel microphone;
[0034] An energy difference between the two-channel audio is calculated, and when the energy difference is greater than a preset threshold and the duration is greater than a preset duration, the two-channel audio is determined to be the noise audio.
[0035] In one embodiment, after the step of inputting the real part and the imaginary part of the complex spectrum matrix of the noise audio collected by the microphone into the sliding converter network to obtain the predicted real part and the predicted imaginary part, the headphone-based noise processing method further includes:
[0036] Inputting the complex spectrum into a complex phase refinement model to obtain the complex spectrum with updated phase;
[0037] Based on the complex spectrum after phase update, the step of calculating a ratio mask between the complex spectrum and the complex frequency spectrum matrix based on the complex spectrum composed of the predicted real part and the predicted imaginary part is performed.
[0038] In addition, to achieve the above-mentioned purpose, the present application also proposes an earphone, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the earphone-based noise processing method as described above.
[0039] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium and stores a computer program. When the computer program is executed by a processor, the steps of the headphone-based noise processing method described above are implemented.
[0040] One or more technical solutions proposed in this application have at least the following technical effects:
[0041] When noise is present in the audio signal currently being collected by the headphones, the audio signal is converted to obtain a complex spectrum matrix. The real and imaginary parts of the spectrum matrix are then input into a sliding converter network as training data to obtain predicted real and imaginary parts. Based on the complex spectra corresponding to the predicted real and imaginary parts, a ratio mask is calculated between the complex spectrum and the complex spectrum matrix. The complex spectrum matrix is then enhanced using the ratio mask, so that an inverse short-time Fourier transform is performed based on the enhanced target complex spectrum to obtain the noise-suppressed frequency spectrum. By calculating the predicted output of the complex spectrum matrix using the trained sliding converter network and combining it with the ratio mask fusion inference path, the spectrum of the noisy audio is reconstructed, avoiding direct noise reduction based on the audio energy threshold and improving the clarity of the audio after noise reduction. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0043] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0044] Figure 1 A flowchart illustrating the first embodiment of the headphone-based noise processing method of the present application;
[0045] Figure 2 This is a brief flowchart of the headphone-based noise processing method of this application;
[0046] Figure 3 A flowchart illustrating a second embodiment of the headphone-based noise processing method of the present application;
[0047] Figure 4 This is a schematic diagram of the model training process of the headphone-based noise processing method of this application;
[0048] Figure 5 Schematic diagram of the device structure of the hardware operating environment involved in the headphone-based noise processing method in the embodiment of the present application.
[0049] The purpose, features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0050] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.
[0051] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.
[0052] TWS earphones are susceptible to wind noise interference in open environments such as streets, cycling, and subway stations, resulting in sound pickup distortion, which seriously affects voice clarity and recognition accuracy.
[0053] Current noise reduction technologies typically suppress noise based on the energy difference between the noise and the target signal, along with a set threshold. However, this energy threshold-based noise suppression approach often treats noise in a one-size-fits-all manner, which can easily cause speech distortion and reduce speech clarity.
[0054] Based on this, the main solution of the embodiment of the present application is: inputting the real part and imaginary part of the complex spectrum matrix of the noise audio collected by the microphone into the sliding converter network to obtain the predicted real part and the predicted imaginary part;
[0055] Calculating a ratio mask between the complex spectrum and the complex spectrum matrix based on the complex spectrum consisting of the predicted real part and the predicted imaginary part;
[0056] enhancing the complex spectrum matrix based on the ratio mask to obtain a target complex spectrum;
[0057] An inverse short-time Fourier transform is performed on the target complex spectrum to obtain a noise suppression frequency.
[0058] Specifically, the predicted output of the complex spectrum matrix is calculated through the trained sliding converter network, and combined with the inference path of ratio mask fusion, the spectrum reconstruction of noisy audio is achieved, avoiding direct noise cutting based on the audio energy threshold, and improving the audio clarity after noise reduction.
[0059] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, mobile phone, etc., or an electronic device or headset capable of implementing the above functions. The following uses headsets as an example to illustrate this embodiment and the following embodiments.
[0060] The embodiment of the present application provides a method for noise processing based on headphones, referring to Figure 1 , Figure 1 This is a flow chart of the first embodiment of the headphone-based noise processing method of the present application.
[0061] In this embodiment, the headphone-based noise processing method includes steps S10 to S40:
[0062] Step S10: input the real part and the imaginary part of the complex spectrum matrix of the noise audio collected by the microphone into the sliding converter network to obtain the predicted real part and the predicted imaginary part.
[0063] It should be noted that the complex spectrum matrix is a complex matrix representing the frequency domain obtained by transforming the audio signal through the Short-Time Fourier Transform (STFT). The real part corresponds to the cosine component of the spectrum, and the imaginary part corresponds to the sine component. The sliding transformer network is a Swin-transformer network.
[0064] In this embodiment, the presence of noise in the audio captured by the microphone can be determined by a preset signal-to-noise ratio threshold, a spectrum entropy detection algorithm, or a dual-channel audio energy difference detection method. When the audio is determined to be noisy audio, it is converted into an auxiliary spectrum matrix and then input into the Swin-transformer network. The headset can perform noise recognition based on the audio captured by a single-channel microphone, or based on the dual-channel audio captured by a dual-channel microphone. The noise can be wind noise when the user is riding, walking, or blowing in the wind, or other noise such as the superposition of human voices, traffic sounds, keyboard sounds, etc., which are not limited in this application.
[0065] Optionally, if the headset does not have a judgment module, it is assumed that all audio signals are noise signals, that is, noise reduction processing is performed on all collected signals.
[0066] As an optional implementation method for determining the presence of noise, during the process of identifying audio collected by a dual-channel microphone array, the dual-channel audio signals collected by the dual microphones can be obtained, and the energy difference between the two-channel audio signals can be calculated. If the energy difference is greater than a preset threshold and lasts longer than a preset duration, the presence of noise in the audio signal is determined. This detection method is typically used to determine whether wind noise is wind noise, but can also be used to determine other noise types.
[0067] Exemplarily, the energy difference between the two-channel audio signals is calculated as follows:
[0068] ,
[0069] Where η(t) is the wind noise discrimination function value. By calculating the energy difference between the two-channel microphones, it is determined whether there is wind noise at the current time t (range: -1 to 1). x1(t) represents the time domain audio signal of microphone channel 1 (usually close to the wind direction source), and x2(t) represents the time domain audio signal of microphone channel 2 (usually far away from the wind direction source). 2 The squared L2 norm of the signal represents the signal energy. ε is a very small positive number to avoid the denominator being zero. θ is the empirical threshold for wind noise discrimination. Wind noise detection is triggered when η(t)>θ and the temporal persistence constraint is met. The current frame is considered to be in a wind noise state. This application does not limit the method of noise determination based on single-channel audio.
[0070] By integrating energy ratios and spectral features, the earphones can adaptively identify wind noise types, such as mild disturbances and sustained strong winds, allowing them to dynamically switch noise reduction strategies and enhance the accuracy and robustness of wind noise processing. The same effect can be achieved for other noise types.
[0071] After determining that there is noise in the signal input from one or both sides of the microphone, the noisy audio is framed and windowed, and then STFT is performed on it to obtain a complex spectrum matrix. The complex matrix is then split into independent real and imaginary matrices, and the real and imaginary matrices are input as input parameters to the Swin-transformer network, and the output is the predicted real part. and the predicted imaginary part Among them, the output parameters also include the amplitude spectrum prediction The amplitude spectrum prediction is used in the training phase to construct the perceptual loss of the training network. It does not participate in the audio reconstruction process in the inference phase and only serves as an auxiliary supervision signal for training.
[0072] Step S20: Calculate a ratio mask between the complex spectrum and the complex frequency spectrum matrix based on the complex spectrum consisting of the predicted real part and the predicted imaginary part.
[0073] In this embodiment, the predicted real and imaginary parts output by the Swin-transformer network can form a new complex spectrum, which is the predicted complex spectrum. After combining the predicted real and imaginary parts to obtain the complex spectrum, a ratio mask between the two complex spectra can be calculated to enhance the spectrum based on the ratio mask.
[0074] Specifically, the predicted real part and the predicted imaginary part Constructing the complex spectrum The formula is as follows:
[0075] .
[0076] The formula for calculating the ratio mask is as follows:
[0077] ,
[0078] in, is the complex spectrum matrix, is the complex spectrum, is the ratio mask.
[0079] For example, the complex spectrum matrix is X=0.5+0.3j, and the complex spectrum is Y=0.8+0.1j. In this case, the calculated ratio mask M={0.8+0.1j} / {0.5+0.3j}≈1.21-0.56j.
[0080] It can be understood that the ratio mask can characterize the amplitude / phase difference between the predicted complex spectrum and the original complex spectrum, thereby maintaining the phase structure of the original spectrum while strengthening noise suppression in the amplitude domain.
[0081] Step S30: enhancing the complex spectrum matrix based on the ratio mask to obtain a target complex spectrum.
[0082] In this embodiment, spectrum enhancement can be performed by multiplying the original complex spectrum matrix element-by-element by a ratio mask to obtain an enhanced target complex spectrum. The product of the ratio mask and the complex spectrum matrix can be calculated to form the target complex spectrum. This target complex spectrum incorporates the phase adjustment of the ratio mask, reducing the mechanical quality of speech caused by neglecting signal phase in traditional noise processing.
[0083] Step S40: performing an inverse short-time Fourier transform on the target complex spectrum to obtain a noise-suppressed frequency spectrum.
[0084] In this embodiment, after obtaining the target complex spectrum, it is reconstructed back into the time domain through the ISTFT (Inverse Short-Time Fourier Transform), resulting in the noise-suppressed audio. This audio is the time-domain audio signal after the noise audio has been denoised. The headphones do not need to perform additional processing or mixing on it and can be directly played to the user through the speakers. Therefore, in the macro-level noise reduction process, the headphones collect audio in real time, perform noise reduction on the audio, and finally output the noise-suppressed audio after noise reduction to the user.
[0085] Exemplarily, the ratio mask M=1.21-0.56j, the complex spectrum matrix is X=0.5+0.3j, and the result of multiplying the two is approximately equal to 0.8+0.1j, which is close to the pure audio spectrum.
[0086] After obtaining the enhanced audio, it is processed through the audio pipeline inside the headset to obtain the target audio signal output by the speaker. The specific audio output process is not limited in this application. It is understood that the enhanced audio is the result of digital domain noise reduction. It is still a digital signal and needs to be further processed to obtain the final signal output by the speaker.
[0087] This embodiment enhances the original complex spectrum matrix through ratio masking, realizing the audio processing path of the signal to be denoised from STFT → SwinTransformer → mask fusion → iSTFT → output, thereby improving the noise reduction effect of audio with wind noise.
[0088] For example, to help understand the implementation process of the headphone-based noise processing method of this embodiment, please refer to Figure 2 , Figure 2 A brief flow chart of a headphone-based noise processing method is provided. Specifically, the headphone microphone collects audio s_n, and then the energy η_t of the audio is analyzed based on the wind noise discrimination module. If the result is that wind noise is present, the signal is converted to STFT to obtain S1_ft, which is then processed by the Swin transformer network to output the predicted spectrum. , then calculate the ratio mask The original spectrum is enhanced based on the ratio mask to obtain the fused spectrum S2_ft = M_ft x S1_ft. The fused spectrum is then reconstructed using iSTFT to obtain the noise-suppressed frequency spectrum s2_n, which is then output. In the absence of wind noise, the earphone outputs the original speech s_n.
[0089] Optionally, in addition to processing signals based on ratio masks, it is also possible to predict a soft mask (amplitude or complex number) based on time-frequency masks such as IRM, PSM, and CIRM, or reconstruct the audio after enhancing the phase based on phase reconstruction such as Griffin-Lim and PhaseNet, or perform end-to-end waveform prediction, such as processing audio through the Demucs series and directly outputting the time domain signal.
[0090] This embodiment provides a headphone-based noise processing method. Using a sliding converter network, it jointly models the real and imaginary parts of the complex spectrum, effectively capturing the nonlinear coupling between noise and speech in the complex plane. A ratio mask is then constructed in the complex domain to precisely quantize the speech components to be retained. This method implements an inference path that combines multiple complex spectrum outputs with ratio mask fusion, achieving spectrum reconstruction. Finally, complex spectrum enhancement simultaneously optimizes amplitude and phase information. This method significantly improves speech clarity and naturalness while suppressing background noise (such as fan noise and keyboard tapping), avoiding the speech clarity issues associated with traditional audio processing.
[0091] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the first embodiment can be referred to the above introduction and will not be described in detail later. Figure 3 Before step S10, the headphone-based noise processing method further includes steps S01 to S04:
[0092] Step S01: Obtain a noise spectrum corresponding to the pre-trained noise audio and a real spectrum corresponding to the reference audio.
[0093] In this embodiment, the training process of the sliding converter network is as follows: Figure 3 As shown, the speech pair input in the training phase is a noisy audio s_noisy containing noise and a clean control audio s_clean without noise. Then, the two audios are subjected to STFT processing to obtain corresponding complex noise spectra S_noisy and S_clean.
[0094] Step S02: input the noise spectrum into the sliding converter network to be trained to obtain a real part prediction, an imaginary part prediction and an amplitude spectrum.
[0095] In this embodiment, please continue to refer to Figure 3 After obtaining the noise spectrum and the real spectrum, the complex spectrum needs to be split into Re (Real Part), Im (Imaginary Part) and Mag (Magnitude). Based on this, the noise spectrum is input into the Swin transformer generator G to be trained, thereby outputting a three-way structure prediction spectrum: real part prediction , imaginary part prediction and the amplitude spectrum Among them, the real part prediction and the imaginary part prediction Composition prediction spectrum , that is, constructing the complex spectrum based on the real part prediction and the imaginary part prediction: .
[0096] Step S03: Calculate a ratio mask between the predicted spectrum and the noise spectrum, and enhance the noise spectrum according to the ratio mask to obtain a fused spectrum.
[0097] In this embodiment, please continue to refer to Figure 3 After obtaining the three-way output structure and constructing the prediction spectrum, calculate the prediction spectrum and the noise spectrum S_noisy, , and then the enhanced spectrum is fused based on the ratio mask to obtain the fused spectrum: .
[0098] Step S04 : Based on the loss function, the generator parameters and the discriminator parameters of the sliding converter network are updated according to the fused spectrum, the predicted spectrum, the true spectrum, and the amplitude spectrum to obtain the trained sliding converter network.
[0099] In this embodiment, after obtaining the fused spectrum, the predicted spectrum, the real spectrum, and the amplitude spectrum, it is necessary to construct a plurality of loss values so as to update and optimize the parameters of the sliding converter network based on the plurality of loss values.
[0100] For details, please refer to Figure 4 In the process of calculating the loss value, the fused spectrum S2 can be processed by STFT to obtain the denoised signal reconstructed in the time domain. , then calculate the first loss value L1 of the denoised signal on the time domain reconstruction error based on the first loss function to enhance the overall accuracy of the spectrum amplitude, and calculate the denoised signal based on the second loss function The second loss value SDR loss on the signal reconstruction quality is calculated based on the amplitude spectrum of the third loss function The third loss value SSIM loss is used to improve the similarity of the spectrum structure. The true spectrum S_clean is input into the discriminator D of the Swin transformer, and the fourth loss value, LSGAN loss, is calculated based on the discriminator between the predicted spectrum and the true spectrum.
[0101] Finally, based on the generator optimizer, the first loss value, the second loss value, the third loss value and the fourth loss value are fused to obtain the target loss value L_total of the joint loss, and the generator parameters G and the discriminator parameters D are updated based on the target loss value, and the training model data is included to obtain the trained sliding converter network.
[0102] The calculation formula for fusing multiple loss values is as follows:
[0103] ,
[0104] in, The norm of the L1 loss value, and therefore, in the process of fusing the loss values, the norm of the first loss value L1, the complement of the second loss value SSIM, and the first weight the second weight of the third loss value SDR the third weight of the fourth loss value LSGAN and then the product of the norm, the complement, and the first weight, the product of the third loss value and the second weight, and the product of the fourth loss value and the third weight are accumulated to obtain the target loss value .
[0105] In the model training process of the embodiment, the complex spectrum modeling method is used to output a predicted spectrum with a three-path structure, and at the same time, a complex spectrum, a ratio mask, and a fused and enhanced spectrum are constructed to realize data denoising. At the same time, the predicted frequency domain signal and the real spectrum of the original input real signal are input into the discriminator, and based on the multi-loss optimization fusion method, the data input and data reconstruction of the discriminator are on the same path to form a closed-loop optimization of model training, thereby improving the model training effect.
[0106] The embodiment provides a noise processing method based on a headset. In the training process of a Swin transformer network, a complex spectrum three-path output, a mask fusion, and a multi-loss optimization are combined to realize spectrum enhancement processing, thereby improving the clarity of the denoised audio. At the same time, the trained model can be deployed in a terminal low-power voice enhancement system such as a headset to realize low-power noise reduction processing.
[0107] Based on the first embodiment of the present application, in the third embodiment of the present application, the same or similar contents as the above first embodiment can be referred to the above introduction, and will not be described in detail. On this basis, step S10 includes steps S11-S13:
[0108] Step S11, windowing and short-time Fourier transform are performed on the noise audio to obtain the complex frequency spectrum matrix.
[0109] After detecting the noise signal, windowing and weighting processing are performed, and then STFT conversion is performed to obtain the input parameters of the Swin transformer. The specific Fourier transform process is not limited herein.
[0110] For example, a noisy speech with a sampling rate of 16 kHz is converted by STFT to obtain a complex matrix with T=120 frames and F=257 frequency points.
[0111] Step S12, the real part and the imaginary part of the complex frequency spectrum matrix of the audio signal are input into the sliding converter network, and based on the complex convolution layer of the multi-scale encoder of the sliding converter network, the local structure features of the complex frequency spectrum matrix are extracted and associated.
[0112] In this embodiment, the core of the sliding converter network includes a multi-stage encoder-decoder structure, and the first stage includes a complex convolution module. Therefore, the real part and the imaginary part of the input can be subjected to complex convolution operation based on the complex convolution layer, so as to output a local feature map capable of representing local structure features. The window of the multi-scale encoder is multi-layered, and the local time-frequency dependence can be modeled based on the self-attention weight of Q / K / V in each window, so as to realize the association of local structure features.
[0113] In step S13, the cross-layer transmission features are obtained based on the decoder, and after the cross-layer transmission features are subjected to up-sampling processing, the predicted real part and the predicted imaginary part are output based on the output layer.
[0114] In this embodiment, the decoder outputs the final predicted real part and the predicted imaginary part by means of up-sampling to restore the resolution, splicing with the encoder features, and convolution fusion.
[0115] Therefore, in the processing process, the output result of the complex convolution layer needs to be directly transmitted to the corresponding layer of the decoder, so as to be subjected to up-sampling in the decoder. After the up-sampling is completed, the features are subjected to splicing processing with the features sent by the encoder, and finally the predicted real part and the predicted imaginary part are obtained.
[0116] In this embodiment, the Swin Transformer models the local area in the spectrogram through the sliding window attention mechanism, extracts the structure mode of the local area and the cross-time frame, including extracting the local structure of the complex spectrum through the complex convolution layer (CConv2d), capturing the cross-frame attention under different windows through the multi-scale encoder (CW-MSA + CSW-MSA), and recovering the spectral resolution through the decoder layer by layer. Since the convolution layer, the multi-scale encoder, and the real part decoder and the imaginary part decoder are directly connected, the predicted real part and the predicted imaginary part can be directly obtained.
[0117] The embodiment provides a noise processing method based on a headset, which accurately extracts the time-frequency local structure through a complex convolution layer, associates the local structure features through a multi-scale encoder, effectively generates a real part prediction and an imaginary part prediction for constructing a final complex spectrum, improves the accuracy of the predicted real part / imaginary part, and further enhances the robustness of complex domain noise reduction.
[0118] Based on the first embodiment of the present application, in the fourth embodiment of the present application, the same or similar contents as the above first embodiment can be referred to the above description, and will not be repeated hereinafter. On this basis, after obtaining the predicted complex spectrum and before applying the mask, due to the global attention mechanism of the sliding converter network, the modeling of the amplitude spectrum is more focused, and the time sequence continuity of the phase information is insufficient, resulting in inter-frame jump error of the predicted phase (such as the breaking of the fundamental frequency trajectory); at the same time, the ratio mask is sensitive to phase distortion, and small phase deviation will be significantly amplified in complex division, causing the enhanced audio to produce water ripple artifacts. Therefore, the complex spectrum needs to be phase recovered or optimized.
[0119] Specifically, after step S10, step S40 of inputting the complex spectrum into a complex phase refinement model to obtain the complex spectrum with updated phase is further included. In this step, the complex phase refinement model models the complex spectrum and predicts the inter-frame error of the current phase of the complex spectrum, and if the error is small, it is not necessary to suppress, otherwise the inter-frame error of the phase is suppressed below the preset error, and then the complex spectrum is updated based on the suppressed phase information.
[0120] After updating the complex spectrum, the processing action of step S20 is performed based on the complex spectrum with updated phase, so as to further improve the clarity of the audio after noise reduction processing.
[0121] The present application provides an earphone, which comprises at least one processor, and a memory connected with the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the earphone-based noise processing method in the above first embodiment.
[0122] Reference will be made to the following description Figure 5 which shows a structural schematic diagram of an earphone suitable for implementing the embodiments of the present application. Figure 5 The earphone shown is only an example, and should not impose any limitation on the functions and use range of the embodiments of the present application.
[0123] As Figure 5As shown, the earphone can include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.) that can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 1002 or programs loaded from a storage device 1003 into a random access memory (RAM) 1004. Various programs and data required for earphone operation are also stored in the random access memory 1004. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems can be connected to the input / output interface 1006: input devices 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; the storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the earphone to communicate wirelessly or wiredly with other devices to exchange data. Although the earphone with various systems is shown in the figure, it should be understood that all the systems shown are not required to be implemented or possessed. More or fewer systems can be alternatively implemented or possessed.
[0124] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network through the communication device, or installed from the storage device 1003, or installed from the read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the methods of the embodiments of the present disclosure are performed.
[0125] The earphone provided by the present application adopts the earphone-based noise processing method in the above-mentioned embodiments, which can solve the technical problem of low speech intelligibility caused by noise suppression based on an energy threshold. Compared with the prior art, the earphone provided by the present application has the same beneficial effects as the earphone-based noise processing method provided by the above-mentioned embodiments, and the other technical features in the earphone are the same as the features disclosed in the previous embodiment method, which will not be repeated here.
[0126] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0127] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
[0128] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, a computer program) stored thereon, wherein the computer-readable program instructions are used to execute the headphone-based noise processing method in the above embodiment.
[0129] The computer-readable storage medium provided herein may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems, or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory, read-only memory, erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including, but not limited to, wires, optical cables, radio frequency (RF), etc., or any suitable combination thereof.
[0130] The computer-readable storage medium may be contained in the earphone, or may exist independently without being assembled into the earphone.
[0131] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by the headset, the headset:
[0132] Input the real part and the imaginary part of the complex spectrum matrix of the noise audio collected by the microphone into the sliding converter network to obtain the predicted real part and the predicted imaginary part;
[0133] Calculating a ratio mask between the complex spectrum and the complex spectrum matrix based on the complex spectrum consisting of the predicted real part and the predicted imaginary part;
[0134] enhancing the complex spectrum matrix based on the ratio mask to obtain a target complex spectrum;
[0135] An inverse short-time Fourier transform is performed on the target complex spectrum to obtain a noise suppression frequency.
[0136] Computer program code for performing the operations of the present application may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0137] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, program segment or part of code, and the module, program segment or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, as well as the combination of boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or can be implemented using a combination of dedicated hardware and computer instructions.
[0138] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.
[0139] The readable storage medium provided by the application is a computer readable storage medium, which stores computer readable program instructions (i.e. computer programs) for executing the above-mentioned earphone-based noise processing method, and can solve the technical problem of low voice intelligibility caused by noise suppression based on an energy threshold. Compared with the prior art, the computer readable storage medium provided by the application has the same beneficial effects as the earphone-based noise processing method provided by the above-mentioned embodiments, and will not be described here.
[0140] The above-mentioned is only part of the embodiments of the application, and does not limit the patent scope of the application. Any equivalent structural transformation, direct / indirect application in other related technical fields based on the technical concept of the application, and the contents of the specification and drawings are included in the patent protection scope of the application.
Claims
1. A noise processing method based on headphones, characterized in that: The headphone-based noise processing method includes: Obtain the noise spectrum corresponding to the pre-trained noise audio and the true spectrum corresponding to the control audio; Inputting the noise spectrum into a sliding converter network to be trained to obtain a real part prediction, an imaginary part prediction, and an amplitude spectrum, wherein the real part prediction and the imaginary part prediction constitute a prediction spectrum; calculating a ratio mask between the predicted spectrum and the noise spectrum, and enhancing the noise spectrum according to the ratio mask to obtain a fused spectrum; Based on the loss function, performing an inverse short-time Fourier transform on the fused spectrum to obtain a denoised signal, and calculating a first loss value of the denoised signal on a time-domain reconstruction error based on a first loss function; Calculating a second loss value of the denoised signal in signal reconstruction quality based on a second loss function; Calculating a third loss value of the amplitude spectrum based on a third loss function; Calculating a fourth loss value between the predicted spectrum and the true spectrum based on the discriminator; fusing the first loss value, the second loss value, the third loss value, and the fourth loss value to obtain a target loss value; Update the generator parameters and the discriminator parameters based on the target loss value to obtain the trained sliding converter network; Input the real part and the imaginary part of the complex spectrum matrix of the noise audio collected by the microphone into the sliding converter network to obtain the predicted real part and the predicted imaginary part; Calculating a ratio mask between the complex spectrum and the complex spectrum matrix based on the complex spectrum consisting of the predicted real part and the predicted imaginary part; enhancing the complex spectrum matrix based on the ratio mask to obtain a target complex spectrum; An inverse short-time Fourier transform is performed on the target complex spectrum to obtain a noise suppression frequency.
2. The headphone-based noise processing method according to claim 1, wherein: The step of fusing the first loss value, the second loss value, the third loss value, and the fourth loss value to obtain a target loss value includes: determining a norm of the first loss value, a complement and a first weight of the second loss value, a second weight of the third loss value, and a third weight of the fourth loss value; Accumulate the norm, the product of the complement and the first weight, the product of the third loss value and the second weight, and the product of the fourth loss value and the third weight to obtain the target loss value.
3. The headphone-based noise processing method according to claim 1, wherein: The step of enhancing the complex spectrum matrix based on the ratio mask to obtain a target complex spectrum comprises: The product of the ratio mask and the complex spectrum matrix is calculated, and the product is set as the target complex spectrum.
4. The headphone-based noise processing method according to claim 1, wherein: The step of inputting the real part and the imaginary part of the complex spectrum matrix of the noise audio collected by the microphone into the sliding converter network to obtain the predicted real part and the predicted imaginary part includes: Performing windowing and short-time Fourier transform on the noise audio to obtain the complex spectrum matrix; Inputting the real and imaginary parts of the complex spectrum matrix of the noisy audio into a sliding converter network, and extracting and correlating local structural features of the complex spectrum matrix based on the complex convolutional layer of the multi-scale encoder of the sliding converter network; The cross-layer transfer features are obtained based on the decoder, and after upsampling processing is performed on the cross-layer transfer features, the predicted real part and the predicted imaginary part are output based on the output layer.
5. The headphone-based noise processing method according to claim 1, wherein: The microphone is a dual-channel microphone. Before the step of inputting the real part and the imaginary part of the complex spectrum matrix of the noise audio collected by the microphone into the sliding converter network to obtain the predicted real part and the predicted imaginary part, the headphone-based noise processing method further includes: Acquire dual-channel audio captured by the dual-channel microphone; An energy difference between the two-channel audio is calculated, and when the energy difference is greater than a preset threshold and the duration is greater than a preset duration, the two-channel audio is determined to be the noise audio.
6. The headphone-based noise processing method according to claim 1, wherein: After the step of inputting the real part and the imaginary part of the complex spectrum matrix of the noise audio collected by the microphone into the sliding converter network to obtain the predicted real part and the predicted imaginary part, the headphone-based noise processing method further includes: Inputting the complex spectrum into a complex phase refinement model to obtain the complex spectrum with updated phase; Based on the complex spectrum after phase update, the step of calculating a ratio mask between the complex spectrum and the complex frequency spectrum matrix based on the complex spectrum composed of the predicted real part and the predicted imaginary part is performed.
7. A headset, characterized in that: The headset comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the headset-based noise processing method according to any one of claims 1 to 6.
8. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the headphone-based noise processing method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Scene perception method, scene perception system, electronic equipment and storage medium
CN117315275A
Target noise extraction and evaluation method and device
CN118197353A
Cited By
A multi-microphone intelligent sports earphone noise reduction system and a noise reduction method thereof
CN122602027A