Voice processing method and device based on noise reduction and reverberation elimination

Through the multi-stage model processing of voice signals, the voice signals are respectively reduced, reverb removal and voiceprint repair, which solves the problem of difficulty in processing noise and reverb simultaneously in the prior art, and achieves high-quality restoration of voice signals.

CN120375847APending Publication Date: 2025-07-25YEALINK (XIAMEN) NETWORK TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510447951.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The prior art is difficult to effectively remove noise and reverb simultaneously, resulting in poor voiceprint reduction of voice signals and insignificant reverb removal.

Method used

The multi-stage model is used to process the voice signal. First, the complex spectrum of the voice signal is reduced by the audio noise reduction model, then the amplitude spectrum and reverb clear model are used to remove the reverb, and finally the low-frequency part is repaired through the voiceprint repair model.

Benefits of technology

The processed voice signal has no obvious fluctuations, the voice listening feeling is improved, and the voiceprint restoration degree is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375847A_ABST
    Figure CN120375847A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of audio processing, and discloses a voice processing method and device based on noise reduction and reverberation elimination, and the method comprises the steps: obtaining a voice signal, and carrying out the short-time frame division, and obtaining a plurality of short-time frame signals; changing the short-time frame signal to obtain a first complex spectrum; inputting the first complex spectrum into an audio noise reduction model to obtain a first complex spectrum mask; obtaining a noise reduction audio signal according to the first complex spectrum and the first complex spectrum mask; inputting the amplitude spectrum of the noise-reduced audio signal into a reverberation removal model to obtain an amplitude spectrum mask; obtaining a second complex spectrum according to the amplitude spectrum and the amplitude spectrum mask; inputting the first complex number spectrum and the second complex number spectrum into a voiceprint repairing model to obtain a third complex number spectrum; and obtaining a target voice signal after noise reduction and reverberation elimination according to the third complex spectrum. According to the invention, the processed target voice signal has no obvious fluctuation, the hearing sense of the voice is obviously improved, and the reduction degree of the voiceprint becomes good.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio processing technologies, and in particular, to a voice processing method and device based on noise reduction and reverberation removal. Background Art

[0002] In the existing technical solutions, the enhancement of voice signals aims to suppress noise and remove reverberation from the input voice. This technology performs denoising and dereverberation simultaneously through a single neural network. The method of this technology is the mapping from the noisy and reverberant audio spectrum to the clean audio spectrum, and it has been widely applied in the field of voice enhancement.

[0003] However, although the existing technical solutions can achieve obvious denoising effects, due to the different natures of noise and reverberation, background noise is an additive signal to the clean voice, while reverberation is a convolution process of the clean voice and the room impulse response, making it difficult for a single neural network to solve them together, resulting in poor restoration of the voiceprint of the reverberant audio and insignificant removal of the reverberation feeling. Summary of the Invention

[0004] This application provides a voice processing method and device based on noise reduction and reverberation removal, which can make the processed target voice signal have no obvious fluctuations, significantly improve the listening experience of the voice, and better restore the voiceprint.

[0005] In a first aspect, an embodiment of this application provides a voice processing method based on noise reduction and reverberation removal, including:

[0006] Obtain a voice signal and perform short-time frame division to obtain a plurality of short-time frame signals;

[0007] Perform a transformation on the short-time frame signal to obtain a first complex spectrum;

[0008] Input the first complex spectrum into an audio noise reduction model to obtain a first complex spectrum mask;

[0009] Obtain a noise-reduced audio signal according to the first complex spectrum and the first complex spectrum mask;

[0010] Input the amplitude spectrum of the noise-reduced audio signal into a reverberation removal model to obtain an amplitude spectrum mask;

[0011] Obtain a second complex spectrum according to the amplitude spectrum and the amplitude spectrum mask;

[0012] Input the first complex spectrum and the second complex spectrum into a voiceprint restoration model to obtain a third complex spectrum;

[0013] Obtain a target voice signal after noise reduction and reverberation removal according to the third complex spectrum.

[0014] Further, the above-mentioned performing a transformation on the short-time frame signal to obtain a first complex spectrum includes:

[0015] Perform a short-time Fourier transform on the short-time frame signal to obtain a first complex spectrum.

[0016] Further, obtaining the noise-reduced audio signal according to the first complex spectrum and the first complex spectrum mask includes:

[0017] Multiply the first complex spectrum by the first complex spectrum mask to obtain a target complex spectrum;

[0018] Perform an inverse short-time Fourier transform on the target complex spectrum to obtain a noise-reduced audio signal.

[0019] Further, obtaining the second complex spectrum according to the amplitude spectrum and the amplitude spectrum mask includes:

[0020] Multiply the amplitude spectrum by the amplitude spectrum mask to obtain a target amplitude spectrum;

[0021] Convert the target amplitude spectrum to obtain a second complex spectrum.

[0022] Further, the method further includes:

[0023] After obtaining the speech signal, remove the DC offset of the speech signal.

[0024] Further, the number of points for the short-time Fourier transform is 512.

[0025] Further, the audio noise reduction model includes an encoder and a decoder with the same dimensions;

[0026] The encoder includes a real part depth convolution, a real part point convolution, an imaginary part depth convolution, and an imaginary part point convolution;

[0027] The decoder includes a real part depth transposed convolution, a real part point transposed convolution, an imaginary part depth transposed convolution, and an imaginary part point transposed convolution.

[0028] In a second aspect, an embodiment of the present application provides a speech processing device based on noise reduction and reverberation cancellation, including:

[0029] An acquisition module, configured to acquire a speech signal and perform short-time frame division to obtain a plurality of short-time frame signals;

[0030] A decomposition module, configured to perform a transform on the short-time frame signal to obtain a first complex spectrum;

[0031] A noise reduction module, configured to input the first complex spectrum into an audio noise reduction model to obtain a first complex spectrum mask;

[0032] A first operation module, configured to obtain a noise-reduced audio signal according to the first complex spectrum and the first complex spectrum mask;

[0033] A reverberation processing module, configured to input the amplitude spectrum of the noise-reduced audio signal into a reverberation cancellation model to obtain an amplitude spectrum mask;

[0034] A second operation module, configured to obtain a second complex spectrum according to the amplitude spectrum and the amplitude spectrum mask;

[0035] A repair module, configured to input the first complex spectrum and the second complex spectrum into a voiceprint repair model to obtain a third complex spectrum;

[0036] A third operation module, configured to obtain a target voice signal after noise reduction and reverberation removal according to the third complex spectrum.

[0037] In a third aspect, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it executes the steps of a voice processing method based on noise reduction and reverberation removal according to any one of the above embodiments.

[0038] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, storing a computer program, and when the computer program is executed by a processor, it implements the steps of a voice processing method based on noise reduction and reverberation removal according to any one of the above embodiments.

[0039] In summary, compared with the prior art, the beneficial effects brought by the technical solution provided by the embodiment of the present application at least include:

[0040] A voice processing method based on noise reduction and reverberation removal provided by an embodiment of the present application processes a voice signal by using a multi-stage model. First, in the first stage, an audio noise reduction model is used to perform noise reduction processing on the complex spectrum of the voice signal. Then, in the second stage, reverberation is removed from the noise-reduced audio signal based on the amplitude spectrum and the reverberation removal model. Finally, in the third stage, the voiceprint of the low-frequency part in the voice signal is repaired based on the complex spectrum and the voiceprint repair model. The present application splits the simultaneous noise reduction and reverberation removal and executes the noise reduction and reverberation removal tasks in multiple stages. In the first stage, the audio noise reduction model will not deteriorate compared with the model that simultaneously reduces noise and removes reverberation; in the second stage, the effect of the reverberation removal model is better than that of the model that simultaneously reduces noise and removes reverberation, so that the processed target voice signal has no obvious fluctuations, the listening experience of the voice is significantly improved, and the restoration degree of the voiceprint becomes better. Description of the Drawings

[0041] Figure 1 It is a flowchart of a voice processing method based on noise reduction and reverberation removal provided by an embodiment of the present application.

[0042] Figure 2 It is an architecture diagram of three models used in the present application provided by an embodiment of the present application.

[0043] Figure 3Structural diagram of a voice processing device based on noise reduction and reverberation cancellation provided by an embodiment of the present application. Detailed implementation manners

[0044] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments.

[0045] Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0046] Please refer to Figure 1 , the embodiments of the present application provide a voice processing method based on noise reduction and reverberation cancellation, including:

[0047] Step S1, obtaining a voice signal and performing short-time frame division to obtain a plurality of short-time frame signals.

[0048] Specifically, after obtaining the voice signal, some preprocessing operations are first performed on the voice signal, such as removing the DC offset of the voice signal, adjusting the sampling rate, audio normalization, etc., and then the voice signal is divided into short-time frames for subsequent processing.

[0049] Step S2, performing a transformation on the short-time frame signal to obtain a first complex spectrum.

[0050] Specifically, the first complex spectrum can be obtained by performing a short-time Fourier transform (STFT, Short-Time Fourier Transform) on the short-time frame signal. The complex spectrum of the voice signal includes the amplitude spectrum and phase spectrum in its frequency domain, where the size of the non-DC part is FFT_n / 2, and FFT_n represents the number of Fourier transform points of the STFT. In the present application, 512 is preferably used as this value.

[0051] Step S3, inputting the first complex spectrum into an audio noise reduction model to obtain a first complex spectrum mask.

[0052] It can be understood that since the complex spectrum can process both amplitude and phase information simultaneously for more accurate signal reconstruction; in addition, in a low signal-to-noise ratio scenario, the complex spectrum can better retain the characteristics of the original signal, thereby achieving effective noise reduction.

[0053] Step S4, obtaining a noise-reduced audio signal according to the first complex spectrum and the first complex spectrum mask.

[0054] Specifically, multiply the first complex spectrum by the first complex spectrum mask to obtain the target complex spectrum; perform the inverse short-time Fourier transform (ISTFT) on the target complex spectrum to obtain the denoised audio signal.

[0055] Step S5: Input the amplitude spectrum of the denoised audio signal into the reverberation removal model to obtain the amplitude spectrum mask.

[0056] Specifically, in the stage of removing reverberation, this application uses the amplitude spectrum as the input of the model for full-band reverberation removal. Compared with the complex spectrum, the amplitude spectrum does not require excessive complex operations and parameter adjustments; and the computing power required by the amplitude spectrum is also smaller, having better real-time performance in the implementation of the algorithm.

[0057] Step S6: Obtain the second complex spectrum according to the amplitude spectrum and the amplitude spectrum mask.

[0058] Specifically, multiply the amplitude spectrum by the amplitude spectrum mask to obtain the target amplitude spectrum; the target amplitude spectrum at this time is the amplitude spectrum after reverberation removal, and then convert the target amplitude spectrum to obtain the second complex spectrum.

[0059] Specifically, please refer to Figure 2 , both the audio denoising model and the reverberation removal model of this application adopt the U-Net structure. When constructing the audio denoising model, both its encoder and decoder are changed to the complex spectrum form. The encoder includes real-part depth convolution, real-part point convolution, imaginary-part depth convolution, and imaginary-part point convolution; the decoder includes real-part depth deconvolution, real-part point deconvolution, imaginary-part depth deconvolution, and imaginary-part point deconvolution. It should be noted that the dimensions of the encoder and decoder are the same to ensure that the dimensions of the input and output can be aligned, and the output can be restored to the normal audio signal after ISTFT.

[0060] When constructing the reverberation removal model, improve the Encoder, RNN (Recurrent Neural Network), and Decoder through the amplitude operation method, and change them to the amplitude-form Encoder, RNN, and Decoder.

[0061] Step S7: Input the first complex spectrum and the second complex spectrum into the voiceprint restoration model to obtain the third complex spectrum.

[0062] The reason for introducing the model input in the denoising stage in this application is to repair the result obtained in the reverberation removal stage, that is, to enable the voiceprint restoration model to obtain the information of the audio before damage through the first complex spectrum, so as to repair the signal after denoising and reverberation removal.

[0063] In the voiceprint restoration stage of this application, the voiceprint restoration model adopted is the complex spectrum U-Net network used in the noise reduction stage. However, the output complex spectrum mask needs to be changed to the complex spectrum directly mapped by the model.

[0064] Step S8: Obtain the target voice signal after noise reduction and reverberation removal based on the third complex spectrum.

[0065] Specifically, after obtaining the target voice signal, some post-processing steps may be required, such as removing artifacts caused by the window function, volume adjustment, and smoothing processing, etc., before output.

[0066] In the above-mentioned embodiment, a voice processing method based on noise reduction and reverberation removal is provided. The voice signal is processed by a multi-stage model. First, in the first stage, an audio noise reduction model is used to perform noise reduction processing on the complex spectrum of the voice signal. Then, in the second stage, based on the amplitude spectrum and the reverberation removal model, the reverberation of the noise-reduced audio signal is removed. Finally, in the third stage, based on the complex spectrum and the voiceprint restoration model, the voiceprint of the low-frequency part in the voice signal is restored. In the above-mentioned embodiment, the simultaneous noise reduction and reverberation removal are split, and the noise reduction and reverberation removal tasks are executed in multiple stages. In the first stage, the audio noise reduction model will not deteriorate compared with the model that performs simultaneous noise reduction and reverberation removal. In the second stage, the effect of the reverberation removal model is better than that of the model that performs simultaneous noise reduction and reverberation removal, making the processed target voice signal have no obvious fluctuations, and the listening experience of the voice is significantly improved, and the restoration degree of the voiceprint becomes better.

[0067] Furthermore, for the product, separating noise reduction and reverberation removal allows external control over whether to perform the reverberation removal operation. In a low-reverberation scenario, turning off the reverberation removal function can obtain high-fidelity audio, avoiding audio damage and sound quality deterioration caused by reverberation removal, thereby increasing the competitiveness of the product. In addition, compared with existing open-source models, the model of this application is a causal model, which can perform real-time inference and can be deployed in conference communication devices to improve the voice quality in actual communication.

[0068] Please refer to Figure 3 , another embodiment of this application provides a voice processing device based on noise reduction and reverberation removal, including:

[0069] An acquisition module 101, configured to acquire a voice signal and perform short-time frame division to obtain a plurality of short-time frame signals.

[0070] A decomposition module 102, configured to perform transformation on the short-time frame signal to obtain a first complex spectrum.

[0071] A noise reduction module 103, configured to input the first complex spectrum into an audio noise reduction model to obtain a first complex spectrum mask.

[0072] The first operation module 104 is configured to obtain a noise-reduced audio signal according to the first complex spectrum and the first complex spectrum mask.

[0073] The reverberation processing module 105 is configured to input the amplitude spectrum of the noise-reduced audio signal into a reverberation removal model to obtain an amplitude spectrum mask.

[0074] The second operation module 106 is configured to obtain a second complex spectrum according to the amplitude spectrum and the amplitude spectrum mask.

[0075] The repair module 107 is configured to input the first complex spectrum and the second complex spectrum into a voiceprint repair model to obtain a third complex spectrum.

[0076] The third operation module 108 is configured to obtain a target speech signal after noise reduction and reverberation removal according to the third complex spectrum.

[0077] For the specific limitations of the voice processing device based on noise reduction and reverberation removal provided in this embodiment, reference may be made to the embodiment of the voice processing method based on noise reduction and reverberation removal in the foregoing text, which will not be elaborated herein. Each module in the above-mentioned voice processing device based on noise reduction and reverberation removal can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to the above-mentioned modules.

[0078] The embodiment of the present application provides a computer device, which may include a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the processor is caused to execute the steps of a voice processing method based on noise reduction and reverberation removal as described in any of the above embodiments.

[0079] For the working process, working details, and technical effects of the computer device provided in this embodiment, reference may be made to the embodiment of the voice processing method based on noise reduction and reverberation removal in the foregoing text, which will not be elaborated herein.

[0080] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of a voice processing method based on noise reduction and reverberation cancellation as described in any of the above embodiments are implemented. Among them, the computer-readable storage medium refers to a carrier for storing data, which may include, but is not limited to, floppy disks, optical discs, hard disks, flash memories, USB flash drives, and / or memory sticks, etc. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. For the working process, working details, and technical effects of the computer-readable storage medium provided in this embodiment, reference may be made to the embodiments of the voice processing method based on noise reduction and reverberation cancellation in the above text, which will not be elaborated here.

[0081] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it may include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application may include non-volatile and / or volatile memories. Non-volatile memories may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM).

[0082] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0083] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. A voice processing method based on noise reduction and reverberation removal, characterized in that, including: obtaining a voice signal and performing short-time frame division to obtain a plurality of short-time frame signals; performing a transformation on the short-time frame signals to obtain a first complex spectrum; inputting the first complex spectrum into an audio noise reduction model to obtain a first complex spectrum mask; obtaining a noise-reduced audio signal according to the first complex spectrum and the first complex spectrum mask; inputting the amplitude spectrum of the noise-reduced audio signal into a reverberation cancellation model to obtain an amplitude spectrum mask; obtaining a second complex spectrum according to the amplitude spectrum and the amplitude spectrum mask; inputting the first complex spectrum and the second complex spectrum into a voiceprint restoration model to obtain a third complex spectrum; obtaining a target voice signal after noise reduction and reverberation cancellation according to the third complex spectrum.

2. The voice processing method based on noise reduction and reverberation cancellation according to claim 1, wherein The performing a transformation on the short-time frame signals to obtain a first complex spectrum includes: performing a short-time Fourier transform on the short-time frame signals to obtain the first complex spectrum.

3. The voice processing method based on noise reduction and reverberation cancellation according to claim 2, wherein The obtaining a noise-reduced audio signal according to the first complex spectrum and the first complex spectrum mask includes: multiplying the first complex spectrum and the first complex spectrum mask to obtain a target complex spectrum; performing an inverse short-time Fourier transform on the target complex spectrum to obtain the noise-reduced audio signal.

4. The voice processing method based on noise reduction and reverberation cancellation according to claim 2, wherein The obtaining a second complex spectrum according to the amplitude spectrum and the amplitude spectrum mask includes: multiplying the amplitude spectrum and the amplitude spectrum mask to obtain a target amplitude spectrum; performing a conversion on the target amplitude spectrum to obtain the second complex spectrum.

5. The voice processing method based on noise reduction and reverberation cancellation according to claim 1, characterized in that, It further includes: after obtaining the voice signal, removing the DC offset of the voice signal.

6. The voice processing method based on noise reduction and reverberation cancellation according to claim 2, characterized in that, The number of points of the short-time Fourier transform is 512.

7. The voice processing method based on noise reduction and reverberation cancellation according to claim 1, characterized in that, The audio noise reduction model includes an encoder and a decoder with the same dimension; The encoder includes a real part depth convolution, a real part point convolution, an imaginary part depth convolution, and an imaginary part point convolution; The decoder includes a real part depth transpose convolution, a real part point transpose convolution, an imaginary part depth transpose convolution, and an imaginary part point transpose convolution.

8. A voice processing device based on noise reduction and reverberation cancellation, characterized in that, including: an obtaining module, configured to obtain a voice signal and perform short-time frame division to obtain a plurality of short-time frame signals; a decomposition module, configured to perform a transformation on the short-time frame signals to obtain a first complex spectrum; a noise reduction module, configured to input the first complex spectrum into an audio noise reduction model to obtain a first complex spectrum mask; a first operation module, configured to obtain a noise-reduced audio signal according to the first complex spectrum and the first complex spectrum mask; a reverberation processing module, configured to input the amplitude spectrum of the noise-reduced audio signal into a reverberation cancellation model to obtain an amplitude spectrum mask; a second operation module, configured to obtain a second complex spectrum according to the amplitude spectrum and the amplitude spectrum mask; a restoration module, configured to input the first complex spectrum and the second complex spectrum into a voiceprint restoration model to obtain a third complex spectrum; a third operation module, configured to obtain a target voice signal after noise reduction and reverberation cancellation according to the third complex spectrum.

9. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the voice processing method based on noise reduction and reverberation cancellation according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the voice processing method based on noise reduction and reverberation cancellation according to any one of claims 1 to 7.