Voice processing method, storage medium and electronic equipment
By identifying the time domain and frequency domain distribution information of the voice signal to determine the noise signal, and performing noise reduction processing driven by the target recognition result, the problem of noise reduction in the prior art damage to the voice signal and the poor effect of eliminating non-stationary noise is solved, and more accurate and effective voice signal noise reduction is achieved.
Patent Information
- Application Number
- CN202311607427.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-27
- Publication Date
- 2025-05-27
AI Technical Summary
The existing speech processing methods are prone to damage the original speech signal during the noise reduction process, and the non-stationary noise is not effective.
By acquiring the time domain and frequency domain distribution information of the original voice signal, it is recognized to determine whether the sub-signal is a noise signal, and then the noise reduction process driven by the target recognition result is performed.
It effectively avoids damage to the original voice signal during the noise reduction process, improves the effect of canceling non-stationary noise, and makes the voice signal after noise reduction closer to the pure voice signal.
Smart Images

Figure CN120048278A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of signal processing, and more particularly, to a voice processing method, a storage medium, and an electronic device. Background Art
[0002] With the rise of interactive live streaming and video conferencing in recent years, the scenarios of real-time communication have become increasingly complex. The voice signals that users can receive during communication usually contain a certain degree of noise signals, such as signals corresponding to air conditioner sounds, signals corresponding to keyboard tapping sounds, signals corresponding to people's coughing sounds, etc. These signals are very likely to affect the user's understanding and perception of the voice signal when receiving the voice signal. Currently, when processing voice signals, simple single-channel voice enhancement technology is usually used. After analyzing the voice signal in the data frequency domain, the analyzed noise signal is denoised. However, this method has great limitations. Most of the noise signals analyzed from the voice signal are incomplete, resulting in a large degree of damage to the original voice signal after denoising, or incomplete elimination of some special non-stationary noises.
[0003] In view of the above problems, no effective solution has been proposed yet. Summary of the Invention
[0004] Embodiments of the present application provide a voice processing method, a storage medium, and an electronic device to at least solve the technical problem that the voice processing method in the related art has limitations, resulting in a large degree of damage to the original voice signal after denoising, or incomplete elimination of some special non-stationary noises.
[0005] According to one aspect of the embodiments of the present application, a voice processing method is provided, including: obtaining an original voice signal; identifying the original voice signal based on the time-domain distribution information of the original voice signal in the time domain and the frequency-domain distribution information of the original voice signal in the frequency domain to obtain a target recognition result corresponding to the original voice signal, where the target recognition result is used to represent whether a sub-signal in the original voice signal is a noise signal; and performing denoising processing on the original voice signal based on the target recognition result to obtain a target voice signal.
[0006] According to another aspect of the embodiments of the present application, a voice processing method is further provided, including: obtaining an original voice signal in a live video; identifying the original voice signal based on the time-domain distribution information of the original voice signal in the time domain and the frequency-domain distribution information of the original voice signal in the frequency domain to obtain a target recognition result corresponding to the original voice signal, where the target recognition result is used to represent whether a sub-signal in the original voice signal is a noise signal; performing denoising processing on the original voice signal based on the target recognition result to obtain a target voice signal; and replacing the original voice signal in the live video with the target voice signal.
[0007] According to another aspect of the embodiments of the present application, a voice processing method is further provided, including: in response to an input instruction acting on an operation interface, displaying a waveform diagram of an original voice signal on the operation interface; in response to a noise reduction instruction acting on the operation interface, displaying a waveform diagram of a target voice signal on the operation interface, where the target voice signal is obtained by performing noise reduction processing on the original voice signal based on a target recognition result corresponding to the original voice signal, the target recognition result is obtained by recognizing the original voice signal based on the time-domain distribution information of the original voice signal in the time domain and the frequency-domain distribution information of the original voice signal in the frequency domain, and the target recognition result is used to characterize whether a sub-signal in the original voice signal is a noise signal.
[0008] According to another aspect of the embodiments of the present application, a voice processing method is further provided, including: obtaining an original voice signal by calling a first interface, where the first interface includes a first parameter, and the parameter value of the first parameter includes the original voice signal; recognizing the original voice signal based on the time-domain distribution information of the original voice signal in the time domain and the frequency-domain distribution information of the original voice signal in the frequency domain to obtain a target recognition result corresponding to the original voice signal, where the target recognition result is used to characterize whether a sub-signal in the original voice signal is a noise signal; performing noise reduction processing on the original voice signal based on the target recognition result to obtain a target voice signal; and outputting the target voice signal by calling a second interface, where the second interface includes a second parameter, and the parameter value of the second parameter includes the target voice signal.
[0009] According to another aspect of the embodiments of the present application, a computer-readable storage medium is further provided. The computer-readable storage medium includes a stored executable program. When the executable program runs, it controls the device where the computer-readable storage medium is located to execute the method of any one of the above.
[0010] According to another aspect of the embodiments of the present application, an electronic device is further provided, including: a memory storing an executable program; and a processor for running the program. When the program runs, it executes the method of any one of the above.
[0011] In an embodiment of the present application, the method includes: obtaining an original speech signal; based on the time-domain distribution information of the original speech signal in the time domain and the frequency-domain distribution information of the original speech signal in the frequency domain, identifying the original speech signal to obtain a target recognition result corresponding to the original speech signal, where the target recognition result is used to characterize whether a sub-signal in the original speech signal is a noise signal; and based on the target recognition result, performing noise reduction processing on the original speech signal to obtain a target speech signal. By using the time-domain distribution information and frequency-domain distribution information of the original speech signal in the time domain and frequency domain to reflect the distribution characteristics of the original speech signal in the time domain and frequency domain, and at the same time combining the time-domain distribution information and frequency-domain distribution information to determine the target recognition result of the original speech signal, the judgment of whether a sub-signal in the original speech signal is a noise signal is made more accurate. Finally, based on the target recognition result, noise reduction processing is performed on the original speech signal, which can avoid damaging the original speech signal during the noise reduction process, ensure that the target speech signal after noise reduction is closer to the pure speech signal, thereby ensuring the accuracy of the obtained target speech signal, improving the effect of noise reduction on the original speech signal, and further solving the technical problems in the related art that the speech processing method has limitations, resulting in a large degree of damage to the original speech signal after noise reduction or incomplete elimination of some special non-stationary noises.
[0012] It is easy to note that the above general description and the following detailed description are only for exemplifying and explaining the present application, and do not constitute a limitation to the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation to the present application. In the drawings:
[0014] Figure 1 is a hardware structure block diagram of a computer terminal (or mobile device) for implementing a speech processing method according to Embodiment 1 of the present application;
[0015] Figure 2 is a schematic diagram of a computer terminal communication method according to Embodiment 1 of the present application;
[0016] Figure 3 is a flowchart of a speech processing method according to Embodiment 1 of the present application
[0017] Figure 4 is a schematic diagram of a convolutional layer and a deconvolutional layer according to Embodiment 1 of the present application;
[0018] Figure 5 is a schematic diagram of a speech processing process according to Embodiment 1 of the present application;
[0019] Figure 6 It is a schematic diagram of a gated convolutional recurrent network according to Embodiment 1 of the present application;
[0020] Figure 7 It is a schematic diagram of a traditional speech processing method according to Embodiment 1 of the present application;
[0021] Figure 8 It is a schematic diagram of the noise reduction effect of a traditional speech processing according to Embodiment 1 of the present application;
[0022] Figure 9 It is a schematic diagram of a traditional gated convolutional recurrent network according to Embodiment 1 of the present application;
[0023] Figure 10 It is a schematic diagram of another traditional speech noise reduction effect according to Embodiment 1 of the present application;
[0024] Figure 11 It is a schematic diagram of a new speech noise reduction effect according to Embodiment 1 of the present application;
[0025] Figure 12 It is a flowchart of a speech processing method according to Embodiment 2 of the present application;
[0026] Figure 13 It is a flowchart of a speech processing method according to Embodiment 3 of the present application;
[0027] Figure 14 It is a flowchart of a speech processing method according to Embodiment 4 of the present application;
[0028] Figure 15 It is a structural block diagram of a speech processing device according to Embodiment 5 of the present application;
[0029] Figure 16 It is a structural block diagram of a speech processing device according to Embodiment 6 of the present application;
[0030] Figure 17 It is a structural block diagram of a speech processing device according to Embodiment 7 of the present application;
[0031] Figure 18 It is a structural block diagram of a speech processing device according to Embodiment 8 of the present application;
[0032] Figure 19 It is a structural block diagram of a computer terminal according to Embodiment 9 of the present application. Detailed implementation manners
[0033] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.
[0034] It should be noted that the terms "first", "second", etc. in the specification, claims and above-mentioned drawings of this application are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily need to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0035] First, some nouns or terms that appear in the process of describing the embodiments of this application are applicable to the following explanations:
[0036] Time-Frequency Domain: (Time-Frequency Domain) is a method or representation for analyzing a signal simultaneously in the time domain and the frequency domain. It can provide information about the changes and distributions of the signal in time and frequency. Time-frequency domain analysis aims to overcome the complementarity between time-domain and frequency-domain analysis.
[0037] Single-channel speech enhancement: refers to a technology that processes the signal of one channel to improve its quality and audibility, aiming to reduce noise, reverberation or other interferences in the speech signal to improve the clarity and intelligibility of the speech.
[0038] Convolutional neural network: a deep learning network model that uses convolutional layers and pooling layers to extract features of input data and performs tasks such as classification or regression through fully connected layers.
[0039] Gated Recurrent Unit: a variant of the Recurrent Neural Network (RNN) used to process sequential data. GRU introduces a gating mechanism on the basis of the recurrent neural network to solve the problem of long-term dependence. It consists of two main parts: the reset gate and the update gate.
[0040] Embodiment 1
[0041] According to an embodiment of the present application, an embodiment of a method for voice processing is further provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0042] The method embodiment provided by Embodiment 1 of the present application can be executed in a mobile terminal, a computer terminal, or a similar computing device. Figure 1 It is a hardware structure block diagram of a computer terminal (or mobile device) for implementing a voice processing method according to Embodiment 1 of the present application. As Figure 1 shown, the computer terminal 10 (or mobile device) may include one or more processors 102 (shown as 102a, 102b,..., 102n in the figure) (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a Universal Serial Bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 may further include more or fewer components than Figure 1 shown, or have a different configuration from Figure 1 shown.
[0043] It should be noted that the above one or more processors 102 and / or other data processing circuits can generally be referred to as "data processing circuits" in this article. The data processing circuit can be embodied as software, hardware, firmware, or any combination thereof, in whole or in part. In addition, the data processing circuit can be a single independent processing module, or be incorporated in whole or in part into any one of the other elements in the computer terminal 10 (or mobile device). As involved in the embodiments of the present application, the data processing circuit is a processor control (such as the selection of a variable resistance terminal path connected to an interface).
[0044] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the voice processing method in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implements the above-mentioned voice processing method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the computer terminal 10 through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0045] The transmission device 106 is used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by a communication provider of the computer terminal 10. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0046] The display may be, for example, a touch-screen liquid crystal display (Liquid Crystal Display, LCD), and the liquid crystal display enables a user to interact with the user interface of the computer terminal 10 (or mobile device).
[0047] Figure 1 The shown hardware structure block diagram can be used not only as an exemplary block diagram of the above computer terminal 10 (or mobile device), but also as an exemplary block diagram of the above server. In an alternative embodiment, Figure 2 is a schematic diagram of a computer terminal communication method according to Embodiment 1 of the present application, Figure 2 which shows, in block diagram form, an embodiment using the above-mentioned Figure 1 shown computer terminal 10 (or mobile device) as [such as a sending end, a receiving end, etc.]. As Figure 2As shown, the computer terminal 10 (or mobile device) can act as a server and be connected to one or more clients [such as security clients, resource clients, game clients, etc.] via a data network connection or an electronic connection. In an alternative embodiment, the above computer terminal 10 (or mobile device) can be [any data server, etc.]. The data network connection can be a local area network connection, a wide area network connection, an Internet connection, or other types of data network connections. The computer terminal 10 (or mobile device) can execute to connect to network services performed by a client (such as a security client) or a group of clients 20. A network client is a network-based user client, such as a social network, cloud resources, email, online payment, or other online applications.
[0048] Under the above operating environment, the present application provides a voice processing method as shown in Figure 3 Figure 5. Figure 3 It is a flowchart of a voice processing method according to Embodiment 1 of the present application. As shown in Figure 3 Figure 6, the method may include the following steps:
[0049] Step S302, obtain an original voice signal.
[0050] The above original voice signal may refer to a single-channel noisy voice, that is, the original voice signal contains noise. Therefore, noise reduction processing is required.
[0051] Currently, the noise existing in the voice signal is mainly divided into two types. One is stationary noise, such as the air conditioner sound and fan sound in the environment, and the other is non-stationary noise, such as the keyboard tapping sound and coughing sound in the environment. For a voice signal containing stationary noise, a traditional noise reduction algorithm can be used to achieve a good noise reduction effect. However, for a voice signal containing non-stationary noise, using a traditional noise reduction algorithm may identify parts of the voice signal that are not noise as noise, and then damage the voice signal or incompletely reduce the noise of the voice signal when reducing the noise of the voice signal. Considering that there are many types of stationary noise and non-stationary noise, stationary noise and non-stationary noise often exist simultaneously in a voice signal. At the same time, considering that multi-channel voice can be divided into multiple single-channel voices, when the voice processing system obtains the original voice signal that needs to be subjected to noise reduction processing, the obtained voice signal can be a single-channel noisy voice signal.
[0052] It should be noted that the above manner of obtaining the original voice signal may include but is not limited to: real-time recording of the voice signal through a recording device, actively selecting the voice signal that needs to be noise-reduced through a voice library, automatically receiving the voice signal that needs to be noise-reduced through a preset port, etc. The specific obtaining manner can be selected by the user according to the actual situation and is not limited here.
[0053] Step S304: Based on the time-domain distribution information of the original speech signal in the time domain and the frequency-domain distribution information of the original speech signal in the frequency domain, identify the original speech signal to obtain the corresponding target recognition result of the original speech signal.
[0054] The target recognition result is used to characterize whether the sub-signal in the original speech signal is a noise signal.
[0055] The above target recognition result can refer to the mask information obtained through network prediction, which can be used to indicate whether to retain or remove the sub-signals included in the original speech signal. That is, if the sub-signal is a noise signal to be removed, the target recognition result can be used to indicate removing the sub-signal, which can be represented by 0; if the sub-signal is a target signal to be retained, the target recognition result can be used to indicate retaining the sub-signal, which can be represented by 1. The above sub-signals can refer to the signals divided from the original speech signal in a preset manner. For example, the original speech signal can be divided according to a preset duration, or evenly divided according to a preset quantity. Usually, the original speech signal contains noise signals to be removed and target signals to be retained. For example, the sound sources included in a piece of original speech signal can include, but are not limited to, air conditioners, fans, keyboards, people, etc. in the environment. The corresponding noise signals to be removed can include air conditioner sounds, fan sounds, keyboard sounds, people's coughing sounds, etc., and the target signals to be retained can include the speech uttered by people. However, since there is no obvious boundary between the noise signals and the target signals in the original speech signal, after obtaining the original speech signal that needs to be denoised, the speech processing system can first identify the sub-signals included in the original speech signal to determine whether the sub-signal is a noise signal to be removed.
[0056] Considering that in the original speech signal, the signal characteristics of the noise signal and the target signal in the time-frequency domain are usually different. Taking the target signal as the speech uttered by the user, the noise signal as the cough sound of the user or the air-conditioning sound emitted by the air conditioner as an example, the duration of the speech uttered by the user is usually greater than the duration of the cough sound of the user and less than the duration of the air-conditioning sound emitted by the air conditioner. The frequency of the speech uttered by the user in a short period is usually less than the frequency of the cough of the user in a short period and greater than the frequency of the air-conditioning sound emitted by the air conditioner in a short period. Therefore, when determining whether a sub-signal in the original speech signal is a noise signal, the speech processing system can obtain the signal characteristics of the original speech signal in the time-frequency domain, analyze the original speech signal to determine whether the corresponding sub-signal is a noise signal. Specifically, the speech processing system can first obtain the time-domain distribution information of the original speech signal in the time domain and the time-domain distribution information of the original speech signal in the frequency domain, and then use the time-domain distribution information and the frequency-domain distribution information to identify and analyze the sub-signals in the original speech signal, so as to obtain a target recognition result that can reflect whether the sub-signal is a noise signal.
[0057] It should be noted that since the same object may emit a noise signal and a target signal at the same time. For example, the user may cough while speaking, or the air conditioner may emit the corresponding speech of the operation instruction when it is running. Therefore, the speech processing system cannot determine whether the corresponding sub-signal is a noise signal based on the identity of the sound-emitting object, but needs to determine all the sub-signals in the original speech signal and analyze the sub-signals. For example, according to the signal characteristics of the sub-signal in the frequency domain and the signal characteristics in the time domain, to determine whether the sub-signal is a noise signal. By removing the sub-signals that are noise signals from the original speech signal, the operation of noise reduction on the original speech signal can be realized.
[0058] Step S306, perform noise reduction processing on the original speech signal based on the target recognition result to obtain a target speech signal.
[0059] The above-mentioned target speech signal may refer to the pure speech signal contained in the original speech signal.
[0060] After determining the target recognition result corresponding to the original speech signal, the speech processing system can, according to the target recognition result, remove the sub-signals belonging to the noise signal from the original speech signal and retain the sub-signals that do not belong to the noise signal to obtain the pure speech signal contained in the original speech signal, that is, the target speech signal.
[0061] In an alternative solution of this embodiment, to avoid the target speech signal determined from the original speech signal may still contain some noise signals, for example, the relatively severe breathing sounds that may accompany the user when speaking, the speech processing system can also re-obtain the time-domain distribution information of the target speech signal in the time domain and the frequency-domain distribution information in the frequency domain, and repeat the above process multiple times to improve the accuracy of the obtained target speech signal, so as to achieve the effect of single-channel speech enhancement for the original speech signal. It should be noted that the above-mentioned process of repeating multiple times can be the process of determining the target recognition result to improve the accuracy of the determined target recognition result, or the process of determining the target recognition result and denoising the original speech signal according to the target recognition result. The specific process of repeating multiple times can be set by the user himself and will not be limited here.
[0062] In the embodiment of the present application, the method includes: obtaining an original speech signal; identifying the original speech signal based on the time-domain distribution information of the original speech signal in the time domain and the frequency-domain distribution information of the original speech signal in the frequency domain to obtain a target recognition result corresponding to the original speech signal, where the target recognition result is used to characterize whether the sub-signal in the original speech signal is a noise signal; performing denoising processing on the original speech signal based on the target recognition result to obtain a target speech signal. By using the time-domain distribution information and frequency-domain distribution information of the original speech signal in the time domain and the frequency domain to reflect the distribution characteristics of the original speech signal in the time domain and the frequency domain, and combining the time-domain distribution information and the frequency-domain distribution information at the same time to determine the target recognition result of the original speech signal, the judgment of whether the sub-signal in the original speech signal is a noise signal is more accurate. Finally, denoising processing is performed on the original speech signal according to the target recognition result, which can avoid damaging the original speech signal during the denoising process, ensure that the denoised target speech signal is closer to the pure speech signal, thus ensuring the accuracy of the obtained target speech signal, improving the denoising effect on the original speech signal, and further solving the technical problems in the related art that the speech processing method has limitations, resulting in a large degree of damage to the original speech signal after denoising or incomplete elimination of some special non-stationary noises.
[0063] In the embodiment of the present application, identifying the original speech signal based on the time-domain distribution information of the original speech signal in the time domain and the frequency-domain distribution information of the original speech signal in the frequency domain to obtain a target recognition result corresponding to the original speech signal includes: using a speech processing model to identify the original speech signal based on the time-domain distribution information and the frequency-domain distribution information to obtain a target recognition result.
[0064] In an alternative solution of this embodiment, a pre-trained speech processing model can be used to process the above-mentioned time-domain distribution information and frequency-domain distribution information, identify sub-signals from the original speech signal, and the corresponding target recognition results of the sub-signals, so as to ensure the accuracy of the obtained target recognition results.
[0065] To ensure that the speech processing model has high precision and thus improve the accuracy of the determined target recognition results, a multi-layer convolutional neural network can be set in the speech processing model, so that the speech processing model can fully analyze the sub-signals in the original speech signal to ensure the accuracy of the obtained target recognition results. At the same time, when training the above-mentioned speech processing model, to improve the training effect, a clean speech signal can be selected as the true value, and perturbations can be added to the clean speech signal to obtain a speech signal containing noise signals. Then, the initial processing model is used to perform preliminary noise reduction processing on the speech signal containing noise signals, and the obtained signal is used as the predicted value. Finally, according to the difference between the predicted value and the true value, the initial processing model is adjusted. To ensure the accuracy of the adjusted processing model, this training process can be repeated multiple times until the number of repetitions reaches a preset number threshold, or the difference between the predicted value and the true value is less than a preset difference threshold, then the adjusted model can be determined as the above-mentioned speech processing model. In addition, the accuracy of the target speech signal obtained each time after using the speech processing model to perform noise reduction on the original speech signal can be recorded in real time. For example, it can be determined whether there is still a noise signal in the target speech signal. If so, the type, data volume, etc. of the noise signal can be obtained, and then the speech processing model can be adjusted in real time according to this information to further improve the accuracy of the speech processing model.
[0066] In the embodiment of the present application, the speech processing model includes: a frequency-domain feature extraction module, an encoder, a gated recurrent module, and a decoder. Using the speech processing model to identify the original speech signal based on the time-domain distribution information and frequency-domain distribution information, and obtain the target recognition result, including: using the frequency-domain feature extraction module to extract the frequency-domain distribution information of the original speech signal to obtain the original amplitude spectrum feature of the original speech signal; using the encoder to perform dimensionality reduction processing on the original amplitude spectrum feature to obtain the first feature corresponding to the original amplitude spectrum feature; using the gated recurrent module to process the first feature based on the time-domain distribution information to obtain the second feature; using the decoder to perform reconstruction processing on the second feature to obtain the target recognition result.
[0067] The above frequency-domain feature extraction module can be used to extract the features of the original speech signal in the frequency domain to improve the accuracy of recognizing the original speech signal. The above gated recurrent module can be used to capture the dependencies between speech sequences with a long time-step distance. To ensure the accuracy of the determined dependencies, the structure of the gated recurrent gate can be an LSTM (Long Short-Term Memory) structure. The above first feature can refer to the feature with lower complexity obtained after dimensionality reduction of the original amplitude spectrum feature of the original speech signal. The above second feature can refer to the adjustment obtained by adjusting the first feature according to the dependencies captured by the gated recurrent module.
[0068] In an alternative solution of this embodiment, to ensure the accuracy of recognizing the original speech signal using the speech processing model, in addition to the common frequency-domain feature extraction module, encoder, and decoder, the above speech processing model can at least include the above gated recurrent module. By using the gated recurrent gate to capture the above dependencies and process the first feature after dimensionality reduction, the fine-grainedness of the obtained second feature can be improved, thereby ensuring the accuracy of the target recognition result obtained after reconstructing the second feature.
[0069] In an alternative solution of this embodiment, when using the speech processing module to recognize the original speech signal based on the time-domain distribution information and frequency-domain distribution information, the above frequency-domain feature extraction module can first be used to extract the frequency-domain distribution information of the original speech signal in the frequency domain. Considering that among the various features of the speech signal, the amplitude spectrum feature can significantly reflect the differences between different sub-signals, the feature obtained by extracting the frequency-domain distribution information of the original speech signal can be the original amplitude spectrum feature of the original speech signal. After obtaining the original amplitude spectrum feature, the above encoder can be further used to encode the original amplitude spectrum feature, that is, perform dimensionality reduction processing to obtain the first feature corresponding to the original amplitude spectrum feature, and use the above gated recurrent module to process the first feature based on the above time-domain distribution information and frequency-domain distribution information to obtain the above second feature. Finally, the obtained second feature can be reconstructed using the above decoder to determine the target recognition result corresponding to the original speech signal.
[0070] In the embodiment of the present application, the frequency-domain feature extraction module includes: a transformation unit and an amplitude processing unit. Using the frequency-domain feature extraction module to extract the frequency-domain distribution information of the original speech signal to obtain the original amplitude spectrum feature of the original speech signal includes: using the transformation unit to perform a short-time Fourier transform on the original speech signal to obtain the frequency-domain distribution information of the original speech signal; using the amplitude processing unit to perform a modulus operation on the frequency-domain distribution information to obtain the original amplitude spectrum feature.
[0071] The above transformation unit can be used to perform a short-time Fourier transform on a speech signal to obtain a frequency-domain signal of the speech signal in the frequency domain.
[0072] In an alternative solution of this embodiment, considering that there may be non-stationary noise signals in the original speech signal, and non-stationary noise signals usually occur in a relatively short period of time, and the non-stationary noise signals corresponding to different times in the original speech signal may be different. Therefore, when using the frequency-domain feature extraction module to extract the original amplitude spectrum features of the original speech signal, the speech processing system can first use the above transformation unit to perform a short-time Fourier transform on the original speech information to determine the frequency-domain signals corresponding to different times of the original speech signal in the frequency domain, that is, the above frequency-domain distribution information. Then, use the above amplitude processing unit to process the frequency-domain distribution information to determine the modulus values of the frequency-domain distribution information at different times, so as to obtain the original amplitude spectrum features corresponding to the original speech signal.
[0073] In an alternative solution of this embodiment, in order to ensure the accuracy of the determined original amplitude spectrum features, when determining the above frequency-domain distribution information, the original speech signal can be subjected to multiple short-time Fourier transforms, and the parameters used in each short-time Fourier transform can be different.
[0074] In the embodiment of the present application, the encoder is composed of multiple convolutional blocks, and the decoder is composed of multiple transposed convolutional blocks. The outputs of the multiple convolutional blocks are respectively input into the corresponding transposed convolutional blocks among the multiple transposed convolutional blocks; using the encoder to perform dimensionality reduction processing on the original amplitude spectrum features to obtain a first feature corresponding to the original amplitude spectrum features, including: using the multiple convolutional blocks to sequentially perform multiple convolutional processes on the original amplitude spectrum features to obtain the first feature; using the decoder to perform reconstruction processing on the second feature to obtain a target recognition result, including: using the multiple transposed convolutional blocks to sequentially perform multiple transposed convolutional processes on the second feature based on the outputs of the multiple convolutional blocks to obtain the target recognition result.
[0075] The above multiple convolutional blocks and multiple transposed convolutional blocks correspond one by one. Multiple convolutional blocks can be set in the encoder, and correspondingly, multiple transposed convolutional blocks can be set in the decoder. After encoding the data using the convolutional blocks in the encoder, the encoded data can be decoded by the transposed convolutional blocks corresponding to the convolutional blocks in the decoder.
[0076] In an alternative solution of this embodiment, when using an encoder to perform dimensionality reduction processing on the original amplitude spectrum features, the speech processing system can use multiple convolutional blocks in the encoder to perform multiple convolutional processes on the above-mentioned original amplitude spectrum features to obtain corresponding first features. When using a decoder to reconstruct the second features, according to the execution order between the above-mentioned multiple convolutional blocks, the corresponding multiple deconvolutional blocks in the decoder can be used to perform deconvolutional processing on the above-mentioned second features in sequence according to the output results of the multiple convolutional blocks to obtain the corresponding target recognition results.
[0077] In the embodiment of the present application, the convolutional block and the deconvolutional block have the same structure. The convolutional block includes: a first convolutional layer, a sigmoid function layer, a second convolutional layer, a product layer, a normalization layer, and an exponential linear unit layer. Among them, the first convolutional layer is used to perform convolutional processing on the input data of the convolutional block to obtain a first convolutional result, the second convolutional layer is used to perform convolutional processing on the input data to obtain a second convolutional result, the sigmoid function layer is used to map the first convolutional result to obtain a mapping result, the product layer is used to multiply the mapping result and the second convolutional result to obtain a product result, the normalization layer is used to perform normalization processing on the product result to obtain a normalized result, and the exponential linear unit layer is used to perform exponential linear processing on the normalized result to obtain the output data of the convolutional block.
[0078] In an alternative solution of this embodiment, to ensure the accuracy of the obtained first features, the above-mentioned convolutional block may include: a first convolutional layer conv1, a sigmoid function layer sigmoid, a second convolutional layer conv2, a product layer product layer, a batch normalization layer batch normalization, and an exponential linear unit layer exponential linear unit. Among them, the first convolutional layer and the second convolutional layer are juxtaposed, that is, the first convolutional layer is used to perform convolutional processing on the input data received by the encoder to obtain a first convolutional result, the second convolutional layer is used to perform convolutional processing on the input data received by the encoder to obtain a second convolutional result, the sigmoid function layer is connected to the first convolutional layer and is used to perform mapping processing on the first convolutional result output by the first convolutional layer to obtain a corresponding mapping result, the product layer is connected to the second convolutional layer and the sigmoid function layer and is used to multiply the second convolutional result output by the second convolutional layer and the mapping result output by the sigmoid function layer to obtain a corresponding product result, the normalization layer is connected to the product layer and is used to perform normalization processing on the product result output by the product layer to obtain a corresponding normalized result, and the exponential linear unit layer is connected to the normalization layer and is used to perform exponential linear processing on the normalized result output by the normalization layer, thereby obtaining the output data of the convolutional block.
[0079] Corresponding to the convolutional block, the deconvolutional block may include: a first deconvolutional layer, a second deconvolutional layer, a sigmoid function layer, a product layer, a normalization layer, and an exponential linear unit layer. Among them, the first deconvolutional layer and the second deconvolutional layer are juxtaposed and correspond to the first convolutional layer and the second convolutional layer respectively. The sigmoid function layer is connected to the first deconvolutional layer, the product layer is connected to the second deconvolutional layer and the sigmoid function layer, the normalization layer is connected to the product layer, and the exponential linear unit layer is connected to the normalization layer.
[0080] For ease of understanding, Figure 4 FIG. 1 is a schematic diagram of a convolutional layer and a deconvolutional layer according to Embodiment 1 of the present application. Among them, the left part represents the schematic diagram of the convolutional layer, and the right part represents the schematic diagram of the deconvolutional layer. 401 represents the first convolutional layer, 402 represents the second convolutional layer, 403 represents the sigmoid function layer, 404 represents the product layer, 405 represents the normalization layer, 406 represents the exponential linear unit layer, 407 represents the first deconvolutional layer, and 408 represents the second deconvolutional layer. Among them, the output of the convolutional layer can be connected to the input of the deconvolutional layer, that is, the exponential linear unit layer in the convolutional layer can be connected to the first deconvolutional layer and the second deconvolutional layer in the deconvolutional layer to perform a deconvolution operation on the output result of the convolutional layer. To ensure the transmission efficiency, the connection method between the convolutional layer and the deconvolutional layer can be a skip connection.
[0081] In the embodiment of the present application, the gated recurrent module includes a plurality of gated recurrent units connected in sequence.
[0082] In an alternative solution of this embodiment, a plurality of gated recurrent units connected in sequence can be set in the above-mentioned gated recurrent module to extract the dependence relationship between the first features multiple times, thereby improving the accuracy of the dependence relationship between the first features extracted, that is, the second features.
[0083] In the embodiment of the present application, the above method further includes: obtaining training data, where the training data includes: a first speech signal and a second speech signal, and the second speech signal is used to represent the signal obtained after denoising the first speech signal; using an initial processing model to perform recognition on the first speech signal based on the time-domain distribution information of the first speech signal in the time domain and the frequency-domain distribution information of the first speech signal in the frequency domain to obtain a training recognition result; performing denoising processing on the first speech signal based on the training recognition result to obtain a predicted speech signal; determining a target loss function value corresponding to the initial processing model based on the second speech signal and the predicted speech signal; and adjusting the model parameters of the initial processing model based on the target loss function value to obtain a speech signal model.
[0084] The above-mentioned first speech signal may refer to a speech signal containing a noise signal. The above-mentioned second speech signal may be a clean speech signal, which is a signal obtained by performing noise reduction processing on the first speech signal. Considering that in the actual process of noise reduction of the speech signal, it may not be possible to completely remove the noise signal in the first speech signal. Therefore, a staff member can first record or select a clean speech signal as the above-mentioned second speech signal, and then add a perturbation to the second speech signal, that is, add a noise signal to the second speech signal to obtain the above-mentioned first speech signal.
[0085] In an alternative solution of this embodiment, in order to ensure the accuracy of the speech signal model, when training the initial processing model, first, training data constructed from the first speech signal and the corresponding second speech signal can be obtained, and the time-domain distribution information of the first speech signal in the time domain and the frequency-domain distribution information of the first speech signal in the frequency domain can be obtained. Then, using the above-mentioned initial processing model, based on the time-domain distribution information and frequency-domain distribution information of the first speech signal, the first speech signal is recognized to determine the corresponding mask information of the first speech signal, that is, the above-mentioned training recognition result. This training recognition result can be used to characterize whether the sub-signal in the first speech signal is a noise signal. After determining the training recognition result, the first speech signal can be further processed for noise reduction based on this training recognition result to obtain a preliminarily noise-reduced predicted speech signal. Based on this predicted speech signal and the corresponding second speech signal, a corresponding target loss function can be established. Finally, according to this target loss function, the model parameters of the above-mentioned initial processing model are adjusted, and the above-mentioned speech signal model can be obtained. To further improve the accuracy of the trained speech signal model, the adjusted initial processing model can be used as a new initial processing model, and the above process can be repeated multiple times to obtain a speech processing model with higher accuracy.
[0086] In the embodiment of this application, determining the target loss function value corresponding to the initial processing model based on the second speech signal and the predicted speech signal includes: determining the first loss function value corresponding to the initial processing model based on the second speech signal and the predicted speech signal; extracting features from the frequency-domain distribution information of the second speech signal to obtain the first amplitude spectrum feature of the second speech signal, and extracting features from the frequency-domain distribution information of the predicted speech signal to obtain the second amplitude spectrum feature of the predicted speech signal; determining the second loss function value corresponding to the initial processing model based on the first amplitude spectrum feature and the second amplitude spectrum feature; performing a weighted sum on the first loss function value and the second loss function value to obtain the target loss function value.
[0087] The above-mentioned first loss function value may refer to the loss function value in the time domain. The above-mentioned second loss function value may refer to the loss function value in the frequency domain.
[0088] In an alternative solution of this embodiment, to ensure the accuracy of the constructed target loss function, when determining the target loss function corresponding to the initial processing model, the first loss function value corresponding to the initial processing model can be determined first according to the predicted speech signal obtained by noise reduction and the corresponding second speech signal. At the same time, the frequency-domain distribution information of the second speech signal is extracted to determine the first amplitude spectrum feature corresponding to the second speech signal, and the frequency-domain distribution information of the predicted speech signal is extracted to determine the second amplitude spectrum feature corresponding to the predicted speech signal. According to the first amplitude spectrum feature and the second amplitude spectrum feature, the second loss function value corresponding to the initial processing model can be constructed. Finally, by integrating the above first loss function and the second loss function, such as weighted processing, a target loss function value with higher accuracy can be obtained. By determining the target loss function value based on the first loss function value in the time domain and the second loss function value in the frequency domain, and adjusting the model parameters of the initial processing model according to the target loss function value, the characteristics of the trained speech processing model in the time-frequency domain can be greatly improved. While retaining the spectral information of the speech signal, the dynamic range problem of the amplitude spectrum of the speech signal is alleviated, making the trained speech processing model pay more attention to the accuracy of the low-amplitude region, thereby improving the ability of the speech processing model to process the original speech signal in the time-frequency domain.
[0089] In an alternative solution of this embodiment, to ensure the accuracy of adjusting the initial processing model using the target loss function, the first loss function value used to construct the target loss function can refer to the loss function value of the predicted speech signal and the second speech signal in the time domain, which can be the Mse loss (Mean Squared Error loss) function value, that is, the square value of the average difference between the predicted speech signal and the second speech signal in the time domain. The second loss function value can refer to the multi-resolution loss function value of the first amplitude spectrum feature and the second amplitude spectrum feature in the frequency domain, which can be the multi-resolution Stft loss (Short-Time Fourier Transform loss) function value, and can be the sum of the amplitude loss values of the first amplitude spectrum feature and the second amplitude spectrum feature at different resolutions.
[0090] In the embodiment of the present application, feature extraction is performed on the frequency-domain distribution information of the second speech signal to obtain the first amplitude spectrum feature of the second speech signal, and feature extraction is performed on the frequency-domain distribution information of the predicted speech signal to obtain the second amplitude spectrum feature of the predicted speech signal, including: performing short-time Fourier transform on the second speech signal using multiple different transformation parameters to obtain the first frequency-domain distribution information of the second speech signal, and performing short-time Fourier transform on the predicted speech signal using multiple different transformation parameters to obtain the second frequency-domain distribution information of the predicted speech signal; performing a modulus operation on the first frequency-domain distribution information to obtain the first amplitude spectrum feature, and performing a modulus operation on the second frequency-domain distribution information to obtain the second amplitude spectrum feature.
[0091] In an alternative solution of this embodiment, when extracting the first amplitude spectrum feature of the second speech signal and the second amplitude spectrum feature of the predicted speech signal, the short-time Fourier transform can be performed on the above-mentioned second speech signal using multiple different transformation parameters according to the process of determining the original amplitude spectrum feature mentioned above to obtain the first frequency-domain distribution information of the corresponding first speech signal, and the modulus value of this first frequency-domain distribution information is determined to obtain the above-mentioned first spectrum feature. At the same time, the short-time Fourier transform is performed on the above-mentioned predicted speech signal using multiple different transformation parameters to obtain the second frequency-domain distribution information of the corresponding predicted speech signal, and the modulus value of this second frequency-domain distribution information is determined to obtain the above-mentioned second spectrum feature.
[0092] In the embodiment of the present application, based on the first amplitude spectrum feature and the second amplitude spectrum feature, the second loss function value corresponding to the initial processing model is determined, including: obtaining the spectral convergence loss of the first amplitude spectrum feature and the second amplitude spectrum feature to obtain the spectral convergence loss function value; obtaining the logarithmic short-time Fourier amplitude loss of the first amplitude spectrum feature and the second amplitude spectrum feature to obtain the amplitude loss function value; performing a weighted sum on the spectral convergence loss function value and the amplitude loss function value to obtain the second loss function value.
[0093] In an alternative solution of this embodiment, when determining the second loss function according to the first amplitude spectrum feature and the second amplitude spectrum feature, the difference value in the frequency domain between the first amplitude spectrum feature and the second amplitude spectrum feature can be obtained first, such as the spectral convergence loss between the two, to obtain the corresponding spectral convergence loss function. At the same time, the loss value generated during the process of obtaining the first amplitude spectrum feature and the second amplitude spectrum feature using the short-time Fourier transform is obtained, that is, the above-mentioned logarithmic short-time Fourier amplitude loss, to obtain the corresponding amplitude loss function value. Finally, the spectral convergence loss function value and the amplitude loss function value are integrated, such as weighted processing, to obtain a second loss function value with higher accuracy.
[0094] In an embodiment of the present application, the above method further includes: outputting an original voice signal and a target recognition result; in response to receiving an adjustment instruction for adjusting the target recognition result, determining an adjusted recognition result corresponding to the adjustment instruction; performing noise reduction processing on the original voice signal based on the adjusted recognition result to obtain a target voice signal; and adjusting the voice processing model based on the adjusted recognition result and the original voice signal.
[0095] In an alternative solution of this embodiment, the voice processing model can also be updated in real time when processing a voice signal containing a noise signal by using the voice processing model, so as to further improve the accuracy of the voice processing model. Specifically, the voice processing system can also output the original voice signal and the target recognition result corresponding to the original voice signal in a preset user interface. The user can view the situation of the noise signal of the original voice signal according to the target recognition result, and adjust the target recognition result according to their own experience or system prompts. The voice processing system can adjust the target recognition result according to the received adjustment instruction to obtain a corresponding adjusted recognition result, and finally adjust the current voice processing model according to the adjusted recognition result and the original voice signal, so as to further improve the accuracy of the voice processing model. While adjusting the voice processing model, the adjusted recognition result can also be used to perform noise reduction processing on the original voice signal to obtain a target voice signal that can meet the user's needs.
[0096] In an embodiment of the present application, performing noise reduction processing on the original voice signal based on the target recognition result to obtain a target voice signal includes: extracting features from the frequency domain distribution information of the original voice signal to obtain the original amplitude spectrum features of the original voice signal; multiplying the original amplitude spectrum features and the target recognition result element by element to obtain target amplitude spectrum features; and performing an inverse short-time Fourier transform on the target amplitude spectrum features to obtain a target voice signal.
[0097] In an alternative solution of this embodiment, when denoising the original speech signal using the target recognition result, the speech processing system can first extract the feature of the frequency domain distribution information of the original speech signal to obtain the original amplitude spectrum feature corresponding to the original speech signal. The feature extraction process can refer to the foregoing process and will not be elaborated here. After obtaining the original amplitude spectrum feature, the speech processing system can perform an element-by-element multiplication operation on the original amplitude spectrum feature and the target recognition result according to the elements included in the original amplitude spectrum feature, so as to obtain the corresponding target amplitude spectrum feature. If the user adjusts the obtained target recognition result to obtain an adjusted recognition result, then the adjusted recognition result can be used for the element-by-element multiplication with the original amplitude spectrum feature, so as to ensure that the obtained target speech signal better meets the user's expectations. When performing the element-by-element multiplication, if the value representing the removal of the sub-signal in the target recognition result is 0, then what can be retained in the target amplitude spectrum feature after multiplication is the amplitude spectrum feature corresponding to the clean speech signal. Finally, performing an inverse short-time Fourier transform on the target amplitude spectrum feature can obtain the target speech signal corresponding to the original speech signal.
[0098] To facilitate the understanding of the above process, Figure 5 FIG. 5 is a schematic diagram of a speech processing process according to Embodiment 1 of the present application. Among them, 501 represents the original speech signal that needs to be denoised, 502 represents the aforementioned transformation unit, which can be used to perform a short-time Fourier transform on the original speech signal to obtain the frequency domain distribution information of the original speech signal in the frequency domain, 503 represents the aforementioned amplitude value processing unit, which is used to perform an operation of taking the modulus value on the frequency domain distribution information to obtain the original amplitude spectrum feature of the original speech signal, 504 represents the gated convolutional recurrent network, which can include the aforementioned encoder, gated recurrent module, and decoder, 505 represents the determined target recognition result, 506 represents multiplying the target recognition result and the above-mentioned original amplitude spectrum feature element by element to obtain the target amplitude spectrum feature, 507 represents the inverse transformation unit, which can be used to perform an inverse short-time Fourier transform on the target amplitude spectrum feature to obtain the above-mentioned target speech signal, and 508 represents the target speech signal after denoising.
[0099] Figure 6Schematic diagram of a gated convolutional recurrent network according to Embodiment 1 of the present application. Among them, 601 represents the encoder, which may include multiple convolutional blocks mentioned above, such as the first convolutional layer, the second convolutional layer, the sigmoid function layer, the product layer, the normalization layer, and the exponential linear unit layer mentioned above, respectively represented by 6011, 6012, 6013, 6014, 6015, and can be used to process the original amplitude spectrum features to obtain the first feature mentioned above. 602 represents the gated recurrent module, which may include multiple gated recurrent units connected in sequence. Taking two gated recurrent units as an example, represented by 6021 and 6022, and can be used to process the first feature to obtain the second feature mentioned above. 603 represents the decoder, which may include deconvolutional blocks corresponding to multiple convolutional blocks, such as the first deconvolutional layer, the second deconvolutional layer, the sigmoid function layer, the product layer, the normalization layer, and the exponential linear unit layer mentioned above, respectively represented by 6033, 6032, 6033, 6034, 6035, and can be used to process the second feature to obtain the target recognition result mentioned above.
[0100] To facilitate understanding of the noise reduction effect of the above speech processing method, the following will illustrate by comparing the execution effects of the speech processing method in the prior art and the speech processing method of the present application.
[0101] First, Figure 7It is a schematic diagram of a traditional speech processing method according to Embodiment 1 of the present application. Among them, 701 represents the input original speech signal, 702 represents a transformation unit for performing short-time Fourier transform, 703 represents a power spectrum calculation unit for determining the original amplitude spectrum characteristics of the original speech signal, 704 represents a noise spectrum estimation unit for determining the mask information in the original amplitude spectrum characteristics, 705 represents a signal scaling unit for adjusting the mask information to obtain a target recognition result, 706 represents determining the target amplitude spectrum characteristics based on the original amplitude spectrum characteristics and the target recognition result, 707 represents an inverse transformation unit for performing inverse short-time Fourier transform on the target amplitude spectrum characteristics to obtain a target speech signal, and 708 represents the output target speech signal. Although the process of the traditional speech processing method is similar to that of the speech processing method proposed in the present application, in the traditional speech processing method, such as the frequency-domain single-channel speech noise reduction algorithm, the core point of the algorithm is to obtain the gain function. It is necessary to first transform the noisy speech, that is, the original speech signal, from the time-domain signal to the frequency-domain signal through short-time Fourier transform, then calculate the power spectrum of the frequency-domain signal, that is, the original amplitude spectrum characteristics, and then estimate the noise variance of the power spectrum. Among them, the method for estimating the noise variance is generally: first determine whether the current frame is a speech frame or a noise frame according to the speech detection module. If it is a noise frame, update the variance of the noise, otherwise, do not update the variance of the noise. After estimating the noise variance, the gain function can be estimated. There are various ways to estimate the gain function. Currently, a better estimation method is to first estimate the probability of speech presence, the a priori signal-to-noise ratio, and the a posteriori signal-to-noise ratio, and then multiply the estimated gain function by the frequency-domain form of the noisy speech to obtain the frequency-domain estimated form of the clean speech. Finally, through inverse Fourier transform, the time-domain estimated form of the clean speech is obtained, that is, the measurable waveform speech is obtained. However, this method does not consider the time-frequency domain characteristics of the original speech signal itself, resulting in the actual gain function being not ideal, and the effect of reducing noise on the original speech signal is also average. Figure 8 It is a schematic diagram of the noise reduction effect of a traditional speech processing according to Embodiment 1 of the present application. Among them, the area outlined by the white square is the area where part of the noise signal is located, as Figure 8 shown. By using the traditional speech processing method to reduce noise on the original speech signal, there will still be a lot of noise signals left.
[0102] And as Figure 9 shown, Figure 9It is a schematic diagram of a traditional gated convolutional recurrent network according to Embodiment 1 of the present application. Among them, 901 represents the encoder, 902 represents the gated convolution, 903 represents the decoder, 904 represents the 1x1 convolutional layer, and 905 represents the fully connected layer. In the traditional gated convolutional recurrent network, the input signal, that is, the original speech signal, needs to pass through an encoder composed of four gated convolution modules, use 1x1 convolution to connect the shallow feature representation and the deep feature representation, and then process the two feature representations by a three-layer time-frequency domain long short-term memory neural network to output two parts, the real part and the imaginary part. Finally, use a decoder composed of four transposed gated convolution modules to decode the above real part and imaginary part respectively and input them into the fully connected layer, and output after being processed by the fully connected layer. However, this is only applicable to the echo cancellation task. In this network, it is necessary to calculate the real part and the imaginary part separately. The network structure is too complex, the number of parameters used is too large, and it is impossible to process the signal in real time. The efficiency of denoising the original semantic signal is low. At the same time, the uneven distribution characteristics of the clean speech pre-noise in the time domain and the frequency domain are not considered, resulting in a poor denoising effect on the original speech signal in the actual denoising task, and even damaging the original speech signal.
[0103] The effect of denoising the speech signal using the traditional speech processing method can be as Figure 9 shown. Figure 9 It is a schematic diagram of the denoising effect of a traditional speech according to Embodiment 1 of the present application. From Figure 9 it can be seen that the denoising effect of the traditional speech processing method is poor. The target speech signal after denoising still contains a lot of noise signals, and the denoising effect is poor, and even the speech quality of the original speech signal will be damaged.
[0104] Finally, Figure 10 It is another schematic diagram of the denoising effect of a traditional speech according to Embodiment 1 of the present application. Figure 11 It is a schematic diagram of the denoising effect of a new speech according to Embodiment 1 of the present application. Figure 10 and Figure 11 are the target speech signals obtained by processing the same original speech signal. By Figure 10 and Figure 11 comparison, it can be seen that through the speech processing method of the present application, it is possible to more accurately distinguish whether the sub-signal in the original speech signal is a noise signal, and when denoising, multiply the original amplitude spectrum feature of the original speech signal and the target recognition result element by element, which can largely ensure the denoising effect on the original speech signal. Compared with the speech processing method in the prior art, there is an obvious improvement in the effect.
[0105] Embodiment 2
[0106] According to an embodiment of the present application, there is also provided a voice processing method. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0107] Figure 12 is a flowchart of a voice processing method according to Embodiment 2 of the present application. As Figure 12 shown, the method may include the following steps:
[0108] Step S1202, obtain the original voice signal in the live video.
[0109] The above original voice signal may refer to the voice signal in the live scene, and may include but not be limited to: the voice signal emitted by the anchor, the played music signal, the sound signal of typing on the keyboard, etc.
[0110] In an alternative solution of this embodiment, considering that in the live scene, there are usually many kinds of sound signals mixed together, resulting in the sound signals related to the live broadcast being drowned out and unable to be clearly heard. Therefore, when the user is watching the live broadcast, the voice processing system can obtain the original voice signal in the live video in real time.
[0111] Step S1204, identify the original voice signal based on the time-domain distribution information of the original voice signal in the time domain and the frequency-domain distribution information of the original voice signal in the frequency domain, and obtain the target recognition result corresponding to the original voice signal.
[0112] Among them, the target recognition result is used to characterize whether the sub-signal in the original voice signal is a noise signal.
[0113] The above target recognition result may refer to a mask obtained by network prediction, and can be used to indicate the retention or removal of the sub-signals included in the original voice signal. That is, if the sub-signal is a noise signal that needs to be removed, the target recognition result can be used to indicate the removal of the sub-signal, which can be represented by 0; if the sub-signal is a target signal that needs to be retained, the target recognition result can be used to indicate the retention of the sub-signal, which can be represented by 1. The above sub-signals may refer to the voice signals corresponding to different sound-emitting objects in the original voice signal.
[0114] Considering that in the original speech signal, the signal characteristics of the noise signal and the target signal are usually different in the time-frequency domain. Taking the speech emitted by the host and the air conditioner sound emitted by the air conditioner as an example, the duration of the speech emitted by the host is usually less than the duration of the air conditioner sound emitted by the air conditioner, and the frequency of the speech emitted by the host in a short period is usually greater than the frequency of the air conditioner sound emitted by the air conditioner in a short period. Therefore, when determining whether a sub-signal in the original speech signal is a noise signal, the speech processing system can obtain the signal characteristics of the original speech signal in the time-frequency domain, analyze the original speech signal to determine whether the corresponding sub-signal is a noise signal. Specifically, the speech processing system can first extract the feature of the frequency-domain distribution information of the original speech signal in the time domain to obtain the time-domain distribution information of the original speech signal in the time domain, and extract and obtain the time-domain distribution information of the original speech signal in the frequency domain by extracting the feature of the frequency-domain distribution information of the original speech signal in the frequency domain. Then, using the time-domain distribution information and the frequency-domain distribution information, identify and analyze the sub-signals in the original speech signal, so as to obtain a target recognition result that can reflect whether the sub-signal is a noise signal.
[0115] Step S1206, perform noise reduction processing on the original speech signal based on the target recognition result to obtain the target speech signal.
[0116] After determining the target recognition result corresponding to the original speech signal, the speech processing system can, according to the target recognition result, determine the sub-signals belonging to the noise signal from the original speech signal and remove the sub-signals to obtain the pure speech signal contained in the original speech signal, that is, the target speech signal.
[0117] Step S1208, replace the original speech signal in the live video with the target speech signal.
[0118] After performing noise reduction on the original speech signal, the speech processing system can replace the original speech signal in the live video with the noise-reduced target speech signal, thereby reducing the impact of the noise signal on the user's viewing of the live broadcast and improving the user experience.
[0119] It should be noted that the preferred implementation schemes involved in the above embodiments of the present application are the same as the schemes, application scenarios, and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.
[0120] Embodiment 3
[0121] According to an embodiment of the present application, a speech processing method is further provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0122] Figure 13 is a flowchart of a voice processing method according to Embodiment 3 of the present application. As Figure 13 shown, the method may include the following steps:
[0123] Step S1302, in response to an input instruction acting on the operation interface, display a waveform diagram of the original voice signal on the operation interface.
[0124] The above input instruction may refer to an instruction for the user to input to display the original voice signal.
[0125] In an alternative solution of this embodiment, to facilitate the user to view the processing of the original voice signal, the voice processing system may also show the noise reduction process to the user on a preset operation interface. Specifically, the user can select the original voice signal to be subjected to noise reduction processing according to needs. When the voice processing system receives the input instruction input by the user, it can display the original voice signal in the preset operation interface.
[0126] Step S1304, in response to a noise reduction instruction acting on the operation interface, display a waveform diagram of the target voice signal on the operation interface.
[0127] Among them, the target voice signal is obtained by performing noise reduction processing on the original voice signal based on the target recognition result corresponding to the original voice signal. The target recognition result is obtained by recognizing the original voice signal based on the time-domain distribution information of the original voice signal in the time domain and the frequency-domain distribution information of the original voice signal in the frequency domain. The target recognition result is used to characterize whether the sub-signal in the original voice signal is a noise signal.
[0128] The above noise reduction instruction may refer to an instruction for removing a noise signal from the original voice signal.
[0129] In an alternative solution of this embodiment, after the user views the selected original voice signal, the user can continue to input the above noise reduction instruction to the voice processing system in the operation interface. After the voice processing system receives the noise reduction instruction, it can perform noise reduction processing on the original voice signal to obtain the target voice signal, and display the target voice signal in the above operation interface to facilitate the user to view.
[0130] Specifically, the speech processing system can first extract the frequency-domain distribution information of the original speech signal in the time domain to obtain the time-domain distribution information of the original speech signal in the time domain, and extract the frequency-domain distribution information of the original speech signal in the frequency domain to obtain the time-domain distribution information of the original speech signal in the frequency domain. Then, the time-domain distribution information and the frequency-domain distribution information are used to identify and analyze the sub-signals in the original speech signal, so as to obtain a target recognition result that can reflect whether the sub-signal is a noise signal. After determining the target recognition result corresponding to the original speech signal, the speech processing system can determine the sub-signals belonging to the noise signal from the original speech signal according to the target recognition result and remove the sub-signals. For example, the original amplitude spectrum feature of the original speech signal and the target recognition result are multiplied element by element to obtain the pure speech signal contained in the original speech signal, that is, the target speech signal.
[0131] It should be noted that the preferred implementation schemes involved in the above embodiments of the present application are the same as the schemes, application scenarios, and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.
[0132] Embodiment 4
[0133] According to an embodiment of the present application, a speech processing method is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0134] Figure 14 is a flowchart of a speech processing method according to Embodiment 4 of the present application, as Figure 14 shown, the method may include the following steps:
[0135] Step S1402, obtain the original speech signal by calling the first interface.
[0136] Among them, the first interface includes a first parameter, and the parameter value of the first parameter includes the original speech signal.
[0137] In an optional solution of this embodiment, the speech processing system can first receive the original speech signal that needs to be denoised through a preset first interface, that is, receive the above-mentioned first parameter.
[0138] Step S1404, identify the original speech signal based on the time-domain distribution information of the original speech signal in the time domain and the frequency-domain distribution information of the original speech signal in the frequency domain to obtain the target recognition result corresponding to the original speech signal.
[0139] Among them, the target recognition result is used to characterize whether the sub-signal in the original speech signal is a noise signal.
[0140] After obtaining the original speech signal, the speech processing system can first extract the frequency-domain distribution information of the original speech signal in the time domain to obtain the time-domain distribution information of the original speech signal in the time domain, and extract the frequency-domain distribution information of the original speech signal in the frequency domain to obtain the time-domain distribution information of the original speech signal in the frequency domain. Then, using the time-domain distribution information and the frequency-domain distribution information, the sub-signals in the original speech signal are identified and analyzed, so as to obtain a target recognition result that can reflect whether the sub-signal is a noise signal.
[0141] Step S1406: Perform noise reduction processing on the original speech signal based on the target recognition result to obtain the target speech signal.
[0142] After determining the target recognition result corresponding to the original speech signal, the speech processing system can, according to the target recognition result, determine the sub-signals belonging to the noise signal from the original speech signal and remove the sub-signals. For example, multiply the original amplitude spectrum feature of the original speech signal and the target recognition result element by element to obtain the pure speech signal contained in the original speech signal, that is, the target speech signal.
[0143] Step S1408: Output the target speech signal by calling the second interface.
[0144] Among them, the second interface includes a second parameter, and the parameter value of the second parameter includes the target speech signal.
[0145] Finally, after obtaining the target speech signal, the speech processing system can output the above-mentioned target speech signal by calling the preset second interface. For example, display the target speech signal in a preset user interface or operation interface for the user to view conveniently.
[0146] It should be noted that the preferred implementation schemes involved in the above embodiments of the present application are the same as the schemes, application scenarios, and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.
[0147] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.
[0148] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0149] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to enable a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of various embodiments of this application.
[0150] Embodiment 5
[0151] According to an embodiment of this application, there is also provided a device for implementing the above voice processing method, and this device can be deployed in a target client. Figure 15 is a structural block diagram of a voice processing device according to Embodiment 5 of this application, as Figure 15 shown. The device 1500 includes: a first signal acquisition module 1502, a first signal recognition module 1504, and a first signal noise reduction module 1506.
[0152] Among them, the first signal acquisition module 1502 is used to acquire an original voice signal; the signal recognition module 1504 is used to recognize the original voice signal based on the time-domain distribution information of the original voice signal in the time domain and the frequency-domain distribution information of the original voice signal in the frequency domain, and obtain a target recognition result corresponding to the original voice signal, where the target recognition result is used to represent whether the sub-signal in the original voice signal is a noise signal; the first signal noise reduction module 1506 is used to perform noise reduction processing on the original voice signal based on the target recognition result to obtain a target voice signal.
[0153] Optionally, the signal recognition module 1504 includes: an information recognition unit, which is used to use a voice processing model to recognize the original voice signal based on the time-domain distribution information and the frequency-domain distribution information to obtain a target recognition result.
[0154] Optionally, the speech processing model includes: a frequency-domain feature extraction module, an encoder, a gated recurrent module, and a decoder. The information recognition unit is further configured to: use the frequency-domain feature extraction module to extract features from the frequency-domain distribution information of the original speech signal to obtain the original amplitude spectrum features of the original speech signal; use the encoder to perform dimensionality reduction processing on the original amplitude spectrum features to obtain a first feature corresponding to the original amplitude spectrum features; use the gated recurrent module to process the first feature based on the time-domain distribution information to obtain a second feature; and use the decoder to perform reconstruction processing on the second feature to obtain the target recognition result.
[0155] Optionally, the frequency-domain feature extraction module includes: a transformation unit and an amplitude processing unit. The information recognition unit is further configured to: use the transformation unit to perform a short-time Fourier transform on the original speech signal to obtain the frequency-domain distribution information of the original speech signal; and use the amplitude processing unit to perform a modulus operation on the frequency-domain distribution information to obtain the original amplitude spectrum features.
[0156] Optionally, the encoder is composed of multiple convolutional blocks, and the decoder is composed of multiple deconvolutional blocks. The outputs of the multiple convolutional blocks are respectively input into the corresponding deconvolutional blocks among the multiple deconvolutional blocks. The information recognition unit is further configured to: use the multiple convolutional blocks to perform multiple convolutional processes on the original amplitude spectrum features in sequence to obtain a first feature; and use the decoder to perform reconstruction processing on the second feature to obtain the target recognition result, including: using the multiple deconvolutional blocks to perform multiple deconvolutional processes on the second feature based on the outputs of the multiple convolutional blocks in sequence to obtain the target recognition result.
[0157] Optionally, the convolutional block and the deconvolutional block have the same structure. The convolutional block includes: a first convolutional layer, a sigmoid function layer, a second convolutional layer, a product layer, a normalization layer, and an exponential linear unit layer. Among them, the first convolutional layer is used to perform convolutional processing on the input data of the convolutional block to obtain a first convolutional result, the second convolutional layer is used to perform convolutional processing on the input data to obtain a second convolutional result, the sigmoid function layer is used to map the first convolutional result to obtain a mapping result, the product layer is used to multiply the mapping result and the second convolutional result to obtain a product result, the normalization layer is used to perform normalization processing on the product result to obtain a normalized result, and the exponential linear unit layer is used to perform exponential linear processing on the normalized result to obtain the output data of the convolutional block.
[0158] Optionally, the gated recurrent module includes a plurality of gated recurrent units connected in sequence.
[0159] Optionally, the above device further includes: a data acquisition module for acquiring training data, where the training data includes: a first voice signal and a second voice signal, and the second voice signal is used to represent the signal obtained after denoising the first voice signal; an information recognition module for using an initial processing model to recognize the first voice signal based on the time-domain distribution information of the first voice signal in the time domain and the frequency-domain distribution information of the first voice signal in the frequency domain to obtain a training recognition result; an information denoising module for denoising the first voice signal based on the training recognition result to obtain a predicted voice signal; a loss function determination module for determining a target loss function value corresponding to the initial processing model based on the second voice signal and the predicted voice signal; and a model adjustment module for adjusting the model parameters of the initial processing model based on the target loss function value to obtain a voice signal model.
[0160] Optionally, the loss function determination module includes: a first loss function value determination unit for determining a first loss function value corresponding to the initial processing model based on the second voice signal and the predicted voice signal; a feature extraction unit for extracting features from the frequency-domain distribution information of the second voice signal to obtain a first amplitude spectrum feature of the second voice signal, and extracting features from the frequency-domain distribution information of the predicted voice signal to obtain a second amplitude spectrum feature of the predicted voice signal; a second loss function value determination unit for determining a second loss function value corresponding to the initial processing model based on the first amplitude spectrum feature and the second amplitude spectrum feature; and a target loss value determination unit for performing a weighted sum on the first loss function value and the second loss function value to obtain a target loss function value.
[0161] Optionally, the feature extraction unit is further configured to: perform short-time Fourier transform on the second voice signal using multiple different transformation parameters to obtain first frequency-domain distribution information of the second voice signal, and perform short-time Fourier transform on the predicted voice signal using multiple different transformation parameters to obtain second frequency-domain distribution information of the predicted voice signal; perform a modulus operation on the first frequency-domain distribution information to obtain a first amplitude spectrum feature, and perform a modulus operation on the second frequency-domain distribution information to obtain a second amplitude spectrum feature.
[0162] Optionally, the second loss function value determination unit is further configured to: obtain a spectral convergence loss between the first amplitude spectrum feature and the second amplitude spectrum feature to obtain a spectral convergence loss function value; obtain a logarithmic short-time Fourier amplitude loss between the first amplitude spectrum feature and the second amplitude spectrum feature to obtain an amplitude loss function value; and perform a weighted sum on the spectral convergence loss function value and the amplitude loss function value to obtain a second loss function value.
[0163] Optionally, the above device further includes: an information output module for outputting the original voice signal and the target recognition result; an information determination module for determining an adjusted recognition result corresponding to an adjustment instruction in response to receiving an adjustment instruction for adjusting the target recognition result; an original signal noise reduction module for performing noise reduction processing on the original voice signal based on the adjusted recognition result to obtain a target voice signal; and a model adjustment module for adjusting the voice processing model based on the adjusted recognition result and the original voice signal.
[0164] Optionally, the first signal noise reduction module 1506 includes: an amplitude spectrum feature extraction unit for extracting features from the frequency domain distribution information of the original voice signal to obtain the original amplitude spectrum feature of the original voice signal; an amplitude spectrum determination unit for multiplying the original amplitude spectrum feature and the target recognition result element by element to obtain a target amplitude spectrum feature; and a voice signal determination unit for performing an inverse short-time Fourier transform on the target amplitude spectrum feature to obtain a target voice signal.
[0165] It should be noted here that the above first signal acquisition module 1502, signal recognition module 1504, and first signal noise reduction module 1506 correspond to steps S302 to S306 in Embodiment 1. The functions and application scenarios of the three modules are the same as those of the corresponding steps, but are not limited to the content disclosed in Embodiment 1 above. It should be noted that the above modules or units may be hardware components or software components stored in a memory (for example, memory 104) and processed by one or more processors (for example, processors 102a, 102b,..., 102n). The above modules may also be part of the device and can run in the computer terminal 10 provided in Embodiment 1.
[0166] It should be noted that the preferred implementation schemes involved in the above embodiments of the present application are the same as the schemes, application scenarios, and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.
[0167] Embodiment 6
[0168] According to an embodiment of the present application, there is also provided a device for implementing the above voice processing method, and the device can be deployed in a target client. Figure 16 is a structural block diagram of a voice processing device according to Embodiment 6 of the present application, as Figure 16 shown. The device 1600 includes: a second signal acquisition module 1602, a second signal recognition module 1604, a second signal noise reduction module 1606, and a signal replacement module 1608.
[0169] Among them, the second signal acquisition module 1602 is used to acquire the original voice signal in the live video; the second signal recognition module 1604 is used to recognize the original voice signal based on the time-domain distribution information of the original voice signal in the time domain and the frequency-domain distribution information of the original voice signal in the frequency domain, and obtain the target recognition result corresponding to the original voice signal, where the target recognition result is used to characterize whether the sub-signal in the original voice signal is a noise signal; the second signal noise reduction module 1606 is used to perform noise reduction processing on the original voice signal based on the target recognition result to obtain the target voice signal; the signal replacement module 1608 is used to replace the original voice signal in the live video with the target voice signal.
[0170] It should be noted here that the above-mentioned second signal acquisition module 1602, second signal recognition module 1604, second signal noise reduction module 1606, and signal replacement module 1608 correspond to steps S1202 to S1208 in Embodiment 2. The examples and application scenarios implemented by the four modules and the corresponding steps are the same, but are not limited to the content disclosed in the above-mentioned Embodiment 1. It should be noted that the above-mentioned module or unit can be a hardware component or a software component stored in a memory (for example, memory 104) and processed by one or more processors (for example, processors 102a, 102b,..., 102n). The above-mentioned module can also be a part of the device and can run in the computer terminal 10 provided in Embodiment 1.
[0171] It should be noted that the preferred implementation schemes involved in the above embodiments of the present application are the same as the schemes, application scenarios, and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.
[0172] Embodiment 7
[0173] According to an embodiment of the present application, there is also provided a device for implementing the above voice processing method, and the device can be deployed in a target client. Figure 17 It is a structural block diagram of a voice processing device according to Embodiment 7 of the present application, as Figure 17 shown. The device 1700 includes: a first waveform display module 1702 and a second waveform display module 1704.
[0174] Among them, the first waveform display module 1702 is configured to display a waveform diagram of an original voice signal on the operation interface in response to an input instruction acting on the operation interface; the second waveform display module 1704 is configured to display a waveform diagram of a target voice signal on the operation interface in response to a noise reduction instruction acting on the operation interface, where the target voice signal is obtained by performing noise reduction processing on the original voice signal based on a target recognition result corresponding to the original voice signal, the target recognition result is obtained by recognizing the original voice signal based on the time-domain distribution information of the original voice signal in the time domain and the frequency-domain distribution information of the original voice signal in the frequency domain, and the target recognition result is used to characterize whether a sub-signal in the original voice signal is a noise signal.
[0175] It should be noted here that the above-mentioned first waveform display module 1702 and second waveform display module 1704 correspond to steps S1302 to S1304 in Embodiment 3. The examples and application scenarios implemented by the two modules and the corresponding steps are the same, but are not limited to the content disclosed in the above-mentioned Embodiment 1. It should be noted that the above-mentioned module or unit may be a hardware component or a software component stored in a memory (for example, memory 104) and processed by one or more processors (for example, processors 102a, 102b,..., 102n), and the above-mentioned module may also be part of a device and may run in the computer terminal 10 provided in Embodiment 1.
[0176] It should be noted that the preferred implementation schemes involved in the above embodiments of the present application are the same as the schemes, application scenarios, and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.
[0177] Embodiment 8
[0178] According to an embodiment of the present application, there is also provided a device for implementing the above-mentioned voice processing method, and the device may be deployed in a target client. Figure 18 It is a structural block diagram of a voice processing device according to Embodiment 8 of the present application, as Figure 18 shown, the device 1800 includes: a third signal acquisition module 1802, a second signal recognition module 1804, a third signal noise reduction module 1806, and a signal output module 1808.
[0179] Among them, the third signal acquisition module 1802 is used to acquire the original voice signal by calling the first interface. The first interface includes a first parameter, and the parameter value of the first parameter includes the original voice signal. The second signal recognition module 1804 is used to recognize the original voice signal based on the time-domain distribution information of the original voice signal in the time domain and the frequency-domain distribution information of the original voice signal in the frequency domain, and obtain a target recognition result corresponding to the original voice signal, where the target recognition result is used to characterize whether the sub-signal in the original voice signal is a noise signal. The third signal noise reduction module 1806 is used to perform noise reduction processing on the original voice signal based on the target recognition result to obtain a target voice signal. The signal output module 1808 is used to output the target voice signal by calling the second interface. The second interface includes a second parameter, and the parameter value of the second parameter includes the target voice signal.
[0180] It should be noted here that the above-mentioned third signal acquisition module 1802, second signal recognition module 1804, third signal noise reduction module 1806, and signal output module 1808 correspond to steps S1402 to S1408 in Embodiment 4. The examples and application scenarios implemented by the two modules and the corresponding steps are the same, but are not limited to the content disclosed in the above-mentioned Embodiment 1. It should be noted that the above-mentioned module or unit can be a hardware component or a software component stored in a memory (for example, memory 104) and processed by one or more processors (for example, processors 102a, 102b,..., 102n). The above-mentioned module can also be a part of the device and can run in the computer terminal 10 provided in Embodiment 1.
[0181] It should be noted that the preferred implementation schemes involved in the above-mentioned embodiments of the present application are the same as the schemes, application scenarios, and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.
[0182] Embodiment 9
[0183] An embodiment of the present application can provide a computer terminal, and the computer terminal can be any computer terminal device in a computer terminal group. Optionally, in this embodiment, the above-mentioned computer terminal can also be replaced with a terminal device such as a mobile terminal.
[0184] Optionally, in this embodiment, the above-mentioned computer terminal can be located in at least one of multiple network devices in a computer network.
[0185] In this embodiment, the above computer terminal may execute the program code of the following steps in the voice processing method: obtaining an original voice signal; recognizing the original voice signal based on the time-domain distribution information of the original voice signal in the time domain and the frequency-domain distribution information of the original voice signal in the frequency domain, to obtain a target recognition result corresponding to the original voice signal, where the target recognition result is used to characterize whether the sub-signal in the original voice signal is a noise signal; performing noise reduction processing on the original voice signal based on the target recognition result to obtain a target voice signal.
[0186] Optionally, Figure 19 is a structural block diagram of a computer terminal according to Embodiment 9 of the present application. As Figure 19 shown, the computer terminal A may include: one or more (only one is shown in the figure) processors 102, a memory 104, a storage controller, and a peripheral interface, where the peripheral interface is connected to a radio frequency module, an audio module, and a display.
[0187] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the voice processing method and device in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implements the above voice processing method. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory may further include a memory remotely disposed relative to the processor, and these remote memories may be connected to the terminal A through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0188] The processor may call the information and application programs stored in the memory through a transmission device to execute the following steps: obtaining an original voice signal; recognizing the original voice signal based on the time-domain distribution information of the original voice signal in the time domain and the frequency-domain distribution information of the original voice signal in the frequency domain, to obtain a target recognition result corresponding to the original voice signal, where the target recognition result is used to characterize whether the sub-signal in the original voice signal is a noise signal; performing noise reduction processing on the original voice signal based on the target recognition result to obtain a target voice signal.
[0189] Optionally, the above processor may also execute the program code of the following steps: recognizing the original voice signal based on the time-domain distribution information of the original voice signal in the time domain and the frequency-domain distribution information of the original voice signal in the frequency domain, to obtain a target recognition result corresponding to the original voice signal, including: using a voice processing model to recognize the original voice signal based on the time-domain distribution information and the frequency-domain distribution information to obtain a target recognition result.
[0190] Optionally, the above-mentioned processor can also execute the program code of the following steps: The voice processing model includes: a frequency-domain feature extraction module, an encoder, a gated recurrent module, and a decoder. The original voice signal is recognized based on the time-domain distribution information and the frequency-domain distribution information by using the voice processing model to obtain a target recognition result, including: using the frequency-domain feature extraction module to extract the frequency-domain distribution information of the original voice signal to obtain the original amplitude spectrum feature of the original voice signal; using the encoder to perform dimensionality reduction processing on the original amplitude spectrum feature to obtain a first feature corresponding to the original amplitude spectrum feature; using the gated recurrent module to process the first feature based on the time-domain distribution information to obtain a second feature; using the decoder to perform reconstruction processing on the second feature to obtain the target recognition result.
[0191] Optionally, the above-mentioned processor can also execute the program code of the following steps: The frequency-domain feature extraction module includes: a transformation unit and an amplitude processing unit. Using the frequency-domain feature extraction module to extract the frequency-domain distribution information of the original voice signal to obtain the original amplitude spectrum feature of the original voice signal, including: using the transformation unit to perform a short-time Fourier transform on the original voice signal to obtain the frequency-domain distribution information of the original voice signal; using the amplitude processing unit to perform a modulus operation on the frequency-domain distribution information to obtain the original amplitude spectrum feature.
[0192] Optionally, the above-mentioned processor can also execute the program code of the following steps: The encoder consists of multiple convolutional blocks, and the decoder consists of multiple deconvolutional blocks. The outputs of the multiple convolutional blocks are respectively input into the corresponding deconvolutional blocks among the multiple deconvolutional blocks; using the encoder to perform dimensionality reduction processing on the original amplitude spectrum feature to obtain a first feature corresponding to the original amplitude spectrum feature, including: using the multiple convolutional blocks to sequentially perform multiple convolutional processes on the original amplitude spectrum feature to obtain the first feature; using the decoder to perform reconstruction processing on the second feature to obtain the target recognition result, including: using the multiple deconvolutional blocks to sequentially perform multiple deconvolutional processes on the second feature based on the outputs of the multiple convolutional blocks to obtain the target recognition result.
[0193] Optionally, the above-mentioned processor can also execute the program code of the following steps: The convolutional block and the deconvolutional block have the same structure. The convolutional block includes: a first convolutional layer, a sigmoid function layer, a second convolutional layer, a product layer, a normalization layer, and an exponential linear unit layer. Among them, the first convolutional layer is used to perform a convolutional process on the input data of the convolutional block to obtain a first convolutional result, the second convolutional layer is used to perform a convolutional process on the input data to obtain a second convolutional result, the sigmoid function layer is used to map the first convolutional result to obtain a mapping result, the product layer is used to multiply the mapping result and the second convolutional result to obtain a product result, the normalization layer is used to perform normalization processing on the product result to obtain a normalized result, and the exponential linear unit layer is used to perform exponential linear processing on the normalized result to obtain the output data of the convolutional block.
[0194] Optionally, the above-mentioned processor may also execute the program code of the following steps: The gated recurrent module includes a plurality of gated recurrent units connected in sequence.
[0195] Optionally, the above-mentioned processor may also execute the program code of the following steps: The above method further includes: obtaining training data, where the training data includes: a first speech signal and a second speech signal, and the second speech signal is used to represent the signal obtained after denoising the first speech signal; using an initial processing model to identify the first speech signal based on the time-domain distribution information of the first speech signal in the time domain and the frequency-domain distribution information of the first speech signal in the frequency domain, to obtain a training recognition result; performing denoising processing on the first speech signal based on the training recognition result to obtain a predicted speech signal; determining a target loss function value corresponding to the initial processing model based on the second speech signal and the predicted speech signal; adjusting the model parameters of the initial processing model based on the target loss function value to obtain a speech signal model.
[0196] Optionally, the above-mentioned processor may also execute the program code of the following steps: Determining a target loss function value corresponding to the initial processing model based on the second speech signal and the predicted speech signal includes: determining a first loss function value corresponding to the initial processing model based on the second speech signal and the predicted speech signal; extracting features from the frequency-domain distribution information of the second speech signal to obtain a first amplitude spectrum feature of the second speech signal, and extracting features from the frequency-domain distribution information of the predicted speech signal to obtain a second amplitude spectrum feature of the predicted speech signal; determining a second loss function value corresponding to the initial processing model based on the first amplitude spectrum feature and the second amplitude spectrum feature; performing a weighted sum on the first loss function value and the second loss function value to obtain the target loss function value.
[0197] Optionally, the above-mentioned processor may also execute the program code of the following steps: Extracting features from the frequency-domain distribution information of the second speech signal to obtain a first amplitude spectrum feature of the second speech signal, and extracting features from the frequency-domain distribution information of the predicted speech signal to obtain a second amplitude spectrum feature of the predicted speech signal includes: performing short-time Fourier transform on the second speech signal using a plurality of different transformation parameters to obtain first frequency-domain distribution information of the second speech signal, and performing short-time Fourier transform on the predicted speech signal using a plurality of different transformation parameters to obtain second frequency-domain distribution information of the predicted speech signal; performing a modulus operation on the first frequency-domain distribution information to obtain the first amplitude spectrum feature, and performing a modulus operation on the second frequency-domain distribution information to obtain the second amplitude spectrum feature.
[0198] Optionally, the above-mentioned processor may also execute the program code of the following steps: determining a second loss function value corresponding to the initial processing model based on the first amplitude spectrum feature and the second amplitude spectrum feature, including: obtaining a spectral convergence loss of the first amplitude spectrum feature and the second amplitude spectrum feature to obtain a spectral convergence loss function value; obtaining a logarithmic short-time Fourier amplitude loss of the first amplitude spectrum feature and the second amplitude spectrum feature to obtain an amplitude loss function value; performing a weighted sum on the spectral convergence loss function value and the amplitude loss function value to obtain the second loss function value.
[0199] Optionally, the above-mentioned processor may also execute the program code of the following steps: The above method further includes: outputting the original speech signal and the target recognition result; in response to receiving an adjustment instruction for adjusting the target recognition result, determining an adjusted recognition result corresponding to the adjustment instruction; performing noise reduction processing on the original speech signal based on the adjusted recognition result to obtain a target speech signal; adjusting the speech processing model based on the adjusted recognition result and the original speech signal.
[0200] Optionally, the above-mentioned processor may also execute the program code of the following steps: performing noise reduction processing on the original speech signal based on the target recognition result to obtain a target speech signal, including: extracting features from the frequency domain distribution information of the original speech signal to obtain the original amplitude spectrum feature of the original speech signal; multiplying the original amplitude spectrum feature and the target recognition result element by element to obtain a target amplitude spectrum feature; performing an inverse short-time Fourier transform on the target amplitude spectrum feature to obtain the target speech signal.
[0201] In the embodiment of the present application, the method is adopted: obtaining the original speech signal; performing recognition on the original speech signal based on the time domain distribution information of the original speech signal in the time domain and the frequency domain distribution information of the original speech signal in the frequency domain to obtain a target recognition result corresponding to the original speech signal, where the target recognition result is used to characterize whether the sub-signal in the original speech signal is a noise signal; performing noise reduction processing on the original speech signal based on the target recognition result to obtain a target speech signal. By using the time domain distribution information and frequency domain distribution information of the original speech signal in the time domain and the frequency domain to reflect the distribution characteristics of the original speech signal in the time domain and the frequency domain, and at the same time combining the time domain distribution information and the frequency domain distribution information to determine the target recognition result of the original speech signal, the judgment of whether the sub-signal in the original speech signal is a noise signal is made more accurate. Finally, noise reduction processing is performed on the original speech signal according to the target recognition result, which can avoid damaging the original speech signal during the noise reduction process, ensure that the target speech signal after noise reduction is closer to the pure speech signal, thereby ensuring the accuracy of the obtained target speech signal, improving the effect of noise reduction on the original speech signal, and further solving the technical problems in the related art that the speech processing method has limitations, resulting in a large degree of damage to the original speech signal after noise reduction or incomplete elimination of some special non-stationary noises.
[0202] Those of ordinary skill in the art can understand that Figure 19 the structure shown is only illustrative, and the computer terminal can also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a handheld computer, and terminal devices such as Mobile Internet Devices (MID), PAD, etc. Figure 19 It does not limit the structure of the above-mentioned electronic device. For example, computer terminal A may also include more or fewer components (such as a network interface, a display device, etc.) than those shown in Figure 19 and may have a different configuration from that shown in Figure 19 .
[0203] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium. The storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, an optical disk, etc.
[0204] Embodiment 10
[0205] An embodiment of the present application also provides a storage medium. Optionally, in this embodiment, the above storage medium can be used to store the program code executed by the voice processing method provided in the first embodiment above.
[0206] Optionally, in this embodiment, the above storage medium can be located in any one of the computer terminals in the computer terminal group in the computer network, or in any one of the mobile terminals in the mobile terminal group.
[0207] Optionally, in this embodiment, the storage medium is set to store program code for performing the following steps: obtaining an original voice signal; performing recognition on the original voice signal based on the time-domain distribution information of the original voice signal in the time domain and the frequency-domain distribution information of the original voice signal in the frequency domain to obtain a target recognition result corresponding to the original voice signal, where the target recognition result is used to characterize whether the sub-signal in the original voice signal is a noise signal; performing noise reduction processing on the original voice signal based on the target recognition result to obtain a target voice signal.
[0208] Optionally, the storage medium is configured to store program code further for performing the following steps: recognizing the original speech signal based on the time-domain distribution information of the original speech signal in the time domain and the frequency-domain distribution information of the original speech signal in the frequency domain, to obtain a target recognition result corresponding to the original speech signal, including: using a speech processing model to recognize the original speech signal based on the time-domain distribution information and the frequency-domain distribution information, to obtain the target recognition result.
[0209] Optionally, the storage medium is configured to store program code further for performing the following steps: the speech processing model includes: a frequency-domain feature extraction module, an encoder, a gated recurrent module, and a decoder; using the speech processing model to recognize the original speech signal based on the time-domain distribution information and the frequency-domain distribution information, to obtain a target recognition result, including: using the frequency-domain feature extraction module to extract features from the frequency-domain distribution information of the original speech signal, to obtain the original amplitude spectrum features of the original speech signal; using the encoder to perform dimensionality reduction processing on the original amplitude spectrum features, to obtain a first feature corresponding to the original amplitude spectrum features; using the gated recurrent module to process the first feature based on the time-domain distribution information, to obtain a second feature; using the decoder to perform reconstruction processing on the second feature, to obtain the target recognition result.
[0210] Optionally, the storage medium is configured to store program code further for performing the following steps: the frequency-domain feature extraction module includes: a transformation unit and an amplitude processing unit; using the frequency-domain feature extraction module to extract features from the frequency-domain distribution information of the original speech signal, to obtain the original amplitude spectrum features of the original speech signal, including: using the transformation unit to perform short-time Fourier transform on the original speech signal, to obtain the frequency-domain distribution information of the original speech signal; using the amplitude processing unit to perform a modulus operation on the frequency-domain distribution information, to obtain the original amplitude spectrum features.
[0211] Optionally, the storage medium is configured to store program code further for performing the following steps: the encoder is composed of multiple convolutional blocks, and the decoder is composed of multiple transposed convolutional blocks, and the outputs of the multiple convolutional blocks are respectively input into the corresponding transposed convolutional blocks among the multiple transposed convolutional blocks; using the encoder to perform dimensionality reduction processing on the original amplitude spectrum features, to obtain a first feature corresponding to the original amplitude spectrum features, including: using the multiple convolutional blocks to sequentially perform multiple convolutional processes on the original amplitude spectrum features, to obtain the first feature; using the decoder to perform reconstruction processing on the second feature, to obtain the target recognition result, including: using the multiple transposed convolutional blocks to sequentially perform multiple transposed convolutional processes on the second feature based on the outputs of the multiple convolutional blocks, to obtain the target recognition result.
[0212] Optionally, the storage medium is configured to store program code further for performing the following steps: The convolutional block and the transposed convolutional block have the same structure. The convolutional block includes: a first convolutional layer, a sigmoid function layer, a second convolutional layer, a product layer, a normalization layer, and an exponential linear unit layer. Among them, the first convolutional layer is used to perform convolutional processing on the input data of the convolutional block to obtain a first convolutional result. The second convolutional layer is used to perform convolutional processing on the input data to obtain a second convolutional result. The sigmoid function layer is used to map the first convolutional result to obtain a mapping result. The product layer is used to multiply the mapping result and the second convolutional result to obtain a product result. The normalization layer is used to perform normalization processing on the product result to obtain a normalized result. The exponential linear unit layer is used to perform exponential linear processing on the normalized result to obtain the output data of the convolutional block.
[0213] Optionally, the storage medium is configured to store program code further for performing the following steps: The gated recurrent module includes a plurality of gated recurrent units connected in sequence.
[0214] Optionally, the storage medium is configured to store program code further for performing the following steps: The above method further includes: obtaining training data, where the training data includes: a first speech signal and a second speech signal, and the second speech signal is used to represent the signal obtained after denoising the first speech signal; using an initial processing model to perform recognition on the first speech signal based on the time-domain distribution information of the first speech signal in the time domain and the frequency-domain distribution information of the first speech signal in the frequency domain to obtain a training recognition result; performing denoising processing on the first speech signal based on the training recognition result to obtain a predicted speech signal; determining a target loss function value corresponding to the initial processing model based on the second speech signal and the predicted speech signal; adjusting the model parameters of the initial processing model based on the target loss function value to obtain a speech signal model.
[0215] Optionally, the storage medium is configured to store program code further for performing the following steps: Determining a target loss function value corresponding to the initial processing model based on the second speech signal and the predicted speech signal includes: determining a first loss function value corresponding to the initial processing model based on the second speech signal and the predicted speech signal; extracting feature information from the frequency-domain distribution information of the second speech signal to obtain a first amplitude spectrum feature of the second speech signal, and extracting feature information from the frequency-domain distribution information of the predicted speech signal to obtain a second amplitude spectrum feature of the predicted speech signal; determining a second loss function value corresponding to the initial processing model based on the first amplitude spectrum feature and the second amplitude spectrum feature; performing a weighted sum on the first loss function value and the second loss function value to obtain the target loss function value.
[0216] Optionally, the storage medium is configured to store program code further for performing the following steps: extracting features from the frequency-domain distribution information of the second speech signal to obtain the first amplitude spectrum feature of the second speech signal, and extracting features from the frequency-domain distribution information of the predicted speech signal to obtain the second amplitude spectrum feature of the predicted speech signal, including: performing short-time Fourier transform on the second speech signal using a plurality of different transformation parameters to obtain the first frequency-domain distribution information of the second speech signal, and performing short-time Fourier transform on the predicted speech signal using a plurality of different transformation parameters to obtain the second frequency-domain distribution information of the predicted speech signal; performing a modulus operation on the first frequency-domain distribution information to obtain the first amplitude spectrum feature, and performing a modulus operation on the second frequency-domain distribution information to obtain the second amplitude spectrum feature.
[0217] Optionally, the storage medium is configured to store program code further for performing the following steps: determining a second loss function value corresponding to the initial processing model based on the first amplitude spectrum feature and the second amplitude spectrum feature, including: obtaining a spectral convergence loss of the first amplitude spectrum feature and the second amplitude spectrum feature to obtain a spectral convergence loss function value; obtaining a logarithmic short-time Fourier amplitude loss of the first amplitude spectrum feature and the second amplitude spectrum feature to obtain an amplitude loss function value; performing a weighted sum on the spectral convergence loss function value and the amplitude loss function value to obtain the second loss function value.
[0218] Optionally, the storage medium is configured to store program code further for performing the following steps: the above method further includes: outputting the original speech signal and the target recognition result; in response to receiving an adjustment instruction for adjusting the target recognition result, determining an adjusted recognition result corresponding to the adjustment instruction; performing noise reduction processing on the original speech signal based on the adjusted recognition result to obtain a target speech signal; and adjusting the speech processing model based on the adjusted recognition result and the original speech signal.
[0219] Optionally, the storage medium is configured to store program code further for performing the following steps: performing noise reduction processing on the original speech signal based on the target recognition result to obtain a target speech signal, including: extracting features from the frequency-domain distribution information of the original speech signal to obtain the original amplitude spectrum feature of the original speech signal; multiplying the original amplitude spectrum feature and the target recognition result element by element to obtain a target amplitude spectrum feature; and performing an inverse short-time Fourier transform on the target amplitude spectrum feature to obtain the target speech signal.
[0220] In the embodiments of the present application, the method includes obtaining an original voice signal; identifying the original voice signal based on the time-domain distribution information of the original voice signal in the time domain and the frequency-domain distribution information of the original voice signal in the frequency domain to obtain a target recognition result corresponding to the original voice signal, where the target recognition result is used to characterize whether a sub-signal in the original voice signal is a noise signal; and performing noise reduction processing on the original voice signal based on the target recognition result to obtain a target voice signal. By using the time-domain distribution information and frequency-domain distribution information of the original voice signal in the time domain and the frequency domain, the distribution characteristics of the original voice signal in the time domain and the frequency domain are reflected, and the time-domain distribution information and the frequency-domain distribution information are combined to determine the target recognition result of the original voice signal, making the judgment on whether a sub-signal in the original voice signal is a noise signal more accurate. Finally, noise reduction processing is performed on the original voice signal according to the target recognition result, which can avoid damaging the original voice signal during the noise reduction process, ensure that the target voice signal after noise reduction is closer to a pure voice signal, thereby ensuring the accuracy of the obtained target voice signal, improving the effect of noise reduction on the original voice signal, and further solving the technical problems in the related art that the voice processing method has limitations, resulting in a large degree of damage to the original voice signal after noise reduction or incomplete elimination of some special non-stationary noises.
[0221] The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments.
[0222] In the above embodiments of the present application, the descriptions of the respective embodiments have their own focuses. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0223] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces, and the indirect coupling or communication connection of units or modules may be in an electrical or other form.
[0224] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0225] In addition, in each embodiment of the present application, each functional unit can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0226] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs that can store program codes.
[0227] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. A speech processing method, It is characterized in that include: Obtaining the original speech signal; Based on the time domain distribution information of the original speech signal in the time domain and the frequency domain distribution information of the original speech signal in the frequency domain, the original speech signal is identified to obtain a target recognition result corresponding to the original speech signal, wherein the target recognition result is used to characterize whether the sub-signal in the original speech signal is a noise signal; The original speech signal is subjected to noise reduction processing based on the target recognition result to obtain a target speech signal.
2. The method according to claim 1, It is characterized in that Based on the time domain distribution information of the original voice signal in the time domain and the frequency domain distribution information of the original voice signal in the frequency domain, the original voice signal is recognized to obtain a target recognition result corresponding to the original voice signal, including: The original speech signal is recognized by using a speech processing model based on the time domain distribution information and the frequency domain distribution information to obtain the target recognition result.
3. The method according to claim 2, It is characterized in that The speech processing model includes: a frequency domain feature extraction module, an encoder, a gated loop module and a decoder. The speech processing model is used to recognize the original speech signal based on the time domain distribution information and the frequency domain distribution information to obtain the target recognition result, including: Using the frequency domain feature extraction module to extract features from the frequency domain distribution information of the original speech signal, to obtain original amplitude spectrum features of the original speech signal; Using the encoder to perform dimensionality reduction processing on the original amplitude spectrum feature to obtain a first feature corresponding to the original amplitude spectrum feature; Using the gated loop module to process the first feature based on the time domain distribution information to obtain a second feature; The decoder is used to reconstruct the second feature to obtain the target recognition result.
4. The method according to claim 3, It is characterized in that The frequency domain feature extraction module includes: a transformation unit and an amplitude processing unit. The frequency domain feature extraction module is used to extract the frequency domain distribution information of the original speech signal to obtain the original amplitude spectrum feature of the original speech signal, including: Using the transformation unit to perform short-time Fourier transform on the original speech signal to obtain frequency domain distribution information of the original speech signal; The amplitude processing unit is used to perform a modulo operation on the frequency domain distribution information to obtain the original amplitude spectrum feature.
5. The method according to claim 3, It is characterized in that The encoder comprises a plurality of convolution blocks, the decoder comprises a plurality of deconvolution blocks, and the outputs of the plurality of convolution blocks are respectively input to corresponding deconvolution blocks among the plurality of deconvolution blocks; Using the encoder to perform dimensionality reduction processing on the original amplitude spectrum feature to obtain a first feature corresponding to the original amplitude spectrum feature, including: using the multiple convolution blocks to sequentially perform multiple convolution processing on the original amplitude spectrum feature to obtain the first feature; The decoder is used to reconstruct the second feature to obtain the target recognition result, including: using the multiple deconvolution blocks to sequentially perform multiple deconvolution processes on the second feature based on the outputs of the multiple convolution blocks to obtain the target recognition result.
6. The method according to claim 5, It is characterized in that The convolution block and the deconvolution block have the same structure, and the convolution block includes: a first convolution layer, a sigmoid function layer, a second convolution layer, a convolution layer, a normalization layer and an exponential linear unit layer, wherein the first convolution layer is used to perform convolution processing on the input data of the convolution block to obtain a first convolution result, the second convolution layer is used to perform convolution processing on the input data to obtain a second convolution result, the sigmoid function layer is used to map the first convolution result to obtain a mapping result, the convolution layer is used to multiply the mapping result and the second convolution result to obtain a product result, the normalization layer is used to normalize the product result to obtain a normalized result, and the exponential linear unit layer is used to perform exponential linear processing on the normalized result to obtain output data of the convolution block.
7. The method according to claim 3, It is characterized in that The gated circulation module comprises a plurality of gated circulation units connected in sequence.
8. The method according to claim 2, It is characterized in that The method further comprises: Acquire training data, wherein the training data includes: a first speech signal and a second speech signal, wherein the second speech signal is used to represent a signal obtained after noise reduction processing is performed on the first speech signal; Using an initial processing model to recognize the first speech signal based on time domain distribution information of the first speech signal in the time domain and frequency domain distribution information of the first speech signal in the frequency domain, to obtain a training recognition result; Performing noise reduction processing on the first speech signal based on the training recognition result to obtain a predicted speech signal; Determining a target loss function value corresponding to the initial processing model based on the second speech signal and the predicted speech signal; The model parameters of the initial processing model are adjusted based on the target loss function value to obtain the speech signal model.
9. The method according to claim 8, It is characterized in that Determining a target loss function value corresponding to the initial processing model based on the second speech signal and the predicted speech signal includes: Determining a first loss function value corresponding to the initial processing model based on the second speech signal and the predicted speech signal; Performing feature extraction on the frequency domain distribution information of the second speech signal to obtain a first amplitude spectrum feature of the second speech signal, and performing feature extraction on the frequency domain distribution information of the predicted speech signal to obtain a second amplitude spectrum feature of the predicted speech signal; Determining a second loss function value corresponding to the initial processing model based on the first amplitude spectrum feature and the second amplitude spectrum feature; A weighted sum is performed on the first loss function value and the second loss function value to obtain the target loss function value.
10. The method according to claim 9, It is characterized in that Extracting features from the frequency domain distribution information of the second speech signal to obtain a first amplitude spectrum feature of the second speech signal, and extracting features from the frequency domain distribution information of the predicted speech signal to obtain a second amplitude spectrum feature of the predicted speech signal, including: Performing a short-time Fourier transform on the second speech signal using a plurality of different transform parameters to obtain first frequency domain distribution information of the second speech signal, and performing a short-time Fourier transform on the predicted speech signal using the plurality of different transform parameters to obtain second frequency domain distribution information of the predicted speech signal; A modulo operation is performed on the first frequency domain distribution information to obtain the first amplitude spectrum feature, and a modulo operation is performed on the second frequency domain distribution information to obtain the second amplitude spectrum feature.
11. The method according to claim 9, It is characterized in that Determining a second loss function value corresponding to the initial processing model based on the first amplitude spectrum feature and the second amplitude spectrum feature includes: Obtaining the spectrum convergence loss of the first amplitude spectrum feature and the second amplitude spectrum feature to obtain a spectrum convergence loss function value; Obtaining the logarithmic short-time Fourier amplitude loss of the first amplitude spectrum feature and the second amplitude spectrum feature to obtain an amplitude loss function value; A weighted sum is performed on the spectral convergence loss function value and the amplitude loss function value to obtain the second loss function value.
12. The method according to claim 2, It is characterized in that The method further comprises: Outputting the original speech signal and the target recognition result; In response to receiving an adjustment instruction for adjusting the target recognition result, determining an adjusted recognition result corresponding to the adjustment instruction; Performing noise reduction processing on the original speech signal based on the adjusted recognition result to obtain the target speech signal; The speech processing model is adjusted based on the adjusted recognition result and the original speech signal.
13. The method according to claim 1, It is characterized in that The method further comprises: performing noise reduction processing on the original speech signal based on the target recognition result to obtain a target speech signal, including: Extracting features of the frequency domain distribution information of the original speech signal to obtain original amplitude spectrum features of the original speech signal; Multiplying the original amplitude spectrum feature and the target recognition result element by element to obtain the target amplitude spectrum feature; Performing short-time inverse Fourier transform on the target amplitude spectrum feature to obtain the target speech signal.
14. A method for speech processing, It is characterized in that include: Get the original voice signal in the live video; Based on the time domain distribution information of the original speech signal in the time domain and the frequency domain distribution information of the original speech signal in the frequency domain, the original speech signal is identified to obtain a target recognition result corresponding to the original speech signal, wherein the target recognition result is used to characterize whether the sub-signal in the original speech signal is a noise signal; Performing noise reduction processing on the original speech signal based on the target recognition result to obtain a target speech signal; The original voice signal in the live video is replaced with the target voice signal.
15. A method for speech processing, It is characterized in that include: In response to an input instruction acting on the operation interface, a waveform diagram of the original speech signal is displayed on the operation interface; In response to a noise reduction instruction applied on the operation interface, a waveform of a target voice signal is displayed on the operation interface, wherein the target voice signal is obtained by performing noise reduction processing on the original voice signal based on a target recognition result corresponding to the original voice signal, and the target recognition result is obtained by identifying the original voice signal based on time domain distribution information of the original voice signal in the time domain and frequency domain distribution information of the original voice signal in the frequency domain, and the target recognition result is used to characterize whether a sub-signal in the original voice signal is a noise signal.
16. A method for speech processing, It is characterized in that include: Acquire an original voice signal by calling a first interface, wherein the first interface includes a first parameter, and a parameter value of the first parameter includes the original voice signal; Based on the time domain distribution information of the original speech signal in the time domain and the frequency domain distribution information of the original speech signal in the frequency domain, the original speech signal is identified to obtain a target recognition result corresponding to the original speech signal, wherein the target recognition result is used to characterize whether the sub-signal in the original speech signal is a noise signal; Performing noise reduction processing on the original speech signal based on the target recognition result to obtain a target speech signal; The target voice signal is output by calling a second interface, wherein the second interface includes a second parameter, and a parameter value of the second parameter includes the target voice signal.
17. A computer-readable storage medium, It is characterized in that The computer-readable storage medium includes a stored executable program, wherein when the executable program is executed, the device where the computer-readable storage medium is located is controlled to execute the method according to any one of claims 1 to 16.
18. An electronic device, It is characterized in that include: A memory storing an executable program; A processor, configured to run the program, wherein the program executes the method according to any one of claims 1 to 16 when running.