Training Method, Speech Scoring Method, Device and Medium for Speech Noise Reduction Model
By introducing pronunciation difference processing layer and content difference processing layer in the speech noise reduction model, the model parameters are updated to improve the noise reduction accuracy, and the problem of speech information loss in the noise reduction processing in the prior art is solved, achieving higher noise reduction accuracy.
Patent Information
- Application Number
- CN202111025632.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-02
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-09-02
AI Technical Summary
The existing voice noise reduction model is prone to losing some voice information during the noise reduction process, resulting in low noise reduction accuracy.
The pronunciation difference processing layer and the content difference processing layer are introduced in the speech noise reduction model. The model parameters are updated to improve the noise reduction accuracy through pronunciation score prediction and content difference analysis.
Through model training based on pronunciation similarity and content differences, the loss of voice information before and after noise reduction is avoided, and the accuracy of noise reduction processing is improved.
Smart Images

Figure CN114283828B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a method for training a voice noise reduction model, a voice scoring method, a device, an electronic device, and a storage medium. Background Art
[0002] Artificial Intelligence (AI) is to use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, and is a theory, method, technology, and application system for perceiving the environment, acquiring knowledge, and using knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.
[0003] Artificial intelligence has been increasingly applied to voice processing. In related technologies, the learning objective of a voice noise reduction model is usually to make the waveform of the denoised voice most similar to the waveform of the pure voice. When learning with the closest waveform as the goal, usually only the voice with a large waveform amplitude can be concerned, while the voice with a small amplitude is directly ignored, resulting in the loss of some voice information during the noise reduction process and low noise reduction accuracy. Summary of the Invention
[0004] Embodiments of this application provide a method, a device, an electronic device, and a storage medium for training a voice noise reduction model, which can improve the noise reduction accuracy of the voice noise reduction model.
[0005] The technical solution of the embodiments of this application is implemented as follows:
[0006] Embodiments of this application provide a method for training a voice noise reduction model. The voice noise reduction model includes: a noise processing layer, a pronunciation difference processing layer, and a content difference processing layer. The method includes:
[0007] Through the noise processing layer, perform noise reduction processing on the voice sample to obtain a target voice sample;
[0008] Through the pronunciation difference processing layer, predict the pronunciation score of the target voice sample to obtain a pronunciation prediction result, where the pronunciation prediction result is used to indicate the pronunciation similarity between the target voice sample and the reference pronunciation corresponding to the voice sample;
[0009] Through the content difference processing layer, determine the content difference between the content of the target voice sample and the content of the voice sample;
[0010] Update the model parameters of the voice noise reduction model based on the pronunciation prediction result and the content difference to obtain a trained voice noise reduction model.
[0011] In the above solution, the pronunciation difference processing layer includes: a pronunciation score loss processing layer;
[0012] The updating of the model parameters of the voice noise reduction model based on the pronunciation prediction result and the content difference includes:
[0013] Determine the difference between the pronunciation prediction result and the sample label corresponding to the voice sample through the pronunciation score loss processing layer, and determine the value of the score loss function based on the difference;
[0014] Update the model parameters of the voice noise reduction model based on the content difference and the value of the score loss function.
[0015] In the above solution, the updating of the model parameters of the voice noise reduction model based on the content difference and the value of the score loss function includes:
[0016] Obtain the first weight value corresponding to the content difference and the second weight value corresponding to the value of the score loss function;
[0017] Combine the first weight value and the second weight value, and determine the value of the loss function of the voice noise reduction model based on the content difference and the value of the score loss function;
[0018] Update the model parameters of the voice noise reduction model based on the value of the loss function.
[0019] In the above solution, the updating of the model parameters of the voice noise reduction model based on the value of the loss function includes:
[0020] When the value of the loss function exceeds the loss threshold, determine the error signal of the voice noise reduction model based on the loss function;
[0021] Backpropagate the error signal in the voice noise reduction model, and update the model parameters of each layer in the voice noise reduction model during the propagation process.
[0022] An embodiment of the present application also provides a voice scoring method, which is applied to a voice noise reduction model. The method includes:
[0023] Present a reference voice text and a voice input function item;
[0024] In response to a trigger operation on the voice input function item, present a voice input interface and present a voice end function item in the voice input interface;
[0025] Receive voice information input based on the voice input interface;
[0026] In response to a trigger operation for the voice end function item, present a pronunciation score indicating the pronunciation similarity between the voice information and the reference pronunciation corresponding to the reference voice text;
[0027] Wherein, the pronunciation score is obtained by predicting the pronunciation score of the target voice information, and the target voice information is obtained by denoising the voice information based on the voice denoising model;
[0028] Wherein, the voice denoising model is trained based on the training method of the above voice denoising model.
[0029] An embodiment of the present application further provides a training device for a voice denoising model. The voice denoising model includes: a noise processing layer, a pronunciation difference processing layer, and a content difference processing layer. The device includes:
[0030] A denoising module for denoising the voice sample through the noise processing layer to obtain a target voice sample;
[0031] A prediction module for predicting the pronunciation score of the target voice sample through the pronunciation difference processing layer to obtain a pronunciation prediction result, where the pronunciation prediction result is used to indicate the pronunciation similarity between the target voice sample and the reference pronunciation corresponding to the voice sample;
[0032] A determination module for determining the content difference between the content of the target voice sample and the content of the voice sample through the content difference processing layer;
[0033] An update module for updating the model parameters of the voice denoising model based on the pronunciation prediction result and the content difference to obtain a trained voice denoising model.
[0034] In the above solution, the noise processing layer includes: a first feature transformation layer, a filtering processing layer, and a second feature transformation layer;
[0035] The denoising module is further configured to perform Fourier transform on the voice sample through the first feature transformation layer to obtain the amplitude spectrum and phase spectrum corresponding to the voice sample;
[0036] Perform filtering processing on the amplitude spectrum through the filtering processing layer to obtain a target amplitude spectrum, and perform phase correction on the phase spectrum to obtain a target phase spectrum;
[0037] Through the second feature transformation layer, multiply the target magnitude spectrum and the target phase spectrum, and perform an inverse Fourier transform on the result of the multiplication to obtain the target speech sample.
[0038] In the above solution, the filtering processing layer includes at least two cascaded sub-filtering processing layers;
[0039] The noise reduction module is further configured to filter the magnitude spectrum through the first-level sub-filtering processing layer to obtain an intermediate magnitude spectrum, and perform phase correction on the phase spectrum to obtain an intermediate phase spectrum;
[0040] Filter the intermediate magnitude spectrum through the non-first-level sub-filtering processing layer to obtain the target magnitude spectrum, and perform phase correction on the intermediate phase spectrum to obtain the target phase spectrum.
[0041] In the above solution, each of the sub-filtering processing layers includes a phase spectrum correction layer and at least two cascaded magnitude spectrum filtering layers;
[0042] The noise reduction module is further configured to filter the magnitude spectrum through the at least two cascaded magnitude spectrum filtering layers to obtain an intermediate magnitude spectrum;
[0043] Perform phase correction on the phase spectrum based on the intermediate magnitude spectrum through the phase spectrum correction layer to obtain an intermediate phase spectrum.
[0044] In the above solution, the second feature transformation layer includes a feature conversion layer and a feature inverse transformation layer;
[0045] The noise reduction module is further configured to convert the target magnitude spectrum into a magnitude spectrum mask through the feature conversion layer, and determine the phase angle corresponding to the target phase spectrum;
[0046] Multiply the target magnitude spectrum, the magnitude spectrum mask, and the phase angle corresponding to the target phase spectrum through the feature inverse transformation layer, and perform an inverse Fourier transform on the result of the multiplication to obtain the target speech sample.
[0047] In the above solution, the content difference processing layer includes: a Fourier transform layer;
[0048] The determination module is further configured to perform a Fourier transform on the target speech sample through the Fourier transform layer to obtain a first magnitude spectrum, and perform a Fourier transform on the speech sample to obtain a second magnitude spectrum;
[0049] Determine the magnitude difference between the first magnitude spectrum and the second magnitude spectrum, and determine the magnitude difference as the content difference between the content of the target speech sample and the content of the speech sample.
[0050] In the above solution, the Fourier transform layer includes at least two sub-Fourier transform layers, and different sub-Fourier transform layers correspond to different transformation scales;
[0051] The determining module is further configured to perform Fourier transforms with corresponding transformation scales on the target voice sample through each of the sub-Fourier transform layers, to obtain first amplitude spectra corresponding to the sub-Fourier transform layers;
[0052] Perform Fourier transforms with corresponding transformation scales on the voice sample through each of the sub-Fourier transform layers, to obtain second amplitude spectra corresponding to the sub-Fourier transform layers;
[0053] The determining module is further configured to determine an intermediate amplitude difference between the first amplitude spectra and the second amplitude spectra corresponding to the sub-Fourier transform layers;
[0054] Perform a summation and averaging process on the intermediate amplitude differences corresponding to the at least two sub-Fourier transform layers, to obtain an average amplitude difference, and use the average amplitude difference as the amplitude difference.
[0055] In the above solution, the content difference processing layer further includes: a power compression processing layer;
[0056] The determining module is further configured to perform compression processing on the first amplitude spectrum through the power compression processing layer, to obtain a first compressed amplitude spectrum, and perform compression processing on the second amplitude spectrum, to obtain a second compressed amplitude spectrum;
[0057] Determine a compressed amplitude difference between the first compressed amplitude spectrum and the second compressed amplitude spectrum, and use the compressed amplitude difference as the amplitude difference.
[0058] In the above solution, the pronunciation difference processing layer includes: a pronunciation scoring loss processing layer;
[0059] The updating module is further configured to determine a difference between the pronunciation prediction result and the sample label corresponding to the voice sample through the pronunciation scoring loss processing layer, and determine a value of a scoring loss function based on the difference;
[0060] Update the model parameters of the voice noise reduction model based on the content difference and the value of the scoring loss function.
[0061] In the above solution, the updating module is further configured to obtain a first weight value corresponding to the content difference, and a second weight value corresponding to the value of the scoring loss function;
[0062] Based on the first weight value and the second weight value, determine the value of the loss function of the voice noise reduction model based on the content difference and the value of the scoring loss function;
[0063] Based on the value of the loss function, update the model parameters of the voice noise reduction model.
[0064] In the above solution, the update module is further configured to, when the value of the loss function exceeds the loss threshold, determine an error signal of the voice noise reduction model based on the loss function;
[0065] Backpropagate the error signal in the voice noise reduction model, and update the model parameters of each layer in the voice noise reduction model during the propagation process.
[0066] In the above solution, the pronunciation difference processing layer further includes: a first feature mapping layer, a second feature mapping layer, and a feature splicing and prediction layer, and the network structure of the first feature mapping layer is different from the network structure of the second feature mapping layer;
[0067] The prediction module is further configured to perform a mapping process on the target voice sample through the first feature mapping layer to obtain a first mapped feature;
[0068] Perform a mapping process on the target voice sample through the second feature mapping layer to obtain a second mapped feature;
[0069] Perform a splicing process on the first mapped feature and the second mapped feature through the feature splicing and prediction layer to obtain a spliced feature, and
[0070] Predict the pronunciation score of the spliced feature to obtain the pronunciation prediction result.
[0071] An embodiment of the present application further provides a voice scoring device, which is applied to a voice noise reduction model. The device includes:
[0072] A first presentation module, configured to present a reference voice text and a voice input function item;
[0073] A second presentation module, configured to, in response to a trigger operation on the voice input function item, present a voice input interface and present a voice end function item in the voice input interface;
[0074] A receiving module, configured to receive voice information input based on the voice input interface;
[0075] A third presentation module, configured to, in response to a trigger operation on the voice end function item, present a pronunciation score indicating the pronunciation similarity between the voice information and the reference pronunciation corresponding to the reference voice text;
[0076] Among them, the pronunciation score is obtained based on the prediction of the pronunciation score of the target speech information, and the target speech information is obtained by performing noise reduction processing on the speech information based on the speech noise reduction model;
[0077] Among them, the speech noise reduction model is trained based on the training method of the above speech noise reduction model.
[0078] An embodiment of the present application further provides an electronic device, including:
[0079] A memory for storing executable instructions;
[0080] A processor, when executing the executable instructions stored in the memory, implements the method provided by the embodiment of the present application.
[0081] An embodiment of the present application further provides a computer-readable storage medium, storing executable instructions, and when the executable instructions are executed by a processor, the method provided by the embodiment of the present application is implemented.
[0082] The embodiment of the present application has the following beneficial effects:
[0083] Applying the embodiment of the present application, a pronunciation difference processing layer and a content difference processing layer are added to the speech noise reduction model. Through the pronunciation difference processing layer, the pronunciation score of the target speech sample after noise reduction processing is predicted to obtain a pronunciation prediction result indicating the pronunciation similarity between the target speech sample and the reference pronunciation corresponding to the speech sample, and the content difference between the content of the target speech sample and the content of the speech sample is determined through the content difference processing layer. Then, based on the pronunciation prediction result and the content difference, the model parameters of the speech noise reduction model are updated to complete model training; thus, training the speech noise reduction model based on the pronunciation similarity and content difference before and after noise reduction can enable the trained speech noise reduction model to avoid the loss of speech information before and after noise reduction and improve the accuracy of noise reduction processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0084] Figure 1 is a schematic structural diagram of a training system 100 of a speech noise reduction model provided by an embodiment of the present application;
[0085] Figure 2 is a schematic structural diagram of an electronic device 500 for implementing the training method of a speech noise reduction model provided by an embodiment of the present application;
[0086] Figure 3 is a schematic flowchart of the training method of a speech noise reduction model provided by an embodiment of the present application;
[0087] Figure 4 is a schematic structural diagram of a speech noise reduction model provided by an embodiment of the present application;
[0088] Figure 5 It is a schematic structural diagram of the noise processing layer provided by an embodiment of the present application;
[0089] Figure 6 It is a schematic structural diagram of the first feature transformation layer provided by an embodiment of the present application;
[0090] Figure 7 It is a schematic structural diagram of the filtering processing layer provided by an embodiment of the present application;
[0091] Figure 8 It is a schematic structural diagram of the sub-filtering processing layer provided by an embodiment of the present application;
[0092] Figure 9 It is a schematic structural diagram of the second feature transformation layer provided by an embodiment of the present application;
[0093] Figure 10 It is a schematic structural diagram of the content difference processing layer provided by an embodiment of the present application;
[0094] Figure 11 It is a schematic structural diagram of the pronunciation difference processing layer adopted by an embodiment of the present application;
[0095] Figure 12 It is a schematic flow diagram of the speech scoring method provided by an embodiment of the present application;
[0096] Figure 13 It is a schematic presentation diagram of the speech scoring process provided by an embodiment of the present application;
[0097] Figure 14 It is a schematic flow diagram of the speech scoring method based on a speech noise reduction model provided by an embodiment of the present application. Detailed implementation manners
[0098] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be construed as limitations on the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.
[0099] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0100] In the following description, the terms "first", "second", and "third" are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first", "second", and "third" can be interchanged in a specific order or sequence when permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0101] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0102] Before further elaborating on the embodiments of this application, the nouns and terms involved in the embodiments of this application are described. The nouns and terms involved in the embodiments of this application are subject to the following explanations.
[0103] 1) Client: An application program running on a terminal for providing various services, such as an instant messaging client, a video playback client.
[0104] 2) In response to: Used to represent the conditions or states on which the executed operations depend. When the dependent conditions or states are met, one or more executed operations can be real-time or can have a set delay; without special instructions, there is no restriction on the execution order of multiple executed operations.
[0105] Based on the above explanations of the nouns and terms involved in the embodiments of this application, the training system of the voice noise reduction model provided by the embodiments of this application is described below. Refer to Figure 1 , Figure 1 FIG. is a schematic diagram of the architecture of the training system 100 of the voice noise reduction model provided by the embodiments of this application. To support an exemplary application, the terminal 400 is connected to the server 200 through the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two, and uses a wireless or wired link to implement data transmission.
[0106] The terminal 400 is configured to send a training request for the corresponding voice noise reduction model to the server 200 in response to a training instruction for the voice noise reduction model; the voice noise reduction model includes: a noise processing layer, a pronunciation difference processing layer, and a content difference processing layer;
[0107] The server 200 is configured to receive and respond to a training request, perform noise reduction processing on a speech sample through a noise processing layer to obtain a target speech sample; perform prediction of a pronunciation score on the target speech sample through a pronunciation difference processing layer to obtain a pronunciation prediction result, where the pronunciation prediction result is used to indicate the pronunciation similarity between the target speech sample and the reference pronunciation corresponding to the speech sample; determine the content difference between the content of the target speech sample and the content of the speech sample through a content difference processing layer; update the model parameters of the speech noise reduction model based on the pronunciation prediction result and the content difference to obtain a trained speech noise reduction model; and return the trained speech noise reduction model to the terminal 400.
[0108] The terminal 400 is configured to receive the trained speech noise reduction model and perform speech noise reduction processing on the input speech information based on the speech noise reduction model, thereby improving the accuracy of speech noise reduction and avoiding loss of some speech information during the noise reduction process.
[0109] In practical applications, the server 200 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal 400 can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart TV, a smart watch, etc., but is not limited thereto. The terminal 400 and the server 200 can be directly or indirectly connected through wired or wireless communication methods, and this application does not make any restrictions here.
[0110] See Figure 2 , Figure 2 is a schematic structural diagram of an electronic device 500 for implementing the training method of the speech noise reduction model provided by an embodiment of this application. In practical applications, the electronic device 500 can be Figure 1 the server or terminal shown in, taking the electronic device 500 as Figure 1 the terminal shown in as an example, the electronic device for implementing the training method of the speech noise reduction model according to the embodiment of this application is described. The electronic device 500 provided by the embodiment of this application includes: at least one processor 510, a memory 550, at least one network interface 520, and a user interface 530. Each component in the electronic device 500 is coupled together through a bus system 540. It can be understood that the bus system 540 is used to implement the connection and communication between these components. In addition to the data bus, the bus system 540 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, in Figure 2 all kinds of buses are labeled as the bus system 540.
[0111] The processor 510 may be an integrated circuit chip with the ability to process signals, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or any conventional processor, etc.
[0112] The user interface 530 includes one or more output devices 531 that enable the presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 530 also includes one or more input devices 532, including user interface components that facilitate user input, such as a keyboard, a mouse, a microphone, a touch screen display, a camera, other input buttons, and controls.
[0113] The memory 550 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disc drives, etc. The memory 550 optionally includes one or more storage devices that are physically remote from the processor 510.
[0114] The memory 550 includes volatile memory or non-volatile memory, and may also include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 550 described in the embodiments of the present application is intended to include any suitable type of memory.
[0115] In some embodiments, the memory 550 is capable of storing data to support various operations. Examples of such data include programs, modules, and data structures, or subsets or supersets thereof, which are described below by way of example.
[0116] The operating system 551 includes system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks;
[0117] The network communication module 552 is used to reach other computing devices via one or more (wired or wireless) network interfaces 520. Exemplary network interfaces 520 include: Bluetooth, Wi-Fi (Wireless Fidelity), and USB (Universal Serial Bus), etc.;
[0118] A presentation module 553 for enabling the presentation of information (e.g., a user interface for operating peripheral devices and displaying content and information) via one or more output devices 531 associated with the user interface 530 (e.g., a display screen, a speaker, etc.);
[0119] An input processing module 554 for detecting and translating one or more user inputs or interactions from one of one or more input devices 532.
[0120] In some embodiments, the training device of the voice noise reduction model provided by the embodiments of the present application may be implemented in software. Figure 2 The training device 555 of the voice noise reduction model stored in the memory 550 is shown, which may be software in the form of a program and a plug-in, etc., and includes the following software modules: a noise reduction module 5551, a prediction module 5552, a determination module 5553, and an update module 5554. These modules are logical, so they can be combined arbitrarily or further split according to the implemented functions. The functions of each module will be described below.
[0121] In other embodiments, the training device of the voice noise reduction model provided by the embodiments of the present application may be implemented in a combination of software and hardware. As an example, the training device of the voice noise reduction model provided by the embodiments of the present application may be a processor in the form of a hardware decoding processor, which is programmed to execute the training method of the voice noise reduction model provided by the embodiments of the present application. For example, the processor in the form of a hardware decoding processor may employ one or more application specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), or other electronic components.
[0122] Based on the above description of the training system and electronic device of the voice noise reduction model provided by the embodiments of the present application, the training method of the voice noise reduction model provided by the embodiments of the present application will be described below. In some embodiments, the training method of the voice noise reduction model provided by the embodiments of the present application may be implemented independently by a server or a terminal, or implemented jointly by a server and a terminal. Below, the training method of the voice noise reduction model provided by the embodiments of the present application will be described by taking the server implementation as an example.
[0123] See Figure 3 , Figure 3It is a schematic flowchart of a method for training a voice noise reduction model provided by an embodiment of the present application. The voice noise reduction model provided by the embodiment of the present application includes: a noise processing layer, a pronunciation difference processing layer, and a content difference processing layer. The method for training the voice noise reduction model provided by the embodiment of the present application includes:
[0124] Step 101: The server performs noise reduction processing on the voice sample through the noise processing layer to obtain a target voice sample.
[0125] Here, the voice noise reduction model includes a noise processing layer, a pronunciation difference processing layer, and a content difference processing layer, and is used to perform voice noise reduction processing on the input voice information. As an example, see Figure 4 , Figure 4 It is a schematic structural diagram of the voice noise reduction model provided by the embodiment of the present application. Here, the voice noise reduction model includes a noise processing layer 410 (i.e., the voice enhancement network EnhanceNet), a pronunciation difference processing layer 420 (i.e., the pronunciation error prediction network PronNet), and a content difference processing layer 430 (i.e., the multi-scale voice similarity measurement network SimilarNet).
[0126] In practical applications, the voice noise reduction model can be constructed based on a machine learning network, such as a convolutional neural network, a deep neural network, etc.; after the initial voice noise reduction model is constructed based on the machine learning network, the voice noise reduction model contains initial model parameters. To improve the noise reduction accuracy of the voice noise reduction model, it is necessary to train the voice noise reduction model to update the model parameters of the voice noise reduction model during the model training process, and obtain a trained voice noise reduction model, so as to perform noise reduction processing on voice information based on the trained voice noise reduction model.
[0127] During the training of the voice noise reduction model, first, training samples for training, that is, voice samples, are obtained. The voice samples can be for certain reference voice texts, and the reference voice texts correspond to corresponding reference pronunciations. After the server obtains the voice samples for training the voice noise reduction model, it performs noise reduction processing on the voice samples through the noise processing layer of the voice noise reduction model, such as filter noise reduction processing, etc., to obtain target voice samples.
[0128] In some embodiments, the noise processing layer includes: a first feature transformation layer, a filtering processing layer, and a second feature transformation layer; the server can perform noise reduction processing on the speech sample through the noise processing layer in the following manner to obtain the target speech sample: through the first feature transformation layer, perform Fourier transform on the speech sample to obtain the amplitude spectrum and phase spectrum corresponding to the speech sample; through the filtering processing layer, perform filtering processing on the amplitude spectrum to obtain the target amplitude spectrum, and perform phase correction on the phase spectrum to obtain the target phase spectrum; through the second feature transformation layer, multiply the target amplitude spectrum and the target phase spectrum, and perform inverse Fourier transform on the result of the multiplication to obtain the target speech sample.
[0129] Here, the above-mentioned noise processing layer includes a first feature transformation layer, a filtering processing layer, and a second feature transformation layer. As an example, refer to Figure 5 , Figure 5 which is a schematic structural diagram of the noise processing layer provided by an embodiment of the present application. Here, the noise processing layer 410 is the speech enhancement network EnhanceNet, including a first feature transformation layer 501 (i.e., the pre-processing network PrevNet), a filtering processing layer 502 (i.e., the cascaded activation network CasNet), and a second feature transformation layer 503 (i.e., the post-processing network PostNet). In practical applications, first, through the first feature transformation layer, perform Fourier transform on the waveform features of the speech sample to obtain the corresponding amplitude spectrum A and phase spectrum P; then, through the filtering processing layer, perform filtering processing on the amplitude spectrum A to obtain the amplitude spectrum A' (i.e., the target amplitude spectrum), and at the same time, through this filtering processing layer, perform phase correction on the phase spectrum P based on the filtered amplitude spectrum A' to obtain the phase spectrum P' (i.e., the target phase spectrum); finally, through the second feature transformation layer, perform inverse Fourier transform processing on the amplitude spectrum A' and the phase spectrum P' to output the transformed waveform, i.e., the target speech sample.
[0130] Next, the processing process of the noise reduction processing layer will be described in detail. First, when the server performs noise reduction processing on the speech sample through the noise processing layer, it first performs Fourier transform on the speech sample through the first feature transformation layer, specifically, performs Fourier transform on the waveform features of the speech sample to obtain the amplitude spectrum and phase spectrum corresponding to the speech sample. As an example, refer to Figure 6 , Figure 6 which is a schematic structural diagram of the first feature transformation layer provided by an embodiment of the present application. Here, the first feature transformation layer 501 is the Figure 5 pre-processing network PrevNet shown in, including a Fourier transform layer 610 and a convolutional layer 620. Through the Fourier transform layer, using the short-time Fourier transform, convert the waveform features of the speech sample into a 2-channel Fourier spectrum (including an amplitude spectrum and a phase spectrum), and further, through the convolutional layer 620, convert it from the 2-channel Fourier spectrum into a 64-channel amplitude spectrum A and a 64-channel phase spectrum P.
[0131] Second, the server then performs filtering processing (i.e., noise reduction processing) on the amplitude spectrum through the filtering processing layer, such as convolutional filtering processing, to obtain the filtered target amplitude spectrum; at the same time, through this filtering processing layer, the phase spectrum is phase-corrected based on the filtered target amplitude spectrum to obtain the target phase spectrum.
[0132] In some embodiments, the filtering processing layer includes at least two cascaded sub-filtering processing layers; the server can perform filtering processing on the amplitude spectrum through the filtering processing layer in the following manner to obtain the target amplitude spectrum and perform phase correction on the phase spectrum to obtain the target phase spectrum: through the first-level sub-filtering processing layer, perform filtering processing on the amplitude spectrum to obtain an intermediate amplitude spectrum and perform phase correction on the phase spectrum to obtain an intermediate phase spectrum; through the non-first-level sub-filtering processing layer, perform filtering processing on the intermediate amplitude spectrum to obtain the target amplitude spectrum and perform phase correction on the intermediate phase spectrum to obtain the target phase spectrum.
[0133] In practical applications, the filtering processing layer includes at least two cascaded sub-filtering processing layers. The server can perform filtering processing on the amplitude spectrum through the first-level sub-filtering processing layer to obtain an intermediate amplitude spectrum and perform phase correction on the phase spectrum to obtain an intermediate phase spectrum; then through the non-first-level sub-filtering processing layer, perform filtering processing on the intermediate amplitude spectrum to obtain the target amplitude spectrum and perform phase correction on the intermediate phase spectrum to obtain the target phase spectrum. Specifically, through the non-first-level sub-filtering processing layer, perform filtering processing on the intermediate amplitude spectrum output by the previous level and perform phase correction on the intermediate phase spectrum output by the previous level, and loop until the processing of the last-level sub-filtering processing layer is completed. The intermediate amplitude spectrum output by the last-level sub-filtering processing layer is used as the target amplitude spectrum, and the intermediate phase spectrum output by the last-level sub-filtering processing layer is used as the target phase spectrum.
[0134] As an example, refer to Figure 7 , Figure 7 which is a schematic structural diagram of the filtering processing layer provided by the embodiments of the present application. Here, the filtering processing layer 502 includes multiple sub-filtering processing layers, and the sub-filtering processing layer is composed of a third-order activation attention network TAB. The amplitude spectrum A and the phase spectrum P output by the first feature transformation layer 501 are filtered to output a 64-channel amplitude spectrum A' (i.e., the target amplitude spectrum) and a phase spectrum P' (i.e., the target phase spectrum).
[0135] In some embodiments, each sub-filter processing layer includes a phase spectrum correction layer and at least two cascaded amplitude spectrum filtering layers; the server can filter the amplitude spectrum through the first-level sub-filter processing layer in the following manner to obtain an intermediate amplitude spectrum and correct the phase of the phase spectrum to obtain an intermediate phase spectrum: filter the amplitude spectrum through at least two cascaded amplitude spectrum filtering layers to obtain an intermediate amplitude spectrum; based on the intermediate amplitude spectrum, correct the phase of the phase spectrum through the phase spectrum correction layer to obtain an intermediate phase spectrum.
[0136] Here, each of the above-mentioned sub-filter processing layers is composed of a phase spectrum correction layer and at least two cascaded amplitude spectrum filtering layers. The server can first filter the amplitude spectrum through at least two cascaded amplitude spectrum filtering layers, such as harmonic filtering, to obtain an intermediate amplitude spectrum; then, through the phase spectrum correction layer, correct the phase of the phase spectrum based on the intermediate amplitude spectrum to obtain an intermediate phase spectrum. In practical applications, the relationship between the intermediate amplitude spectrum and the intermediate phase spectrum is:
[0137]
[0138] where Conv is a convolution operation; Tanh is a hyperbolic tangent function operation (converting the input value to a value between -1 and 1); represents dot product, represents concatenation, A’ is the intermediate amplitude spectrum, P is the phase spectrum, and P’ is the intermediate phase spectrum.
[0139] As an example, refer to Figure 8 Figure 8 which is a schematic structural diagram of the sub-filter processing layer provided by an embodiment of the present application. Here, the sub-filter processing layer includes an amplitude spectrum filtering network 810 (i.e., a third-order amplitude spectrum enhancement network AmpNet) and 1 phase spectrum correction layer 820 (i.e., a first-order phase spectrum correction network PhaseNet), as shown in Figure 8 in FIG. A, for filtering the amplitude spectrum A to obtain an intermediate amplitude spectrum A’; the amplitude spectrum filtering network 810 includes a plurality of cascaded amplitude spectrum filtering layers, as shown in Figure 8 in FIG. B, which are 3 cascaded amplitude spectrum filtering layers (i.e., harmonic enhancers H); among them, the structure of each amplitude spectrum filtering layer is as shown in Figure 8 in FIG. C, including two linear processing layers Linear-F and two convolutional layers Conv1*1, for performing harmonic filtering on the amplitude spectrum.
[0140] Thirdly, finally, through the second feature transformation layer, multiply the target amplitude spectrum and the target phase spectrum. In practical applications, it can be to calculate the dot product of the target amplitude spectrum and the target phase spectrum, and then perform an inverse Fourier transform on the result of the dot product to obtain the target speech sample.
[0141] In some embodiments, the second feature transformation layer includes a feature conversion layer and a feature inverse transformation layer; the server can multiply the target amplitude spectrum and the target phase spectrum through the second feature transformation layer in the following manner, and perform an inverse Fourier transform on the result of the multiplication to obtain the target speech sample: convert the target amplitude spectrum into an amplitude spectrum mask through the feature conversion layer, and determine the phase angle corresponding to the target phase spectrum; multiply the target amplitude spectrum, the amplitude spectrum mask, and the phase angle corresponding to the target phase spectrum through the feature inverse transformation layer, and perform an inverse Fourier transform on the result of the multiplication to obtain the target speech sample.
[0142] In practical applications, the second feature transformation layer includes a feature conversion layer and a feature inverse transformation layer. Specifically, the server can convert the target amplitude spectrum into an amplitude spectrum mask through the feature conversion layer, and determine the phase angle corresponding to the target phase spectrum; multiply the target amplitude spectrum, the amplitude spectrum mask, and the phase angle corresponding to the target phase spectrum through the feature inverse transformation layer, and perform an inverse Fourier transform on the result of the multiplication to obtain the target speech sample.
[0143] As an example, refer to Figure 9 , Figure 9 is a schematic structural diagram of the second feature transformation layer provided by an embodiment of the present application. Here, the second feature transformation layer 503 includes a feature conversion layer, which is composed of multiple convolutional layers; it also includes a feature inverse transformation layer. Convert the target amplitude spectrum (i.e., amplitude spectrum A') output by the filtering processing layer 502 into an amplitude spectrum mask M, convert the target phase spectrum (i.e., phase spectrum P') into a phase angle Ω, and then convert it into a waveform output through an inverse Fourier transform, that is, obtain the denoised target speech sample. Specifically, perform a dot product calculation on the dot product result of the target amplitude spectrum and the amplitude spectrum mask, and the phase angle Ω, and perform an inverse short-time Fourier transform (iSTFT) on the obtained result to convert it into a waveform output, that is, obtain the denoised target speech sample.
[0144] Step 102: Predict the pronunciation score of the target speech sample through the pronunciation difference processing layer to obtain a pronunciation prediction result.
[0145] Among them, the pronunciation prediction result is used to indicate the pronunciation similarity between the target speech sample and the reference pronunciation corresponding to the speech sample.
[0146] Here, the target speech sample is the speech sample after noise reduction processing. Predict the pronunciation score of the target speech sample through the pronunciation difference processing layer to obtain a pronunciation prediction result, that is, predict the pronunciation score. The pronunciation prediction result is used to indicate the pronunciation similarity between the target speech sample and the reference pronunciation corresponding to the speech sample.
[0147] In some embodiments, the pronunciation difference processing layer further includes: a first feature mapping layer, a second feature mapping layer, and a feature concatenation and prediction layer. The network structure of the first feature mapping layer is different from that of the second feature mapping layer. The server can predict the pronunciation score of the target speech sample through the pronunciation difference processing layer in the following manner to obtain a pronunciation prediction result: perform mapping processing on the target speech sample through the first feature mapping layer to obtain a first mapped feature; perform mapping processing on the target speech sample through the second feature mapping layer to obtain a second mapped feature; perform concatenation processing on the first mapped feature and the second mapped feature through the feature concatenation and prediction layer to obtain a concatenated feature, and predict the pronunciation score of the concatenated feature to obtain a pronunciation prediction result.
[0148] Here, in practical applications, the first feature mapping layer can be constructed based on a Transformer network, and the second feature mapping layer can be constructed based on a Time-Delay Neural Network (TDNN).
[0149] Step 103: Determine the content difference between the content of the target speech sample and the content of the speech sample through the content difference processing layer.
[0150] After predicting the pronunciation prediction result corresponding to the target speech sample through the pronunciation difference processing layer, determine the content difference between the content of the target speech sample and the content of the speech sample through the content difference processing layer. Here, the content difference mainly includes the difference in speech information quantity.
[0151] In some embodiments, the content difference processing layer includes: a Fourier transform layer. The server can determine the content difference between the content of the target speech sample and the content of the speech sample through the content difference processing layer in the following manner: perform Fourier transform on the target speech sample through the Fourier transform layer to obtain a first amplitude spectrum, and perform Fourier transform on the speech sample to obtain a second amplitude spectrum; determine the amplitude difference between the first amplitude spectrum and the second amplitude spectrum, and determine the amplitude difference as the content difference between the content of the target speech sample and the content of the speech sample.
[0152] Here, the content difference processing layer includes: a Fourier transform layer; the server can perform a Fourier transform on the target speech sample through the Fourier transform layer to obtain a first amplitude spectrum, and perform a Fourier transform on the speech sample to obtain a second amplitude spectrum; determine the amplitude difference between the first amplitude spectrum and the second amplitude spectrum. Specifically, it can be to calculate the first average amplitude of the first amplitude spectrum and the second average amplitude of the second amplitude spectrum, and then determine the amplitude difference between the first average amplitude and the second average amplitude as the amplitude difference between the first amplitude spectrum and the second amplitude spectrum; thus, determine the amplitude difference between the first amplitude spectrum and the second amplitude spectrum as the content difference between the content of the target speech sample and the content of the speech sample.
[0153] In some embodiments, the Fourier transform layer includes at least two sub-Fourier transform layers, and different sub-Fourier transform layers correspond to different transformation scales; the server can perform a Fourier transform on the target speech sample through the Fourier transform layer in the following manner to obtain a first amplitude spectrum, and perform a Fourier transform on the speech sample to obtain a second amplitude spectrum: perform Fourier transforms with corresponding transformation scales on the target speech sample through each sub-Fourier transform layer to obtain the first amplitude spectrum corresponding to each sub-Fourier transform layer; perform Fourier transforms with corresponding transformation scales on the speech sample through each sub-Fourier transform layer to obtain the second amplitude spectrum corresponding to each sub-Fourier transform layer;
[0154] Correspondingly, the server can determine the amplitude difference between the first amplitude spectrum and the second amplitude spectrum in the following manner: determine the intermediate amplitude difference between the first amplitude spectrum and the second amplitude spectrum corresponding to each sub-Fourier transform layer; perform a summation and averaging process on the intermediate amplitude differences corresponding to at least two sub-Fourier transform layers to obtain an average amplitude difference, and use the average amplitude difference as the amplitude difference.
[0155] In some embodiments, the content difference processing layer further includes: a power compression processing layer; the server can determine the amplitude difference between the first amplitude spectrum and the second amplitude spectrum in the following manner: perform a compression process on the first amplitude spectrum through the power compression processing layer to obtain a first compressed amplitude spectrum, and perform a compression process on the second amplitude spectrum to obtain a second compressed amplitude spectrum; determine the compressed amplitude difference between the first compressed amplitude spectrum and the second compressed amplitude spectrum, and use the compressed amplitude difference as the amplitude difference.
[0156] As an example, see Figure 10 , Figure 10It is a schematic structural diagram of the content difference processing layer provided by an embodiment of the present application. Here, the content difference processing layer 430 includes Fourier transform layers of three scales and a power compression processing layer. The analysis window sizes of the three scales are 256 points, 512 points, and 1024 points respectively. Under the three window length conditions, the STFT magnitude spectra of the speech sample and the target speech sample after noise reduction are calculated respectively. Then, the calculated magnitude spectra are subjected to 0.3 - power compression to obtain the compressed magnitude spectra. The average magnitude difference is calculated through the compressed magnitude spectra of the speech sample and the target speech sample after noise reduction, and the calculated average magnitude difference is used as the magnitude difference at the corresponding scale. Finally, the average value of the magnitude differences at the three scales is used as the final content difference.
[0157] Step 104: Based on the pronunciation prediction result and the content difference, update the model parameters of the speech noise reduction model to obtain a trained speech noise reduction model.
[0158] Here, after the server predicts the pronunciation prediction result corresponding to the speech sample based on the pronunciation difference processing layer and determines the content difference between the content of the speech sample and the content of the target speech sample based on the content difference processing layer, to avoid losing some speech information during the noise reduction process, at this time, based on the pronunciation prediction result and the content difference, the model parameters of the speech noise reduction model are updated, thereby obtaining a trained speech noise reduction model.
[0159] In some embodiments, the pronunciation difference processing layer includes: a pronunciation score loss processing layer; the server can update the model parameters of the speech noise reduction model based on the pronunciation prediction result and the content difference in the following manner: through the pronunciation score loss processing layer, determine the difference between the pronunciation prediction result and the sample label corresponding to the speech sample, and determine the value of the score loss function based on the difference; based on the content difference and the value of the score loss function, update the model parameters of the speech noise reduction model.
[0160] Here, the pronunciation difference processing layer further includes a pronunciation score loss processing layer, which is used to determine the value of the score loss function based on the difference between the pronunciation prediction result and the sample label corresponding to the speech sample. The sample label is the true pronunciation score corresponding to the speech sample. In practical applications, the value of the pronunciation loss function can be calculated by the following formula:
[0161]
[0162] Where, is the value of the pronunciation loss function, p >= 1, x t is the true pronunciation score, is the pronunciation prediction result output by the pronunciation difference processing layer.
[0163] After determining the value of the scoring loss function, update the model parameters of the speech noise reduction model based on the value of the scoring loss function and the content difference.
[0164] As an example, refer to Figure 11 , Figure 11 FIG. is a schematic structural diagram of the pronunciation difference processing layer adopted by the embodiments of the present application. Here, the pronunciation difference processing layer 420 (i.e., the pronunciation error prediction network PronNet) is composed of a first feature mapping layer (i.e., the TDNN network), a second feature mapping layer (i.e., the Transformer network), a feature splicing and prediction layer (i.e., the linear fusion layer Linear), and a pronunciation scoring loss processing layer. The pronunciation scoring loss processing layer includes a pronunciation similarity scoring loss Lp. Among them, the number of layers of the TDNN network is greater than 3 layers, the number of hidden layer nodes is greater than 128, the number of output layer nodes is equal to the number of phonemes, and the output activation function adopts the Sigmoid function; the number of encoding layers of the Transformer network is greater than 6 layers, the number of decoding layers is greater than 4 layers, the number of attention heads is greater than 4, and the number of hidden nodes is greater than 128. The pronunciation similarity scoring loss Lp is calculated by the following formula:
[0165]
[0166] where p >= 1, x t is the true pronunciation score, is the pronunciation score predicted by the pronunciation error prediction network.
[0167] In some embodiments, the server can update the model parameters of the speech noise reduction model based on the content difference and the value of the scoring loss function in the following manner: obtain the first weight value corresponding to the content difference and the second weight value corresponding to the value of the scoring loss function; combine the first weight value and the second weight value, and determine the value of the loss function of the speech noise reduction model based on the content difference and the value of the scoring loss function; update the model parameters of the speech noise reduction model based on the value of the loss function.
[0168] Here, the first weight value corresponding to the content difference and the second weight value corresponding to the value of the scoring loss function can be preset in advance. At this time, when updating the model parameters of the speech noise reduction model based on the content difference and the value of the scoring loss function, the server first obtains the first weight value corresponding to the content difference and the second weight value corresponding to the value of the scoring loss function; then combines the first weight value and the second weight value, and determines the value of the loss function of the speech noise reduction model based on the content difference and the value of the scoring loss function. Specifically, it can be to perform weighted processing on the content difference and the value of the scoring loss function based on the first weight value and the second weight value, and use the obtained result as the value of the loss function of the speech noise reduction model; finally, update the model parameters of the speech noise reduction model based on the value of the loss function of the speech noise reduction model.
[0169] In some embodiments, the server may update the model parameters of the voice noise reduction model based on the value of the loss function in the following manner: when the value of the loss function exceeds the loss threshold, determine the error signal of the voice noise reduction model based on the loss function; backpropagate the error signal in the voice noise reduction model, and update the model parameters of each layer in the voice noise reduction model during the propagation process.
[0170] Here, when the server updates the model parameters of the voice noise reduction model based on the value of the loss function of the voice noise reduction model, it determines whether the value of the loss function exceeds the loss threshold. When the value of the loss function exceeds the loss threshold, it determines the error signal of the voice noise reduction model based on the loss function, and backpropagates the error signal in the voice noise reduction model. Thus, during the backpropagation process of the error information, the model parameters of each layer in the voice noise reduction model are updated until the loss function converges. The model parameters of the voice noise reduction model obtained when it converges are used as the model parameters of the trained voice noise reduction model.
[0171] Applying the above embodiments of the present application, a pronunciation difference processing layer and a content difference processing layer are added to the voice noise reduction model. Through the pronunciation difference processing layer, the pronunciation score of the denoised target voice sample is predicted, and a pronunciation prediction result indicating the pronunciation similarity between the target voice sample and the reference pronunciation corresponding to the voice sample is obtained. And through the content difference processing layer, the content difference between the content of the target voice sample and the content of the voice sample is determined. Thus, based on the pronunciation prediction result and the content difference, the model parameters of the voice noise reduction model are updated to complete the model training; training the voice noise reduction model based on the pronunciation similarity and content difference before and after noise reduction in this way can enable the trained voice noise reduction model to avoid the loss of voice information before and after noise reduction and improve the accuracy of the noise reduction process.
[0172] Based on the above description of the training method of the voice noise reduction model provided by the embodiments of the present application, the voice scoring method provided by the embodiments of the present application is described below. This voice scoring method is applied to the voice noise reduction model, and this voice noise reduction model is trained based on the above training method of the voice noise reduction model.
[0173] In some embodiments, the voice scoring method provided by the embodiments of the present application can be implemented independently by the server or the terminal, or jointly implemented by the server and the terminal. The voice scoring method provided by the embodiments of the present application is described below taking the implementation by the terminal as an example. Refer to Figure 12 , Figure 12 is a schematic flowchart of the voice scoring method provided by the embodiments of the present application. The voice scoring method provided by the embodiments of the present application includes:
[0174] Step 201: The terminal presents the reference voice text and the voice input function item.
[0175] Here, the terminal is provided with a client for voice scoring. By running the client, a reference voice text and a voice input function item are presented.
[0176] Step 202: In response to a trigger operation on the voice input function item, a voice input interface is presented, and a voice end function item is presented in the voice input interface.
[0177] When a trigger operation on the voice input function item is received, in response to this trigger operation, a voice input interface is presented, and at the same time, a voice end function item is presented in the voice input interface. At this time, the user can input corresponding voice information according to the reference voice text based on this voice input interface.
[0178] Step 203: The voice information input based on the voice input interface is received.
[0179] Step 204: In response to a trigger operation on the voice end function item, a pronunciation score indicating the pronunciation similarity between the voice information and the reference pronunciation corresponding to the reference voice text is presented.
[0180] The terminal receives the voice information input based on this voice input interface. When a trigger operation on the voice end function item is received, in response to this trigger operation, a pronunciation score indicating the pronunciation similarity between the voice information and the reference pronunciation corresponding to the reference voice text is presented. In practical applications, this pronunciation score can be identified in various ways such as numbers and graphics.
[0181] Among them, this pronunciation score is obtained by predicting the pronunciation score of the target voice information, and the target voice information is obtained by performing noise reduction processing on the voice information based on a voice noise reduction model; among them, this voice noise reduction model is trained based on the training method of the above voice noise reduction model.
[0182] As an example, refer to Figure 13 , Figure 13 is a schematic diagram showing the presentation of the voice scoring process provided by the embodiments of the present application. Here, taking the voice scoring method provided by the embodiments of the present application being applied to the scenario of character dubbing as an example, the terminal displays multiple selectable dubbing characters in the dubbing interface, including "Character 1, Character 2, Character 3, and Character 4", and corresponding dubbing entrances, which can be represented by character images, such as Figure 13 shown in Figure A; when a trigger operation on the dubbing entrance corresponding to "Character 2" is received, the reference voice text (i.e., the character lines) "Hello everyone, I'm your good friend XXX" corresponding to "Character 2", and the voice input function item "Start Dubbing" are presented, such as Figure 13 shown in Figure B;
[0183] In response to a trigger operation for the voice input function item "Start Dubbing", a voice input interface is presented, and the voice end function item "End Dubbing" is presented in the voice input interface, as shown in Figure 13 Figure C in; when voice information input based on the voice input interface is received, in response to a trigger operation for the voice end function item "End Dubbing", a pronunciation score indicating the pronunciation similarity between the received voice information and the reference pronunciation corresponding to the reference voice text "Hello everyone, I'm your good friend XXX" is presented, that is, "90 points, great!", as shown in Figure 13 Figure D in.
[0184] In practical applications, the voice scoring method provided by the embodiments of the present application can also be applied to the scenario of singing scoring. Specifically, when a user sings and selects a song to sing, the terminal presents the reference voice text (i.e., the lyrics) corresponding to the song and the voice input function item; in response to a trigger operation for the voice input function item, a voice input interface is presented to collect the user's singing voice information, and the voice end function item is presented in the voice input interface; when the singing voice information input based on the voice input interface is received, in response to a trigger operation for the voice end function item, a pronunciation score indicating the pronunciation similarity between the singing voice information and the reference pronunciation corresponding to the reference voice text is presented.
[0185] Applying the above embodiments of the present application, a pronunciation difference processing layer and a content difference processing layer are added to the voice noise reduction model. Through the pronunciation difference processing layer, a pronunciation score prediction of the denoised target voice sample is performed to obtain a pronunciation prediction result indicating the pronunciation similarity between the target voice sample and the reference pronunciation corresponding to the voice sample, and the content difference between the content of the target voice sample and the content of the voice sample is determined through the content difference processing layer. Thus, based on the pronunciation prediction result and the content difference, the model parameters of the voice noise reduction model are updated to complete model training; thus, training the voice noise reduction model based on the pronunciation similarity and content difference before and after noise reduction can enable the trained voice noise reduction model to avoid the loss of voice information before and after noise reduction and improve the accuracy of the noise reduction process. Thereby further improving the prediction accuracy of the pronunciation score.
[0186] The exemplary application of the embodiments of the present application in an actual application scenario will be described below.
[0187] In related technologies, voice enhancement solutions all belong to pure acoustic prediction solutions. The prediction goal is usually to make the waveform of the enhanced voice most similar to that of the clean voice. However, for computer-aided language teaching, the waveform of the enhanced voice being closest to that of the clean voice is not the best solution. In practical applications, when learning with the goal of the closest waveform, usually only the recovery degree of vowels with large amplitudes is concerned, while the recovery degree of consonants with small amplitudes is ignored, which easily causes phenomena such as fricative loss, plosive loss of explosion, and lack of aspiration segments in aspirated sounds. As a result, the accuracy of pronunciation score prediction is affected due to the voice noise reduction process.
[0188] Based on this, the embodiments of the present application provide a training method for a voice noise reduction model. A pronunciation error prediction network (i.e., the above-mentioned pronunciation difference processing layer) and a multi-scale voice similarity metric network (i.e., the above-mentioned content difference processing layer) are introduced into the voice noise reduction model to explicitly penalize the pronunciation error information of the enhanced voice. At the same time, a voice enhancement network that can mutually integrate and promote spectral harmonic information, phase information, and amplitude information is proposed, which is mainly reflected in the detailed design of the cascaded activation network CasNet, including the structure of multiple harmonic enhancers H, and using the amplitude spectrum to assist the phase spectrum for phase estimation.
[0189] Next, the application scenario of the training method for the voice noise reduction model provided by the embodiments of the present application will be described first. Refer to Figure 13 , which is mainly applied to the role dubbing evaluation function. Here, 1) Click the start dubbing button to start following the role's lines; 2) Click the end dubbing button to end following the role's lines; 3) The screen presents the pronunciation evaluation result of the voice of the collected role dubbing to the user, as Figure 13 shown in the pronunciation evaluation result of the voice of the role dubbing, which is represented by a score, that is, 90 points.
[0190] Next, the voice scoring method provided by the embodiments of the present application will be described in detail. Refer to Figure 14 , Figure 14 is a schematic flowchart of the voice scoring method based on the voice noise reduction model provided by the embodiments of the present application, including: 1) The user opens the voice scoring client, the screen displays the text to follow, clicks the start recording button displayed on the client, and follows the sentence based on the text to follow;
[0191] 2) The client sends the audio information collected during the following process and the text to follow to the server side;
[0192] 3) The server side sends the audio information to the voice noise reduction model for voice noise reduction processing;
[0193] 4) After the voice noise reduction model performs noise reduction processing on the audio information, it inputs the denoised audio information into the speech recognition model.
[0194] 5) The speech recognition model performs speech recognition on the noise-reduced audio information and extracts basic acoustic features to obtain the recognized text and acoustic features (such as pronunciation accuracy, pronunciation fluency, pronunciation prosody, etc.).
[0195] 6) The speech recognition model inputs the results of speech recognition (i.e., the recognized text and acoustic features) into the evaluation model;
[0196] 7) The evaluation model predicts the pronunciation score based on the recognized text and acoustic features, outputs the pronunciation score, and returns the pronunciation score to the server side;
[0197] 8) The server side receives the pronunciation score and returns the pronunciation score to the client side so that the user can view the final pronunciation score on the client side.
[0198] Next, the speech noise reduction model provided by the embodiments of the present application will be further described in detail. See Figure 4 , the speech noise reduction model includes a speech enhancement network EnhanceNet (i.e., the noise processing layer), a pronunciation error predictor PronNet (i.e., the pronunciation difference processing layer), and a multi-scale speech similarity metric network SimilarNet (i.e., the content difference processing layer).
[0199] Specifically, the training process of the speech noise reduction model can be as follows: The speech enhancement network EnhanceNet performs speech enhancement processing (i.e., noise reduction processing) on the collected original speech, and then inputs the noise-reduced target speech into the pronunciation error prediction network PronNet and the multi-scale speech similarity metric network SimilarNet respectively; the pronunciation similarity score loss is obtained through the pronunciation error prediction network PronNet, and the speech similarity loss (i.e., the loss of the content contained in the speech before and after noise reduction) is obtained through the multi-scale speech similarity metric network SimilarNet; the loss of the speech noise reduction model is determined based on the pronunciation similarity score loss and the speech similarity loss, and thus the gradient is backpropagated based on the loss of the speech noise reduction model to update the model parameters of the speech noise reduction model, thereby realizing the model training of the speech noise reduction model.
[0200] See Figure 5 , here, the speech enhancement network EnhanceNet includes a pre-processing network PrevNet (i.e., the first feature transformation layer), a post-processing network PostNet (i.e., the second feature transformation layer), and a cascaded activation network CasNet (i.e., the filtering processing layer).
[0201] Among them, the above pre-processing network PrevNet is composed of a Fourier transform layer and multiple convolutional layers. See Figure 6This preprocessing network PrevNet (i.e., the first feature transformation layer) uses the STFT transformation through the Fourier transform layer to convert the waveform of the original speech into a 2-channel Fourier spectrum, and then converts it from the 2-channel Fourier spectrum into a 64-channel amplitude spectrum A and a 64-channel phase spectrum P through a convolutional layer.
[0202] Among them, the above-mentioned cascaded activation network CasNet (i.e., the filtering processing layer) is composed of a cascade of multiple third-order activation attention modules TAB (i.e., sub-filtering processing layers), see Figure 7 Here, the cascaded activation network CasNet filters the 64-channel amplitude spectrum A and phase spectrum P output by the preprocessing network PrevNet through a convolutional layer, and outputs a 64-channel amplitude spectrum A' and a phase spectrum P'.
[0203] See Figure 8 As shown in Figure A of , the third-order attention module TAB (i.e., the sub-filtering processing layer) in the cascaded activation network CasNet consists of 1 third-order amplitude spectrum enhancement network AmpNet and 1 first-order phase spectrum correction network PhaseNet. Among them, the amplitude spectrum enhancement network AmpNet (i.e., the amplitude spectrum filtering network) enhances the 64-channel amplitude spectrum A output by the preprocessing network to obtain the amplitude spectrum A'. The phase spectrum correction layer PhaseNet receives two inputs, one from the enhanced amplitude spectrum A', and the other is the phase spectrum itself P. The relationship between the output phase spectrum P' and the two inputs is: Among them, Conv is the convolution operation; Tanh is the hyperbolic tangent function operation (converting the input value to between -1 and 1); represents the dot product, represents the concatenation.
[0204] Furthermore, the amplitude spectrum enhancement network AmpNet is composed of 3 levels of harmonic enhancers H (i.e., amplitude spectrum filtering layers) (as shown in Figure B of ), and the composition method of the harmonic enhancer H is as shown in Figure C of Figure 8 See Figure 8 See
[0205] Among them, see Figure 9 The above-mentioned postprocessing network PostNet (i.e., the second feature transformation layer) consists of multiple convolutions, converts the 64-channel amplitude spectrum A' output by the cascaded activation network CasNet into a 1-channel amplitude mask M, converts the 64-channel phase spectrum P' into a 2-channel phase angle Ω, and then converts it into a waveform output through the inverse Fourier transform, that is, the target speech after noise reduction is obtained.
[0206] See Figure 11, the above pronunciation error prediction network PronNet is composed of a TDNN network (i.e., the second feature mapping layer), a Transformer network (i.e., the first feature mapping layer), a linear fusion layer Linear (i.e., the feature concatenation and prediction layer), and a pronunciation score loss processing layer. Among them, the number of layers of the TDNN network is greater than 3, the number of hidden layer nodes is greater than 128, the number of output layer nodes is equal to the number of phonemes, and the output activation function uses the Sigmoid function; the number of encoding layers of the Transformer network is greater than 6, the number of decoding layers is greater than 4, the number of attention heads is greater than 4, and the number of hidden nodes is greater than 128. The pronunciation similarity score loss Lp of the pronunciation score loss processing layer is calculated using the following formula:
[0207]
[0208] where p >= 1, x t is the true pronunciation score, is the pronunciation score predicted by the pronunciation error prediction network.
[0209] See Figure 10 , the above multi-scale speech similarity metric network SimilarNet includes a Fourier transform layer, which consists of 3 different Fourier transform scales. The analysis window sizes of the three scales are 256 points, 512 points, and 1024 points respectively. Under the three window length conditions, after calculating the STFT spectra of the original speech sample and the denoised target speech sample respectively, then through the power compression processing layer, the calculated STFT spectra are subjected to 0.3 power compression to obtain the CompressSTFT spectra. The average amplitude loss is calculated through the CompressSTFT spectra of the original speech sample and the denoised target speech sample, and the calculated average amplitude loss is used as the speech similarity loss under the corresponding scale. Finally, the average value of the speech similarity losses under the 3 scales is used as the final speech similarity loss (i.e., the content difference).
[0210] In some other embodiments, PrevNet and PostNet proposed in this application can adopt a variety of different implementation schemes. Among them, PrevNet only needs to transform the waveform signal into 2-channel time-frequency features, and then transform the 2-channel time-frequency features into high-channel time-frequency features. It is found in the implementation process of this application that the higher the number of channels, the better the performance. The design of PostNet is similar, and it can also adopt structures such as BLSTM, GRU, or Transformer to implement the conversion from high-channel features to 2-channel time-frequency domain, and then from the time-frequency domain to the waveform signal.
[0211] Applying the above embodiments of the present application, in the pronunciation evaluation scenario, a pronunciation error network and a multi-scale speech similarity metric network are introduced into the speech noise reduction network. While reducing the noise of the speech, the influence of the noise reduction process on pronunciation evaluation is reduced, and the pronunciation evaluation error caused by noise reduction is greatly reduced. Especially for the features of consonants such as fricatives, plosives, and aspirated sounds, after introducing the pronunciation error network, the error evaluation rates of these three sounds are relatively reduced by 23.5%.
[0212] Next, the implementation of the training device 555 of the speech noise reduction model provided by the embodiments of the present application as an exemplary structure of software modules will be continued. In some embodiments, as Figure 2 shown, the software modules stored in the training device 555 of the speech noise reduction model in the memory 550 may include:
[0213] A noise reduction module 5551, configured to perform noise reduction processing on the speech sample through the noise processing layer to obtain a target speech sample;
[0214] A prediction module 5552, configured to predict the pronunciation score of the target speech sample through the pronunciation difference processing layer to obtain a pronunciation prediction result, where the pronunciation prediction result is used to indicate the pronunciation similarity between the target speech sample and the reference pronunciation corresponding to the speech sample;
[0215] A determination module 5553, configured to determine the content difference between the content of the target speech sample and the content of the speech sample through the content difference processing layer;
[0216] An update module 5554, configured to update the model parameters of the speech noise reduction model based on the pronunciation prediction result and the content difference to obtain a trained speech noise reduction model.
[0217] In some embodiments, the noise processing layer includes: a first feature transformation layer, a filtering processing layer, and a second feature transformation layer;
[0218] The noise reduction module 5551 is further configured to perform Fourier transform on the speech sample through the first feature transformation layer to obtain the amplitude spectrum and phase spectrum corresponding to the speech sample;
[0219] Perform filtering processing on the amplitude spectrum through the filtering processing layer to obtain a target amplitude spectrum, and perform phase correction on the phase spectrum to obtain a target phase spectrum;
[0220] Multiply the target amplitude spectrum and the target phase spectrum through the second feature transformation layer, and perform inverse Fourier transform on the result of the multiplication to obtain the target speech sample.
[0221] In some embodiments, the filtering processing layer includes at least two cascaded sub-filtering processing layers;
[0222] The noise reduction module 5551 is further configured to filter the amplitude spectrum through the first-level sub-filtering processing layer to obtain an intermediate amplitude spectrum, and perform phase correction on the phase spectrum to obtain an intermediate phase spectrum;
[0223] Filter the intermediate amplitude spectrum through non-first-level sub-filtering processing layers to obtain the target amplitude spectrum, and perform phase correction on the intermediate phase spectrum to obtain the target phase spectrum.
[0224] In some embodiments, each of the sub-filtering processing layers includes a phase spectrum correction layer and at least two cascaded amplitude spectrum filtering layers;
[0225] The noise reduction module 5551 is further configured to filter the amplitude spectrum through the at least two cascaded amplitude spectrum filtering layers to obtain an intermediate amplitude spectrum;
[0226] Based on the intermediate amplitude spectrum, perform phase correction on the phase spectrum through the phase spectrum correction layer to obtain an intermediate phase spectrum.
[0227] In some embodiments, the second feature transformation layer includes a feature conversion layer and a feature inverse transformation layer;
[0228] The noise reduction module 5551 is further configured to convert the target amplitude spectrum into an amplitude spectrum mask through the feature conversion layer, and determine the phase angle corresponding to the target phase spectrum;
[0229] Multiply the target amplitude spectrum, the amplitude spectrum mask, and the phase angle corresponding to the target phase spectrum through the feature inverse transformation layer, and perform an inverse Fourier transform on the result of the multiplication to obtain the target speech sample.
[0230] In some embodiments, the content difference processing layer includes: a Fourier transform layer;
[0231] The determination module 5553 is further configured to perform a Fourier transform on the target speech sample through the Fourier transform layer to obtain a first amplitude spectrum, and perform a Fourier transform on the speech sample to obtain a second amplitude spectrum;
[0232] Determine the amplitude difference between the first amplitude spectrum and the second amplitude spectrum, and determine the amplitude difference as the content difference between the content of the target speech sample and the content of the speech sample.
[0233] In some embodiments, the Fourier transform layer includes at least two sub-Fourier transform layers, and different sub-Fourier transform layers correspond to different transformation scales;
[0234] The determining module 5553 is further configured to perform Fourier transforms with corresponding transformation scales on the target voice sample through each of the sub-Fourier transform layers, to obtain first amplitude spectra corresponding to the sub-Fourier transform layers;
[0235] Perform Fourier transforms with corresponding transformation scales on the voice sample through each of the sub-Fourier transform layers, to obtain second amplitude spectra corresponding to the sub-Fourier transform layers;
[0236] The determining module 5553 is further configured to determine an intermediate amplitude difference between the first amplitude spectra and the second amplitude spectra corresponding to the sub-Fourier transform layers;
[0237] Perform a summation and averaging process on the intermediate amplitude differences corresponding to the at least two sub-Fourier transform layers, to obtain an average amplitude difference, and use the average amplitude difference as the amplitude difference.
[0238] In some embodiments, the content difference processing layer further includes: a power compression processing layer;
[0239] The determining module 5553 is further configured to perform compression processing on the first amplitude spectrum through the power compression processing layer, to obtain a first compressed amplitude spectrum, and perform compression processing on the second amplitude spectrum, to obtain a second compressed amplitude spectrum;
[0240] Determine a compressed amplitude difference between the first compressed amplitude spectrum and the second compressed amplitude spectrum, and use the compressed amplitude difference as the amplitude difference.
[0241] In some embodiments, the pronunciation difference processing layer includes: a pronunciation scoring loss processing layer;
[0242] The updating module 5554 is further configured to determine a difference between the pronunciation prediction result and a sample label corresponding to the voice sample through the pronunciation scoring loss processing layer, and determine a value of a scoring loss function based on the difference;
[0243] Update model parameters of the voice noise reduction model based on the content difference and the value of the scoring loss function.
[0244] In some embodiments, the updating module 5554 is further configured to obtain a first weight value corresponding to the content difference and a second weight value corresponding to the value of the scoring loss function;
[0245] Combine the first weight value and the second weight value, and determine a value of a loss function of the voice noise reduction model based on the content difference and the value of the scoring loss function;
[0246] Update the model parameters of the speech noise reduction model based on the value of the loss function.
[0247] In some embodiments, the updating module 5554 is further configured to, when the value of the loss function exceeds a loss threshold, determine an error signal of the speech noise reduction model based on the loss function;
[0248] Backpropagate the error signal in the speech noise reduction model, and update the model parameters of each layer in the speech noise reduction model during the propagation process.
[0249] In some embodiments, the pronunciation difference processing layer further includes: a first feature mapping layer, a second feature mapping layer, and a feature splicing and prediction layer, and the network structure of the first feature mapping layer is different from the network structure of the second feature mapping layer;
[0250] The prediction module 5552 is further configured to perform a mapping process on the target speech sample through the first feature mapping layer to obtain a first mapped feature;
[0251] Perform a mapping process on the target speech sample through the second feature mapping layer to obtain a second mapped feature;
[0252] Perform a splicing process on the first mapped feature and the second mapped feature through the feature splicing and prediction layer to obtain a spliced feature, and
[0253] Predict the pronunciation score of the spliced feature to obtain the pronunciation prediction result.
[0254] Applying the above embodiments of the present application, a pronunciation difference processing layer and a content difference processing layer are added to the speech noise reduction model. Through the pronunciation difference processing layer, the pronunciation score of the target speech sample after noise reduction processing is predicted to obtain a pronunciation prediction result indicating the pronunciation similarity between the target speech sample and the reference pronunciation corresponding to the speech sample, and the content difference between the content of the target speech sample and the content of the speech sample is determined through the content difference processing layer. Then, based on the pronunciation prediction result and the content difference, the model parameters of the speech noise reduction model are updated to complete model training; thus, training the speech noise reduction model based on the pronunciation similarity and content difference before and after noise reduction can prevent the loss of speech information before and after noise reduction in the trained speech noise reduction model and improve the accuracy of the noise reduction processing.
[0255] Next, the speech scoring device provided by the embodiments of the present application is continued to be described. The speech scoring device provided by the embodiments of the present application is applied to a speech noise reduction model. The speech scoring device provided by the embodiments of the present application includes:
[0256] A first presentation module, configured to present a reference speech text and a speech input function item;
[0257] A second presentation module, configured to present a voice input interface in response to a trigger operation on the voice input function item, and present a voice end function item in the voice input interface;
[0258] A receiving module, configured to receive voice information input based on the voice input interface;
[0259] A third presentation module, configured to present a pronunciation score indicating the pronunciation similarity between the voice information and the reference pronunciation corresponding to the reference voice text in response to a trigger operation on the voice end function item;
[0260] Wherein, the pronunciation score is predicted based on the pronunciation scoring of the target voice information, and the target voice information is obtained by performing noise reduction processing on the voice information based on the voice noise reduction model;
[0261] Wherein, the voice noise reduction model is trained based on the training method of the above voice noise reduction model.
[0262] Applying the above embodiments of the present application, a pronunciation difference processing layer and a content difference processing layer are added to the voice noise reduction model. Through the pronunciation difference processing layer, the pronunciation scoring of the denoised target voice sample is predicted to obtain a pronunciation prediction result indicating the pronunciation similarity between the target voice sample and the reference pronunciation corresponding to the voice sample, and the content difference between the content of the target voice sample and the content of the voice sample is determined through the content difference processing layer. Thus, based on the pronunciation prediction result and the content difference, the model parameters of the voice noise reduction model are updated to complete the model training; thus, training the voice noise reduction model based on the pronunciation similarity and content difference before and after noise reduction can enable the trained voice noise reduction model to avoid the loss of voice information before and after noise reduction, improve the accuracy of the noise reduction processing, and thus further improve the prediction accuracy of the pronunciation score.
[0263] An embodiment of the present application further provides an electronic device, where the electronic device includes:
[0264] A memory, configured to store executable instructions;
[0265] A processor, configured to implement the method provided by the embodiment of the present application when executing the executable instructions stored in the memory.
[0266] An embodiment of the present application further provides a computer program product or a computer program, where the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method provided by the embodiment of the present application.
[0267] The embodiments of the present application further provide a computer-readable storage medium storing executable instructions, which, when executed by a processor, implement the method for training a voice noise reduction model provided by the embodiments of the present application.
[0268] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disc, or CD-ROM; or may be various devices including one or any combination of the above memories.
[0269] In some embodiments, the executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as an independent program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0270] As an example, the executable instructions may or may not correspond to a file in the file system, may be stored as part of a file that stores other programs or data, for example, stored in one or more scripts in a Hyper Text Markup Language (HTML) document, stored in a single file dedicated to the program in question, or stored in multiple cooperating files (for example, files that store one or more modules, subroutines, or portions of code).
[0271] As an example, the executable instructions may be deployed to execute on one computing device, or on multiple computing devices located at one location, or alternatively, on multiple computing devices distributed at multiple locations and interconnected by a communication network.
[0272] As described above, the above are only embodiments of the present application and are not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are all included in the protection scope of the present application.
Claims
1. A training method for a voice noise reduction model, characterized in that, The voice noise reduction model includes: a noise processing layer, a pronunciation difference processing layer, and a content difference processing layer. The method includes: Through the noise processing layer, perform noise reduction processing on the voice sample to obtain a target voice sample; Through the pronunciation difference processing layer, predict the pronunciation score of the target voice sample to obtain a pronunciation prediction result, which is used to indicate the pronunciation similarity between the target voice sample and the reference pronunciation corresponding to the voice sample; Through the content difference processing layer, determine the amplitude difference between the first amplitude spectrum of the target voice sample and the second amplitude spectrum of the voice sample, and use the amplitude difference as the content difference between the content of the target voice sample and the content of the voice sample; Based on the pronunciation prediction result and the content difference, update the model parameters of the voice noise reduction model to obtain a trained voice noise reduction model.
2. The method according to claim 1, characterized in that The noise processing layer includes: a first feature transformation layer, a filtering processing layer, and a second feature transformation layer; The step of performing noise reduction processing on the voice sample through the noise processing layer to obtain a target voice sample includes: Through the first feature transformation layer, perform Fourier transform on the voice sample to obtain the amplitude spectrum and phase spectrum corresponding to the voice sample; Through the filtering processing layer, perform filtering processing on the amplitude spectrum to obtain a target amplitude spectrum, and perform phase correction on the phase spectrum to obtain a target phase spectrum; Through the second feature transformation layer, multiply the target amplitude spectrum and the target phase spectrum, and perform inverse Fourier transform on the multiplied result to obtain the target voice sample.
3. The method according to claim 2, wherein The filtering processing layer includes at least two cascaded sub-filtering processing layers; The step of performing filtering processing on the amplitude spectrum through the filtering processing layer to obtain a target amplitude spectrum and performing phase correction on the phase spectrum to obtain a target phase spectrum includes: Through the first-level sub-filtering processing layer, perform filtering processing on the amplitude spectrum to obtain an intermediate amplitude spectrum, and perform phase correction on the phase spectrum to obtain an intermediate phase spectrum; Through the non-first-level sub-filtering processing layer, perform filtering processing on the intermediate amplitude spectrum to obtain the target amplitude spectrum, and perform phase correction on the intermediate phase spectrum to obtain the target phase spectrum.
4. The method according to claim 3, wherein Each of the sub-filtering processing layers includes a phase spectrum correction layer and at least two cascaded amplitude spectrum filtering layers; The step of performing filtering processing on the amplitude spectrum through the first-level sub-filtering processing layer to obtain an intermediate amplitude spectrum and performing phase correction on the phase spectrum to obtain an intermediate phase spectrum includes: Through the at least two cascaded amplitude spectrum filtering layers, perform filtering processing on the amplitude spectrum to obtain an intermediate amplitude spectrum; Through the phase spectrum correction layer, perform phase correction on the phase spectrum based on the intermediate amplitude spectrum to obtain an intermediate phase spectrum.
5. The method according to claim 2, characterized in that The second feature transformation layer includes a feature transformation layer and a feature inverse transformation layer; The step of multiplying the target amplitude spectrum and the target phase spectrum through the second feature transformation layer and performing inverse Fourier transform on the multiplied result to obtain the target voice sample includes: Through the feature transformation layer, convert the target amplitude spectrum into an amplitude spectrum mask, and determine the phase angle corresponding to the target phase spectrum; Through the feature inverse transformation layer, multiply the target amplitude spectrum, the amplitude spectrum mask, and the phase angle corresponding to the target phase spectrum, and perform an inverse Fourier transform on the result of the multiplication to obtain the target speech sample.
6. The method according to claim 1, wherein The content difference processing layer includes: a Fourier transform layer; Before determining the amplitude difference between the first amplitude spectrum of the target speech sample and the second amplitude spectrum of the speech sample, the method further includes: Through the Fourier transform layer, perform a Fourier transform on the target speech sample to obtain a first amplitude spectrum, and perform a Fourier transform on the speech sample to obtain a second amplitude spectrum.
7. The method according to claim 6, wherein The Fourier transform layer includes at least two sub-Fourier transform layers, and different sub-Fourier transform layers correspond to different transformation scales; The step of performing a Fourier transform on the target speech sample through the Fourier transform layer to obtain a first amplitude spectrum, and performing a Fourier transform on the speech sample to obtain a second amplitude spectrum includes: Through each of the sub-Fourier transform layers, perform a Fourier transform on the target speech sample with a corresponding transformation scale to obtain a first amplitude spectrum corresponding to each of the sub-Fourier transform layers; Through each of the sub-Fourier transform layers, perform a Fourier transform on the speech sample with a corresponding transformation scale to obtain a second amplitude spectrum corresponding to each of the sub-Fourier transform layers; The step of determining the amplitude difference between the first amplitude spectrum and the second amplitude spectrum includes: Determine the intermediate amplitude difference between the first amplitude spectrum and the second amplitude spectrum corresponding to each of the sub-Fourier transform layers; Perform a summation and averaging process on the intermediate amplitude differences corresponding to the at least two sub-Fourier transform layers to obtain an average amplitude difference, and use the average amplitude difference as the amplitude difference.
8. The method according to claim 6, characterized in that, The content difference processing layer further includes: a power compression processing layer; The step of determining the amplitude difference between the first amplitude spectrum and the second amplitude spectrum includes: Through the power compression processing layer, perform a compression process on the first amplitude spectrum to obtain a first compressed amplitude spectrum, and perform a compression process on the second amplitude spectrum to obtain a second compressed amplitude spectrum; Determine the compressed amplitude difference between the first compressed amplitude spectrum and the second compressed amplitude spectrum, and use the compressed amplitude difference as the amplitude difference.
9. The method according to claim 1, wherein The pronunciation difference processing layer includes: a pronunciation scoring loss processing layer; Updating the model parameters of the speech noise reduction model based on the pronunciation prediction result and the content difference includes: Through the pronunciation scoring loss processing layer, determine the difference between the pronunciation prediction result and the sample label corresponding to the speech sample, and determine the value of the scoring loss function based on the difference; Update the model parameters of the speech noise reduction model based on the content difference and the value of the scoring loss function.
10. The method according to claim 9, characterized in that The step of updating the model parameters of the speech noise reduction model based on the content difference and the value of the scoring loss function includes: Obtain a first weight value corresponding to the content difference and a second weight value corresponding to the value of the scoring loss function; Combine the first weight value and the second weight value, and determine the value of the loss function of the voice noise reduction model based on the content difference and the value of the scoring loss function; Update the model parameters of the voice noise reduction model based on the value of the loss function.
11. The method according to claim 9, characterized in that, The pronunciation difference processing layer further includes: a first feature mapping layer, a second feature mapping layer, and a feature splicing and prediction layer, and the network structure of the first feature mapping layer is different from that of the second feature mapping layer; The method of predicting the pronunciation score of the target voice sample through the pronunciation difference processing layer to obtain a pronunciation prediction result includes: Perform a mapping process on the target voice sample through the first feature mapping layer to obtain a first mapped feature; Perform a mapping process on the target voice sample through the second feature mapping layer to obtain a second mapped feature; Perform a splicing process on the first mapped feature and the second mapped feature through the feature splicing and prediction layer to obtain a spliced feature, and Predict the pronunciation score of the spliced feature to obtain the pronunciation prediction result.
12. A voice scoring method, characterized in that, The method is applied to a voice noise reduction model, and the method includes: Present a reference voice text and a voice input function item; In response to a trigger operation on the voice input function item, present a voice input interface and present a voice end function item in the voice input interface; Receive voice information input based on the voice input interface; In response to a trigger operation on the voice end function item, present a pronunciation score indicating the pronunciation similarity between the voice information and the reference pronunciation corresponding to the reference voice text; Wherein, the pronunciation score is obtained by predicting the pronunciation score of the target voice information, and the target voice information is obtained by performing noise reduction processing on the voice information through the voice noise reduction model; Wherein, the voice noise reduction model is trained based on the training method of the voice noise reduction model according to any one of claims 1-11.
13. A training device for a voice noise reduction model, characterized in that, The voice noise reduction model includes: a noise processing layer, a pronunciation difference processing layer, and a content difference processing layer, and the device includes: A noise reduction module for performing noise reduction processing on a voice sample through the noise processing layer to obtain a target voice sample; A prediction module for predicting the pronunciation score of the target voice sample through the pronunciation difference processing layer to obtain a pronunciation prediction result, and the pronunciation prediction result is used to indicate the pronunciation similarity between the target voice sample and the reference pronunciation corresponding to the voice sample; A determination module for determining the amplitude difference between the first amplitude spectrum of the target voice sample and the second amplitude spectrum of the voice sample through the content difference processing layer, and using the amplitude difference as the content difference between the content of the target voice sample and the content of the voice sample; An update module for updating the model parameters of the voice noise reduction model based on the pronunciation prediction result and the content difference to obtain a trained voice noise reduction model.
14. An electronic device, characterized in that, The electronic device includes: A memory for storing executable instructions; A processor, when executing the executable instructions stored in the memory, implements the method according to any one of claims 1 to 12.
15. A computer-readable storage medium, characterized in that, Stored with executable instructions, when the executable instructions are executed, they are used to implement the method according to any one of claims 1 to 12.
16. A computer program product, comprising computer instructions, characterized in that, When the computer instructions are executed by a processor, the method according to any one of claims 1 to 12 is implemented.
Citation Information
Patent Citations
Method for transforming a noisy audio signal to an enhanced audio signal
CN107077860A
Depth learning voice reinforcement method based on absolute hearing threshold value
CN109671446A