Speech restoration model training method and device, electronic equipment and storage medium

By combining the initial speech repair model and the audio semantic content extraction model during the training process, and adjusting the model parameters, the problem of unclear semantics and poor audio effects in the existing technology is solved, and the high semantic accuracy and high sound quality of the speech repair model is achieved.

CN120279924APending Publication Date: 2025-07-08NETEASE (HANGZHOU) NETWORK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410022485.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-08
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing speech repair model cannot accurately and completely extract the semantic content information in the audio samples to be repaired, resulting in poor repair results, unclear semantic information, low semantic intelligibility, and poor audio effects.

Method used

By obtaining the initial speech repair model to be trained, the audio samples to be repaired, and the audio semantic content extraction model, the initial speech repair model is used for speech processing, the first semantic content representation and model repair audio are obtained, and the second semantic content representation is obtained in combination with the audio semantic content extraction model, the model parameters are adjusted based on the difference, and the speech repair model is obtained.

Benefits of technology

It improves the semantic accuracy and clarity of the output of the speech repair model, improves the sound quality of the audio, and ensures that the semantic information of the repaired audio is accurate, intelligible, and has good audio effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279924A_ABST
    Figure CN120279924A_ABST
Patent Text Reader

Abstract

The invention discloses a voice restoration model training method and device and electronic equipment. The method comprises the steps of obtaining a to-be-trained initial voice restoration model, a to-be-restored audio sample and an audio semantic content extraction model; inputting the to-be-repaired audio sample into a to-be-trained initial voice repair model for voice processing to obtain a first semantic content representation used for representing the audio semantic content of the to-be-repaired audio sample and a model repair audio; inputting the to-be-repaired audio sample into an audio semantic content extraction model for semantic extraction processing to obtain a second semantic content representation used for representing the audio semantic content of the to-be-repaired audio sample; and based on the difference between the model repair audio and the to-be-repaired audio sample and the difference between the first semantic content representation and the second semantic content representation, adjusting model parameters of the initial voice repair model to obtain a voice repair model. Through the method, the restored audio output by the voice restoration model obtained through training is high in semantic accuracy, good in definition and good in tone quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent voice, and particularly to a training method, device, electronic device, and computer-readable storage medium for a voice repair model. Background Art

[0002] As a technology for improving or restoring damaged voice segments and voice segments contaminated by noise, voice repair technology can achieve noise elimination of voice segments or improvement of voice distortion conditions. Therefore, voice repair technology is widely used in fields such as film and television drama dubbing, privacy information protection, and personalized voice synthesis.

[0003] However, with the development of voice repair technology, due to the increasingly diverse user requirements, higher requirements are put forward for voice repair technology. In existing voice repair technology, when the voice repair model generation model repairs audio, the model often cannot accurately and completely extract the semantic content information in the audio sample to be repaired, resulting in the repair effect of the model repair audio output by the model being unable to be guaranteed. Therefore, the existing voice repair model has the following defects: the semantic information of the model repair audio output by the model is not clear, the semantic intelligibility is low, and the audio effect is poor. Therefore, how to obtain a voice repair model such that the model repair audio output by the model has high semantic accuracy, good clarity, and good sound quality has become an urgent technical problem to be solved currently. Summary of the Invention

[0004] Embodiments of the present application provide a training method, device, electronic device, and computer-readable storage medium for a voice repair model to solve the technical problem that the semantic information of the model repair audio output by the existing voice repair model is not clear, the semantic intelligibility is low, and the audio effect is poor.

[0005] Embodiments of the present application provide a training method for a voice repair model, and the method includes:

[0006] Embodiments of the present application further provide a training device for a voice repair model, and the device includes:

[0007] A data acquisition unit configured to acquire an initial voice repair model to be trained, an audio sample to be repaired, and an audio semantic content extraction model.

[0008] A sample processing unit configured to input the audio sample to be repaired into the initial voice repair model to be trained for voice processing to obtain a first semantic content representation representing the audio semantic content of the audio sample to be repaired and a model repair audio.

[0009] A characterization extraction unit, configured to input the audio sample to be repaired into the audio semantic content extraction model for semantic extraction processing, so as to obtain a second semantic content characterization for representing the audio semantic content of the audio sample to be repaired.

[0010] A parameter adjustment unit, configured to adjust the model parameters of the initial speech repair model based on the difference between the model-repaired audio and the audio sample to be repaired, and the difference between the first semantic content characterization and the second semantic content characterization, so as to obtain a speech repair model.

[0011] An embodiment of the present application further provides an electronic device, including a processor and a memory; wherein, the memory is used to store one or more computer instructions, and the one or more computer instructions are executed by the processor to implement the above method.

[0012] An embodiment of the present application further provides a computer-readable storage medium, on which one or more computer instructions are stored, and the instructions are executed by a processor to implement the above method.

[0013] Compared with the prior art, the embodiments of the present application have the following advantages:

[0014] In the training method of the speech repair model provided by the embodiment of the present application, during the training process of the initial speech repair model, an initial speech repair model to be trained, an audio sample to be repaired, and an audio semantic content extraction model are obtained; the audio sample to be repaired is input into the initial speech repair model to be trained for speech processing, so as to obtain a first semantic content characterization for representing the audio semantic content of the audio sample to be repaired and model-repaired audio. The audio sample to be repaired is input into the audio semantic content extraction model for semantic extraction processing, so as to obtain a second semantic content characterization for representing the audio semantic content of the audio sample to be repaired. Based on the obtained first semantic content characterization, second semantic content characterization, and model-repaired audio, the model parameters of the initial speech repair model are continuously adjusted based on the difference between the model-repaired audio and the audio sample to be repaired, and the difference between the first semantic content characterization and the second semantic content characterization. The generated repaired audio of the obtained speech repair model has high semantic accuracy, good intelligibility, and good audio quality. Description of the Drawings

[0015] Figure 1 is a schematic diagram of an application scenario provided by an embodiment of the present application;

[0016] Figure 2 is a flowchart of a training method of a speech repair model provided by an embodiment of the present application;

[0017] Figure 3It is a schematic diagram of the principle for training an initial voice repair model provided by an embodiment of the present application;

[0018] Figure 4 It is a unit block diagram of a training device for a voice repair model provided by an embodiment of the present application;

[0019] Figure 5 It is a schematic diagram of the logical structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0020] Many specific details are set forth in the following description in order to provide a thorough understanding of the present application. However, the present application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the connotation of the present application. Therefore, the present application is not limited by the specific implementations disclosed below.

[0021] First, some technical terms related to the present application are explained:

[0022] A voice repair model is usually used to improve or restore damaged voice segments or voice segments contaminated by noise. The above-mentioned voice repair model is also called a voice encoder, which is a device or software for converting an audio signal into a digital representation. This conversion process allows for efficient and accurate voice transmission through a digital data network (such as the Internet). The main role of a voice encoder is to convert analog voice into a digital signal and compress it for transmission over the network. The voice encoder converts the continuous analog voice signal into a discrete digital data stream through steps such as sampling, quantization, and encoding of the audio signal.

[0023] The basic architecture of a voice encoder generally includes an encoder, a quantizer, and an encoder. The encoder is responsible for encoding the input voice audio into a hidden representation; the quantizer then converts the hidden representation into a smaller amount of data; and the encoder takes the quantized data output by the quantizer as input for audio restoration.

[0024] Next, in order to facilitate the understanding of the training method for the voice repair model provided by the embodiments of the present application, before introducing the embodiments of the present application, the background of the embodiments of the present application is first introduced.

[0025] Voice repair technology can restore or correct damaged or lost voice signals, and thus achieve noise cancellation of voice segments or improve voice distortion. Therefore, voice repair technology is widely used in fields such as film and television drama dubbing, privacy information protection, and personalized voice synthesis.

[0026] In the existing speech restoration technology, speech restoration models can be divided into parallel corpus and non-parallel corpus types from the perspective of the type of data used; from the perspective of model structure, they can be divided into two-segment models and end-to-end models. However, from the perspective of the type of data used, the existing speech restoration models, on the one hand, use the parallel corpus method. It requires a data set of two or more people speaking the same content. This data set is expensive to collect and difficult to use in practice. Limited by the above-mentioned defects of the parallel corpus method, the non-parallel corpus method has become a research focus. The non-parallel corpus method has the characteristics of being easy to collect, and the speech restoration model of the non-parallel corpus can decouple the timbre features and semantic features in the audio information, and perform content restoration based on the semantic features. Therefore, the speech restoration model trained using the non-parallel corpus method is currently the mainstream. On the other hand, from the perspective of model structure: in the two-segment model, the acoustic model generates the spectrum, and the vocoder restores the audio from the spectrum. Since the two-segment model has the problem of feature mismatch that cannot be avoided when the two models are docked, and the sound quality of the generated model-repaired audio cannot be guaranteed, the end-to-end model has become a hot topic in current research.

[0027] From the above content, we can see that in the existing speech restoration technology, based on the non-parallel corpus method and the end-to-end model, when the speech restoration model corresponding to the non-parallel corpus method and the end-to-end model generates the model restoration audio, the model is often unable to accurately and completely extract the semantic content information in the audio sample to be restored, which leads to the restoration effect of the model restoration audio output by the model cannot be guaranteed. Therefore, how to train a speech restoration model so that the restoration audio output by the model has high semantic accuracy, good clarity and good sound quality has become a technical problem that needs to be solved urgently.

[0028] In view of the above problems existing in the prior art, the present application provides a method for training a speech restoration model. The speech restoration model obtained by the training method of the speech restoration model of this embodiment generates a model restoration audio with high audio semantic accuracy, good clarity and good sound quality.

[0029] After the background introduction of the above content, those skilled in the art can understand the problems existing in the prior art. Next, the application scenario of the training method of the speech repair model of the present application is described in detail. The training method of the speech repair model provided in the embodiment of the present application can be applied to the communication field or other related technical fields with speech repair model training requirements.

[0030] In the following, the application scenario of the training method of the speech restoration model according to the embodiment of the present application is firstly described by examples.

[0031] Figure 1 Schematic diagram of an application scenario of the training method of the speech restoration model provided in the first embodiment of the present application. Figure 1As shown in the figure, in this application scenario, there are a client 101 and a server 102. Among them, the client 101 and the server 102 are connected through network communication.

[0032] Taking Figure 1 as an example for detailed description, in the application background of training a voice repair model, the client 101 is used to obtain the audio sample to be repaired for model training as the initial voice repair model. The server 102 is used to set the initial voice repair model and the audio semantic content extraction model, and perform model training based on the above-mentioned audio sample to be repaired from the client 101 to obtain the voice repair model.

[0033] It should be noted that Figure 1 is a schematic diagram of the application scenario of a method for training a voice repair model provided by an embodiment of the present application. The embodiments of the present application do not limit Figure 1 the devices included therein, nor do they limit the number of the client 101 and the server 102. For example, in the application scenario satisfying Figure 1 shown, there may also be a data storage device. The data storage device may be an external memory relative to the client 101 and the server 102, or may be an internal memory integrated in the client 101 and the server 102. The client 101 may be a variety of devices with communication functions such as a smart phone, a smart bracelet, a tablet computer, a wearable device, a multimedia player, an e-reader, etc., and an application program (Application, APP) with an audio signal acquisition function is correspondingly installed on the device; the server 102 may be a single server or a cluster composed of several servers, or may be a cloud computing service center. In the embodiments of the present application, Figure 1 the number of the client 101 and the server 102 in

[0034] After introducing the application scenario of the embodiments of the present application, the present application also provides a method for training a voice repair model, as well as a device, an electronic device, and a computer-readable storage medium corresponding to the above method. The following provides embodiments to describe the above method, device, electronic device, and computer-readable storage medium in detail.

[0035] The second embodiment of the present application provides a method for training a voice repair model. Figure 2 is a flowchart of a method for training a voice repair model provided by an embodiment of the present application. The following combines Figure 2 to describe the method provided in this embodiment in detail. The embodiments involved in the following description are used to explain the principle of the method and are not limitations for actual use.

[0036] As Figure 2As shown in the figure, the training method of the voice repair model provided in this embodiment includes the following steps S201 to S204:

[0037] Step S201: Obtain an initial voice repair model to be trained, an audio sample to be repaired, and an audio semantic content extraction model.

[0038] The function of this step is to obtain an initial voice repair model, an audio sample to be repaired, and an audio semantic content extraction model for model training. Among them, the training materials for the initial voice repair model to perform model training include: the audio sample to be repaired.

[0039] The initial voice repair model is an algorithm or system capable of repairing damaged or degraded voice signals. This model can improve the voice quality by extracting useful information from the damaged voice audio signal and integrating it. Usually, the voice repair model performs voice repair on the audio sample to be repaired as the original input data, aiming to improve the intelligibility and audibility of the voice.

[0040] It should be understood that most of the existing voice repair models adopt the parallel corpus method, that is, the electronic larynx audio signal and the normal natural voice audio signal with the same voice content information are jointly used as the input data of the voice repair model to repair the electronic larynx audio signal. Usually, in this parallel corpus method, it is difficult to obtain the input data, resulting in a high difficulty in model repair and the inability to achieve real-time repair of the audio signal. Therefore, the voice repair model of the embodiment of this application adopts the non-parallel corpus method. In the non-parallel corpus method, the voice repair model can decouple the audio sample to be repaired to obtain the semantic features of the audio signal, and repair the audio signal according to the semantic features.

[0041] The audio sample to be repaired is a sample for training the initial voice repair model. In the embodiment of this application, in the application background of voice repair, the above-mentioned audio sample to be repaired can be used as the original input sample data for model training to train the initialized voice repair model. Specifically, the above-mentioned audio sample to be repaired can be a normal audio segment, or an audio segment contaminated by noise, distorted, or with missing sound.

[0042] In an embodiment of the present application, as a feasible implementation manner, the audio sample to be repaired is an electrolarynx audio signal. Here, an explanation of the electrolarynx audio signal is provided. Electrolarynx speech is a type of speech generated by an electrolarynx device (Electrolarynx, EL). Since electrolarynx audio signals usually produce speech with low naturalness, the speech has a high degree of mechanization, poor naturalness, and poor intelligibility and audibility. By using electrolarynx speech as the input data for the initial speech repair model, drawing on the characteristics of high mechanization and poor naturalness of electrolarynx speech itself, the training effect on the initial speech repair model is more significant.

[0043] In an embodiment of the present application, as another feasible implementation manner, the audio sample to be repaired can also be any audio signal collected in a real-world scenario. Any form of the above-mentioned speech audio signal can increase the randomness of model training during the training process of the initial speech repair model, thereby making the application of the trained speech repair model more extensive. The above examples of the audio sample to be repaired are for illustration purposes only and are not actual limitations.

[0044] The audio semantic content extraction model is an algorithm or system capable of extracting the semantic content in an audio sample. This model can extract meaningful semantic information from audio data that conforms to human understanding. It is usually applied to scenarios such as speech recognition, speech emotion analysis, keyword detection, topic recognition, and environmental sound understanding.

[0045] In an embodiment of the present application, the audio semantic content extraction model includes one of the following models: the HuBERT model, the Wav2Vec model, and the WavLM model. Among them, the HuBERT model is fully called the Hidden Unit BERT model, which is a self-supervised semantic recognition model; the Wav2Vec model is a deep learning model for automatic speech recognition. This model focuses on learning useful speech feature representations from raw audio signals through self-supervised learning without the need for a large amount of manually annotated transcript data. The WavLM model; the WavLM model is a self-supervised speech model based on the Transformer architecture. For ease of understanding, the HuBERT model is used as an example in the subsequent illustrations of the embodiments of the present application.

[0046] Through the above steps, the audio sample to be repaired for training the initial speech repair model and the audio semantic content extraction model (HuBERT model) are obtained.

[0047] Step S202: Input the audio sample to be repaired into the initial speech repair model to be trained for speech processing, and obtain a first semantic content representation representing the audio semantic content of the audio sample to be repaired and a model-repaired audio.

[0048] The function of this step is that based on the audio sample to be repaired, the initial speech repair model repairs the above-mentioned audio sample to be repaired, and obtains the first semantic content representation and the model-repaired audio.

[0049] In the embodiment of the present application, as a feasible implementation manner, inputting the audio sample of the model to be repaired into the initial speech repair model to obtain the first semantic content representation and the model-repaired audio for representing the audio semantic content of the audio sample to be repaired can be implemented in the following manner:

[0050] Step S202-1: Obtain the timbre representation corresponding to the audio sample to be repaired; the timbre representation is used for speech repair of the audio sample to be repaired;

[0051] Step S202-2: Input the audio sample to be repaired into the encoding quantization module in the initial speech repair model to be trained, and obtain the first semantic content representation;

[0052] Step S202-3: Input the first semantic content representation and the timbre representation corresponding to the audio sample to be repaired into the decoding module in the initial speech repair model to be trained, and generate the model-repaired audio.

[0053] The above-mentioned first semantic content representation refers to the semantic content features extracted by the initial speech repair model according to the audio sample to be repaired. It can also be understood that after the initial speech repair model decouples the above-mentioned audio sample to be repaired, the semantic content features corresponding to the audio content information for representing the audio sample to be repaired are obtained.

[0054] The timbre representation corresponding to the above-mentioned audio sample to be repaired includes: the sound quality feature of the audio sample to be repaired, and / or, the volume feature of the audio sample to be repaired, and / or, the emotional feature of the audio sample to be repaired. The sound quality feature of the audio sample to be repaired refers to the features for representing the pitch, frequency, and sense of space of the audio sample to be repaired. The volume feature of the audio sample to be repaired refers to the feature for representing the loudness or intensity of the audio sample to be repaired. The emotional feature of the audio sample to be repaired refers to the features for representing the intonation, speech rate, and rhythm of the audio sample to be repaired; with the timbre features of the above-mentioned audio sample to be repaired, the basic sound attributes of the audio sample to be repaired can be manifested. It can also be understood that the timbre representation corresponding to the specified timbre during the model training process of the initial speech repair model. It should be understood that the timbre information expressed by different specified timbre representations is different. Usually, in order to facilitate the distinction of the timbre representations corresponding to different audio samples to be repaired, the timbre representation corresponding to the audio sample to be repaired is represented by Speaker ID.

[0055] To facilitate the understanding of the process of obtaining the first semantic content representation and the model-repaired audio in the above-mentioned initial speech repair model, reference may be made to the schematic diagram of the principle of model training for the initial speech repair model provided in the embodiments of the present application. As Figure 3 shown, the data flow of the audio sample to be repaired is divided into three directions. One of the data flows is for the initial speech repair model, another data flow is for the audio semantic content extraction model (HuBERT model), and the other data flow is for the discriminant model (discriminator). It should be noted that when the audio sample to be repaired flows to the discriminator, specifically, it means that the audio sample to be repaired and the model-repaired audio are jointly input into the discriminator.

[0056] When the data flow of the audio sample to be repaired is for the initial speech repair model, through the coordinated work of the internal encoder, quantizer (encoding and quantization module), and decoder (decoding module) of the model, the first semantic content representation and the model-repaired audio are obtained.

[0057] As described above, the basic architecture of the speech repair model includes: an encoder (Encoder), a quantizer (quantizer), and a decoder (Decoder). In the embodiments of the present application, after the audio sample to be repaired is input into the initial speech repair model, it needs to be encoded by the encoder of the model first, and then can be input into the quantizer of the model for conversion processing, and finally input into the decoder of the model for the reconstruction of the audio signal to achieve the speech repair of the audio sample to be repaired.

[0058] For the convenience of understanding, the encoder, quantizer, and decoder involved in the above process will be described in detail respectively below.

[0059] The encoder in the above-mentioned initial speech repair model, as one of the component modules of the initial speech repair model, can convert the audio sample to be repaired input into the initial speech repair model into a representation form suitable for internal model calculation through the encoder. It can also be understood that by means of the encoder in the initial speech repair model, the audio sample to be repaired can be converted into a feature vector recognizable and calculable by the model. Usually, the encoder, as a key module of the speech repair model, can assist the model in correctly understanding and processing the audio sample to be repaired as the original input sample data. In specific implementation, the above-mentioned audio sample to be repaired is encoded by the encoder and converted into the first audio feature. In the embodiments of the present application, the first audio feature is represented by the vector α.

[0060] The encoder of the above initial voice repair model is connected to a quantizer. With the help of the quantizer, data discretization processing can be performed on the first audio feature α obtained by encoding the above audio sample to be repaired. It should be understood that in the embodiment of the present application, the role of the quantizer of the initial voice repair model is to decouple the first audio feature α to separate information such as semantic features and electronic larynx timbre features in the audio sample to be repaired (electronic larynx audio signal). It can also be understood that through the quantizer of the initial voice repair model, the semantic information in the first audio feature α is extracted, and then the first semantic content representation β corresponding to the audio sample to be repaired is obtained.

[0061] In the embodiment of the present application, as a feasible implementation manner, the above quantizer can also be a residual quantizer. The residual quantizer can convert the first audio feature α into discrete features, which is convenient for network transmission. It can also be understood that the continuous analog signal is converted into a discrete digital signal, thereby effectively reducing the complexity in the data processing process of the initial voice repair model.

[0062] The quantizer of the above initial voice repair model is connected to a decoder. The decoder is a module used to decode the feature vector generated by the encoder and restore and reconstruct the original audio signal as much as possible. Specifically, the decoder will generate a sound as similar as possible to the original voice signal by using probability models, waveform reconstruction, feature optimization, etc. on the basis of the feature vector provided by the encoder. In the embodiment of the present application, the decoder is used to perform feature vector decoding on the first semantic content representation β and the electronic larynx timbre representation corresponding to the audio sample to be repaired (electronic larynx audio signal), and then reconstruct the audio sample to generate a model-repaired audio. It should be understood that the purpose of the decoder to decode and restore the first semantic content representation β and the timbre representation corresponding to the audio sample to be repaired is to make the generated model-repaired audio as close as possible to the above audio sample to be repaired. It can also be understood that audio synthesis is performed on the first semantic content representation β and the timbre representation corresponding to the specified audio sample to be repaired.

[0063] Through the above steps, a model-repaired audio is generated according to the first semantic content representation corresponding to the audio sample to be repaired and the timbre representation corresponding to the audio sample to be repaired.

[0064] Step S203: Input the audio sample to be repaired into the audio semantic content extraction model for semantic extraction processing to obtain a second semantic content representation for representing the audio semantic content of the audio sample to be repaired.

[0065] The function of this step is to perform semantic extraction processing according to the audio semantic content extraction model to obtain the second semantic content representation of the audio sample to be repaired.

[0066] In an embodiment of the present application, as a feasible implementation manner, the above-mentioned input of the audio sample to be repaired into the audio semantic content extraction model for semantic extraction processing to obtain the second semantic content representation for representing the audio semantic content of the audio sample to be repaired can be implemented as follows:

[0067] Step S203-1: Obtain the HuBERT model for training the initial speech repair model.

[0068] Step S203-2: Input the audio sample to be repaired into the HuBERT model for semantic distillation processing to obtain the second semantic content representation for representing the audio semantic content of the audio sample to be repaired.

[0069] Among them, as mentioned above, the full name of the HuBERT model is the Hidden Unit BERT model, which is a self-supervised semantic recognition model. As a self-supervised semantic extraction model, this model is used to automatically extract audio content information from the audio sample to be repaired input into the model. In an embodiment of the present application, the above-mentioned HuBERT model can also be replaced by the Wav2Vec model, the WavLM model, etc. For the sake of easy understanding, the HuBERT model is used as an example in all embodiments of the present application.

[0070] It should be understood that the working principle of the HuBERT model is as follows: Adopting the principle of self-supervised learning, using unlabeled speech data to train a deep neural network to automatically learn how to extract useful features from the original audio signal. Specifically, first, the HuBERT model divides the audio signal into a series of short-time windows, and then encodes each window to obtain a series of hidden units. These hidden units can be regarded as a high-level abstraction of the original audio signal, containing rich semantic feature information. Secondly, the HuBERT model uses a clustering algorithm to group these hidden units to obtain a set of "phoneme" labels. These labels can be used for subsequent speech recognition tasks, such as recognizing the intention of the speaker or recognizing keywords in the speech. Since these labels are obtained through self-supervised learning, no additional manually labeled data is required, greatly reducing the training cost. In the training method of the speech repair model in an embodiment of the present application, the above-mentioned HuBERT model is a directly applicable model obtained by training with a large number of audio samples.

[0071] As described in step S202 above, the data flow of the audio sample to be repaired is divided into 3 directions. One of the data flows is the initial speech repair model, another data flow is the audio semantic content extraction model (HuBERT model), and there is also a data flow to the discriminant model (discriminator). Next, continue to refer to Figure 3The data flow of the audio sample to be repaired is described in detail for the audio semantic content extraction model (HuBERT model).

[0072] As Figure 3 shown, the above HuBERT model can be used to extract the second semantic content representation according to the audio sample to be repaired. For example, Hubert features. In the embodiments of the present application, the second semantic content representation is represented by the vector γ. The process of the HuBERT model extracting the second semantic content representation according to the audio sample to be repaired can refer to the introduction of the working principle of the HuBERT model above and will not be elaborated here. It should be understood that compared with the first semantic content representation β, since the acquisition source of the second semantic content representation γ is the HuBERT model and the acquisition source of the first semantic content representation β is the initial speech repair model, the second semantic content representation γ expresses the semantic information in the audio sample to be repaired more comprehensively and accurately. In the embodiments of the present application, since the initial speech repair model can obtain the first semantic content representation β of the audio sample to be repaired, on this basis, the semantic content representation (second semantic content representation γ) of the audio sample to be repaired is obtained through the HuBERT model, and the model parameters of the initial speech repair model can be adjusted according to the second semantic content representation γ and the first semantic content representation β.

[0073] Through the above steps, the audio sample to be repaired is input into the HuBERT model, and then the second semantic content representation γ for comprehensively and accurately representing the semantic information of the audio sample to be repaired is obtained.

[0074] Step S204, based on the difference between the model-repaired audio and the audio sample to be repaired, and the difference between the first semantic content representation and the second semantic content representation, adjust the model parameters of the initial speech repair model to obtain a speech repair model.

[0075] The function of this step is to obtain a speech repair model obtained by adjusting the model parameters of the initial speech repair model.

[0076] Among them, based on the difference between the model-repaired audio and the audio sample to be repaired, and the difference between the first semantic content representation and the second semantic content representation, adjusting the model parameters of the initial speech repair model to obtain a speech repair model is carried out as follows: adjusting the corresponding parameters of the encoding quantization module in the initial speech repair model with the goal of reducing the semantic loss value between the second semantic content representation and the first semantic content representation, and adjusting the corresponding parameters of the decoding module in the initial speech repair model with the goal of reducing the audio loss value between the model-repaired audio and the audio sample to be repaired, to obtain a speech repair model.

[0077] On the one hand, in the embodiments of the present application, the difference between the model-repaired audio and the audio sample to be repaired is obtained in the following manner:

[0078] S204-1-1, obtain a discriminant model for audio recognition and judgment processing;

[0079] S204-1-2, jointly input the model-repaired audio and the audio sample to be repaired into the discriminant model for audio recognition and judgment processing to obtain a discriminant result on whether the model-repaired audio is the audio sample to be repaired;

[0080] S204-1-3, if the discriminant result is no, calculate the audio loss value between the model-repaired audio and the audio sample to be repaired, and use the audio loss value as the difference between the model-repaired audio and the audio sample to be repaired.

[0081] For ease of understanding, continue to refer to the Figure 3 illustration. As Figure 3 shown, the data flow of the audio sample to be repaired is divided into 3 directions. One of the data flows is to the initial speech repair model, another data flow is to the audio semantic content extraction model (HuBERT model), and there is also a data flow to the discriminant model (discriminator). When the above-mentioned audio sample to be repaired flows to the discriminator, the data input to the discriminator specifically refers to the audio sample to be repaired and the model-repaired audio.

[0082] For ease of understanding, the discriminator involved in the above process will be described in detail next.

[0083] The discriminator, as a machine learning model, is used as the discriminator in a generative adversarial network (GAN for short), and learns by taking the audio sample to be repaired and the model-repaired audio as two neural networks and using the mutual game of the two neural networks. It can also be understood that the audio sample to be repaired and the model-repaired audio are jointly input into the discriminator to enable the discriminator to distinguish the differences between them and obtain a discriminant result; if the discriminant result of the discriminator is that it cannot distinguish between the audio sample to be repaired and the model-repaired audio, the model parameters of the initial speech repair model are not adjusted; if the discriminant result of the discriminator is that the model-repaired audio is not the audio sample to be repaired, the model parameters of the initial speech repair model are adjusted.

[0084] The above model-repaired audio can be understood as randomly sampled samples in the discriminator, and the audio to be repaired can be understood as real sampled samples. The role of the discriminator is to distinguish the model-repaired audio from the audio samples to be repaired as much as possible. By the mutual confrontation between two neural networks and continuously adjusting the model parameters of the initial speech repair model, the ultimate goal is to make the discriminator unable to determine whether the model-repaired audio generated by the initial speech repair model is real.

[0085] In specific implementation, when the audio to be repaired and the model-repaired audio are jointly input into the discriminator, by calculating the audio loss value between the model-repaired audio and the repaired audio samples, according to the audio loss value, a judgment result on whether the model-repaired audio is the repaired audio sample is obtained; and the model parameters of the initial speech repair model are adjusted according to the judgment result. In the embodiments of the present application, the specific implementation manner of calculating the audio loss value is the mean squared error loss function MSE Loss (Mean Squared Error Loss). As a common function, MSE Loss evaluates the performance of the model itself by calculating the mean squared error between the predicted result and the actual result. Specifically, given a true label y and the model prediction value The calculation formula for the loss value of MSELoss is as follows:

[0086]

[0087] Among them, in the above formula, N represents the number of samples; the audio samples to be repaired are used as the true label y; the model-repaired audio is used as the model prediction value The calculated loss value L represents the audio loss value between the audio samples to be repaired and the model-repaired audio. The smaller the audio loss value, the closer the model prediction result is to the actual result.

[0088] Among them, the model parameters of the initial speech repair model include: the corresponding parameters of the encoder, the corresponding parameters of the quantizer, and the corresponding parameters of the decoder. When the discriminator determines whether the model-repaired audio is the audio sample to be repaired, it can adjust the corresponding parameters of the above encoder, the preset parameters of the quantizer, and the corresponding parameters of the decoder with the help of the discrimination result. When the discrimination result shows that the model-repaired audio and the audio sample to be repaired cannot be distinguished, the adjustment of the above model parameters is stopped. It should be understood that in the actual application process, the adjustment of the model parameters of the initial speech repair model is carried out synchronously according to the audio loss value between the model-repaired audio and the audio sample to be repaired and the semantic loss value between the second semantic content representation and the first semantic content representation.

[0089] On the other hand, in the embodiments of the present application, the difference between the above-mentioned first semantic content representation and the second semantic content representation is obtained in the following manner: calculate the semantic loss value between the second semantic content representation and the first semantic content representation, and use the semantic loss value as the difference between the first semantic content representation and the second semantic content representation.

[0090] For ease of understanding, continue to refer to Figure 3 for the schematic illustration. As Figure 3 shown, the data flow of the audio sample to be repaired is divided into three directions. One of the data flows is to the initial speech repair model, another data flow is to the audio semantic content extraction model (HuBERT model), and the other data flow is to the discriminant model (discriminator). After the above-mentioned audio sample to be repaired flows to the HuBERT model, the HuBERT model outputs a second semantic content representation, such as Hubert features, which comprehensively and accurately expresses the semantic information in the audio sample to be repaired. The second semantic content representation is represented by the vector γ. Since the first semantic content representation β corresponding to the audio sample to be repaired is obtained in the foregoing step S202, by obtaining the semantic loss value of the first semantic content representation β relative to the second semantic content representation γ, the model parameters of the initial speech repair model are adjusted.

[0091] The above process involves calculating the semantic loss value between the second semantic content representation γ and the first semantic content representation β. The specific illustration of this process is as follows:

[0092] This process is a process of semantic distillation. Using the second semantic content representation γ as the target orientation and the HuBERT model as the teacher model, the student model composed of the encoder and quantizer of the initial speech repair model is guided to perform model learning; a preset loss function is used to measure the model performance of the student model composed of the encoder and quantizer of the initial speech repair model. That is, by calculating the semantic loss value between the first semantic content representation β generated by the student model composed of the encoder and quantizer of the initial speech repair model and the second semantic content representation γ as the target orientation, the corresponding parameters of the above-mentioned encoder and quantizer are adjusted respectively.

[0093] For the convenience of understanding the solution, a detailed description of semantic distillation processing is given here. Semantic Distillation is a machine learning method whose goal is to convert an original model into a more lightweight model while maintaining its performance. This technique is achieved by transferring knowledge from a large and complex model (teacher model) to a small and concise model (student model). Similar to traditional knowledge distillation methods, the goal of semantic distillation processing is to enable the student model (composed of the initial speech repair model encoder and quantizer) to imitate the features of the teacher model (HuBERT model); its model focuses on obtaining semantic information. This means that the student model (composed of the initial speech repair model encoder and quantizer) not only needs to use the second semantic content representation of the output data of the teacher model (HuBERT model) as the target guidance, but also needs to understand the semantic information of the audio sample to be repaired as the input data and imitate the behavior of the teacher model. Through the above processing, the knowledge of the large language model (teacher model - HuBERT model) can be transferred to the small model (student model - composed of the initial speech repair model encoder and quantizer), so that the latter can run on devices with limited resources.

[0094] In the embodiment of the present application, the above calculation of the semantic loss value between the second semantic content representation γ and the first semantic content representation β can also be implemented by the following loss function MSE Loss. Specifically, given a true label y and the model prediction value The formula for calculating the loss value of MSELoss is as follows:

[0095]

[0096] Among them, in the above formula, N represents the number of samples; the second semantic content representation γ is used as the true label y; the first semantic content representation β is used as the model prediction value of the student model The smaller the loss value L of MSELoss, the closer the model prediction result is to the actual result.

[0097] Through the above process, the semantic loss value between the second semantic content representation γ and the first semantic content representation γ, and the audio loss value between the model-repaired audio and the audio sample to be repaired are calculated. The above data is used as the basis for adjusting the model parameters of the initial speech repair model: guiding by reducing the audio loss value and the semantic loss value, the model parameters of the initial speech repair model are adjusted; during the adjustment process of the model parameters of the initial speech repair model, if both the audio loss value and the semantic loss value are less than their respective preset thresholds, the initial speech repair model corresponding to the current model parameters is used as the speech repair model.

[0098] In the encoder and quantizer of the embodiments of the present application, corresponding parameters are set. The above parameters can also be understood as initialized device parameters. The preset parameters of the quantizer are used to control the extraction of semantic features. Therefore, if one attempts to obtain a first semantic content representation with high accuracy and good descriptiveness, it is necessary to adjust the parameters corresponding to the quantizer. It should be understood that during the speech repair process, the extraction of the first semantic content representation will directly affect subsequent speech processing steps. Therefore, based on obtaining a first semantic content representation with high accuracy and clear and accurate semantic expression by adjusting the parameters corresponding to the encoder and quantizer, can the repair effect of the model repair audio repaired according to the first semantic content representation and the timbre representation corresponding to the audio sample to be repaired be ensured.

[0099] As mentioned above, the corresponding parameters in the quantizer are used to control the extraction of semantic features. If one attempts to obtain a first semantic content representation with high accuracy and good descriptiveness, it is necessary to obtain high-accuracy parameters corresponding to the quantizer. With the help of the second semantic content representation γ obtained by the HuBERT model, since the second semantic content representation γ represents the semantic information of the audio sample to be repaired more comprehensively and accurately than the first semantic content representation β. Correspondingly, the semantic loss value calculated according to the second semantic content representation and the first semantic content representation is more instructive for adjusting the preset parameters in the quantizer. Therefore, during the process of adjusting the model parameters of the initial speech repair model, the model parameters of the initial speech repair model can be adjusted with the guidance of reducing the audio loss value. When the audio loss values of the above two are less than the audio preset threshold, the adjustment of the model parameters is ended. On this basis, during the process of adjusting the model parameters of the initial speech repair model, the model parameters of the initial speech repair model are adjusted with the guidance of reducing the semantic loss value. When both the above semantic loss value and audio loss value are less than their respective preset thresholds, the adjustment of the model parameters is ended. The above process can be understood as adjusting the model parameters of the initial speech repair model according to both the audio loss value and the semantic loss value. In the embodiments of the present application, the above preset threshold can be the threshold for model initialization or the threshold preset according to training requirements. This embodiment is for illustration and not an actual limitation.

[0100] Through the above steps of the embodiments of the present application, a speech repair model obtained by training the initial speech repair model is obtained. The speech repair model has the following advantages: The semantics of the model output model repair audio are clear, the intelligibility is high, and the sound quality is good.

[0101] In an embodiment of the present application, as a feasible implementation manner, the audio sample to be repaired for model training may also be randomly adjusted according to the following manner. For example, randomly adjust the fundamental frequency of the audio signal of the audio sample to be repaired, the mean value and range of the formants of the audio signal, and randomly adjust the speech rate of the audio signal in proportion. It should be understood that by using the above data augmentation method, the diversity of the training samples used for training the initial speech repair model can be ensured. With a large number of training samples, the repair effect of the audio signal by the speech repair model obtained after training can be achieved. Further, different types of audio samples to be repaired can also improve the diversity of the model's recognition of rhythm, which helps the model to improve its robustness.

[0102] In an embodiment of the present application, as a feasible implementation manner, the above speech repair model can also be used for repairing audio samples. Specifically, during implementation, obtain the audio sample for speech repair; input the audio sample for speech repair into the speech repair model to obtain the repaired audio sample. Of course, the speech repair model of this embodiment can be applied to the repair processing of speech samples in real-time communication. It should be understood that during the application process of the speech repair model, the model is composed of an encoder, a quantizer, and a decoder, and neither the HuBERT model and discriminator described above are set. Through the above speech repair model, on the premise of ensuring a high similarity between the converted audio and the original audio, the audio signal with strong mechanical feeling and poor naturalness is repaired into natural speech audio with high sound quality.

[0103] For the training method of the speech repair model provided by the embodiment of the present application, during the training process of the initial speech repair model, obtain the initial speech repair model to be trained, the audio sample to be repaired, and the audio semantic content extraction model; input the audio sample to be repaired into the initial speech repair model to be trained for speech processing to obtain the first semantic content representation representing the audio semantic content of the audio sample to be repaired and the model-repaired audio. Input the audio sample to be repaired into the audio semantic content extraction model for semantic extraction processing to obtain the second semantic content representation representing the audio semantic content of the audio sample to be repaired. On the basis of obtaining the first semantic content representation, the second semantic content representation, and the model-repaired audio, continuously adjust the model parameters of the initial speech repair model based on the difference between the model-repaired audio and the audio sample to be repaired and the difference between the first semantic content representation and the second semantic content representation. The obtained speech repair model has high semantic accuracy, good intelligibility, and good sound quality for the repaired audio generated by the model itself.

[0104] The above second embodiment provides a method for training a voice repair model. Correspondingly, an embodiment of the present application also provides a device for training a voice repair model. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple. For the details of the relevant technical features, please refer to the corresponding description of the method embodiment provided above. The following description of the device embodiment is only illustrative. As Figure 4 shown, it is a block diagram of the units of the voice repair model training device 400 provided in this embodiment, including:

[0105] A data acquisition unit 401, configured to acquire an initial voice repair model to be trained, an audio sample to be repaired, and an audio semantic content extraction model.

[0106] A sample processing unit 402, configured to input the audio sample to be repaired into the initial voice repair model to be trained for voice processing, and obtain a first semantic content representation for representing the audio semantic content of the audio sample to be repaired and a model-repaired audio.

[0107] A characterization extraction unit 403, configured to input the audio sample to be repaired into the audio semantic content extraction model for semantic extraction processing, and obtain a second semantic content representation for representing the audio semantic content of the audio sample to be repaired.

[0108] A parameter adjustment unit 404, configured to adjust the model parameters of the initial voice repair model based on the difference between the model-repaired audio and the audio sample to be repaired, and the difference between the first semantic content representation and the second semantic content representation, to obtain a voice repair model.

[0109] The above embodiment provides a device for training a voice repair model. In addition, an embodiment of the present application also provides an electronic device. Since the electronic device embodiment is basically similar to the method embodiment, the description is relatively simple. For the details of the relevant technical features, please refer to the corresponding description of the method embodiment provided above. The following description of the electronic device embodiment is only illustrative. The electronic device embodiment is as follows: Please refer to Figure 5 Understand this embodiment, Figure 5 It is a schematic diagram of the electronic device provided in this embodiment.

[0110] As Figure 5 shown, Figure 5 It is a schematic diagram of an electronic device provided in an embodiment of the present application.

[0111] In this embodiment, an optional hardware structure of the electronic device 500 can be as Figure 5As shown, it includes: at least one processor 501, at least one memory 502, and at least one communication bus 505; the memory 502 contains a program 503 and data 504.

[0112] The bus 505 can be a communication device that transfers data between components inside the electronic device 500, such as an internal bus (e.g., a CPU-memory bus, where the processor is the central processing unit, abbreviated as CPU), an external bus (e.g., a universal serial bus port, a peripheral component interconnect express port), etc.

[0113] In addition, the electronic device also includes: at least one network interface 506, at least one peripheral interface 507. The network interface 506 provides wired or wireless communication related to an external network 508 (e.g., the Internet, an intranet, a local area network, a mobile communication network, etc.); in some embodiments, the network interface 506 can include any combination of network interface controllers (abbreviated as NIC), radio frequency (abbreviated as RF) modules, repeaters, transceivers, modems, routers, gateways, any combination of wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near field communication (abbreviated as NFC) adapters, cellular network chips, etc.

[0114] The peripheral interface 507 is used to connect to peripherals, and the peripherals can be, for example, peripheral 1 ( Figure 5 in the figure as 509), peripheral 2 ( Figure 5 in the figure as 510), and peripheral 3 ( Figure 5 in the figure as 511). Peripherals are peripheral devices, and peripheral devices can include, but are not limited to, cursor control devices (e.g., a mouse, a touchpad, or a touch screen), a keyboard, a display (e.g., a cathode ray tube display, a liquid crystal display, a light emitting diode display), a video input device (e.g., a camera or an input interface communicatively coupled to a video archive), etc.

[0115] The processor 501 may be a CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application.

[0116] The memory 502 may include a high-speed RAM (Random Access Memory) memory, and may also include a non-volatile memory, such as at least one disk memory.

[0117] Among them, the processor 501 invokes the programs and data stored in the memory 502 to execute the method of the second embodiment of this application.

[0118] Fifth Embodiment

[0119] Corresponding to the method of the second embodiment of this application, the fifth embodiment of this application also provides a computer storage medium. The computer storage medium stores a computer program, and this computer program is run by a processor to execute the method of the second embodiment of this application.

[0120] Although this application is disclosed above with preferred embodiments, it is not used to limit this application. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of the present invention. Therefore, the protection scope of this application should be subject to the scope defined by the claims of this application.

[0121] The embodiments of this application may involve the use of user data. In practical applications, it is possible to use user-specific personal data in the solutions described in this article within the scope permitted by applicable laws and regulations (for example, with the user's explicit consent, giving the user a practical notice, etc.) and in compliance with the requirements of the applicable laws and regulations of the country where it is located. In the above embodiments, a training method for a voice repair model, as well as a device and an electronic device corresponding to the above method, are provided. In addition, the embodiments of this application also provide a computer-readable storage medium for implementing the training method of the above voice repair model. The embodiments of the computer-readable storage medium provided by this application are described relatively simply. For the corresponding descriptions of the relevant parts, please refer to the corresponding descriptions of the above method embodiments. The embodiments described below are merely illustrative.

[0122] The computer-readable storage medium provided in this embodiment stores computer instructions, and when these instructions are executed by a processor, the steps of the above method embodiments are implemented.

[0123] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.

[0124] The memory may include non-permanent memory in the computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of a computer-readable medium.

[0125] 1. A computer-readable medium includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information accessible by a computing device. As defined herein, a computer-readable medium does not include transitory media such as modulated data signals and carrier waves.

[0126] 2. Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0127] Although the present invention is disclosed above in preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of the present invention. Therefore, the protection scope of the present invention should be determined by the scope defined in the claims of the present invention.

[0128] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject.

Claims

1. A training method for a voice repair model, characterized in that, The method includes: Obtaining an initial speech repair model to be trained, an audio sample to be repaired, and an audio semantic content extraction model; Inputting the audio sample to be repaired into the initial speech repair model to be trained for speech processing, to obtain a first semantic content representation for representing the audio semantic content of the audio sample to be repaired and a model-repaired audio; Inputting the audio sample to be repaired into the audio semantic content extraction model for semantic extraction processing, to obtain a second semantic content representation for representing the audio semantic content of the audio sample to be repaired; Adjusting the model parameters of the initial speech repair model based on the difference between the model-repaired audio and the audio sample to be repaired, and the difference between the first semantic content representation and the second semantic content representation, to obtain a speech repair model.

2. The training method of the voice restoration model according to claim 1, wherein The inputting the audio sample of the to-be-repaired model into the initial speech repair model to be trained for speech processing, to obtain a first semantic content representation for representing the audio semantic content of the audio sample to be repaired and a model-repaired audio, includes: Obtaining a timbre representation corresponding to the audio sample to be repaired; the timbre representation is used for speech repair of the audio sample to be repaired; Inputting the audio sample to be repaired into an encoding quantization module in the initial speech repair model to be trained, to obtain the first semantic content representation; Jointly inputting the first semantic content representation and the timbre representation corresponding to the audio sample to be repaired into a decoding module in the initial speech repair model to be trained, to generate the model-repaired audio.

3. The training method of the voice repair model according to claim 2, wherein The audio semantic content extraction model includes one of the following models: HuBERT model, Wav2Vec model, WavLM model.

4. The training method of the voice repair model according to claim 3, wherein The inputting the audio sample to be repaired into the audio semantic content extraction model for semantic extraction processing, to obtain a second semantic content representation for representing the audio semantic content of the audio sample to be repaired, includes: Inputting the audio sample to be repaired into the HuBERT model for semantic distillation processing, to obtain a second semantic content representation for representing the audio semantic content of the audio sample to be repaired.

5. The training method of the voice repair model according to claim 2, characterized in that The difference between the model-repaired audio and the audio sample to be repaired is obtained in the following manner: Obtaining a discriminant model for audio recognition judgment processing; Jointly inputting the model-repaired audio and the audio sample to be repaired into the discriminant model for audio recognition judgment processing, to obtain a discriminant result as to whether the model-repaired audio is the audio sample to be repaired; If the discriminant result is negative, calculating an audio loss value between the model-repaired audio and the audio sample to be repaired, and using the audio loss value as the difference between the model-repaired audio and the audio sample to be repaired.

6. The training method of the voice repair model according to claim 4, wherein The difference between the first semantic content representation and the second semantic content representation is obtained in the following manner: Calculating a semantic loss value between the second semantic content representation and the first semantic content representation, and using the semantic loss value as the difference between the first semantic content representation and the second semantic content representation.

7. The training method of the voice repair model according to claim 6, characterized in that Adjusting the model parameters of the initial speech restoration model based on the differences between the audio restored based on the model and the audio sample to be restored, and between the first semantic content representation and the second semantic content representation, to obtain a speech restoration model, including: Adjusting the corresponding parameters of the encoding quantization module in the initial speech restoration model with the orientation of reducing the semantic loss value between the second semantic content representation and the first semantic content representation, and adjusting the corresponding parameters of the decoding module in the initial speech restoration model with the orientation of reducing the audio loss value between the audio restored by the model and the audio sample to be restored, to obtain a speech restoration model.

8. The training method of the voice repair model according to claim 2, characterized in that The timbre representation corresponding to the audio sample to be restored includes: the sound quality feature of the audio sample to be restored, and / or, the volume feature of the audio sample to be restored, and / or, the emotion feature of the audio sample to be restored.

9. A training device for a voice repair model, characterized in that, Including: A data acquisition unit configured to acquire an initial speech restoration model to be trained, an audio sample to be restored, and an audio semantic content extraction model. A sample processing unit configured to input the audio sample to be restored into the initial speech restoration model to be trained for speech processing, to obtain a first semantic content representation representing the audio semantic content of the audio sample to be restored and an audio restored by the model. A representation extraction unit configured to input the audio sample to be restored into the audio semantic content extraction model for semantic extraction processing, to obtain a second semantic content representation representing the audio semantic content of the audio sample to be restored. A parameter adjustment unit configured to adjust the model parameters of the initial speech restoration model based on the differences between the audio restored by the model and the audio sample to be restored, and between the first semantic content representation and the second semantic content representation, to obtain a speech restoration model.

10. An electronic device, characterized in that, Including a processor and a memory; wherein, The memory is used to store one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the method according to any one of claims 1-8.

11. A computer-readable storage medium having one or more computer instructions stored thereon, characterized in that, The instruction is executed by the processor to implement the method according to any one of claims 1-8.