Speech restoration model training method and device, and speech restoration method and device
By acquiring descriptive information and feature mappings from damaged speech data, a speech restoration model is trained, which solves the problem of insufficient generalization in existing technologies and achieves more accurate and broader speech data restoration.
Patent Information
- Application Number
- CN202410626045.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-20
- Publication Date
- 2025-11-21
AI Technical Summary
Existing speech restoration models lack generalization ability when faced with damaged speech data caused by unknown noise, reverberation environment, or device parameters, resulting in poor restoration effects.
By acquiring the damage description information of damaged speech data, a speech restoration feature mapping relationship is established using text encoding and speech prior network. The speech restoration model is then trained by combining speakerprint features and restoration tasks to improve the model's generalization and restoration accuracy.
The ability of the speech restoration model to restore damaged speech data under unknown noise and reverberation environments has been enhanced, improving the restoration effect and generalization.
Smart Images

Figure CN120998208A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech technology, and more specifically, to a training method for a speech restoration model, a speech restoration method, and an apparatus in the field of speech technology. Background Technology
[0002] During the recording, transmission, and storage of voice data, issues such as noise, reverberation, low resolution, and clipping may occur, resulting in damage to the voice data.
[0003] In related technologies, when repairing damaged speech data, it is necessary to determine the type of damage to the damaged speech data from noise, reverberation, and low resolution, and then determine the corresponding repair method based on the damage type. For example, if the damage type is noise, the repair method can be noise reduction; if the damage type is reverberation, the repair method can be dereverberation; if the damage type is low resolution, the repair method can be super-resolution; and if the damage type is clipping, the repair method can be clipping repair.
[0004] When repairing damaged speech data using repair methods determined by the type of damage, it is necessary to train models corresponding to different repair methods. During training, known noise data, reverberation environments, or device parameters are typically used. This can limit the generalization ability of the pre-trained repair model, resulting in poor repair performance when the damaged speech data is caused by unknown noise data, reverberation environments, or device parameters.
[0005] Therefore, improving the restoration effect of damaged voice data is an urgent problem to be solved. Summary of the Invention
[0006] This application provides a speech restoration method, speech restoration device, electronic device, and storage medium. The method can improve the generalization and restoration capability of the speech restoration model, as well as the restoration effect on damaged speech data.
[0007] Firstly, a training method for a voice restoration model is provided, which is applied to electronic devices and includes:
[0008] In cases where damaged speech data exists in the original speech data, damage description information of the damaged speech data is obtained; wherein, the damage description information is used to indicate the cause of damage to the damaged speech data; based on the damage description information, speech restoration features for repairing the damaged speech data are obtained; the speech restoration model is trained by using the speech features corresponding to the damaged speech data as input data, the speech restoration features as prior information, and the sample speech features as target output data, thereby obtaining the trained speech restoration model; wherein, the sample speech features are used to represent the speech features when the damaged speech data is not damaged.
[0009] In this embodiment, since the damage description information of damaged speech data can represent the cause of the damage, during the training of the speech restoration model, using the speech restoration features obtained through the damage description information as prior information to guide the training of the speech restoration model can improve the ability of the trained speech restoration model to restore damaged speech data based on the damage description information, making the restored speech data more accurate, thereby improving the restoration effect of the trained speech restoration model on damaged speech data.
[0010] Furthermore, since the damage description information of damaged speech data is not limited by the training set of the speech restoration model, the damage description information can be used to represent the cause of damage to any damaged speech data during the training process of the speech restoration model. The speech restoration features obtained through the damage description information for repairing damaged speech data can be used as prior information to guide the training of the speech restoration model, thereby making the trained speech restoration model have stronger generalization ability for repairing damaged speech data.
[0011] In conjunction with the first aspect, in certain implementations of the first aspect, the aforementioned speech restoration features for repairing damaged speech data, based on the damage description information, include:
[0012] Obtain the text corresponding to the damaged description information; encode the text in a pre-trained text-speech model to obtain the encoded text features; input the text features into a speech prior network model; in the speech prior network model, map the speech restoration features corresponding to the text features through a pre-trained mapping relationship.
[0013] In this embodiment of the application, by describing the damage description information corresponding to the damaged speech data through text, the problem of deviation in the damage description information caused by the interference of surrounding environmental noise when describing the damage description information through speech or other means can be avoided. Thus, during the training process of the speech restoration model, the accuracy of obtaining the damage description information can be improved. Furthermore, with more accurate damage description information, the trained speech restoration model can also restore the damaged speech data more accurately.
[0014] In conjunction with the first aspect and the above implementation methods, in some implementations of the first aspect, the method further includes:
[0015] Extracting target voiceprint features from the original speech data; the above-mentioned training of the speech restoration model using the speech features corresponding to the damaged speech data as input data, the speech restoration features as prior information, and the sample speech features as target output data to obtain the trained speech restoration model includes: training the speech restoration model using the speech features corresponding to the damaged speech data and the target voiceprint features as input data, the speech restoration features as prior information, and the sample speech features as target output data to obtain the trained speech restoration model.
[0016] In this embodiment of the application, during the training process of the speech restoration model, the voiceprint features in the original speech data are introduced, which enables the trained speech restoration model to have the function of fusing voiceprint features. This allows the speech data restored by the trained speech restoration model to be consistent with the voiceprint information in the original speech data, further improving the restoration effect of the trained speech restoration model on damaged speech data.
[0017] In conjunction with the first aspect and the above implementation methods, in some implementations of the first aspect, the method further includes:
[0018] The task of obtaining damaged speech data includes: using the speech features corresponding to the damaged speech data and the target voiceprint features as input data, the speech restoration features as prior information, and the sample speech features as the target output data to train the speech restoration model, thereby obtaining the trained speech restoration model.
[0019] In this embodiment, since the causes of damage to different damaged speech data may vary, using the same restoration method to restore speech data with different causes may result in significant deviations in the restored speech data. Therefore, a restoration task corresponding to the damaged speech data can be obtained and introduced into the training process of the speech restoration model. This allows the speech restoration model to be trained based on the restoration task corresponding to the damaged speech data, making the restoration function of the trained speech restoration model more closely aligned with the causes of damage to the damaged speech data and more targeted. This leads to accurate restoration of the damaged speech data and improves the accuracy of damaged speech data restoration.
[0020] In conjunction with the first aspect and the above-described implementation methods, in some implementation methods of the first aspect, the acquisition of damage description information of damaged speech data includes:
[0021] If multiple damaged speech data are acquired, obtain the damage description information for each damaged speech data in the multiple damaged speech data; or, if multiple damaged speech data are acquired, select a preset number of target damaged speech data in the multiple damaged speech data; obtain the damage description information for the target damaged speech data.
[0022] In this embodiment, multiple damaged speech data can be used as a sample training set, and the damage description information of each damaged speech data in the multiple damaged speech data can be used as prior information to train the speech restoration model. This makes the speech restoration model more accurate in restoring damaged speech data after being trained on a large number of sample damaged speech data, thereby improving the restoration effect of the trained speech restoration model on damaged speech data.
[0023] Alternatively, a preset number of target damaged speech data can be selected from multiple damaged speech data, and the damage description information corresponding to the target damaged speech data can be used as prior information to train the speech restoration model. This can avoid the problem that the trained speech restoration model cannot restore the damaged speech data when it cannot obtain the damage description information corresponding to the damaged speech data, thus improving the generalization of the trained speech restoration model.
[0024] Secondly, a method for voice restoration is provided, which is applied to electronic devices and includes:
[0025] If damaged speech data exists in the original speech data, it is determined whether damage description information of the damaged speech data has been obtained. The damage description information is used to indicate the cause of the damage to the speech data. If damage description information is obtained, speech restoration features for repairing the damaged speech data are obtained based on the damage description information. The speech restoration features are used as prior information and the damaged speech data are input into a pre-trained speech restoration model for repair, resulting in repaired speech data. The pre-trained speech restoration model is a restoration model trained using the speech features corresponding to the damaged speech data as input data, the speech features for repairing the damaged speech data obtained from the damage description information of the damaged speech data as training prior information, and the speech features of the undamaged speech data as target output data.
[0026] In this embodiment of the application, if the original speech data contains damaged speech data, and if the descriptive information corresponding to the damaged speech data can be obtained, the damaged descriptive information can be used as prior information to guide the pre-trained speech restoration model to restore the damaged speech data when the pre-trained speech restoration model restores the damaged speech data. This can make the pre-trained speech restoration model restore the damaged speech data more accurately, thereby making the restored speech data more accurate.
[0027] In conjunction with the second aspect, in some implementations of the second aspect, the aforementioned speech restoration features are used as prior information and input into a pre-trained speech restoration model along with the damaged speech data for restoration, resulting in restored speech data, including:
[0028] If a task to repair damaged speech data is obtained, the prior information, damaged speech data and repair task are input into a pre-trained speech repair model for repair, and the repaired speech data is obtained.
[0029] In this embodiment, since the causes of damage to different damaged speech data may vary, using the same restoration method to restore speech data with different causes may result in significant deviations in the restored speech data. Therefore, when the pre-trained speech restoration model restores damaged speech data, a restoration task corresponding to the damaged speech data can be introduced. This allows the pre-trained speech restoration model to restore the damaged speech data based on the corresponding restoration task, making the restoration more accurate and resulting in more accurate restored speech data.
[0030] In conjunction with the second aspect and the above implementation methods, in some implementations of the second aspect, the method further includes:
[0031] If the damaged description information is not obtained, the damaged speech data is input into a pre-trained speech restoration model for restoration to obtain the restored speech data.
[0032] In this embodiment, because prior information can be omitted during the training of the pre-trained speech restoration model for some samples of damaged speech data, the pre-trained speech restoration model can still restore the damaged speech data even when it cannot obtain the corresponding damage description information, thus improving the restoration effect of the damaged speech data.
[0033] Thirdly, a training device for a speech restoration model is provided, characterized in that the device is configured in an electronic device and includes:
[0034] The model is used to obtain damage description information of damaged speech data when damaged speech data exists in the original speech data; the damage description information is used to indicate the cause of damage to the damaged speech data.
[0035] A processing model is used to obtain speech restoration features for repairing damaged speech data based on the damaged description information;
[0036] The training model is used to train the speech restoration model by taking the speech features corresponding to the damaged speech data as input data, the speech restoration features as prior information, and the sample speech features as target output data, and obtain the trained speech restoration model; wherein, the sample speech features are used to represent the speech features when the damaged speech data is not damaged.
[0037] Fourthly, a speech restoration apparatus is provided, characterized in that the apparatus is configured in an electronic device and includes:
[0038] The model is used to determine whether damage description information of the damaged speech data has been obtained when damaged speech data exists in the original speech data; the damage description information is used to indicate the cause of the damage to the damaged speech data.
[0039] The processing model is used to obtain speech restoration features for repairing damaged speech data based on the damage description information if the damage description information is obtained.
[0040] The repair model is used to input the damaged speech data into a pre-trained speech repair model as prior information to repair the speech data.
[0041] The pre-trained speech restoration model is trained by using the speech features corresponding to the damaged speech data as input data, the speech features obtained from the damage description information corresponding to the damaged speech data as training prior information, and the speech features of the damaged speech data when it is undamaged as target output data.
[0042] Fifthly, an electronic device is provided, including a memory and a processor. The memory is used to store executable program code, and the processor is used to call and run the executable program code from the memory, causing the electronic device to perform the methods described in the first aspect or any one of the first aspects, or the second aspect or any possible implementation thereof.
[0043] In a sixth aspect, a computer program product is provided, comprising: computer program code, which, when run on a computer, causes the computer to perform the methods described in the first aspect or any one of the first aspects, or the second aspect or any possible implementation thereof.
[0044] In a seventh aspect, a computer-readable storage medium is provided that stores computer program code, which, when executed on a computer, causes the computer to perform the methods described in the first aspect or any one of the first aspects, or in the second aspect or any possible implementation thereof. Attached Figure Description
[0045] Figure 1 This is a schematic diagram of a speech restoration scenario in related technologies.
[0046] Figure 2 This is a schematic diagram of a voice restoration model provided in an embodiment of this application.
[0047] Figure 3 This is a flowchart illustrating a training method for a speech restoration model provided in an embodiment of this application.
[0048] Figure 4 This is a flowchart illustrating another training method for a speech restoration model provided in this application embodiment.
[0049] Figure 5 This is a flowchart illustrating a voice restoration method provided in an embodiment of this application.
[0050] Figure 6 This is a schematic diagram of the structure of the training device for the speech restoration model provided in this application embodiment.
[0051] Figure 7 This is a schematic diagram of the speech restoration device provided in the embodiments of this application.
[0052] Figure 8 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0053] The technical solutions in this application will be clearly and thoroughly described below with reference to the accompanying drawings. In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B. "And / or" in the text is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Furthermore, in the description of the embodiments of this application, "multiple" refers to two or more than two.
[0054] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature.
[0055] Voice data may be damaged during recording, transmission, and storage due to noise, reverberation, low resolution, and clipping. Figure 1 Taking noise as an example, this paper introduces methods for repairing damaged speech data in related technologies.
[0056] Figure 1 This is a schematic diagram of a speech restoration scenario in related technologies.
[0057] For example, such as Figure 1 As shown, Figure 1 It includes a microphone 110, voice data 120, environmental noise data (hereinafter referred to as "noise data") 130, and an electronic device 140. The microphone 110 can be used to collect voice data 120 and noise data 130, and transmit the collected voice data 120 and noise data 130 to the electronic device 140 through a wireless network or a wired connection, so that the electronic device 140 can process the received voice data 120 and noise data 130.
[0058] For example, when the microphone 110 collects voice data 120, it may collect noise data 130 (e.g., car horn). The presence of noise data 130 will cause the voice data 120 collected by the microphone 110 to have poor clarity or some missing data, which will damage the voice data 120, that is, there is damaged voice data, thereby affecting the integrity and clarity of the voice data 120.
[0059] To repair the damaged voice data 120, the received damaged voice data 120 can be repaired using a voice repair model set in the electronic device 140. Currently, a model with noise reduction function can be used to denoise the noise data 130 in the damaged voice data 120 to repair the damaged voice data 120.
[0060] In addition to noise, speech data is also affected by reverberation, low resolution, and clipping. To repair damaged speech data, the type of damage can be determined from reverberation and low resolution, and a corresponding repair model can be used to repair the damaged speech data based on this type. For example, if the damage type is reverberation, a model with reverberation removal capabilities can be used to repair the damaged speech data; if the damage type is low resolution, a model with super resolution capabilities can be used; and if the damage type is clipping, a model with clipping repair capabilities can be used.
[0061] Using different restoration models to restore damaged speech data results in models with limited functionality. To improve the functionality of restoration models, multi-functional restoration models can be constructed, enabling multiple restoration functions such as noise reduction, dereverberation, super-resolution, and clipping restoration within a unified model. However, the training process of a unified restoration model typically uses known noise data, reverberation environments, or device parameters, which can limit the generalization ability of the unified restoration model. Consequently, when restoring damaged speech data caused by unknown noise data, reverberation environments, or device parameters, the unified restoration model cannot effectively restore the damaged speech data, resulting in poor restoration performance. Therefore, to improve the restoration effect on damaged speech data, this application proposes a speech restoration model training method, a speech restoration method, and an apparatus.
[0062] In addition, the following explanations are provided for the technical terms mentioned above:
[0063] Noise reduction involves identifying and processing noise from speech acquisition devices (e.g., microphones) and the environment to preserve key information in speech data and reduce noise interference, thereby improving the clarity and intelligibility of the speech data. Common noise reduction methods include, but are not limited to, time-domain or frequency-domain filter-based noise reduction models and machine learning-based noise reduction models.
[0064] De-reverberation: Eliminating reverberation effects in speech data to improve its clarity. Reverberation is caused by sound reflection and refraction in enclosed spaces, affecting the clarity of acquired speech data. Common de-reverberation methods include, but are not limited to, digital signal de-reverberation processing techniques and convolutional neural networks (CNNs).
[0065] Super-resolution: Upgrading low-resolution speech data to high-resolution speech data can improve the spatial or frequency resolution of the speech data, thereby enhancing its granularity and clarity. Common super-resolution methods include, but are not limited to, interpolation, generative adversarial networks (GANs), and convolutional neural networks.
[0066] Amplitude clipping correction: This involves identifying and repairing distortion in speech data caused by clipping, ensuring the speech data remains within a certain dynamic range. Clipping occurs when the amplitude of the speech data signal is too large during acquisition or transmission, exceeding the maximum amplitude range that the speech acquisition or transmission equipment can handle. This results in the loss of the excess amplitude, causing speech data distortion. Common clipping correction methods include, but are not limited to, rescaling and restoring the speech data.
[0067] It should be noted that electronic device 140 can be a smart device with voice data repair capabilities, including but not limited to: personal computers, tablets, handheld devices, in-vehicle devices, wearable devices, computing devices, or other processing devices connected to a wireless modem. Electronic devices may have different names in different networks, such as: user equipment, access electronic device, user unit, user station, mobile station, mobile station, remote station, remote electronic device, mobile device, user electronic device, electronic device, wireless communication device, user agent or user device, cellular phone, cordless phone, electronic device in a 5G network or future evolved network, etc. The embodiments in this application are not limited in scope.
[0068] The following is combined Figures 2 to 5 The training method of the speech restoration model provided in the embodiments of this application is described in detail.
[0069] Figure 2 This is a schematic diagram of a voice restoration model provided in an embodiment of this application.
[0070] For example, such as Figure 2 As shown, the large-scale speech restoration model 200 includes a speech coding model 210, a voiceprint feature extraction model 220, a text-speech model 230, a speech prior network model 240, a diffusion model 250, and a vocode model 260. The text description and damaged speech data serve as the input to the large-scale speech restoration model 200, and the restored speech data serves as the output of the large-scale speech restoration model 200.
[0071] Furthermore, the speech coding model 210, the voiceprint feature extraction model 220, the text-speech model 230, the speech prior network model 240, the diffusion model 250, and the vocode model 260 can be used as sub-models in the large speech restoration model 200. The speech coding model 210, the voiceprint feature extraction model 220, the text-speech model 230, the speech prior network model 240, and the vocode model 260 are all pre-trained. In this embodiment, the restoration function of the diffusion model 250 is mainly trained. For ease of understanding, the diffusion model 250 will be referred to as the "speech restoration model" below.
[0072] Optionally, the speech coding model 210, the voiceprint feature extraction model 220, the text speech model 230, the speech prior network model 240, the diffusion model 250 and the vocode model 260 can be trained simultaneously to obtain a large speech restoration model 200 with speech restoration function.
[0073] Speech coding model 210 can be used to receive input damaged speech data and, using a self-supervised speech representation model with contrastive learning and masked language learning capabilities, discretize the continuous speech data into a finite number of distinguishable speech features. These speech features are then learned by solving a masked prediction task, thereby obtaining speech features with contextual information. Finally, the speech features with contextual information are transmitted to diffusion model 250.
[0074] The voiceprint feature extraction model 220 can be used to receive input damaged speech data and extract voiceprint features from the damaged speech data using methods such as residual network models or local and global feature fusion; wherein, the voiceprint features include, but are not limited to, pitch, stress, pronunciation order and prosody. The extracted voiceprint features are then transmitted to the diffusion model 250.
[0075] The text-to-speech model 230 can receive text descriptions related to the causes of damage in damaged speech data; for example, "the noise is a car horn" or "an echo produced in a 10-square-meter room." This text description can be the cause of damage corresponding to the damaged speech data actively input by the user, or it can be generated by analyzing the causes of damage in the damaged speech data and converting the analysis results into a corresponding text description. Then, this text description is input into a pre-defined text encoder, and the encoded text features are transmitted to the speech prior network model 240. The text-to-speech model 230 can include a text encoder and a speech encoder, which can be connected by linear projection in a joint multimodal space. It learns the similarity between the text description and the data pairs composed of repair data used to repair the damaged speech data using cosine similarity, establishing a mapping relationship between text features and speech features used to repair the damaged speech data (which can be called "speech repair features"). Furthermore, during the training of the text-to-speech model 230, corresponding text descriptions can be generated based on different speech repair features, and the loss function of the text-to-speech model 230 can be represented by mean squared error.
[0076] The speech prior network model 240 can be used to receive text features output by the text-speech model 230, and map these text features to corresponding speech restoration features through a pre-established mapping relationship between text features and speech restoration features, thereby obtaining the speech restoration features corresponding to the text description. Furthermore, the speech restoration features corresponding to the text description are transmitted to the diffusion model 250. Alternatively, the speech prior network model 240 can also be used to receive input damaged speech data, analyze the causes of the damage to the damaged speech data, and map the corresponding speech restoration features based on the causes of the damage, thereby obtaining the speech restoration features corresponding to the causes of the damage in the damaged speech data.
[0077] The diffusion model 250 can receive the context-informed speech features output by the speech coding model 210, the voiceprint features extracted by the voiceprint feature extraction model 220, and the speech restoration features mapped by the speech prior network model 240. It then restores the context-informed speech features based on the received voiceprint features and the restored speech restoration features, and transmits the restored Mel-Frequency Spectrum (MFC) to the vocoder model 260. The diffusion model 250 can be trained using a U-Net neural network and / or a Denoising Diffusion Probabilistic Model (DDPM).
[0078] The vocoder model 260 can be used to receive the MFC output by the diffusion model 250 and decode the MFC using its own set adversarial neural network to output the repaired speech data. The vocoder model 260 includes, but is not limited to, channel vocoders, formant vocoders, homomorphic vocoders, linear prediction vocoders, and phoneme vocoders.
[0079] It should be noted that the speech restoration model 200, including the speech input model 210, speech coding model 210, voiceprint feature extraction model 220, text input model 240, text-speech model 230, speech prior network model 240, diffusion model 250, and vocode model 260, is trained using at least hundreds of thousands of speech data points. Furthermore, training the sub-models within the speech restoration model 200 with at least hundreds of thousands of speech data points ensures that the trained speech restoration model 200 possesses strong generalization capabilities.
[0080] Figure 3 This is a flowchart illustrating a training method for a speech restoration model provided in an embodiment of this application. The method 300 can be... Figure 1 The electronic device 140 in the middle performs, or, Figure 2 The large-scale speech restoration model is running for 200 seconds.
[0081] For example, such as Figure 3 As shown, the method 300 includes the following implementation process:
[0082] S310: If damaged speech data exists in the original speech data, obtain the damage description information of the damaged speech data.
[0083] The damage description information is used to indicate the reason for the damage to the speech data. For example, if the original speech data contains damaged speech data caused by a car horn, then "the noise is a car horn" can be used as the damage description information.
[0084] Optionally, the raw voice data is used to represent all audio data collected by the voice acquisition device (e.g., a microphone) during the process of acquiring user voice data.
[0085] Optionally, the damaged description information may be presented in forms including but not limited to text, audio, and images.
[0086] For example, during the recording, transmission, and storage of raw speech data, issues such as noise, reverberation, low resolution, and clipping may occur, resulting in poor clarity or missing data, thus damaging the speech data. To ensure the integrity and clarity of the original speech data, it is necessary to repair the damaged speech data using a speech restoration model. To improve the restoration performance and generalization ability of the speech restoration model, damage description information can be obtained from the damaged speech data and used in the training for repairing the damaged speech data.
[0087] Optionally, the above-mentioned acquisition of damage description information of damaged speech data includes: if multiple damaged speech data are acquired, acquiring damage description information of each damaged speech data in the multiple damaged speech data; or, if multiple damaged speech data are acquired, selecting a preset number of target damaged speech data in the multiple damaged speech data; and acquiring damage description information of the target damaged speech data.
[0088] For example, to improve the restoration performance of a speech restoration model, it can be trained using multiple damaged speech datasets. Therefore, it is necessary to obtain the damage description information for each damaged speech dataset and use this information in the restoration training of the corresponding damaged speech dataset.
[0089] For example, in order to make the speech restoration model have good generalization ability, during the training process of the speech restoration model, it is possible to determine whether to add prior information in the restoration training of a preset number of damaged speech data (which can be called "target damaged speech data").
[0090] The preset quantity can be randomly selected from the damaged speech data sample set by setting a percentage.
[0091] For example, if there are 100 damaged speech data samples in the damaged speech data sample set used for training a speech restoration model (which can be called "sample damaged speech data"), and the percentage is set to 20%, then any 20 sample damaged speech data samples can be selected from the 100 samples. The order of these 20 sample damaged speech data samples in the damaged speech data sample set can be 1-20, 5-15, 30-40, or 75-95. That is, there is no restriction on the position of the sample damaged speech data samples in the damaged speech data sample set, as long as the number of sample damaged speech data samples finally selected meets the set percentage in the damaged speech data sample set.
[0092] For example, when adding prior information to the damaged speech data samples ranked 1 to 20 in the damaged speech data sample set, the damage description information corresponding to each of the 1 to 20 damaged speech data samples can be obtained, and the 1 to 20 damage description information can be used in the repair training of the corresponding damaged speech data samples to improve the repair effect of the trained speech repair model on the damaged speech data.
[0093] Alternatively, when adding prior information to the damaged speech data samples ranked 1 to 20 in the damaged speech data sample set, it can be determined that no prior information should be added to the remaining damaged speech data samples ranked 21 to 100 in the damaged speech data sample set. Furthermore, without adding prior information to the damaged speech data samples ranked 21 to 100, the speech restoration model can be trained. This allows the trained speech restoration model to restore damaged speech data without the need for damage description information, thereby improving the generalization of the trained speech restoration model.
[0094] In this embodiment, multiple damaged speech data can be used as a sample training set, and the damage description information of each damaged speech data in the multiple damaged speech data can be used as prior information to train the speech restoration model. This makes the speech restoration model more accurate in restoring damaged speech data after being trained on a large number of sample damaged speech data, thereby improving the restoration effect of the trained speech restoration model on damaged speech data.
[0095] Alternatively, a preset number of target damaged speech data can be selected from multiple damaged speech data, and the damage description information corresponding to the target damaged speech data can be used as prior information to train the speech restoration model. This can avoid the problem that the trained speech restoration model cannot restore the damaged speech data when it cannot obtain the damage description information corresponding to the damaged speech data, thus improving the generalization of the trained speech restoration model.
[0096] S320: Based on the damage description information, obtain speech restoration features for repairing damaged speech data.
[0097] For example, in the training process of speech restoration models, known noise data, reverberation environments, or device parameters are typically used for training. This can limit the generalization ability of a uniform restoration model, resulting in poor restoration performance when repairing damaged speech data caused by unknown noise data, reverberation environments, or device parameters. Therefore, to improve the restoration effect of the trained speech restoration model, the damage description information of the damaged speech data can be used to guide the training of the speech restoration model, thereby improving the generalization ability of the trained speech restoration model. This allows the trained speech restoration model to achieve better restoration results when repairing damaged speech data caused by unknown noise data, reverberation environments, or device parameters.
[0098] Optionally, the damage description information can be input into the speech prior network model, and the cause of the damage to the damaged speech data can be analyzed through the speech prior network model. Based on the speech features corresponding to the cause of the damage, the speech restoration features corresponding to the cause of the damage in the damaged speech data can be obtained.
[0099] In one possible implementation, the above-mentioned method of obtaining speech restoration features for repairing damaged speech data based on damaged description information includes: obtaining the text corresponding to the damaged description information; encoding the text in a pre-trained text-speech model to obtain encoded text features; inputting the text features into a speech prior network model; and mapping the speech restoration features corresponding to the text features in the speech prior network model through a pre-trained mapping relationship.
[0100] For example, when obtaining damage description information through methods such as speech, the information may be affected by ambient noise, leading to inaccuracies. Therefore, to obtain more accurate damage description information, the damage description information corresponding to the damaged speech data can be described in text. This text is then input into a text-to-speech model, where a text encoder encodes the text to obtain contextualized text features.
[0101] Furthermore, the text features are input into the speech prior network model. The speech prior network model maps the text features to the corresponding speech restoration features through the pre-trained mapping relationship between text features and speech restoration features, thereby obtaining the speech restoration features corresponding to the text.
[0102] In this embodiment of the application, by describing the damage description information corresponding to the damaged speech data through text, the problem of deviation in the damage description information caused by the interference of surrounding environmental noise when describing the damage description information through speech or other means can be avoided. Thus, during the training process of the speech restoration model, the accuracy of obtaining the damage description information can be improved. Furthermore, with more accurate damage description information, the trained speech restoration model can also restore the damaged speech data more accurately.
[0103] S330 uses the speech features corresponding to the damaged speech data as input data, the speech restoration features as prior information, and the sample speech features as target output data to train the speech restoration model, thus obtaining the trained speech restoration model.
[0104] Among them, the sample speech features are used to represent the speech features of the damaged speech data when it is not damaged, that is, the speech features of the damaged speech data when it is not damaged are used as the ground value when training the speech restoration model.
[0105] For example, during the training of the speech restoration model, the speech features corresponding to the damaged speech data can be used as input data, and the speech restoration features can be used as prior information to input into the speech restoration model to obtain the restored speech features. The restored speech features are then compared with the sample speech features. When there is a deviation between the restored speech features and the sample speech features, the speech restoration model is iteratively trained using this deviation until the restored speech features are the same as the sample speech features, or the deviation between the restored speech features and the sample speech features is within the allowable error range. This results in the trained speech restoration model, which can then restore damaged speech data.
[0106] Furthermore, the repaired speech features can be output to the vocode model in the form of Mel, so that the vocode model can decode the Mel to obtain the repaired speech data.
[0107] exist Figure 3 In method 300, since the damage description information of the damaged speech data can represent the cause of the damage, during the training of the speech restoration model, the speech restoration features obtained through the damage description information are used as prior information to guide the training of the speech restoration model. This can improve the ability of the trained speech restoration model to restore damaged speech data based on the damage description information, making the restored speech data more accurate, thereby improving the restoration effect of the trained speech restoration model on damaged speech data.
[0108] Furthermore, since the damage description information of damaged speech data is not limited by the training set of the speech restoration model, the damage description information can be used to represent the cause of damage to any damaged speech data during the training process of the speech restoration model. The speech restoration features obtained through the damage description information for repairing damaged speech data can be used as prior information to guide the training of the speech restoration model, thereby making the trained speech restoration model have stronger generalization ability for repairing damaged speech data.
[0109] In one possible implementation, target voiceprint features are extracted from the original speech data; the above-mentioned speech features corresponding to the damaged speech data are used as input data, speech restoration features are used as prior information, and sample speech features are used as target output data to train the speech restoration model, thereby obtaining the trained speech restoration model.
[0110] For example, after repairing damaged speech data using a trained speech restoration model, it is necessary not only to ensure the integrity and clarity of the restored speech data, but also to ensure that the restored speech data is consistent with the voiceprint information in the original speech data, so that the restored speech data has a very high similarity to the original speech data. Therefore, during the training process of the speech restoration model, voiceprint features from the original speech data (which can be called "target voiceprint features") can be introduced.
[0111] Optionally, the damaged speech data is input into a voiceprint feature extraction model, which extracts voiceprint features from the damaged speech data to obtain target voiceprint features. During the training of the speech restoration model, while using the speech features corresponding to the damaged speech data as input data, the speech restoration features as prior information, and the sample speech features as ground truth to train the speech restoration model, the target voiceprint features can also be introduced as input data into the speech restoration model to jointly train the model, resulting in the trained speech restoration model.
[0112] In this embodiment of the application, during the training process of the speech restoration model, the voiceprint features in the original speech data are introduced, which enables the trained speech restoration model to have the function of fusing voiceprint features. This allows the speech data restored by the trained speech restoration model to be consistent with the voiceprint information in the original speech data, further improving the restoration effect of the trained speech restoration model on damaged speech data.
[0113] In one possible implementation, the task of repairing damaged speech data is obtained; the above-mentioned speech features corresponding to the damaged speech data and target speaker features are used as input data, speech repair features are used as prior information, and sample speech features are used as target output data to train a speech repair model to obtain a trained speech repair model, including: using the speech features corresponding to the damaged speech data, target speaker features and repair task as input data, speech repair features as prior information, and sample speech features as target output data to train a speech repair model to obtain a trained speech repair model.
[0114] Optionally, the restoration task includes, but is not limited to, one or more of noise reduction, dereverberation, super-resolution and clipping restoration.
[0115] For example, when the damage to damaged speech data is due to noise, only denoising is needed to complete the restoration; déreverberation, super-resolution, and clipping restoration are not required. Therefore, to accurately restore damaged speech data, restoration tasks that the speech restoration model needs to perform can be introduced during the training process, providing direction for the model's training.
[0116] Optionally, during the training process of the speech restoration model, while using the speech features corresponding to the damaged speech data and the target speaker features as input data, the speech restoration features as prior information, and the sample speech features as ground truth to train the speech restoration model, one or more restoration tasks can also be input into the speech restoration model to train the speech restoration model together, thus obtaining the trained speech restoration model.
[0117] Specifically, inputting a single repair task into the speech restoration model can make the trained speech restoration model more accurate in repairing damaged speech data; or, inputting multiple repair tasks into the speech restoration model can enable the trained speech restoration model to perform more repair functions on damaged speech data, making the trained speech restoration model more generalizable in repairing damaged speech data.
[0118] In this embodiment, since the causes of damage to different damaged speech data may vary, using the same restoration method to restore speech data with different causes may result in significant deviations in the restored speech data. Therefore, a restoration task corresponding to the damaged speech data can be obtained and introduced into the training process of the speech restoration model. This allows the speech restoration model to be trained based on the restoration task corresponding to the damaged speech data, making the restoration function of the trained speech restoration model more closely aligned with the causes of damage to the damaged speech data and more targeted. This leads to accurate restoration of the damaged speech data and improves the accuracy of damaged speech data restoration.
[0119] Figure 4 This is a flowchart illustrating another training method for a speech restoration model provided in this application embodiment.
[0120] For example, such as Figure 4 As shown, the method 400 includes the following implementation process:
[0121] S410, acquire raw voice data.
[0122] For example, all voice data collected by the microphone can be acquired, and this voice data can be used as raw voice data.
[0123] S420: Determine if there is any damaged speech data in the original speech data. If yes, proceed to S430 or S470. If no, do not proceed. Figure 4 Any step in it.
[0124] For example, it is determined whether there are problems such as noise, reverberation, low resolution and clipping in the original voice data obtained in S410 during the recording, transmission and storage process, which may result in poor clarity of the original voice data or partial data loss, thus causing damaged voice data in the original voice data.
[0125] For example, if it is determined that there is no damaged voice data in the original voice data, then the original voice data is complete and clear, and there is no need for repair. The original voice data can be directly output to the user.
[0126] S430: Obtain the text corresponding to the damage description information of the damaged voice data.
[0127] For example, if it is determined through S420 that there is damaged speech data in the original speech data, in order to make the trained speech restoration model have a better restoration effect on the damaged speech data, the text corresponding to the damage description information of the damaged speech data can be obtained, the damaged description information corresponding to the damaged speech data can be described by the text, and the damaged description information can be used to guide the training of the speech restoration model, thereby improving the generalization of the trained speech restoration model.
[0128] S440 encodes the text in a pre-trained text-to-speech model to obtain encoded text features.
[0129] For example, the text corresponding to the damaged descriptive information can be input into a pre-trained text-to-speech model (i.e., Figure 2 In the text-to-speech model 230, the text is encoded by the text encoder in the pre-trained text-to-speech model to obtain text features with context.
[0130] S450, input the text features into the speech prior network model; in the speech prior network model, the speech restoration features corresponding to the text features are mapped through the pre-trained mapping relationship.
[0131] For example, the text features obtained in S440 are input into the speech prior network model. The speech prior network model maps the text features to the corresponding speech restoration features through the pre-trained mapping relationship between the text features and the speech restoration features, thereby obtaining the speech restoration features corresponding to the text.
[0132] S460 extracts voiceprint features from the raw speech data.
[0133] For example, in order to ensure that the voice data restored by the trained voice restoration model is consistent with the voiceprint information in the original voice data, voiceprint features can be extracted from the original voice data using a voiceprint feature extraction model, and these voiceprint features can be used in the training process of the voice restoration model.
[0134] S470, task of acquiring and repairing damaged voice data.
[0135] For example, if it is determined through S420 that there is damaged speech data in the original speech data, in order to accurately repair the damaged speech data, a repair task for the damaged speech data can be obtained and input into the speech repair model for training.
[0136] S480 uses the speech features, speaker features, and restoration task corresponding to the damaged speech data as input data, the speech restoration features as prior information, and the speech features when the damaged speech data is undamaged as ground truth to train the speech restoration model, thus obtaining the trained speech restoration model.
[0137] For example, during the training of a speech restoration model, speech features corresponding to damaged speech data can be obtained through a speech coding model.
[0138] Furthermore, the speech features corresponding to the damaged speech data, the target speaker features, and the restoration task can all be used as input data, the speech restoration features as prior information, and the sample speech features as ground truth to train the speech restoration model, thus obtaining the trained speech restoration model.
[0139] It should be noted that, Figure 4 All steps are in Figure 3 The corresponding embodiments are described in detail, and will not be repeated here.
[0140] Figure 5 This is a flowchart illustrating a voice restoration method provided in an embodiment of this application. The method 300 can be... Figure 1 The electronic device 140 in the middle performs, or, Figure 2 The large-scale speech restoration model is running for 200 seconds.
[0141] For example, such as Figure 5 As shown, the method 500 includes the following implementation process:
[0142] S510, if damaged speech data exists in the original speech data, determine whether to obtain the damaged speech data description information.
[0143] Among them, the damage description information is used to indicate the reason for the damage to the damaged voice data.
[0144] For example, during the training of the speech restoration model, a predetermined number of damaged speech samples are randomly selected, and prior information is incorporated into the restoration training of these damaged speech samples. That is, prior information is incorporated into the restoration training of a portion of the damaged speech samples, while prior information is not incorporated into the restoration training of another portion of the damaged speech samples. This ensures that the pre-trained speech restoration model (such as...) Figure 2 The diffusion model 250 or the large speech restoration model 200 shown can restore damaged speech data with or without prior information. Therefore, when damaged speech data exists in the obtained original speech data, it is necessary to determine whether the damage description information of the damaged speech data has been obtained.
[0145] S520: If the damage description information is obtained, the speech restoration features used to repair the damaged speech data are obtained based on the damage description information.
[0146] For example, when obtaining the damage description information corresponding to the damaged speech data, the damage description information can be input into the speech prior network model. The speech prior network model can analyze the damage cause of the damaged speech data. Based on the speech restoration features (which can be called "speech restoration features") corresponding to the damage cause mapping, the speech restoration features corresponding to the damage cause in the damaged speech data can be obtained.
[0147] Optionally, in order to obtain more accurate damage description information, the damage description information corresponding to the damaged speech data can be described by text, and the text can be input into a text speech model. The text encoder in the text speech model can encode the text to obtain text features with context.
[0148] Furthermore, the text features are input into the speech prior network model. The speech prior network model maps the text features to the corresponding speech restoration features through the pre-trained mapping relationship between text features and speech restoration features, thereby obtaining the speech restoration features corresponding to the text.
[0149] S530 uses the speech restoration features as prior information and inputs the damaged speech data into a pre-trained speech restoration model for restoration, resulting in restored speech data.
[0150] The pre-trained speech restoration model is trained by using the speech features corresponding to the damaged speech data as input data, the speech features obtained from the damage description information corresponding to the damaged speech data as training prior information, and the speech features of the damaged speech data when it is undamaged as target output data.
[0151] It should be noted that the training process of the pre-trained speech restoration model is in Figures 2 to 4 The corresponding embodiments have been described in detail and will not be repeated here.
[0152] For example, during the process of repairing damaged speech data using a pre-trained speech restoration model, the repair speech features can be used as prior information and input into the pre-trained speech restoration model along with the damaged speech data to obtain the repaired speech data. Furthermore, because the speaker fingerprint features corresponding to the original speech data are introduced during the training process of the speech restoration model, the pre-trained speech restoration model possesses the function of fusing speaker fingerprint features. This ensures that the repaired speech data remains consistent with the speaker fingerprint information in the original speech data, further improving the repair effect of damaged speech data.
[0153] Optionally, the damaged speech data can be first input into a speech coding model to obtain the speech features corresponding to the damaged speech data, and then the speech features corresponding to the damaged speech data and the repaired speech features, which are used as prior information, can be input into a pre-trained speech repair model for repair to obtain the repaired speech features.
[0154] Furthermore, the repaired speech features are output to the vocode model in the form of Mel, so that the vocode model can decode the Mel to obtain the repaired speech data.
[0155] In such Figure 5 In the method 500 shown, if the original speech data contains damaged speech data and the descriptive information corresponding to the damaged speech data can be obtained, the damaged descriptive information can be used as prior information to guide the pre-trained speech restoration model to restore the damaged speech data when the pre-trained speech restoration model restores the damaged speech data. This can make the pre-trained speech restoration model restore the damaged speech data more accurately, thereby making the restored speech data more accurate.
[0156] In one possible implementation, the speech restoration features are used as prior information and the damaged speech data are input into a pre-trained speech restoration model for restoration to obtain the restored speech data. This includes: if a restoration task for damaged speech data is obtained, the prior information, the damaged speech data, and the restoration task are input into a pre-trained speech restoration model for restoration to obtain the restored speech data.
[0157] For example, the causes of damage to different damaged speech data may vary, leading to different restoration requirements. For instance, damaged speech data A requires noise reduction, damaged speech data B requires dredging, damaged speech data C requires super-resolution restoration, and damaged speech data D requires amplitude clipping restoration. Using the same restoration method for speech data with different causes may result in significant deviations in the restored speech data. Therefore, to improve the accuracy of damaged speech data restoration, a restoration task can be obtained for the damaged speech data. This task can be input into a pre-trained speech restoration model during the restoration process, making the model's restoration more closely aligned with the cause of damage and more targeted, thus achieving precise restoration.
[0158] Optionally, during the repair process of damaged speech data by the pre-trained speech restoration model, the prior information corresponding to the damaged speech data, the speech features of the damaged speech data, and the repair task can all be input into the pre-trained speech restoration model for repair to obtain the repaired speech features.
[0159] Furthermore, the repaired speech features are output to the vocode model in the form of Mel, so that the vocode model can decode the Mel to obtain the repaired speech data.
[0160] It should be noted that the repair task can be actively input by the user based on the cause of the damage to the damaged voice data, or it can be obtained by the voice repair model after analyzing the cause of the damage to the damaged voice data.
[0161] In this embodiment, since the causes of damage to different damaged speech data may vary, using the same restoration method to restore speech data with different causes may result in significant deviations in the restored speech data. Therefore, when the pre-trained speech restoration model restores damaged speech data, a restoration task corresponding to the damaged speech data can be introduced. This allows the pre-trained speech restoration model to restore the damaged speech data based on the corresponding restoration task, making the restoration more accurate and resulting in more accurate restored speech data.
[0162] Optionally, if the damaged description information is not obtained, the damaged speech data is input into a pre-trained speech restoration model for restoration to obtain the restored speech data.
[0163] For example, since a pre-trained speech restoration model can repair damaged speech data even without incorporating prior information, when the corresponding damage description information for the damaged speech data is not available, the damaged speech data can be input into the pre-trained speech restoration model for repair, resulting in repaired speech data. Thus, even without obtaining damage description information, damaged speech data can still be repaired; that is, the pre-trained speech restoration model can repair any damaged speech data, regardless of whether the corresponding damage description information for the damaged speech data is available.
[0164] Optionally, the damaged speech data can be first input into a speech coding model to obtain the speech features corresponding to the damaged speech data, and then the speech features corresponding to the damaged speech data can be input into a pre-trained speech restoration model for restoration to obtain the restored speech features.
[0165] Furthermore, the repaired speech features are output to the vocode model in the form of Mel, so that the vocode model can decode the Mel to obtain the repaired speech data.
[0166] In this embodiment, because prior information can be omitted during the training of the pre-trained speech restoration model for some samples of damaged speech data, the pre-trained speech restoration model can still restore the damaged speech data even when it cannot obtain the corresponding damage description information, thus improving the restoration effect of the damaged speech data.
[0167] It should be understood that the above examples are provided to help those skilled in the art understand the embodiments of this application, and are not intended to limit the embodiments of this application to the specific values or scenarios exemplified. Those skilled in the art can obviously make various equivalent modifications or variations based on the above examples, and such modifications or variations also fall within the scope of the embodiments of this application.
[0168] The above text combined Figures 1 to 4 The training method of the speech restoration model provided in the embodiments of this application is described in detail, and combined with Figure 5 The speech restoration method provided in the embodiments of this application is described in detail below; the following will be combined with Figure 6 and Figure 8 The apparatus embodiments of this application are described in detail below. It should be understood that the apparatus in the embodiments of this application can perform the various methods described in the foregoing embodiments of this application, that is, the specific working processes of the various products described below can be referred to the corresponding processes in the foregoing method embodiments.
[0169] Figure 6 This is a schematic diagram of the structure of the training device for the speech restoration model provided in this application embodiment.
[0170] For example, such as Figure 6 As shown, the device 600 is configured in an electronic device and includes:
[0171] Model 610 is used to obtain damage description information of damaged speech data when damaged speech data exists in the original speech data; wherein, the damage description information is used to indicate the cause of damage to the damaged speech data;
[0172] Processing model 620 is used to obtain speech restoration features for repairing damaged speech data based on the damaged description information;
[0173] The training model 630 is used to train the speech restoration model by taking the speech features corresponding to the damaged speech data as input data, the speech restoration features as prior information, and the sample speech features as target output data, and obtain the trained speech restoration model; wherein, the sample speech features are used to represent the speech features when the damaged speech data is not damaged.
[0174] In one possible implementation, the processing model 620 is specifically used to: obtain the text corresponding to the damaged description information; encode the text in a pre-trained text-speech model to obtain encoded text features; input the text features into a speech prior network model; and map the speech restoration features corresponding to the text features in the speech prior network model through a pre-trained mapping relationship.
[0175] In one possible implementation, the processing model 620 is further used to: extract target voiceprint features from the original speech data; the training model 630 is specifically used to: train the speech restoration model by taking the speech features corresponding to the damaged speech data and the target voiceprint features as input data, the speech restoration features as prior information, and the sample speech features as target output data, to obtain the trained speech restoration model.
[0176] In one possible implementation, the acquisition model 610 is further used to: acquire the restoration task of damaged speech data; the training model 630 is specifically used to: train the speech restoration model by taking the speech features corresponding to the damaged speech data, the target speaker features and the restoration task as input data, the speech restoration features as prior information, and the sample speech features as target output data, so as to obtain the trained speech restoration model.
[0177] In one possible implementation, the acquisition model 610 is specifically used for: if multiple damaged speech data are acquired, acquiring the damage description information of each damaged speech data in the multiple damaged speech data; or, if multiple damaged speech data are acquired, selecting a preset number of target damaged speech data in the multiple damaged speech data; and acquiring the damage description information of the target damaged speech data.
[0178] Figure 7 This is a schematic diagram of the speech restoration device provided in the embodiments of this application.
[0179] For example, such as Figure 7 As shown, the device 700 is configured in an electronic device and includes:
[0180] Model 710 is used to determine whether damage description information of the damaged speech data is obtained when damaged speech data exists in the original speech data; wherein, the damage description information is used to indicate the reason for the damage of the damaged speech data.
[0181] The processing model 720 is used to obtain speech restoration features for repairing damaged speech data based on the damaged description information if the damaged description information is obtained.
[0182] Repair model 730 is used to input the speech repair features as prior information and the damaged speech data into a pre-trained speech repair model for repair, so as to obtain the repaired speech data.
[0183] The pre-trained speech restoration model is trained by using the speech features corresponding to the damaged speech data as input data, the speech features obtained from the damage description information corresponding to the damaged speech data as training prior information, and the speech features of the damaged speech data when it is undamaged as target output data.
[0184] In one possible implementation, the repair model 730 is specifically used to: if a repair task for damaged speech data is obtained, input the prior information, the damaged speech data and the repair task into a pre-trained speech repair model for repair, and obtain the repaired speech data.
[0185] In one possible implementation, the repair model 730 is also used to: if the damaged description information is not obtained, input the damaged speech data into the pre-trained speech repair model for repair, and obtain the repaired speech data.
[0186] It should be noted that the aforementioned device 600 or device 700 is embodied in the form of a functional model. The term "model" here can be implemented in software and / or hardware, without specific limitations.
[0187] For example, a “model” can be a software program, hardware circuitry, or a combination of both that implements the above-described functions. Hardware circuitry may include application-specific integrated circuits (ASICs), electronic circuitry, a processor (e.g., a shared processor, a proprietary processor, or a group processor) and memory for executing one or more software or firmware programs, integrated logic circuitry, and / or other suitable components that support the described functions.
[0188] Therefore, the models of the various examples described in the embodiments of this application can be implemented using electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0189] Figure 8 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application.
[0190] For example, such as Figure 8 As shown, the electronic device 800 includes a memory 810 and a processor 820, wherein the memory 810 stores executable program code 8101, and the processor 820 is used to call and execute the executable program code 8101 to perform a voice repair method.
[0191] This application can divide the electronic device into functional models based on the above method examples. For example, it can correspond to various functional models, or it can integrate two or more functions into one processing model. The integrated model can be implemented in hardware. It should be noted that the model division in this embodiment is illustrative and is only a logical functional division. In actual implementation, there may be other division methods.
[0192] When each functional model is divided according to its corresponding function, the electronic device may include: model acquisition, model processing, model training, and model repair. It should be noted that all relevant content of each step involved in the above method embodiments can be referenced from the functional description of the corresponding functional model, and will not be repeated here.
[0193] The electronic device provided in this application is used to perform the training of the above-mentioned speech restoration model or the speech restoration method, and thus can achieve the same effect as the above-mentioned implementation method.
[0194] When using integrated units, an electronic device may include a processing model and a storage model. The processing model is used to control and manage the actions of the electronic device. The storage model is used to support the execution of program code and data by the electronic device.
[0195] The processing model can be a processor or controller that can implement or execute various exemplary logic blocks, models, and circuits shown in conjunction with the disclosure of this application. The processor can also be a combination of functions that implement computation, such as a combination of one or more microprocessors, a combination of digital signal processing (DSP) and microprocessors, etc., and the storage model can be a memory.
[0196] This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the methods described in the foregoing embodiments. The computer-readable storage medium may include, but is not limited to, any type of disk, including floppy disks, optical disks, DVDs (Digital Video Discs), CD-ROMs (Compact Disc Read-Only Memory), microdrives, magneto-optical disks, ROMs (Read-Only Memory), RAMs (Random Access Memory), EPROMs (Erasable Programmable Read-Only Memory), EEPROMs (Electrically Erasable Programmable Read Only Memory), DRAMs (Dynamic Random Access Memory), VRAMs (Video Random Access Memory), flash memory devices, magnetic cards or optical cards, nanosystems (including molecular memory ICs), or any type of medium or device suitable for storing instructions and / or data.
[0197] This application also provides a computer program product that, when run on a computer, causes the computer to perform the aforementioned related steps to implement a voice restoration model training or voice restoration method as described in the above embodiments.
[0198] In addition, the electronic device provided in the embodiments of this application may specifically be a chip, component or model. The electronic device may include a connected processor and a memory. The memory is used to store instructions. When the electronic device is running, the processor may call and execute the instructions to make the chip perform a training of a speech restoration model or a speech restoration method in the above embodiments.
[0199] The electronic devices, computer-readable storage media, computer program products or chips provided in this application are all used to perform the corresponding methods provided above. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects in the corresponding methods provided above, and will not be repeated here.
[0200] Through the above description of the embodiments, those skilled in the art can understand that, for the sake of convenience and brevity, only the division of the above functional models is used as an example. In actual applications, the above functions can be assigned to different functional models as needed, that is, the internal structure of the device can be divided into different functional models to complete all or part of the functions described above.
[0201] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of models or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0202] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A training method for a speech restoration model, characterized in that, Applied to electronic devices, the method includes: In the case where damaged speech data exists in the original speech data, damage description information of the damaged speech data is obtained; wherein, the damage description information is used to indicate the reason for the damage to the damaged speech data; Based on the damage description information, speech restoration features for repairing the damaged speech data are obtained; The speech restoration model is trained by using the speech features corresponding to the damaged speech data as input data, the speech restoration features as prior information, and the sample speech features as target output data, to obtain the trained speech restoration model. The sample speech features are used to represent the speech features of the damaged speech data when it is undamaged.
2. The method according to claim 1, characterized in that, The step of obtaining speech restoration features for repairing the damaged speech data based on the damaged description information includes: Obtain the text corresponding to the damaged description information; In a pre-trained text-to-speech model, the text is encoded to obtain encoded text features; The text features are input into the speech prior network model; In the speech prior network model, the speech restoration features corresponding to the text features are mapped through a pre-trained mapping relationship.
3. The method according to claim 1, characterized in that, The method further includes: Extract the target voiceprint features from the original speech data; The step of training a speech restoration model by using the speech features corresponding to the damaged speech data as input data, the speech restoration features as prior information, and the sample speech features as target output data to obtain the trained speech restoration model includes: The speech restoration model is trained by using the speech features corresponding to the damaged speech data and the target voiceprint features as the input data, the speech restoration features as the prior information, and the sample speech features as the target output data, to obtain the trained speech restoration model.
4. The method according to claim 3, characterized in that, The method further includes: The task of repairing the damaged voice data is to acquire the data. The step of training the speech restoration model by using the speech features corresponding to the damaged speech data and the target speaker features as the input data, the speech restoration features as the prior information, and the sample speech features as the target output data to obtain the trained speech restoration model includes: The speech restoration model is trained by using the speech features corresponding to the damaged speech data, the target voiceprint features, and the restoration task as the input data, the speech restoration features as the prior information, and the sample speech features as the target output data, to obtain the trained speech restoration model.
5. The method according to any one of claims 1 to 4, characterized in that, The process of obtaining the damage description information of the damaged voice data includes: If multiple damaged speech data are obtained, obtain the damage description information for each damaged speech data in the multiple damaged speech data; or... If multiple damaged voice data are obtained, a preset number of target damaged voice data are selected from the multiple damaged voice data; damage description information of the target damaged voice data is obtained.
6. A method for voice restoration, characterized in that, Applied to electronic devices, the method includes: If damaged speech data exists in the original speech data, determine whether to obtain damage description information of the damaged speech data; wherein, the damage description information is used to indicate the reason for the damage to the damaged speech data; If the damage description information is obtained, speech restoration features for repairing the damaged speech data are obtained based on the damage description information. The speech restoration features are used as prior information and input into a pre-trained speech restoration model along with the damaged speech data to restore the speech data, thereby obtaining the restored speech data. The pre-trained speech restoration model is a restoration model trained by using the speech features corresponding to the damaged speech data of the sample as input data, the speech features obtained from the damage description information corresponding to the damaged speech data of the sample as prior information for training, and the speech features of the damaged speech data of the sample when it is undamaged as target output data.
7. The method according to claim 6, characterized in that, The process of using the speech restoration features as prior information and inputting the damaged speech data into a pre-trained speech restoration model for restoration to obtain restored speech data includes: If a repair task for the damaged speech data is obtained, the prior information, the damaged speech data, and the repair task are input into the pre-trained speech repair model for repair, thereby obtaining the repaired speech data.
8. The method according to claim 6, characterized in that, The method further includes: If the damaged description information is not obtained, the damaged speech data is input into the pre-trained speech restoration model for restoration to obtain the restored speech data.
9. A training device for a speech restoration model, characterized in that, Configured in an electronic device, the device includes: A model is used to obtain damage description information of the damaged speech data when damaged speech data exists in the original speech data; wherein, the damage description information is used to indicate the cause of the damage to the damaged speech data; A processing model is used to obtain speech restoration features for repairing the damaged speech data based on the damaged description information; The training model is used to train the speech restoration model by taking the speech features corresponding to the damaged speech data as input data, the speech restoration features as prior information, and the sample speech features as target output data, to obtain the trained speech restoration model; wherein, the sample speech features are used to represent the speech features of the damaged speech data when it is not damaged.
10. A device for voice restoration, characterized in that, Configured in an electronic device, the device includes: An acquisition model is used to determine whether to acquire damage description information of the damaged speech data when damaged speech data exists in the original speech data; wherein, the damage description information is used to indicate the cause of damage to the damaged speech data; A processing model is used to obtain, based on the damage description information, speech restoration features for repairing the damaged speech data if the damage description information is obtained; The repair model is used to input the speech repair features as prior information and the damaged speech data into a pre-trained speech repair model to repair the damaged speech data, thereby obtaining the repaired speech data. The pre-trained speech restoration model is trained by using the speech features corresponding to the damaged speech data as input data, the speech features obtained from the damage description information corresponding to the damaged speech data as training prior information, and the speech features of the damaged speech data when it is undamaged as target output data.
11. An electronic device, characterized in that, The electronic device includes: Memory, used to store executable program code; A processor for calling and running the executable program code from the memory, causing the electronic device to perform the method as claimed in claims 1 to 5, or any one of claims 6 to 8.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed, implements the method as claimed in claims 1 to 5, or any one of claims 6 to 8.