Text-based Speech Editing Method, System, Electronic Device, and Storage Medium

Voice context modeling is performed through BERT, combining text and speech encoder to generate predicted Mel spectrum, which solves the problems of incoherent and poor generalization of speech editing in the prior art, and achieves a more natural and coherent speech editing effect.

CN115966196BActive Publication Date: 2025-06-24AISPEECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211696422.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-28
Publication Date
2025-06-24
Estimated Expiration
2042-12-28

AI Technical Summary

Technical Problem

The existing text-based speech editing methods have problems of incoherence and poor generalization. The TTS model is difficult to synthesize speakers and rhythms that conform to the original speech. The end-to-end method highly relies on speaker characteristics and has low prediction efficiency.

Method used

BERT is used for speech context modeling, and the text representation and speech characteristics of the edited text are determined through text encoder and speech encoder, and combined with the joint network to generate predicted Mel spectrum to realize speech editing.

Benefits of technology

By capturing the rich voice context information of the recorded audio, voice in the editing area that is more in line with the original audio is generated, which avoids the unnatural and discontinuous voice generated by the splicing method, and improves the quality and generalization of voice editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115966196B_ABST
    Figure CN115966196B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides a text-based speech editing method, system, electronic device, and storage medium. The method includes: inputting the editing text into a text encoder, determining the speech duration corresponding to the modified part in the editing text, and determining the text representation of the editing text based on the speech duration and the phoneme encoding of the editing text; inputting the speech duration and the speech before the editing text is modified into a speech encoder, masking the corresponding modified part in the speech before the modification based on the speech duration, to obtain an acoustic representation, a hidden representation with masked context, and a Mel spectrogram with masked regions; inputting the text representation, the masked acoustic representation, and the hidden representation into a joint network to obtain the predicted Mel spectrogram corresponding to the masked region. The embodiments of the present invention enable the model to utilize the context information of the original speech, thereby predicting the speech in the editing area that is more consistent with the original audio, and can also avoid the unnatural and discontinuous phenomena of the speech generated by the splicing method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent voice, and particularly to a text-based voice editing method, system, electronic device, and storage medium. Background Art

[0002] In a text-based voice editing method and system, when there is recorded audio and corresponding text, the user only needs to edit the text, and the system can output the corresponding edited voice according to the edited text. The text-based voice editing method is closely related to the TTS (Text-to-Speech) synthesis model. Currently, there are mainly two types of text-based voice editing methods and systems, one is the splicing-based method, and the other is the end-to-end method. The neural network-based TTS model generally takes words or phonemes as input, generates the Mel spectrogram, and then the vocoder generates the voice, or the TTS model directly generates the voice.

[0003] In a text-based voice editing method and system, for a system based on the splicing method, the voice segment of the editing area in the text is often synthesized by the TTS model or selected from existing voice data, and then the obtained voice segment is inserted into the corresponding area of the original voice. Usually, in order to make the obtained voice segment closer to the speaker of the original voice, the VC (Voice Conversion) model is used. In addition, in order to make the spliced voice more coherent and smooth, pitch-shifting and time-stretching technologies are also used.

[0004] For the end-to-end text-based voice editing method, a neural network model is usually used to predict the voice of the editing area in the text. This type of method inputs the edited text and directly outputs the edited voice or Mel spectrogram.

[0005] In the process of implementing the present invention, the inventors found that there are at least the following problems in the related technologies:

[0006] The TTS model cannot be directly used for text-based voice editing tasks. It is difficult for the TTS model to synthesize the voice of the editing area that conforms to the speaker and prosody of the original voice, which requires a large amount of voice data similar to the original voice and its corresponding text data to fine-tune the TTS model;

[0007] In a text-based speech editing method and system, for a system based on a splicing method, there are obvious intervals in the generated speech between the editing area and the non-editing area, which is an incoherence phenomenon caused by directly splicing the speech; moreover, each part such as the speech segment generation part, the conversion part, and the splicing part is separated, so the utilization of the audio features of the original speech is less, and thus there is a large difference between the speech corresponding to the text in the editing area and the original speech.

[0008] End-to-end text-based speech editing methods use the extracted speaker features to make the predicted speech in the editing area conform to the original speech, which makes these methods highly dependent on the speaker features and thus have poor speaker generalization; some methods use an encoder-decoder architecture and use an autoregressive manner during decoding, resulting in low prediction efficiency. Summary of the Invention

[0009] In order to at least solve the problems of incoherence and poor generalization of the speech generated by the existing text-based speech editing.

[0010] In a first aspect, an embodiment of the present invention provides a text-based speech editing method, including:

[0011] Input the editing text into a text encoder, determine a first speech duration corresponding to the modified part in the editing text and a second speech duration corresponding to the entire editing text, and determine a text representation of the editing text based on the second speech duration and the phoneme encoding of the editing text;

[0012] Input the first speech duration and the speech before the editing text is modified into a speech encoder, cover the corresponding modified part in the speech before modification based on the first speech duration, and obtain a covered acoustic representation, a hidden representation with covered context, and a Mel spectrogram with a covered area, where the length of the text representation is the same as the length of the Mel spectrogram with the covered area;

[0013] Input the text representation, the covered acoustic representation, and the hidden representation with covered context into a joint network to obtain a predicted Mel spectrogram corresponding to the covered area, and obtain the speech after the editing text is modified based on the Mel spectrogram with the covered area and the predicted Mel spectrogram.

[0014] In a second aspect, an embodiment of the present invention provides a text-based speech editing system, including:

[0015] A text encoding program module for inputting an edited text into a text encoder, determining a first speech duration corresponding to a modified part in the edited text and a second speech duration corresponding to the whole edited text, and determining a text representation of the edited text based on the second speech duration and the phoneme encoding of the edited text;

[0016] A speech encoding program module for inputting the first speech duration and the speech before the edited text into a speech encoder, masking the part corresponding to the modified part in the speech before the modification based on the first speech duration to obtain a masked acoustic representation, a hidden representation with masked context, and a Mel spectrogram with masked regions, wherein the length of the text representation is the same as the length of the Mel spectrogram with masked regions;

[0017] A speech editing program module for inputting the text representation, the masked acoustic representation, and the hidden representation with masked context into a joint network to obtain a predicted Mel spectrogram corresponding to the masked region, and obtaining the speech after the modification of the edited text based on the Mel spectrogram with masked regions and the predicted Mel spectrogram.

[0018] In a third aspect, an electronic device is provided, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the steps of the text-based speech editing method according to any embodiment of the present invention.

[0019] In a fourth aspect, an embodiment of the present invention provides a storage medium, on which a computer program is stored, characterized in that when the program is executed by a processor, the steps of the text-based speech editing method according to any embodiment of the present invention are implemented.

[0020] The beneficial effects of the embodiments of the present invention are as follows: Based on the speech context modeling of BERT, the speech in the editing area can capture rich speech context information of the recorded audio during prediction, including features such as the speaker, environment, and pitch. This enables the model to make good use of the context information of the original speech, thereby predicting the speech in the editing area that is more in line with the original audio, and also avoiding the unnatural and discontinuous phenomena of the speech generated by the splicing method. Description of the Drawings

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0022] Figure 1 is a flowchart of a text-based speech editing method provided by an embodiment of the present invention;

[0023] Figure 2 is a schematic diagram of the overall structure of a text-based speech editing method provided by an embodiment of the present invention;

[0024] Figure 3 is the MCD evaluation result of a text-based speech editing method provided by an embodiment of the present invention for different baseline models;

[0025] Figure 4 is the MOS score of a text-based speech editing method provided by an embodiment of the present invention for speaker visible / speaker invisible;

[0026] Figure 5 is a schematic diagram of the editing process of a text-based speech editing method provided by an embodiment of the present invention;

[0027] Figure 6 is a schematic diagram of the structure of a text-based speech editing system provided by an embodiment of the present invention;

[0028] Figure 7 is a schematic diagram of the structure of an embodiment of an electronic device for text-based speech editing provided by an embodiment of the present invention. Detailed implementation manners

[0029] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0030] As Figure 1 shown is a flowchart of a text-based speech editing method provided by an embodiment of the present invention, including the following steps:

[0031] S11: Input the edited text into a text encoder, determine the first speech duration corresponding to the modified part in the edited text and the second speech duration corresponding to the whole edited text, and determine the text representation of the edited text based on the second speech duration and the phoneme encoding of the edited text;

[0032] S12: Input the first speech duration and the speech before the edited text is modified into a speech encoder. Based on the first speech duration, cover the part of the speech before modification corresponding to the modified part to obtain a covered acoustic representation, a hidden representation with covered context, and a Mel spectrogram with covered regions, where the length of the text representation is the same as the length of the Mel spectrogram with covered regions;

[0033] S13: Input the text representation, the covered acoustic representation, and the hidden representation with covered context into a joint network to obtain a predicted Mel spectrogram corresponding to the covered region. Based on the Mel spectrogram with covered regions and the predicted Mel spectrogram, obtain the speech after the edited text is modified.

[0034] In this embodiment, the model proposed by this method can be regarded as an effective combination of a non-autoregressive TTS model FastSpeech2 and a BERT (Bidirectional Encoder Representations from Transformers) with excellent context modeling ability. The model of this method can be called BEdit-TTS (TEXT-BASED SPEECH EDITING SYSTEM WITH BIDIRECTIONAL TRANSFORMERS). As Figure 2 shown, this method includes the following parts: a text encoder, a speech encoder, and a joint network. The input of a traditional TTS model is text, and the target output is a Mel spectrogram. In the model of this method, the text and the real speech after being covered are used as inputs, and the target output is the Mel spectrogram of the covered region. Prepare the text and the speech corresponding to this text in advance. When the user uses it, edit the target words in the text to obtain the edited text.

[0035] For step S11, the purpose of the text encoder is to extract the text representation in the edited text modified by the user from the input edited text.

[0036] Specifically, the text encoder includes: a phoneme embedding block, an encoder, a duration predictor, and a length regulator, where

[0037] The phoneme embedding block is used to determine the phoneme embedding of the edited text;

[0038] The encoder is used to determine the text representation of the edited text according to the phoneme embedding and the corresponding positional encoding;

[0039] The duration predictor is used to determine a first speech duration corresponding to a modified part in the edited text and a second speech duration corresponding to the entire edited text;

[0040] The length adjuster is used to adjust the length of the text representation according to the second speech duration, so that the length of the adjusted text representation is consistent with the length of the Mel spectrogram with a covered area.

[0041] In this embodiment, as Figure 2 shown in the text encoder in the lower left part, the design of this part is similar to the FastSpeech2 structure, including a phoneme embedding block, an encoder, a duration predictor, and a length adjuster. Through this part, the phoneme embedding block determines the corresponding phoneme embedding from the input edited text. Then, the encoder encodes the phoneme embedding and the corresponding position encoding to obtain the text representation of the edited text modified by the user. However, since the content in the edited text has been modified, in order to match the length of the subsequent steps and acoustic features, it is necessary to adjust its length in advance. The duration predictor is used to determine the speech duration of the modified part by the user and the speech duration corresponding to the entire edited text.

[0042] Through this part, a text representation consistent with the length of the corresponding speech Mel spectrogram can be extracted, but this representation does not contain some key acoustic features, such as the speaker, prosody, etc. In order to extract these acoustic features, techniques such as speaker encoding are usually used to generate a speaker representation, but in speech editing, it is very difficult to accurately extract these features. Therefore, this method uses a separate speech encoder to extract acoustic information.

[0043] For step S12, the existing speech corresponding to the text before the user's modification is input into the speech encoder, as Figure 2 shown in the speech encoder in the lower right part. This module takes real speech as input, aiming to learn rich acoustic information, including the speaker, prosody, channel effect, etc.

[0044] As an embodiment, the speech encoder includes: a masking operation block and a conversion encoder, where

[0045] the masking operation block is used to receive the speech before the modification of the edited text and the corresponding Mel spectrogram, mask the corresponding modified part in the speech before the modification of the edited text according to the first speech duration corresponding to the modified part in the edited text, and obtain a masked acoustic representation and a Mel spectrogram with a masked area;

[0046] the conversion encoder is used to transcode the Mel spectrogram with the masked area to obtain a hidden representation with masked context.

[0047] In this embodiment, the duration of the masked region determined by the text encoder is used to mask the corresponding modified parts in the speech and Mel spectrogram, obtaining the speech and Mel spectrogram with the modified parts masked. Specifically, a fixed value such as 0 or 1 can be used as the masking value for the masking process. Then, the conversion encoder calculates the continuous and dense hidden representations in the masked Mel spectrogram.

[0048] For step S13, as Figure 2 shown in the upper part of the joint network, the purpose of this model is to effectively combine the text information extracted by the text encoder and the acoustic information extracted by the speech encoder to generate pitch and energy features and Mel spectrogram features similar to real speech.

[0049] Specifically, the joint network includes: a pitch-energy converter and a Mel spectrogram decoder, where

[0050] The pitch-energy converter is used to generate pitch and energy features simulating real speech and predicted Mel spectrogram features according to the received text representation, masked acoustic representation, and the hidden representation with masked context;

[0051] The Mel spectrogram decoder is used to determine the predicted Mel spectrogram corresponding to the masked region according to the pitch and energy features simulating real speech and the predicted Mel spectrogram features.

[0052] In this embodiment, in order to predict the speech of the masked part (i.e., the speech of the target word modified by the user), this module fuses text information and acoustic information, and uses a converter model to predict the pitch and energy of the masked region. Then, the text information, acoustic information, and the predicted pitch and energy are fused and input into the Mel spectrogram decoder to predict the Mel spectrogram features of the masked region. The Mel spectrogram decoder is implemented using the feed-forward transformer in FastSpeech2. The obtained Mel spectrogram features of the masked region are concatenated with the masked Mel spectrogram features to obtain the speech after text editing.

[0053] As an embodiment, the text encoder is trained by the modified text data and the speech data corresponding to the modified text data, and includes:

[0054] Inputting the modified text data into the text encoder to obtain the predicted speech duration of the speech corresponding to the modified text data;

[0055] Using a Gaussian mixture-hidden Markov model to determine the true speech duration of the speech data;

[0056] Train the duration predictor and the length regulator in the text encoder based on the loss between the true speech duration and the predicted speech duration.

[0057] In this embodiment, during training, the duration predictor only takes text information as input and does not consider the durations of adjacent regions in the masked area. In this case, when the true speech speed is too fast or too slow, the predicted speech speed in the masked area may be inconsistent. To obtain accurate results that conform to the context phoneme durations, when predicting the duration, the predicted duration is processed as follows:

[0058]

[0059] where, is the adjusted duration of phoneme x i M is the index set of unmasked phonemes, represents the true duration of unmasked phoneme x j d′ j represents the duration of unmasked phoneme x j predicted by the duration predictor. The true duration of a phoneme can be obtained by forced alignment of the phoneme and speech using a Gaussian mixture-hidden Markov model (GMM-HMM). Train the duration predictor and the length regulator in the text encoder based on the loss between the true speech duration and the predicted speech duration until convergence.

[0060] As an embodiment, the modified part in the edited text includes: replacement, insertion, and deletion of target words in the edited text.

[0061] In this embodiment, if the user replaces a target word in the text, the user replaces a certain word in the original text with the target word, then inputs the replaced text into the model text encoder, and inputs the Mel spectrogram, pitch, and energy features of the original speech into the speech encoder. The duration predictor predicts the duration of the target word based on the text information, and then in the speech encoder, adjusts the length of the masked area to be consistent with the duration of the target word; the pitch, energy features, and Mel spectrogram features after masking are transformed into hidden representations by the transformer encoder in the speech encoder, and then the above-described inference process is performed. Finally, splice the Mel spectrogram of the target word output by the model with the original masked Mel spectrogram, and then the vocoder transforms the spliced complete Mel spectrogram into speech.

[0062] If the user inserts the target word in the text, the insertion operation is similar to the replacement operation. The user inserts the target word into a certain position in the text, then inputs the edited text into the text encoder, and inputs the Mel spectrum, pitch, and energy features of the original speech into the speech encoder. The duration predictor predicts the duration of the target word based on the text information, and then inserts the masked area with the same length as the predicted duration of the target word into the corresponding position of the Mel spectrum, pitch, and energy features of the original speech in the speech encoder; then the prediction is performed, and the generated Mel spectrum of the target word is inserted into the corresponding position of the Mel spectrum of the original speech, and the edited Mel spectrum is converted into speech using the vocoder.

[0063] If the user deletes the position of the target word in the text, there are two ways to delete the word. One is that the user deletes the target word in the text, and then the system deletes the Mel spectrum at the corresponding position, and then the vocoder converts the edited Mel spectrum into speech; the other is to use the replacement operation, that is, to replace the target word and the adjacent words with the words adjacent to the target word; in this way, the task of deleting the target word in the speech is completed through the replacement operation.

[0064] It can be seen from this implementation that the BERT-based speech context modeling enables the speech in the editing area to capture the rich speech context information of the recorded audio when predicting the speech, including features such as the speaker, environment, and pitch. This enables the model to make good use of the context information of the original speech, thereby predicting the speech in the editing area that is more consistent with the original audio, and can also avoid the unnatural and discontinuous speech caused by the splicing method.

[0065] The method is described in detail in experiments. The experiments are conducted on two English datasets, HiFiTTS and LibriTTS. The HiFiTTS dataset contains approximately 292 hours of audio data and corresponding text, with a total of 10 speakers. This method randomly selects 30 sentences for each speaker to form a speaker-visible test set, and uses the remaining HiFiTTS data as a training set to train the model. In addition, 8-9 sentences are randomly selected from the clean test set (test-clean) of LibriTTS, which contains a total of 39 speakers, for each speaker as a speaker-invisible test set to test the performance of the model.

[0066] All the speech audio sampling rates are 16 kHz. The original audio is extracted into 80-dimensional log-Mel filterbanks (Fbank) features, with a frame length of 50 ms and a frame shift of 12.5 used in the configuration. The G2P (grapheme-to-phoneme) and forced alignment information are both obtained by building a GMM-HMM model on Kaldi. The model proposed by this method is built on the Espnet tool, and HiFiTTS trained with the same training set is used as the vocoder.

[0067] There are three baseline models compared in this method. One is to use the TTS model to generate the whole speech of the edited text; the second is to use the TTS model to only synthesize the speech of the target word and insert the synthesized speech into the corresponding position of the original speech; the third is to use the TTS model to synthesize the whole speech of the edited text, but cut out the speech of the target word from it and then insert it into the corresponding position of the original speech.

[0068] This method uses two evaluation metrics, objective and subjective.

[0069] Objective evaluation experiment: The objective evaluation metric is the average MCD (Mel-cepstral distance) of the DTW (Dynamic time warping) path, and the lower the MCD, the higher the similarity. In the objective experiment, we randomly cover 1-4 words in each sentence and calculate the MCD between the target word and the whole sentence of the edited speech. To avoid the influence of the speech synthesized by the vocoder, this method cuts out the target word part in the speech synthesized by the proposed system and the baseline system and then inserts it into the corresponding position of the original speech. In this way, the experiment is more fair, and the experimental results are as Figure 3 shown in the MCD evaluation results of the model in the HiFiTTS test set (speaker visible) and the LibriTTS test set (speaker invisible) in the figure.

[0070] The BEdit-TTS model proposed by this method obtains the lowest MCD for both the speech of only the target word and the whole sentence speech in the speaker visible and speaker invisible test sets, which indicates that the speech synthesized by the BEdit-TTS model has better human perception and higher naturalness.

[0071] Subjective evaluation experiment: The subjective evaluation index is MOS (Mean Opinion Score). In the experiment, 15 sentences were randomly selected from the speaker-visible test set for substitution and insertion operations respectively. Since the baseline model did not use any speaker adaptation technology, for the speaker-invisible test set, 15 sentences were randomly selected for reconstruction operations and compared with the real speech. A total of 15 people participated in the scoring in this experiment. The participants needed to listen to all the audio and give scores. Before each sentence test, the participants were informed of the edited area. The scores ranged from 1 to 5, where 1 means very poor, 2 means poor, 3 means average, 4 means good, and 5 means very good. The experimental results are as Figure 3 shown.

[0072] The edited speech generated by BEdit-TTS synthesis after substitution and insertion operations on the speaker-visible test set obtained the highest MOS score, which indicates that the generated speech has high naturalness and quality, and the speech characteristics of the generated edited area conform to the characteristics of the original speech. At the same time, the boundary between the edited area and the non-edited area is also smooth and natural. After the BEdit-TTS model performs speech reconstruction operations on the covered area in the speaker-invisible test set, the generated speech has a score relatively close to the real speech, which shows that the model of this method still has good performance in the case of speaker invisibility.

[0073] In addition, with the substitution operation of the model of this method, VC (voice cloning) can be realized. The specific process is as follows: Voice cloning is achieved by repeatedly applying the substitution operation. Given a speech and the corresponding text as the original speech and the original text, and a target text. First, the target text and the original text are divided into the same number of parts, and each part corresponds to one or more words. Then, we replace the corresponding part of the original text with the part of the target text, and the model repeatedly performs the substitution operation with the replaced text and the mel spectrogram, pitch, and energy features of the original speech until all parts of the original text are replaced by the corresponding parts of the target text. Finally, the vocoder is used to convert the mel spectrogram that has undergone multiple editing operations into the speech of the target text. The corresponding process is as Figure 5 shown.

[0074] Generally speaking, this method proposes a new text-based speech editing model called BEdit TTS to simplify various operations on recorded audio, including substitution, insertion, and deletion. For the speech editing task, the requirements for the speech quality and acoustic consistency of the synthesized speech are equally important. To this end, BEdit TTS aims to integrate the advantages of neural TTS in high-fidelity audio generation and the advantages of BERT in context modeling. The experimental results show that the model proposed by this method can generate speech with good quality and highly similar to the recorded audio.

[0075] As shown Figure 6 in the structural schematic diagram of a text-based speech editing system provided by an embodiment of the present invention. The system can execute the text-based speech editing method described in any of the above embodiments and is configured in a terminal.

[0076] A text-based speech editing system 10 provided in this embodiment includes: a text encoding program module 11, a speech encoding program module 12, and a speech editing program module 13.

[0077] Among them, the text encoding program module 11 is used to input the edited text into a text encoder, determine a first speech duration corresponding to a modified part in the edited text and a second speech duration corresponding to the entire edited text, and determine a text representation of the edited text based on the second speech duration and the phoneme encoding of the edited text; the speech encoding program module 12 is used to input the first speech duration and the speech before the edited text is modified into a speech encoder, cover the corresponding modified part in the speech before modification based on the first speech duration, and obtain a covered acoustic representation, a hidden representation with covered context, and a Mel spectrogram with a covered area, where the length of the text representation is the same as the length of the Mel spectrogram with the covered area; the speech editing program module 13 is used to input the text representation, the covered acoustic representation, and the hidden representation with covered context into a joint network, obtain a predicted Mel spectrogram corresponding to the covered area, and obtain the speech after the edited text is modified based on the Mel spectrogram with the covered area and the predicted Mel spectrogram.

[0078] An embodiment of the present invention also provides a non-volatile computer storage medium. The computer storage medium stores computer-executable instructions, and the computer-executable instructions can execute the text-based speech editing method in any of the above method embodiments;

[0079] As an implementation, the non-volatile computer storage medium of the present invention stores computer-executable instructions, and the computer-executable instructions are set as:

[0080] Input the edited text into a text encoder, determine a first speech duration corresponding to a modified part in the edited text and a second speech duration corresponding to the entire edited text, and determine a text representation of the edited text based on the second speech duration and the phoneme encoding of the edited text;

[0081] Input the first speech duration and the speech before the edited text is modified into a speech encoder, cover the part of the speech before the modification corresponding to the modified part based on the first speech duration, and obtain a covered acoustic representation, a hidden representation with covered context, and a Mel spectrogram with covered regions, where the length of the text representation is the same as the length of the Mel spectrogram with covered regions;

[0082] Input the text representation, the covered acoustic representation, and the hidden representation with covered context into a joint network to obtain a predicted Mel spectrogram corresponding to the covered region, and obtain the speech after the edited text is modified based on the Mel spectrogram with covered regions and the predicted Mel spectrogram.

[0083] As a non-volatile computer-readable storage medium, it can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the method in the embodiments of the present invention. One or more program instructions are stored in the non-volatile computer-readable storage medium, and when executed by a processor, execute the text-based speech editing method in any of the above method embodiments.

[0084] Figure 7 It is a schematic diagram of the hardware structure of an electronic device for the text-based speech editing method provided in another embodiment of the present application, as Figure 7 shown. The device includes:

[0085] One or more processors 710 and a memory 720, Figure 7 Taking one processor 710 as an example. The device for the text-based speech editing method may further include: an input device 730 and an output device 740.

[0086] The processor 710, the memory 720, the input device 730, and the output device 740 may be connected through a bus or other means, Figure 7 Taking the connection through a bus as an example.

[0087] The memory 720, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the text-based speech editing method in the embodiments of the present application. The processor 710 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions, and modules stored in the memory 720, that is, implements the text-based speech editing method in the above method embodiments.

[0088] The memory 720 may include a program storage area and a data storage area. The program storage area may store an operating system and application programs required for at least one function. The data storage area may store data and the like. In addition, the memory 720 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 720 may optionally include a memory remotely disposed relative to the processor 710, and these remote memories may be connected to the mobile device through a network. Examples of the above networks include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0089] The input device 730 may receive input digital or character information. The output device 740 may include a display device such as a display screen.

[0090] The one or more modules are stored in the memory 720 and, when executed by the one or more processors 710, perform the text-based voice editing method in any of the above method embodiments.

[0091] The above product may execute the method provided in the embodiments of the present application, and has corresponding functional modules and beneficial effects for executing the method. Technical details not described in detail in this embodiment may be referred to the method provided in the embodiments of the present application.

[0092] The non-volatile computer-readable storage medium may include a program storage area and a data storage area. The program storage area may store an operating system and application programs required for at least one function. The data storage area may store data created according to the use of the device and the like. In addition, the non-volatile computer-readable storage medium may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the non-volatile computer-readable storage medium may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the device through a network. Examples of the above networks include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0093] An embodiment of the present invention further provides an electronic device, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the steps of the text-based voice editing method according to any embodiment of the present invention.

[0094] The electronic device in the embodiments of the present application exists in various forms, including but not limited to:

[0095] (1) Mobile communication devices: These devices are characterized by having mobile communication functions and mainly aim to provide voice and data communications. Such terminals include: smart phones, multimedia phones, functional phones, and low-end phones, etc.

[0096] (2) Ultra-mobile personal computer devices: These devices belong to the category of personal computers, have computing and processing functions, and generally also have the feature of mobile Internet access. Such terminals include: PDA, MID, and UMPC devices, etc., such as tablet computers.

[0097] (3) Portable entertainment devices: These devices can display and play multimedia content. Such devices include: audio and video players, handheld game consoles, e-books, and intelligent toys and portable vehicle navigation devices.

[0098] (4) Other electronic devices with data processing functions.

[0099] In this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include" and "comprise", not only include those elements, but also include other elements not expressly listed, or also include elements inherent to such a process, method, article, or device. Without further limitation, elements defined by the statement "including..." do not preclude the existence of additional identical elements in the process, method, article, or device that includes the said elements.

[0100] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative effort.

[0101] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course also by hardware. Based on such an understanding, the above technical solutions, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0102] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A text-based speech editing method, comprising: Inputting the editing text into a text encoder to determine a first speech duration corresponding to the modified part in the editing text and a second speech duration corresponding to the entire editing text, and determining a text representation of the editing text based on the second speech duration and the phoneme encoding of the editing text; Inputting the first speech duration and the speech before the editing text is modified into a speech encoder, masking the corresponding modified part in the speech before the modification based on the first speech duration to obtain a masked acoustic representation, a hidden representation with masked context, and a Mel spectrogram with a masked region, wherein the length of the text representation is the same as the length of the Mel spectrogram with the masked region; Inputting the text representation, the masked acoustic representation, and the hidden representation with masked context into a joint network to obtain a predicted Mel spectrogram corresponding to the masked region, and obtaining the speech after the editing text is modified based on the Mel spectrogram with the masked region and the predicted Mel spectrogram.

2. The method according to claim 1, wherein, The text encoder includes: a phoneme embedding block, an encoder, a duration predictor, and a length regulator, wherein The phoneme embedding block is used to determine the phoneme embedding of the editing text; The encoder is used to determine the text representation of the editing text according to the phoneme embedding and the corresponding position encoding; The duration predictor is used to determine a first speech duration corresponding to the modified part in the editing text and a second speech duration corresponding to the entire editing text; The length regulator is used to adjust the length of the text representation according to the second speech duration so that the length of the adjusted text representation is the same as the length of the Mel spectrogram with the masked region.

3. The method according to claim 1, wherein The speech encoder includes: a masking operation block and a conversion encoder, wherein The masking operation block is used to receive the speech before the editing text is modified and the corresponding Mel spectrogram, mask the corresponding modified part in the speech before the editing text is modified according to the first speech duration corresponding to the modified part in the editing text to obtain a masked acoustic representation and a Mel spectrogram with a masked region; The conversion encoder is used to transcode the Mel spectrogram with the masked region to obtain a hidden representation with masked context.

4. The method according to claim 1, wherein The joint network includes: a pitch energy converter and a Mel spectrogram decoder, wherein The pitch energy converter is used to generate a pitch energy feature and a predicted Mel spectrogram feature that simulate real speech according to the received text representation, masked acoustic representation, and hidden representation with masked context; The Mel spectrogram decoder is used to determine a predicted Mel spectrogram corresponding to the masked region according to the pitch energy feature and predicted Mel spectrogram feature that simulate real speech.

5. The method according to claim 2, wherein, The text encoder is trained by modified text data and speech data corresponding to the modified text data, and includes: Inputting the modified text data into the text encoder to obtain a predicted speech duration of the speech corresponding to the modified text data; Using a Gaussian mixture-hidden Markov model to determine the real speech duration of the speech data; Train the duration predictor and the length regulator in the text encoder based on the loss between the true speech duration and the predicted speech duration.

6. The method according to claim 1, wherein, The modified part in the edited text includes: replacement, insertion, and deletion of target words in the edited text.

7. A text-based speech editing system, comprising: A text encoding program module, configured to input the edited text into a text encoder, determine a first speech duration corresponding to the modified part in the edited text and a second speech duration corresponding to the entire edited text, and determine a text representation of the edited text based on the second speech duration and the phoneme encoding of the edited text; A speech encoding program module, configured to input the first speech duration and the speech before the modification of the edited text into a speech encoder, mask the part corresponding to the modified part in the speech before the modification based on the first speech duration, to obtain a masked acoustic representation, a hidden representation with masked context, and a Mel spectrogram with masked regions, wherein the length of the text representation is the same as the length of the Mel spectrogram with masked regions; A speech editing program module, configured to input the text representation, the masked acoustic representation, and the hidden representation with masked context into a joint network, obtain a predicted Mel spectrogram corresponding to the masked region, and obtain the speech after the modification of the edited text based on the Mel spectrogram with masked regions and the predicted Mel spectrogram.

8. The system according to claim 7, wherein, The text encoder is obtained by training with modified text data and speech data corresponding to the modified text data, and includes: Input the modified text data into the text encoder to obtain the predicted speech duration of the speech corresponding to the modified text data; Use a Gaussian mixture-hidden Markov model to determine the true speech duration of the speech data; Train the duration predictor and the length regulator in the text encoder based on the loss between the true speech duration and the predicted speech duration.

9. An electronic device, comprising: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps of the method according to any one of claims 1-6.

10. A storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • BERT-based text classification method and device, computer equipment and storage medium

    CN112328786A

  • Speech synthesis method and system for new tone generation

    CN112802448A