Cross-sentence conditional coherence speech editing method, system and terminal

Through the speech editing method that is inter-sentenced and coherent, the overall reasoning generates and edits Mel's spectrum, which solves the problems of boundary discontinuity of editing areas and changes in tone context in the existing system, and realizes high-fidelity reconstruction and efficient resource utilization.

CN116189653BActive Publication Date: 2025-07-22SHANGHAI TECH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310146999.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2023-01-13
Filing Date
2023-02-21
Publication Date
2025-07-22
Estimated Expiration
2043-02-21

AI Technical Summary

Technical Problem

The existing voice editing system adopts partial reasoning methods, which makes the boundary discontinuity of the editing area and the changes in tone and context difficult to deal with, and cannot efficiently save time or computing resources.

Method used

A cross-sentence condition-coherent speech editing method is adopted, and the overall reasoning of the variational autoencoder and decoder is used to generate and edit Mel spectrograms using audio features and context semantic information, including phoneme conversion, context information capture, context embedding and editing modules, and the loss function is trained to optimize the editing process.

Benefits of technology

High-fidelity reconstruction of the original waveform unmodified area is achieved, avoiding incoherence in splicing and no additional resources are consumed, and improving the naturalness and rhythmic consistency of speech editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116189653B_ABST
    Figure CN116189653B_ABST
Patent Text Reader

Abstract

The cross-sentence conditional coherence voice editing method, system and terminal of the present invention use a voice editing model with a variational autoencoder and a decoder that take the audio features and context semantic information in the voice input information as conditional inputs, obtain the corresponding edited Mel spectrogram according to the voice information to be edited, and can reconstruct the unmodified area of the original waveform with high fidelity. By using global inference instead of partial inference, the incoherence at the joints caused by splicing can be completely avoided. In addition, compared with the existing partial inference editing system, the global inference method of the present invention does not consume additional resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech editing, and particularly to a speech editing method, system and terminal with cross-sentence conditional coherence. Background Art

[0002] Speech editing can be applied to various fields with personalized speech requirements and higher requirements for speech naturalness, including video production on social media, game and movie dubbing, etc. Traditional speech editing tools allow users to perform functions such as denoising, adjusting volume, cutting, copying and pasting waveforms. Among them, when the audio text to be edited needs to be modified, the traditional speech editing work will be relatively cumbersome. Especially when new words that are not in the audio transcription text appear, only the corresponding segment needs to be re-recorded and then spliced with the original audio. However, changes in the recording environment and the state of the speaker may both result in differences in background noise, loudness and pitch rhythm between the re-recorded speech segment and the original speech, and the listening experience after splicing will not be natural enough.

[0003] In order to reduce the workload of audio recorders and post-production, speech editing based on text transcription is an emerging audio editing technology. This technology can synthesize speech that matches the pitch and timbre of the original audio according to the text changed by the content editor. Therefore, instead of editing the original audio, the editing burden can be reduced by modifying the text transcription corresponding to the original audio.

[0004] However, when existing speech editing systems based on text transcription are reasoning, they all adopt partial reasoning rather than overall reasoning. Specifically, the direct output of existing editing systems is a complete waveform or Mel spectrogram, which corresponds to the edited full-sentence transcription text. However, in order to improve the similarity with the original audio, existing methods need to additionally intercept the segment that must be modified and then insert it into the original waveform or Mel spectrogram.

[0005] For example, previous work has overcome the prosody mismatch problem caused by directly connecting audio in different scenarios based on the digital signal processing (DSP) part. A method that uses a neural network to predict prosody information and integrates the TD-PSOLA algorithm, denoising, and dereverberation to achieve prosody modification. Although the above system supports cut, copy, and paste operations, it cannot insert or replace new words that do not exist in the speech data of the same speaker. In recent years, research has applied text-to-speech (TTS) systems to synthesize missing inserted words. VoCo uses comparable TTS speech synthesis to insert words and then uses a voice conversion (VC) model to convert them into voices suitable for the target speaker. EditSpeech proposes a partial inference and bidirectional fusion method to achieve smooth transitions at the editing boundaries. CampNet performs masked training on a Transformer-based context-aware neural network to improve the quality of edited speech. Recently, a pre-training method for perceiving aligned acoustic text (A3T) has been proposed. This framework can reconstruct masked acoustic signals with high quality through text input and acoustic text alignment during training and can be directly applied to speech editing. As mentioned above, all speech editing systems to date have adopted a partial inference method. Therefore, it is difficult to avoid incoherence at the splicing point and unable to handle changes in the tone context after text modification. Although this editing method preserves the original audio as much as possible, it also causes the following potential problems:

[0006] Problem 1: The partial inference artificially inserts the predicted acoustic features of the editing area into the corresponding positions of the original waveform. Therefore, discontinuities near the boundaries of the editing area are almost inevitable to a certain extent. At the same time, the direct output of existing speech editing systems based on partial inference is still the entire sentence audio including context segments. Therefore, compared with global inference, it does not save more time or computing resources.

[0007] Problem 2: After the text is modified, the pitch and prosody will also change accordingly. That is to say, we should not blindly pursue making the audio corresponding to the modified text sound exactly the same as the original audio. A special example is that when a general question can be modified into a declarative sentence, partial inference can hardly handle the change in tone because this approach directly uses the original audio segment. Summary of the Invention

[0008] In view of the above-mentioned shortcomings of the prior art, the purpose of the present invention is to provide a cross-sentence conditionally coherent speech editing method, system, and terminal to solve the above technical problems in the prior art.

[0009] To achieve the above and other related objectives, the present invention provides a cross-sentence conditional coherence speech editing method, which includes: obtaining speech input information to be edited; wherein, the speech input information includes: an initial Mel spectrogram, the current transcribed text sentence and the same number of text sentences before and after it; based on a mask-trained speech editing model, obtaining a corresponding edited Mel spectrogram according to the speech information to be edited; wherein, the speech editing model includes: a variational autoencoder that takes the audio features and context semantic information in the speech input information as conditional inputs, and a decoder.

[0010] In an embodiment of the present invention, the variational autoencoder includes: a phoneme conversion module for converting the input current transcribed text sentence into phoneme sequence information; a context information capture module for respectively capturing context information for each sentence pair recombined by the current transcribed text sentence and the same number of text sentences before and after it, and generating BERT embedding information corresponding to each sentence pair; a context embedding module connecting the phoneme conversion module and the context information capture module for obtaining cross-sentence representation output data and phoneme durations based on the phoneme sequence information, target speaker feature information, and each BERT embedding information; an editing module connecting the context embedding module for generating corresponding edited speech data based on the initial Mel spectrogram, cross-sentence representation output data, and phoneme durations and outputting the same for the decoder to decode to obtain a corresponding edited Mel spectrogram.

[0011] In an embodiment of the present invention, the context embedding module includes: an encoding sub-module for encoding the phoneme sequence information and target speaker feature information; a fusion sub-module connecting the encoding sub-module for fusing the encoded phoneme sequence information, target speaker feature information, and each BERT embedding information to obtain and output cross-sentence representation output data; a time prediction sub-module connecting the fusion sub-module for performing time prediction and adjustment based on the cross-sentence representation output data to output phoneme durations.

[0012] In an embodiment of the present invention, the time prediction sub-module includes: a duration predictor for obtaining predicted phoneme durations based on cross-sentence representation output data; a duration adjuster connecting the duration predictor for adjusting based on the predicted phoneme durations to obtain phoneme durations.

[0013] In an embodiment of the present invention, the editing module includes: a replacement processing sub-module, configured to perform replacement processing on the initial Mel spectrogram based on a deletion indicator corresponding to a target deletion position area and an addition indicator corresponding to a target addition position area to obtain corresponding mean sequence processing data and variance sequence processing data; a context statement processing sub-module, configured to obtain corresponding cross-statement mean sequence data and cross-statement variance sequence data based on two one-dimensional convolutional modules according to cross-statement characterization output data and phoneme durations; an editing output sub-module, connected to the replacement processing sub-module and the context statement processing sub-module, configured to obtain editing parameters based on the mean sequence processing data, the variance sequence processing data, the cross-statement mean sequence data, and the cross-statement variance sequence data, so as to generate corresponding edited speech data and output it.

[0014] In an embodiment of the present invention, the replacement processing sub-module includes: a deletion editing unit, configured to modify the Mel spectrogram based on a deletion indicator corresponding to a target deletion position area, and obtain first mean sequence data and first variance sequence data based on two one-dimensional convolutional modules; an addition editing unit, connected to the deletion editing unit, configured to modify the first mean sequence data and the first variance sequence data based on an addition indicator corresponding to a target addition position area to obtain mean sequence processing data and variance sequence processing data.

[0015] In an embodiment of the present invention, the addition editing unit includes: a first processing sub-unit, configured to insert the first mean sequence data and the first variance sequence data into sequences having the same length as the target addition position area respectively based on the addition indicator corresponding to the target position area to generate second mean sequence data and second variance sequence data; a second processing sub-unit, connected to the first processing sub-unit, configured to perform one-dimensional convolution on the second mean sequence data and the second variance sequence data to obtain mean sequence processing data and variance sequence processing data.

[0016] In an embodiment of the present invention, the speech editing model is obtained by mask training using a loss function; wherein, the loss function includes: a non-mask loss function and a mask loss function.

[0017] To achieve the above and other related objectives, the present invention provides a cross-sentence conditionally coherent speech editing system, which includes: an acquisition module for acquiring speech input information to be edited; wherein, the speech input information includes: an initial Mel spectrogram, the current transcribed text sentence, and the same number of text sentences before and after it; an editing module connected to the acquisition module for obtaining a corresponding edited Mel spectrogram based on a speech editing model trained with masks according to the speech information to be edited; wherein, the speech editing model includes: a variational autoencoder and a decoder that take the audio features and context semantic information in the speech input information as conditional inputs.

[0018] To achieve the above and other related objectives, the present invention provides a cross-sentence conditionally coherent speech editing terminal, including: one or more memories and one or more processors; the one or more memories for storing computer programs; the one or more processors connected to the memories for running the computer programs to execute the cross-sentence conditionally coherent speech editing method.

[0019] As described above, the present invention is a cross-sentence conditionally coherent speech editing method, system, and terminal, having the following beneficial effects: By using a speech editing model with a variational autoencoder and a decoder that take the audio features and context semantic information in the speech input information as conditional inputs, the present invention can obtain a corresponding edited Mel spectrogram according to the speech information to be edited, and can reconstruct the unmodified region of the original waveform with high fidelity. By using global inference instead of partial inference, the incoherence at the splicing joints caused by splicing can be completely avoided. In addition, compared with existing partial inference editing systems, the global inference method of the present invention does not consume additional resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 It shows a schematic flowchart of the cross-sentence conditionally coherent speech editing method in an embodiment of the present invention.

[0021] Figure 2 It shows a schematic structural diagram of the variational autoencoder in an embodiment of the present invention.

[0022] Figure 3 It shows a schematic structural diagram of the context embedding module in an embodiment of the present invention.

[0023] Figure 4 It shows a schematic structural diagram of the editing module in an embodiment of the present invention.

[0024] Figure 5 It shows a schematic structural diagram of the speech editing model in an embodiment of the present invention.

[0025] Figure 6It shows a schematic structural diagram of a cross - sentence conditional coherence voice editing system in an embodiment of the present invention.

[0026] Figure 7 It shows a schematic structural diagram of a cross - sentence conditional coherence voice editing terminal in an embodiment of the present invention. Detailed implementation manners

[0027] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0028] It should be noted that in the following description, reference is made to the accompanying drawings, which describe several embodiments of the present invention. It should be understood that other embodiments can also be used, and mechanical composition, structure, electrical, and operational changes can be made without departing from the spirit and scope of the present invention. The following detailed description should not be considered restrictive, and the scope of the embodiments of the present invention is only defined by the claims of the published patent. The terms used here are only for describing specific embodiments and are not intended to limit the present invention. Spatially related terms, such as "upper", "lower", "left", "right", "below", "beneath", "lower part", "above", "upper part", etc., can be used in the text to facilitate the description of the relationship between one element or feature shown in the figure and another element or feature.

[0029] Throughout the specification, when it is said that a part is "connected" to another part, this includes not only the case of "direct connection", but also the case of "indirect connection" with other elements placed in between. In addition, when it is said that a certain part "includes" a certain constituent element, unless there is a particularly contrary record, it does not mean excluding other constituent elements, but means that other constituent elements can also be included.

[0030] The first, second, and third terms mentioned therein are used to illustrate various parts, components, regions, layers, and / or segments, but are not limited thereto. These terms are only used to distinguish one part, component, region, layer, or segment from other parts, components, regions, layers, or segments. Therefore, the first part, component, region, layer, or segment described below can refer to the second part, component, region, layer, or segment within the scope not exceeding the present invention.

[0031] Furthermore, as used herein, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. It should be further understood that the terms "comprising", "including" indicate the presence of the stated features, operations, elements, components, items, kinds, and / or groups, but do not preclude the presence, occurrence, or addition of one or more other features, operations, elements, components, items, kinds, and / or groups. The terms "or" and "and / or" used herein are to be construed as inclusive, meaning either one or any combination. Thus, "A, B, or C" or "A, B, and / or C" means "any of the following: A; B; C; A and B; A and C; B and C; A, B, and C". An exception to this definition occurs only when the combination of elements, functions, or operations is inherently mutually exclusive in some way.

[0032] Audio editing is designed to enable users to perform selection, cutting, copying, and pasting operations based on voice recordings. Some advanced neural network-based editing systems perform partial inference rather than whole, that is, only generate new words that need to be replaced or inserted, which usually results in the prosody of the edited part being inconsistent with the previous and subsequent speech.

[0033] The present invention provides a cross-sentence conditionally coherent speech editing method, system, and terminal. Through a speech editing model with a variational autoencoder and a decoder that takes audio features and context semantic information in the speech input information as conditional inputs, an edited Mel spectrogram corresponding to the speech information to be edited is obtained, and the unmodified region of the original waveform can be reconstructed with high fidelity. By using whole inference instead of partial inference, the incoherence at the splicing joints caused by splicing can be completely avoided. In addition, compared with existing partial inference editing systems, the whole inference method of the present invention does not consume additional resources.

[0034] The following will be described in detail with reference to the accompanying drawings for the embodiments of the present invention so that those skilled in the technical field of the present invention can easily implement it. The present invention can be embodied in many different forms and is not limited to the embodiments described herein.

[0035] As Figure 1 A schematic flowchart showing a cross-sentence conditionally coherent speech editing method in an embodiment of the present invention is presented.

[0036] Step S1: Obtain the speech input information to be edited.

[0037] Specifically, the speech input information includes: an initial Mel spectrogram x extracted from the original waveform i , the current transcribed text sentence u i and the text sentences with the same target number L before and after it, that is, a total of 2L + 1 sentences are jointly formed: u i-L , u i-L+1,..., u i-1 , u i ,...u i+L-1 , u i+L 。

[0038] Step S2: Based on the speech editing model trained with a mask, obtain the corresponding edited Mel spectrogram according to the speech information to be edited.

[0039] Specifically, the speech editing model includes: a variational autoencoder and a decoder that take the audio features and context semantic information in the speech input information as conditional inputs; specifically, the speech input information is input into the variational autoencoder to extract the audio features and context semantic information for encoding, and then input into the decoder for corresponding decoding to obtain the corresponding edited Mel spectrogram.

[0040] In one embodiment, as Figure 2 shown, the variational autoencoder includes:

[0041] Phoneme conversion module 1, used to convert the input current transcribed text statement u i into phoneme sequence information p i ; specifically, by using a text-to-phoneme conversion tool (G2P), the transcribed text statement u i is converted into phonemes p i . In addition, the start and end times of each phoneme can be extracted by Montreal forced alignment.

[0042] Context information capture module 2, used to capture context information for each sentence pair recombined by the current transcribed text statement and the same number of text statements before and after it, and generate BERT embedding information corresponding to each sentence pair;

[0043] Specifically, the 2L + 1 sentences: u i-L , u i-L+1 ,..., u i-1 , u i ,...u i+L-1 , u i+L are recombined into 2L sentence pairs: that is, [(u i-L , u i-L+1 ),..., (u i-1 , u i ), (u i+L-1 , u i+L )]; and the pre-trained language model BERT is used to capture context information, thereby generating 2L BERT embeddings, that is, [b -L , b -L+1 ,..., b L-1 .

[0044] The context embedding module 3, which is connected to the phoneme conversion module 1 and the context information capture module 2, is used to obtain the cross-sentence representation output data H i , the target speaker feature information Si, and each BERT embedding information i and the phoneme duration D' i ; wherein, the target speaker feature information Si can be set according to specific requirements.

[0045] The editing module 4, which is connected to the context embedding module 3, is used to generate corresponding edited speech data and output it based on the initial Mel spectrogram x i , the cross-sentence representation output data H i and the phoneme duration D' i so that the decoder can decode it to obtain the corresponding edited Mel spectrogram.

[0046] In one embodiment, the context embedding module 3 includes:

[0047] An encoding sub-module, which is used to encode the phoneme sequence information p i and the target speaker feature information Si;

[0048] A fusion sub-module, which is connected to the encoding sub-module, is used to fuse the encoded phoneme sequence information, target speaker feature information, and each BERT embedding information to obtain the cross-sentence representation output data H i and output it;

[0049] A time prediction sub-module, which is connected to the fusion sub-module, is used to perform time prediction and adjustment based on the cross-sentence representation output data H i to output the phoneme duration D' i .

[0050] In one embodiment, the time prediction sub-module includes:

[0051] A duration predictor, which is used to obtain the predicted phoneme duration D based on the cross-sentence representation output data H i ; i

[0052] A duration regulator, which is connected to the duration predictor, is used to perform adjustment based on the predicted phoneme duration D i to obtain the phoneme duration D' i .

[0053] In a specific embodiment, in order to more effectively utilize the speaker information contained in the original audio and make the speaking rate of the edited area consistent with that of the unedited area, similar to the methods of EditSpeech and A3T, we use the original audio and the predicted duration of the unedited area audio to further finely adjust the phoneme duration of the edited area.

[0054]

[0055] where ∑D Unedit and respectively represent the sum of the true and predicted durations of all phonemes in the unedited area, represents the duration of the adjusted edited area. When inferring the mel spectrogram subsequently, the true time D Unedit is used for the unedited area, and the finely adjusted is used for the edited area to reconstruct an audio close to the real speaking rhythm.

[0056] In a specific embodiment, as Figure 3 shown, the Transformer encoder and the speaker encoder are respectively used to encode the phoneme sequence information p i and the target speaker feature information Si; and the multi-head attention layer is used to further capture the context semantic information by using each BERT embedding information, and then the captured context semantic information is fused with the encoded phoneme sequence information p i and the target speaker feature information Si, and then the cross-sentence representation output data H i is obtained by using the linear mapping layer; and then the predicted phoneme duration D i is obtained through the duration predictor based on the cross-sentence representation output data H i , and finally the phoneme duration D' i is obtained through adjustment by the duration regulator obtained by the trained model.

[0057] In one embodiment, the editing module 4 aims to overcome the defect that the existing speech editing system cannot restore the unmodified audio part, so it has to splice the modified part with the original mel spectrogram or audio. The specific implementation details are as follows:

[0058] The editing module 4 includes:

[0059] A replacement processing sub-module, which is used to process the initial mel spectrogram x iPerform replacement processing to obtain the corresponding mean sequence processed data μ' and variance sequence processed data σ'; specifically, since the replacement operation in editing can be regarded as deletion before addition, we can use only two flags to indicate the positions for deleting and adding corresponding content, denoted as Flag del and Flag add . Denote the corresponding transcribed text of the original speech as [u a , u b , u c . Correspondingly, the phonemes corresponding to the text can be denoted as p i = [p a , p b , p c . Denote the Mel spectrogram of the original speech as x i = [x a , x b , x c .

[0060] Deletion indicator. The deletion process enables the user to delete a segment of the speech waveform associated with a specific set of words. The target statement to be synthesized after deletion is [u a , u c , where u b is the deleted part. By comparing the utterances before and after editing, we can obtain the corresponding deletion indicator Flag del = [0 a , 1 b , 0 c . This indicator is a 0-1 sequence for subsequent indication of the editing of the Mel spectrogram.

[0061] Addition indicator. Different from the deletion operation, the target synthesized speech after insertion or replacement is based on the edited text [u a , u b’ , u c , where u b’ replaces the content of u b . The insertion process can be regarded as a special case. Similar to the deletion operation, we can obtain the addition indicator Flag add = [0 a , 1 b’ , 0 c .

[0062] Context statement processing sub-module, which is used to obtain the corresponding cross-statement mean sequence data μ i and cross-statement variance sequence data σ i based on two one-dimensional convolutional modules according to the cross-statement representation output data H prior and the phoneme duration D' prior ;

[0063] An editing output sub-module, connected to the replacement processing sub-module and the context statement processing sub-module, is used to process data μ' according to the mean sequence, process data σ' according to the variance sequence, and process cross-statement mean sequence data μ prior and cross-statement variance sequence data σ prior to obtain an editing parameter Z, so as to generate corresponding edited voice data and output it.

[0064] Specifically, the way to obtain the editing parameter Z includes: The editing module can complete sampling from the estimated prior and is re-parameterized as:

[0065]

[0066] where are element addition and multiplication operations. z prior Sampling from the specific statement prior, with the input being the cross-statement representation output H and the duration D, and the re-parameterization is as follows:

[0067]

[0068] where, μ prior and σ prior are learned from the statement-specific prior module, and ∈ is sampled from the standard Gaussian.

[0069] In one embodiment, the replacement processing sub-module includes:

[0070] A deletion editing unit, which is used to modify the mel spectrogram based on the deletion indicator in the corresponding target deletion position area, and obtain the first mean sequence data and the first variance sequence data based on two one-dimensional convolution modules; specifically, based on the deletion indicator Flag del , x i is modified to [x a , x c , and the mean μ and variance σ are learned through two one-dimensional convolutions.

[0071] An addition editing unit, connected to the deletion editing unit, is used to modify the first mean sequence data and the first variance sequence data based on the addition indicator in the corresponding target addition position area, so as to obtain the mean sequence processed data and the variance sequence processed data.

[0072] In a specific implementation, the addition editing unit includes:

[0073] The first processing subunit is configured to insert the first mean sequence data and the first variance sequence data into sequences having the same length as the target addition position area respectively based on the addition indicator corresponding to the target position area, so as to generate second mean sequence data and second variance sequence data; specifically, according to the addition indicator Flag add , add 0 and 1 sequences at the corresponding positions of μ and σ, that is and where the lengths of the 0 sequence and the 1 sequence are the same as b'. Therefore, different from the module in CUCVAE where the input is the unmodified audio before editing, in order to be closer to the audio editing scenario, the speech generated in the editing area in the present invention is sampled from the context sentence prior, while the audio in the area that does not need to be edited is sampled from the real audio and the context sentence prior.

[0074] The second processing subunit is connected to the first processing subunit and is configured to perform one-dimensional convolution on the second mean sequence data and the second variance sequence data to obtain mean sequence processed data and variance sequence processed data. And in order to make the editing boundary more coherent, the second mean sequence data and the second variance sequence data and further generate mean sequence processed data and variance sequence processed data μ' and σ' through one-dimensional convolution.

[0075] In a specific embodiment, as Figure 4 shown, based on the deletion indicator Flag del , x i is modified to [x a , x c , and the mean μ i and the variance σ i are learned through two one-dimensional convolutions. Then, based on the addition indicator Flag add , 0 and 1 sequences are added at the corresponding positions of μ and σ through the upsampling layer, that is and Then, mean sequence processed data and variance sequence processed data μ' and σ' are generated through one-dimensional convolution, realizing that the speech generated in the editing area is sampled from the context sentence prior, while the audio in the area that does not need to be edited is sampled from the real audio and the context sentence prior. Based on two one-dimensional convolution modules, according to the cross-sentence representation output data H i and the phoneme duration D' i , the corresponding cross-sentence mean sequence data μ prior and the cross-sentence variance sequence data σ prior are learned from the sentence-specific prior module, and then z prior sampled from the specific sentence prior is calculated., then sample from the estimated prior and reparameterize to obtain the editing parameter z, generate the corresponding edited speech data and output it.

[0076] In one embodiment, to reconstruct the waveform of the masked part, the common loss function of the acoustic model is the mean absolute error (MAE) between the reconstructed mel spectrogram and the original mel spectrogram. And similar to the approach in BERT, to make the system more focused on the masked part, only the loss of the masked region is calculated. However, during the inference process, the system needs to synthesize coherent audio that conforms to the context of the modified text. Thus, in the scenario of speech editing, it is not appropriate to ignore the loss of the non-masked part.

[0077] Therefore, to balance the two purposes of being close to the original audio and generating new audio that conforms to the context rhythm, we consider the loss of the non-masked part, and to make the system more focused on the generation quality of the editing region, we increase the loss weight of the masked part.

[0078] That is, the speech editing model is obtained by masked training using a loss function; wherein, the loss function includes: a non-masked loss function and a masked loss function.

[0079] In a preferred embodiment, the loss function is as follows:

[0080]

[0081] wherein, the loss weights of the masked part and the non-masked part are λ. and respectively represent the prediction and the target true mel spectrogram of the i-th frame.

[0082] In the experiment, the weight λ is set to 1.5. With this weight, the generated edited audio can be coherent while also retaining the prosodic features of the original audio.

[0083] To better illustrate the above cross-sentence conditional coherent speech editing method, the present invention provides the following specific embodiments.

[0084] Embodiment 1: A cross-sentence conditional coherent speech editing model. As Figure 5 is the framework structure diagram of the speech editing model.

[0085] The model includes: a variational autoencoder with masked training and a decoder;

[0086] wherein, the variational autoencoder includes: using a text-to-phoneme conversion tool (G2P), a pre-trained language model BERT, an editing module and a context sentence embedding module;

[0087] The input of the model includes the mel spectrogram x extracted from the original waveformi and the current transcribed text sentence u i and the sentences before and after it are used as input. By using a text-to-phoneme conversion tool (G2P), the transcribed text sentence u i is converted to phonemes p i . Meanwhile, 2L + 1 sentences are recombined into 2L sentence pairs, namely [(u i-L , u i-L+1 ),..., (u i-1 , u i ), (u i+L-1 , u i+L )], and the pre-trained language model BERT is used to capture context information, thereby generating 2L BERT embeddings, namely [b -L , b -L+1 ,..., b L-1 .

[0088] The context sentence embedding module respectively encodes the phoneme sequence information p i and the target speaker feature information Si by using a Transformer encoder and a speaker encoder; and further captures the context semantic information through a multi-head attention layer by using each BERT embedding information, and then fuses the captured context semantic information with the encoded phoneme sequence information p i and the target speaker feature information Si, and then uses a linear mapping layer to obtain the cross-sentence representation output data H i ; and then obtains the predicted phoneme duration D i through a duration predictor based on the cross-sentence representation output data H i , and finally adjusts it through a duration regulator obtained by a training model to obtain the phoneme duration D' i .

[0089] The editing module, based on the deletion indicator Flag del , x i is modified to [x a , x c , and the mean μ i and variance σ i are learned through two one-dimensional convolutions. Then, based on the addition indicator Flag add , a 0 and 1 sequence is added to the corresponding positions of μ and σ through an upsampling layer, that is and Then, a mean sequence processing data and a variance sequence processing data μ' and σ' are generated through one-dimensional convolution, realizing that the generated speech in the editing area is sampled from the context sentence prior, while the audio in the area that does not need to be edited is sampled from the real audio and the context sentence prior. Based on two one-dimensional convolution modules, according to the cross-sentence representation output data Hi and the phoneme duration D′ i learn the corresponding cross-sentence mean sequence data μ from the sentence-specific prior module prior and the cross-sentence variance sequence data σ prior , and then calculate z sampled from the specific sentence prior prior , then sample from the estimated prior and reparameterize to obtain the editing parameter z, generate the corresponding edited speech data and output it.

[0090] The speech editing model also uses a loss function that takes into account the loss of the unmasked part for model training. The loss deviation set in the training phase can prompt the system to focus more on the masked segments during training, that is, the part of the mel spectrogram that needs to be reconstructed.

[0091] Then, the editing data output by the variational autoencoder is decoded by the decoder to obtain the corresponding edited mel spectrogram, and then converted into audio data by the vocoder and output.

[0092] The model performance proposed in this embodiment is evaluated through qualitative listening tests and quantitative measurements. In the subjective listening test, we selected 15 synthetic audio files for subjective listening tests, recruited 20 volunteers to conduct subjective opinion scoring evaluations (MOS) with a maximum score of 5 for the naturalness of the speech samples, and provided a 95% confidence interval and p-value. In addition to subjective opinion scoring, the F0 frame error (FFE), mel cepstral distortion (MCD), and word error rate (WER) were all evaluated on 512 test samples. Experimental results of various editing operations on the multi-person dataset (LibriTTS) show that the system proposed in this embodiment is superior to some inference methods in terms of naturalness and prosodic consistency.

[0093] According to the MOS naturalness scores shown in Table 1, our overall inference model has an advantage of about 0.5 points in naturalness compared to the worst method. The gap in the replacement operation is more significant because it is difficult for the speech editing system based on partial inference to handle intonation conversion. At the same time, since the partial inference system based on the mel spectrogram highly depends on the accuracy of MFA, its performance in the deletion operation is relatively worse, especially when short words are deleted, and its performance may be worse than the manually refined deletion method based on the waveform. In addition, the partial inference based on the waveform has relatively low naturalness MOS scores in insertion and replacement because it involves inserting new words and there is disharmony between the original audio and the generated audio.

[0094] The MOS similarity score indicates that the partial inference based on waveform is the most similar to the original waveform, which is reasonable and as expected. At the same time, the system performance of the overall inference in this embodiment is also quite close to that of the partial inference, with the maximum difference being only 0.2. In terms of objective reconstruction performance, the result of this embodiment is close to that of the editing system based on partial inference, proving that our editing module can reconstruct the mel spectrogram with high quality.

[0095] Table 1 Objective and subjective results of the partial inference system based on mel spectrogram or waveform and the overall inference (this embodiment)

[0096]

[0097] The MOS similarity score indicates that the partial inference based on waveform is the most similar to the original waveform, which is reasonable and as expected. At the same time, the system performance of the overall inference in this embodiment is also quite close to that of the partial inference, with the maximum difference being only 0.2. In terms of objective reconstruction performance, the result of this embodiment is close to that of the editing system based on partial inference, proving that our editing module can reconstruct the mel spectrogram with high quality.

[0098] The p-values of the significance analysis in Table 2 more intuitively show that, except in the deletion case, the edited audio of this embodiment is significantly superior to the two partial inference systems in terms of naturalness. There is no significant difference in the deletion operation because the overall inference based on waveform is through artificial fine positioning, so it performs better in the deletion operation. However, when inserting or replacing, once new words need to be inserted, this method will be significantly worse than the system proposed in this embodiment. In addition, there is no significant difference between this embodiment and the partial inference in terms of similarity to the real audio. It further shows that this embodiment can reconstruct the acoustic features of the original audio with high fidelity, and the prosody of the synthesized speech conforms to the semantic context after editing.

[0099] Table 2 Results of the significance analysis of using the overall inference (this embodiment) and the two partial inference systems in terms of naturalness and similarity MOS scores

[0100]

[0101] Therefore, the system proposed in this embodiment aims to utilize the reconstruction ability of the conditional variational autoencoder based on context information to reconstruct the unchanged part of the audio with high fidelity, and according to the edited transcription text, can synthesize new audio with the same prosody rhythm as the original audio. During the training process, part of the audio is randomly masked to simulate the effect of audio editing. During inference, the system can use the speaker information, context features, and the mel spectrogram of the unedited segments in the original speech to reconstruct the speech with high quality.

[0102] Similar to the principle of the above embodiments, the present invention provides a cross-sentence conditional coherent speech editing system.

[0103] The following provides specific embodiments in conjunction with the accompanying drawings:

[0104] As Figure 6 Show a schematic structural diagram of a cross-sentence conditional coherent speech editing system in an embodiment of the present invention.

[0105] The system includes:

[0106] An acquisition module 61, configured to acquire speech input information to be edited; wherein, the speech input information includes: an initial Mel spectrogram, a current transcribed text sentence, and the same number of text sentences before and after it.

[0107] An editing module 62, connected to the acquisition module 61, configured to obtain a corresponding edited Mel spectrogram according to the speech information to be edited based on a speech editing model trained by masking.

[0108] Wherein, the speech editing model includes: a variational autoencoder and a decoder that take the audio features and context semantic information in the speech input information as conditional inputs.

[0109] Since the implementation principle of this cross-sentence conditional coherent speech editing system has been described in the foregoing embodiments, it will not be repeated here.

[0110] In one embodiment, the system includes: acquiring speech input information to be edited; wherein, the speech input information includes: an initial Mel spectrogram, a current transcribed text sentence, and the same number of text sentences before and after it; a speech editing model trained by masking, obtaining a corresponding edited Mel spectrogram according to the speech information to be edited; wherein, the speech editing model includes: a variational autoencoder and a decoder that take the audio features and context semantic information in the speech input information as conditional inputs.

[0111] In one embodiment, the variational autoencoder includes: a phoneme conversion module for converting the input current transcribed text statement into phoneme sequence information; a context information capture module for respectively capturing context information for each statement pair recombined from the current transcribed text statement and the same number of text statements before and after it, and generating BERT embedding information corresponding to each statement pair; a context embedding module connected to the phoneme conversion module and the context information capture module for obtaining cross-statement representation output data and phoneme durations based on the phoneme sequence information, target speaker feature information, and each BERT embedding information; and an editing module connected to the context embedding module for generating corresponding edited speech data based on the initial Mel spectrogram, cross-statement representation output data, and phoneme durations and outputting the same for the decoder to decode to obtain the corresponding edited Mel spectrogram.

[0112] In one embodiment, the context embedding module includes: an encoding sub-module for encoding the phoneme sequence information and target speaker feature information; a fusion sub-module connected to the encoding sub-module for fusing the encoded phoneme sequence information, target speaker feature information, and each BERT embedding information to obtain and output cross-statement representation output data; and a time prediction sub-module connected to the fusion sub-module for performing time prediction and adjustment based on the cross-statement representation output data to output phoneme durations.

[0113] In one embodiment, the time prediction sub-module includes: a duration predictor for obtaining predicted phoneme durations based on cross-statement representation output data; and a duration adjuster connected to the duration predictor for adjusting based on the predicted phoneme durations to obtain phoneme durations.

[0114] In one embodiment, the editing module includes: a replacement processing sub-module for performing replacement processing on the initial Mel spectrogram based on a deletion indicator corresponding to a target deletion position area and an addition indicator corresponding to a target addition position area to obtain corresponding mean sequence processing data and variance sequence processing data; a context statement processing sub-module for obtaining corresponding cross-statement mean sequence data and cross-statement variance sequence data based on two one-dimensional convolutional modules according to the cross-statement representation output data and phoneme durations; and an editing output sub-module connected to the replacement processing sub-module and the context statement processing sub-module for obtaining editing parameters based on the mean sequence processing data, variance sequence processing data, cross-statement mean sequence data, and cross-statement variance sequence data to generate corresponding edited speech data and output the same.

[0115] In one embodiment, the replacement processing sub-module includes: a deletion editing unit configured to modify the Mel spectrogram based on a deletion indicator corresponding to a target deletion position area, and obtain first mean sequence data and first variance sequence data based on two one-dimensional convolutional modules; an addition editing unit connected to the deletion editing unit, configured to modify the first mean sequence data and the first variance sequence data based on an addition indicator corresponding to a target addition position area to obtain mean sequence processed data and variance sequence processed data.

[0116] In one embodiment, the addition editing unit includes: a first processing subunit configured to insert the first mean sequence data and the first variance sequence data into sequences having the same length as the target addition position area respectively based on the addition indicator corresponding to the target position area to generate second mean sequence data and second variance sequence data; a second processing subunit connected to the first processing subunit, configured to perform one-dimensional convolution on the second mean sequence data and the second variance sequence data to obtain mean sequence processed data and variance sequence processed data.

[0117] In one embodiment, the speech editing model is obtained by masked training using a loss function; wherein, the loss function includes: a non-masked loss function and a masked loss function.

[0118] As Figure 7 FIG. shows a schematic structural diagram of a cross-sentence conditionally coherent speech editing terminal 70 in an embodiment of the present invention.

[0119] The cross-sentence conditionally coherent speech editing terminal 70 includes: a memory 71 and a processor 72. The memory 71 is configured to store a computer program; the processor 72 runs the computer program to implement as Figure 1 the cross-sentence conditionally coherent speech editing method as described above.

[0120] Optionally, the number of the memories 71 can be one or more, and the number of the processors 72 can be one or more, and Figure 7 one is taken as an example in both cases.

[0121] Optionally, the processor 72 in the cross-sentence conditionally coherent speech editing terminal 70 will, according to the steps as Figure 1 described above, load instructions corresponding to the processes of one or more application programs into the memory 71, and the processor 72 runs the application programs stored in the first memory 71, so as to implement various functions in Figure 1 the cross-sentence conditionally coherent speech editing method as described above.

[0122] Optionally, the memory 71 may include, but is not limited to, high-speed random access memory and non-volatile memory. For example, one or more disk storage devices, flash memory devices, or other non-volatile solid-state storage devices; the processor 72 may include, but is not limited to, a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0123] Optionally, the processor 72 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0124] The present invention also provides a computer-readable storage medium storing a computer program, and when the computer program runs, it implements the cross-sentence conditional coherence voice editing method as Figure 1 shown. The computer-readable storage medium may include, but is not limited to, a floppy disk, an optical disk, a CD-ROM (compact disc read-only memory), a magneto-optical disk, a ROM (read-only memory), a RAM (random access memory), an EPROM (erasable programmable read-only memory), an EEPROM (electrically erasable programmable read-only memory), a magnetic card or an optical card, a flash memory, or other types of media / machine-readable media suitable for storing machine-executable instructions. The computer-readable storage medium may be a product not connected to a computer device, or a component already connected to and used by a computer device.

[0125] In summary, the cross-sentence conditional coherence voice editing method, system, and terminal of the present invention use a voice editing model with a variational autoencoder and a decoder that take the audio features and context semantic information in the voice input information as conditional inputs, and obtain the corresponding edited Mel spectrogram according to the voice information to be edited, and can reconstruct the unmodified area of the original waveform with high fidelity. By using global inference instead of partial inference, the incoherence at the splicing joints caused by splicing can be completely avoided. In addition, compared with the existing partial inference editing system, the global inference method of the present invention does not consume additional resources. Therefore, the present invention effectively overcomes various shortcomings in the prior art and has high industrial utilization value.

[0126] The above embodiments are only used to exemplarily illustrate the principles and effects of the present invention, rather than to limit the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes completed by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed by the present invention should still be covered by the claims of the present invention.

Claims

1. A cross-sentence conditional coherence voice editing method, characterized in that, The method includes: Obtaining voice input information to be edited; wherein, the voice input information includes: an initial Mel spectrogram, a current transcribed text sentence, and the same number of text sentences before and after it; Based on a voice editing model trained with masking, obtaining a corresponding edited Mel spectrogram according to the voice information to be edited; Wherein, the voice editing model includes: a variational autoencoder and a decoder that take audio features and context semantic information in the voice input information as conditional inputs; The variational autoencoder includes: A phoneme conversion module for converting the input current transcribed text sentence into phoneme sequence information; A context information capture module for respectively capturing context information for each sentence pair recombined by the current transcribed text sentence and the same number of text sentences before and after it, and generating BERT embedding information corresponding to each sentence pair; A context embedding module connected to the phoneme conversion module and the context information capture module for obtaining cross-sentence representation output data and phoneme durations based on the phoneme sequence information, target speaker feature information, and each BERT embedding information; An editing module connected to the context embedding module for generating corresponding edited voice data based on the initial Mel spectrogram, cross-sentence representation output data, and phoneme durations and outputting it for the decoder to decode to obtain a corresponding edited Mel spectrogram.

2. The method for cross-sentence conditional coherence speech editing according to claim 1, wherein The context embedding module includes: An encoding sub-module for encoding the phoneme sequence information and target speaker feature information; A fusion sub-module connected to the encoding sub-module for fusing the encoded phoneme sequence information, target speaker feature information, and each BERT embedding information to obtain and output cross-sentence representation output data; A time prediction sub-module connected to the fusion sub-module for performing time prediction and adjustment based on the cross-sentence representation output data to output phoneme durations.

3. The method for cross-sentence conditional coherence speech editing according to claim 2, wherein The time prediction sub-module includes: A duration predictor for obtaining predicted phoneme durations based on cross-sentence representation output data; A duration adjuster connected to the duration predictor for adjusting based on the predicted phoneme durations to obtain phoneme durations.

4. The method for cross-sentence conditional coherence-based speech editing according to claim 1, characterized in that, The editing module includes: A replacement processing sub-module for performing replacement processing on the initial Mel spectrogram based on a deletion indicator in a corresponding target deletion position area and an addition indicator in a corresponding target addition position area to obtain corresponding mean sequence processing data and variance sequence processing data; A context sentence processing sub-module for obtaining corresponding cross-sentence mean sequence data and cross-sentence variance sequence data based on two one-dimensional convolutional modules according to cross-sentence representation output data and phoneme durations; An editing output sub-module connected to the replacement processing sub-module and the context sentence processing sub-module for obtaining editing parameters based on the mean sequence processing data, variance sequence processing data, cross-sentence mean sequence data, and cross-sentence variance sequence data to generate corresponding edited voice data and output it.

5. The method for cross-sentence conditional coherent speech editing according to claim 4, wherein The replacement processing sub-module includes: A deletion editing unit, configured to modify a Mel spectrogram based on a deletion indicator corresponding to a target deletion position area, and obtain first mean sequence data and first variance sequence data based on two one-dimensional convolutional modules; An addition editing unit, connected to the deletion editing unit, configured to modify the first mean sequence data and the first variance sequence data based on an addition indicator corresponding to a target addition position area to obtain mean sequence processed data and variance sequence processed data.

6. The method for cross-sentence conditional coherence speech editing according to claim 5, wherein The addition editing unit includes: A first processing subunit, configured to insert the first mean sequence data and the first variance sequence data into sequences having the same length as the target addition position area respectively based on the addition indicator corresponding to the target position area to generate second mean sequence data and second variance sequence data; A second processing subunit, connected to the first processing subunit, configured to perform one-dimensional convolution on the second mean sequence data and the second variance sequence data to obtain mean sequence processed data and variance sequence processed data.

7. The cross-sentence conditional coherence speech editing method according to claim 1, characterized in that, The speech editing model is obtained by mask training using a loss function; wherein, the loss function includes: a non-mask loss function and a mask loss function.

8. A cross-sentence conditional coherence voice editing system, characterized in that, The system includes: An acquisition module, configured to acquire speech input information to be edited; wherein, the speech input information includes: an initial Mel spectrogram, a current transcription text statement, and text statements with the same number of the same targets before and after the current transcription text statement; An editing module, connected to the acquisition module, configured to obtain a corresponding edited Mel spectrogram according to the speech information to be edited based on a speech editing model trained by masking; Wherein, the speech editing model includes: a variational autoencoder and a decoder that take audio features and context semantic information in the speech input information as conditional inputs; The variational autoencoder includes: A phoneme conversion module, configured to convert an input current transcription text statement into phoneme sequence information; A context information capture module, configured to capture context information for each statement pair recombined by the current transcription text statement and text statements with the same number of the same targets before and after the current transcription text statement respectively, and generate BERT embedding information corresponding to each statement pair; A context embedding module, connected to the phoneme conversion module and the context information capture module, configured to obtain cross-statement representation output data and phoneme durations based on the phoneme sequence information, target speaker feature information, and each BERT embedding information; An editing module, connected to the context embedding module, configured to generate corresponding edited speech data based on the initial Mel spectrogram, cross-statement representation output data, and phoneme durations and output the edited speech data for the decoder to decode to obtain a corresponding edited Mel spectrogram.

9. A cross-sentence conditionally coherent speech editing terminal, characterized in that, Includes: One or more memories and one or more processors; The one or more memories are configured to store computer programs; The one or more processors, connected to the memories, are configured to run the computer programs to execute the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Cross-statement speech synthesis method, system and equipment based on variational automatic encoder

    CN114566141A

  • KR20220105043A