A voice editing method, device and electronic equipment

CN121506157BActive Publication Date: 2026-08-11BEIJING YUNSHANG TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-02
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0005]本发明的目的在于提供一种语音编辑方法、装置和电子设备,用于解决现有技术中语音编辑结果存在编辑精度较低、无法保证语音编辑结果的与原用户的音色、语速及语气风格的统一性的问题

Benefits of technology

本发明实施例所提供的一种语音编辑方法,通过获取待合成文本、原始文本、以及原始音频的原始频谱特征,其中,原始文本以及原始频谱特征是基于时间戳进行对齐的。根据待合成文本、原始文本、原始频谱特征以及时间戳,获取待合成文本对应的重构掩码频谱特征、以及重构掩码频谱特征对应的噪声。将重构掩码频谱特征、噪声、待合成文本以及原始频谱特征输入至训练好的语音编辑模型中,获取待合成文本的目标频谱特征,其中,语音编辑模型包括:文本特征提取模块、声音特征提取模块、频谱特征预测模块以及目标频谱特征拼接模块,语音编辑模型是基于预设训练集对声音特征提取模块的权重参数进行微调确定的。将目标频谱特征通过声码器进行波形重建,获取待合成文本的语音编辑结果。这样,在综合考虑待合成文本、原始文本以及原始音频的原始频谱特征的基础上,获取重构掩码频谱特征,避免了现有技术过度依赖人工经验和判断所导致的编辑精度较低问题。此外,进一步在语音编辑模型中引入声音特征提取模块,能够精准获取用户的声音特征,从而确保生成的语音编辑结果与用户的音色、语速及语气风格保持高度一致,显著提升了语音编辑结果的精度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121506157B_ABST
    Figure CN121506157B_ABST
Patent Text Reader

Abstract

This invention discloses a speech editing method, apparatus, and electronic device. It acquires the original spectral features of the text to be synthesized, the original text, and the original audio. Based on the text to be synthesized, the original text, the original spectral features, and a timestamp, it obtains the reconstructed mask spectral features and corresponding noise corresponding to the text to be synthesized. The reconstructed mask spectral features, noise, the text to be synthesized, and the original spectral features are input into a trained speech editing model to obtain target spectral features. The speech editing model includes a text feature extraction module, a sound feature extraction module, a spectral feature prediction module, and a target spectral feature concatenation module. The speech editing model is determined by fine-tuning the weight parameters of the sound feature extraction module based on a preset training set. The target spectral features are reconstructed using a vocoder to obtain the speech editing result. This ensures that the speech editing result is consistent with the user's timbre, speech rate, and tone style, and improves the accuracy of the speech editing result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech synthesis technology, and in particular to a speech editing method, apparatus, and electronic device. Background Technology

[0002] Voice editing technology is developing rapidly and is widely used in fields such as intelligent assistants and audiobook dubbing. The goal of voice editing technology is to modify, optimize, or synthesize speech content through techniques such as editing, splicing, and pitch adjustment to obtain the desired voice editing result.

[0003] In existing technologies, speech editing techniques are mainly divided into two categories: waveform editing and feature editing. Specifically, waveform editing refers to directly cutting, pasting, and other operations on audio waveforms to obtain the speech editing result. Feature editing, on the other hand, is based on modifying acoustic features such as Mel-Frequency Cepstral Coefficients (MFCCs) to replace or add speech content.

[0004] However, based on existing technology, waveform editing relies heavily on human experience and judgment, resulting in low editing accuracy. During feature editing, when replacing or adding speech content, it is also difficult to ensure the naturalness and expressiveness of the newly generated speech, and it is impossible to guarantee the consistency of the speech editing result with the original user's timbre, speech rate and tone style. Summary of the Invention

[0005] The purpose of this invention is to provide a voice editing method, apparatus, and electronic device to solve the problems in the prior art where the voice editing results have low editing accuracy and cannot guarantee the consistency of the voice editing results with the original user's timbre, speech rate, and tone style.

[0006] To achieve the above objectives, in a first aspect, embodiments of the present invention provide a voice editing method, comprising: Obtain the original spectral features of the text to be synthesized, the original text, and the original audio, wherein the original text and the original spectral features are aligned based on timestamps; Based on the text to be synthesized, the original text, the original spectral features, and the timestamp, obtain the reconstructed mask spectral features corresponding to the text to be synthesized, and the noise corresponding to the reconstructed mask spectral features; The reconstructed mask spectral features, the noise, the text to be synthesized, and the original spectral features are input into a trained speech editing model to obtain the target spectral features of the text to be synthesized. The speech editing model includes a text feature extraction module, a sound feature extraction module, a spectral feature prediction module, and a target spectral feature concatenation module. The speech editing model is determined by fine-tuning the weight parameters of the sound feature extraction module based on a preset training set. The target spectral features are reconstructed using a vocoder to obtain the speech editing result of the text to be synthesized.

[0007] In one embodiment, before obtaining the reconstructed mask spectral features corresponding to the text to be synthesized and the noise corresponding to the reconstructed mask spectral features based on the text to be synthesized, the original text, the original spectral features, and the timestamp, the method further includes: The original text and the original spectral features are input into an automatic speech recognition model, and the automatic speech recognition model aligns the original text and the original spectral features based on the timestamp.

[0008] In one embodiment, obtaining the reconstructed mask spectral features corresponding to the text to be synthesized and the noise corresponding to the reconstructed mask spectral features based on the text to be synthesized, the original text, the original spectral features, and the timestamp includes: Obtain multiple pairs of differing subtexts of the text to be synthesized and the original text, each pair of differing subtexts including: a first differing subtext in the original text and a corresponding second differing subtext in the text to be synthesized; The original text, each of the first difference sub-texts and each of the corresponding second difference sub-texts are segmented into words to obtain multiple first words corresponding to the original text, multiple second words corresponding to each of the first difference sub-texts and multiple third words corresponding to each of the second difference sub-texts. An upsampling operation is performed on the plurality of first word segments, the plurality of second word segments, the plurality of third word segments, and the original spectral features to obtain the target spectral feature length corresponding to the text to be synthesized; Based on the target spectral feature length, the original spectral feature length, and the original spectral feature, a reconstruction mask operation is performed based on the timestamp to obtain the reconstructed mask spectral feature, wherein the length of the reconstructed mask spectral feature is the same as the target spectral feature length. The noise is obtained based on the reconstructed mask spectral features, wherein the length of the noise is the same as the length of the target spectral features.

[0009] In one embodiment, the step of performing a reconstruction mask operation based on the timestamp, according to the target spectral feature length, the original spectral feature length, and the original spectral feature, to obtain the reconstructed mask spectral features includes: Determine whether the length of the target spectral feature is greater than the length of the original spectral feature; If the target spectral feature length is greater than the original spectral feature length, then a padding and reconstruction mask operation is performed on the original spectral feature based on the timestamp to obtain the reconstructed mask spectral feature.

[0010] In one embodiment, the method further includes: If the length of the target spectral feature is not greater than the length of the original spectral feature, then the original spectral feature is cropped and reconstructed using the timestamp to obtain the reconstructed mask spectral feature.

[0011] In one embodiment, the step of inputting the reconstructed mask spectral features, the noise, the text to be synthesized, and the original spectral features into a trained speech editing model to obtain the target spectral features of the text to be synthesized includes: The text to be synthesized is input into the text feature extraction module, and the text features of the text to be synthesized are obtained through the text feature extraction module. The original spectral features are input into the sound feature extraction module, and sound features are obtained through the sound feature extraction module; The text features, the reconstructed mask spectral features, the noise, and the sound features are input into the spectral feature prediction module, and the predicted spectral features corresponding to the text to be synthesized are obtained through the spectral feature prediction module. The predicted spectral features and the original spectral features are input into the target spectral feature stitching module, and the target spectral features are obtained through the target spectral feature stitching module.

[0012] In one embodiment, the sound feature extraction module includes an encoding layer and an upsampling layer. The original spectral features are input into the sound feature extraction module, and sound features are obtained through the sound feature extraction module by: The original spectral features are input into the coding layer for encoding processing to obtain encoded features; The encoded features are input into the upsampling layer for upsampling processing to obtain the sound features.

[0013] In one embodiment, the step of inputting the predicted spectral features and the original spectral features into the target spectral feature stitching module, and obtaining the target spectral features through the target spectral feature stitching module, includes: In the original spectral features, first predicted spectral sub-features corresponding to multiple first differential sub-texts are determined respectively; In the predicted spectral features, determine the second predicted spectral sub-features corresponding to multiple second differential sub-texts respectively; The target spectrum feature splicing module is obtained by replacing multiple first predicted spectrum sub-features in the original spectrum features with multiple second predicted spectrum sub-features.

[0014] Secondly, embodiments of the present invention provide a voice editing device, comprising: The acquisition module is used to acquire the original spectral features of the text to be synthesized, the original text, and the original audio, wherein the original text and the original spectral features are aligned based on timestamps; The reconstruction mask spectral feature and noise acquisition module is used to acquire the reconstruction mask spectral feature corresponding to the text to be synthesized and the noise corresponding to the reconstruction mask spectral feature based on the text to be synthesized, the original text, the original spectral feature and the timestamp; The target spectral feature acquisition module is used to input the reconstructed mask spectral features, the noise, the text to be synthesized, and the original spectral features into a trained speech editing model to obtain the target spectral features of the text to be synthesized. The speech editing model includes a text feature extraction module, a sound feature extraction module, a spectral feature prediction module, and a target spectral feature concatenation module. The speech editing model is determined by fine-tuning the weight parameters of the sound feature extraction module based on a preset training set. The speech editing result acquisition module is used to reconstruct the waveform of the target spectral features through a vocoder to obtain the speech editing result of the text to be synthesized.

[0015] Thirdly, embodiments of the present invention provide an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the method described in the first aspect.

[0016] The technical solution provided by the embodiments of the present invention has the following advantages compared with the prior art: This invention provides a speech editing method that acquires the original spectral features of the text to be synthesized, the original text, and the original audio, wherein the original text and the original spectral features are aligned based on timestamps. Based on the text to be synthesized, the original text, the original spectral features, and the timestamps, the method acquires the reconstructed mask spectral features corresponding to the text to be synthesized, as well as the noise corresponding to the reconstructed mask spectral features. The reconstructed mask spectral features, noise, the text to be synthesized, and the original spectral features are input into a trained speech editing model to obtain the target spectral features of the text to be synthesized. The speech editing model includes a text feature extraction module, a sound feature extraction module, a spectral feature prediction module, and a target spectral feature concatenation module. The speech editing model is determined by fine-tuning the weight parameters of the sound feature extraction module based on a preset training set. The target spectral features are reconstructed using a vocoder to obtain the speech editing result of the text to be synthesized. Thus, by comprehensively considering the original spectral features of the text to be synthesized, the original text, and the original audio, the method acquires the reconstructed mask spectral features, avoiding the low editing accuracy problem caused by the excessive reliance on human experience and judgment in existing technologies. Furthermore, by introducing a voice feature extraction module into the voice editing model, the user's voice features can be accurately obtained, thereby ensuring that the generated voice editing results are highly consistent with the user's timbre, speech rate, and tone style, significantly improving the accuracy of the voice editing results. Attached Figure Description

[0017] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this invention, illustrate exemplary embodiments of the invention and are used to explain the invention, but do not constitute an undue limitation of the invention. In the drawings: Figure 1 A flowchart illustrating a voice editing method provided in an embodiment of the present invention; Figure 2 A schematic diagram of a speech editing model provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of a voice editing device provided in an embodiment of the present invention. Detailed Implementation

[0018] To facilitate a clear description of the technical solutions in the embodiments of the present invention, the terms "first" and "second" are used to distinguish identical or similar items with essentially the same function and effect. For example, the first threshold and the second threshold are merely used to distinguish different thresholds and do not limit their order. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and that the terms "first" and "second" are not necessarily different.

[0019] It should be noted that in this invention, the terms "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in this invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0020] In this invention, "at least one" means one or more, and "more than one" means two or more. "And / or" describes the relationship between the associated objects, indicating that three relationships can exist.

[0021] like Figure 1 As shown, Figure 1 This is a flowchart illustrating a voice editing method provided in an embodiment of the present invention, which specifically includes the following steps: S10: Obtain the original spectral features of the text to be synthesized, the original text, and the original audio.

[0022] The original text and original spectral features are aligned based on timestamps. The original spectral features refer to audio features commonly used in speech signal processing. They are formed by mapping frequency domain information to the Mel scale, which simulates the nonlinear frequency distribution perceived by human hearing, and combining logarithmic operations.

[0023] The text to be synthesized refers to the text for which the voice editing result needs to be obtained, and the original text refers to the text that is similar to the text to be synthesized in terms of text content and form. For example, the text to be synthesized may be "the AAAA that everyone respects, a bit CC", and the original text may be "the BB that everyone admires, a bit DD", but it is not limited to this. This invention does not specifically limit it, and those skilled in the art can set it according to the actual situation.

[0024] Specifically, the original spectral features of the text to be synthesized, the original text, and the original audio corresponding to the original text are obtained.

[0025] S11: Based on the text to be synthesized, the original text, the original spectral features, and the timestamp, obtain the reconstructed mask spectral features corresponding to the text to be synthesized, and the noise corresponding to the reconstructed mask spectral features.

[0026] Specifically, after obtaining the text to be synthesized, the original text, and the original spectral features, the reconstruction mask spectral features corresponding to the text to be synthesized and the noise corresponding to the reconstruction mask spectral features are obtained based on the text to be synthesized, the original text, the original spectral features, and the timestamp.

[0027] Optionally, based on the above embodiments, in some embodiments of the present invention, the method further includes the following step before performing S11: S20: Input the original text and original spectral features into the automatic speech recognition model, and align the original text and original spectral features based on the timestamp through the automatic speech recognition model.

[0028] Specifically, the original text and original spectral features are input into the Automatic Speech Recognition (ASR) model, which then aligns the original text and original spectral features based on timestamps.

[0029] In this way, this example can achieve a precise frame-level mapping relationship between each word in the original text and the original spectral features.

[0030] Optionally, based on the above embodiments, in some embodiments of the present invention, one implementation of S11 may be: S111: Obtain multiple differential subtext pairs of the text to be synthesized and the original text.

[0031] Each difference text pair includes: a first difference text in the original text, and a corresponding second difference text in the text to be synthesized.

[0032] For example, following the above embodiments, the text to be synthesized may be, for example, “AAAA, which is respected by all, is a bit CC”, and the original text may be, for example, “BB, which is admired by thousands, is a bit DD”. Then the difference subtext pairs are: the first difference subtext “admired by thousands” in the original text corresponds to the second difference subtext “respected by all” in the text to be synthesized; or, the first difference subtext “BB” in the original text corresponds to the second difference subtext “AAAA” in the text to be synthesized; or, the first difference subtext “DD” in the original text corresponds to the second difference subtext “CC” in the text to be synthesized. However, this invention is not limited to these, and those skilled in the art can set them according to the actual situation.

[0033] Specifically, a greedy difference algorithm is used to identify multiple differential subtext pairs between the text to be synthesized and the original text.

[0034] S112: Perform word segmentation on the original text, each first difference subtext, and each corresponding second difference subtext to obtain multiple first words corresponding to the original text, multiple second words corresponding to each first difference subtext, and multiple third words corresponding to each second difference subtext.

[0035] Specifically, after obtaining multiple pairs of differing subtexts between the text to be synthesized and the original text, word segmentation is performed on the original text, each first differing subtext included in each differing subtext pair, and each corresponding second differing subtext, in order to obtain multiple first words corresponding to the original text, multiple second words corresponding to each first differing subtext, and multiple third words corresponding to each second differing subtext.

[0036] For example, following the above embodiments, if the original text is, for example, “Baidu, which is admired by thousands, is a little bit DD”, then the multiple first segment words corresponding to the original text “Baidu, which is admired by thousands, is a little bit DD” are “Baidu, which is admired by thousands”, “of”, “BB”, “a little bit”, and “DD”; the first difference subtext “BB” in the original text, the multiple second segment words corresponding to the first difference subtext “BB” are “B”, “B”; the corresponding second difference subtext in the text to be synthesized is “AAAA”, then the multiple third segment words corresponding to the second difference subtext “AAAA” are “A”, “A”, “A”, “A”, but not limited thereto. The present invention is not specifically limited, and those skilled in the art can set it according to the actual situation.

[0037] S113: Perform upsampling operations on multiple first-word segments, multiple second-word segments, multiple third-word segments, and the original spectral features to obtain the target spectral feature length corresponding to the text to be synthesized.

[0038] Specifically, upsampling is performed on multiple first-word segments, multiple second-word segments, multiple third-word segments, and the original spectral features to obtain the target spectral feature length corresponding to the text to be synthesized.

[0039] S114: Based on the target spectral feature length, the original spectral feature length, and the original spectral feature, perform a reconstruction mask operation based on the timestamp to obtain the reconstructed mask spectral features.

[0040] The length of the reconstructed mask spectral feature is the same as the length of the target spectral feature.

[0041] Optionally, based on the above embodiments, in some embodiments of the present invention, one implementation of S114 may be: S1141: Determine whether the target spectral feature length is greater than the original spectral feature length.

[0042] S1142: If the target spectral feature length is greater than the original spectral feature length, then the original spectral feature is padded and reconstructed using a mask based on the timestamp to obtain the reconstructed mask spectral feature.

[0043] Specifically, compare the length of the target spectral feature with the length of the original spectral feature to determine whether the length of the target spectral feature is greater than the length of the original spectral feature. When it is determined that the length of the target spectral feature is greater than the length of the original spectral feature, perform a padding and reconstruction mask operation on the original spectral feature based on the time stamp, so as to obtain the length of the target spectral feature corresponding to the text to be synthesized.

[0044] Exemplarily, continuing with the above embodiment, for the text to be synthesized "AAAA, who is respected by everyone, is a bit CC", the length of the target spectral feature corresponding to this text to be synthesized is 14. For the original text "BB, who is admired by thousands of people, is a bit DD", the length of the original spectral feature corresponding to this original text is 12. It is determined that the length of the target spectral feature is greater than the length of the original spectral feature. Based on the time stamp, perform a padding and reconstruction mask operation behind the original spectral feature corresponding to the first word segmentation "BB". For example, two lengths of 0 can be padded, so as to obtain a reconstructed mask spectral feature with a length of 14. However, this is not limited to this. The present invention does not specifically limit it, and those skilled in the art can set it according to the actual situation.

[0045] Optionally, based on the above embodiment, in some embodiments of the present invention, another implementation manner of S114 may be: S1143: If the length of the target spectral feature is not greater than the length of the original spectral feature, perform a cropping and reconstruction mask operation on the original spectral feature based on the time stamp to obtain a reconstructed mask spectral feature.

[0046] Specifically, compare the length of the target spectral feature with the length of the original spectral feature to determine whether the length of the target spectral feature is greater than the length of the original spectral feature. When it is determined that the length of the target spectral feature is not greater than the length of the original spectral feature, perform a cropping and reconstruction mask operation on the original spectral feature based on the time stamp, so as to obtain the length of the target spectral feature corresponding to the text to be synthesized.

[0047] S115: Obtain noise according to the reconstructed mask spectral feature.

[0048] Among them, the length of the noise is the same as the length of the target spectral feature.

[0049] Specifically, after obtaining the reconstructed mask spectral feature, generate noise with the same length as the length of the target spectral feature according to the reconstructed mask spectral feature.

[0050] S12: Input the reconstructed mask spectral feature, noise, text to be synthesized, and original spectral feature into the trained speech editing model to obtain the target spectral feature of the text to be synthesized.

[0051] Among them, as Figure 2As shown, the speech editing model includes: a text feature extraction module 10, a sound feature extraction module 20, a spectrum feature prediction module 30, and a target spectrum feature splicing module 40.

[0052] The speech editing model is determined by fine-tuning the weight parameters of the voice feature extraction module based on a preset training set.

[0053] Optionally, based on the above embodiments, in some embodiments of the present invention, the specific implementation method of determining the speech editing model by fine-tuning the weight parameters of the sound feature extraction module based on a preset training set can be: Obtain a preset training set, which includes: multiple training texts, training audio corresponding to each training text, and user identifier corresponding to each training text.

[0054] Input the preset training set into the speech editing model, freeze the weight parameters of the speech editing model including the text feature extraction module, the spectral feature prediction module, and the target spectral feature concatenation module, and fine-tune only the weight parameters of the sound feature extraction module until the loss function converges to obtain the trained speech editing model.

[0055] In this way, this embodiment can fine-tune only the weight parameters of the sound feature extraction module in the speech editing model by using a preset training set, so as to ensure that the random initialization parameters of the sound feature extraction module can converge effectively, thereby extracting accurate sound features, reducing computational complexity, and improving training efficiency.

[0056] Specifically, after obtaining the reconstructed mask spectral features and the noise corresponding to the reconstructed mask spectral features, the reconstructed mask spectral features, noise, text to be synthesized, and original spectral features are input into the trained speech editing model, and the target spectral features of the text to be synthesized are obtained through the speech editing model.

[0057] Optionally, based on the above embodiments, in some embodiments of the present invention, S12 may be implemented as follows: S121: Input the text to be synthesized into the text feature extraction module, and obtain the text features of the text to be synthesized through the text feature extraction module.

[0058] Specifically, the text to be synthesized is input into the text feature extraction module, which then performs feature extraction on the text to be synthesized to obtain the text features of the text to be synthesized.

[0059] S122: Input the original spectral features into the sound feature extraction module, and obtain the sound features through the sound feature extraction module.

[0060] Voice features include information such as timbre and style. This voice feature extraction module is based on the Conformer encoder. Specifically, it compresses the input vector through a cross-attention mechanism and extracts it into a fixed-dimensional latent vector set. This vector set contains high-dimensional information such as the speaker's timbre and style.

[0061] Specifically, the original spectral features are input into the sound feature extraction module, which then performs feature extraction processing on the original spectral features to obtain the sound features of the original spectral features.

[0062] Optionally, based on the above embodiments, continue to refer to... Figure 2 As shown, the sound feature extraction module 20 includes an encoding layer and an upsampling layer. Based on this, in some embodiments of the present invention, one implementation of S122 can be: S1221: Input the original spectral features into the coding layer for encoding processing to obtain the encoded features.

[0063] Specifically, the original spectral features are input into the coding layer, and the coding layer encodes the original spectral features to form the encoded features.

[0064] S1222: Input the encoded features into the upsampling layer for upsampling processing to obtain sound features.

[0065] Specifically, the encoded features are used as input to the upsampling layer, and the upsampling layer performs upsampling processing on the encoded features to obtain the original spectral features of the sound features.

[0066] S123: Input the text features, reconstructed mask spectral features, noise, and sound features into the spectral feature prediction module, and obtain the predicted spectral features corresponding to the text to be synthesized through the spectral feature prediction module.

[0067] Specifically, after obtaining the text features of the text to be synthesized and the sound features of the original spectral features, the text features, the reconstructed mask spectral features, noise, and sound features are input into the spectral feature prediction module. The spectral feature prediction module performs prediction to obtain the predicted spectral features corresponding to the text to be synthesized.

[0068] It should be noted that this spectral feature prediction module is a text-to-speech (TTS) model.

[0069] S124: Input the predicted spectral features and the original spectral features into the target spectral feature stitching module, and obtain the target spectral features through the target spectral feature stitching module.

[0070] Specifically, after obtaining the predicted spectral features, the predicted spectral features and the original spectral features are input into the target spectral feature stitching module. The target spectral feature stitching module stitches the predicted spectral features and the original spectral features together to obtain the target spectral features.

[0071] Optionally, based on the above embodiments, in some embodiments of the present invention, one implementation of S124 may be: S1241: Determine the first predicted spectral sub-features corresponding to multiple first differential sub-texts in the original spectral features.

[0072] S1242: Determine the second predicted spectral sub-features corresponding to multiple second differential sub-texts in the predicted spectral features.

[0073] S1243: Replace multiple first predicted spectrum sub-features in the original spectrum features with multiple second predicted spectrum sub-features to obtain the target spectrum feature splicing module.

[0074] Specifically, in the original spectral features, the first predicted spectral sub-features corresponding to the multiple first differing sub-texts in the original text are determined. In the predicted spectral features, the second predicted spectral sub-features corresponding to the multiple second differing sub-texts in the text to be synthesized are determined. After obtaining the multiple first predicted spectral sub-features and the multiple second predicted spectral sub-features, the multiple second predicted spectral sub-features are used to replace the multiple first predicted spectral sub-features corresponding to each other in the original spectral features to obtain the target spectral feature concatenation module.

[0075] Thus, this embodiment retains the first non-differentiated sub-text in the original spectral features and only uses the second predicted spectral sub-features corresponding to the multiple second differentiated sub-texts to replace the multiple first predicted spectral sub-features in the original spectral features, thereby ensuring the consistency of the speech editing result with the original speaker's timbre, speech rate and tone style.

[0076] S13: Reconstruct the waveform of the target spectral features using a vocoder to obtain the speech editing result of the text to be synthesized.

[0077] Specifically, the obtained target spectral features are reconstructed using a vocoder to obtain the speech editing result of the text to be synthesized.

[0078] Thus, the speech editing method provided in this embodiment obtains the original spectral features of the text to be synthesized, the original text, and the original audio, where the original text and original spectral features are aligned based on timestamps. Based on the text to be synthesized, the original text, the original spectral features, and the timestamps, the reconstructed mask spectral features corresponding to the text to be synthesized, and the noise corresponding to the reconstructed mask spectral features, are obtained. The reconstructed mask spectral features, noise, the text to be synthesized, and the original spectral features are input into a trained speech editing model to obtain the target spectral features of the text to be synthesized. The speech editing model includes a text feature extraction module, a sound feature extraction module, a spectral feature prediction module, and a target spectral feature concatenation module. The speech editing model is determined by fine-tuning the weight parameters of the sound feature extraction module based on a preset training set. The target spectral features are reconstructed using a vocoder to obtain the speech editing result of the text to be synthesized. In this way, by comprehensively considering the original spectral features of the text to be synthesized, the original text, and the original audio, the reconstructed mask spectral features are obtained, avoiding the problem of low editing accuracy caused by excessive reliance on human experience and judgment in existing technologies. Furthermore, by introducing a voice feature extraction module into the voice editing model, the user's voice features can be accurately obtained, thereby ensuring that the generated voice editing results are highly consistent with the user's timbre, speech rate, and tone style, significantly improving the accuracy of the voice editing results.

[0079] In one embodiment, such as Figure 3 As shown, Figure 3 A schematic diagram of a speech editing device provided in an embodiment of the present invention includes: an acquisition module 10, a reconstructed mask spectral features and noise acquisition module 11, a target spectral features acquisition module 12, and a speech editing result acquisition module 13.

[0080] The acquisition module 10 is used to acquire the original spectral features of the text to be synthesized, the original text, and the original audio, wherein the original text and the original spectral features are aligned based on timestamps.

[0081] The reconstruction mask spectral feature and noise acquisition module 11 is used to acquire the reconstruction mask spectral features corresponding to the text to be synthesized and the noise corresponding to the reconstruction mask spectral features based on the text to be synthesized, the original text, the original spectral features and the timestamp.

[0082] The target spectral feature acquisition module 12 is used to input the reconstructed mask spectral features, noise, text to be synthesized, and original spectral features into the trained speech editing model to obtain the target spectral features of the text to be synthesized. The speech editing model includes a text feature extraction module, a sound feature extraction module, a spectral feature prediction module, and a target spectral feature concatenation module. The speech editing model is determined by fine-tuning the weight parameters of the sound feature extraction module based on a preset training set.

[0083] The speech editing result acquisition module 13 is used to reconstruct the waveform of the target spectral features through a vocoder to obtain the speech editing result of the text to be synthesized.

[0084] In the above embodiments, the acquisition module acquires the original spectral features of the text to be synthesized, the original text, and the original audio, wherein the original text and the original spectral features are aligned based on timestamps. The reconstruction mask spectral feature and noise acquisition module acquires the reconstruction mask spectral features corresponding to the text to be synthesized and the noise corresponding to the reconstruction mask spectral features based on the text to be synthesized, the original text, the original spectral features, and the timestamps. The target spectral feature acquisition module inputs the reconstruction mask spectral features, noise, the text to be synthesized, and the original spectral features into a trained speech editing model to acquire the target spectral features of the text to be synthesized. The speech editing model includes a text feature extraction module, a sound feature extraction module, a spectral feature prediction module, and a target spectral feature concatenation module. The speech editing model is determined by fine-tuning the weight parameters of the sound feature extraction module based on a preset training set. The speech editing result acquisition module reconstructs the waveform of the target spectral features using a vocoder to acquire the speech editing result of the text to be synthesized. Thus, by comprehensively considering the original spectral features of the text to be synthesized, the original text, and the original audio, the reconstruction mask spectral features are acquired, avoiding the problem of low editing accuracy caused by excessive reliance on human experience and judgment in existing technologies. Furthermore, by introducing a voice feature extraction module into the voice editing model, the user's voice features can be accurately obtained, thereby ensuring that the generated voice editing results are highly consistent with the user's timbre, speech rate, and tone style, significantly improving the accuracy of the voice editing results.

[0085] Specific limitations regarding the voice editing device can be found in the limitations of the voice editing method above, and will not be repeated here. Each module in the aforementioned server can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in hardware or independently of the processor in the computer device, or stored in software in the memory of the computer device, so that the processor can call and execute the corresponding operations of each module.

[0086] This invention provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it can implement the voice editing method provided in this invention. For example, when the processor executes the computer program, it can implement... Figure 1 The technical solutions of any of the method embodiments shown are similar in implementation principle and technical effect, and will not be described again here.

[0087] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. Any references to memory, databases, or other media used in the embodiments provided by this invention can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static random access memory (SRAM) and dynamic random access memory (DRAM), etc.

[0088] Although the invention has been described herein in conjunction with various embodiments, those skilled in the art will understand and implement other variations of the disclosed embodiments by reviewing the accompanying drawings, the disclosure, and the appended claims in carrying out the claimed invention. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality. A single processor or other unit can implement several functions listed in the claims. While different dependent claims may recite certain measures, this does not mean that these measures cannot be combined to produce good results.

[0089] Although the invention has been described in conjunction with specific features and embodiments, it is obvious that various modifications and combinations can be made therein without departing from the spirit and scope of the invention. Accordingly, this specification and drawings are merely exemplary descriptions of the invention as defined by the appended claims, and are considered to cover any and all modifications, variations, combinations, or equivalents within the scope of the invention. Clearly, those skilled in the art can make various alterations and modifications to the invention without departing from its spirit and scope. Thus, if such modifications and modifications of the invention fall within the scope of the claims and their equivalents, the invention is also intended to include such modifications and modifications.

Claims

1. A voice editing method, characterized in that, include: Obtain the original spectral features of the text to be synthesized, the original text, and the original audio, wherein the original text and the original spectral features are aligned based on timestamps; Obtain multiple pairs of differing subtexts of the text to be synthesized and the original text, each pair of differing subtexts including: a first differing subtext in the original text and a corresponding second differing subtext in the text to be synthesized; The original text, each of the first difference sub-texts and each of the corresponding second difference sub-texts are segmented into words to obtain multiple first words corresponding to the original text, multiple second words corresponding to each of the first difference sub-texts and multiple third words corresponding to each of the second difference sub-texts. An upsampling operation is performed on the plurality of first word segments, the plurality of second word segments, the plurality of third word segments, and the original spectral features to obtain the target spectral feature length corresponding to the text to be synthesized; Based on the target spectral feature length, the original spectral feature length, and the original spectral feature, a reconstruction mask operation is performed based on the timestamp to obtain the reconstructed mask spectral feature, wherein the length of the reconstructed mask spectral feature is the same as the target spectral feature length. Based on the reconstructed mask spectral features, noise is obtained, wherein the length of the noise is the same as the length of the target spectral features; The reconstructed mask spectral features, the noise, the text to be synthesized, and the original spectral features are input into a trained speech editing model to obtain the target spectral features of the text to be synthesized. The speech editing model includes a text feature extraction module, a sound feature extraction module, a spectral feature prediction module, and a target spectral feature concatenation module. The speech editing model is determined by fine-tuning the weight parameters of the sound feature extraction module based on a preset training set. The target spectral features are reconstructed using a vocoder to obtain the speech editing result of the text to be synthesized.

2. The method according to claim 1, characterized in that, Before obtaining the reconstructed mask spectral features corresponding to the text to be synthesized and the noise corresponding to the reconstructed mask spectral features based on the text to be synthesized, the original text, the original spectral features, and the timestamp, the method further includes: The original text and the original spectral features are input into an automatic speech recognition model, and the automatic speech recognition model aligns the original text and the original spectral features based on the timestamp.

3. The method according to claim 1, characterized in that, The step of performing a reconstruction mask operation based on the timestamp, according to the target spectral feature length, the original spectral feature length, and the original spectral feature, to obtain the reconstructed mask spectral feature includes: Determine whether the length of the target spectral feature is greater than the length of the original spectral feature; If the target spectral feature length is greater than the original spectral feature length, then a padding and reconstruction mask operation is performed on the original spectral feature based on the timestamp to obtain the reconstructed mask spectral feature.

4. The method according to claim 3, characterized in that, The method further includes: If the length of the target spectral feature is not greater than the length of the original spectral feature, then the original spectral feature is cropped and reconstructed using the timestamp to obtain the reconstructed mask spectral feature.

5. The method according to claim 4, characterized in that, The step of inputting the reconstructed mask spectral features, the noise, the text to be synthesized, and the original spectral features into the trained speech editing model to obtain the target spectral features of the text to be synthesized includes: The text to be synthesized is input into the text feature extraction module, and the text features of the text to be synthesized are obtained through the text feature extraction module. The original spectral features are input into the sound feature extraction module, and sound features are obtained through the sound feature extraction module; The text features, the reconstructed mask spectral features, the noise, and the sound features are input into the spectral feature prediction module, and the predicted spectral features corresponding to the text to be synthesized are obtained through the spectral feature prediction module. The predicted spectral features and the original spectral features are input into the target spectral feature stitching module, and the target spectral features are obtained through the target spectral feature stitching module.

6. The method according to claim 5, characterized in that, The sound feature extraction module includes an encoding layer and an upsampling layer. The original spectral features are input into the sound feature extraction module, and sound features are obtained through the sound feature extraction module, including: The original spectral features are input into the coding layer for encoding processing to obtain coded features; The encoded features are input into the upsampling layer for upsampling processing to obtain the sound features.

7. The method according to claim 5, characterized in that, The step of inputting the predicted spectral features and the original spectral features into the target spectral feature stitching module, and obtaining the target spectral features through the target spectral feature stitching module, includes: In the original spectral features, first predicted spectral sub-features corresponding to multiple first differential sub-texts are determined respectively; In the predicted spectral features, determine the second predicted spectral sub-features corresponding to multiple second differential sub-texts respectively; The target spectrum feature splicing module is obtained by replacing multiple first predicted spectrum sub-features in the original spectrum features with multiple second predicted spectrum sub-features.

8. A voice editing device, characterized in that, include: The acquisition module is used to acquire the original spectral features of the text to be synthesized, the original text, and the original audio, wherein the original text and the original spectral features are aligned based on timestamps; The reconstruction mask spectral features and noise acquisition module is used to acquire multiple pairs of differing subtexts of the text to be synthesized and the original text, each pair of differing subtexts including: a first differing subtext in the original text and a corresponding second differing subtext in the text to be synthesized; The original text, each of the first difference sub-texts and each of the corresponding second difference sub-texts are segmented into words to obtain multiple first words corresponding to the original text, multiple second words corresponding to each of the first difference sub-texts and multiple third words corresponding to each of the second difference sub-texts. An upsampling operation is performed on the plurality of first word segments, the plurality of second word segments, the plurality of third word segments, and the original spectral features to obtain the target spectral feature length corresponding to the text to be synthesized; Based on the target spectral feature length, the original spectral feature length, and the original spectral feature, a reconstruction mask operation is performed based on the timestamp to obtain the reconstructed mask spectral feature, wherein the length of the reconstructed mask spectral feature is the same as the target spectral feature length. Based on the reconstructed mask spectral features, noise is obtained, wherein the length of the noise is the same as the length of the target spectral features; The target spectral feature acquisition module is used to input the reconstructed mask spectral features, the noise, the text to be synthesized, and the original spectral features into a trained speech editing model to obtain the target spectral features of the text to be synthesized. The speech editing model includes a text feature extraction module, a sound feature extraction module, a spectral feature prediction module, and a target spectral feature concatenation module. The speech editing model is determined by fine-tuning the weight parameters of the sound feature extraction module based on a preset training set. The speech editing result acquisition module is used to reconstruct the waveform of the target spectral features through a vocoder to obtain the speech editing result of the text to be synthesized.

9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the speech editing method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Text-based voice editing method and system, electronic equipment and storage medium

    CN115966196A