Speech synthesis model training method, speech synthesis method, electronic device, and storage medium
By segmented detection and splicing text and audio data, and using context to train the speech synthesis model, the problem of speech inconsistent with the situation in the existing model is solved, and the smoothness and emotional compliance of speech synthesis are achieved, which is suitable for a variety of existing model frameworks.
Patent Information
- Application Number
- PCT/CN2024/141147
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-29
- Filing Date
- 2024-12-20
- Publication Date
- 2025-07-03
AI Technical Summary
The existing speech synthesis model cannot effectively capture the context information of the text, resulting in the synthetic speech inconsistent with the situation, the speech connection is not smooth, and the additional prompt word processing module introduced increases the difficulty and cost of data labeling.
By obtaining coherent text and audio data, segmented text and audio bands that meet the conditions, the context trained speech synthesis model, extract rhythm, rhythm and emotional characteristics, and generate speech that conforms to the situation.
The generated voice is more in line with the situation, the voice connection is smooth, which improves training efficiency and reduces data labeling costs, and is suitable for a variety of existing model frameworks.
Smart Images

Figure CN2024141147_03072025_PF_FP_ABST
Abstract
Description
Speech synthesis model training method, speech synthesis method, electronic device and storage medium
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This disclosure claims priority to Chinese patent application number 2023118701147, filed with the Chinese Patent Office on December 29, 2023, entitled “Speech Synthesis Model Training Method, Speech Synthesis Method, Electronic Device and Storage Medium,” the entire contents of which are incorporated by reference into this disclosure. Technical Field
[0003] The present disclosure relates to the technical field of multimedia content processing, and in particular to a speech synthesis model training method, a speech synthesis method, a speech synthesis model training device, a speech synthesis device, an electronic device, and a storage medium. Background Art
[0004] Speech synthesis technology converts text into speech, and it can be achieved through neural network models. Current text-to-speech models only convert the input text into speech. This makes it difficult for the generated speech to align with the semantics, rhythm, and emotion of the overall text. This can cause the speaker's voice to be inconsistent with the context and lead to choppy pronunciation.
[0005] The description of the background technology is intended to help understand the relevant technology in the relevant field, and does not mean that the background technology content is admitted to be prior art. Summary of the Invention
[0006] Therefore, the present disclosure provides a speech synthesis model training method, a speech synthesis method, and related electronic devices and storage media. Through the solutions of the present disclosure, context-based text understanding can be used to capture the rhythm, rhythm, and emotion of the current scene, making the generated speech more realistic.
[0007] In a first aspect, the present disclosure provides a method for training a speech synthesis model, comprising the following steps:
[0008] Acquiring initial training data, wherein the initial training data includes coherent text and corresponding coherent audio;
[0009] selecting, from the continuous text, a plurality of first text segments, a plurality of second text segments located before and adjacent to the first text segments, and a plurality of third text segments located after and adjacent to the first text segments, and acquiring, from the continuous audio, a plurality of first audio segments, a second audio segment, and a third audio segment corresponding to the plurality of first text segments, the second text segments, and the third text segments, respectively;
[0010] detecting whether the associated first text segment, second text segment, third text segment, first audio segment, second audio segment, and third audio segment meet a splicing condition;
[0011] splicing the first text segment, the second text segment, the third text segment, the first audio segment, the second audio segment, and the third audio segment that meet the splicing condition in a text order in an alternating manner of text and audio to obtain combined training data, thereby generating a combined training data set including a plurality of the combined training data;
[0012] An initial speech synthesis model is trained based on the combined training data set to obtain a trained speech synthesis model, during which the prosody, rhythm and / or emotional features in the second text segment and / or the third text segment are extracted.
[0013] In some embodiments of the present disclosure, detecting whether the associated first text segment, second text segment, third text segment, first audio segment, second audio segment, and third audio segment meet a splicing condition includes:
[0014] Obtaining a combined text length of the first text segment, the second text segment, and the third text segment, and obtaining a combined audio duration of the first audio segment, the second audio segment, and the third audio segment. When the combined text length is less than a length threshold and the combined audio duration is less than a duration threshold, the text and speech corresponding to the first text segment, the second text segment, the third text segment, the first audio segment, the second audio segment, and the third audio segment meet a first splicing condition.
[0015] In some embodiments of the present disclosure, detecting whether the associated first text segment, second text segment, third text segment, first audio segment, second audio segment, and third audio segment meet a splicing condition includes:
[0016] Obtain voiceprint feature similarity between the first audio segment, the second audio segment, and the third audio segment. When the voiceprint feature similarity is greater than a voiceprint similarity threshold, the text and speech corresponding to the first text segment, the second text segment, the third text segment, the first audio segment, the second audio segment, and the third audio segment meet a second splicing condition.
[0017] In some embodiments of the present disclosure, detecting whether the associated first text segment, second text segment, third text segment, first audio segment, second audio segment, and third audio segment meet a splicing condition includes:
[0018] When the interval between utterance time points of adjacent audio segments in the first audio segment, the second audio segment, and the third audio segment is less than an interval threshold, the text and speech corresponding to the first text segment, the second text segment, the third text segment, the first audio segment, the second audio segment, and the third audio segment meet a third splicing condition.
[0019] In some embodiments of the present disclosure, the combined training data is obtained by splicing text and audio in an alternating manner according to the text order, including:
[0020] Performing word segmentation processing on the first text segment, the second text segment, and the third text segment to generate a first text sequence, a second text sequence, and a third text sequence; performing vector conversion on the first text sequence, the second text sequence, and the third text sequence to generate a first text vector, a second text vector, and a third text vector;
[0021] Discretizing the first, second, and third audio segments and extracting audio features to generate a first, second, and third audio feature sequences; and performing vector conversion on the first, second, and third audio feature sequences to generate a first, second, and third audio vector.
[0022] The second text vector, the second audio feature sequence, the first text vector, the first audio feature sequence, the third text vector, and the third audio feature sequence are vector-concatenated in the order of the second text vector, the second audio feature sequence, the first text vector, the first audio feature sequence, the third text vector, and the third audio feature sequence to generate the combined training data.
[0023] In some embodiments of the present disclosure, the initial speech synthesis model is trained based on the combined training data set to obtain a trained speech synthesis model, including:
[0024] Iteratively input the combined training data in the combined training data set into the initial speech synthesis model, obtain a first output audio segment corresponding to the first text segment, calculate a first loss value based on the first output audio segment and the audio features of the first audio segment, adjust the model parameters of the initial speech synthesis model based on the first loss value until the initial speech synthesis model converges, and use the converged initial speech synthesis model as the trained speech synthesis model.
[0025] In some embodiments of the present disclosure, the initial speech synthesis model is trained based on the combined training data set to obtain a trained speech synthesis model, including:
[0026] Inputting the combined training data in the combined training data set into the initial speech synthesis model, obtaining a second output audio segment corresponding to the second text segment, calculating a second loss value based on audio features of the second output audio segment and the second audio segment, adjusting model parameters of the initial speech synthesis model based on the second loss value, iteratively updating until the initial speech synthesis model converges, and using the converged initial speech synthesis model as the trained speech synthesis model;
[0027] In some embodiments of the present disclosure, the initial speech synthesis model is trained based on the combined training data set to obtain a trained speech synthesis model, including:
[0028] The combined training data in the combined training data set is iteratively input into the initial speech synthesis model to obtain a third output audio segment corresponding to the third text segment, and a third loss value is calculated based on the third output audio segment and the audio features of the third audio segment. The model parameters of the initial speech synthesis model are adjusted according to the third loss value until the initial speech synthesis model converges, and the converged initial speech synthesis model is used as the trained speech synthesis model.
[0029] In some embodiments of the present disclosure, the initial speech synthesis model is trained based on the combined training data set to obtain a trained speech synthesis model, including:
[0030] The combined training data in the combined training data set is iteratively input into the initial speech synthesis model to obtain a first output audio segment corresponding to the first text segment, a second output audio segment corresponding to the second text segment, and a third output audio segment corresponding to the third text segment; a first loss value is calculated based on the audio features of the first output audio segment and the first audio segment, a second loss value is calculated based on the audio features of the second output audio segment and the second audio segment, a third loss value is calculated based on the audio features of the third output audio segment and the third audio segment, and a fourth loss value is obtained based on the first loss value, the second loss value, and the third loss value; model parameters of the initial speech synthesis model are adjusted according to the fourth loss value until the initial speech synthesis model converges, and the converged initial speech synthesis model is used as the trained speech synthesis model.
[0031] In a second aspect, the present disclosure provides a speech synthesis method, which may include:
[0032] Acquire a target text, and divide the target text into a plurality of target text segments according to the text order;
[0033] Inputting a first target text segment from the plurality of target text segments into a trained target speech synthesis model to obtain a first target audio segment;
[0034] splicing the first target text segment and the first target audio segment to obtain a text-audio alternating sequence;
[0035] The following steps are executed in a loop until a last target audio segment corresponding to a last target text segment is obtained: obtaining a next target text segment, adding the next target text segment to the text-audio alternating sequence to update the text-audio alternating sequence; inputting the text-audio alternating sequence into the trained speech synthesis model to obtain a next target audio segment corresponding to the next target text segment and extracted from the prosody, rhythm, and / or emotional acoustic features of other target text segments, and adding the next target audio segment to the text-audio alternating sequence to update the text-audio alternating sequence;
[0036] The target audio segments are spliced together to generate the target audio.
[0037] In some other embodiments of the present disclosure, another feasible speech synthesis method is provided, which may include:
[0038] Acquire a target text, and divide the target text into a plurality of target text segments according to the text order;
[0039] Inputting a last target text segment among the multiple target text segments into a trained target speech synthesis model to obtain a last target audio segment;
[0040] splicing the last target text segment and the last target audio segment to obtain a text-audio alternating sequence;
[0041] The following steps are executed in a loop until a first target audio segment corresponding to a first target text segment is obtained: obtaining a previous target text segment, adding the previous target text segment to the text-audio alternating sequence to update the text-audio alternating sequence; inputting the text-audio alternating sequence into the trained target speech synthesis model to obtain a previous target audio segment corresponding to the previous target text segment and having prosody, rhythm, and / or emotional acoustic features of other target text segments extracted, and adding the previous target audio segment to the end of the text-audio alternating sequence to update the text-audio alternating sequence;
[0042] Each target audio segment is spliced together to generate the target audio.
[0043] In a third aspect, an embodiment of the present disclosure provides a speech synthesis model, including an encoder and a decoder, wherein the speech synthesis model is trained by the training method in any embodiment of the present disclosure.
[0044] In a fourth aspect, the present disclosure provides a speech synthesis model training device, comprising a training set data acquisition module, a preprocessing module, a splicing detection module, a vector splicing module, and a training module, wherein:
[0045] The training set data acquisition module is configured to acquire training set data, wherein the training set data includes coherent text and corresponding coherent audio;
[0046] The preprocessing module is configured to select, from the continuous text, a plurality of first text segments, a plurality of second text segments located before and adjacent to the first text segments, and a plurality of third text segments located after and adjacent to the first text segments, and obtain, from the continuous audio, a plurality of first audio segments, a second audio segment, and a third audio segment corresponding to the plurality of first text segments, the second text segments, and the third text segments, respectively;
[0047] The splicing detection module is configured to detect whether the texts and voices corresponding to the plurality of first text segments, second text segments, third text segments, first audio segments, second audio segments, and third audio segments meet a splicing condition;
[0048] The vector splicing module is configured to splice a plurality of first text segments, second text segments, third text segments, first audio segments, second audio segments, and third audio segments that meet the splicing condition in a text order in an alternating manner of text and audio to generate a combined training data set, wherein the combined training data set includes a plurality of combined training data;
[0049] The training module is configured to train the initial speech synthesis model according to the combined training data set to obtain a trained speech synthesis model.
[0050] In a fifth aspect, an embodiment of the present disclosure provides a speech synthesis device, comprising a target text acquisition module, an initial generation module, a text-audio splicing module, a loop generation module and an audio splicing module, wherein:
[0051] The target text acquisition module is configured to acquire a target text and divide the target text into a plurality of target text segments according to the text order;
[0052] The initial generation module is configured to obtain a first target text segment from the multiple target text segments, input a vector corresponding to the first target text segment into a trained target speech synthesis model, and obtain a first target audio segment;
[0053] The text-audio splicing module is configured to generate a text-audio alternating sequence based on the splicing of the first target text segment and the first target audio segment;
[0054] The loop generation module is configured to cyclically obtain a next target text segment, add the next target text segment to the end of the text-audio alternating sequence to update the text-audio alternating sequence; input a vector corresponding to the text-audio alternating sequence into the trained target speech synthesis model, obtain a next target audio segment, and add the next target audio segment to the end of the text-audio alternating sequence to update the text-audio alternating sequence, until a last target audio segment corresponding to the last target text segment is obtained;
[0055] The audio splicing module is configured to splice the target audio segments to generate target audio.
[0056] In a sixth aspect, an embodiment of the present disclosure provides a computer-readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the method of any embodiment of the present disclosure is implemented.
[0057] In a seventh aspect, an embodiment of the present disclosure provides an electronic device comprising: a processor and a memory storing a computer program, wherein the processor is configured to execute any method of the embodiment of the present disclosure when running the computer program.
[0058] This technical solution provides a speech synthesis model training method, which forms a sample group of the previous and next paragraphs of the training text and the corresponding training speech data. After determining whether they can be spliced together, the sample group is converted into a vector and input into the initial speech synthesis model. The initial speech synthesis model is trained to generate a trained speech synthesis model. During the speech synthesis process, the generated trained speech synthesis model can use the context of the text to understand the text and capture the rhythm and emotion in the scene. At the same time, it can also understand the speech corresponding to the context and obtain the melody and emotion of the scene. Therefore, the generated speech information finally outputted is more in line with the situation and can be smoothly connected with the speech in the context. The shorter audio in the training data is spliced together to lengthen it, thereby improving training efficiency and making full use of data. This technical solution also provides a speech recognition method, which divides the target text into multiple segments, first uses a trained speech synthesis model to perform speech synthesis on the first segment of text, and then combines the first generated audio segment with the first segment of text to form a text-audio sequence. For subsequent text segments, speech synthesis is iteratively performed, and the current text segment to be recognized is added to the text-audio sequence to obtain the corresponding audio segment. The newly obtained audio segment is also added to the text-audio sequence to provide context text and audio for the next text segment, until the last text segment generates an audio segment, and then the generated audio segments are spliced together to finally generate the target audio. In some embodiments of the present disclosure, speech synthesis can also be started from the last segment, and then the generated audio segments are spliced together until the first text segment generates an audio segment to generate the target audio, and then the generated audio segments are spliced together to finally generate the target audio. In the speech synthesis process of the embodiments of the present disclosure, contextual text understanding and speech understanding are utilized, and the generated audio can be accompanied by the rhythm, melody and emotion of the current scene, so that the speech synthesis is more in line with the situation and enhances the realism of the character's speech or dialogue. The solution in the embodiment of the present disclosure can use the existing speech synthesis model framework to perform speech model training and speech synthesis. The technology is mature and reliable, easy to operate, and has high practical application value.
[0059] Other optional features and technical effects of the embodiments of the present disclosure are partially described below, and partially can be understood by reading this document. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use. It should be noted that the drawings described below only cover some embodiments of the present disclosure. A person skilled in the art can derive other drawings based on these drawings without having to engage in creative work. The purpose of these drawings is to better illustrate the technical details to facilitate understanding of the embodiments of the present disclosure, including:
[0061] FIG1 is a schematic diagram showing the data flow of the speech synthesis model training method according to an embodiment of the present disclosure;
[0062] FIG2 is a schematic flow chart showing a method for training a speech synthesis model according to an embodiment of the present disclosure;
[0063] FIG3 shows a schematic diagram of dividing text segments and audio segments in the speech synthesis model training method according to an embodiment of the present disclosure;
[0064] FIG4 is a schematic diagram showing the process of vectorizing text segments and audio segments in the speech synthesis model training method according to an embodiment of the present disclosure;
[0065] FIG5 is a schematic diagram showing a flow chart of model training in the speech synthesis model training method according to an embodiment of the present disclosure;
[0066] FIG6 shows another schematic diagram of the model training process in the speech synthesis model training method according to an embodiment of the present disclosure;
[0067] FIG7 shows another schematic diagram of the model training process in the speech synthesis model training method according to an embodiment of the present disclosure;
[0068] FIG8 shows another schematic diagram of a process of model training in the speech synthesis model training method according to an embodiment of the present disclosure;
[0069] FIG9 is a schematic diagram showing a workflow of a speech synthesis model obtained by training a speech synthesis model using the speech synthesis model training method according to an embodiment of the present disclosure;
[0070] FIG10 is a schematic flow chart showing a speech synthesis method according to an embodiment of the present disclosure;
[0071] FIG11 is a schematic flow chart showing a method for cyclically generating target audio in a speech synthesis method according to an embodiment of the present disclosure;
[0072] FIG12 shows an exemplary structural diagram of a speech synthesis model training device according to an embodiment of the present disclosure;
[0073] FIG13 shows an exemplary structural diagram of a speech synthesis device according to an embodiment of the present disclosure;
[0074] FIG14 shows a schematic diagram of an exemplary structure of an electronic device capable of implementing the method according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0075] The following detailed description of exemplary embodiments of the present disclosure is provided, and illustrations of the exemplary embodiments are shown in the accompanying drawings. When referring to the drawings, unless otherwise specified, identical numbers or symbols in different figures represent identical or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with one or more embodiments of this specification. Instead, they are merely examples of some aspects of the apparatus and methods covered by one or more embodiments of this specification, as detailed in the claims of this disclosure.
[0076] As used in this specification, the term "including" and its variations are intended to be broadly inclusive, meaning "including but not limited to" the listed items. Unless otherwise stated, the term "or" means "and / or," the term "based on" means relying on, or at least partially relying on, the terms "an example embodiment" and "an embodiment" refer to at least one example embodiment, and the term "another embodiment" refers to at least one different embodiment. The terms "first," "second," and so on may refer to different or the same items. Other explicit and implicit definitions may be included below.
[0077] In the embodiments of the present disclosure, the speech synthesis process is to generate speech audio corresponding to the target text segment based on the input target text segment and the text content of the context and the audio content of the context. It can be implemented through a speech synthesis (Text to Speech, TTS) model or a large language model.
[0078] In the embodiments of the present disclosure, "prosody" is one of the attributes of speech, including the pitch, rhythm, etc. "Semantics" is one of the attributes of speech, identifying the content of speech.
[0079] In the embodiments of the present disclosure, "model" has the conventional meaning in the field of machine learning. For example, the model can be a machine learning or deep learning model, such as a machine learning or deep learning model that includes the above-mentioned network or is composed of the above-mentioned network.
[0080] In the embodiments of the present disclosure, “loss function” and “loss value” have conventional meanings in the field of machine learning.
[0081] In the disclosed embodiments, "embedding" refers to the step of data processing, which converts discrete text and other data into continuous data vectors so that the model can perform calculations on them.
[0082] In the embodiments of the present disclosure, a "phoneme" is the smallest unit of speech divided according to the natural properties of speech, and one pronunciation action forms a phoneme. A "phoneme sequence" is a sequence composed of several phonemes.
[0083] The embodiments of the present disclosure provide a speech synthesis model training method, a speech synthesis method, a speech synthesis model, a training device, a synthesis device, a storage medium, and an electronic device. The method, device / model can be implemented with the aid of one or more computers. In some embodiments, the device / model can be implemented by software, hardware, or a combination of software and hardware. In some embodiments, the electronic device or computer can be implemented by the computer described herein or other electronic devices that can perform the corresponding functions.
[0084] The inventors of the present disclosure have found that the speech of normal people is affected by many factors, not only by the text to be expressed, but also by the situation and context. However, in the current TTS model, the model cannot obtain the contextual information of the speech segment to be synthesized, resulting in the synthesized speech not being consistent with the situation, causing the connection with the context to be not smooth. At the same time, the model cannot control the tone and other attributes of the speech according to demand during synthesis. Some improvement schemes propose to introduce additional inputs in addition to text into the input content of the speech synthesis process, such as style descriptions corresponding to the text, and a prompt word processing module is added to some TTS models. However, the inventors found that this speech synthesis method requires the creation of a special training data set, in which each data needs to contain audio, text and style descriptions. The data annotation is difficult, and it will result in additional costs in actual application and promotion, making it difficult to expand the data scale to improve the effect. At the same time, the newly added prompt word processing module requires additional fine-tuning of a BERT model, resulting in a longer link and increased errors.
[0085] In this regard, the solution of the embodiments of the present disclosure provides a speech synthesis model training method and device, a speech synthesis method and device, a speech synthesis model, a storage medium and an electronic device, and proposes a method for speech synthesis using context, which controls the rhythm of the speech through the context, thereby controlling the emotion of the speech, and also makes the generated speech more coherent.
[0086] As shown in FIG1 , the model training process in the embodiment of the present disclosure is to select a coherent text and corresponding audio from the training set, and divide it into three parts: the preceding text, the current text, and the following text according to the part to be synthesized. The text of each part is converted into a digital sequence through word segmentation and then converted into training data that can be input into the model using embedding. Similarly, the corresponding audio data is discretized and converted into training data through embedding. The three parts of data are tested to see whether they meet the conditions for splicing: (1) the total length cannot be too long; (2) each part needs to come from the same speaker; (3) the position of each part in the original audio cannot be too far apart and must be coherent. If at least one of the above splicing conditions is met, the text and audio of the three parts are spliced together and separated by special symbols; then the spliced training data is input into the model for training to obtain the corresponding speech synthesis model. In the embodiment of the present disclosure, if there is noise in each audio segment, the noise can be retained as the ambient sound of the input model.
[0087] In the embodiment shown in FIG2 , a speech synthesis model training method is provided, and a trained speech synthesis model is obtained through training.
[0088] As shown in FIG2 , the training method in the embodiment of the present disclosure includes the following steps.
[0089] S110: Acquire initial training data, wherein the initial training data includes coherent text and corresponding coherent audio.
[0090] In the embodiment of the present disclosure, the initial training data may be obtained from conventional training data, which may include a plurality of coherent texts and a plurality of coherent audios corresponding thereto, without the need for additional annotation of the training data.
[0091] S120: Selecting, from the continuous text, a plurality of first text segments, a plurality of second text segments located before and adjacent to the first text segments, and a plurality of third text segments located after and adjacent to the first text segments, and obtaining, from the continuous audio, a plurality of first audio segments, second audio segments, and third audio segments corresponding to the plurality of first text segments, second text segments, and third text segments, respectively.
[0092] In the embodiment of the present disclosure, contextual text data and voice data are selected and input into the model for training, thereby improving the scene understanding ability of the speech synthesis model.
[0093] In some embodiments of the present disclosure, as shown in FIG3 , the present disclosure selects a first text, a second text preceding the first text, and a third text following the first text from a coherent text, and correspondingly extracts a first audio, a second audio, and a third audio from the coherent speech.
[0094] In the embodiment of the present disclosure, the audio segment end line is divided based on the end of the text pronunciation in the corresponding text segment.
[0095] S130: Detecting whether the associated first text segment, second text segment, third text segment, first audio segment, second audio segment, and third audio segment meet a splicing condition.
[0096] In the embodiment of the present disclosure, in order to ensure the continuity of the spliced data, the embodiment of the present disclosure performs splicing detection on the selected continuous multiple segments of text and audio.
[0097] In the embodiment of the present disclosure, when the data are spliced and input into the model, the data can be converted into a continuous data vector by "embedding" the discrete data so that the model can calculate it.
[0098] In some embodiments of the present disclosure, for example, in step S130 above, it is possible to detect whether three text segments or three audio segments meet a first splicing condition. The first splicing condition may be a threshold for the overall length of the three text segments or three audio segments. Specifically, the combined text length of the first, second, and third text segments may be obtained, and the combined audio duration of the first, second, and third audio segments may be obtained. When the combined text length is less than the length threshold and the combined audio duration is less than the duration threshold, the text and speech corresponding to the first, second, and third text segments, the first, second, and third audio segments meet the first splicing condition.
[0099] In a specific embodiment of the present disclosure, with reference to Figure 3, the length threshold can be set to 50 words and the duration threshold to 30 seconds. The second text segment is "I went to the riverbank on holiday", the text length is 8 words, and the duration of the corresponding second audio is 3.5 seconds. The first text segment is "I saw many children flying kites", the text length is 10 words, and the duration of the corresponding first audio is 6.1 seconds. The third text segment is "Long fuses", the text length is 8 words, and the duration of the corresponding third audio is 5.0 seconds. The total text length is 8+10+8=26 words, which is less than the length threshold (50 words). The total sample duration is 3.5+6.1+5.0=14.6 seconds, which is less than the duration threshold (30 seconds). The text and speech corresponding to the first text segment, the second text segment, the third text segment, the first audio segment, the second audio segment, and the third audio segment meet the first splicing condition.
[0100] In other embodiments of the present disclosure, to ensure the coherence of the model's output speech, the audio input to the training model must be spoken by the same person. For example, voiceprint features can be used to determine whether the audio is spoken by the same person. In some embodiments of the present disclosure, for example, in step S130 above, the voiceprint feature similarity between the first audio segment, the second audio segment, and the third audio segment can be obtained. When the voiceprint feature similarity is greater than a voiceprint similarity threshold, the text and speech corresponding to the first text segment, the second text segment, the third text segment, the first audio segment, the second audio segment, and the third audio segment meet the second splicing condition.
[0101] In some embodiments of the present disclosure, after discrete sampling of the first, second, and third audio segments, the time domain signals can be converted into frequency domain signals, and frequency features can be extracted. Based on the similarity of the frequency features, it can be determined whether the voices belong to the same person. In the embodiments of the present disclosure, a voiceprint similarity threshold is set, for example, 80%. When the voiceprint feature similarity between any two audio segments is greater than 80%, the first, second, and third audio segments are considered to be the voices of the same person. The text and speech corresponding to the first, second, and third text segments, the first, second, and third audio segments meet the second splicing condition.
[0102] In some embodiments of the present disclosure, in order to determine whether multiple audio segments are emitted by the same person, identification can also be based on established voiceprint models (such as Gaussian mixture models (GMM)) or deep learning models (such as CNN and RNN).
[0103] In some embodiments of the present disclosure, for example, in the above-mentioned step S130, in order to improve the coherence of the audio output by the speech synthesis model and avoid the selected training audio segments being too far apart and lacking coherence characteristics, the embodiments of the present disclosure also detect the intervals between the audio segments when detecting the splicing conditions. When the distance between the utterance time points of adjacent audio segments in the first audio segment, the second audio segment, and the third audio segment is less than the interval threshold, the text and speech corresponding to the first text segment, the second text segment, the third text segment, the first audio segment, the second audio segment, and the third audio segment meet the third splicing condition.
[0104] During the training sample generation process, when the speaker reads the text continuously, there will generally be a short pause after the sentence break. In order to ensure the listener's coherent experience, the time between the sentence breaks should not be too long, otherwise the listener experience will be poor. In the embodiment of the present disclosure, the distance between the utterance time points of adjacent audio segments can be understood as the interval from the end of the utterance of the previous audio segment to the start of the utterance of the next audio segment. In the embodiment of the present disclosure, based on the pause after the sentence break, there is often a blank sound period before each audio segment starts to speak. The shaded part in Figure 3 is the blank sound period, and the duration of the blank sound period is the distance between the utterance time points of adjacent audio segments. In some embodiments of the present disclosure, the interval threshold can be selected as 2 to 4 seconds.
[0105] In a specific embodiment of the present disclosure, referring to FIG3 , the interval threshold can be set to 3 seconds. The second text segment is “Walk on the riverbank on holiday”, the text length is 8 words, the corresponding second audio duration is 3.5 seconds, and there is no blank sound period; the first text segment is “See many children flying kites”, the text length is 10 words, the corresponding first audio duration is 6.1 seconds, and the blank sound period is 0.8 seconds; the third text segment is “One by one long fuse”, the text length is 8 words, the corresponding third audio duration is 5 seconds, and the blank sound period is 1.4 seconds. The time interval between the utterance time points of the second audio segment and the first audio segment is 0.8 seconds, and the time interval between the utterance time points of the first audio segment and the third audio segment is 1.4 seconds, both of which are less than the interval threshold. Therefore, the text and speech corresponding to the first text segment, the second text segment, the third text segment, the first audio segment, the second audio segment, and the third audio segment meet the third splicing condition.
[0106] In the embodiment of the present disclosure, when detecting whether the splicing conditions are met, the above detection conditions may be used alone or in combination.
[0107] S140: Splicing the first text segment, the second text segment, the third text segment, the first audio segment, the second audio segment, and the third audio segment that meet the splicing conditions in a text order in an alternating manner of text and audio to obtain combined training data, thereby generating a combined training dataset including a plurality of the combined training data.
[0108] In an embodiment of the present disclosure, when combined training data is obtained by splicing text and audio alternately, the text and audio data can be converted into continuous text and audio data vectors through "embedding" right discrete data conversion.
[0109] In some embodiments of the present disclosure, as shown in FIG4 , the process of obtaining combined training data by splicing text and audio in an alternating manner according to the text order includes:
[0110] S141: Perform word segmentation processing on the first text segment, the second text segment, and the third text segment to generate a first text sequence, a second text sequence, and a third text sequence; perform vector conversion on the first text sequence, the second text sequence, and the third text sequence to generate a first text vector, a second text vector, and a third text vector.
[0111] In the embodiments of the present disclosure, in order to perform word segmentation on a text segment to generate a text sequence, a word segmentation algorithm can be used for processing, such as a method based on string matching or a rule-based method (such as forward maximum matching word segmentation algorithm, reverse maximum matching word segmentation algorithm, statistical method (such as Jieba segmentation, etc.)) for word segmentation. In some embodiments of the present disclosure, before the word segmentation, the text in the text segment can also be subjected to text normalization processing, including converting the character set (such as uniformly using Unicode encoding), eliminating case differences, and removing irrelevant characters (such as HTML tags).
[0112] In some embodiments of the present disclosure, the vector conversion process of the sequence can be processed by using embedding, and the embedding processing process can include:
[0113] a. Construct a vocabulary
[0114] Word frequency statistics: Calculate the frequency of each token appearing in the entire dataset.
[0115] Vocabulary screening: Screen out the words in the vocabulary based on frequency and relevance. Remove high-frequency stop words (such as "de", "he") and low-frequency words.
[0116] Index assignment: Assign a unique numerical index to each token in the vocabulary.
[0117] b. Token encoding
[0118] One-Hot encoding: Convert each token into a high-dimensional sparse vector, where only the index position corresponding to the token is 1 and the rest are 0.
[0119] Limit the vector size: To avoid too high dimensions, which may limit the size of the vocabulary, use special tokens to replace uncommon words, such as ` <unk>`(Unknown).
[0120] c. Learning Dense Vector Representation
[0121] Choose a model: Use pre-trained models like Word2Vec, GloVe, FastText, or BERT, or train a custom model for your specific task.
[0122] Training process: Learn the vector representation of a token through its context. For example, in Word2Vec, this may be achieved by predicting the context of a word (CBOW model), or by using a word to predict its context (Skip-gram model).
[0123] Vector features: These vectors capture complex relationships between words, such as semantic similarity, grammatical rules, etc.
[0124] d. Vector space optimization
[0125] Dimensionality reduction: Dimensionality reduction is performed using techniques such as principal component analysis (PCA) or t-SNE to optimize the vector space so that similar words are closer in the vector space.
[0126] Refinement training: Further fine-tuning is performed on a specific task to optimize the performance of the word vector in that task.
[0127] S142: Discretize the first audio segment, the second audio segment, and the third audio segment and extract audio features to generate a first audio feature sequence, a second audio feature sequence, and a third audio feature sequence; perform vector conversion on the first audio feature sequence, the second audio feature sequence, and the third audio feature sequence to generate a first audio vector, a second audio vector, and a third audio vector.
[0128] In embodiments of the present disclosure, the audio segments may be discretized. Specifically, the audio segments may be sampled, frequency features may be extracted, and spectral features may be processed to generate an audio feature sequence. The audio features may include Mel-spectrum spectra, Mel-spectrum cepstral coefficients, etc. In some embodiments of the present disclosure, such as in step S142 above, embedding may be used to discretize the first, second, and third audio segments and extract audio features.
[0129] S143: Perform vector concatenation on the second text vector, the second audio feature sequence, the first text vector, the first audio feature sequence, the third text vector, and the third audio feature sequence in the order of the second text vector, the second audio feature sequence, the first text vector, the first audio feature sequence, the third text vector, and the third audio feature sequence to generate the combined training data.
[0130] In one embodiment of the present disclosure, for example, the first text vector is T1, the second text vector is T2, the third text vector is T3, the first audio vector is A1, the second audio vector is A2, and the third audio vector is A3, then the generated combined training vector is {T2, A2, T1, A1, T3, A3}, and each vector is separated by a special symbol.
[0131] S150: Training the initial speech synthesis model according to the combined training data set to obtain a trained speech synthesis model.
[0132] In an embodiment of the present disclosure, prosody, rhythm and / or emotional features may be extracted from the second text segment and / or the third text segment during the training.
[0133] The disclosed embodiment can be trained by inputting both the text and voice audio of the context into the model. It can use the context of the text to understand the text and capture the rhythm and emotion in the scene. It can also understand the voice corresponding to the context and obtain the rhythm and emotion of the scene. The generated voice information finally outputted is more in line with the situation and can be smoothly connected with the voice of the context.
[0134] In an embodiment of the present disclosure, for example, in the above-mentioned step S150, during the training process, relevant rhythm, melody and / or emotional features in the second text segment and / or the third text segment are extracted and used for training the speech synthesis model, so that the trained speech synthesis model can integrate the context, that is, the above-mentioned features of the second text segment and / or the third text segment to control the rhythm of the output speech and control the emotion of the speech.
[0135] In the embodiments of the present disclosure, the speech synthesis model includes multiple speech synthesis models (such as Tacotron, WaveNet, Fastspeech, etc.). It can adapt to multiple existing models without adjusting the model structure, and the method in the embodiments of the present disclosure can be used, which has high practical application value.
[0136] In some embodiments of the present disclosure, as shown in FIG5 , training the initial speech synthesis model according to the target training vector set to generate a trained speech synthesis model may include:
[0137] S151a: Iteratively inputting the combined training data in the combined training data set into the initial speech synthesis model to obtain a first output audio segment corresponding to the first text segment;
[0138] S152a: Calculate a first loss value according to the first output audio segment and the audio feature of the first audio segment;
[0139] S153a: Adjust the model parameters of the initial speech synthesis model according to the first loss value until the initial speech synthesis model converges, and use the converged initial speech synthesis model as the trained speech synthesis model.
[0140] In the disclosed embodiments, the loss values include, but are not limited to, Mel-Spectrogram Loss, Duration Loss, CTC Loss, and Angular Margin Loss. Mel-Spectrogram Loss is used to compare the mel-spectrograms of the generated audio and the input audio to ensure the accuracy of the sound features; Duration Loss ensures that the predicted pronunciation duration is consistent with the actual duration for attention-based models; CTC Loss handles misaligned input and output sequences for speech recognition models; and Angular Margin Loss is used to enhance the model's ability to identify speaker features in speaker verification and recognition tasks.
[0141] In some embodiments of the present disclosure, the loss value can be calculated based on the middle audio segment among the second audio segment, the first audio segment, and the third audio segment, that is, the first audio segment and the corresponding first output audio segment. This can make full use of the front part of the text and speech and the back part of the text and speech to fully understand the rhythm and emotion of the middle part of the text, and the model's understanding of semantics and scenes is more complete.
[0142] In some embodiments of the present disclosure, the loss value can be calculated using the first audio segment of the second, first, and third audio segments. This allows the emotion and context of the subsequent speech and text to be used to train the speech synthesis model, making the speech output more coherent. Specifically, as shown in FIG6 , the training process may include:
[0143] S151b: Input the combined training data in the combined training data set into the initial speech synthesis model to obtain a second output audio segment corresponding to the second text segment.
[0144] S152b: Calculate a second loss value according to the second output audio segment and the audio features of the second audio segment.
[0145] S153b: Adjust the model parameters of the initial speech synthesis model according to the second loss value, iteratively update until the initial speech synthesis model converges, and use the converged initial speech synthesis model as the trained speech synthesis model.
[0146] In some embodiments of the present disclosure, the loss value can also be calculated using the second audio segment, the first audio segment, and the third audio segment. This can be used to train the speech synthesis model based on the emotion and scene of the preceding speech and text, making the speech output more coherent. Specifically, as shown in FIG7 , the training process may include:
[0147] S151c: Iteratively input the combined training data in the combined training data set into the initial speech synthesis model to obtain a third output audio segment corresponding to the third text segment.
[0148] S152c: Calculate a third loss value according to the third output audio segment and the audio features of the third audio segment.
[0149] S153c: Adjust the model parameters of the initial speech synthesis model according to the third loss value until the initial speech synthesis model converges, and use the converged initial speech synthesis model as the trained speech synthesis model.
[0150] In the embodiment of the present disclosure, the target training vector is iteratively or cyclically taken out from the target training vector set and input into the initial speech synthesis model, the loss value is calculated, and the model parameters are adjusted until the model converges, that is, the training ends when the loss value is less than the set threshold.
[0151] As shown in FIG9 , the speech synthesis model 500 in the embodiment of the present disclosure can be abstracted into an encoder 510 and a decoder 520 . During the training process, the corresponding loss value is calculated, and a discriminator can also be set to perform automatic parameter adjustment training.
[0152] The initial speech synthesis model in the embodiment of the present disclosure outputs audio segments during the training process, including a first output audio segment, a second output audio segment, and a third output audio segment. According to the need to calculate the loss value, the audio segment involved in calculating the loss value is selected from the output audio segments.
[0153] In addition to the neural network part, the speech synthesis model in the embodiment of the present disclosure can also include a frequency domain to time domain processing module. After the neural network part outputs the audio segment, it is converted into a sound signal through the frequency domain to time domain processing module.
[0154] During the operation of the speech synthesis model disclosed in the present invention, the text and audio sequences first pass through the embedding module, enter the encoder, and then are output to the decoder. Finally, the output audio segment is converted from the frequency domain to the time domain signal to generate the final sound signal.
[0155] In some embodiments of the present disclosure, to further improve the accuracy of the speech output model, each segment of the output audio is included in the loss value calculation process to improve the model output's understanding of the context. As shown in FIG8 , the training process may further include:
[0156] S151d: Iteratively input the combined training data in the combined training data set into the initial speech synthesis model to obtain a first output audio segment corresponding to the first text segment, a second output audio segment corresponding to the second text segment, and a third output audio segment corresponding to the third text segment.
[0157] S152d: Calculate a first loss value based on the first output audio segment and the audio features of the first audio segment, calculate a second loss value based on the second output audio segment and the audio features of the second audio segment, calculate a third loss value based on the third output audio segment and the audio features of the third audio segment, and obtain a fourth loss value based on the first loss value, the second loss value, and the third loss value.
[0158] In the embodiment of the present disclosure, when calculating the fourth loss value, weights of the first, second, and third loss values can be set. For example, if the first loss value is L1, the second loss value is L2, and the third loss value is L3, then the fourth loss value is L4 = a*L1+b*L2+c*L3, where a, b, and c are the weights of the respective loss values.
[0159] In the process of calculating the fourth loss value, the disclosed embodiment can adjust the weight of each loss value according to the degree of attention. For example, if the trained model generates the latter part of the speech based more on the text and speech in the front and middle parts, the weight of the third loss value setting becomes larger; if the middle part of the speech is generated more based on the text and speech in the front and back parts, the weight of the first loss value setting becomes larger; if the front part of the speech is generated more based on the text and speech in the middle and back parts, the weight of the second loss value setting becomes larger.
[0160] S153d: Adjust the model parameters of the initial speech synthesis model according to the fourth loss value until the initial speech synthesis model converges, and use the converged initial speech synthesis model as the trained speech synthesis model.
[0161] In the embodiments of the present disclosure, the target speech synthesis model can be a TTS model. During training, supervised learning is performed using a large number of annotated speech-text pairs. Hyperparameters such as the learning rate and batch size are adjusted, and techniques such as dropout and batch normalization are applied to prevent overfitting. The model can also be fine-tuned and adjusted based on its performance. For example, performance on a validation set can be monitored and the model can be fine-tuned to adjust the model depth, number of hidden units, and attention mechanism.
[0162] In the disclosed embodiments, the target speech synthesis model can be a large language model. Training can be performed using a large-scale speech dataset, enabling self-supervised learning and adjustment to handle speech synthesis in different scenarios, such as classroom speech synthesis, recitation speech synthesis, or conversational speech synthesis. After the initial model is generated, model evaluation and tuning can be performed. Performance can be evaluated using an independent validation set, and model parameters such as the number of layers, number of hidden units, and learning rate can be adjusted to optimize performance.
[0163] The speech synthesis model training method in the embodiment of the present disclosure forms a sample group with the previous and next paragraphs of the training text and the corresponding training speech data. After determining whether they can be spliced together, the sample group is converted into a vector and input into the initial speech synthesis model. The initial speech synthesis model is trained to generate a target speech synthesis model. During the speech synthesis process, the generated target speech synthesis model can use the context of the text to understand the text and capture the rhythm and emotion in the scene. At the same time, it can also understand the speech corresponding to the context and obtain the melody and emotion of the scene, so that the generated speech information finally output is more in line with the situation and can be smoothly connected with the speech in the context.
[0164] As shown in FIG10 , the embodiment of the present disclosure provides a speech synthesis method, including:
[0165] S210: Acquire a target text, and divide the target text into multiple target text segments according to the text order.
[0166] In some embodiments of the present disclosure, dividing the target text into multiple target text segments according to the text order includes: performing sentence segmentation processing on the target text to obtain multiple target text segments.
[0167] In a specific embodiment of the present disclosure, the target text segment is, for example, "I went to the riverbank on holiday and saw many children flying kites. Long strings were tied to the sky at one end and to the ground at the other. The children and the kites were swinging between the sky and the ground, and even my heart was dazed by the swing, as if I had returned to my childhood." Based on sentence segmentation, the target text segment is divided into a first text segment T1 ("I went to the riverbank on holiday"), a second text segment T2 ("I saw many children flying kites"), a third text segment T3 ("long strings"), a fourth text segment T4 ("one end tied to the sky"), a fifth text segment T5 ("one end tied to the ground"), a sixth text segment T6 ("the children and the kites were swinging between the sky and the ground"), a seventh text segment T7 ("even my heart was dazed by the swing"), and an eighth text segment T8 ("as if I had returned to my childhood").
[0168] S220: Inputting a first target text segment among the multiple target text segments into a trained target speech synthesis model to obtain a first target audio segment.
[0169] In one embodiment of the present disclosure, for example, a first text segment T1 ("Walk to the riverbank on holiday") is input into a target speech synthesis model to obtain a first target audio segment A1.
[0170] In some embodiments of the present disclosure, the trained target speech synthesis model is obtained by executing any training method of the present disclosure.
[0171] In the embodiment of the present disclosure, the first audio segment may not be generated using the target speech synthesis model. Instead, the sound produced by the user after reading the first text segment may be recorded to generate the first audio segment so that subsequent speech synthesis can carry the characteristics of the user's voice, such as the user's emotions, accent, and other information.
[0172] S230: Obtain a text-audio alternating sequence by splicing the first target text segment and the first target audio segment.
[0173] In one embodiment of the present disclosure, for example, referring to the above example, a text-audio alternating sequence {T1, A1} is generated by splicing.
[0174] S240: Loop the following steps until the last target audio segment corresponding to the last target text segment is obtained: obtain the next target text segment, and add the next target text segment to the end of the text-audio alternating sequence to update the text-audio alternating sequence; input the text-audio alternating sequence into the trained speech synthesis model to obtain the next target audio segment corresponding to the next target text segment, and add the next target audio segment to the text-audio alternating sequence to update the text-audio alternating sequence.
[0175] In an embodiment of the present disclosure, for example, in the above-mentioned step S240, the next target audio segment corresponding to the next target text segment obtained extracts the rhythm, melody and / or emotional acoustic features of other target text segments. The other target text segments refer to text segments other than the currently corresponding next target text segment, that is, other text segments other than the text segment corresponding to the current audio segment. Preferably, the other text segments may be or include the previous target text segment of the next target text segment. In an embodiment of the present disclosure, the rhythm of the next target audio segment can be controlled by the context, especially the rhythm, melody and / or emotional acoustic features of the "previous context", thereby controlling the emotion of the speech.
[0176] In one embodiment of the present disclosure, for example, continuing with the above example, a second text segment is added to the text-audio alternating sequence {T1, A1}, the text-audio alternating sequence is updated to {T1, A1, T2}, the updated text-audio alternating sequence {T1, A1, T2} is input into the trained speech synthesis model to obtain the second target audio segment A2, and then based on the next text segment T3, the text-audio alternating sequence is updated to {T1, A1, T2, A2, T3}, and the trained speech synthesis model is input to obtain the corresponding third target audio segment A3, and the above steps are repeated until the eighth target audio segment A8 is generated.
[0177] In some embodiments of the present disclosure, as shown in Figure 11, text 1 is input when speech synthesis starts, the speech model generates audio 1, and spells it into a text-audio sequence, then a new text segment is input in sequence, spelled to the end of the sequence, and then the entire sequence is input into the trained speech synthesis model to synthesize the next audio segment, thereby making full use of context information during synthesis, when audio n is generated, audio n is added to the text-audio sequence, and the text-audio sequence is updated to {text 1, audio 1, text 2...audio n}, when text n+1 is input, text n+1 is added to the text-audio sequence, and the text-audio sequence is updated to {text 1, audio 1, text 2...audio n, text n+1}, the updated text-audio sequence is input into the trained speech synthesis model, audio n+1 is obtained, and the above steps are repeated until the audio corresponding to all text segments is generated.
[0178] S250: Splicing the target audio segments to generate target audio.
[0179] In an embodiment of the present disclosure, for example, in the above step S250, each target audio segment may be spliced together according to the text order to generate the target audio.
[0180] In one embodiment of the present disclosure, for example, continuing with the above example, the target audio segments finally obtained are {A1, A2, A3, A4, A5, A6, A7, A8}, which are concatenated to generate the target audio A_target.
[0181] In another alternative embodiment of the present disclosure, a speech synthesis method is also provided, in which the speech is synthesized by reversing the text order, and the synthesized speech can obtain the context, especially the prosody, rhythm and / or emotional acoustic features in the "below".
[0182] Specifically, the method may include:
[0183] Acquire a target text, and divide the target text into a plurality of target text segments according to the text order;
[0184] Inputting a last target text segment among the multiple target text segments into a trained target speech synthesis model to obtain a last target audio segment;
[0185] splicing the last target text segment and the last target audio segment to obtain a text-audio alternating sequence;
[0186] The following steps are looped and executed until a first target audio segment corresponding to a first target text segment is obtained: obtaining a previous target text segment, and adding the previous target text segment to the text-audio alternating sequence to update the text-audio alternating sequence; inputting the text-audio alternating sequence into the trained target speech synthesis model to obtain a previous target audio segment corresponding to the previous target text segment, and adding the previous target audio segment to the end of the text-audio alternating sequence to update the text-audio alternating sequence;
[0187] The target audio segments are spliced together to generate the target audio. The target audio segments can be spliced together in the order of the text to generate the target audio.
[0188] In an embodiment of the present disclosure, for example, in the above-mentioned embodiment, the previous target audio segment corresponding to the previous target text segment obtained extracts the rhythm, melody and / or emotional acoustic features of other target text segments. The other target text segments refer to text segments other than the current corresponding previous target text segment, that is, other text segments other than the text segment corresponding to the current audio segment used to synthesize. Preferably, the other text segments may be the next target text segment including the previous target text segment. In an embodiment of the present disclosure, the rhythm of the previous target audio segment can be controlled by the context, especially the rhythm, melody and / or emotional acoustic features of the "hereafter", thereby controlling the emotion of the speech.
[0189] In some embodiments of the present disclosure, the step of splicing target audio segments to generate target audio further includes: performing smoothing processing at the connections between the audio segments.
[0190] In an embodiment of the present disclosure, when the target speech synthesis model is TTS (Text-to-Speech) or a large speech model, after the target audio segments are generated, smoothing processing can be performed when the target audio segments are spliced to ensure the auditory coherence and naturalness of the synthesized audio.
[0191] In the embodiment of the present disclosure, the smoothing process is performed when the target audio segments are spliced as follows.
[0192] 1. Accurate selection of splicing points
[0193] Selection based on prosodic features: Automatically analyze the prosodic features of speech (pitch, intensity, rhythm, etc.) and select points that are prosodic and naturally connected as splicing points.
[0194] 2. Digital Signal Processing (DSP)
[0195] Time alignment and pitch correction: Use digital signal processing techniques such as WSOLA (Waveform Similarity Overlap-Add) to time-stretch / compress and pitch-match audio clips to ensure coherence at the splice points.
[0196] 3. Crossfade
[0197] Smooth Transition: Implements crossfades at splice points, making one clip gradually fade out while the other gradually increases, creating a smooth transition.
[0198] 4. Maintaining spectrum continuity
[0199] Spectral matching: Use an equalizer or other spectral processing tools to ensure that the spectral characteristics of the audio before and after the splicing remain consistent, especially in the low and high frequency ranges.
[0200] 5. Harmonic and Resonant Peak Matching
[0201] Harmonic Matching: Analyzes and matches the harmonic structure before and after the splicing point to ensure tonal coherence.
[0202] Formant Matching Adjustment: Fine-tune formants to keep the sound consistent and natural.
[0203] 6. Ambient noise and background sound processing
[0204] Noise smoothing: If background noise exists, adjust the consistency of the noise through technical means, or use noise reduction technology to eliminate the noise.
[0205] 7. Post-mixing and mastering
[0206] Dynamic range processing: Use compressors and limiters to adjust dynamic range and ensure consistent volume across your audio.
[0207] Fine-tuning: Make fine adjustments and optimizations to the spliced audio, including equalization, reverb, etc.
[0208] 8. Quality inspection and feedback
[0209] Auditory inspection: Perform a detailed auditory inspection after splicing to ensure there are no obvious splicing marks.
[0210] Objective evaluation: Use objective evaluation tools (such as PESQ, Perceptual Evaluation of Speech Quality) to evaluate the quality of the spliced audio. Adjust the parameters of the previous steps based on the evaluation quality.
[0211] In the embodiment of the present disclosure, during the speech synthesis process, the text and audio sequences first pass through the embedding module, enter the encoder, and then are output to the decoder. Finally, the output audio segment (expressed using frequency features) is generated and converted from the frequency domain to the time domain signal to generate the final sound signal.
[0212] The speech recognition method in the embodiment of the present disclosure divides the target text into multiple segments, first performs speech synthesis on the first segment of text using the target speech synthesis model, then combines the first generated audio segment with the first segment of text to form a text-audio sequence, and iteratively performs speech synthesis on subsequent text segments. The current text segment to be recognized is added to the text-audio sequence to obtain the corresponding audio segment, and the latest obtained audio segment is also added to the text-audio sequence to provide context text and audio for the next text segment, until the last text segment generates an audio segment, and then the generated audio segments are spliced to finally generate the target audio. In the speech synthesis process, the contextual text understanding and speech understanding are utilized, and the generated audio can be accompanied by the rhythm, melody and emotion of the current scene, so that the speech synthesis is more in line with the situation.
[0213] As shown in FIG12 , the embodiment of the present disclosure provides a speech synthesis model training device 1200, including a training set data acquisition module 1210, a preprocessing module 1220, a splicing detection module 1230, a vector splicing module 1240 and a training module 1250, wherein:
[0214] The training set data acquisition module 1210 is configured to acquire training set data, wherein the training set data includes coherent text and corresponding coherent audio.
[0215] The preprocessing module 1220 is configured to select multiple first text segments, multiple second text segments located before and adjacent to the first text segments, and multiple third text segments located after and adjacent to the first text segments from the coherent text, and obtain multiple first audio segments, second audio segments, and third audio segments corresponding to the multiple first text segments, second text segments, and third text segments, respectively, from the coherent audio.
[0216] The splicing detection module 1230 is configured to detect whether the texts and voices corresponding to the first text segment, the second text segment, the third text segment, the first audio segment, the second audio segment, and the third audio segment meet a splicing condition.
[0217] The vector splicing module 1240 is configured to splice multiple first text segments, second text segments, third text segments, first audio segments, second audio segments, and third audio segments that meet the splicing conditions in text order, in an alternating manner of text and audio, to generate a combined training dataset, wherein the combined training dataset includes multiple combined training data.
[0218] The training module 1250 is configured to train the initial speech synthesis model according to the combined training data set to obtain a trained speech synthesis model.
[0219] In some embodiments of the present disclosure, the splicing detection module 1230 is further configured to:
[0220] Obtaining a combined text length of the first text segment, the second text segment, and the third text segment, and obtaining a combined audio duration of the first audio segment, the second audio segment, and the third audio segment; when the combined text length is less than a length threshold and the combined audio duration is less than a duration threshold, the text and speech corresponding to the first text segment, the second text segment, the third text segment, the first audio segment, the second audio segment, and the third audio segment meet a first splicing condition;
[0221] In some other embodiments of the present disclosure, the splicing detection module 1230 is further configured to:
[0222] Obtaining a voiceprint feature similarity between the first audio segment, the second audio segment, and the third audio segment; when the voiceprint feature similarity is greater than a voiceprint similarity threshold, the text and speech corresponding to the first text segment, the second text segment, the third text segment, the first audio segment, the second audio segment, and the third audio segment meet a second splicing condition;
[0223] In some other embodiments of the present disclosure, the splicing detection module 1230 is further configured to:
[0224] When the interval between utterance time points of adjacent audio segments in the first audio segment, the second audio segment, and the third audio segment is less than an interval threshold, the text and speech corresponding to the first text segment, the second text segment, the third text segment, the first audio segment, the second audio segment, and the third audio segment meet a third splicing condition.
[0225] In some embodiments of the present disclosure, the vector stitching module 1240 is configured to:
[0226] Performing word segmentation processing on the first text segment, the second text segment, and the third text segment to generate a first text sequence, a second text sequence, and a third text sequence; performing vector conversion on the first text sequence, the second text sequence, and the third text sequence to generate a first text vector, a second text vector, and a third text vector;
[0227] Discretizing the first, second, and third audio segments and extracting audio features to generate a first, second, and third audio feature sequences; and performing vector conversion on the first, second, and third audio feature sequences to generate a first, second, and third audio vector.
[0228] The second text vector, the second audio feature sequence, the first text vector, the first audio feature sequence, the third text vector, and the third audio feature sequence are vector-concatenated in the order of the second text vector, the second audio feature sequence, the first text vector, the first audio feature sequence, the third text vector, and the third audio feature sequence to generate the combined training data.
[0229] In some embodiments of the present disclosure, the target speech synthesis model includes a large language model.
[0230] In some embodiments of the present disclosure, the training module 1250 is configured to:
[0231] Iteratively inputting the combined training data in the combined training data set into the initial speech synthesis model, obtaining a first output audio segment corresponding to the first text segment, calculating a first loss value based on the first output audio segment and audio features of the first audio segment, adjusting model parameters of the initial speech synthesis model based on the first loss value until the initial speech synthesis model converges, and using the converged initial speech synthesis model as the trained speech synthesis model;
[0232] In some other embodiments of the present disclosure, the training module 1250 is further configured to:
[0233] Inputting the combined training data in the combined training data set into the initial speech synthesis model, obtaining a second output audio segment corresponding to the second text segment, calculating a second loss value based on audio features of the second output audio segment and the second audio segment, adjusting model parameters of the initial speech synthesis model based on the second loss value, iteratively updating until the initial speech synthesis model converges, and using the converged initial speech synthesis model as the trained speech synthesis model;
[0234] In some other embodiments of the present disclosure, the training module 1250 is further configured to:
[0235] The combined training data in the combined training data set is iteratively input into the initial speech synthesis model to obtain a third output audio segment corresponding to the third text segment, and a third loss value is calculated based on the third output audio segment and the audio features of the third audio segment. The model parameters of the initial speech synthesis model are adjusted according to the third loss value until the initial speech synthesis model converges, and the converged initial speech synthesis model is used as the trained speech synthesis model.
[0236] In some embodiments of the present disclosure, the training module 1250 may also be configured to:
[0237] The combined training data in the combined training data set is iteratively input into the initial speech synthesis model to obtain a first output audio segment corresponding to the first text segment, a second output audio segment corresponding to the second text segment, and a third output audio segment corresponding to the third text segment; a first loss value is calculated based on the audio features of the first output audio segment and the first audio segment, a second loss value is calculated based on the audio features of the second output audio segment and the second audio segment, a third loss value is calculated based on the audio features of the third output audio segment and the third audio segment, and a fourth loss value is obtained based on the first loss value, the second loss value, and the third loss value; model parameters of the initial speech synthesis model are adjusted according to the fourth loss value until the initial speech synthesis model converges, and the converged initial speech synthesis model is used as the trained speech synthesis model.
[0238] As shown in FIG13 , the embodiment of the present disclosure provides a speech synthesis device 1300, which includes a target text acquisition module 1310, an initial generation module 1320, a text-audio splicing module 1330, a loop generation module 1340, and an audio splicing module 1350, wherein:
[0239] The target text acquisition module 410 is configured to acquire a target text and divide the target text into a plurality of target text segments according to the text order;
[0240] The initial generation module 1320 is configured to obtain a first target text segment from the plurality of target text segments, input a vector corresponding to the first target text segment into a trained target speech synthesis model, and obtain a first target audio segment;
[0241] The text-audio splicing module 1330 is configured to generate a text-audio alternating sequence based on the first target text segment and the first target audio segment;
[0242] The loop generation module 1340 is configured to cyclically obtain the next target text segment, add the next target text segment to the end of the text-audio alternating sequence to update the text-audio alternating sequence; input the vector corresponding to the text-audio alternating sequence into the trained target speech synthesis model, obtain the next target audio segment, and add the next target audio segment to the end of the text-audio alternating sequence to update the text-audio alternating sequence, until the last target audio segment corresponding to the last target text segment is obtained;
[0243] The audio splicing module 1350 is configured to splice target audio segments to generate target audio.
[0244] In some embodiments of the present disclosure, the trained target speech synthesis model is obtained by executing the training method in any one of the embodiments.
[0245] In some embodiments of the present disclosure, the target text acquisition module 1310 is configured to:
[0246] The target text is segmented to obtain multiple target text segments.
[0247] In some embodiments of the present disclosure, the audio splicing module 1350 is further configured to:
[0248] Smooths the joins between audio segments.
[0249] FIG14 shows a schematic diagram of an electronic device 1400 that can be used to implement the method or realize the embodiments of the present disclosure. In some embodiments, the number of electronic devices may be more or less than the number shown. In some embodiments, a single or multiple electronic devices may be used for implementation. In some embodiments, cloud-based or distributed electronic devices may also be used for implementation.
[0250] As shown in Figure 14, the electronic device 1400 includes a processor 1410 and a memory 1420. The processor is used to execute programs stored in the memory, which can implement the methods, steps or functions described in the above embodiments when executed by a computer. The processor 1410 may include various types of processors, such as a central processing unit (CPU), a graphics processing unit (GPU), a neural network processor (NPU), a digital signal processor (DSP), etc. The processor 1410 and the memory 1420 are interconnected via a bus 1430. An input / output (I / O) interface and the like can also be connected to the bus 1430.
[0251] Although not shown in the figures, the embodiments of the present disclosure also provide a storage medium storing computer programs configured to execute speech synthesis model training and speech synthesis methods related to any embodiment at runtime.
[0252] It should be understood that all operations in the method described above are merely exemplary, and the present disclosure is not limited to any operation in the method or the order of these operations, but should cover all other equivalent transformations under the same or similar concept.
[0253] It should also be understood that all modules in the above-described device can be implemented in various ways. These modules can be implemented as hardware, software, or a combination thereof. In addition, any module in these modules can be further divided into submodules or combined together in function.
[0254] Processor has been described in conjunction with various devices and methods.These processors can be implemented using electronic hardware, computer software or its arbitrary combination.Whether these processors are implemented as hardware or software will depend on specific application and the overall design constraint imposed on the system.As an example, the processor provided in this disclosure, any part of the processor or any combination of processors can be implemented as microprocessor, microcontroller, digital signal processor (DSP), field programmable gate array (FPGA), programmable logic device (PLD), state machine, gate logic, discrete hardware circuit and other suitable processing components configured for performing the various functions described in this disclosure.The function of the processor provided in this disclosure, any part of the processor or any combination of processors can be implemented as software performed by microprocessor, microcontroller, DSP or other suitable platform.
[0255] Software should be broadly considered to mean instructions, instruction sets, codes, code segments, program codes, programs, subroutines, software modules, applications, software applications, software packages, routines, subroutines, objects, running threads, processes, functions, etc. Software can reside in a computer-readable medium. A computer-readable medium can include, for example, a memory, which can be, for example, a magnetic storage device (e.g., a hard disk, a floppy disk, a magnetic stripe), an optical disk, a smart card, a flash memory device, a random access memory (RAM), a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), a register, or a removable disk. Although the memory is shown as being separate from the processor in various aspects provided in the present disclosure, the memory can also be located inside the processor (e.g., a cache or register).
[0256] The above description is provided to enable any person skilled in the art to implement the various aspects described herein. Various modifications to these aspects will be apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects. Therefore, the claims are not intended to be limited to the aspects shown herein. All structural and functional equivalents of the elements of the various aspects described in this disclosure that are known or to be known to those skilled in the art are expressly incorporated herein by reference and are intended to be covered by the claims. Industrial Applicability
[0257] By using the above solution, contextual text and speech understanding are utilized in the speech synthesis process. The generated audio can be accompanied by the rhythm, melody, and emotion of the current scene, making the speech synthesis more consistent with the situation and enhancing the realism of the character's speech or dialogue.< / unk>
Claims
1. A method for training a speech synthesis model, characterized in that Including the following steps: Obtain initial training data, where the initial training data includes coherent text and corresponding coherent audio; Select multiple first text segments, multiple second text segments adjacent to and before the first text segments, and multiple third text segments adjacent to and after the first text segments from the coherent text, and obtain multiple first audio segments, second audio segments, and third audio segments corresponding to the multiple first text segments, second text segments, and third text segments respectively from the coherent audio; Detect whether the associated first text segments, second text segments, third text segments, first audio segments, second audio segments, and third audio segments meet the splicing conditions; Splice the first text segments, second text segments, third text segments, first audio segments, second audio segments, and third audio segments that meet the splicing conditions in the text order, in an alternating manner of text and audio, to obtain combined training data, and generate a combined training data set including multiple pieces of the combined training data; Train an initial speech synthesis model according to the combined training data set to obtain a trained speech synthesis model. During the training, prosody, pitch, and / or emotional features in the second text segment and / or the third text segment are extracted.
2. The method according to claim 1, characterized in that, The detecting whether the associated first text segments, second text segments, third text segments, first audio segments, second audio segments, and third audio segments meet the splicing conditions includes: Obtain the combined text length of the first text segment, second text segment, and third text segment, and obtain the combined audio duration of the first audio segment, second audio segment, and third audio segment. When the combined text length is less than a length threshold and the combined audio duration is less than a duration threshold, the text and speech corresponding to the first text segment, second text segment, third text segment, first audio segment, second audio segment, and third audio segment meet the first splicing condition.
3. The method according to claim 1, characterized in that The detecting whether the associated first text segments, second text segments, third text segments, first audio segments, second audio segments, and third audio segments meet the splicing conditions includes: Obtain the acoustic feature similarity between the first audio segment, second audio segment, and third audio segment. When the acoustic feature similarity is greater than an acoustic similarity threshold, the text and speech corresponding to the first text segment, second text segment, third text segment, first audio segment, second audio segment, and third audio segment meet the second splicing condition.
4. The method according to claim 1, wherein The detecting whether the associated first text segments, second text segments, third text segments, first audio segments, second audio segments, and third audio segments meet the splicing conditions includes: When the time point spacing between adjacent audio segments in the first audio segment, second audio segment, and third audio segment is less than an interval threshold, the text and speech corresponding to the first text segment, second text segment, third text segment, first audio segment, second audio segment, and third audio segment meet the third splicing condition.
5. The method according to claim 1, wherein The splicing in the text order, in an alternating manner of text and audio, to obtain combined training data includes: Tokenize the first text segment, second text segment, and third text segment to generate a first text sequence, a second text sequence, and a third text sequence, and perform vector transformation on the first text sequence, second text sequence, and third text sequence to generate a first text vector, a second text vector, and a third text vector; Perform discretization processing on the first audio segment, second audio segment, and third audio segment and extract audio features to generate a first audio feature sequence, a second audio feature sequence, and a third audio feature sequence, and perform vector transformation on the first audio feature sequence, second audio feature sequence, and third audio feature sequence to generate a first audio vector, a second audio vector, and a third audio vector; Perform vector concatenation on the second text vector, second audio feature sequence, first text vector, first audio feature sequence, third text vector, and third audio feature sequence in the order of the second text vector, second audio feature sequence, first text vector, first audio feature sequence, third text vector, and third audio feature sequence to generate the combined training data.
6. The method according to claim 1, characterized in that Training the initial speech synthesis model according to the combined training dataset to obtain a trained speech synthesis model, including: Iteratively input the combined training data in the combined training dataset into the initial speech synthesis model, obtain a first output audio segment corresponding to the first text segment, calculate a first loss value according to the audio features of the first output audio segment and the first audio segment, and adjust the model parameters of the initial speech synthesis model according to the first loss value until the initial speech synthesis model converges, and use the converged initial speech synthesis model as the trained speech synthesis model.
7. The method according to claim 1, characterized in that, Training the initial speech synthesis model according to the combined training dataset to obtain a trained speech synthesis model, including: Input the combined training data in the combined training dataset into the initial speech synthesis model, obtain a second output audio segment corresponding to the second text segment, calculate a second loss value according to the audio features of the second output audio segment and the second audio segment, and adjust the model parameters of the initial speech synthesis model according to the second loss value, and iteratively update until the initial speech synthesis model converges, and use the converged initial speech synthesis model as the trained speech synthesis model.
8. The method according to claim 1, characterized in that, Training the initial speech synthesis model according to the combined training dataset to obtain a trained speech synthesis model, including: Iteratively input the combined training data in the combined training dataset into the initial speech synthesis model, obtain a third output audio segment corresponding to the third text segment, calculate a third loss value according to the audio features of the third output audio segment and the third audio segment, and adjust the model parameters of the initial speech synthesis model according to the third loss value until the initial speech synthesis model converges, and use the converged initial speech synthesis model as the trained speech synthesis model.
9. The method according to claim 1, characterized in that Training the initial speech synthesis model according to the combined training dataset to obtain a trained speech synthesis model, including: Iteratively input the combined training data in the combined training dataset into the initial speech synthesis model to obtain a first output audio segment corresponding to the first text segment, a second output audio segment corresponding to the second text segment, and a third output audio segment corresponding to the third text segment; calculate a first loss value according to the audio features of the first output audio segment and the first audio segment, calculate a second loss value according to the audio features of the second output audio segment and the second audio segment, calculate a third loss value according to the audio features of the third output audio segment and the third audio segment, and obtain a fourth loss value according to the first loss value, the second loss value, and the third loss value; adjust the model parameters of the initial speech synthesis model according to the fourth loss value until the initial speech synthesis model converges, and use the converged initial speech synthesis model as the trained speech synthesis model.
10. A voice synthesis method, characterized in that, It includes the following steps: Obtain a target text, and divide the target text into multiple target text segments in the text order; Input the first target text segment in the multiple target text segments into the trained target speech synthesis model to obtain a first target audio segment; Concatenate the first target text segment and the first target audio segment to obtain a text-audio alternating sequence; Loop and execute the following steps until the last target audio segment corresponding to the last target text segment is obtained: obtain the next target text segment, and add the next target text segment to the text-audio alternating sequence to update the text-audio alternating sequence; input the text-audio alternating sequence into the trained speech synthesis model to obtain the next target audio segment corresponding to the next target text segment, and add the next target audio segment to the text-audio alternating sequence to update the text-audio alternating sequence; Perform a concatenation process on each target audio segment to generate a target audio.
11. A voice synthesis method, characterized in that, It includes the following steps: Obtain a target text, and divide the target text into multiple target text segments in the text order; Input the last target text segment in the multiple target text segments into the trained target speech synthesis model to obtain a last target audio segment; Concatenate the last target text segment and the last target audio segment to obtain a text-audio alternating sequence; Loop and execute the following steps until the first target audio segment corresponding to the first target text segment is obtained: obtain the previous target text segment, and add the previous target text segment to the text-audio alternating sequence to update the text-audio alternating sequence; input the text-audio alternating sequence into the trained target speech synthesis model to obtain the previous target audio segment corresponding to the previous target text segment, and add the previous target audio segment to the end of the text-audio alternating sequence to update the text-audio alternating sequence; Perform a concatenation process on each target audio segment to generate a target audio.
12. The method according to claim 10 or 11, characterized in that The trained speech synthesis model is obtained by executing the method described in any one of claims 1 to 9.
13. The method according to claim 10 or 11, characterized in that The dividing the target text into multiple target text segments in the text order includes: Perform a sentence-breaking process on the target text to obtain multiple target text segments.
14. The method according to claim 10 or 11, characterized in that, Splicing the respective target audio segments to generate a target audio further includes: Performing smoothing processing at the connection points of the respective audio segments.
15. A voice synthesis model, characterized in that, Comprising an encoder and a decoder, wherein the speech synthesis model is trained by the method according to any one of claims 1 to 9.
16. A voice synthesis model training device, characterized in that, Comprising a training set data acquisition module, a preprocessing module, a splicing detection module, a vector splicing module, and a training module, wherein, The training set data acquisition module is configured to acquire training set data, wherein the training set data includes coherent text and corresponding coherent audio; The preprocessing module is configured to select multiple first text segments, multiple second text segments adjacent to and in front of the first text segments, and multiple third text segments adjacent to and behind the first text segments from the coherent text, and acquire multiple first audio segments, second audio segments, and third audio segments corresponding to the multiple first text segments, second text segments, and third text segments respectively from the coherent audio; The splicing detection module is configured to detect whether the text and speech corresponding to the multiple first text segments, second text segments, third text segments, first audio segments, second audio segments, and third audio segments meet the splicing conditions; The vector splicing module is configured to splice the multiple first text segments, second text segments, third text segments, first audio segments, second audio segments, and third audio segments that meet the splicing conditions in the text order in an alternating manner of text and audio to generate a combined training data set, wherein the combined training data set includes multiple combined training data; The training module is configured to train the initial speech synthesis model according to the combined training data set to obtain a trained speech synthesis model.
17. A voice synthesis device, characterized in that, Comprising a target text acquisition module, an initial generation module, a text-audio splicing module, a cyclic generation module, and an audio splicing module, wherein, The target text acquisition module is configured to acquire a target text and divide the target text into multiple target text segments in the text order; The initial generation module is configured to acquire the first target text segment among the multiple target text segments, input the vector corresponding to the first target text segment into the trained speech synthesis model to obtain a first target audio segment; The text-audio splicing module is configured to splice the first target text segment and the first target audio segment to generate a text-audio alternating sequence; The cyclic generation module is configured to cyclically acquire the next target text segment, add the next target text segment to the end of the text-audio alternating sequence to update the text-audio alternating sequence; input the vector corresponding to the text-audio alternating sequence into the trained speech synthesis model to obtain the next target audio segment, and add the next target audio segment to the end of the text-audio alternating sequence to update the text-audio alternating sequence until the last target audio segment corresponding to the last target text segment is acquired; The audio splicing module is configured to splice the respective target audio segments to generate a target audio.
18. An electronic device, characterized in that, Comprising a processor and a memory storing a computer program, the processor being configured to execute the method according to any one of claims 1 to 14 when running the computer program.
19. A storage medium, characterized in that, The storage medium stores a computer program, the computer program being configured to execute the method according to any one of claims 1 to 14 when being run.
Citation Information
Patent Citations
Text-based audio generation method and device
CN113192484A
Speech synthesis model, model training method and speech synthesis method
CN113920977A
Spliced voice generation method and device, electronic equipment and storage medium
CN115602146A
Speech synthesis model training method and device, electronic equipment and storage medium
CN116312474A
Speech synthesis model training method, speech synthesis method, electronic equipment and storage medium
CN118116364A
Cited By
Audio processing method and device
CN120877767A
Speech synthesis method and device, electronic equipment and storage medium
CN120895024A
Voice cloning method and device based on TTS voice technology, and storage medium
CN121053960A