Speech synthesis method and device
By splicing coherent text and audio bands according to specific conditions, a combined training data set is generated and used to train a speech synthesis model, the problem of inconsistent speech and text semantics in the prior art is solved, and a more context-compatible speech synthesis is achieved.
Patent Information
- Application Number
- CN202411879546.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2043-12-29
AI Technical Summary
The current text generation pronunciation model is difficult to match the overall text semantics, rhythms, and emotions, resulting in the generated pronunciation that does not conform to the situation and the pronunciation is not smooth.
By obtaining the initial training data, including coherent text and audio, selecting multiple text segments and corresponding audio segments, detecting whether these paragraphs meet the splicable conditions, and splicing the text and audio segments that meet the conditions in the alternating text and audio in text order, generating a combined training data set for training the speech synthesis model.
Through context understanding, the rhythm and emotions in the scene are captured, so that the generated voice is more in line with the real scene, and the voice connection with the context is smooth, improving the authenticity and coherence of speech synthesis.
Smart Images

Figure CN119400155B_ABST
Abstract
Description
[0001] This application is a divisional application. The application number of the original application is 202311870114.7, the application date is December 29, 2023, and the name of the invention is "Speech synthesis model training method, speech synthesis method, electronic device and storage medium". Technical Field
[0002] The present application relates to the technical field of multimedia content processing, and in particular to a speech synthesis method, a speech synthesis device, an electronic device and a storage medium. Background Art
[0003] Speech synthesis technology is a technology that converts text into speech. It can achieve text-to-speech function through neural network models. In the current text-to-speech model, only the current text input is converted into speech. The generated speech is difficult to match the semantics, rhythm, and emotion of the overall text, resulting in the speaker's voice not being in line with the situation and the pronunciation not being connected smoothly.
[0004] The description of the background technology is intended to help understand the relevant technology in the relevant field, and does not constitute an admission that the content of the background technology belongs to the prior art. Summary of the invention
[0005] Therefore, the embodiments of the present application intend to provide a speech synthesis method and related electronic equipment and storage medium. Through the scheme of the embodiments of the present application, the rhythm, melody and emotion of the current scene can be captured based on the text understanding of the context, so that the generated speech is more in line with the real scene.
[0006] In a first aspect, an embodiment of the present application provides a speech synthesis model training method, comprising the following steps:
[0007] Acquire initial training data, wherein the initial training data includes coherent text and corresponding coherent audio;
[0008] Selecting from the continuous text a plurality of first text segments, a plurality of second text segments located before the first text segment and adjacent to the first text segment, and a plurality of third text segments located after the first text segment and adjacent to the first text segment, and acquiring from the continuous audio a plurality of first audio segments, second audio segments, and third audio segments corresponding to the plurality of first text segments, second text segments, and third text segments, respectively;
[0009] Detecting whether the associated first text segment, second text segment, third text segment, first audio segment, second audio segment, and third audio segment meet a splicing condition;
[0010] splicing the first text segment, the second text segment, the third text segment, the first audio segment, the second audio segment and the third audio segment that meet the splicing condition in a text order in an alternating manner of text and audio to obtain combined training data, so as to generate a combined training data set including a plurality of the combined training data;
[0011] The initial speech synthesis model is trained according to the combined training data set to obtain a trained speech synthesis model, during which the prosody, rhythm and / or emotional features in the second text segment and / or the third text segment are extracted.
[0012] In some embodiments of the present application, the detecting whether the associated first text segment, second text segment, third text segment, first audio segment, second audio segment, and third audio segment meet the splicing condition includes:
[0013] The combined text length of the first text segment, the second text segment, and the third text segment is obtained, and the combined audio duration of the first audio segment, the second audio segment, and the third audio segment is obtained. When the combined text length is less than a length threshold and the combined audio duration is less than a duration threshold, the text and speech corresponding to the first text segment, the second text segment, the third text segment, the first audio segment, the second audio segment, and the third audio segment meet a first splicing condition.
[0014] In some embodiments of the present application, the detecting whether the associated first text segment, second text segment, third text segment, first audio segment, second audio segment, and third audio segment meet the splicing condition includes:
[0015] The voiceprint feature similarity between the first audio segment, the second audio segment and the third audio segment is obtained. When the voiceprint feature similarity is greater than a voiceprint similarity threshold, the text and speech corresponding to the first text segment, the second text segment, the third text segment, the first audio segment, the second audio segment and the third audio segment meet a second splicing condition.
[0016] In some embodiments of the present application, the detecting whether the associated first text segment, second text segment, third text segment, first audio segment, second audio segment, and third audio segment meet the splicing condition includes:
[0017] When the interval between the utterance time points of adjacent audio segments in the first audio segment, the second audio segment and the third audio segment is less than the interval threshold, the text and speech corresponding to the first text segment, the second text segment, the third text segment, the first audio segment, the second audio segment and the third audio segment meet the third splicing condition.
[0018] In some embodiments of the present application, the combined training data is obtained by splicing text and audio alternately in the order of text, including:
[0019] Performing word segmentation processing on the first text segment, the second text segment, and the third text segment to generate a first text sequence, a second text sequence, and a third text sequence, and performing vector conversion on the first text sequence, the second text sequence, and the third text sequence to generate a first text vector, a second text vector, and a third text vector;
[0020] Discretize the first audio segment, the second audio segment, and the third audio segment and extract audio features to generate a first audio feature sequence, a second audio feature sequence, and a third audio feature sequence; perform vector conversion on the first audio feature sequence, the second audio feature sequence, and the third audio feature sequence to generate a first audio vector, a second audio vector, and a third audio vector;
[0021] The second text vector, the second audio feature sequence, the first text vector, the first audio feature sequence, the third text vector, and the third audio feature sequence are vector-concatenated in the order of the second text vector, the second audio feature sequence, the first text vector, the first audio feature sequence, the third text vector, and the third audio feature sequence to generate the combined training data.
[0022] In some embodiments of the present application, the initial speech synthesis model is trained according to the combined training data set to obtain a trained speech synthesis model, including:
[0023] Iteratively input the combined training data in the combined training data set into the initial speech synthesis model, obtain a first output audio segment corresponding to the first text segment, calculate a first loss value based on the first output audio segment and the audio features of the first audio segment, adjust the model parameters of the initial speech synthesis model based on the first loss value until the initial speech synthesis model converges, and use the converged initial speech synthesis model as the trained speech synthesis model.
[0024] In some embodiments of the present application, the initial speech synthesis model is trained according to the combined training data set to obtain a trained speech synthesis model, including:
[0025] Inputting the combined training data in the combined training data set into the initial speech synthesis model, obtaining a second output audio segment corresponding to the second text segment, calculating a second loss value according to the audio features of the second output audio segment and the second audio segment, adjusting the model parameters of the initial speech synthesis model according to the second loss value, iteratively updating until the initial speech synthesis model converges, and using the converged initial speech synthesis model as the trained speech synthesis model;
[0026] In some embodiments of the present application, the initial speech synthesis model is trained according to the combined training data set to obtain a trained speech synthesis model, including:
[0027] Iteratively input the combined training data in the combined training data set into the initial speech synthesis model, obtain a third output audio segment corresponding to the third text segment, calculate a third loss value based on the third output audio segment and the audio features of the third audio segment, adjust the model parameters of the initial speech synthesis model according to the third loss value until the initial speech synthesis model converges, and use the converged initial speech synthesis model as the trained speech synthesis model.
[0028] In some embodiments of the present application, the initial speech synthesis model is trained according to the combined training data set to obtain a trained speech synthesis model, including:
[0029] Iteratively input the combined training data in the combined training data set into the initial speech synthesis model to obtain a first output audio segment corresponding to the first text segment, a second output audio segment corresponding to the second text segment, and a third output audio segment corresponding to the third text segment; calculate a first loss value according to the audio features of the first output audio segment and the first audio segment, calculate a second loss value according to the audio features of the second output audio segment and the second audio segment, calculate a third loss value according to the audio features of the third output audio segment and the third audio segment, and obtain a fourth loss value according to the first loss value, the second loss value, and the third loss value; adjust the model parameters of the initial speech synthesis model according to the fourth loss value until the initial speech synthesis model converges, and use the converged initial speech synthesis model as the trained speech synthesis model.
[0030] In a second aspect, the embodiments of the present application provide a speech synthesis method, which may include:
[0031] Acquire a target text, and divide the target text into a plurality of target text segments according to the text order;
[0032] Inputting a first target text segment among the multiple target text segments into a trained target speech synthesis model to obtain a first target audio segment;
[0033] splicing the first target text segment and the first target audio segment to obtain a text-audio alternating sequence;
[0034] The following steps are executed in a loop until the last target audio segment corresponding to the last target text segment is obtained: obtaining a next target text segment, adding the next target text segment to the text-audio alternating sequence to update the text-audio alternating sequence; inputting the text-audio alternating sequence into the trained speech synthesis model to obtain a next target audio segment corresponding to the next target text segment and extracting the rhythm, cadence and / or emotional acoustic features of other target text segments, and adding the next target audio segment to the text-audio alternating sequence to update the text-audio alternating sequence;
[0035] Each target audio segment is spliced together to generate the target audio.
[0036] In some other embodiments of the present application, another feasible speech synthesis method is provided, which may include:
[0037] Acquire a target text, and divide the target text into a plurality of target text segments according to the text order;
[0038] Inputting a last target text segment among the multiple target text segments into a trained target speech synthesis model to obtain a last target audio segment;
[0039] splicing the last target text segment and the last target audio segment to obtain a text-audio alternating sequence;
[0040] The following steps are executed in a loop until the first target audio segment corresponding to the first target text segment is obtained: the previous target text segment is obtained, and the previous target text segment is added to the text-audio alternating sequence to update the text-audio alternating sequence; the text-audio alternating sequence is input into the trained target speech synthesis model to obtain the previous target audio segment corresponding to the previous target text segment and from which the rhythm, cadence and / or emotional acoustic features of other target text segments are extracted, and the previous target audio segment is added to the end of the text-audio alternating sequence to update the text-audio alternating sequence;
[0041] Each target audio segment is spliced together to generate the target audio.
[0042] In a third aspect, an embodiment of the present application provides a speech synthesis model, including an encoder and a decoder, wherein the speech synthesis model is trained by the training method in any embodiment of the present application.
[0043] In a fourth aspect, an embodiment of the present application provides a speech synthesis model training device, including a training set data acquisition module, a preprocessing module, a splicing detection module, a vector splicing module and a training module, wherein:
[0044] The training set data acquisition module is configured to acquire training set data, wherein the training set data includes coherent text and corresponding coherent audio;
[0045] The preprocessing module is configured to select a plurality of first text segments, a plurality of second text segments located before the first text segment and adjacent to the first text segment, and a plurality of third text segments located after the first text segment and adjacent to the first text segment from the continuous text, and obtain a plurality of first audio segments, second audio segments, and third audio segments corresponding to the plurality of first text segments, second text segments, and third text segments, respectively, from the continuous audio;
[0046] The splicing detection module is configured to detect whether the texts and voices corresponding to the plurality of first text segments, second text segments, third text segments, first audio segments, second audio segments and third audio segments meet a splicing condition;
[0047] The vector concatenation module is configured to concatenate a plurality of first text segments, second text segments, third text segments, first audio segments, second audio segments, and third audio segments that meet the concatenation condition in a text order in an alternating manner of text and audio to generate a combined training data set, wherein the combined training data set includes a plurality of combined training data;
[0048] The training module is configured to train the initial speech synthesis model according to the combined training data set to obtain a trained speech synthesis model.
[0049] In a fifth aspect, an embodiment of the present application provides a speech synthesis device, including a target text acquisition module, an initial generation module, a text-audio splicing module, a loop generation module and an audio splicing module, wherein:
[0050] The target text acquisition module is configured to acquire a target text and divide the target text into a plurality of target text segments according to the text order;
[0051] The initial generation module is configured to obtain a first target text segment from the multiple target text segments, input a vector corresponding to the first target text segment into a trained target speech synthesis model, and obtain a first target audio segment;
[0052] The text-audio splicing module is configured to generate a text-audio alternating sequence by splicing the first target text segment and the first target audio segment;
[0053] The loop generation module is configured to loop to obtain the next target text segment, add the next target text segment to the end of the text-audio alternating sequence to update the text-audio alternating sequence; input the vector corresponding to the text-audio alternating sequence into the trained target speech synthesis model, obtain the next target audio segment, add the next target audio segment to the end of the text-audio alternating sequence to update the text-audio alternating sequence, until the last target audio segment corresponding to the last target text segment is obtained;
[0054] The audio splicing module is configured to splice each target audio segment to generate a target audio.
[0055] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, it implements any method of the embodiment of the present application.
[0056] In a seventh aspect, an embodiment of the present application provides an electronic device, comprising: a processor and a memory storing a computer program, wherein the processor is configured to execute any method of the embodiment of the present application when running the computer program.
[0057] The technical scheme provides a speech synthesis model training method, which forms a sample group with the previous paragraph of the training text and the next paragraph of the training text and the corresponding training speech data. After determining whether they can be spliced, the sample group is converted into a vector and input into the initial speech synthesis model, the initial speech synthesis model is trained, and a trained speech synthesis model is generated. During the speech synthesis process, the generated trained speech synthesis model can use the context of the text to understand the text and capture the rhythm and emotion in the scene. At the same time, it can also understand the speech corresponding to the context and obtain the melody and emotion of the scene, so that the generated speech information finally outputted is more in line with the situation and can be smoothly connected with the speech in the context. The shorter audio in the training data is spliced together to lengthen it, thereby improving the training efficiency and making full use of the data. This technical solution also provides a speech recognition method, which divides the target text into multiple segments, first uses the trained speech synthesis model to perform speech synthesis on the first segment of text, then combines the generated first audio segment with the first segment of text to form a text-audio sequence, and iterates the speech synthesis for subsequent text segments. The current text segment to be recognized is added to the text-audio sequence, and the corresponding audio segment is obtained. The latest audio segment obtained is also added to the text-audio sequence, and the context text and audio are provided for the next segment of text, until the last text segment generates an audio segment, and then the generated audio segments are spliced to finally generate the target audio. In some embodiments of the present application, speech synthesis can also be started from the last segment, and then the generated audio segments are spliced until the first text segment generates an audio segment to generate the target audio, and then the generated audio segments are spliced to finally generate the target audio. In the speech synthesis process of the embodiment of the present application, the text understanding and speech understanding of the context are utilized, and the generated audio can be accompanied by the rhythm, rhythm and emotion of the current scene, so that the speech synthesis is more in line with the scene, and the realism of the character's speech or dialogue is improved. The solution in the embodiment of the present application can utilize the existing speech synthesis model framework for speech model training and speech synthesis. The technology is mature and reliable, easy to operate, and has high practical application value.
[0058] Other optional features and technical effects of the embodiments of the present application are partially described below, and partially can be understood by reading this document. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use are briefly introduced below. It should be clear that the drawings in the following description only cover some embodiments of the present application, and ordinary technicians in this field can obtain other drawings based on these drawings without performing creative work. The purpose of these drawings is to better illustrate the technical details to help understand the implementation methods of the present application, among which:
[0060] Figure 1A schematic diagram showing data flow of a speech synthesis model training method according to an embodiment of the present application;
[0061] Figure 2 A schematic flow chart showing a method for training a speech synthesis model according to an embodiment of the present application is shown;
[0062] Figure 3 A schematic diagram showing the division of text segments and audio segments in the speech synthesis model training method according to an embodiment of the present application is shown;
[0063] Figure 4 A schematic diagram showing the process of vectorizing text segments and audio segments in the speech synthesis model training method according to an embodiment of the present application;
[0064] Figure 5 A schematic diagram showing the flow of model training in the speech synthesis model training method of an embodiment of the present application;
[0065] Figure 6 Another schematic diagram showing the model training process in the speech synthesis model training method according to an embodiment of the present application;
[0066] Figure 7 Another schematic diagram showing the model training process in the speech synthesis model training method according to an embodiment of the present application;
[0067] Figure 8 Another schematic diagram showing the model training process in the speech synthesis model training method according to an embodiment of the present application;
[0068] Fig. 9 A schematic diagram showing the workflow of training a speech synthesis model by the speech synthesis model training method according to an embodiment of the present application;
[0069] Fig.10 A schematic flow chart showing a speech synthesis method according to an embodiment of the present application;
[0070] Fig.11 A schematic flow chart showing the cyclic generation of target audio in the speech synthesis method of an embodiment of the present application;
[0071] Fig.12 An exemplary structural diagram showing a speech synthesis model training device according to an embodiment of the present application;
[0072] Fig.13 An exemplary structural diagram showing a speech synthesis device according to an embodiment of the present application;
[0073] Fig.14 An exemplary structural diagram of an electronic device capable of implementing the method according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0074] The exemplary embodiments of the present application will be described in detail below, and the illustrations of the examples have been shown in the accompanying drawings. When referring to the drawings, unless otherwise specified, the same numbers or marks in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with one or more embodiments of this specification. Instead, they are only examples of some aspects of the devices and methods covered by one or more embodiments of this specification as detailed in the claims of this application.
[0075] The term "including" and its variations used in this specification are all intended to be broadly inclusive, i.e., "including but not limited to" the items listed. Unless otherwise stated, the term "or" means "and / or", the term "based on" means relying on, or at least partially relying on, the terms "an example embodiment" and "an embodiment" refer to at least one example embodiment, and the term "another embodiment" means at least one different embodiment. The terms "first", "second", etc. may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0076] In an embodiment of the present application, the speech synthesis process generates speech audio corresponding to the target text segment based on the input target text segment and context text content and context audio content. It can be implemented through a speech synthesis (Text to Speech, TTS) model or a large language model.
[0077] In the embodiment of the present application, "prosody" is one of the attributes of speech, including the pitch, rhythm, etc. of speech. "Semantics" is one of the attributes of speech, identifying the content of speech.
[0078] In the embodiments of the present application, "model" has the conventional meaning in the field of machine learning. For example, the model can be a machine learning or deep learning model, such as a machine learning or deep learning model including the above-mentioned network or composed of the above-mentioned network.
[0079] In the embodiments of the present application, “loss function” and “loss value” have conventional meanings in the field of machine learning.
[0080] In the embodiment of the present application, "embedding" refers to the step of data processing, which converts discrete data such as text into continuous data vectors so that the model can calculate it.
[0081] In the embodiment of the present application, a "phoneme" is the smallest unit of speech divided according to the natural properties of speech, and one pronunciation action forms a phoneme. A "phoneme sequence" is a sequence composed of a number of phonemes.
[0082] The embodiments of the present application provide a speech synthesis model training method and device, a speech synthesis method, a speech synthesis model, a training device, a generating device, a storage medium and an electronic device. The method, device / model can be implemented with the aid of one or more computers. In some embodiments, the device / model can be implemented by software, hardware, or a combination of software and hardware. In some embodiments, the electronic device or computer can be implemented by a computer described herein or other electronic device that can realize the corresponding function.
[0083] The inventor of the present application found that the speech of a normal person is affected by many factors, not only depending on the text to be expressed, but also by the situation and context. However, in the current TTS model, the model cannot obtain the context information of the speech segment to be synthesized, resulting in the synthesized speech not being consistent with the situation, causing the connection with the context to be unsmooth, and at the same time, the tone and other attributes of the speech cannot be controlled according to demand during model synthesis. Some improvement schemes propose to introduce additional inputs other than text into the input content of the speech synthesis process, such as corresponding to the style description of the text, and a prompt word processing module is added to some TTS models. However, the inventor found that this speech synthesis method requires the creation of a special training data set, in which each data needs to contain audio, text and style description, and the data annotation is difficult. In actual application and promotion, it will lead to additional costs, and it is difficult to expand the data scale to improve the effect. At the same time, the newly added prompt word processing module requires additional fine-tuning of a BERT model, resulting in a longer link and increased errors.
[0084] In this regard, the embodiments of the present application provide a speech synthesis model training method and device, a speech synthesis method and device, a speech synthesis model, a storage medium and an electronic device, and propose a method for speech synthesis using context, which controls the rhythm of speech through context, thereby controlling the emotion of speech, and also makes the generated speech more coherent.
[0085] like Figure 1As shown, the model training process in the embodiment of the present application is to select a coherent text and corresponding audio from the training set, divide it into three parts: the previous text, the current text, and the following text according to the part to be synthesized, convert the text of each part into a digital sequence through word segmentation, and then use embedding to convert it into training data that can be input into the model. Similarly, the corresponding audio data is discretized and converted into training data through embedding; check whether the three parts of data meet the conditions for splicing: (1) the total length cannot be too long; (2) each part needs to come from the same speaker; (3) the position of each part in the original audio cannot be too far apart and needs to be coherent. If at least one of the above splicing conditions is met, the text and audio of the three parts are spliced together and separated by special symbols; then the spliced training data is input into the model for training to obtain the corresponding speech synthesis model. In the embodiment of the present application, if there is noise in each audio segment, the noise can be retained as the ambient sound of the input model.
[0086] Figure 2 In the illustrated embodiment, a speech synthesis model training method is provided, and a trained speech synthesis model is obtained through training.
[0087] like Figure 2 As shown, the training method in the embodiment of the present application includes the following steps.
[0088] S110: Acquire initial training data, wherein the initial training data includes coherent text and corresponding coherent audio.
[0089] In an embodiment of the present application, the initial training data may be obtained from conventional training data, and the conventional training data may include a plurality of coherent texts and a plurality of coherent audios corresponding thereto, without the need for additional annotation of the training data.
[0090] S120: Selecting from the continuous text a plurality of first text segments, a plurality of second text segments located before the first text segment and adjacent to the first text segment, and a plurality of third text segments located after the first text segment and adjacent to the first text segment, and acquiring from the continuous audio a plurality of first audio segments, second audio segments, and third audio segments corresponding to the plurality of first text segments, second text segments, and third text segments, respectively.
[0091] In an embodiment of the present application, contextual text data and voice data are selected and input into the model for training, thereby improving the scene understanding ability of the speech synthesis model.
[0092] In some embodiments of the present application, Figure 3As shown, the present application selects a first text, a second text which is the preceding text of the first text, and a third text which is the following text of the first text from a coherent text, and correspondingly extracts a first audio, a second audio, and a third audio from the coherent speech.
[0093] In the embodiment of the present application, the division basis of the audio segment end line is the end of the text pronunciation in the corresponding text segment.
[0094] S130: Detect whether the associated first text segment, second text segment, third text segment, first audio segment, second audio segment and third audio segment meet a splicing condition.
[0095] In an embodiment of the present application, in order to ensure the continuity of the spliced data, the embodiment of the present application performs splicing detection on multiple selected continuous segments of text and audio.
[0096] In an embodiment of the present invention, when the data are concatenated and input into the model, the data can be converted into a continuous data vector by "embedding" the discrete data so that the model can calculate it.
[0097] In some embodiments of the present application, for example, in the above step S130, it can be detected whether the three text segments or the three audio segments meet the first splicing condition, and the first splicing condition can be the overall length threshold of the three text segments or the three audio segments. Specifically, the combined text length of the first text segment, the second text segment, and the third text segment can be obtained, and the combined audio duration of the first audio segment, the second audio segment, and the third audio segment can be obtained. When the combined text length is less than the length threshold and the combined audio duration is less than the duration threshold, the text and speech corresponding to the first text segment, the second text segment, the third text segment, the first audio segment, the second audio segment, and the third audio segment meet the first splicing condition.
[0098] In a specific embodiment of the present application, continue to refer to Figure 3 , the length threshold can be set to 50 words, and the duration threshold can be set to 30 seconds. Among them, the second text segment is "I went to the river beach on holidays", the text length is 8 words, and the corresponding second audio duration is 3.5 seconds. The first text segment is "I saw many children flying kites", the text length is 10 words, and the corresponding first audio duration is 6.1 seconds. The third text segment is "long leads", the text length is 8 words, and the corresponding third audio duration is 5.0 seconds. Then the total text length is 8+10+8=26 words, which is less than the length threshold (50 words), and the total sample duration is 3.5+6.1+5.0=14.6 seconds, which is less than the duration threshold (30 seconds). Then the text and speech corresponding to the first text segment, the second text segment, the third text segment, the first audio segment, the second audio segment, and the third audio segment meet the first splicing condition.
[0099] In other embodiments of the present application, in order to ensure the coherence of the model output speech, the audio input to the training model needs to be spoken by the same person, for example, the voiceprint feature can be used to determine whether it is the same person speaking. In some embodiments of the present application, for example, in the above step S130, the voiceprint feature similarity between the first audio segment, the second audio segment, and the third audio segment can be obtained. When the voiceprint feature similarity is greater than the voiceprint similarity threshold, the text and speech corresponding to the first text segment, the second text segment, the third text segment, the first audio segment, the second audio segment, and the third audio segment meet the second splicing condition.
[0100] In some embodiments of the present application, the first audio segment, the second audio segment, and the third audio segment can be discretely sampled, and the time domain signal can be converted into a frequency domain signal to extract the frequency feature. According to the similarity of the frequency feature, it is determined whether the same person is speaking. The embodiment of the present application sets a voiceprint approximation threshold, which can be 80%, for example. When the voiceprint feature similarity between two audios is greater than 80%, the first audio, the second audio, and the third audio are the same person speaking. The text and speech corresponding to the first text segment, the second text segment, the third text segment, the first audio segment, the second audio segment, and the third audio segment meet the second splicing condition.
[0101] In some embodiments of the present application, in order to determine whether multiple audio segments are emitted by the same person, identification can also be based on an established voiceprint model (such as a Gaussian mixture model (GMM)) or a deep learning model (such as CNN and RNN).
[0102] In some embodiments of the present application, for example, in the above-mentioned step S130, in order to improve the coherence of the audio output by the speech synthesis model and avoid the audio segments selected for training being too far apart and lacking coherence characteristics, the embodiments of the present application also detect the intervals between the audio segments when detecting the splicing conditions. When the interval between the utterance time points of adjacent audio segments in the first audio segment, the second audio segment, and the third audio segment is less than the interval threshold, the text and speech corresponding to the first text segment, the second text segment, the third text segment, the first audio segment, the second audio segment, and the third audio segment meet the third splicing condition.
[0103] In the training sample generation process, when the speaker reads the text continuously, there will generally be a short pause after the sentence break. In order to ensure the listener's coherent experience, the time between the sentence breaks should not be too long, otherwise the listener experience will be poor. In the embodiment of the present application, the interval between the sounding time points of adjacent audio segments can be understood as the interval from the end of the sounding of the previous audio segment to the start of the sounding of the next audio segment. In the embodiment of the present application, based on the pause after the sentence break, there is often a blank sound period before each audio segment starts to sound, such as Figure 3The shaded portion in is a blank sound period, and the duration of the blank sound period is the interval between the sounding time points of adjacent audio segments. In some embodiments of the present application, the interval threshold may be selected to be 2 to 4 seconds.
[0104] In a specific embodiment of the present application, continue to refer to Figure 3 , the interval threshold can be set to 3 seconds, where the second text segment is "Go to the river beach on holidays", the text length is 8 words, the corresponding second audio duration is 3.5 seconds, and there is no blank sound period; the first text segment is "See many children flying kites", the text length is 10 words, the corresponding first audio duration is 6.1, and the blank sound period is 0.8 seconds; the third text segment is "One by one long lead", the text length is 8 words, the corresponding third audio duration is 5 seconds, and the blank sound period is 1.4 seconds. The interval between the utterance time points of the second audio segment and the first audio segment is 0.8 seconds, and the interval between the utterance time points of the first audio segment and the third audio segment is 1.4 seconds, both of which are less than the interval threshold, then the text and speech corresponding to the first text segment, the second text segment, the third text segment, the first audio segment, the second audio segment, and the third audio segment meet the third splicing condition.
[0105] In the embodiment of the present application, when detecting whether the splicing condition is met, the above detection conditions can be used alone or in combination.
[0106] S140: splicing the first text segment, the second text segment, the third text segment, the first audio segment, the second audio segment and the third audio segment that meet the splicing condition in a text order in an alternating manner of text and audio to obtain combined training data, so as to generate a combined training data set including a plurality of the combined training data.
[0107] In an embodiment of the present invention, when combined training data is obtained by splicing text and audio in an alternating manner, the text and audio data can be converted into continuous text and audio data vectors by "embedding" discrete data.
[0108] In some embodiments of the present application, Figure 4 As shown, the process of splicing the combined training data in an alternating manner of text and audio in the order of text includes:
[0109] S141: performing word segmentation processing on the first text segment, the second text segment and the third text segment to generate a first text sequence, a second text sequence and a third text sequence, and performing vector conversion on the first text sequence, the second text sequence and the third text sequence to generate a first text vector, a second text vector and a third text vector.
[0110] In the embodiments of the present application, in order to perform word segmentation on a text segment to generate a text sequence, a word segmentation algorithm can be used for processing. For example, a method based on string matching or a rule-based method (such as forward maximum matching word segmentation algorithm, reverse maximum matching word segmentation algorithm, statistical-based method (such as Jieba word segmentation, etc.)) can be used for word segmentation. In some embodiments of the present application, before the word segmentation, text normalization processing can also be performed on the text in the text segment, including converting the character set (such as uniformly using Unicode encoding), eliminating case differences, and removing irrelevant characters (such as HTML tags).
[0111] In some embodiments of the present application, the vector conversion process of the sequence can be processed by using embedding, and the embedding processing process may include:
[0112] a. Construct a vocabulary
[0113] Word frequency statistics: Calculate the frequency of each token appearing in the entire dataset.
[0114] Vocabulary screening: Screen out the words in the vocabulary based on frequency and relevance. Remove high-frequency stop words (such as "de", "he") and low-frequency words.
[0115] Index assignment: Assign a unique numerical index to each token in the vocabulary.
[0116] b. Token encoding
[0117] One-Hot encoding: Convert each token into a high-dimensional sparse vector, where only the index position corresponding to the token is 1, and the rest are 0.
[0118] Limit the vector size: To avoid too high a dimension, which may limit the size of the vocabulary, use special tokens to replace uncommon words, such as ` <unk>`(Unknown).
[0119] c. Learning Dense Vector Representation
[0120] Choose a model: Use pre-trained models like Word2Vec, GloVe, FastText, or BERT, or train a custom model for a specific task.
[0121] Training process: Learn the vector representation of a token through context. For example, in Word2Vec, this may be achieved by predicting the context of a word (CBOW model), or by using a word to predict its context (Skip-gram model).
[0122] Vector properties: These vectors capture complex relationships between words, such as semantic similarity, grammatical regularities, etc.
[0123] d. Vector space optimization
[0124] Dimensionality reduction: Dimensionality reduction is performed through techniques such as principal component analysis (PCA) or t-SNE to optimize the vector space so that similar words are closer in the vector space.
[0125] Refinement training: Further fine-tuning is performed on a specific task to optimize the performance of the word vector in that task.
[0126] S142: Discretize the first audio segment, the second audio segment, and the third audio segment and extract audio features to generate a first audio feature sequence, a second audio feature sequence, and a third audio feature sequence; perform vector conversion on the first audio feature sequence, the second audio feature sequence, and the third audio feature sequence to generate a first audio vector, a second audio vector, and a third audio vector.
[0127] In an embodiment of the present application, the audio segment may be discretized. Specifically, the frequency features may be extracted after sampling the audio segment, and the spectral features may be processed to generate an audio feature sequence. The audio features may include Mel spectrum, Mel spectrum cepstral coefficients, etc. In some embodiments of the present application, for example, in the above step S142, embedding may also be used to discretize the first audio segment, the second audio segment, and the third audio segment and extract audio features.
[0128] S143: Perform vector concatenation on the second text vector, the second audio feature sequence, the first text vector, the first audio feature sequence, the third text vector, and the third audio feature sequence in the order of the second text vector, the second audio feature sequence, the first text vector, the first audio feature sequence, the third text vector, and the third audio feature sequence to generate the combined training data.
[0129] In one embodiment of the present application, for example, the first text vector is T1, the second text vector is T2, the third text vector is T3, the first audio vector is A1, the second audio vector is A2, and the third audio vector is A3, then the generated combined training vector is {T2, A2, T1, A1, T3, A3}, and each vector is separated by a special symbol.
[0130] S150: Training an initial speech synthesis model according to the combined training data set to obtain a trained speech synthesis model.
[0131] In an embodiment of the present application, the prosody, rhythm and / or emotion features in the second text segment and / or the third text segment may be extracted during the training.
[0132] The embodiments of the present application can be trained by inputting the text and voice audio of the context into the model together. It can use the context of the text to understand the text and capture the rhythm and emotion in the scene. It can also understand the voice corresponding to the context and obtain the melody and emotion of the scene. The generated voice information finally outputted is more in line with the situation and can be smoothly connected with the voice of the context.
[0133] In an embodiment of the present application, for example, in the above-mentioned step S150, during the training process, the relevant rhythm, melody and / or emotional features in the second text segment and / or the third text segment are extracted and used for training the speech synthesis model, so that the trained speech synthesis model can integrate the context, that is, the above-mentioned features of the second text segment and / or the third text segment to control the rhythm of the output speech and control the emotion of the speech.
[0134] In the embodiments of the present application, the speech synthesis model includes a variety of speech synthesis models (such as Tacotron, WaveNet, Fastspeech, etc.). It can adapt to a variety of existing models without adjusting the model structure, and the method in the embodiments of the present application can be used, which has a high practical promotion and application value.
[0135] In some embodiments of the present application, Figure 5 As shown, the training of the initial speech synthesis model according to the target training vector set to generate a trained speech synthesis model may include:
[0136] S151a: iteratively inputting the combined training data in the combined training data set into the initial speech synthesis model to obtain a first output audio segment corresponding to the first text segment;
[0137] S152a: Calculate a first loss value according to the first output audio segment and the audio feature of the first audio segment;
[0138] S153a: Adjust the model parameters of the initial speech synthesis model according to the first loss value until the initial speech synthesis model converges, and use the converged initial speech synthesis model as the trained speech synthesis model.
[0139] In the embodiment of the present application, the loss values include but are not limited to Mel-Spectrogram Loss, DurationLoss, CTC Loss, Angular Margin Loss, etc. Mel-Spectrogram Loss is used to compare the Mel-spectrograms of the generated audio and the input audio to ensure the accuracy of the sound features; Duration Loss ensures that the predicted pronunciation duration is consistent with the actual duration for the attention-based model; CTC Loss handles misaligned input and output sequences for the speech recognition model; Angular Margin Loss is used to enhance the model's ability to identify speaker features in speaker verification and recognition tasks.
[0140] In some embodiments of the present application, the loss value can be calculated based on the middle audio segment among the second audio segment, the first audio segment, and the third audio segment, that is, the first audio segment and the corresponding first output audio segment, so that the front part of the text and speech and the back part of the text and speech can be fully utilized to fully understand the rhythm and emotion of the middle part of the text, and the model's understanding of semantics and scenes is more complete.
[0141] In some embodiments of the present application, the loss value can be calculated by using the first audio segment of the second audio segment, the first audio segment, and the third audio segment, and the emotion and scene of the latter speech and text can be used to train the speech synthesis model, so that the speech output is more coherent. Figure 6 As shown, the training process may include:
[0142] S151b: Input the combined training data in the combined training data set into the initial speech synthesis model to obtain a second output audio segment corresponding to the second text segment.
[0143] S152b: Calculate a second loss value according to the second output audio segment and the audio features of the second audio segment.
[0144] S153b: Adjust the model parameters of the initial speech synthesis model according to the second loss value, iteratively update until the initial speech synthesis model converges, and use the converged initial speech synthesis model as the trained speech synthesis model.
[0145] In some embodiments of the present application, the loss value can be calculated by using the latter audio segment of the second audio segment, the first audio segment, and the third audio segment, and the emotion and scene of the previous speech and text can be used to train the speech synthesis model, so that the speech output is more coherent. Figure 7 As shown, the training process may include:
[0146] S151c: Iteratively input the combined training data in the combined training data set into the initial speech synthesis model to obtain a third output audio segment corresponding to the third text segment.
[0147] S152c: Calculate a third loss value according to the third output audio segment and the audio features of the third audio segment.
[0148] S153c: Adjust the model parameters of the initial speech synthesis model according to the third loss value until the initial speech synthesis model converges, and use the converged initial speech synthesis model as the trained speech synthesis model.
[0149] In the embodiment of the present application, the target training vector is iteratively or cyclically taken out from the target training vector set and input into the initial speech synthesis model, the loss value is calculated, and the model parameters are adjusted until the model converges, that is, the training ends when the loss value is less than the set threshold.
[0150] like Fig. 9 As shown, the speech synthesis model 500 in the embodiment of the present application can be abstracted into an encoder 510 and a decoder 520. During the training process, the corresponding loss value is calculated, and the discriminator can also be set to perform automatic parameter adjustment training.
[0151] The initial speech synthesis model in the embodiment of the present application outputs audio segments including a first output audio segment, a second output audio segment, and a third output audio segment during the training process. According to the need to calculate the loss value, an audio segment involved in calculating the loss value is selected from the output audio segments.
[0152] In addition to the neural network part, the speech synthesis model in the embodiment of the present application may also include a frequency domain to time domain processing module. After the neural network part outputs the audio segment, it is converted into a sound signal through the frequency domain to time domain processing module.
[0153] During the operation of the speech synthesis model in this application, the text and audio sequences first pass through the embedding module, enter the encoder, and then are output to the decoder. Finally, the output audio segment is converted from the frequency domain to the time domain signal to generate the final sound signal.
[0154] In some embodiments of the present application, in order to further improve the accuracy of the speech output model output, the output audio segments of each segment are involved in the calculation process of the loss value to improve the degree of context understanding of the model output, such as Figure 8 As shown, the training process may further include:
[0155] S151d: Iteratively input the combined training data in the combined training data set into the initial speech synthesis model to obtain a first output audio segment corresponding to the first text segment, a second output audio segment corresponding to the second text segment, and a third output audio segment corresponding to the third text segment.
[0156] S152d: Calculate a first loss value based on the first output audio segment and the audio features of the first audio segment, calculate a second loss value based on the second output audio segment and the audio features of the second audio segment, calculate a third loss value based on the third output audio segment and the audio features of the third audio segment, and obtain a fourth loss value based on the first loss value, the second loss value and the third loss value.
[0157] In the process of calculating the fourth loss value in the embodiment of the present application, the weights of the first loss value, the second loss value and the third loss value can be set. For example, the first loss value is L1, the second loss value is L2, and the third loss value is L3, then the fourth loss value is L4 = a*L1+b*L2+c*L3, where a, b, c are the weights of each loss value.
[0158] In the process of calculating the fourth loss value, the embodiment of the present application can adjust the weights of each loss value according to the attention. For example, if the trained model generates the latter part of the speech more based on the text and speech in the front and middle, the weight of the third loss value setting becomes larger; if the middle part of the speech is generated more based on the text and speech in the front and back, the weight of the first loss value setting becomes larger; if the front part of the speech is generated more based on the text and speech in the middle and back, the weight of the second loss value setting becomes larger.
[0159] S153d: Adjust the model parameters of the initial speech synthesis model according to the fourth loss value until the initial speech synthesis model converges, and use the converged initial speech synthesis model as the trained speech synthesis model.
[0160] The target speech synthesis model in the embodiment of the present application can be a TTS model. During the training process, a large number of annotated speech-text pairs are used for supervised learning, and hyperparameters such as learning rate and batch size are adjusted. Techniques such as Dropout and BatchNormalization are applied to prevent overfitting. The model can also be carefully tuned and adjusted according to the model performance. For example, the performance on the validation set can be monitored and the model can be tuned. The depth of the model, the number of hidden units, and the attention mechanism can be adjusted.
[0161] The target speech synthesis model in the embodiment of the present application can be a large language model, which is trained using a large-scale speech data set during the training process, and can be trained by self-supervised learning, and the large language model can be adjusted to handle speech synthesis in different scenarios, such as classroom speech synthesis, recitation speech synthesis, or conversational speech synthesis. After the initial model is generated, the model can be evaluated and tuned, and an independent validation set can be used for performance evaluation, and model parameters such as the number of layers, the number of hidden units, the learning rate, etc. can be adjusted to optimize performance.
[0162] The speech synthesis model training method in the embodiment of the present application forms a sample group with the previous paragraph of the training text, the next paragraph of the text and the corresponding training speech data. After determining whether they can be spliced, the sample group is converted into a vector and input into the initial speech synthesis model, the initial speech synthesis model is trained, and a target speech synthesis model is generated. During the speech synthesis process, the generated target speech synthesis model can use the context of the text to understand the text and capture the rhythm and emotion in the scene. At the same time, it can also understand the speech corresponding to the context and obtain the rhythm and emotion of the scene, so that the generated speech information finally outputted is more in line with the situation and can be smoothly connected with the speech in the context.
[0163] like Fig.10 As shown, the embodiment of the present application provides a speech synthesis method, including:
[0164] S210: Acquire a target text, and divide the target text into a plurality of target text segments according to the text sequence.
[0165] In some embodiments of the present application, dividing the target text into multiple target text segments according to the text order includes: performing sentence segmentation processing on the target text to obtain multiple target text segments.
[0166] In a specific embodiment of the present application, the target text segment is, for example, "I went to the river beach on holidays and saw many children flying kites. Long strings were tied to the sky at one end and the ground at the other end. The children and the kites were swinging between the sky and the ground, and even my heart was swinging in a trance, as if I had returned to my childhood." According to sentence segmentation, the target text segment is divided into a first text segment T1 ("I went to the river beach on holidays"), a second text segment T2 ("I saw many children flying kites"), a third text segment T3 ("long strings"), a fourth text segment T4 ("one end tied to the sky"), a fifth text segment T5 ("one end tied to the ground"), a sixth text segment T6 ("the children and the kites were swinging between the sky and the ground"), a seventh text segment T7 ("even my heart was swinging in a trance"), and an eighth text segment T8 ("as if I had returned to my childhood").
[0167] S220: Inputting a first target text segment among the multiple target text segments into a trained target speech synthesis model to obtain a first target audio segment.
[0168] In one embodiment of the present application, for example, the first text segment T1 ("Go to the river beach on holidays") is input into the target speech synthesis model to obtain the first target audio segment A1.
[0169] In some embodiments of the present application, the trained target speech synthesis model is obtained by executing any training method of the present application.
[0170] In the embodiment of the present application, the first audio segment may not be generated using the target speech synthesis model. The sound emitted by the user after reading the first text segment may be recorded to generate the first audio segment so that subsequent speech synthesis can carry the characteristics of the user's voice, such as the user's emotions, accent and other information.
[0171] S230: Obtain a text-audio alternating sequence by splicing the first target text segment and the first target audio segment.
[0172] In one embodiment of the present application, for example, referring to the above example, a text-audio alternating sequence {T1, A1} is generated by splicing.
[0173] S240: Loop through the following steps until the last target audio segment corresponding to the last target text segment is obtained: obtain the next target text segment, and add the next target text segment to the end of the text-audio alternating sequence to update the text-audio alternating sequence; input the text-audio alternating sequence into the trained speech synthesis model to obtain the next target audio segment corresponding to the next target text segment, and add the next target audio segment to the text-audio alternating sequence to update the text-audio alternating sequence.
[0174] In an embodiment of the present application, for example, in the above-mentioned step S240, the next target audio segment corresponding to the next target text segment obtained extracts the rhythm, melody and / or emotional acoustic features of other target text segments. Among them, the other target text segments refer to text segments other than the currently corresponding next target text segment, that is, other text segments other than the text segment corresponding to the current audio segment for synthesizing. Preferably, the other text segments may be or include the previous target text segment of the next target text segment. In an embodiment of the present application, the rhythm of the next target audio segment can be controlled by the context, especially the rhythm, melody and / or emotional acoustic features of the "previous context", thereby controlling the emotion of the speech.
[0175] In one embodiment of the present application, for example, continuing with the above example, a second text segment is added to the text-audio alternating sequence {T1, A1}, the text-audio alternating sequence is updated to {T1, A1, T2}, the updated text-audio alternating sequence {T1, A1, T2} is input into the trained speech synthesis model to obtain the second target audio segment A2, and then based on the next text segment T3, the text-audio alternating sequence is updated to {T1, A1, T2, A2, T3}, and the trained speech synthesis model is input to obtain the corresponding third target audio segment A3, and the above steps are executed repeatedly until the eighth target audio segment A8 is generated.
[0176] In some embodiments of the present application, Fig.11 As shown, when the speech synthesis starts, text 1 is input, the speech model generates audio 1, and it is pieced into a text-audio sequence. Then, a new text segment is input in sequence, and it is pieced to the end of the sequence. Then, the entire sequence is input into the trained speech synthesis model to synthesize the next audio segment, thereby making full use of context information during synthesis. After audio n is generated, audio n is added to the text-audio sequence, and the text-audio sequence is updated to {text 1, audio 1, text 2...audio n}. When text n+1 is input, text n+1 is added to the text-audio sequence, and the text-audio sequence is updated to {text 1, audio 1, text 2...audio n, text n+1}. The updated text-audio sequence is input into the trained speech synthesis model to obtain audio n+1. The above steps are repeated until the audio corresponding to all text segments is generated.
[0177] S250: splicing the target audio segments to generate target audio.
[0178] In an embodiment of the present application, for example, in the above step S250, each target audio segment may be concatenated according to the text order to generate the target audio.
[0179] In one embodiment of the present application, for example, continuing with the above example, the target audio segments finally obtained are {A1, A2, A3, A4, A5, A6, A7, A8}, which are concatenated to generate the target audio A_target.
[0180] In another alternative embodiment of the present application, a speech synthesis method is also provided, in which speech is synthesized in reverse order of the text, and the synthesized speech can obtain the context, especially the rhythm, cadence and / or emotional acoustic features in the "below". Specifically, the method may include:
[0181] Acquire a target text, and divide the target text into a plurality of target text segments according to the text order;
[0182] Inputting a last target text segment among the multiple target text segments into a trained target speech synthesis model to obtain a last target audio segment;
[0183] splicing the last target text segment and the last target audio segment to obtain a text-audio alternating sequence;
[0184] The following steps are executed in a loop until a first target audio segment corresponding to a first target text segment is obtained: a previous target text segment is obtained, and the previous target text segment is added to the text-audio alternating sequence to update the text-audio alternating sequence; the text-audio alternating sequence is input into the trained target speech synthesis model to obtain a previous target audio segment corresponding to the previous target text segment, and the previous target audio segment is added to the end of the text-audio alternating sequence to update the text-audio alternating sequence;
[0185] Each target audio segment is spliced together to generate the target audio. In particular, each target audio segment can be spliced together in the order of the text to generate the target audio.
[0186] In an embodiment of the present application, for example, in the above embodiment, the previous target audio segment corresponding to the previous target text segment obtained extracts the rhythm, melody and / or emotional acoustic features of other target text segments. Among them, the other target text segments refer to text segments other than the current corresponding previous target text segment, that is, other text segments other than the text segment corresponding to the current audio segment used to synthesize. Preferably, the other text segments can be or include the next target text segment of the previous target text segment. In an embodiment of the present application, the rhythm of the previous target audio segment can be controlled by the context, especially the rhythm, melody and / or emotional acoustic features of the "below", thereby controlling the emotion of the speech.
[0187] In some embodiments of the present application, the step of splicing the target audio segments to generate the target audio further includes: performing smoothing processing at the connection points of the audio segments.
[0188] In an embodiment of the present application, when the target speech synthesis model is TTS (Text-to-Speech) or a large speech model, after the target audio segment is generated, smoothing processing can be performed when the target audio segments are spliced to ensure the auditory coherence and naturalness of the synthesized audio.
[0189] In the embodiment of the present application, the smoothing process is performed when the target audio segments are spliced as follows.
[0190] 1. Accurate selection of splicing points
[0191] Selection based on prosodic features: Automatically analyze the prosodic features of speech (pitch, intensity, rhythm, etc.) and select points that are similar in rhythm or naturally connected to each other as splicing points.
[0192] 2. Digital Signal Processing (DSP)
[0193] Time alignment and pitch correction: Use digital signal processing techniques such as WSOLA (Waveform Similarity Overlap-Add) to time-stretch / compress and pitch-match audio clips to ensure coherence at splice points.
[0194] 3. Crossfade
[0195] Smooth Transition: Implement crossfading at the splice point so that one clip gradually fades out while the other gradually increases, creating a smooth transition.
[0196] 4. Maintain spectrum continuity
[0197] Spectral matching: Use an equalizer or other spectral processing tool to ensure that the spectral characteristics of the audio before and after the splicing remain consistent, especially in the low and high frequency ranges.
[0198] 5. Harmonic and Resonant Peak Matching
[0199] Harmonic Matching: Analyze and match the harmonic structure before and after the splicing point to ensure timbre continuity.
[0200] Formant Matching Adjustment: Fine-tune formants to keep your sound consistent and natural.
[0201] 6. Ambient noise and background sound processing
[0202] Noise smoothing: If background noise is present, adjust the consistency of the noise through technical means, or use noise reduction technology to eliminate the noise.
[0203] 7. Post-mixing and mastering
[0204] Dynamic range processing: Use compressors and limiters to adjust dynamic range and ensure consistent volume across your audio.
[0205] Minor adjustments: Make minor adjustments and optimizations to the spliced audio, including equalization, reverb, etc.
[0206] 8. Quality inspection and feedback
[0207] Auditory inspection: Perform a detailed auditory inspection after splicing to ensure there are no obvious splicing marks.
[0208] Objective evaluation: Use objective evaluation tools (such as PESQ, Perceptual Evaluation of Speech Quality) to evaluate the quality of the spliced audio. Adjust the parameters of the previous steps based on the evaluation quality.
[0209] In the embodiment of the present application, during the speech synthesis process, the text and audio sequences first pass through the embedding module, enter the encoder, and then are output to the decoder. Finally, the output audio segment (expressed using frequency characteristics) is generated and converted from the frequency domain to the time domain signal to generate the final sound signal.
[0210] The speech recognition method in the embodiment of the present application divides the target text into multiple segments, first uses the target speech synthesis model to perform speech synthesis on the first segment of text, then combines the generated first audio segment with the first segment of text to form a text-audio sequence, and iteratively performs speech synthesis on subsequent text segments, adds the current text segment to be recognized to the text-audio sequence, obtains the corresponding audio segment, and also adds the latest obtained audio segment to the text-audio sequence to provide context text and audio for the next text segment, until the last text segment generates an audio segment, and then splices the generated audio segments to finally generate the target audio. In the speech synthesis process, the text understanding and speech understanding of the context are utilized, and the generated audio can be accompanied by the rhythm, melody and emotion of the current scene, and the speech synthesis is more in line with the situation.
[0211] like Fig.12 As shown, the embodiment of the present application provides a speech synthesis model training device 1200, including a training set data acquisition module 1210, a preprocessing module 1220, a splicing detection module 1230, a vector splicing module 1240 and a training module 1250, wherein:
[0212] The training set data acquisition module 1210 is configured to acquire training set data, wherein the training set data includes coherent text and corresponding coherent audio.
[0213] The preprocessing module 1220 is configured to select a plurality of first text segments, a plurality of second text segments located before the first text segment and adjacent to the first text segment, and a plurality of third text segments located after the first text segment and adjacent to the first text segment from the coherent text, and obtain a plurality of first audio segments, second audio segments, and third audio segments corresponding to the plurality of first text segments, second text segments, and third text segments, respectively, from the coherent audio.
[0214] The splicing detection module 1230 is configured to detect whether the texts and voices corresponding to the plurality of first text segments, second text segments, third text segments, first audio segments, second audio segments and third audio segments meet a splicing condition.
[0215] The vector splicing module 1240 is configured to splice a plurality of first text segments, a second text segment, a third text segment, a first audio segment, a second audio segment and a third audio segment that meet the splicing conditions in a text order and in an alternating manner of text and audio to generate a combined training data set, wherein the combined training data set includes a plurality of combined training data.
[0216] The training module 1250 is configured to train the initial speech synthesis model according to the combined training data set to obtain a trained speech synthesis model.
[0217] In some embodiments of the present application, the splicing detection module 1230 is further configured to:
[0218] Obtaining a combined text length of the first text segment, the second text segment, and the third text segment, and obtaining a combined audio duration of the first audio segment, the second audio segment, and the third audio segment; when the combined text length is less than a length threshold and the combined audio duration is less than a duration threshold, the text and speech corresponding to the first text segment, the second text segment, the third text segment, the first audio segment, the second audio segment, and the third audio segment meet a first splicing condition;
[0219] In some other embodiments of the present application, the splicing detection module 1230 is further configured to:
[0220] Acquire the voiceprint feature similarity between the first audio segment, the second audio segment, and the third audio segment. When the voiceprint feature similarity is greater than a voiceprint similarity threshold, the text and speech corresponding to the first text segment, the second text segment, the third text segment, the first audio segment, the second audio segment, and the third audio segment meet a second splicing condition.
[0221] In some other embodiments of the present application, the splicing detection module 1230 is further configured to:
[0222] When the interval between the utterance time points of adjacent audio segments in the first audio segment, the second audio segment and the third audio segment is less than the interval threshold, the text and speech corresponding to the first text segment, the second text segment, the third text segment, the first audio segment, the second audio segment and the third audio segment meet the third splicing condition.
[0223] In some embodiments of the present application, the vector stitching module 1240 is configured to:
[0224] Performing word segmentation processing on the first text segment, the second text segment, and the third text segment to generate a first text sequence, a second text sequence, and a third text sequence, and performing vector conversion on the first text sequence, the second text sequence, and the third text sequence to generate a first text vector, a second text vector, and a third text vector;
[0225] Discretize the first audio segment, the second audio segment, and the third audio segment and extract audio features to generate a first audio feature sequence, a second audio feature sequence, and a third audio feature sequence; perform vector conversion on the first audio feature sequence, the second audio feature sequence, and the third audio feature sequence to generate a first audio vector, a second audio vector, and a third audio vector;
[0226] The second text vector, the second audio feature sequence, the first text vector, the first audio feature sequence, the third text vector, and the third audio feature sequence are vector-concatenated in the order of the second text vector, the second audio feature sequence, the first text vector, the first audio feature sequence, the third text vector, and the third audio feature sequence to generate the combined training data.
[0227] In some embodiments of the present application, the target speech synthesis model includes a large language model.
[0228] In some embodiments of the present application, the training module 1250 is configured to:
[0229] Iteratively inputting the combined training data in the combined training data set into the initial speech synthesis model, obtaining a first output audio segment corresponding to the first text segment, calculating a first loss value according to the first output audio segment and the audio features of the first audio segment, adjusting a model parameter of the initial speech synthesis model according to the first loss value until the initial speech synthesis model converges, and using the converged initial speech synthesis model as the trained speech synthesis model;
[0230] In some other embodiments of the present application, the training module 1250 is further configured to:
[0231] Inputting the combined training data in the combined training data set into the initial speech synthesis model, obtaining a second output audio segment corresponding to the second text segment, calculating a second loss value according to the audio features of the second output audio segment and the second audio segment, adjusting the model parameters of the initial speech synthesis model according to the second loss value, iteratively updating until the initial speech synthesis model converges, and using the converged initial speech synthesis model as the trained speech synthesis model;
[0232] In some other embodiments of the present application, the training module 1250 is further configured to:
[0233] Iteratively input the combined training data in the combined training data set into the initial speech synthesis model, obtain a third output audio segment corresponding to the third text segment, calculate a third loss value based on the third output audio segment and the audio features of the third audio segment, adjust the model parameters of the initial speech synthesis model according to the third loss value until the initial speech synthesis model converges, and use the converged initial speech synthesis model as the trained speech synthesis model.
[0234] In some embodiments of the present application, the training module 1250 may also be configured as:
[0235] Iteratively input the combined training data in the combined training data set into the initial speech synthesis model to obtain a first output audio segment corresponding to the first text segment, a second output audio segment corresponding to the second text segment, and a third output audio segment corresponding to the third text segment; calculate a first loss value according to the audio features of the first output audio segment and the first audio segment, calculate a second loss value according to the audio features of the second output audio segment and the second audio segment, calculate a third loss value according to the audio features of the third output audio segment and the third audio segment, and obtain a fourth loss value according to the first loss value, the second loss value, and the third loss value; adjust the model parameters of the initial speech synthesis model according to the fourth loss value until the initial speech synthesis model converges, and use the converged initial speech synthesis model as the trained speech synthesis model.
[0236] like Fig.13 As shown, the embodiment of the present application provides a speech synthesis device 1300, including a target text acquisition module 1310, an initial generation module 1320, a text-audio splicing module 1330, a loop generation module 1340 and an audio splicing module 1350, wherein:
[0237] The target text acquisition module 410 is configured to acquire a target text and divide the target text into a plurality of target text segments according to the text order;
[0238] The initial generation module 1320 is configured to obtain a first target text segment from the multiple target text segments, input a vector corresponding to the first target text segment into a trained target speech synthesis model, and obtain a first target audio segment;
[0239] The text-audio splicing module 1330 is configured to generate a text-audio alternating sequence according to the splicing of the first target text segment and the first target audio segment;
[0240] The loop generation module 1340 is configured to loop to obtain the next target text segment, add the next target text segment to the end of the text-audio alternating sequence to update the text-audio alternating sequence; input the vector corresponding to the text-audio alternating sequence into the trained target speech synthesis model, obtain the next target audio segment, add the next target audio segment to the end of the text-audio alternating sequence to update the text-audio alternating sequence, until the last target audio segment corresponding to the last target text segment is obtained;
[0241] The audio splicing module 1350 is configured to splice each target audio segment to generate a target audio.
[0242] In some embodiments of the present application, the trained target speech synthesis model is obtained by executing the training method in any one of the embodiments.
[0243] In some embodiments of the present application, the target text acquisition module 1310 is configured to:
[0244] The target text is segmented to obtain a plurality of target text segments.
[0245] In some embodiments of the present application, the audio splicing module 1350 is further configured to:
[0246] Smooths the joins between audio segments.
[0247] Fig.14 A schematic diagram of an electronic device 1400 that can be used to implement the method of an embodiment of the present application or to implement an embodiment of the present application is shown. In some embodiments, the number of electronic devices may be more or less than the number shown. In some embodiments, a single or multiple electronic devices may be used for implementation. In some embodiments, cloud or distributed electronic devices may also be used for implementation.
[0248] like Fig.14 As shown, the electronic device 1400 includes a processor 1410 and a memory 1420. The processor is used to execute programs stored in the memory, which can implement the methods, steps or functions described in the above embodiments when executed by the computer. The processor 1410 may include various types of processors, such as a central processing unit (CPU), a graphics processing unit (GPU), a neural network processor (NPU), a digital signal processor (DSP), etc. The processor 1410 and the memory 1420 are interconnected via a bus 1430. The bus 1430 may also be connected to an input / output (I / O) interface, etc.
[0249] Although not shown in the figure, an embodiment of the present application also provides a storage medium, which stores computer programs, which are configured to execute speech synthesis model training and speech synthesis methods related to any embodiment at runtime.
[0250] It should be understood that all operations in the method described above are merely exemplary, and the present disclosure is not limited to any operation in the method or the order of these operations, but should cover all other equivalent transformations under the same or similar concept.
[0251] It should also be understood that all modules in the above described device can be implemented in various ways. These modules can be implemented as hardware, software, or a combination thereof. In addition, any module in these modules can be further divided into submodules or combined together functionally.
[0252] Processors have been described in conjunction with various devices and methods. These processors can be implemented using electronic hardware, computer software or any combination thereof. Whether these processors are implemented as hardware or software will depend on specific application and the overall design constraints imposed on the system. As an example, the processor provided in the present disclosure, any part of the processor or any combination of processors can be implemented as a microprocessor, a microcontroller, a digital signal processor (DSP), a field programmable gate array (FPGA), a programmable logic device (PLD), a state machine, a gate logic, a discrete hardware circuit, and other suitable processing components configured to perform the various functions described in the present disclosure. The function of the processor provided in the present disclosure, any part of the processor or any combination of processors can be implemented as software performed by a microprocessor, a microcontroller, a DSP or other suitable platforms.
[0253] Software should be broadly considered to represent instructions, instruction sets, codes, code segments, program codes, programs, subroutines, software modules, applications, software applications, software packages, routines, subroutines, objects, running threads, processes, functions, etc. Software can reside in a computer-readable medium. A computer-readable medium can include, for example, a memory, which can be, for example, a magnetic storage device (e.g., a hard disk, a floppy disk, a magnetic stripe), an optical disk, a smart card, a flash memory device, a random access memory (RAM), a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), a register, or a removable disk. Although the memory is shown as being separated from the processor in the various aspects provided in the present disclosure, the memory can also be located inside the processor (e.g., a cache or a register).
[0254] The above description is provided to enable any person skilled in the art to implement the various aspects described herein. Various modifications to these aspects are apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects. Therefore, the claims are not intended to be limited to the aspects shown herein. All structural and functional equivalents of elements of various aspects described in this disclosure that are known or to be known to those skilled in the art will be expressly included herein by reference and are intended to be covered by the claims.< / unk>
Claims
1. A speech synthesis method, characterized in that: The steps include: Acquire a target text, and divide the target text into a plurality of target text segments according to the text order; Inputting a last target text segment among the multiple target text segments into a trained speech synthesis model to obtain a last target audio segment; splicing the last target text segment and the last target audio segment to obtain a text-audio alternating sequence; The following steps are executed in a loop until a first target audio segment corresponding to a first target text segment is obtained: a previous target text segment is obtained, and the previous target text segment is added to the text-audio alternating sequence to update the text-audio alternating sequence; the text-audio alternating sequence is input into the trained speech synthesis model to obtain a previous target audio segment corresponding to the previous target text segment, and the previous target audio segment is added to the end of the text-audio alternating sequence to update the text-audio alternating sequence; Splicing each target audio segment to generate the target audio; The training process of the speech synthesis model includes the following steps: Acquire initial training data, wherein the initial training data includes coherent text and corresponding coherent audio; Selecting from the continuous text a plurality of first text segments, a plurality of second text segments located before the first text segment and adjacent to the first text segment, and a plurality of third text segments located after the first text segment and adjacent to the first text segment, and acquiring from the continuous audio a plurality of first audio segments, second audio segments, and third audio segments corresponding to the plurality of first text segments, second text segments, and third text segments, respectively; Detecting whether the associated first text segment, second text segment, third text segment, first audio segment, second audio segment, and third audio segment meet a splicing condition; splicing the first text segment, the second text segment, the third text segment, the first audio segment, the second audio segment and the third audio segment that meet the splicing condition in a text order in an alternating manner of text and audio to obtain combined training data, so as to generate a combined training data set including a plurality of the combined training data; According to the combined training data set, the initial speech synthesis model is trained to obtain a trained speech synthesis model. During the training, the rhythm, melody and / or emotional features in the second text segment and / or the third text segment are extracted and used for the training of the speech synthesis model.
2. The method according to claim 1, characterized in that The combined training data is obtained by splicing text and audio alternately in the order of text, including: Performing word segmentation processing on the first text segment, the second text segment, and the third text segment to generate a first text sequence, a second text sequence, and a third text sequence, and performing vector conversion on the first text sequence, the second text sequence, and the third text sequence to generate a first text vector, a second text vector, and a third text vector; Discretize the first audio segment, the second audio segment, and the third audio segment and extract audio features to generate a first audio feature sequence, a second audio feature sequence, and a third audio feature sequence; perform vector conversion on the first audio feature sequence, the second audio feature sequence, and the third audio feature sequence to generate a first audio vector, a second audio vector, and a third audio vector; The second text vector, the second audio feature sequence, the first text vector, the first audio feature sequence, the third text vector, and the third audio feature sequence are vector-concatenated in the order of the second text vector, the second audio feature sequence, the first text vector, the first audio feature sequence, the third text vector, and the third audio feature sequence to generate the combined training data.
3. The method according to claim 1, characterized in that The detecting whether the associated first text segment, the second text segment, the third text segment, the first audio segment, the second audio segment, and the third audio segment meet the splicing condition includes: The combined text length of the first text segment, the second text segment, and the third text segment is obtained, and the combined audio duration of the first audio segment, the second audio segment, and the third audio segment is obtained. When the combined text length is less than a length threshold and the combined audio duration is less than a duration threshold, the text and speech corresponding to the first text segment, the second text segment, the third text segment, the first audio segment, the second audio segment, and the third audio segment meet a first splicing condition.
4. The method according to claim 1, characterized in that The detecting whether the associated first text segment, the second text segment, the third text segment, the first audio segment, the second audio segment, and the third audio segment meet the splicing condition includes: The voiceprint feature similarity between the first audio segment, the second audio segment and the third audio segment is obtained. When the voiceprint feature similarity is greater than a voiceprint similarity threshold, the text and speech corresponding to the first text segment, the second text segment, the third text segment, the first audio segment, the second audio segment and the third audio segment meet a second splicing condition.
5. The method according to claim 1, characterized in that The detecting whether the associated first text segment, the second text segment, the third text segment, the first audio segment, the second audio segment, and the third audio segment meet the splicing condition includes: When the interval between the utterance time points of adjacent audio segments in the first audio segment, the second audio segment and the third audio segment is less than the interval threshold, the text and speech corresponding to the first text segment, the second text segment, the third text segment, the first audio segment, the second audio segment and the third audio segment meet the third splicing condition.
6. The method according to claim 1, characterized in that The step of training the initial speech synthesis model according to the combined training data set to obtain a trained speech synthesis model comprises: Iteratively input the combined training data in the combined training data set into the initial speech synthesis model, obtain a first output audio segment corresponding to the first text segment, calculate a first loss value based on the first output audio segment and the audio features of the first audio segment, adjust the model parameters of the initial speech synthesis model based on the first loss value until the initial speech synthesis model converges, and use the converged initial speech synthesis model as the trained speech synthesis model.
7. The method according to claim 1, characterized in that The step of training the initial speech synthesis model according to the combined training data set to obtain a trained speech synthesis model comprises: The combined training data in the combined training data set is input into the initial speech synthesis model, a second output audio segment corresponding to the second text segment is obtained, a second loss value is calculated according to the audio features of the second output audio segment and the second audio segment, model parameters of the initial speech synthesis model are adjusted according to the second loss value, and the initial speech synthesis model is iteratively updated until the initial speech synthesis model converges, and the converged initial speech synthesis model is used as the trained speech synthesis model.
8. The method according to claim 1, characterized in that The step of training the initial speech synthesis model according to the combined training data set to obtain a trained speech synthesis model comprises: Iteratively input the combined training data in the combined training data set into the initial speech synthesis model, obtain a third output audio segment corresponding to the third text segment, calculate a third loss value based on the third output audio segment and the audio features of the third audio segment, adjust the model parameters of the initial speech synthesis model according to the third loss value until the initial speech synthesis model converges, and use the converged initial speech synthesis model as the trained speech synthesis model.
9. The method according to claim 1, characterized in that: The step of training the initial speech synthesis model according to the combined training data set to obtain a trained speech synthesis model comprises: Iteratively input the combined training data in the combined training data set into the initial speech synthesis model to obtain a first output audio segment corresponding to the first text segment, a second output audio segment corresponding to the second text segment, and a third output audio segment corresponding to the third text segment; calculate a first loss value according to the audio features of the first output audio segment and the first audio segment, calculate a second loss value according to the audio features of the second output audio segment and the second audio segment, calculate a third loss value according to the audio features of the third output audio segment and the third audio segment, and obtain a fourth loss value according to the first loss value, the second loss value, and the third loss value; adjust the model parameters of the initial speech synthesis model according to the fourth loss value until the initial speech synthesis model converges, and use the converged initial speech synthesis model as the trained speech synthesis model.
10. The method according to claim 1, characterized in that The step of dividing the target text into a plurality of target text segments according to the text order comprises: The target text is segmented to obtain a plurality of target text segments.
11. The method according to claim 1, characterized in that: The step of splicing the target audio segments to generate the target audio further includes: Smooths the joins between audio segments.
12. A speech synthesis device, characterized in that: It includes a target text acquisition module, an initial generation module, a text-audio splicing module, a loop generation module and an audio splicing module, wherein: The target text acquisition module is configured to acquire a target text and divide the target text into a plurality of target text segments according to the text order; The initial generation module is configured to input the last target text segment among the multiple target text segments into the trained speech synthesis model to obtain the last target audio segment; The text-audio splicing module is configured to obtain a text-audio alternating sequence by splicing the last target text segment and the last target audio segment; The loop generation module is configured to cyclically obtain a previous target text segment, add the previous target text segment to the text-audio alternating sequence to update the text-audio alternating sequence; input the text-audio alternating sequence to the trained speech synthesis model to obtain a previous target audio segment corresponding to the previous target text segment, add the previous target audio segment to the end of the text-audio alternating sequence to update the text-audio alternating sequence, until a first target audio segment corresponding to a first target text segment is obtained; The audio splicing module is configured to splice each target audio segment to generate a target audio; The training process of the speech synthesis model includes the following steps: Acquire initial training data, wherein the initial training data includes coherent text and corresponding coherent audio; Selecting from the continuous text a plurality of first text segments, a plurality of second text segments located before the first text segment and adjacent to the first text segment, and a plurality of third text segments located after the first text segment and adjacent to the first text segment, and acquiring from the continuous audio a plurality of first audio segments, second audio segments, and third audio segments corresponding to the plurality of first text segments, second text segments, and third text segments, respectively; Detecting whether the associated first text segment, second text segment, third text segment, first audio segment, second audio segment, and third audio segment meet a splicing condition; splicing the first text segment, the second text segment, the third text segment, the first audio segment, the second audio segment and the third audio segment that meet the splicing condition in a text order in an alternating manner of text and audio to obtain combined training data, so as to generate a combined training data set including a plurality of the combined training data; According to the combined training data set, the initial speech synthesis model is trained to obtain a trained speech synthesis model. During the training, the rhythm, melody and / or emotional features in the second text segment and / or the third text segment are extracted and used for the training of the speech synthesis model.
Citation Information
Patent Citations
Speech synthesis model training method, speech synthesis method, electronic equipment and storage medium
CN118116364A
Speech synthesis method and device
CN119541457A