Speech synthesis model training method and device, computer equipment and storage medium

By using real acoustic markers and duration information in the speech synthesis model training, the sound loss and multitone problems that the speech synthesis model arise when converting semantic features are solved, and the accuracy of generating speech is improved.

CN120032622APending Publication Date: 2025-05-23GUANGZHOU QUYAN NETWORK TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510265519.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

Existing speech synthesis models may experience unstable problems such as sound loss or multitone in the process of converting semantic features into predicting acoustic semantic features, resulting in low accuracy in generating speech.

Method used

By obtaining the sample speech signal and its corresponding real acoustic marks and duration information, the sample speech signal and text data are input into the speech synthesis model to be trained. The speech synthesis model is trained to improve the accuracy of generating speech by using the difference between the predicted acoustic marks and the real acoustic marks, as well as the difference between the predicted duration information and the real time information.

Benefits of technology

By considering the duration information of real acoustic marks, the model can better learn to predict the alignment of phoneme duration and pause duration in acoustic marks, reduce the occurrence of unstable problems such as sound loss and polyphonics, and thus improve the accuracy of speech synthesis model generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032622A_ABST
    Figure CN120032622A_ABST
Patent Text Reader

Abstract

The invention relates to a speech synthesis model training method and device, computer equipment and a storage medium. The method comprises the following steps: acquiring a sample voice signal, and acquiring sample text data corresponding to the sample voice signal, and a real acoustic mark and real duration information corresponding to the sample voice signal; inputting the sample voice signal and the sample text data into a to-be-trained voice synthesis model, and obtaining a predicted acoustic mark and predicted duration information corresponding to the predicted acoustic mark through the voice synthesis model; and according to the difference between the predicted acoustic mark and the real acoustic mark and the difference between the predicted duration information and the real duration information, training a speech synthesis model to obtain a trained speech synthesis model. By adopting the method, unstable problems such as voice loss and multiple voices can be reduced, so that the voice generation accuracy of the voice synthesis model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of speech synthesis technology, and in particular to a speech synthesis model training method, apparatus, computer equipment, storage medium and computer program product. Background Art

[0002] With the development of speech synthesis technology, a technology for synthesizing speech using a speech synthesis model has emerged. The speech synthesis model can include a large language model. For example, the tag features can be first obtained by inputting text and corresponding speech signals, and then the tag features are input into the large language model module for converting acoustic tags. After the large language model outputs the predicted acoustic tags, the predicted acoustic semantic features are used to output the predicted speech signal corresponding to the input text.

[0003] However, in the above process, the large language model module of the conversion mark may have unstable problems such as missing sounds or polyphony when converting semantic features into predicted acoustic semantic features. Therefore, the accuracy of speech generated by the currently trained speech synthesis model is low. Summary of the invention

[0004] Based on this, it is necessary to provide a speech synthesis model training method, device, computer equipment, computer-readable storage medium and computer program product that can improve the accuracy of speech generated by the speech synthesis model in order to solve the above-mentioned technical problems.

[0005] In a first aspect, the present application provides a speech synthesis model training method, comprising:

[0006] Acquire a sample speech signal, and acquire sample text data corresponding to the sample speech signal, as well as a real acoustic marker and real duration information corresponding to the sample speech signal; the real duration information is used to characterize the phoneme duration of each phoneme contained in the real acoustic marker, and the pause duration of the real acoustic marker;

[0007] Inputting the sample speech signal and the sample text data into a speech synthesis model to be trained, and obtaining a predicted acoustic marker and predicted duration information corresponding to the predicted acoustic marker through the speech synthesis model; the predicted duration information is used to characterize the phoneme duration of each phoneme included in the predicted acoustic marker and the pause duration of the predicted acoustic marker;

[0008] The speech synthesis model is trained according to the difference between the predicted acoustic marker and the real acoustic marker, and the difference between the predicted duration information and the real duration information, so as to obtain a trained speech synthesis model.

[0009] In one embodiment, the speech synthesis model includes: a tag extraction module and a large language model module for converting acoustic tags; the step of inputting the sample speech signal and the sample text data into the speech synthesis model to be trained, and obtaining predicted acoustic tags and predicted duration information corresponding to the predicted acoustic tags through the speech synthesis model, includes: inputting the sample speech signal and the sample text data into the tag extraction module to obtain sample semantic tags and sample acoustic tags; inputting the sample acoustic tags and the sample semantic tags into the large language model module, and outputting the predicted acoustic tags and the predicted duration information corresponding to the predicted acoustic tags through the large language model module.

[0010] In one embodiment, the predicted duration information is represented by a coding sequence; the predicted acoustic label includes multiple sub-predicted acoustic labels, which correspond to different speech signal frames respectively; the sample acoustic label and the sample semantic label are input into the large language model module, and the predicted acoustic label and the predicted duration information corresponding to the predicted acoustic label are output through the large language model module, including: inputting the sample acoustic label and the sample semantic label into the large language model module to obtain the sub-predicted acoustic label corresponding to each speech signal frame, and the predicted acoustic label category code corresponding to each speech signal frame; the sub-predicted acoustic labels are combined into the predicted acoustic label in the order of each speech signal frame, and the predicted acoustic label category codes are arranged in the order of each speech signal frame to obtain a predicted acoustic label category code sequence, and the predicted acoustic label category code sequence is used as the predicted duration information.

[0011] In one of the embodiments, the real duration information is represented by a coding sequence; the real acoustic label and the real duration information corresponding to the sample speech signal are obtained in the following manner: the sample speech signal is input into a pre-constructed speech segmenter, and the sub-real acoustic labels corresponding to each speech signal frame and the real acoustic label category codes corresponding to each speech signal frame are obtained through the speech segmenter; the sub-real acoustic labels are combined into the real acoustic label in the order of the speech signal frames, and the real acoustic label category codes are arranged in the order of the speech signal frames to obtain a real acoustic label category code sequence, and the real acoustic label category code sequence is used as the real duration information.

[0012] In one of the embodiments, obtaining sample text data corresponding to the sample voice signal includes: obtaining phoneme information of each phoneme contained in the sample voice signal, and pause information corresponding to each pause contained in the sample voice signal; generating original text data corresponding to the sample voice signal based on each phoneme; and adding the pause information to the original text data to obtain the sample text data.

[0013] In one embodiment, the pause information includes the pause position of each of the pauses and the pause duration of each of the pauses; adding the pause information to the original text data to obtain the sample text data includes: obtaining a pause mark corresponding to each of the pauses according to the pause duration of each of the pauses; adding the pause mark corresponding to each of the pauses to the text position in the original text data that matches the pause position of each of the pauses to obtain the sample text data.

[0014] In a second aspect, the present application also provides a speech synthesis model training device, comprising:

[0015] A training data acquisition module, used to acquire a sample speech signal, and acquire sample text data corresponding to the sample speech signal, and real duration information corresponding to the sample speech signal; the real duration information is used to characterize the phoneme duration of each phoneme contained in the sample speech signal, and the pause duration of the sample speech signal;

[0016] A prediction information acquisition module, used for inputting the sample speech signal and the sample text data into a speech synthesis model to be trained, and obtaining a predicted speech signal and predicted duration information corresponding to the predicted speech signal through the speech synthesis model; the predicted duration information is used for representing the phoneme duration of each phoneme contained in the predicted speech signal, and the pause duration of the predicted speech signal;

[0017] The synthesis model training module is used to train the speech synthesis model according to the difference between the predicted speech signal and the sample speech signal, and the difference between the predicted duration information and the actual duration information, so as to obtain a trained speech synthesis model.

[0018] In a third aspect, the present application further provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the method described in any one of the embodiments of the first aspect when executing the computer program.

[0019] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the method described in any one of the embodiments of the first aspect.

[0020] In a fifth aspect, the present application further provides a computer program product, comprising a computer program, which, when executed by a processor, implements the steps of the method described in any one of the embodiments of the first aspect.

[0021] The speech synthesis model training method, apparatus, computer equipment, storage medium and computer program product described above obtain sample speech signals, and obtain sample text data corresponding to the sample speech signals, as well as real acoustic labels and real duration information corresponding to the sample speech signals; the real duration information is used to characterize the phoneme duration of each phoneme contained in the real acoustic label, and the pause duration of the real acoustic label; the sample speech signals and sample text data are input into the speech synthesis model to be trained, and the predicted acoustic labels and predicted duration information corresponding to the predicted acoustic labels are obtained through the speech synthesis model; the predicted duration information is used to characterize the phoneme duration of each phoneme contained in the predicted acoustic label, and the pause duration of the predicted acoustic label; the speech synthesis model is trained according to the difference between the predicted acoustic label and the real acoustic label, and the difference between the predicted duration information and the real duration information, so as to obtain a trained speech synthesis model. The present application collects sample speech signals and corresponding text data when training a speech synthesis model, and can also obtain the real acoustic markers corresponding to the sample speech signals, as well as the real duration information for characterizing the phoneme duration of each phoneme contained in the real acoustic marker and the pause duration. Therefore, when the sample speech signal and the sample text data are input into the speech synthesis model to be trained, the model can output the predicted acoustic marker and the corresponding predicted duration information for characterizing the phoneme duration of each phoneme contained in the predicted acoustic marker and the pause duration. The model is then trained using the difference between the predicted acoustic marker and the real acoustic marker, as well as the difference between the predicted duration information and the real duration information. Compared with the existing training method of the speech synthesis model, the present application further considers the difference between the real duration information of the real acoustic marker and the predicted duration information of the predicted acoustic marker. Therefore, the alignment of the phoneme duration and the pause duration in the predicted acoustic marker can be better learned, and the occurrence of unstable problems such as missing sounds and polyphony can be reduced, thereby improving the accuracy of speech generated by the speech synthesis model. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related technologies, the drawings required for use in the embodiments or the related technical descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0023] Figure 1 A flowchart of a speech synthesis model training method in one embodiment;

[0024] Figure 2 A schematic diagram of a process for obtaining predicted acoustic markers and predicted duration information in one embodiment;

[0025] Figure 3 A schematic diagram of a process for obtaining predicted duration information in one embodiment;

[0026] Figure 4 A schematic diagram of a process for obtaining sample text data in one embodiment;

[0027] Figure 5 A schematic diagram of the structure of a text-to-speech system model in an embodiment;

[0028] Figure 6 is a structural block diagram of a speech synthesis model training device in one embodiment;

[0029] Figure 7 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0030] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0031] In one embodiment, Figure 1 As shown, a speech synthesis model training method is provided. This embodiment uses the method applied to a server as an example for illustration. It can be understood that the method can also be applied to a terminal, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0032] Step S101, obtain a sample speech signal, and obtain sample text data corresponding to the sample speech signal, as well as real acoustic markers and real duration information corresponding to the sample speech signal; the real duration information is used to characterize the phoneme duration of each phoneme contained in the real acoustic marker, and the pause duration of the real acoustic marker.

[0033] Among them, the sample voice signal is the voice signal used to train the speech synthesis model, and the sample text data is the text content data corresponding to the sample voice signal. For example, the sample voice signal can be the voice of "The weather is very good today", then the sample text data corresponding to the sample voice signal can be the text information recording "The weather is very good today".

[0034] The real acoustic marker can be the acoustic marker extracted from the sample speech signal, that is, the Acoustic token information extracted from the sample speech signal, and the real duration information is the speech duration corresponding to the real acoustic marker, which can be used to represent the phoneme duration of each phoneme contained in the real acoustic marker, as well as the pause duration contained in the real acoustic marker.

[0035] For example, a sample speech signal may be “Today, the weather is very good”, then the speech signal is composed of 6 phonemes and a pause between today and weather, and the real duration information may be the phoneme duration used to characterize each phoneme and the pause duration of the pause.

[0036] Specifically, when the server is training a speech synthesis model, it can first collect sample speech signals for training, and obtain text data for training based on the sample speech signals as sample text data, and can further extract the acoustic marker corresponding to the sample speech signal as a real acoustic marker, as well as extract the phoneme duration of each phoneme contained in the real acoustic marker, and the real duration information of the pause duration of the real acoustic marker.

[0037] Step S102, input the sample speech signal and the sample text data into the speech synthesis model to be trained, and obtain the predicted acoustic marker and the predicted duration information corresponding to the predicted acoustic marker through the speech synthesis model; the predicted duration information is used to characterize the phoneme duration of each phoneme contained in the predicted acoustic marker, as well as the pause duration of the predicted acoustic marker.

[0038] The speech synthesis model to be trained refers to a speech synthesis model that needs to be trained. The model can be used to output a predicted speech signal. The predicted speech signal can be obtained through a corresponding acoustic marker, that is, a predicted acoustic marker, and the predicted duration information is the speech duration corresponding to the predicted acoustic marker. Similar to the real duration information, the predicted duration information can be used to characterize the phoneme duration of each phoneme contained in the predicted acoustic marker, as well as the real duration information of the pause duration of the predicted acoustic marker.

[0039] Specifically, the speech synthesis model to be trained may also first generate a predicted acoustic marker, and then use the predicted acoustic marker to generate a predicted speech signal. When the server performs model training, it may first input the collected sample speech signal and sample text data into the speech synthesis model to be trained, thereby obtaining the predicted acoustic marker and the predicted duration information corresponding to the predicted acoustic marker through the model.

[0040] Step S103, training a speech synthesis model according to the difference between the predicted acoustic marker and the actual acoustic marker, and the difference between the predicted duration information and the actual duration information, to obtain a trained speech synthesis model.

[0041] Finally, the server can use the difference between the predicted acoustic markers and the actual acoustic markers, as well as the difference between the predicted duration information and the actual duration information, to construct a loss function to train the speech synthesis model, so that the speech synthesis model can not only learn the difference between the acoustic markers, but also learn the difference in the duration corresponding to the acoustic markers. In this way, the output of the acoustic markers can be better learned, and the alignment of the phoneme duration and pause duration in the acoustic markers can be ensured. Therefore, the occurrence of unstable problems such as dropped sounds, polyphony, and truncation can be reduced, thereby improving the accuracy of speech generated by the speech synthesis model.

[0042] In the above-mentioned speech synthesis model training method, a sample speech signal is obtained, and sample text data corresponding to the sample speech signal, as well as real acoustic markers and real duration information corresponding to the sample speech signal are obtained; the real duration information is used to characterize the phoneme duration of each phoneme contained in the real acoustic marker, and the pause duration of the real acoustic marker; the sample speech signal and the sample text data are input into the speech synthesis model to be trained, and the predicted acoustic marker and the predicted duration information corresponding to the predicted acoustic marker are obtained through the speech synthesis model; the predicted duration information is used to characterize the phoneme duration of each phoneme contained in the predicted acoustic marker, and the pause duration of the predicted acoustic marker; the speech synthesis model is trained according to the difference between the predicted acoustic marker and the real acoustic marker, and the difference between the predicted duration information and the real duration information, so as to obtain a trained speech synthesis model. The present application collects sample speech signals and corresponding text data when training a speech synthesis model, and can also obtain the real acoustic markers corresponding to the sample speech signals, as well as the real duration information for characterizing the phoneme duration of each phoneme contained in the real acoustic marker and the pause duration. Therefore, when the sample speech signal and the sample text data are input into the speech synthesis model to be trained, the model can output the predicted acoustic marker and the corresponding predicted duration information for characterizing the phoneme duration of each phoneme contained in the predicted acoustic marker and the pause duration. The model is then trained using the difference between the predicted acoustic marker and the real acoustic marker, as well as the difference between the predicted duration information and the real duration information. Compared with the existing training method of the speech synthesis model, the present application further considers the difference between the real duration information of the real acoustic marker and the predicted duration information of the predicted acoustic marker. Therefore, the alignment of the phoneme duration and the pause duration in the predicted acoustic marker can be better learned, and the occurrence of unstable problems such as missing sounds and polyphony can be reduced, thereby improving the accuracy of speech generated by the speech synthesis model.

[0043] Furthermore, if Figure 2 As shown, step S102 may further include:

[0044] Step S201: input the sample speech signal and the sample text data into a tag extraction module to obtain a sample semantic tag and a sample acoustic tag.

[0045] In this embodiment, the speech synthesis model can be composed of two parts, namely a module for extracting the markup token of the input information, namely the markup extraction module, and a large language model module for converting the markup token into an acoustic markup, wherein the sample semantic markup refers to the markup token corresponding to the sample text data, which can be the Semantic token corresponding to the sample text data, and the sample acoustic markup is the markup token converted from the sample speech signal, for example, it can be the Acoustic token converted from the sample speech signal.

[0046] Specifically, after the server obtains the sample speech signal and the sample text data, it can first input the sample speech signal and the sample text data into a tag extraction module included in the speech synthesis model, and the tag extraction module extracts the sample semantic tag and the sample acoustic tag.

[0047] For example, the tag extraction module may include a text encoder and a speech segmenter, wherein the text encoder is used to generate semantic tags and the speech segmenter is used to generate acoustic tags. The server may input sample text data into the text encoder and input sample speech signals into the speech segmenter, thereby outputting sample semantic tags through the text encoder and outputting sample acoustic tags through the speech segmenter.

[0048] Step S202: input the sample acoustic tag and the sample semantic tag into the large language model module, and output the predicted acoustic tag and the predicted duration information corresponding to the predicted acoustic tag through the large language model module.

[0049] The large language model module can adopt the large language model of text2token, which can be used to convert the input token into a predicted acoustic token. After the server obtains the sample semantic token and the sample acoustic token through the token extraction module, the sample semantic token and the sample acoustic token can be input into the large language model module of text2token, so that the large language model module outputs the predicted acoustic token and the duration information corresponding to the predicted acoustic token, that is, the predicted duration information.

[0050] In this embodiment, the speech synthesis model may include a tag extraction module and a large language model module for converting acoustic tags, wherein the tag extraction module is used to extract semantic tags and acoustic tags, and the large language model module is used to output predicted acoustic tags and predicted duration information. In this way, the accuracy of the output of predicted acoustic tags and predicted duration information can be improved.

[0051] Furthermore, the prediction duration information is represented by a coding sequence; the prediction acoustic marker includes multiple sub-prediction acoustic markers, each corresponding to different speech signal frames; Figure 3 As shown, step S202 may further include:

[0052] Step S301: Input the sample acoustic label and the sample semantic label into the large language model module to obtain the sub-prediction acoustic label corresponding to each speech signal frame and the prediction acoustic label category code corresponding to each speech signal frame.

[0053] In this embodiment, the predicted acoustic marker can be composed of multiple sub-predicted acoustic markers, and each sub-predicted acoustic marker corresponds to a different speech signal frame. For example, the predicted acoustic marker can be composed of sub-predicted acoustic marker 1 corresponding to speech signal frame 1, sub-predicted acoustic marker 2 corresponding to speech signal frame 2, and sub-predicted acoustic marker 3 corresponding to speech signal frame 3. The predicted acoustic marker category code is used to characterize the phoneme category code corresponding to each sub-predicted acoustic marker. For example, if the category code of sub-predicted acoustic marker 1 is 0, it means that sub-predicted acoustic marker 1 belongs to the first phoneme, and the category code of sub-predicted acoustic marker 2 is 1, it means that sub-predicted acoustic marker 2 belongs to the second phoneme, and so on.

[0054] Specifically, after the server inputs the sample acoustic label and the sample semantic label into the large language model module, it can output a predicted acoustic label composed of multiple sub-predicted acoustic labels corresponding to each speech signal frame, and a predicted acoustic label category code corresponding to each speech signal frame.

[0055] For example, the output of the large model text-to-token is: (token _id1, 0), (token_id2, 0), (token_id3, 0), (token_id4, 1), (token_id5, 1), (token_id6, 2)..., where token _id1 in (token _id1, 0) indicates that the sub-predicted acoustic marker corresponding to the first speech signal frame is token _id1, and 0 indicates that the predicted acoustic marker category corresponding to the first speech signal frame is encoded as code 0, which can represent that the acoustic marker belongs to the first phoneme. Similarly, token _id2 in (token_id2, 0) indicates that the sub-predicted acoustic marker corresponding to the second speech signal frame is token _id2, and 0 indicates that the predicted acoustic marker category corresponding to the second speech signal frame is encoded as code 0, that is, the acoustic marker also belongs to the first phoneme. (token_id4, 1) indicates that the sub-predicted acoustic marker corresponding to the fourth speech signal frame is token _id4, and 1 means that the predicted acoustic marker category code corresponding to the fourth speech signal frame is code 1, that is, the acoustic marker belongs to the second phoneme. From the above predicted acoustic marker category code, it can be known that the predicted acoustic marker category codes corresponding to the first three speech signal frames are the same, that is, the first three speech signal frames all belong to the first phoneme, that is, the first phoneme lasts for a total of 3 speech signal frames. Similarly, the predicted acoustic marker category codes corresponding to the 4th and 5th speech signal frames are the same, that is, the two speech signal frames 4 and 5 belong to the second phoneme, that is, the second phoneme lasts for a total of 2 speech signal frames.

[0056] Step S302, compose each sub-prediction acoustic marker into a predicted acoustic marker in the order of each speech signal frame, and arrange each predicted acoustic marker category code in the order of each speech signal frame to obtain a predicted acoustic marker category code sequence, and use the predicted acoustic marker category code sequence as the predicted duration information.

[0057] After obtaining each sub-prediction acoustic marker, the sub-prediction acoustic markers can be combined into a predicted acoustic marker in the order of each speech signal frame. At the same time, the predicted acoustic marker category codes are arranged in the order of each speech signal frame to obtain a predicted acoustic marker category code sequence, which is used as the predicted duration information.

[0058] Taking the output of the large model text-to-token as: (token _id1, 0), (token_id2, 0), (token_id3, 0), (token_id4, 1), (token_id5, 1), (token_id6, 2)... as an example, the output predicted acoustic tag category coding sequence is 000112..., which means that the first phoneme lasts for 3 frames, and the second phoneme lasts for 2 frames. Since the frame length of each speech signal frame is fixed, taking the 50HZ model as an example, the length of a frame is 20ms, then it can be known that the duration of the first phoneme is 60ms, and the duration of the second phoneme is 40ms. Therefore, the predicted duration information can be represented by the coding sequence.

[0059] In this embodiment, the large language model module can output the sub-prediction acoustic markers corresponding to each speech signal frame and the prediction acoustic marker category codes corresponding to each speech signal frame, so that the sub-prediction acoustic markers are combined into a prediction acoustic marker in the order of the speech signal frames, and the prediction acoustic marker category codes are arranged at the same time to obtain a prediction acoustic marker category code sequence, which is used as the prediction duration information. In this way, the accuracy of the output of the prediction acoustic marker and the prediction duration information can be improved.

[0060] In addition, the real duration information is represented by a coding sequence; step S101 may further include: inputting the sample speech signal into a pre-constructed speech segmenter, obtaining the sub-real acoustic markers corresponding to each speech signal frame and the real acoustic marker category codes corresponding to each speech signal frame through the speech segmenter; composing the sub-real acoustic markers into a real acoustic marker in the order of each speech signal frame, and arranging the real acoustic marker category codes in the order of each speech signal frame to obtain a real acoustic marker category code sequence, and using the real acoustic marker category code sequence as the real duration information.

[0061] The pre-built speech segmenter can be a pre-trained speech segmenter, which can be used to extract the acoustic marker corresponding to the sample speech signal, that is, the real acoustic marker, which is similar to the predicted acoustic marker. In this embodiment, the real acoustic marker can also be composed of sub-real acoustic markers corresponding to multiple speech signal frames. At the same time, the real duration information is also similar to the predicted duration information, and both are represented by means of a coding sequence. The coding sequence can be composed of real acoustic marker category codes corresponding to different speech signal frames.

[0062] Specifically, after obtaining the sample speech signal, the server can also input the sample speech signal into a pre-built speech segmenter, and the speech segmenter can output the sub-real acoustic markers corresponding to each speech signal frame, and the real acoustic marker category code used to characterize the phoneme category corresponding to each sub-real acoustic marker. Afterwards, the server can compose the real acoustic markers according to the order of the speech signal frames, and arrange the real acoustic marker category codes according to the order of the speech signal frames to obtain a real acoustic marker category code sequence, so as to use the real acoustic marker category code sequence as the real duration information.

[0063] In this embodiment, a pre-built speech segmenter can be used to output the sub-real acoustic markers and the real acoustic marker category codes corresponding to each speech signal frame in the sample speech signal, so that the real acoustic markers are composed of the sub-real acoustic markers, and the real acoustic marker category codes are arranged in the order of each speech signal frame to obtain the real acoustic marker category code sequence as the real duration information. This method can also improve the accuracy of the output of the real acoustic marker and the real duration information.

[0064] In one embodiment, Figure 4 As shown, step S101 may further include:

[0065] Step S401: Acquire phoneme information of each phoneme contained in the sample speech signal and pause information corresponding to each pause contained in the sample speech signal.

[0066] In this embodiment, the voice signal may be composed of phonemes and pauses that may exist between phonemes. For example, the sample voice signal may be: "The weather is very good today", where there is a pause between "today" and "the weather is very good", and the pause information is used to describe the relevant information of the pause in the above voice signal. Specifically, after obtaining the sample voice signal, the server can obtain the phoneme information of each phoneme contained in the sample voice signal, and the pause information corresponding to each pause, for example, the phoneme information and pause information of the sample voice signal can be obtained by ASR or MFA.

[0067] Step S402, generating original text data corresponding to the sample speech signal according to each phoneme information;

[0068] Step S403: adding the pause information to the original text data to obtain sample text data.

[0069] The original text data is the text data directly generated by each phoneme information. After the server obtains the phoneme information of each phoneme, it can use the above phoneme information to generate the initial text data as the original text data. The original text data only contains the text itself, and does not carry the pause-related information. After that, the server adds the pause information to the original text data before using it as sample text data. Therefore, the currently generated sample text data can carry pause information. Therefore, when using the sample text data for model inference, the generation of random pauses can be reduced by controlling the silent segments, thereby improving the accuracy of the voice signal output.

[0070] In this embodiment, when obtaining sample text data, the phoneme information of each phoneme contained in the sample speech signal and the pause information of the pauses can be first obtained, so as to generate original text data using the phoneme information, and then add the pause information to the original text data, so that the sample text data carries the pause information. In this way, when using the sample text data for model inference, the generation of random pauses can be reduced by controlling the silent segments, thereby improving the accuracy of the speech signal output.

[0071] Furthermore, the pause information includes the pause position of each pause and the pause duration of each pause; step S403 may further include: obtaining a pause mark corresponding to each pause according to the pause duration of each pause; adding the pause mark corresponding to each pause to the text position in the original text data that matches the pause position of each pause to obtain sample text data.

[0072] In this embodiment, the pause information of each pause may include two parts, namely, the position of each pause in the speech signal, i.e., the pause position, and the duration corresponding to each pause, i.e., the pause duration. For example, for the sample speech signal "Today is a good weather", if there is a pause between "Today" and "The weather is good", then the pause position is between "Today" and "The weather is good", and if there is a pause of 60ms between "Today" and "The weather is good", then the pause duration is 60ms.

[0073] The pause marker is identification information used to characterize the pause duration, and the identification information can be matched with the pause duration. In this embodiment, after obtaining the stop loss information, the pause marker used for the pause can be determined based on the pause duration, and then the pause marker can be added to the original text data to obtain sample text data.

[0074] For example, the original text data generated according to the phoneme information is "Today the weather is very good", and the pause information indicates that there is a pause of 60ms between "Today" and "The weather is very good", that is, the pause duration is 60ms. In this case, if the pause duration represented by the pause mark comma is set in advance to 60ms, then a comma can be added between "Today" and "The weather is very good" in the original text data, thereby generating the final sample text data "Today, the weather is very good".

[0075] In this embodiment, a pause mark can be obtained according to the pause duration of the pause, and the pause mark can be added to the text position matching the original text data with each pause position, so as to obtain sample text data. In this way, the accuracy of obtaining sample text data can be improved.

[0076] In one embodiment, a text-to-speech system based on a large language model is also provided, which can be applied to Figure 5 In the text-to-speech system model shown in the figure, S represents the input start mark, x-vec is the timbre information, the Text text obtains the Semantic token information through BPE encoding, that is, the semantic tag, T represents the text-to-speech mark, the speech is converted into the Acoustic token information through the Speech tokenizer module, that is, the acoustic tag, S+Semantic token+T+Acoustic token are jointly used as the input of the large model Text-to-token LM, and the output is the Acoustic token information predicted by the model, that is, the predicted acoustic tag.

[0077] In order to reduce the random pause problem that is prone to occur after model training, this embodiment obtains the phoneme duration and pause information of the training audio through ASR or MFA, and divides the pause information into 8 levels according to 0~60ms, 60~120ms, 120~180ms, 180~240ms, 240~300ms, 300~360ms, 360~420ms and greater than 420ms, and combines it into the text in the form of punctuation marks. Since the punctuation marks of the text and the pause positions of the actual audio may not correspond one to one, the punctuation marks here need to be based on the duration information as the standard. For example, the text corresponding to the audio is: "The weather is very good today, I want to go out and play." The punctuation marks judged according to the pause duration of the audio may be: "Today, the weather is very good and I want to go out and play." This involves a rhythm problem, that is, everyone's way of speaking may be different. Regarding the punctuation mark, it can be defined as follows: a colon ',' represents 60ms, two colons represent 120ms, and so on; or different lengths can be defined as different punctuation marks.

[0078] As for unstable issues such as lost sound and multiple sounds, Figure 5 In the example, the Acoustic token information predicted by the model is converted into a mel spectrum after Flow Match, and the mel spectrum is hifigan to generate PCM audio information. The AcousticToken information and the mel spectrum information have a one-to-one correspondence in the time dimension. Therefore, it is inferred that the causes of unstable problems such as lost sounds and polyphony may be related to the Acoustic Token information generated by the large model Text-to-token, that is, lost sounds may be caused by a certain input character not predicting the corresponding Acoustic token information, and polyphony may be caused by a certain input character predicting multiple identical Acoustic token information. Therefore, this embodiment modifies the output of the large model text-to-token into two parts, one part represents the Acoustic token information, and the other part is the duration information of the text corresponding to the Acoustic token. For example, taking the 50HZ model as an example, the length of one frame is 20ms. Based on the phoneme duration information of the training audio obtained by ASR or MFA, the following input information can be obtained: ", [3] The weather [2] is [4] very good [3] today [3],

[12] I [2] want [4] to go out [2] and play [3].

[15] ". The information in brackets is the audio duration corresponding to the Chinese characters and punctuation marks. 3 represents 3 frames, i.e. 60ms, and punctuation marks represent the silent part.

[0079] Acoustic token has 4096 states, represented by 0~4095. For the 50HZ model, the duration of each acoustic token is fixed at 20ms. Therefore, we can perform duration encoding on the input text, starting from 0 and increasing the sequence number according to the length of the text. Take the sentence ", [3] today [2] day [3] weather [2] is [4] good [3]" as an example. ", [3]" represents "000", "today [2]" represents "11", and "day [3]" represents "222", so the final duration information of this text is output as "00011222333445555666". This sequence just corresponds to the duration information of the Acoustic token, that is, the output of the large model text-to-token is: (token _id1, 0), (token_id2, 0), (token_id3, 0), (token_id4, 1), (token_id5, 1), (token_id6, 2)... Finally, this embodiment calculates the loss between the duration prediction information generated by the model and the actual duration information, so as to better fit the model in the correspondence between the Semantic token and the Acoustic token, and reduce the occurrence of unstable problems such as sound loss, polyphony, and truncation.

[0080] In this embodiment, the phoneme duration of the training data is obtained through forced alignment methods such as ASR and MFA, and the phoneme duration information is correspondingly encoded, and the corresponding loss is performed with the Acoustic token information generated by the text2token of the large model part to obtain the residual, thereby forcing the model to better fit the alignment information of the Semantic token and the Acoustic token, reducing unstable situations such as lost sound and multiple sounds, and at the same time, the silent segment information of the training data is obtained through VAD, ASR, MFA, etc., and after classifying the silent segments by their length, they are encoded into Semantic token information, merged with the Semantic token information of the text, and input into the model for training. During reasoning, the control of silent segments can reduce the generation of random pauses and improve the user's listening experience.

[0081] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0082] Based on the same inventive concept, the embodiment of the present application also provides a speech synthesis model training device for implementing the speech synthesis model training method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more speech synthesis model training device embodiments provided below can refer to the limitations on the speech synthesis model training method above, and will not be repeated here.

[0083] In one embodiment, Figure 6 As shown, a speech synthesis model training device is provided, including: a training data acquisition module 601, a prediction information acquisition module 602 and a synthesis model training module 603, wherein:

[0084] The training data acquisition module 601 is used to acquire a sample speech signal, and acquire sample text data corresponding to the sample speech signal, and real duration information corresponding to the sample speech signal; the real duration information is used to characterize the phoneme duration of each phoneme contained in the sample speech signal, and the pause duration of the sample speech signal;

[0085] The prediction information acquisition module 602 is used to input the sample speech signal and the sample text data into the speech synthesis model to be trained, and obtain the predicted speech signal and the prediction duration information corresponding to the predicted speech signal through the speech synthesis model; the prediction duration information is used to characterize the phoneme duration of each phoneme contained in the predicted speech signal and the pause duration of the predicted speech signal;

[0086] The synthesis model training module 603 is used to train the speech synthesis model according to the difference between the predicted speech signal and the sample speech signal, and the difference between the predicted duration information and the actual duration information, so as to obtain a trained speech synthesis model.

[0087] In one embodiment, the speech synthesis model includes: a tag extraction module and a large language model module for converting acoustic tags; a prediction information acquisition module 602, further used to input the sample speech signal and the sample text data into the tag extraction module to obtain sample semantic tags and sample acoustic tags; input the sample acoustic tags and the sample semantic tags into the large language model module, and output the predicted acoustic tags and the predicted duration information corresponding to the predicted acoustic tags through the large language model module.

[0088] In one embodiment, the predicted duration information is represented by a coding sequence; the predicted acoustic label includes multiple sub-predicted acoustic labels, which correspond to different speech signal frames respectively; the prediction information acquisition module 602 is further used to input the sample acoustic label and the sample semantic label into the large language model module to obtain the sub-predicted acoustic label corresponding to each speech signal frame, and the predicted acoustic label category code corresponding to each speech signal frame; the sub-predicted acoustic labels are composed of the predicted acoustic label in the order of each speech signal frame, and the predicted acoustic label category codes are arranged in the order of each speech signal frame to obtain a predicted acoustic label category code sequence, and the predicted acoustic label category code sequence is used as the predicted duration information.

[0089] In one embodiment, the real duration information is represented by a coding sequence; the training data acquisition module 601 is further used to input the sample speech signal into a pre-constructed speech segmenter, and obtain the sub-real acoustic labels corresponding to each speech signal frame and the real acoustic label category code corresponding to each speech signal frame through the speech segmenter; the sub-real acoustic labels are composed of the real acoustic labels in the order of the speech signal frames, and the real acoustic label category codes are arranged in the order of the speech signal frames to obtain a real acoustic label category code sequence, and the real acoustic label category code sequence is used as the real duration information.

[0090] In one embodiment, the training data acquisition module 601 is further used to obtain the phoneme information of each phoneme contained in the sample speech signal, and the pause information corresponding to each pause contained in the sample speech signal; generate the original text data corresponding to the sample speech signal according to the phoneme information; add the pause information to the original text data to obtain the sample text data.

[0091] In one embodiment, the pause information includes the pause position of each of the pauses and the pause duration of each of the pauses; the training data acquisition module 601 is further used to obtain the pause mark corresponding to each of the pauses according to the pause duration of each of the pauses; the pause mark corresponding to each of the pauses is added to the text position in the original text data that matches the pause position of each of the pauses to obtain the sample text data.

[0092] Each module in the above-mentioned speech synthesis model training device can be implemented in whole or in part by software, hardware and their combination. Each of the above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.

[0093] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 7 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store sample voice signals. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a speech synthesis model training method is implemented.

[0094] Those skilled in the art will understand that Figure 7 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0095] In one embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above method embodiments when executing the computer program.

[0096] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0097] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0098] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0099] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but are not limited to this.

[0100] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0101] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A speech synthesis model training method, characterized in that: The method comprises: Acquire a sample speech signal, and acquire sample text data corresponding to the sample speech signal, as well as a real acoustic marker and real duration information corresponding to the sample speech signal; the real duration information is used to characterize the phoneme duration of each phoneme contained in the real acoustic marker, and the pause duration of the real acoustic marker; Inputting the sample speech signal and the sample text data into a speech synthesis model to be trained, and obtaining a predicted acoustic marker and predicted duration information corresponding to the predicted acoustic marker through the speech synthesis model; the predicted duration information is used to characterize the phoneme duration of each phoneme included in the predicted acoustic marker and the pause duration of the predicted acoustic marker; The speech synthesis model is trained according to the difference between the predicted acoustic marker and the real acoustic marker, and the difference between the predicted duration information and the real duration information, so as to obtain a trained speech synthesis model.

2. The method according to claim 1, characterized in that The speech synthesis model includes: a tag extraction module and a large language model module for converting acoustic tags; the sample speech signal and the sample text data are input into the speech synthesis model to be trained, and the predicted acoustic tags and the predicted duration information corresponding to the predicted acoustic tags are obtained through the speech synthesis model, including: Inputting the sample speech signal and the sample text data into the tag extraction module to obtain a sample semantic tag and a sample acoustic tag; The sample acoustic label and the sample semantic label are input into the large language model module, and the predicted acoustic label and the predicted duration information corresponding to the predicted acoustic label are output through the large language model module.

3. The method according to claim 2, characterized in that The predicted duration information is represented by a coding sequence; the predicted acoustic marker includes a plurality of sub-predicted acoustic markers, each corresponding to a different speech signal frame; The step of inputting the sample acoustic tag and the sample semantic tag into the large language model module, and outputting the predicted acoustic tag and the predicted duration information corresponding to the predicted acoustic tag through the large language model module includes: Inputting the sample acoustic label and the sample semantic label into the large language model module to obtain the sub-prediction acoustic label corresponding to each speech signal frame and the prediction acoustic label category code corresponding to each speech signal frame; The sub-prediction acoustic labels are combined into the prediction acoustic label in the order of the speech signal frames, and the prediction acoustic label category codes are arranged in the order of the speech signal frames to obtain a prediction acoustic label category code sequence, and the prediction acoustic label category code sequence is used as the prediction duration information.

4. The method according to claim 3, characterized in that The real duration information is represented by a coding sequence; the real acoustic marker and the real duration information corresponding to the sample speech signal are obtained by: Inputting the sample speech signal into a pre-built speech segmenter, and obtaining, through the speech segmenter, sub-real acoustic labels corresponding to each speech signal frame, and real acoustic label category codes corresponding to each speech signal frame; The sub-real acoustic labels are combined into the real acoustic label in the order of the speech signal frames, and the real acoustic label category codes are arranged in the order of the speech signal frames to obtain a real acoustic label category code sequence, and the real acoustic label category code sequence is used as the real duration information.

5. The method according to any one of claims 1 to 4, characterized in that: The obtaining of sample text data corresponding to the sample voice signal includes: Acquire phoneme information of each phoneme contained in the sample speech signal, and pause information corresponding to each pause contained in the sample speech signal; Generating original text data corresponding to the sample speech signal according to each of the phoneme information; The pause information is added to the original text data to obtain the sample text data.

6. The method according to claim 5, characterized in that The pause information includes the pause position of each pause and the pause duration of each pause; the adding of the pause information into the original text data to obtain the sample text data includes: According to the pause duration of each of the pauses, obtaining a pause identifier corresponding to each of the pauses; The pause marks corresponding to the pauses are added to the text positions in the original text data that match the pause positions of the pauses, so as to obtain the sample text data.

7. A speech synthesis model training device, characterized in that: The device comprises: A training data acquisition module, used to acquire a sample speech signal, and acquire sample text data corresponding to the sample speech signal, and real duration information corresponding to the sample speech signal; the real duration information is used to characterize the phoneme duration of each phoneme contained in the sample speech signal, and the pause duration of the sample speech signal; A prediction information acquisition module, used for inputting the sample speech signal and the sample text data into a speech synthesis model to be trained, and obtaining a predicted speech signal and predicted duration information corresponding to the predicted speech signal through the speech synthesis model; the predicted duration information is used for representing the phoneme duration of each phoneme contained in the predicted speech signal, and the pause duration of the predicted speech signal; The synthesis model training module is used to train the speech synthesis model according to the difference between the predicted speech signal and the sample speech signal, and the difference between the predicted duration information and the actual duration information, so as to obtain a trained speech synthesis model.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Cited By

  • Speech synthesis method and device

    CN120564695A