Front-end model generation method and device, front-end processing method and device, equipment and medium

By generating a second training sample containing linguistic feature information of multiple preset dimension samples and using this sample to train a multi-task front-end model, the problem of mismatch in the training data of multi-task model and large difference in data volume in the prior art is solved, and more efficient model training and lower manual annotation cost are achieved.

CN120164447APending Publication Date: 2025-06-17BEIJING XIAOMI MOBILE SOFTWARE CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311725140.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-14
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

In the prior art, the multi-task front-end model has a large difference in training data of different preset dimensions during training, resulting in poor performance on each task and relies on manual annotation, which increases cost and workload.

Method used

By obtaining a first training sample of a plurality of preset dimensions, a second training sample is generated using a pre-trained benchmark model, which contains sample linguistic feature information of a plurality of preset dimensions. Then, the second training sample is trained to generate a multi-task front-end model, and the manual annotation workload is reduced through automatic annotation.

Benefits of technology

It effectively solves the problem of large differences in training data in different preset dimensions, improves the accuracy of multi-task front-end models, reduces manual labeling costs, and improves the generation efficiency of speech synthesis front-end models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164447A_ABST
    Figure CN120164447A_ABST
Patent Text Reader

Abstract

The invention relates to a speech synthesis front-end model generation method, a speech synthesis front-end processing method, a speech synthesis front-end processing device, speech synthesis front-end processing equipment and a medium. The generation efficiency of a speech synthesis front-end model is improved. The method comprises the steps of obtaining first training samples corresponding to a plurality of preset dimensions, wherein the first training sample corresponding to each preset dimension comprises a first sample text of the preset dimension and sample linguistic feature information of the preset dimension corresponding to the first sample text; according to a pre-trained first reference model and the first training sample, a second training sample is generated, and the second training sample comprises first sample texts of multiple preset dimensions and sample linguistic feature information of multiple preset dimensions corresponding to each first sample text; and training a pre-trained second reference model by using the second training sample to generate a multi-task front-end model, the speech synthesis front-end model including the multi-task front-end model, and the model scale of the first reference model being greater than that of the second reference model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of natural language processing, and particularly to a method for generating a front-end model for speech synthesis, a front-end processing method for speech synthesis, an apparatus, a device, and a medium. Background Art

[0002] Speech synthesis technology, also known as text-to-speech (TTS) technology, is a technology that can convert any text into audible speech. Speech synthesis technology is widely used in many fields such as intelligent voice assistants, accessibility technologies, entertainment, and education. In a speech synthesis task, it generally includes two parts: a front end and a back end. The speech synthesis front end mainly analyzes the text and extracts the linguistic feature information required by the back-end module. The speech synthesis back end then generates a speech waveform through a certain method based on the analysis result of the speech synthesis front end.

[0003] In recent years, with the development of deep learning and neural network technologies, the speech synthesis front-end framework is a multi-task learning framework based on a neural network model, and uses the neural network model to perform front-end processing of speech synthesis, improving the naturalness and accuracy of speech synthesis. Summary of the Invention

[0004] To overcome the problems existing in the related art, the present disclosure provides a method for generating a front-end model for speech synthesis, a front-end processing method for speech synthesis, an apparatus, a device, and a medium.

[0005] According to a first aspect of an embodiment of the present disclosure, there is provided a method for generating a front-end model for speech synthesis, the generating method including:

[0006] Obtaining first training samples corresponding to respective multiple preset dimensions, where each first training sample corresponding to a preset dimension includes a first sample text of the preset dimension and sample linguistic feature information of the first sample text corresponding to the preset dimension, and the sample linguistic feature information of the preset dimension refers to the feature information of the first sample text described from the preset dimension and required for synthesizing speech corresponding to the first sample text in a speech synthesis back end;

[0007] Generating second training samples according to a pre-trained first benchmark model and the first training samples, the second training samples including the first sample texts of the multiple preset dimensions and the sample linguistic feature information of each first sample text corresponding to the multiple preset dimensions;

[0008] Training the pre-trained second reference model using the second training sample to generate a multi-task front-end model, the speech synthesis front-end model includes the multi-task front-end model, and the model scale of the first reference model is larger than that of the second reference model.

[0009] According to a second aspect of the embodiments of the present disclosure, there is provided a method for processing a speech synthesis front-end, the processing method including:

[0010] Obtaining a target text to be converted;

[0011] Inputting the target text to be converted into the multi-task front-end model to obtain linguistic feature information of multiple dimensions corresponding to the target text output by the multi-task front-end model. The multi-task front-end model is generated according to the method for generating a speech synthesis front-end model described in the first aspect of the present disclosure. The linguistic feature information of multiple dimensions is used to input into a speech synthesis back-end model to generate speech corresponding to the target text.

[0012] According to a third aspect of the embodiments of the present disclosure, there is provided a device for generating a speech synthesis front-end model, the generating device including:

[0013] A first obtaining module, configured to obtain first training samples corresponding to respective multiple preset dimensions. Wherein, the first training sample corresponding to each preset dimension includes the first sample text of the dimension and the sample linguistic feature information of the dimension corresponding to the first sample text. Wherein, the sample linguistic feature information of the preset dimension refers to the feature information of the first sample text described from the preset dimension and required when synthesizing speech corresponding to the first sample text in a speech synthesis back-end.

[0014] A first generating module, configured to generate second training samples according to a pre-trained first reference model and the first training samples. The second training samples include the first sample texts of the multiple preset dimensions and the sample linguistic feature information of multiple dimensions corresponding to each first sample text.

[0015] A first training module, configured to train a pre-trained second reference model using the second training samples to generate a multi-task front-end model. The speech synthesis front-end model includes the multi-task front-end model, and the model scale of the first reference model is larger than that of the second reference model.

[0016] According to a fourth aspect of the embodiments of the present disclosure, there is provided a device for processing a speech synthesis front-end, the processing device including:

[0017] A second obtaining module, configured to obtain a target text to be converted;

[0018] An input module, configured to input the target text to be converted into a multi-task front-end model, and obtain linguistic feature information of multiple dimensions corresponding to the target text output by the multi-task front-end model. The multi-task front-end model is generated according to the method for generating a speech synthesis front-end model described in the first aspect of the present disclosure. The linguistic feature information of multiple dimensions is used to input a speech synthesis back-end model to generate speech corresponding to the target text.

[0019] According to a fifth aspect of the embodiments of the present disclosure, there is provided an electronic device, including:

[0020] A processor;

[0021] A memory for storing instructions executable by the processor;

[0022] Wherein, the processor is configured to execute the executable instructions to implement the steps of the method for generating a speech synthesis front-end model provided in the first aspect of the present disclosure, and / or to implement the steps of the speech synthesis front-end processing method provided in the second aspect of the present disclosure.

[0023] According to a sixth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, on which computer program instructions are stored. When the program instructions are executed by a processor, the steps of the method for generating a speech synthesis front-end model provided in the first aspect of the present disclosure, and / or the steps of the speech synthesis front-end processing method provided in the second aspect of the present disclosure are implemented.

[0024] By adopting the above technical solution, a second training sample is automatically generated according to a pre-trained first benchmark model and a first training sample. The second training sample includes first sample texts of multiple preset dimensions and sample linguistic feature information of multiple preset dimensions corresponding to each first sample text. In this way, each first sample text has sample linguistic feature information of multiple preset dimensions, effectively solving the problem that the training data of different preset dimensions vary greatly when training a multi-task front-end model. And, since each first sample text has sample linguistic feature information of multiple preset dimensions, when training with each first sample text, all tasks will be optimized as much as possible, alleviating the seesaw phenomenon to a certain extent, enabling the multi-task model to learn better effects for each task, and improving the accuracy of the generated multi-task front-end model. In addition, the second training sample for training the multi-task front-end model is obtained by automatic annotation, reducing the manual annotation workload and cost, and improving the generation efficiency of the speech synthesis front-end model.

[0025] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure and, together with the specification, are used to explain the principles of the present disclosure.

[0027] Figure 1 It is a flowchart of a voice synthesis shown according to an exemplary embodiment.

[0028] Figure 2 It is a flowchart of a method for generating a voice synthesis front-end model shown according to an exemplary embodiment.

[0029] Figure 3 It is a schematic diagram of a method for generating a voice synthesis front-end model shown according to an exemplary embodiment.

[0030] Figure 4 It is a flowchart of another method for generating a voice synthesis front-end model shown according to an exemplary embodiment.

[0031] Figure 5 It is a schematic diagram of another method for generating a voice synthesis front-end model shown according to an exemplary embodiment.

[0032] Figure 6 It is a flowchart of a method for generating a text regularization model shown according to an exemplary embodiment.

[0033] Figure 7 It is a schematic diagram of a method for generating a text regularization model shown according to an exemplary embodiment.

[0034] Figure 8 It is a flowchart of a voice synthesis front-end processing method shown according to an exemplary embodiment.

[0035] Figure 9 It is a schematic diagram of a voice synthesis front-end processing method shown according to an exemplary embodiment.

[0036] Figure 10 It is a flowchart of a method for generating a target text shown according to an exemplary embodiment.

[0037] Figure 11 It is a flowchart of another method for generating a target text shown according to an exemplary embodiment.

[0038] Figure 12 It is a block diagram of a device for generating a voice synthesis front-end model shown according to an exemplary embodiment.

[0039] Figure 13 It is a block diagram of a voice synthesis front-end processing device shown according to an exemplary embodiment.

[0040] Figure 14It is a block diagram of an electronic device shown according to an exemplary embodiment. Detailed implementation manners

[0041] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0042] It should be noted that all actions of obtaining signals, information, or data in the present disclosure are carried out on the premise of complying with the corresponding data protection regulations and policies of the country where the device is located and obtaining the authorization given by the owner of the corresponding device.

[0043] In recent years, with the development of deep learning and neural network technologies, especially the rapid development of pre-trained language models, pre-trained models can be used to generate speech synthesis front-end models. Among them, the advantage of pre-trained language models is that they can make full use of a large amount of unlabeled text data for pre-training, and fine-tune the pre-trained multi-task model on specific speech synthesis front-end tasks (such as phoneme conversion, prosody analysis, etc.) to make it adapt to specific tasks. It should be understood that for English synthesis front-ends, multi-tasks refer to tasks such as part-of-speech prediction, pitch accent prediction, phrase accent prediction, and boundary tones. For Chinese synthesis front-ends, multi-tasks refer to tasks such as word segmentation part-of-speech prediction, prosody structure prediction, and polyphonic characters.

[0044] In the related art, the multi-task front-end model is mostly generated in the following ways. Method 1: Use BERT (Bidirectional Encoder Representations from Transformers) and TinyBERT as the basic encoders, and divide the output of BERT into a word segmentation part-of-speech prediction layer, a prosody structure prediction layer, and a polyphonic character prediction layer to achieve multi-task prediction. During training, first perform joint fine-tuning of multiple sub-tasks of the front-end on the general BERT model, and each task has its own training set. After training the BERT teacher model, then use the BERT teacher model to distill the TinyBERT student model, and use the TinyBERT student model as the final multi-task front-end model. Method 2: Divide the front-end tasks into two categories: disambiguation tasks and sequence labeling tasks. Among them, polyphonic character prediction and text regularization are disambiguation tasks, and prosody structure prediction and word segmentation part-of-speech prediction are sequence labeling tasks. The two types of tasks share the classification layer respectively. After the text passes through the pre-trained RoBERTa, the encoding is obtained and input into the corresponding prediction layer, and the label of the task is output through the mask corresponding to the task.

[0045] In the related art, since the training data for each task is independent, there is a large gap in the amount of training data between different tasks. Moreover, when training models with independent databases between tasks, negative transfer and the seesaw phenomenon may occur. Especially when the training data of different tasks is independent of each other, there will be a more serious seesaw phenomenon, and the multi-task model cannot learn the best effect of each task. In addition, the solutions in the related art completely rely on manual annotation. Since the data annotation formats and contents of each task are different and require relatively professional linguistic knowledge, it leads to a very large annotation cost, and issues such as the amount of annotation also need to be considered.

[0046] In view of this, the present disclosure provides a method for generating a speech synthesis front-end model, a speech synthesis front-end processing method, an apparatus, a device, and a medium to solve the problems of mismatched training data annotation and large differences in the amount of data in the multi-task model in the related art. In addition, through automatic annotation, the manual annotation workload and cost are reduced, and the generation efficiency of the speech synthesis front-end model is improved.

[0047] First, the process of speech synthesis will be described. Figure 1 is a flowchart of a speech synthesis shown according to an exemplary embodiment. As Figure 1 shown, first, the text is input into the front-end model to obtain the linguistic feature information of each multi-task or multiple dimensions output by the front-end model. For example, the multi-tasks may include a word segmentation and part-of-speech prediction task, a prosody structure prediction task, and a polyphone task. Correspondingly, the multiple dimensions may include a word segmentation and part-of-speech dimension, a prosody structure dimension, and a polyphone dimension.

[0048] The linguistic feature information of each dimension may include multiple linguistic categories. The linguistic categories included in the linguistic feature information of each dimension may be determined according to the types of linguistic characteristics of that dimension. For example, the linguistic feature information of the prosody structure dimension may include syllables, prosodic words, prosodic phrases, etc.

[0049] Next, the linguistic feature information of each of the multiple dimensions is input into the acoustic model to obtain an intermediate representation. Finally, the intermediate representation is input into the vocoder to obtain the speech output by the vocoder. Among them, the intermediate representation may be a Mel spectrogram.

[0050] Figure 2 is a flowchart of a method for generating a speech synthesis front-end model shown according to an exemplary embodiment. The speech synthesis front-end model may be the front-end model shown as Figure 1 shown. As Figure 2 shown, the method for generating the speech synthesis front-end model may include the following steps.

[0051] In step S21, first training samples corresponding to each of the multiple preset dimensions are obtained.

[0052] In the present disclosure, the first training sample corresponding to each preset dimension includes the first sample text of the preset dimension and the sample linguistic feature information of the preset dimension corresponding to the first sample text.

[0053] The sample linguistic feature information of the preset dimension refers to the feature information describing the text from this dimension, which can synthesize the speech corresponding to the text at the back-end of speech synthesis. Exemplarily, the multiple preset dimensions may include a word segmentation part-of-speech dimension, a prosody structure dimension, and a polyphonic character dimension. That is to say, in step S21, the first training sample corresponding to the word segmentation part-of-speech dimension is obtained, and the first training sample includes the first sample text of the word segmentation part-of-speech dimension and the sample linguistic feature information of the word segmentation part-of-speech dimension corresponding to the first sample text. The first training sample corresponding to the prosody structure dimension is obtained, and the first training sample includes the first sample text of the prosody structure dimension and the sample linguistic feature information of the prosody structure dimension corresponding to the first sample text. And, the first training sample corresponding to the polyphonic character dimension is obtained, and the first training sample includes the first sample text of the polyphonic character dimension and the sample linguistic feature information of the polyphonic character dimension corresponding to the first sample text.

[0054] It should be understood that the number of the first sample texts of each preset dimension may be the same or different, and the first sample texts of each preset dimension may be the same text or different texts, and the present disclosure does not limit this.

[0055] In step S22, according to the pre-trained first benchmark model and the first training sample, a second training sample is generated.

[0056] The second training sample includes the first sample texts of multiple preset dimensions and the sample linguistic feature information of multiple preset dimensions corresponding to each first sample text.

[0057] In the present disclosure, according to the pre-trained first benchmark model and the first training sample, the sample linguistic feature information of multiple preset dimensions corresponding to each first sample text is automatically generated. That is to say, in the second training sample, each first sample text corresponds to the sample linguistic feature information of multiple preset dimensions. For example, the first sample text corresponds to the sample linguistic feature information of the word segmentation part-of-speech dimension, the sample linguistic feature information of the prosody structure dimension, and the sample linguistic feature information of the polyphonic character dimension.

[0058] In step S23, the pre-trained second benchmark model is trained using the second training sample to generate a multi-task front-end model, and the speech synthesis front-end model includes the multi-task front-end model.

[0059] Among them, the first reference model can be a large pre-trained Chinese BERT model, and the second reference model can be a small pre-trained Chinese BERT model. That is to say, the model scale of the second reference model is smaller than that of the first reference model. In this way, a small multi-task front-end model can be obtained, and the model scale of the generated multi-task front-end model can be reduced.

[0060] Adopting the above technical solution, the second training sample is automatically generated according to the pre-trained first reference model and the first training sample. The second training sample includes first sample texts of multiple preset dimensions and sample linguistic feature information of multiple preset dimensions corresponding to each first sample text. In this way, each first sample text has sample linguistic feature information of multiple preset dimensions, effectively solving the problem that the training data of different preset dimensions vary greatly when training the multi-task front-end model. And because each first sample text has sample linguistic feature information of multiple preset dimensions, when each first sample text is used for training, all tasks will be optimized as much as possible, alleviating the seesaw phenomenon to a certain extent, enabling the multi-task model to learn better effects for each task, and improving the accuracy of the generated multi-task front-end model. In addition, the second training sample obtained by automatic annotation is used to train the multi-task front-end model, reducing the manual annotation workload and cost, and improving the generation efficiency of the speech synthesis front-end model.

[0061] It should be understood that training the pre-trained second reference model with the second training sample to generate the multi-task front-end model means obtaining it by knowledge distillation of the pre-trained second reference model. Knowledge distillation can be regarded as model compression, where a trained, larger model (i.e., the second reference model in the present disclosure) is gradually used to teach a smaller model (i.e., the multi-task front-end model in the present disclosure) what to do. By attempting to replicate the output of each layer of the larger model (not just the final output), the smaller model is trained to learn the accurate behavior of the larger model, enabling the smaller model to also possess the accuracy of the larger model.

[0062] In this way, through knowledge distillation of the second reference model, on the basis of not losing the accuracy of the original model, the compression of the original model scale is realized, so that the obtained multi-task front-end model not only has the excellent reasoning ability of the reference model but also has a smaller model scale. In the process of using the speech synthesis front-end model to process the text to extract linguistic features from the text, the reasoning speed can be effectively improved, the reasoning time can be reduced, the text processing efficiency can be improved, and then the speech synthesis efficiency can be improved, and the response speed can be improved.

[0063] To enable those skilled in the art to better understand the method for generating the speech synthesis front-end model provided by the present disclosure, the above steps will be described in detail with examples below.

[0064] In one embodiment, Figure 2 Step S22 in generating the second training sample according to the pre-trained first reference model and the first training sample may include: First, for each preset dimension, taking the first sample text of the preset dimension as the model input parameter, and taking the sample linguistic feature information of the preset dimension corresponding to the first sample text as the model output parameter, training the pre-trained first reference model to generate a teacher model corresponding to the preset dimension.

[0065] Figure 3 is a schematic diagram of a method for generating a speech synthesis front-end model shown according to an exemplary embodiment. As Figure 3 shown, taking the first sample text of the word segmentation part-of-speech dimension in the first training sample corresponding to the word segmentation part-of-speech dimension as the model input parameter, and taking the sample linguistic feature information of the word segmentation part-of-speech dimension corresponding to the first sample text as the model output parameter, training or fine-tuning the large pre-trained Chinese BERT model to obtain a teacher model corresponding to the word segmentation part-of-speech dimension, simply referred to as the word segmentation part-of-speech teacher model. Taking the first sample text of the prosodic structure dimension as the model input parameter, and taking the sample linguistic feature information of the prosodic structure dimension corresponding to the first sample text in the first training sample corresponding to the prosodic structure dimension as the model output parameter, training or fine-tuning the large pre-trained Chinese BERT model to obtain a teacher model corresponding to the prosodic structure dimension, simply referred to as the prosodic structure teacher model. Similarly, taking the first sample text of the polyphone dimension in the first training sample corresponding to the polyphone dimension as the model input parameter, and taking the sample linguistic feature information of the polyphone dimension corresponding to the first sample text as the model output parameter, training or fine-tuning the large pre-trained Chinese BERT model to obtain a teacher model corresponding to the polyphone dimension, simply referred to as the polyphone teacher model.

[0066] Finally, generate the second training sample according to the teacher model corresponding to each preset dimension and the first training sample corresponding to each of the multiple preset dimensions.

[0067] Among them, the specific implementation manner of generating the second training sample according to the teacher model corresponding to each preset dimension and the first training samples corresponding to multiple preset dimensions may be: for each preset dimension, input the first sample text of this preset dimension into the teacher models corresponding to the other preset dimensions except this preset dimension respectively, and obtain the sample linguistic feature information of the other preset dimensions corresponding to the first sample text output by the teacher models of the other preset dimensions; splice the first sample text of the preset dimension, the sample linguistic feature information of the preset dimension corresponding to the first sample text, and the sample linguistic feature information of the other preset dimensions corresponding to the first sample text to obtain the sample linguistic feature information of multiple preset dimensions corresponding to the first sample text.

[0068] Exemplarily, as Figure 3 shown, input the first sample text of the word segmentation part-of-speech dimension into the teacher model corresponding to the prosodic structure dimension, and obtain the sample linguistic feature information of the prosodic structure dimension corresponding to the first sample text output by the teacher model of the prosodic structure dimension. Input the first sample text of the word segmentation part-of-speech dimension into the teacher model corresponding to the polyphone dimension, and obtain the sample linguistic feature information of the polyphone dimension corresponding to the first sample text output by the teacher model of the polyphone dimension. That is, input the first sample text of the word segmentation part-of-speech dimension into the prosodic structure teacher model and the polyphone teacher model respectively, and obtain the sample linguistic feature information of the prosodic structure dimension and the sample linguistic feature information of the polyphone dimension respectively.

[0069] Similarly, for the prosodic structure dimension, the sample linguistic feature information of the word segmentation part-of-speech dimension and the sample linguistic feature information of the polyphone dimension corresponding to the first sample text of the prosodic structure dimension can also be obtained. For the polyphone dimension, the sample linguistic feature information of the word segmentation part-of-speech dimension and the sample linguistic feature information of the prosodic structure dimension corresponding to the first sample text of the polyphone dimension can also be obtained.

[0070] For example, if the first sample text in the word segmentation part-of-speech dimension is text: "We work hard together to strive for a better tomorrow", and the sample linguistic feature information in the word segmentation part-of-speech dimension is segpos: ["B-r","E-r","B-d","E-d","B-a","E-a","B-p","E-p","B-v","E-v","S-u","B-t","E-t","S-c","B-v","E-v"], then input the first sample text text in the word segmentation part-of-speech dimension into the prosodic structure teacher model, and the sample linguistic feature information in the prosodic structure dimension is obtained as pph: ["B","I","I","I","I","I","B","I","I","I","I","I","I","I","I","I"]. Similarly, input the first sample text text in the word segmentation part-of-speech dimension into the polyphone teacher model, and the sample linguistic feature information in the polyphone dimension is obtained as g2p: offsets: [2,6,7,9,15,8,10], pinyins: ["yi4","wei4","le5","hao3","dou4","geng4","de5"], where offsets represents the position where the polyphone is located, and pinyins represents the pronunciation of the polyphone. Finally, splice the first sample text text, the sample linguistic feature information segpos in the word segmentation part-of-speech dimension, the sample linguistic feature information pw and pph in the prosodic structure dimension, and the sample linguistic feature information g2p in the polyphone dimension to obtain the sample linguistic feature information of multiple preset dimensions corresponding to the first sample text. Among them, the sample linguistic feature information of multiple preset dimensions corresponding to the first sample text is as follows, where pw represents the sample linguistic feature information of prosodic words, and pph represents the sample linguistic feature information of prosodic phrases:

[0071] "text": "We work hard together to strive for a better tomorrow",

[0072] "segpos": ["B-r","E-r","B-d","E-d","B-a","E-a","B-p","E-p","B-v","E-v","S-u","B-t","E-t","S-c","B-v","E-v"],

[0073] "g2p": {

[0074] "offsets": [2,6,7,9,15,8,10],

[0075] "pinyins": ["yi4","wei4","le5","hao3","dou4","geng4","de5"]

[0076] },

[0077] "pw": ["B", "I", "B", "I", "B", "I", "B", "I", "B", "I", "I", "B", "I", "B", "B", "I"],

[0078] "pph": ["B", "I", "I", "I", "I", "I", "B", "I", "I", "I", "I", "I", "I", "I", "I", "I"]

[0079] First, it should be understood that the first training sample in the word segmentation part-of-speech dimension only includes the first sample text "text" and "segpos", the first training sample in the prosodic structure dimension only includes the first sample text "text", "pw" and "pph", and the first training sample in the polyphone dimension only includes the first sample text "text" and "g2p".

[0080] Second, it should be understood that after obtaining the second training sample, it can be stored in JSON format or other formats.

[0081] In this way, according to the above solution, the first training sample is labeled by the teacher model corresponding to a single preset dimension to obtain the second training sample for training the multi-task front-end model, so that each first sample text has sample linguistic feature information of multiple preset dimensions, effectively solving the problem that the training data of different preset dimensions vary greatly when training the multi-task front-end model. In addition, by using the teacher model corresponding to a single preset dimension, that is, the single-task model, to label and obtain the second training sample, the labeling workload and cost are reduced.

[0082] In practical applications, when labeling the sample linguistic feature information of different preset dimensions in a sample text, there may be conflicts. For example, the conflict between the word segmentation boundary and the prosodic structure boundary. Therefore, in order to further improve the accuracy of the generated second training sample, the generated second training sample can also be subjected to conflict verification. When it is determined that there is no conflict, the pre-trained second benchmark model is trained to generate the multi-task front-end model. Exemplarily, Figure 2 Step S23 in it may further include: performing conflict verification on the sample linguistic feature information of multiple preset dimensions corresponding to each first sample text in the second training sample; when the verification result is that there is no conflict, using the second training sample to train the pre-trained second benchmark model to generate the multi-task front-end model.

[0083] For example, the first sample text is "I love scenic spot a in City A". Suppose the sample linguistic feature information in the dimension of word segmentation part-of-speech divides "scenic spot a in City A" into a phrase, and the sample linguistic feature information in the dimension of prosodic structure has a pause after "City A" and a pause after "of", that is, the prosodic words are "City A", "of", and "scenic spot a". Thus, there is a conflict between the sample linguistic feature information in the dimension of word segmentation part-of-speech and the sample linguistic feature information in the dimension of prosodic structure corresponding to the first sample text "I love scenic spot a in City A".

[0084] In addition, when there is a conflict, the conflict problem can also be repaired. For example, continuing with the above example, the sample linguistic feature information in the dimension of word segmentation part-of-speech can be modified to divide "City A" into a phrase, "of" into a phrase, and "scenic spot a" into a phrase. Or, the sample linguistic feature information in the dimension of prosodic structure can be modified to take "scenic spot a in City A" as a prosodic word.

[0085] After that, as Figure 3 shown, the small pre-trained Chinese BERT model is trained, that is, fine-tuned, using the second training sample without conflict to obtain a small multi-task front-end model.

[0086] Thus, after generating the second training sample, conflict detection can also be performed on the sample linguistic feature information in multiple preset dimensions corresponding to each first sample text in the second training sample. When there is no conflict, the second training sample is used to train the pre-trained second benchmark model to generate a multi-task front-end model, thereby further improving the accuracy of the generated multi-task front-end model.

[0087] In one embodiment, the first sample text is a text in the first business scenario. Figure 4 It is a flowchart of another method for generating a speech synthesis front-end model shown according to an exemplary embodiment. As Figure 4 shown, in addition to including steps S21 - S23, the method for generating the speech synthesis front-end model further includes the following steps.

[0088] In step S41, a third training sample in the newly added second business scenario is obtained.

[0089] Among them, the third training sample at least includes a second sample text in the second business scenario. The second business scenario can be a scenario different from the first business scenario or the same scenario, and the present disclosure does not make specific limitations thereon. For example, the first business scenario can be an intelligent voice assistant business scenario, and the second business scenario can be a navigation business scenario.

[0090] In step S42, a fourth training sample is generated according to the second sample text and the teacher model corresponding to each preset dimension.

[0091] Among them, the fourth training sample includes the second sample text and the sample linguistic feature information of multiple preset dimensions corresponding to each second sample text.

[0092] In step S43, the pre-trained second benchmark model is trained according to the second training sample and the fourth training sample to update the multi-task front-end model.

[0093] Adopting the above technical solution, for the third training sample in the newly added second business scenario, the teacher model corresponding to each preset dimension can be used to obtain the fourth training sample, so that the second sample text corresponds to the sample linguistic feature information of the above multiple preset dimensions. Then, according to the obtained fourth training sample and Figure 2 the second training sample obtained in Figure 2 the pre-trained second benchmark model is trained to obtain a new multi-task front-end model, and the multi-task front-end model shown in

[0094] is updated with the new multi-task front-end model, so that the updated multi-task front-end model can process the text in the first business scenario and the text in the second business scenario. In this way, it is convenient to migrate the multi-task front-end model to other business scenarios and expand the scope of use.

[0095] Next, the specific implementation manner of generating the fourth training sample according to the second sample text and the teacher model corresponding to each preset dimension in step S42 will be described.

[0096] In one implementation, the third training sample does not include the sample linguistic feature information of at least one preset dimension among the above-mentioned multiple preset dimensions. That is to say, the third training sample does not include the sample linguistic feature information of any dimension, or the dimension of the sample linguistic feature information included in the third training sample does not belong to the above-mentioned multiple preset dimensions. Accordingly, the specific implementation of step S42 is as follows: If the third training sample does not include the sample linguistic feature information of at least one preset dimension among the multiple preset dimensions, the second sample text is input into the teacher models respectively corresponding to each preset dimension, and the sample linguistic feature information of the preset dimension corresponding to the second sample text output by each teacher model is obtained; the second sample text and the sample linguistic feature information of each preset dimension corresponding to the second sample text are concatenated to obtain the fourth training sample.

[0097] Exemplarily, the second sample text is input into Figure 3 the word segmentation and part-of-speech teacher model, prosody structure teacher model, and polyphone teacher model shown in the figure, and the sample linguistic feature information of the word segmentation and part-of-speech dimension, prosody structure dimension, and polyphone dimension corresponding to the second sample text are obtained respectively. Then, the second sample text, the sample linguistic feature information of the word segmentation and part-of-speech dimension, the sample linguistic feature information of the prosody structure dimension, and the sample linguistic feature information of the polyphone dimension are concatenated to obtain the fourth training sample.

[0098] In another implementation, the third training sample includes the sample linguistic feature information of at least one preset dimension among the above-mentioned multiple preset dimensions. Accordingly, the specific implementation of step S42 is as follows: If the third training sample includes the sample linguistic feature information of the first target dimension, the second sample text is input into the teacher models corresponding to the other preset dimensions except the first target dimension among the multiple preset dimensions, and the sample linguistic feature information of the other preset dimensions corresponding to the second sample text output by the teacher models of the other preset dimensions is obtained, where the first target dimension is at least one preset dimension among the multiple preset dimensions; the second sample text, the sample linguistic feature information of the first target dimension corresponding to the second sample text, and the sample linguistic feature information of the other preset dimensions corresponding to the second sample text are concatenated to obtain the fourth training sample.

[0099] Exemplarily, as Figure 3As shown, assuming that the first target dimension is the word segmentation part-of-speech dimension, the second sample text is respectively input into the prosody structure teacher model and the polyphone teacher model to obtain the sample linguistic feature information of the prosody structure dimension and the sample linguistic feature information of the polyphone dimension corresponding to the first sample text. Then, the second sample text, the sample linguistic feature information of the word segmentation part-of-speech dimension corresponding to the second sample text, the sample linguistic feature information of the prosody structure dimension, and the sample linguistic feature information of the polyphone dimension are concatenated to obtain the fourth training sample.

[0100] In addition, if the third training sample includes the sample linguistic feature information of at least one preset dimension among the above multiple preset dimensions, the generation method may further include:

[0101] Training the first benchmark model according to the third training sample and the first training sample corresponding to the first target dimension to obtain an updated teacher model corresponding to the first target dimension.

[0102] Exemplarily, assuming that the first target dimension is the word segmentation part-of-speech dimension, the first benchmark model is trained according to the second sample text of the word segmentation part-of-speech dimension in the third training sample, the sample linguistic feature information of the word segmentation part-of-speech dimension corresponding to the second sample text, and the first training sample of the word segmentation part-of-speech dimension to obtain an updated teacher model corresponding to the word segmentation part-of-speech dimension.

[0103] Figure 5 It is a schematic diagram of another method for generating a speech synthesis front-end model shown according to an exemplary embodiment. As Figure 5 shown, assuming that the third training sample in the newly added second business scenario includes the sample linguistic feature information of the word segmentation part-of-speech dimension, the second sample text included in the third training sample is respectively input into the prosody structure teacher model and the polyphone teacher model for annotation. After annotation, conflict detection is performed on the sample linguistic feature information of multiple preset dimensions. When the detection result indicates that there is no conflict, the fourth training sample is directly obtained. When there is a conflict, the conflict problem is repaired to obtain the fourth training sample. Then, the second training sample and the fourth training sample are used to fine-tune the small pre-trained Chinese BERT model to obtain an updated post-task front-end model.

[0104] In addition, as Figure 5 shown, the large pre-trained Chinese BERT model is also fine-tuned according to the third training sample and the first training sample of the word segmentation part-of-speech dimension to obtain an updated part-of-speech teacher model.

[0105] In addition, considering Chinese Text Normalization, the main problem to be solved is to convert non-standard texts (such as abbreviations, numbers, special symbols, etc.) into natural speech expression forms, so as to achieve more natural and fluent speech synthesis in the TTS system. The Chinese text normalization technology has gone through an evolution process from rule-based methods, statistical methods to deep learning methods based on Transformer. Early text normalization mainly relied on manually formulated rules to perform character-by-character conversion on the input text. Although this method is simple and intuitive, it is difficult to handle complex and diverse language phenomena, and a large amount of manpower is required for maintenance. With the development of big data and the application of machine learning technology, statistical methods have gradually gained attention in the field of text normalization. The statistical-based method uses labeled data for training and automatically learns normalization rules through techniques such as classifiers or sequence labeling. This method alleviates the burden of manual rule compilation and maintenance to a certain extent, but still has limitations in dealing with long-distance dependencies and complex structures. In related technologies, the sentence context encoder has been changed from BiRNN to a large-scale pre-trained BERT model. First, BERT is used to encode the sentence context, then the parts that need to be normalized are marked by Tagger RNN, and finally, the text is normalized through an independent BiRNN-Attention-Decoder RNN structure. Alternatively, for the English text normalization scheme based on one-way editing, the normalization position is located by inserting normalization marks in the input text. However, since there is no space as the default boundary of words in Chinese text, the smallest unit of Chinese text is a character. If normalization marks are inserted between each character, the length of the input sequence will double, increasing the difficulty of understanding the text context. Moreover, Chinese characters themselves are generally not the objects of text normalization. Therefore, adding normalization marks before and after Chinese characters is of little significance. In addition, since the characters to be normalized are relatively few in Chinese text, using Tagger RNN will result in omissions and incorrect matches, and retraining the model is required for both adding and deleting recognition targets, making manual intervention difficult.

[0106] Therefore, based on the above considerations, the method for generating a speech synthesis front-end model provided by the present disclosure can also be used to generate a text normalization model. Exemplarily, the speech synthesis front-end model further includes a text normalization model. Figure 6 It is a flowchart of a method for generating a text normalization model shown according to an exemplary embodiment. As Figure 6 shown, the method for generating a text normalization model may include the following steps.

[0107] In step S61, obtain text normalization training samples.

[0108] In the present disclosure, text regularization training samples include: sample text to be regularized with regularization marks, target sample fields, and regularized target sample text corresponding to the sample text to be regularized. The target sample fields include regularization marks corresponding to the sample characters to be regularized and target regularized sample characters corresponding to the sample characters to be regularized.

[0109] The sample text to be regularized with regularization marks can be generated in the following way:

[0110] Firstly, a sample text to be regularized is identified from the sample text by using a preset recognition rule, and sample characters to be regularized and position information of the sample characters to be regularized are identified from the sample text to be regularized.

[0111] The sample text includes sample texts to be regularized that need to be regularized and also includes texts that do not need to be regularized. Therefore, the sample texts to be regularized are identified from the sample texts by using pre-set recognition rules. The recognition rules are used to define which texts belong to the sample texts to be regularized. For example, texts containing characters such as abbreviations, numbers, and special symbols are defined as sample texts to be regularized. The sample characters to be regularized and the position information of the characters to be regularized are identified from the sample texts to be regularized. The pre-set recognition rules may be weighted finite state transition machine WFST recognition rules.

[0112] Afterwards, a regularization mark is added at the position represented by the position information of the sample character to be regularized, so as to generate the sample text to be regularized with the regularization mark.

[0113] For example, assuming that the sample text to be regularized is "I shout 123, let's drink 50kg together", then through the pre-set recognition rules, the sample characters to be regularized are identified as "123" and "50kg", and the position of the sample characters "123" to be regularized in the sample text to be regularized is position 1, and the position of the sample characters "50kg" to be regularized in the sample text to be regularized is position 3. Among them, the regularization mark can be identified as [pos x], where x represents position x. Therefore, a regularization mark [pos 1] is added before "123", a regularization mark [pos 2] is added after "123", and a regularization mark [pos 3] is added before "50kg". It should be understood that the starting position of the text is position 0. In order to indicate the starting position of the text, a regularization mark [pos 0] is usually added at the beginning of the text. For example, the sample text to be regularized with regularization marks is "[pos 0] I shout [pos 1] 123 [pos2], let's drink [pos 3] 50kg together".

[0114] In this way, by means of the preset recognition rules, a to-be-regularized sample text with regularized marks is generated, which can ensure no omission during the recognition of the to-be-regularized sample text, improving the recognition efficiency. In addition, it is very convenient to manually modify the recognition rules.

[0115] Among them, the target sample field can be generated in the following way:

[0116] Align the to-be-regularized sample text with the target sample text, identify the target regularized sample characters corresponding to the characters of the to-be-regularized sample from the target sample text, and splice the regularized marks corresponding to the to-be-regularized sample characters and the target regularized sample characters corresponding to the to-be-regularized sample characters to generate the target sample field.

[0117] It should be understood that the regularized target sample text corresponding to the to-be-regularized sample text can be automatically generated or manually labeled. For example, existing text regularization techniques can be used to regularize the to-be-regularized sample text to generate the regularized target sample text corresponding to the to-be-regularized sample text. In addition, in order to ensure the accuracy of the determined target sample text, after obtaining the target sample text by using the existing text regularization techniques, manual quality inspection can also be carried out to determine whether the generated target sample text is accurate.

[0118] Exemplarily, align the to-be-regularized sample text "I shout 123, drink 50kg together" with the target sample text "I shout one two three, drink fifty kilograms together", and determine the target regularized sample characters corresponding to the characters of the to-be-regularized sample. For example, the target regularized sample characters corresponding to "123" are "one two three", and the target regularized sample characters corresponding to "50kg" are "fifty kilograms". Then, splice the regularized mark [pos 1] corresponding to "123" and the corresponding target regularized sample characters "one two three", and the regularized mark [pos 3] corresponding to "50kg" and the corresponding target regularized sample characters "fifty kilograms" to obtain the target sample field: [pos 1]one two three[pos 3]fifty kilograms.

[0119] In step S62, use the to-be-regularized sample text with regularized marks and the target sample field as the model input parameters, and use the target sample text as the model output parameter to train the initial text regularization model to obtain the trained text regularization model.

[0120] In the present disclosure, the initial text regularization model is an encoder-decoder structure based on Transformer, and among them, a cross-attention mechanism is further introduced in the decoder.

[0121] In one embodiment, the initial text regularization model includes an encoder, a decoder, and a splicing unit. The specific implementation of step S62 is as follows: input the text sample to be regularized with regularization tags into the encoder to obtain the context semantic information of the text sample to be regularized output by the encoder; input the context semantic information and the target sample field into the decoder to obtain the regularization result of the characters of the sample to be regularized output by the decoder; input the regularization result and the text sample to be regularized with regularization tags into the splicing unit to obtain the spliced text output by the splicing unit; train the encoder and the decoder according to the spliced text and the target sample text to obtain the trained text regularization model.

[0122] Figure 7 FIG. is a schematic diagram of a text regularization model generated according to an exemplary embodiment. As Figure 7 shown, first, input the text sample to be regularized with regularization tags, "[pos 0] I shout [pos 1] 123 [pos 2], drink together [pos3] 50 kg" into the encoder to obtain the context semantic information output by the encoder, where the encoder is a miniature pre-trained Chinese BERT model. Then, input the context semantic information and the target sample field "[pos 1] one, two, three [pos 3] fifty kilograms" into the decoder to obtain the regularization result output by the decoder. After that, input the regularization result and the text sample to be regularized with regularization tags into the splicing unit to obtain the spliced text output by the splicing unit. Finally, train the encoder and the decoder according to the spliced text and the target sample text to obtain the trained text regularization model. For example, use the cross-entropy loss function to calculate the error between the spliced text and the target sample text, and adjust the parameters of the encoder and the decoder according to the error.

[0123] In this way, training in the above manner can obtain a text regularization model for regularizing text, which can solve the text regularization problem and reduce the amount of calculation.

[0124] It should be understood that the text regularization training samples can be automatically generated in the above manner or manually labeled. In the method of automatically generating text regularization training samples, manual quality inspection can also be performed on the generated text regularization training samples, and training can be carried out only when it is ensured that the training samples are correct. In this way, the performance of the trained text regularization model is further improved.

[0125] In addition, in order to reduce the scale of the speech synthesis front-end model, the multi-task front-end model and the text regularization model share the vocabulary and the pre-trained embedding vectors, thereby reducing the memory occupied by the speech synthesis front-end model.

[0126] In the present disclosure, the encoder is a pre-trained third reference model, the model scale of the first reference model is larger than that of the second reference model, and the model scale of the second reference model is larger than that of the third reference model.

[0127] In this way, a text regularization model with a smaller memory footprint can be used, further reducing the model scale of the speech synthesis front-end model, and thus reducing the memory occupied by the speech synthesis front-end model.

[0128] Based on the same inventive concept, the present disclosure also provides a speech synthesis front-end processing method. Figure 8 It is a flowchart of a speech synthesis front-end processing method shown according to an exemplary embodiment. As Figure 8 shown, the speech synthesis front-end processing method may include the following steps.

[0129] In step S81, a target text to be converted is obtained.

[0130] The target text to be converted refers to the text that needs to be converted into the corresponding speech.

[0131] In step S82, the target text to be converted is input into the multi-task front-end model, and multiple-dimensional linguistic feature information corresponding to the target text output by the multi-task front-end model is obtained.

[0132] In the present disclosure, the multi-task front-end model is generated according to the generation method of the speech synthesis front-end model provided by the present disclosure. The multiple-dimensional linguistic feature information corresponding to the target text output by the multi-task front-end model is used to input into the speech synthesis back-end model to generate the speech corresponding to the target text.

[0133] By adopting the above technical solution, the multi-task front-end model with higher accuracy is used to identify the multiple-dimensional linguistic feature information corresponding to the target text, improving the accuracy of the identified linguistic feature information, and thus improving the efficiency of speech synthesis.

[0134] Next, a complete embodiment is used to describe the speech synthesis front-end processing method provided by the present disclosure. Figure 9 It is a schematic diagram of a speech synthesis front-end processing method shown according to an exemplary embodiment. As Figure 9 shown, first, an original text to be converted is obtained. Then, the original text is input into the text regularization model to obtain the target text, that is, through the text regularization model, numbers, special letters, and symbols are converted into spoken text representations. After that, the target text is input into the multi-task front-end model to obtain multiple-dimensional linguistic feature information.

[0135] Exemplarily, as Figure 9As shown, the linguistic feature information in multiple dimensions output by the multi-task front-end model includes, but is not limited to, the linguistic feature information in the word segmentation and part-of-speech dimension, the linguistic feature information in the prosodic structure dimension, and the linguistic feature information in the polyphone dimension. Among them, the linguistic feature information in the prosodic structure dimension may include the linguistic feature information of prosodic words and the linguistic feature information of prosodic phrases. The linguistic feature information in the polyphone dimension includes the linguistic feature information of the stress position and the linguistic feature information of the polyphone pronunciation. It should be understood that different types of linguistic feature information are output through different output layers. For example, the linguistic feature information in the word segmentation and part-of-speech dimension, the linguistic feature information of prosodic words, the linguistic feature information of prosodic phrases, the linguistic feature information of the stress position, and the linguistic feature information of the polyphone pronunciation are all output through different output layers.

[0136] Preferably, the output layers of the linguistic feature information in the word segmentation and part-of-speech dimension, the linguistic feature information of prosodic words, the linguistic feature information of prosodic phrases, and the linguistic feature information of the stress position each include a linear layer and a CRF (conditional random fields) layer, where the CRF layer is used to record the transition relationships between various types of linguistic feature information. The output layer of the linguistic feature information of the stress position and the linguistic feature information of the polyphone pronunciation only includes a linear layer.

[0137] The following describes the specific implementation manner for obtaining the target text to be converted.

[0138] The specific implementation manner for obtaining the target text to be converted is as follows:

[0139] Obtain the original text to be converted, and identify whether the original text contains the characters to be regularized;

[0140] If it is determined that the original text contains the characters to be regularized, then determine the position information of the characters to be regularized;

[0141] Add a regularization mark at the position represented by the position information of the characters to be regularized to generate a text to be regularized with a regularization mark;

[0142] Input the text to be regularized into the text regularization model to obtain the target text to be converted. The text regularization model is generated according to the generation method of the speech synthesis front-end model provided by the present disclosure.

[0143] It should be understood that the specific implementation manner for generating the text to be regularized with a regularization mark is similar to the manner for generating the text sample to be regularized with a regularization mark, and will not be elaborated here.

[0144] In one embodiment, the text regularization model includes an encoder, a decoder, and a splicing unit. The specific implementation of obtaining the target text to be converted by inputting the text to be regularized into the text regularization model is as follows: Input the text to be regularized into the encoder to obtain the context semantic information of the text to be regularized output by the encoder; input the context semantic information and the regularization markers corresponding to the characters to be regularized into the decoder to obtain the target characters corresponding to the characters to be regularized output by the decoder; input the target characters and the text to be regularized into the splicing unit to obtain the target text to be converted output by the splicing unit.

[0145] Exemplarily, Figure 10 is a flowchart of generating a target text shown according to an exemplary embodiment. As Figure 10 shown, first, input the original text "I shout 456, drink 30 kg together" into the recognition rule unit, which is used to process the original text into the text to be regularized "[pos 0]I shout[pos 1]456[pos 2], drink[pos 3]30 kg". Then, input the text to be regularized "[pos 0]I shout[pos 1]456[pos 2], drink[pos 3]30 kg" into the encoder to obtain the context semantic information output by the encoder. Then, input the context semantic information and the regularization markers "[pos 1]" and "[pos 3]" corresponding to the characters to be regularized into the decoder together to obtain the regularization result "four five six thirty kilograms" output by the decoder. After that, input the regularization result "four five six thirty kilograms" and the text sample to be regularized with regularization markers "[pos 0]I shout[pos 1]456[pos 2], drink[pos 3]30 kg" into the splicing unit to obtain the spliced text output by the splicing unit, that is, generate the target text to be converted "I shout four five six, drink thirty kilograms".

[0146] It should be understood that the splicing unit can obtain the regularization result of each position, replace the characters to be regularized with target characters according to the position, and obtain the target text.

[0147] In addition, the speech synthesis front-end model further includes a candidate character generation unit, and the processing method further includes: inputting the characters to be regularized into the candidate character generation unit to obtain multiple candidate characters output by the candidate character generation unit; correspondingly, inputting the context semantic information and the regularization markers corresponding to the characters to be regularized into the decoder to obtain the target characters corresponding to the characters to be regularized output by the decoder, including: inputting the context semantic information, the regularization markers corresponding to the characters to be regularized, and multiple candidate characters into the decoder, so that the decoder determines the target characters corresponding to the characters to be regularized from the multiple candidate characters according to the context semantic information and the regularization markers corresponding to the characters to be regularized, and outputs the target characters.

[0148] Figure 11is another flowchart for generating target text shown according to an exemplary embodiment. As Figure 11 shown, the candidate character generation unit is respectively connected to the recognition and regularization unit and the decoder. Among them, in addition to outputting the text to be regularized "[pos 0] I shout [pos 1] 456 [pos 2], drink together [pos 3] 30kg", the recognition and regularization unit also outputs the characters to be regularized "456" and "30kg". Then, the text to be regularized is input into the encoder, and the regularized characters are input into the candidate character generation unit to obtain multiple candidate characters output by the candidate character generation unit. After that, the context semantic information output by the encoder, the regularization markers corresponding to the characters to be regularized, and the multiple candidate characters are input into the decoder, so that the decoder determines the target characters corresponding to the characters to be regularized from the multiple candidate characters according to the context semantic information and the regularization markers corresponding to the characters to be regularized, and outputs the target characters.

[0149] For example, for the character to be regularized "456", the multiple candidate characters generated by the candidate character generation unit are respectively "four five six" and "four hundred and fifty-six". The decoder determines from the multiple candidate characters the target character corresponding to the character to be regularized "456" as "four five six" according to the context semantic information and the regularization marker corresponding to the character to be regularized, and then outputs the target character "four five six". Similarly, for the character to be regularized "30kg", the decoder determines from the multiple candidate characters the target character corresponding to the character to be regularized "30kg" as "thirty kilograms" according to the context semantic information and the regularization marker corresponding to the character to be regularized, and then outputs the target character "thirty kilograms". That is, the target characters output by the decoder are "four five six thirty kilograms".

[0150] Adopting the above technical solution, multiple candidate characters are generated by the candidate character generation unit to ensure that the regularized result output by the decoder will not have imperceptible errors.

[0151] Based on the same inventive concept, the present disclosure also provides a generating device for a speech synthesis front-end model. Figure 12 is a block diagram of a generating device for a speech synthesis front-end model shown according to an exemplary embodiment. As Figure 12 shown, the generating device 1200 for the speech synthesis front-end model may include:

[0152] The first acquisition module 1201 is configured to acquire first training samples corresponding to multiple preset dimensions. Among them, the first training sample corresponding to each preset dimension includes the first sample text of the dimension and the sample linguistic feature information of the dimension corresponding to the first sample text. The sample linguistic feature information of the preset dimension refers to the feature information of the first sample text described from the preset dimension and required for synthesizing the speech corresponding to the first sample text in the speech synthesis backend;

[0153] The first generation module 1202 is configured to generate a second training sample according to a pre-trained first benchmark model and the first training sample, where the second training sample includes first sample texts of the multiple preset dimensions and sample linguistic feature information of the multiple preset dimensions corresponding to each of the first sample texts;

[0154] The first training module 1203 is configured to train a pre-trained second benchmark model by using the second training sample to generate a multi-task front-end model, where the speech synthesis front-end model includes the multi-task front-end model, and the model scale of the first benchmark model is larger than that of the second benchmark model.

[0155] Optionally, the first generation module 1202 includes:

[0156] The first training sub-module is configured to, for each preset dimension, use the first sample text of the preset dimension as a model input parameter and the sample linguistic feature information of the preset dimension corresponding to the first sample text as a model output parameter to train the pre-trained first benchmark model to generate a teacher model corresponding to the preset dimension;

[0157] The generation sub-module is configured to generate a second training sample according to the teacher model corresponding to each preset dimension and the first training samples corresponding to the multiple preset dimensions.

[0158] Optionally, the generation sub-module is configured to: for each preset dimension, input the first sample text of the preset dimension into the teacher models corresponding to the other preset dimensions except the preset dimension respectively to obtain the sample linguistic feature information of the other preset dimensions corresponding to the first sample text output by the teacher models of the other preset dimensions;

[0159] Concatenate the first sample text of the preset dimension, the sample linguistic feature information of the preset dimension corresponding to the first sample text, and the sample linguistic feature information of the other preset dimensions corresponding to the first sample text to obtain the sample linguistic feature information of the multiple preset dimensions corresponding to the first sample text.

[0160] Optionally, the first training module 1203 includes:

[0161] The verification sub-module is configured to perform conflict verification on the sample linguistic feature information of the multiple preset dimensions corresponding to each first sample text in the second training sample;

[0162] The first training sub-module is configured to train a pre-trained second benchmark model using the second training sample when the verification result indicates no conflict, so as to generate a multi-task front-end model.

[0163] Optionally, the first sample text is text in a first business scenario; the generating device 1200 of the speech synthesis front-end model may further include:

[0164] A second obtaining module, configured to obtain a third training sample in a newly added second business scenario, where the third training sample at least includes second sample text in the second business scenario;

[0165] A second generating module, configured to generate a fourth training sample according to the second sample text and teacher models corresponding to each preset dimension, where the fourth training sample includes the second sample text and sample linguistic feature information of the plurality of preset dimensions corresponding to each second sample text;

[0166] A second training module, configured to train a pre-trained second benchmark model according to the second training sample and the fourth training sample, so as to update the multi-task front-end model.

[0167] Optionally, the second generating module is configured to: if the third training sample does not include sample linguistic feature information of at least one preset dimension among the plurality of preset dimensions, input the second sample text into teacher models corresponding to each preset dimension respectively, and obtain sample linguistic feature information of the preset dimension corresponding to the second sample text output by the teacher models corresponding to each preset dimension;

[0168] Concatenate the second sample text and the sample linguistic feature information of each preset dimension corresponding to the second sample text to obtain a fourth training sample.

[0169] Optionally, the second generating module is configured to: if the third training sample includes sample linguistic feature information of a first target dimension, input the second sample text into teacher models corresponding to other preset dimensions except the first target dimension among the plurality of preset dimensions, and obtain sample linguistic feature information of the other preset dimensions corresponding to the second sample text output by the teacher models corresponding to the other preset dimensions, where the first target dimension is at least one preset dimension among the plurality of preset dimensions;

[0170] Concatenate the second sample text, the sample linguistic feature information of the first target dimension corresponding to the second sample text, and the sample linguistic feature information of the other preset dimensions corresponding to the second sample text to obtain a fourth training sample.

[0171] Optionally, the generating device 1200 of the speech synthesis front-end model may further include:

[0172] A third training module, configured to train the first reference model according to the third training sample and the first training sample corresponding to the first target dimension, so as to obtain the updated teacher model corresponding to the first target dimension.

[0173] Optionally, the speech synthesis front-end model further includes a text regularization model; the generating device 1200 of the speech synthesis front-end model may further include:

[0174] A third obtaining module, configured to obtain text regularization training samples, where the text regularization training samples include: to-be-regularized sample texts with regularization marks, target sample fields, and the regularized target sample texts corresponding to the to-be-regularized sample texts, the target sample fields include the regularization marks corresponding to the to-be-regularized sample characters and the target regularized sample characters corresponding to the to-be-regularized sample characters, and the regularization marks are used to mark the position information of the to-be-regularized sample characters;

[0175] A fourth training module, configured to use the to-be-regularized sample text with the regularization mark and the target sample field as model input parameters, and use the target sample text as a model output parameter to train an initial text regularization model to obtain a trained text regularization model.

[0176] Optionally, the original text regularization model includes an encoder, a decoder, and a splicing unit; the fourth training module is configured to:

[0177] Input the to-be-regularized sample text with the regularization mark into the encoder to obtain the context semantic information of the to-be-regularized sample text output by the encoder;

[0178] Input the context semantic information and the target sample field into the decoder to obtain the regularization result of the to-be-regularized sample characters output by the decoder;

[0179] Input the regularization result and the to-be-regularized sample text with the regularization mark into the splicing unit to obtain the spliced text output by the splicing unit;

[0180] Train the encoder and the decoder according to the spliced text and the target sample text to obtain a trained text regularization model.

[0181] Optionally, the to-be-regularized sample text with the regularization mark is generated in the following manner:

[0182] Identify the sample text to be regularized from the sample text according to the pre-set identification rules, identify the sample characters to be regularized from the sample text to be regularized, and the position information of the sample characters to be regularized;

[0183] Add a regularization mark at the position characterized by the position information of the sample characters to be regularized to generate a sample text to be regularized with a regularization mark;

[0184] The target sample field is generated by the following method:

[0185] Align the sample text to be regularized with the target sample text, identify the target regularized sample characters corresponding to the sample characters to be regularized from the target sample text, and splice the regularization mark corresponding to the sample characters to be regularized and the target regularized sample characters corresponding to the sample characters to be regularized to generate the target sample field.

[0186] Optionally, the encoder is a pre-trained third reference model, and the model scale of the second reference model is larger than that of the third reference model.

[0187] Based on the same inventive concept, the present disclosure also provides a voice synthesis front-end processing device. Figure 13 It is a block diagram of a voice synthesis front-end processing device shown according to an exemplary embodiment. As Figure 13 shown, the voice synthesis front-end processing device 1300 may include:

[0188] A fourth acquisition module 1301, configured to acquire a target text to be converted;

[0189] An input module 1302, configured to input the target text to be converted into a multi-task front-end model to obtain multiple-dimensional linguistic feature information corresponding to the target text output by the multi-task front-end model. The multi-task front-end model is generated according to the generation method of the voice synthesis front-end model provided by the present disclosure. The multiple-dimensional linguistic feature information is used to input a voice synthesis back-end model to generate the voice corresponding to the target text.

[0190] Optionally, the fourth acquisition module 1301 is configured to:

[0191] Acquire the original text to be converted and identify whether the original text contains characters to be regularized;

[0192] If it is determined that the original text contains the characters to be regularized, determine the position information of the characters to be regularized;

[0193] Add a regularization mark at the position characterized by the position information of the characters to be regularized to generate a regularized text with a regularization mark;

[0194] Input the text to be regularized into the text regularization model to obtain the target text to be converted, where the text regularization model is generated according to the method for generating the speech synthesis front-end model provided in this disclosure.

[0195] Optionally, the text regularization model includes an encoder, a decoder, and a splicing unit; the fourth acquisition module 1301 is configured to: input the text to be regularized into the encoder to obtain the context semantic information of the text to be regularized output by the encoder;

[0196] Input the context semantic information and the regularization mark corresponding to the text to be regularized into the decoder to obtain the target character corresponding to the text to be regularized output by the decoder;

[0197] Input the target character and the text to be regularized into the splicing unit to obtain the target text to be converted output by the splicing unit.

[0198] Optionally, the speech synthesis front-end model further includes a candidate character generation unit; the speech synthesis front-end processing device 1300 may further include:

[0199] A second input module, configured to input the text to be regularized into the candidate character generation unit to obtain a plurality of candidate characters output by the candidate character generation unit;

[0200] The fourth acquisition module 1301 is configured to: input the context semantic information, the regularization mark corresponding to the text to be regularized, and the plurality of candidate characters into the decoder, so that the decoder determines the target character corresponding to the text to be regularized from the plurality of candidate characters according to the context semantic information and the regularization mark corresponding to the text to be regularized, and outputs the target character.

[0201] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0202] This disclosure also provides a computer-readable storage medium, on which computer program instructions are stored, and when the program instructions are executed by a processor, the steps of the method for generating the speech synthesis front-end model and / or the speech synthesis front-end processing method provided in this disclosure are implemented.

[0203] This disclosure also provides an electronic device, including: a processor;

[0204] A memory for storing instructions executable by the processor;

[0205] Wherein, the processor is configured to execute the executable instructions to implement the steps of the method for generating the speech synthesis front-end model provided by the present disclosure, and / or to implement the steps of the method for processing the speech synthesis front-end provided by the present disclosure.

[0206] Figure 14 is a block diagram of an electronic device shown according to an exemplary embodiment. For example, the electronic device 1400 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0207] Referring to Figure 14 , the electronic device 1400 may include one or more of the following components: a processing component 1402, a memory 1404, a power supply component 1406, a multimedia component 1408, an audio component 1410, an input / output interface 1412, a sensor component 1414, and a communication component 1416.

[0208] The processing component 1402 generally controls the overall operation of the electronic device 1400, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 1402 may include one or more processors 1420 to execute instructions to complete the method for generating the speech synthesis front-end model, and / or to implement all or part of the steps of the method for processing the speech synthesis front-end provided by the present disclosure. In addition, the processing component 1402 may include one or more modules to facilitate the interaction between the processing component 1402 and other components. For example, the processing component 1402 may include a multimedia module to facilitate the interaction between the multimedia component 1408 and the processing component 1402.

[0209] The memory 1404 is configured to store various types of data to support the operation of the electronic device 1400. Examples of such data include instructions for any application or method operating on the electronic device 1400, contact data, phone book data, messages, pictures, videos, etc. The memory 1404 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.

[0210] The power supply component 1406 provides power to various components of the electronic device 1400. The power supply component 1406 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 1400.

[0211] The multimedia component 1408 includes a screen that provides an output interface between the electronic device 1400 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of the touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operations. In some embodiments, the multimedia component 1408 includes a front camera and / or a rear camera. When the electronic device 1400 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.

[0212] The audio component 1410 is configured to output and / or input audio signals. For example, the audio component 1410 includes a microphone (MIC) that is configured to receive external audio signals when the electronic device 1400 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 1404 or transmitted via the communication component 1416. In some embodiments, the audio component 1410 further includes a speaker for outputting audio signals.

[0213] The input / output interface 1412 provides an interface between the processing component 1402 and a peripheral interface module, and the peripheral interface module can be a keyboard, a click wheel, buttons, etc. These buttons can include but are not limited to: a home button, a volume button, a power button, and a lock button.

[0214] The sensor assembly 1414 includes one or more sensors for providing status assessments of various aspects for the electronic device 1400. For example, the sensor assembly 1414 can detect the on / off state of the electronic device 1400, the relative positioning of components, such as the display and keypad of the electronic device 1400. The sensor assembly 1414 can also detect a change in the position of the electronic device 1400 or a component of the electronic device 1400, the presence or absence of user contact with the electronic device 1400, the orientation or acceleration / deceleration of the electronic device 1400, and the temperature change of the electronic device 1400. The sensor assembly 1414 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 1414 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 1414 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0215] The communication component 1416 is configured to facilitate communication between the electronic device 1400 and other devices in a wired or wireless manner. The electronic device 1400 can access a wireless network based on communication standards, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 1416 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 1416 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0216] In an exemplary embodiment, the electronic device 1400 can be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the method for generating the voice synthesis front-end model, and / or for implementing the voice synthesis front-end processing method provided by the present disclosure.

[0217] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 1404 including instructions. The above instructions can be executed by a processor 1420 of an electronic device 1400 to complete the method for generating the voice synthesis front-end model, and / or to implement the voice synthesis front-end processing method provided by the present disclosure. For example, the non-transitory computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0218] In another exemplary embodiment, a computer program product is also provided. The computer program product includes a computer program that can be executed by a programmable device. The computer program has a code portion for executing the above method for generating the voice synthesis front-end model and / or for implementing the voice synthesis front-end processing method provided by the present disclosure when executed by the programmable device.

[0219] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the present disclosure. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.

[0220] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. A method for generating a front-end model of speech synthesis, characterized in that, The generation method includes: Obtaining first training samples corresponding to multiple preset dimensions. Each first training sample corresponding to a preset dimension includes the first sample text of the preset dimension and the sample linguistic feature information of the preset dimension corresponding to the first sample text. The sample linguistic feature information of the preset dimension refers to the feature information of the first sample text required for synthesizing the voice corresponding to the first sample text in the voice synthesis backend as described by the preset dimension. Generating second training samples according to a pre-trained first benchmark model and the first training samples. The second training samples include the first sample texts of the multiple preset dimensions and the sample linguistic feature information of the multiple preset dimensions corresponding to each first sample text. Training a pre-trained second benchmark model with the second training samples to generate a multi-task front-end model. The voice synthesis front-end model includes the multi-task front-end model. The model scale of the first benchmark model is larger than that of the second benchmark model.

2. The generating method according to claim 1, characterized in that, The generating second training samples according to a pre-trained first benchmark model and the first training samples includes: For each preset dimension, using the first sample text of the preset dimension as the model input parameter and the sample linguistic feature information of the preset dimension corresponding to the first sample text as the model output parameter to train the pre-trained first benchmark model to generate a teacher model corresponding to the preset dimension. Generating second training samples according to the teacher models corresponding to each preset dimension and the first training samples corresponding to the multiple preset dimensions.

3. The generating method according to claim 2, characterized in that, The generating second training samples according to the teacher models corresponding to each preset dimension and the first training samples corresponding to the multiple preset dimensions includes: For each preset dimension, inputting the first sample text of the preset dimension into the teacher models corresponding to the other preset dimensions except the preset dimension to obtain the sample linguistic feature information of the other preset dimensions corresponding to the first sample text output by the teacher models of the other preset dimensions. Concatenating the first sample text of the preset dimension, the sample linguistic feature information of the preset dimension corresponding to the first sample text, and the sample linguistic feature information of the other preset dimensions corresponding to the first sample text to obtain the sample linguistic feature information of the multiple preset dimensions corresponding to the first sample text.

4. The generating method according to claim 1, characterized in that, The training a pre-trained second benchmark model with the second training samples to generate a multi-task front-end model includes: Performing conflict verification on the sample linguistic feature information of the multiple preset dimensions corresponding to each first sample text in the second training samples. When the verification result shows no conflict, training the pre-trained second benchmark model with the second training samples to generate a multi-task front-end model.

5. The generating method according to claim 2, characterized in that, The first sample text is a text in a first business scenario. The generation method further includes: Obtain the third training samples in the newly added second service scenario, where the third training samples at least include the second sample texts in the second service scenario; Generate fourth training samples according to the second sample texts and the teacher models corresponding to each preset dimension, where the fourth training samples include the second sample texts and the sample linguistic feature information of the multiple preset dimensions corresponding to each of the second sample texts; Train the pre-trained second benchmark model according to the second training samples and the fourth training samples to update the multi-task front-end model.

6. The generating method according to claim 5, characterized in that, The generating the fourth training samples according to the second sample texts and the teacher models corresponding to each preset dimension includes: If the third training samples do not include the sample linguistic feature information of at least one preset dimension among the multiple preset dimensions, input the second sample texts into the teacher models corresponding to each preset dimension respectively to obtain the sample linguistic feature information of the preset dimension corresponding to the second sample texts output by the teacher models corresponding to each preset dimension; Concatenate the second sample texts and the sample linguistic feature information of each preset dimension corresponding to the second sample texts to obtain the fourth training samples.

7. The generating method according to claim 5, characterized in that, The generating the fourth training samples according to the second sample texts and the teacher models corresponding to each preset dimension includes: If the third training samples include the sample linguistic feature information of the first target dimension, input the second sample texts into the teacher models corresponding to the other preset dimensions except the first target dimension among the multiple preset dimensions to obtain the sample linguistic feature information of the other preset dimensions corresponding to the second sample texts output by the teacher models corresponding to the other preset dimensions, where the first target dimension is at least one preset dimension among the multiple preset dimensions; Concatenate the second sample texts, the sample linguistic feature information of the first target dimension corresponding to the second sample texts, and the sample linguistic feature information of the other preset dimensions corresponding to the second sample texts to obtain the fourth training samples.

8. The generating method according to claim 7, characterized in that, The generating method further includes: Train the first benchmark model according to the third training samples and the first training samples corresponding to the first target dimension to obtain the updated teacher model corresponding to the first target dimension.

9. The generating method according to any one of claims 1-8, characterized in that, The speech synthesis front-end model further includes a text regularization model; the generating method further includes: Obtain text regularization training samples, where the text regularization training samples include: the to-be-regularized sample texts with regularization marks, target sample fields, and the target sample texts after regularization corresponding to the to-be-regularized sample texts, the target sample fields include the regularization marks corresponding to the to-be-regularized sample characters and the target regularized sample characters corresponding to the to-be-regularized sample characters, and the regularization marks are used to mark the position information of the to-be-regularized sample characters; Use the to-be-regularized sample texts with regularization marks and the target sample fields as model input parameters, and use the target sample texts as model output parameters to train the initial text regularization model to obtain the trained text regularization model.

10. The generating method according to claim 9, characterized in that, The original text regularization model includes an encoder, a decoder, and a splicing unit; Training the initial text regularization model by using the to-be-regularized sample text with regularization marks and the target sample field as model input parameters and the target sample text as the model output parameter, including: Inputting the to-be-regularized sample text with regularization marks into the encoder to obtain the context semantic information of the to-be-regularized sample text output by the encoder; Inputting the context semantic information and the target sample field into the decoder to obtain the regularization result of the characters of the to-be-regularized sample output by the decoder; Inputting the regularization result and the to-be-regularized sample text with regularization marks into the splicing unit to obtain the spliced text output by the splicing unit; Training the encoder and the decoder according to the spliced text and the target sample text to obtain the trained text regularization model.

11. The generation method according to claim 9, wherein, The to-be-regularized sample text with regularization marks is generated in the following manner: Identifying the to-be-regularized sample text from the sample text through a preset recognition rule, and identifying the characters of the to-be-regularized sample and the position information of the characters of the to-be-regularized sample from the to-be-regularized sample text; Adding a regularization mark at the position characterized by the position information of the characters of the to-be-regularized sample to generate the to-be-regularized sample text with regularization marks; The target sample field is generated in the following manner: Aligning the to-be-regularized sample text with the target sample text, identifying the target regularized sample characters corresponding to the characters of the to-be-regularized sample from the target sample text, and splicing the regularization marks corresponding to the characters of the to-be-regularized sample and the target regularized sample characters corresponding to the characters of the to-be-regularized sample to generate the target sample field.

12. The generation method according to claim 10, wherein, The encoder is a pre-trained third reference model, and the model scale of the second reference model is larger than that of the third reference model.

13. A method for front-end processing of speech synthesis, wherein, The processing method includes: Obtaining a target text to be converted; Inputting the target text to be converted into a multi-task front-end model to obtain multi-dimensional linguistic feature information corresponding to the target text output by the multi-task front-end model, where the multi-task front-end model is generated according to the generation method of the speech synthesis front-end model described in any one of claims 1-12, and the multi-dimensional linguistic feature information is used to input into a speech synthesis back-end model to generate speech corresponding to the target text.

14. The method for front-end processing of speech synthesis according to claim 13, wherein, The obtaining of the target text to be converted includes: Obtaining the original text to be converted and identifying whether the original text contains characters to be regularized; If it is determined that the original text contains the characters to be regularized, determining the position information of the characters to be regularized; Adding a regularization mark at the position characterized by the position information of the characters to be regularized to generate a to-be-regularized text with regularization marks; Inputting the to-be-regularized text into a text regularization model to obtain the target text to be converted, where the text regularization model is generated according to the generation method of the speech synthesis front-end model described in any one of claims 9-12.

15. The method for front-end processing of speech synthesis according to claim 14, wherein, The text regularization model includes an encoder, a decoder, and a splicing unit; The inputting the text to be regularized into the text regularization model to obtain a target text to be converted includes: Inputting the text to be regularized into the encoder to obtain the context semantic information of the text to be regularized output by the encoder; Inputting the context semantic information and the regularization mark corresponding to the text to be regularized into the decoder to obtain the target character corresponding to the text to be regularized output by the decoder; Inputting the target character and the text to be regularized into the splicing unit to obtain the target text to be converted output by the splicing unit.

16. The method for front-end processing of speech synthesis according to claim 15, wherein, The speech synthesis front-end model further includes a candidate character generation unit; the processing method further includes: Inputting the text to be regularized into the candidate character generation unit to obtain a plurality of candidate characters output by the candidate character generation unit; The inputting the context semantic information and the regularization mark corresponding to the text to be regularized into the decoder to obtain the target character corresponding to the text to be regularized output by the decoder includes: Inputting the context semantic information, the regularization mark corresponding to the text to be regularized, and the plurality of candidate characters into the decoder, so that the decoder determines the target character corresponding to the text to be regularized from the plurality of candidate characters according to the context semantic information and the regularization mark corresponding to the text to be regularized, and outputs the target character.

17. A device for generating a front-end model of speech synthesis, wherein, The generating device includes: A first acquisition module, configured to acquire first training samples corresponding to a plurality of preset dimensions, wherein each first training sample corresponding to a preset dimension includes a first sample text of the dimension and sample linguistic feature information of the dimension corresponding to the first sample text, and the sample linguistic feature information of the preset dimension refers to the feature information of the first sample text described from the preset dimension and required for synthesizing the speech corresponding to the first sample text in the speech synthesis backend; A first generation module, configured to generate second training samples according to a pre-trained first benchmark model and the first training samples, the second training samples including the first sample texts of the plurality of preset dimensions and the sample linguistic feature information of the plurality of preset dimensions corresponding to each first sample text; A first training module, configured to train a pre-trained second benchmark model by using the second training samples to generate a multi-task front-end model, the speech synthesis front-end model includes the multi-task front-end model, and the model scale of the first benchmark model is larger than the model scale of the second benchmark model.

18. A front-end processing device for speech synthesis, characterized in that, The processing device includes: A second acquisition module, configured to acquire a target text to be converted; An input module, configured to input the target text to be converted into a multi-task front-end model, and obtain linguistic feature information of multiple dimensions corresponding to the target text output by the multi-task front-end model. The multi-task front-end model is generated according to the generation method of the speech synthesis front-end model described in any one of claims 1-12. The linguistic feature information of multiple dimensions is used to input a speech synthesis back-end model to generate speech corresponding to the target text.

19. An electronic device, characterized in that, Comprising: A processor; A memory for storing processor-executable instructions; Wherein, the processor is configured to execute the executable instructions to implement the steps of the generation method of the speech synthesis front-end model described in any one of claims 1-12, and / or to implement the steps of the speech synthesis front-end processing method described in any one of claims 13-16.

20. A computer-readable storage medium, on which computer program instructions are stored, characterized in that, When the program instructions are executed by the processor, the steps of the generation method of the speech synthesis front-end model described in any one of claims 1-12, and / or the steps of the speech synthesis front-end processing method described in any one of claims 13-16 are implemented.