Speech synthesis method and device based on few-sample learning, equipment and medium

Through the speech synthesis method based on few sample learning, the speech synthesis problem caused by the scarcity of data in the Chinese civil aviation land and air call scenario is solved, and the effect of generating realistic voice under limited data is achieved. It is suitable for civil aviation land and air call simulation training, improving the authenticity and efficiency of voice interaction.

CN119964545AInactive Publication Date: 2025-05-09SHENZHEN POLYTECHNIC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510021659.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-05-09
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the Chinese Civil Aviation Land-Air call scenario, due to the scarcity of data, existing speech synthesis technologies are difficult to generate realistic speech that meets the requirements, which affects the training effect and requires a high degree of speaker adaptability, which increases the difficulty of speech synthesis based on few sample learning.

Method used

A speech synthesis method based on few-sample learning is proposed. By obtaining the training data set of the few-sample speech synthesis model, including training speech information and text information of multiple speakers, the speech synthesis model is trained, and the speech with the acoustic characteristics of the target speech human being is generated under the input of the application text information and the target speaker information.

Benefits of technology

Quickly generate realistic speech with targeted speaker acoustic features under limited data samples, reduces dependence on a large amount of training data, improves the flexibility and efficiency of speech synthesis, and is suitable for civil aviation land-air call simulation training, achieving more realistic and efficient voice interaction training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964545A_ABST
    Figure CN119964545A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and discloses a speech synthesis method and device based on few-sample learning, equipment and a medium, and the method comprises the steps: obtaining a training data set of a few-sample speech synthesis model, the training data set comprises a plurality of sections of training voice information of a plurality of speakers and training text information corresponding to each section of training voice; training a few-sample speech synthesis model based on the training data set; acquiring application text information of the to-be-synthesized voice; obtaining target speaker information of the to-be-synthesized voice; and inputting the application text information and the target speaker information into at least a sample speech synthesis model, outputting application speech information corresponding to the application text information, and making a sound by the application speech information according to the acoustic characteristics of the target speaker. According to the invention, the vivid voice with the acoustic features of the target speaker can be quickly generated under limited data samples, the dependence on a large amount of training data is reduced, and the flexibility and efficiency of voice synthesis are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a speech synthesis method, device, equipment and medium based on few-sample learning. Background Art

[0002] In the field of civil aviation, voice interaction between ground and air is an indispensable part of ensuring flight safety. The traditional training method for air traffic controllers (ATCO) relies heavily on manual simulation, that is, experienced pilots or simulators repeat the ATCO's instructions face to face for simulation training. However, this training method has many shortcomings.

[0003] First, the manual simulation training method consumes a lot of resources. Each training session requires the organization of professional personnel to participate, and the preparation of corresponding training venues and equipment, which increases the overall training cost. At the same time, since the training process needs to be conducted face-to-face, there are also geographical restrictions.

[0004] Secondly, the particularity of Chinese civil aviation ground-to-air conversations makes text-to-speech pairing data scarce. Compared with general speech synthesis, ground-to-air conversations in the civil aviation field have more stringent specifications and terminology requirements, which makes high-quality text-to-speech pairing data more difficult to obtain. Existing speech synthesis technologies often rely on a large amount of training data to generate realistic speech, but in the Chinese civil aviation ground-to-air conversation scenario, due to the scarcity of data, it is difficult for the model to fully learn and generate speech that meets the requirements, thus affecting the training effect. In addition, ATCOs need to interact with pilots with different characteristics in actual work, which requires the speech synthesis system to have a high degree of speaker adaptability. This further increases the difficulty of speech synthesis based on few-sample learning.

[0005] Therefore, there is an urgent need to develop a new speech synthesis solution for few-shot learning to overcome the above shortcomings. Summary of the invention

[0006] Based on this, it is necessary to address the problem that the existing speech synthesis based on few-sample learning is difficult to generate realistic speech, and propose a speech synthesis method, device, equipment and medium method based on few-sample learning.

[0007] A first aspect of the present invention provides a speech synthesis method based on few-sample learning, the method comprising:

[0008] Obtaining a training data set for a few-sample speech synthesis model, the training data set comprising a plurality of training speech information segments of a plurality of speakers and training text information corresponding to each training speech segment;

[0009] Training the few-sample speech synthesis model based on the training data set;

[0010] Get the application text information of the speech to be synthesized;

[0011] Obtain target speaker information of the speech to be synthesized;

[0012] The application text information and the target speaker information are input into the few-sample speech synthesis model, and application speech information corresponding to the application text information is output, where the application speech information is voiced using the acoustic features of the target speaker.

[0013] Furthermore, the step of training the few-sample speech synthesis model based on the training data set includes:

[0014] Construct an initial speech synthesis model and set initial model parameters;

[0015] The training speech information and training text information corresponding to each speaker in the training data set are used as a task training data, and each task training data corresponds to a task; a plurality of task training data are randomly sampled from the training data set as a first task set, and the remaining task training data are used as a second task set;

[0016] Performing inner loop optimization, the inner loop optimization comprising inputting the training data of each task in the first task set into the initial speech synthesis model for training, adjusting the initial model parameters by using a gradient descent method, and obtaining the optimal task parameters corresponding to each task;

[0017] performing outer loop optimization, wherein the outer loop optimization includes randomly selecting a task from the second task set as a new task, calculating a loss function gradient of the new task relative to each of the optimal task parameters, and updating the initial model parameters to optimized model parameters according to the loss function gradient;

[0018] Through multiple iterative optimizations of inner and outer loops, final model parameters that meet preset generalization requirements are obtained, and an initial speech synthesis model using the final model parameters is used as the few-sample speech synthesis model.

[0019] Furthermore, the step of inputting the training data of each task in the first task set into the initial speech synthesis model for training, adjusting the initial model parameters by using the gradient descent method, and obtaining the optimal task parameters corresponding to each task includes:

[0020] randomly selecting task training data corresponding to a task from the first task set, taking the selected task as the current task, taking the task training data corresponding to the current task as the current task training data, and taking the speaker corresponding to the current task as the current speaker;

[0021] Dividing the current task training data into a support set and a query set according to a preset ratio, wherein the support set contains a first number of groups of training voice information and corresponding training text information, and the query set contains a second number of groups of training voice information and corresponding training text information;

[0022] Preprocessing the training speech information in the current task data to extract acoustic features of the current speaker, wherein the acoustic features include speaker-level features, sentence-level features, and phoneme-level features;

[0023] Taking the acoustic features in the support set and the corresponding training text information as input, training the initial speech synthesis model to learn the acoustic features of the current speaker, adjusting the initial model parameters by using a gradient descent method to obtain a temporary speech synthesis model, and saving the current model parameters;

[0024] Randomly selecting training text information from a query set as test text information, and inputting it into the temporary speech synthesis model to output test speech information, and evaluating the test speech information according to the training speech information corresponding to the test text information in the query set to obtain an evaluation result;

[0025] Determine whether the temporary speech synthesis model is qualified according to the evaluation result;

[0026] If so, the current model parameters corresponding to the temporary speech synthesis model are used as the optimal task parameters.

[0027] Furthermore, after the step of obtaining the target speaker information of the speech to be synthesized, the method further includes:

[0028] Determining whether the target speaker is included in the training data set;

[0029] If not, a request command to obtain the target speaker's voice is issued;

[0030] receiving training speech information of a target speaker returned based on the request instruction;

[0031] Generating corresponding training text information according to the training voice information of the target speaker;

[0032] The training speech information and corresponding training text information of the target speaker are added to the training data set, and the few-sample speech synthesis model is updated.

[0033] Furthermore, the step of adding the training speech information and the corresponding training text information of the target speaker to the training data set to update the few-sample speech synthesis model includes:

[0034] adding the training speech information and the corresponding training text information of the target speaker to the first task set of the training data set, and performing the inner loop optimization on the few-sample speech synthesis model to update the few-sample speech synthesis model; or,

[0035] The training speech information and corresponding training text information of the target speaker are added to the second task set of the training data set, and the outer loop optimization is performed on the few-sample speech synthesis model to update the few-sample speech synthesis model.

[0036] Furthermore, the step of constructing an initial speech synthesis model includes:

[0037] The architecture of the initial speech synthesis model is set, and the initial speech synthesis model includes a phoneme embedding module, an encoder module, an acoustic conditional modeling module, a variable adapter module and a Mel-spectrogram decoding module; wherein the input of the initial speech synthesis model is text information, the text information is converted into a phoneme embedding vector through the phoneme embedding module, the phoneme embedding vector is converted into an encoding vector through the encoder module, the acoustic characteristics of the target speaker are generated for the encoding vector through the acoustic conditional modeling module, the acoustic characteristics of the target speaker and the encoding vector are converted into an acoustic feature vector through the variable adapter module, and the acoustic feature vector is converted into synthetic speech information through the Mel-spectrogram decoding module.

[0038] Furthermore, the step of obtaining application text information of the speech to be synthesized includes:

[0039] Determining whether the text in the application text information is text;

[0040] If it is text, traverse the application text information to find out whether there is a space between adjacent texts;

[0041] If there is no space, insert a space between adjacent words where there is no space;

[0042] According to a preset dictionary, the characters in the application text information are sequentially converted into corresponding phonemes to obtain phoneme text information, and the phoneme text information is used as new application text information.

[0043] A second aspect of the present invention provides a speech synthesis device based on few-sample learning, the device comprising:

[0044] A training data acquisition module, used to acquire a training data set for a few-sample speech synthesis model, wherein the training data set includes a plurality of training speech information segments of a plurality of speakers and training text information corresponding to each training speech segment;

[0045] A model training module, used for training the few-sample speech synthesis model based on the training data set;

[0046] A text acquisition module is used to acquire application text information of the speech to be synthesized;

[0047] A speaker determination module is used to obtain the target speaker information of the speech to be synthesized;

[0048] The speech synthesis module is used to input the application text information and the target speaker information into the few-sample speech synthesis model, and output application speech information corresponding to the application text information, wherein the application speech information is voiced with the acoustic features of the target speaker.

[0049] A third aspect of the present invention provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the following steps:

[0050] Obtaining a training data set for a few-sample speech synthesis model, the training data set comprising a plurality of training speech information segments of a plurality of speakers and training text information corresponding to each training speech segment;

[0051] Training the few-sample speech synthesis model based on the training data set;

[0052] Get the application text information of the speech to be synthesized;

[0053] Obtain target speaker information of the speech to be synthesized;

[0054] The application text information and the target speaker information are input into the few-sample speech synthesis model, and application speech information corresponding to the application text information is output, where the application speech information is voiced using the acoustic features of the target speaker.

[0055] A computer-readable storage medium stores a computer program, which, when executed by a processor, causes the processor to perform the following steps:

[0056] Obtaining a training data set for a few-sample speech synthesis model, the training data set comprising a plurality of training speech information segments of a plurality of speakers and training text information corresponding to each training speech segment;

[0057] Training the few-sample speech synthesis model based on the training data set;

[0058] Get the application text information of the speech to be synthesized;

[0059] Obtain target speaker information of the speech to be synthesized;

[0060] The application text information and the target speaker information are input into the few-sample speech synthesis model, and application speech information corresponding to the application text information is output, where the application speech information is voiced using the acoustic features of the target speaker.

[0061] The few-sample speech synthesis method, device, equipment and medium of the present invention, the method trains a few-sample speech synthesis model; then uses the few-sample speech synthesis model to generate application speech information corresponding to the application text information according to the application text information and the target speaker information of the speech to be synthesized, and the application speech information is pronounced with the acoustic characteristics of the target speaker. By training the few-sample speech synthesis model, the present invention can quickly generate realistic speech with the acoustic characteristics of the target speaker with limited data samples, reduces the dependence on a large amount of training data, and improves the flexibility and efficiency of speech synthesis. In civil aviation ground-to-air call simulation training, this method enables ATCO and pilots to conduct more realistic and efficient speech interaction training in different scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0063] in:

[0064] Figure 1 is an application environment diagram of a speech synthesis method based on few-sample learning in one embodiment;

[0065] Figure 2 is a flowchart of a speech synthesis method based on few-sample learning in one embodiment;

[0066] Figure 3 is a structural block diagram of a speech synthesis device based on few-sample learning in one embodiment;

[0067] Figure 4 FIG. 4 is a structural block diagram of a computer device in one embodiment. DETAILED DESCRIPTION

[0068] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0069] Figure 1 FIG. 1 is an application environment diagram of a speech synthesis method based on few-sample learning in one embodiment. Figure 1 , the speech synthesis method based on few-sample learning is applied to a speech synthesis system based on few-sample learning. The speech synthesis system based on few-sample learning includes a terminal 110 and a server 120. The terminal 110 and the server 120 are connected via a network, and the terminal 110 can be a desktop terminal or a mobile terminal, and the mobile terminal can be at least one of a mobile phone, a tablet computer, a laptop computer, etc. The server 120 can be implemented as an independent server or a server cluster composed of multiple servers. The terminal 110 is used for speech synthesis and interaction with the user, and the server 120 is used for training a few-sample speech synthesis model and storing related data.

[0070] The method of the embodiment of the present invention is particularly suitable for speech synthesis fields with a small number of samples, such as the civil aviation field for Chinese land-air conversations. Since the speech data of civil aviation land-air conversations is relatively limited, traditional speech synthesis systems are difficult to accurately capture the speech characteristics of different speakers.

[0071] like Figure 2 As shown, in one embodiment, a speech synthesis method based on few-sample learning is provided. The method can be applied to both a terminal and a server. This embodiment is illustrated by applying to a terminal. The speech synthesis method based on few-sample learning specifically includes the following steps:

[0072] S1: Obtain a training data set for a few-sample speech synthesis model, wherein the training data set includes a plurality of training speech information segments of a plurality of speakers and training text information corresponding to each training speech segment;

[0073] S2: Training the few-sample speech synthesis model based on the training data set;

[0074] S3: Obtain application text information of the speech to be synthesized;

[0075] S4: Obtaining target speaker information of the speech to be synthesized;

[0076] S5: Input the application text information and the target speaker information into the few-sample speech synthesis model, and output application speech information corresponding to the application text information, wherein the application speech information is voiced using the acoustic features of the target speaker.

[0077] In this embodiment, in the above step S1, taking the field of civil aviation air-ground communication as an example, the training data set is a corpus recorded and constructed by professionals in a quiet environment with reference to actual air-ground communication recording data (i.e., training voice information) and related teaching materials (i.e., training text information) under the guidance of civil aviation air traffic controllers and air-ground communication course instructors. The training samples of this training data set are scarce and the cost of obtaining them is high. This training data set is a repetition data set, and a sentence usually consists of instructions sent by air traffic controllers and instructions repeated by pilots. As a semi-artificial professional language, Chinese civil aviation air-ground communication exhibits unique voice and language characteristics of the industry, involving aspects such as pronunciation of terms, wording and sentence formation, and communication methods. Unlike other languages, in order to reduce the uncertainty caused by ambiguous pronunciation and terms, and to avoid safety hazards caused by unclear and effective information exchange between the two parties, Chinese air-ground communication complies with the civil aviation radio communication standards and has the following characteristics:

[0078] (1) Special pronunciation of words, such as the number 0 is pronounced as "dong", the number 1 is pronounced as "yao", and the number 7 is pronounced as "guai"; standard terms are used with single meanings, such as "approach" and "departure";

[0079] (2) The sentences are relatively long, usually including the control instructions sent by the air traffic controller and the pilot's repetition of the instructions;

[0080] (3) The sentence structure is concise and relatively fixed. Air traffic controllers and pilots need to communicate in the format of "the other party's number-their own number-the content of the conversation", for example, air traffic controller (South three five into runway two), pilot (two south three five);

[0081] (4) From an emotional perspective, the speaker's emotions are relatively stable, without showing any obvious emotions of joy, anger, sorrow or fear;

[0082] (5) From a stylistic perspective, it is more like a news report. Because the speaker is trained, there are no stylistic features like talk shows or commentary (interjections such as "ah" and "um" or obvious breathing or pauses);

[0083] (6) There are obvious differences in pitch, loudness, and speaking speed between different speakers, which is more obvious among announcers of different genders. Specifically, the speaker vectors are far apart.

[0084] In the above step S2, the few-sample speech synthesis model is trained by using the collected training data set. During the training process, the few-sample speech synthesis model will learn how to convert text information into speech information with specific speaker acoustic characteristics.

[0085] In the above steps S3-S4, during the dialogue process using the synthesized few-sample speech synthesis model, application text information and target speaker information are obtained to synthesize speech with the acoustic characteristics of the target speaker.

[0086] The application text information may be, for example, an ATCO instruction, a pilot's reply, or other text content that needs to be converted into speech.

[0087] The target speaker information may be, for example, information such as the speaker's identity, which can be associated with the acoustic features of the corresponding speaker from the few-sample speech synthesis model. The above steps S3 and S4 may be performed simultaneously, or S3 may be performed first and then S4. In other embodiments, S4 may be performed first and then S3, and the present invention does not make any special limitation on this.

[0088] In the above step S5, the application text information and the target speaker information are input into the trained few-sample speech synthesis model, and the model outputs the application speech information voiced with the acoustic characteristics of the target speaker.

[0089] The few-sample speech synthesis method of this embodiment can quickly generate realistic speech with the acoustic characteristics of the target speaker under limited data samples by training the few-sample speech synthesis model, reducing the dependence on a large amount of training data and improving the flexibility and efficiency of speech synthesis. In civil aviation ground-to-air call simulation training, this method can enable ATCOs and pilots to conduct more realistic and efficient speech interaction training in different scenarios.

[0090] In a specific embodiment, the step S2 of training the few-sample speech synthesis model based on the training data set includes:

[0091] S21: construct an initial speech synthesis model and set initial model parameters;

[0092] S22: taking the training speech information and training text information corresponding to each speaker in the training data set as a task training data, where each task training data corresponds to a task; randomly sampling a number of task training data from the training data set as a first task set, and the remaining task training data as a second task set;

[0093] S23: performing inner loop optimization, wherein the inner loop optimization includes inputting the training data of each task in the first task set into the initial speech synthesis model for training, and adjusting the initial model parameters by using a gradient descent method to obtain the optimal task parameters corresponding to each task;

[0094] S24: Execute outer loop optimization, wherein the outer loop optimization includes randomly selecting a task from the second task set as a new task, calculating a loss function gradient of the new task relative to each of the optimal task parameters, and updating the initial model parameters to optimized model parameters according to the loss function gradient;

[0095] S25: Through multiple iterative optimizations of inner and outer loops, final model parameters that meet preset generalization requirements are obtained, and an initial speech synthesis model using the final model parameters is used as the few-sample speech synthesis model.

[0096] In this embodiment, in the above step S21, a suitable basic model architecture is selected for training a few-sample speech synthesis model, such as a Transformer model based on deep learning, and parameters such as weights and biases of the model are initialized.

[0097] In the above step S22, the training voice information and training text information of each speaker are regarded as an independent task, and the tasks in the training data set are divided into a first task set for inner loop training optimization; the other part is a second task set for outer loop training optimization. Specifically, the training data set is traversed to create a task training data item for each speaker, including its corresponding training voice information and training text information; several items are randomly sampled from the task training data set as the first task set for inner loop training. The remaining task training data is used as the second task set for outer loop training and verification.

[0098] In the above step S23, the inner loop optimization is used to enable the model to achieve optimal performance on each task and find the optimal task parameters corresponding to each task. Exemplarily, each task training data item in the first task set is traversed; each task training data item is input into the initial speech synthesis model for forward propagation and back propagation; the initial parameters of the model are adjusted using a gradient descent method (such as SGD, Adam, etc.) to minimize the loss function on each task; and the optimal task parameters obtained after each task training is completed are recorded.

[0099] In a specific embodiment, the step S23 of inputting the training data of each task in the first task set into the initial speech synthesis model for training, adjusting the initial model parameters by using the gradient descent method, and obtaining the optimal task parameters corresponding to each task respectively includes:

[0100] S231: randomly selecting task training data corresponding to a task from the first task set, taking the selected task as the current task, taking the task training data corresponding to the current task as the current task training data, and taking the speaker corresponding to the current task as the current speaker;

[0101] S232: Dividing the current task training data into a support set and a query set according to a preset ratio, wherein the support set contains a first number of groups of training voice information and corresponding training text information, and the query set contains a second number of groups of training voice information and corresponding training text information;

[0102] S233: preprocessing the training speech information in the current task data to extract acoustic features of the current speaker, where the acoustic features include speaker-level features, sentence-level features, and phoneme-level features;

[0103] S234: using the acoustic features in the support set and the corresponding training text information as input, training the initial speech synthesis model to learn the acoustic features of the current speaker, adjusting the initial model parameters by using a gradient descent method to obtain a temporary speech synthesis model, and saving the current model parameters;

[0104] S235: randomly selecting training text information from the query set as test text information, and inputting the training text information into the temporary speech synthesis model to output test speech information, and evaluating the test speech information according to the training speech information corresponding to the test text information in the query set to obtain an evaluation result;

[0105] S236: judging whether the temporary speech synthesis model is qualified according to the evaluation result;

[0106] S237: If yes, the current model parameters corresponding to the temporary speech synthesis model are used as the optimal task parameters.

[0107] In the above step S24, the outer loop optimization is used to evaluate the performance of the model on the new task and update the initial model parameters according to the gradient of the loss function to improve the generalization ability of the model, so that the model can quickly adapt to and learn unseen tasks, and achieve rapid learning and adaptation to new tasks in the case of few samples, which is of great significance for solving challenges in fields such as small sample learning and transfer learning. Exemplarily, the following steps can be used to achieve this: randomly select a task from the second task set as a new task; use the optimal task parameters obtained in the inner loop optimization to perform forward propagation on the new task and calculate the loss function; calculate the gradient according to the loss function, that is, the derivative of the loss function with respect to the initial model parameters; use the gradient descent method to update the initial model parameters to minimize the loss function on the new task. This enables the model to perform well on unseen tasks.

[0108] In the above step S25, the model parameters are gradually adjusted through multiple iterations of inner loop and outer loop optimization until the model meets the preset generalization requirements. Specifically, the number of iterations or the iteration stop condition (such as the convergence of the loss function, etc.) is set, and the inner loop and outer loop optimization steps are repeated until the number of iterations or the iteration stop condition is met; after each iteration, the performance of the model on the validation set is evaluated to monitor the training progress and generalization ability of the model; when the model meets the preset generalization requirements, the iteration is stopped, and the initial speech synthesis model using the final model parameters is used as the few-sample speech synthesis model.

[0109] Through the above steps, this embodiment can use limited training data to quickly train a few-sample speech synthesis model with good generalization ability. The model can provide ATCO and pilots with a realistic and efficient voice interaction experience in civil aviation air-land call simulation training.

[0110] In a specific embodiment, after the step S4 of obtaining the target speaker information of the speech to be synthesized, the method further includes:

[0111] S6: Determine whether the target speaker is included in the training data set;

[0112] S7: If not, issuing a request instruction to obtain the target speaker's voice;

[0113] S8: receiving training speech information of the target speaker returned based on the request instruction;

[0114] S9: Generate corresponding training text information according to the training voice information of the target speaker;

[0115] S10: Add the training speech information and corresponding training text information of the target speaker to the training data set, and update the few-sample speech synthesis model.

[0116] In this embodiment, in the above step S6, it is checked whether the target speaker is already included in the training data set. If yes, the speech for the target speaker is directly generated by the few-sample speech synthesis model; if no, the model can be supplemented by obtaining a small amount of speech of the target speaker in the subsequent steps. Specifically, the unique identifier (such as ID, name, etc.) of the target speaker can be extracted; the training data set is traversed to check whether there is a data item that matches the identifier of the target speaker; if a match is found, the target speaker is already in the training data set; otherwise, the target speaker is not in the training data set.

[0117] In the above steps S7 and S8, if the target speaker is not in the training data set, a request instruction will be triggered to obtain the speech data of the target speaker. Specifically, a request instruction including the target speaker identifier can be generated; the request instruction is sent to a preset data management system, server or the target speaker himself; the received data stream or message queue is monitored; when a message including the speech data of the target speaker is detected, it is extracted and saved as the training speech information of the target speaker.

[0118] In the above step S9, if the target speaker provides text content synchronized with the voice data, the text content is directly used as the training text information. If the target speaker does not provide text content synchronized with the voice data, the voice data can be converted into text content using speech recognition technology as the training text information.

[0119] In the above step S10, the newly acquired training speech information of the target speaker and the corresponding training text information are added to the training data set, and the few-sample speech synthesis model is updated. When the few-sample speech synthesis model is updated, the target speaker information already exists in the model, and the updated few-sample speech synthesis model is used to perform step S5, and the application text information and the target speaker information are input into the model to generate the corresponding application speech information.

[0120] This embodiment can dynamically obtain the speech data and text information of the target speaker when the target speaker is not in the training data set, and quickly update the small sample speech synthesis model, thereby improving the flexibility and practicality of civil aviation land-to-air call simulation training.

[0121] In a specific embodiment, the step S10 of adding the training speech information and the corresponding training text information of the target speaker to the training data set and updating the few-sample speech synthesis model includes:

[0122] S101: adding the training speech information and the corresponding training text information of the target speaker to the first task set of the training data set, and performing the inner loop optimization on the few-sample speech synthesis model to update the few-sample speech synthesis model; or,

[0123] S102: Add the training speech information and corresponding training text information of the target speaker to the second task set of the training data set, and perform the outer loop optimization on the few-sample speech synthesis model to update the few-sample speech synthesis model.

[0124] This embodiment provides two methods for fine-tuning the few-sample speech synthesis model when a new task needs to be learned.

[0125] Among them, in step S101, the training voice information of the target speaker and its corresponding training text information are added to the first task set, and the optimal task parameters of the task corresponding to the target speaker are found by executing the inner loop optimization step S23. The few-sample speech synthesis model is quickly iterated and trained by the task data added to the first task set, so as to quickly adjust the model parameters in a short time to learn the acoustic features of the newly added speaker without large-scale retraining of the entire model.

[0126] In the above step S102, the training speech information of the target speaker and its corresponding training text information are added to the second task set, and the step of outer loop optimization S24 is executed, that is, the loss function gradient of the task corresponding to the target speaker relative to the optimal task parameters of the few-sample speech synthesis model is calculated, and the model parameters of the entire model are adjusted according to the loss function gradient, thereby improving the long-term learning ability of the model, ensuring that the model maintains good generalization ability, and reducing the risk of overfitting.

[0127] In a specific embodiment, the step S21 of constructing an initial speech synthesis model includes:

[0128] S211: Setting the architecture of the initial speech synthesis model, the initial speech synthesis model includes a phoneme embedding module, an encoder module, an acoustic conditional modeling module, a variable adapter module and a Mel-spectrogram decoding module; wherein the input of the initial speech synthesis model is text information, the text information is converted into a phoneme embedding vector through the phoneme embedding module, the phoneme embedding vector is converted into an encoding vector through the encoder module, the acoustic features of the target speaker are generated for the encoding vector through the acoustic conditional modeling module, the acoustic features of the target speaker and the encoding vector are converted into an acoustic feature vector through the variable adapter module, and the acoustic feature vector is converted into synthetic speech information through the Mel-spectrogram decoding module.

[0129] In this embodiment, the phoneme embedding module converts the input text information (such as characters, words, sentences, phoneme sequences, etc.) into a phoneme embedding vector. Phoneme is the basic unit of speech, which can usually be completed using a pre-trained phoneme embedding table or by learning a phoneme embedding matrix. It can also be converted into a phoneme text by pre-processing Chinese character texts, and then input into the phoneme embedding module for phoneme embedding. The encoder module receives the phoneme embedding vector as input and converts it into a coding vector, extracting semantic and contextual information in the text during the coding process.

[0130] The acoustic conditional modeling module generates acoustic features for the target speaker. During the model training phase, the acoustic conditional modeling module uses the speaker embedding vector as input to extract the speaker's acoustic features. Acoustic features include speaker-level features, sentence-level features, and phoneme-level features. Speaker-level features include the speaker's voice characteristics, pronunciation habits, intonation, timbre, etc., which provide the model with the ability to generalize to different speakers' speech. The sentence level extracts a sentence-level feature vector sequence from the reference speech through an acoustic encoder. The phoneme level uses another acoustic encoder to extract a phoneme-level feature vector sequence from the target speech. The sentence-level feature vector sequence and the phoneme-level feature vector sequence are introduced into the mel-spectrogram decoder as input, representing the global and local acoustic conditions, respectively. This helps the decoder predict speech under different acoustic conditions, thereby improving the generalization performance of the model and avoiding overfitting to specific acoustic conditions in the training data. During the use of the model, the acoustic conditional modeling module uses the target speaker's identifier as input, matches the target speaker, and generates corresponding acoustic features for the encoding vector.

[0131] The variable adapter module combines the acoustic features of the target speaker with the encoding vector generated by the encoder and converts it into an acoustic feature vector, which enables the model to adapt to different speaker styles. The Mel spectrum decoding module receives the acoustic feature vector as input and converts it into a Mel spectrum, which is then converted into the final synthesized speech information.

[0132] Under the above model framework, the specific execution process of step S5 (i.e., using the updated few-sample speech synthesis model to generate the synthesized speech information of the target speaker according to the input text information) is as follows:

[0133] Input text information and target speaker information; pass the text information to the phoneme embedding module to generate the corresponding phoneme embedding vector; input the phoneme embedding vector to the encoder module, and generate the encoding vector through the encoding process; input the encoding vector and the target speaker information to the acoustic conditional modeling module, and the acoustic conditional modeling module generates the corresponding acoustic features according to the target speaker information (such as speaker identification, etc.); the variable adapter module combines the acoustic features and the encoding vector to generate an acoustic feature vector; the Mel spectrum decoding module converts the acoustic feature vector into a Mel spectrum, and finally converts it into the final synthetic speech information with the style of the target speaker.

[0134] In a specific embodiment, the step S3 of obtaining application text information of the speech to be synthesized includes:

[0135] S31: Determine whether the text in the application text information is text;

[0136] S32: If it is text, traverse the application text information to find out whether there is a space between adjacent texts;

[0137] S33: if there is no space, insert a space between adjacent characters where there is no space;

[0138] S34: According to a preset dictionary, the characters in the application text information are sequentially converted into corresponding phonemes to obtain phoneme text information, and the phoneme text information is used as new application text information.

[0139] In this embodiment, the accuracy and speed of text-to-speech matching are improved by unifying the text intervals of the application text information and converting the text into phonemes.

[0140] like Figure 3 As shown, in one embodiment, a speech synthesis device based on few-sample learning is provided, the device comprising:

[0141] A training data acquisition module 10 is used to acquire a training data set for a few-sample speech synthesis model, wherein the training data set includes a plurality of training speech information segments of a plurality of speakers and training text information corresponding to each training speech segment;

[0142] A model training module 20, configured to train the few-sample speech synthesis model based on the training data set;

[0143] A text acquisition module 30, used to acquire application text information of the speech to be synthesized;

[0144] A speaker determination module 40 is used to obtain target speaker information of the speech to be synthesized;

[0145] The speech synthesis module 50 is used to input the application text information and the target speaker information into the few-sample speech synthesis model, and output application speech information corresponding to the application text information, wherein the application speech information is voiced with the acoustic features of the target speaker.

[0146] In one embodiment, the model training module 20 includes:

[0147] A model building unit, used to build an initial speech synthesis model and set initial model parameters;

[0148] A task division unit is used to use the training speech information and training text information corresponding to each speaker in the training data set as a task training data, and each task training data corresponds to a task; randomly sample a number of task training data from the training data set as a first task set, and the remaining task training data as a second task set;

[0149] An inner loop unit, used to perform inner loop optimization, wherein the inner loop optimization includes inputting the training data of each task in the first task set into the initial speech synthesis model for training, and adjusting the initial model parameters by using a gradient descent method to obtain the optimal task parameters corresponding to each task;

[0150] an outer loop unit, configured to perform outer loop optimization, wherein the outer loop optimization comprises randomly selecting a task from the second task set as a new task, calculating a loss function gradient of the new task relative to each of the optimal task parameters, and updating the initial model parameters to optimized model parameters according to the loss function gradient;

[0151] The optimization iteration unit is used to obtain final model parameters that meet the preset generalization requirements through multiple iterative optimizations of inner loops and outer loops, and the initial speech synthesis model using the final model parameters is used as the few-sample speech synthesis model.

[0152] In one embodiment, the inner circulation unit comprises:

[0153] A task selection subunit, configured to randomly select task training data corresponding to a task from the first task set, use the selected task as the current task, use the task training data corresponding to the current task as the current task training data, and use the speaker corresponding to the current task as the current speaker;

[0154] A data set division subunit, used to divide the current task training data into a support set and a query set according to a preset ratio, wherein the support set contains a first number of groups of training voice information and corresponding training text information, and the query set contains a second number of groups of training voice information and corresponding training text information;

[0155] A feature extraction subunit, used to pre-process the training speech information in the current task data, and extract the acoustic features of the current speaker, wherein the acoustic features include speaker-level features, sentence-level features, and phoneme-level features;

[0156] The task training subunit is used to take the acoustic features in the support set and the corresponding training text information as input, train the initial speech synthesis model to learn the acoustic features of the current speaker, adjust the initial model parameters by using the gradient descent method, obtain a temporary speech synthesis model, and save the current model parameters;

[0157] A query test subunit, configured to randomly select training text information from a query set as test text information, input the training text information into the temporary speech synthesis model to output test speech information, and evaluate the test speech information according to the training speech information corresponding to the test text information in the query set to obtain an evaluation result;

[0158] An evaluation subunit, configured to determine whether the temporary speech synthesis model is qualified according to the evaluation result;

[0159] The optimal task parameter determination subunit is used to take the current model parameters corresponding to the temporary speech synthesis model as the optimal task parameters if qualified.

[0160] In a specific embodiment, the device further comprises:

[0161] A target speaker determination module, used to determine whether the target speaker is included in the training data set;

[0162] A request instruction module, used for issuing a request instruction to obtain the speech of the target speaker if it is not included in the training data set;

[0163] A training voice receiving module, used for receiving the training voice information of the target speaker returned based on the request instruction;

[0164] A training text generation module, used to generate corresponding training text information according to the training voice information of the target speaker;

[0165] The model update training module is used to add the training speech information and corresponding training text information of the target speaker to the training data set to update the few-sample speech synthesis model.

[0166] In a specific embodiment, the model update training module includes:

[0167] An inner loop updating unit, configured to add the training speech information and the corresponding training text information of the target speaker to the first task set of the training data set, and perform the inner loop optimization on the few-sample speech synthesis model to update the few-sample speech synthesis model;

[0168] An outer loop updating unit is used to add the training speech information and corresponding training text information of the target speaker to the second task set of the training data set, and perform the outer loop optimization on the few-sample speech synthesis model to update the few-sample speech synthesis model.

[0169] In a specific embodiment, the model building unit includes:

[0170] An architecture construction subunit is used to set the architecture of the initial speech synthesis model, wherein the initial speech synthesis model includes a phoneme embedding module, an encoder module, an acoustic conditional modeling module, a variable adapter module and a Mel-spectrogram decoding module; wherein the input of the initial speech synthesis model is text information, the text information is converted into a phoneme embedding vector through the phoneme embedding module, the phoneme embedding vector is converted into an encoding vector through the encoder module, the acoustic features of the target speaker are generated for the encoding vector through the acoustic conditional modeling module, the acoustic features of the target speaker and the encoding vector are converted into an acoustic feature vector through the variable adapter module, and the acoustic feature vector is converted into synthetic speech information through the Mel-spectrogram decoding module.

[0171] In a specific embodiment, the text acquisition module 30 includes:

[0172] A text determination unit, used to determine whether the text in the application text information is text;

[0173] A space traversal unit, used for traversing the application text information if it is text, and finding whether there is a space between adjacent texts;

[0174] A space insertion unit is used to insert a space between adjacent characters where no space exists if no space exists;

[0175] The phoneme conversion unit is used to convert the characters in the application text information into corresponding phonemes in sequence according to a preset dictionary to obtain phoneme text information, and use the phoneme text information as new application text information.

[0176] The few-sample speech synthesis device of this embodiment can quickly generate realistic speech with the acoustic characteristics of the target speaker under limited data samples by training the few-sample speech synthesis model, reducing the dependence on a large amount of training data and improving the flexibility and efficiency of speech synthesis. In civil aviation ground-to-air call simulation training, this method can enable ATCO and pilots to conduct more realistic and efficient speech interaction training in different scenarios.

[0177] Figure 4 FIG. 1 shows an internal structure diagram of a computer device in an embodiment. The computer device may be a terminal or a server. Figure 4As shown, the computer device includes a processor, a memory and a network interface connected via a system bus. The memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium of the computer device stores an operating system and may also store a computer program. When the computer program is executed by the processor, the processor may implement a speech synthesis method based on few-sample learning. The internal memory may also store a computer program. When the computer program is executed by the processor, the processor may implement a speech synthesis method based on few-sample learning. Those skilled in the art will understand that Figure 4 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present invention, and does not constitute a limitation on the computer device to which the solution of the present invention is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0178] In one embodiment, a computer device is provided, comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the following steps:

[0179] Obtaining a training data set for a few-sample speech synthesis model, the training data set comprising a plurality of training speech information segments of a plurality of speakers and training text information corresponding to each training speech segment;

[0180] Training the few-sample speech synthesis model based on the training data set;

[0181] Get the application text information of the speech to be synthesized;

[0182] Obtain target speaker information of the speech to be synthesized;

[0183] The application text information and the target speaker information are input into the few-sample speech synthesis model, and application speech information corresponding to the application text information is output, where the application speech information is voiced using the acoustic features of the target speaker.

[0184] The few-sample speech synthesis method of this embodiment can quickly generate realistic speech with the acoustic characteristics of the target speaker under limited data samples by training the few-sample speech synthesis model, reducing the dependence on a large amount of training data and improving the flexibility and efficiency of speech synthesis. In civil aviation ground-to-air call simulation training, this method can enable ATCOs and pilots to conduct more realistic and efficient speech interaction training in different scenarios.

[0185] In one embodiment, a computer-readable storage medium is provided, storing a computer program, wherein when the computer program is executed by a processor, the processor performs the following steps:

[0186] Obtaining a training data set for a few-sample speech synthesis model, the training data set comprising a plurality of training speech information segments of a plurality of speakers and training text information corresponding to each training speech segment;

[0187] Training the few-sample speech synthesis model based on the training data set;

[0188] Get the application text information of the speech to be synthesized;

[0189] Obtain target speaker information of the speech to be synthesized;

[0190] The application text information and the target speaker information are input into the few-sample speech synthesis model, and application speech information corresponding to the application text information is output, where the application speech information is voiced using the acoustic features of the target speaker.

[0191] The few-sample speech synthesis method of this embodiment can quickly generate realistic speech with the acoustic characteristics of the target speaker under limited data samples by training the few-sample speech synthesis model, reducing the dependence on a large amount of training data and improving the flexibility and efficiency of speech synthesis. In civil aviation ground-to-air call simulation training, this method can enable ATCOs and pilots to conduct more realistic and efficient speech interaction training in different scenarios.

[0192] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by instructing the relevant hardware through a computer program, and the program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided by the present invention can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synch li nk) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0193] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0194] The above-mentioned embodiments only express several implementation methods of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the patent of the present invention. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.

Claims

1. A speech synthesis method based on few-sample learning, characterized in that: The method comprises: Obtaining a training data set for a few-sample speech synthesis model, the training data set comprising a plurality of training speech information segments of a plurality of speakers and training text information corresponding to each training speech segment; Training the few-sample speech synthesis model based on the training data set; Get the application text information of the speech to be synthesized; Obtaining target speaker information of the speech to be synthesized; The application text information and the target speaker information are input into the few-sample speech synthesis model, and application speech information corresponding to the application text information is output, where the application speech information is voiced using the acoustic features of the target speaker.

2. The speech synthesis method based on few-sample learning according to claim 1, characterized in that: The step of training the few-sample speech synthesis model based on the training data set comprises: Construct an initial speech synthesis model and set initial model parameters; The training speech information and training text information corresponding to each speaker in the training data set are used as a task training data, and each task training data corresponds to a task; a plurality of task training data are randomly sampled from the training data set as a first task set, and the remaining task training data are used as a second task set; Performing inner loop optimization, the inner loop optimization comprising inputting the training data of each task in the first task set into the initial speech synthesis model for training, adjusting the initial model parameters by using a gradient descent method, and obtaining the optimal task parameters corresponding to each task; performing outer loop optimization, wherein the outer loop optimization includes randomly selecting a task from the second task set as a new task, calculating a loss function gradient of the new task relative to each of the optimal task parameters, and updating the initial model parameters to optimized model parameters according to the loss function gradient; Through multiple iterative optimizations of inner and outer loops, final model parameters that meet preset generalization requirements are obtained, and an initial speech synthesis model using the final model parameters is used as the few-sample speech synthesis model.

3. The speech synthesis method based on few-sample learning according to claim 2, characterized in that: The step of inputting the training data of each task in the first task set into the initial speech synthesis model for training, and adjusting the initial model parameters by using the gradient descent method to obtain the optimal task parameters corresponding to each task, comprises: randomly selecting task training data corresponding to a task from the first task set, taking the selected task as the current task, taking the task training data corresponding to the current task as the current task training data, and taking the speaker corresponding to the current task as the current speaker; Dividing the current task training data into a support set and a query set according to a preset ratio, wherein the support set contains a first number of groups of training voice information and corresponding training text information, and the query set contains a second number of groups of training voice information and corresponding training text information; Preprocessing the training speech information in the current task data to extract acoustic features of the current speaker, wherein the acoustic features include speaker-level features, sentence-level features, and phoneme-level features; Taking the acoustic features in the support set and the corresponding training text information as input, training the initial speech synthesis model to learn the acoustic features of the current speaker, adjusting the initial model parameters by using a gradient descent method to obtain a temporary speech synthesis model, and saving the current model parameters; Randomly selecting training text information from a query set as test text information, and inputting it into the temporary speech synthesis model to output test speech information, and evaluating the test speech information according to the training speech information in the query set corresponding to the test text information to obtain an evaluation result; Determine whether the temporary speech synthesis model is qualified according to the evaluation result; If so, the current model parameters corresponding to the temporary speech synthesis model are used as the optimal task parameters.

4. The speech synthesis method based on few-sample learning according to claim 2, characterized in that: After the step of obtaining the target speaker information of the speech to be synthesized, the method further includes: Determining whether the target speaker is included in the training data set; If not, a request command to obtain the target speaker's voice is issued; receiving training speech information of a target speaker returned based on the request instruction; Generating corresponding training text information according to the training voice information of the target speaker; The training speech information and corresponding training text information of the target speaker are added to the training data set, and the few-sample speech synthesis model is updated.

5. The speech synthesis method based on few-sample learning according to claim 4, characterized in that: The step of adding the training speech information and the corresponding training text information of the target speaker to the training data set and updating the few-sample speech synthesis model comprises: adding the training speech information and the corresponding training text information of the target speaker to the first task set of the training data set, and performing the inner loop optimization on the few-sample speech synthesis model to update the few-sample speech synthesis model; or, The training speech information and corresponding training text information of the target speaker are added to the second task set of the training data set, and the outer loop optimization is performed on the few-sample speech synthesis model to update the few-sample speech synthesis model.

6. The speech synthesis method based on few-sample learning according to claim 2, characterized in that: The step of constructing an initial speech synthesis model comprises: The architecture of the initial speech synthesis model is set, and the initial speech synthesis model includes a phoneme embedding module, an encoder module, an acoustic conditional modeling module, a variable adapter module and a Mel-spectrogram decoding module; wherein the input of the initial speech synthesis model is text information, the text information is converted into a phoneme embedding vector through the phoneme embedding module, the phoneme embedding vector is converted into an encoding vector through the encoder module, the acoustic characteristics of the target speaker are generated for the encoding vector through the acoustic conditional modeling module, the acoustic characteristics of the target speaker and the encoding vector are converted into an acoustic feature vector through the variable adapter module, and the acoustic feature vector is converted into synthetic speech information through the Mel-spectrogram decoding module.

7. The speech synthesis method based on few-sample learning according to claim 1, characterized in that: The step of obtaining application text information of the speech to be synthesized includes: Determining whether the text in the application text information is text; If it is text, traverse the application text information to find out whether there is a space between adjacent texts; If there is no space, insert a space between adjacent words where there is no space; According to a preset dictionary, the characters in the application text information are sequentially converted into corresponding phonemes to obtain phoneme text information, and the phoneme text information is used as new application text information.

8. A speech synthesis device based on few-sample learning, characterized in that: The device comprises: A training data acquisition module, used to acquire a training data set for a few-sample speech synthesis model, wherein the training data set includes a plurality of training speech information segments of a plurality of speakers and training text information corresponding to each training speech segment; A model training module, used for training the few-sample speech synthesis model based on the training data set; A text acquisition module is used to acquire application text information of the speech to be synthesized; A speaker determination module is used to obtain the target speaker information of the speech to be synthesized; The speech synthesis module is used to input the application text information and the target speaker information into the few-sample speech synthesis model, and output application speech information corresponding to the application text information, wherein the application speech information is voiced with the acoustic features of the target speaker.

9. A computer-readable storage medium, characterized in that: A computer program is stored, and when the computer program is executed by a processor, the processor executes the steps of speech synthesis based on few-sample learning as described in any one of claims 1 to 7.

10. A computer device, characterized in that: The device includes a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the speech synthesis method based on few-sample learning as described in any one of claims 1 to 7.