Data processing method, speech synthesis method and related equipment

By introducing instruction text with speech rate indication information during the training process of the speech synthesis model and utilizing an improved bimodal data processing model and speech synthesis technology, the limitations of the speech synthesis model in speech rate adjustment are overcome, the flexibility and naturalness of speech synthesis are achieved, and the user experience is improved.

CN120690169APending Publication Date: 2025-09-23MASHANG CONSUMER FINANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510286434.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing intelligent speech synthesis models have significant limitations in speech rate adjustment and cannot meet users' needs for intelligent and flexible speech synthesis.

Method used

By introducing instruction text containing speaking rate indication information when training the speech synthesis model, the model can learn to parse and respond to these instructions, thereby achieving flexible adjustment of speaking rate. By using an improved bimodal data processing model and speech synthesis technology, combining a large text model with discrete speech features, realistic speech is generated.

Benefits of technology

The speech synthesis model is implemented to instantly adjust the speech speed according to the user's speech speed instructions, improving the user's experience satisfaction during the intelligent voice interaction process, making the interaction more natural and smooth, and meeting the user's personalized needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120690169A_ABST
    Figure CN120690169A_ABST
Patent Text Reader

Abstract

The invention relates to artificial intelligence, and provides a data processing method, a speech synthesis method and related equipment. According to the method, first sample data can be encoded by using a model to obtain a first encoding vector, and the first sample data comprises first voice data; on the basis of the first coding vector, predicting the first voice data by using the model to obtain first predicted voice; determining a first loss value of the model based on the first voice data, the first predicted voice and a preset first loss function; and training the model according to the first loss value. The method can improve the accuracy of voice generation by the model and the flexibility of voice speed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and specifically to a data processing method, a speech synthesis method, and related equipment. Background Art

[0002] In intelligent voice interaction scenarios, users expect to have natural and fluent conversations with smart devices. However, the speech synthesis models used by these smart devices typically synthesize speech at a pre-set fixed speaking rate. This results in significant limitations in the speed adjustment of these speech synthesis methods, making them unable to meet users' demands for intelligent and flexible speech synthesis. Summary of the Invention

[0003] The present application provides a data processing method, a speech synthesis method and related equipment to solve the problem that the above-mentioned speech synthesis method usually relies on a fixed speaking rate, has great limitations in speech rate adjustment, and cannot meet users' needs for intelligent and flexible speech synthesis.

[0004] A first aspect of an embodiment of the present application provides a data processing method, comprising: encoding first sample data using a model to obtain a first encoding vector, wherein the first sample data includes first speech data; based on the first encoding vector, predicting the first speech data using the model to obtain a first predicted speech; determining a first loss value of the model based on the first speech data, the first predicted speech and a preset first loss function; and training the model according to the first loss value.

[0005] A second aspect of an embodiment of the present application provides a speech synthesis method, which includes: recognizing a user's input speech to obtain a recognized text; if the recognized text contains preset keywords, determining an instruction text corresponding to the preset keywords, wherein the instruction text is used to indicate the user's expected speaking speed; based on the instruction text and the text to be output, generating speech discrete features using a speech generation model; synthesizing the speech discrete features using speech synthesis technology, and outputting synthesized speech data.

[0006] A third aspect of an embodiment of the present application provides a chip system, which is applied to a computer device. The chip system includes one or more processors, which are used to call computer instructions to enable the computer device to input training data into the chip system and execute the steps in the method provided in the first or second aspect above.

[0007] A fourth aspect of an embodiment of the present application provides a data processing device, which includes: an encoding unit for encoding first sample data using a model to obtain a first encoding vector, wherein the first sample data includes first speech data; a prediction unit for predicting the first speech data using the model based on the first encoding vector to obtain a first predicted speech; a determination unit for determining a first loss value of the model based on the first speech data, the first predicted speech and a preset first loss function; and a training unit for training the model according to the first loss value.

[0008] A fifth aspect of an embodiment of the present application provides a computer device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method provided in the first or second aspect above when executing the computer program.

[0009] A sixth aspect of an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the method provided in the first or second aspect are implemented.

[0010] A seventh aspect of the embodiments of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the steps in the method provided in the first or second aspect above.

[0011] In the data processing method of this embodiment, the instruction text containing speech speed indication information can be integrated into the training process when training the speech synthesis model as an important reference for model learning. Through this method, the model can learn to parse and respond to the speech speed instructions in these instruction texts, thereby having the ability to generate speech at the corresponding speed according to the instructions. In the application stage, the model can adjust the speech speed of its output speech in real time according to the user's speech speed instructions, achieving a high degree of flexibility in speech speed adjustment. In this way, the user's experience satisfaction in the intelligent voice interaction process can be significantly improved, making the interaction more natural and smooth, in line with the user's personalized expectations. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0013] Figure 1 This is a flowchart of a pre-training method for a model provided in one embodiment of the present application; Figure 2 is an example diagram of second sample data provided by an embodiment of the present application; Figure 3 This is an example diagram of data distribution of the TAE layer provided by an embodiment of the present application; Figure 4 is a flow chart of a data processing method provided by an embodiment of the present application; Figure 5 This is a flowchart of a method for constructing first sample data provided by an embodiment of the present application; Figure 6 is an example diagram of first sample data provided by an embodiment of the present application; Figure 7 is a flowchart of a method for predicting first voice data provided by an embodiment of the present application; Figure 8 This is a flowchart of a speech synthesis method provided by an embodiment of the present application; Figure 9 is a functional module diagram of a data processing device provided by an embodiment of the present application; Figure 10 It is a structural diagram of a computer device for implementing a data processing method provided in one embodiment of the present application. DETAILED DESCRIPTION

[0014] In order to more clearly understand the above-mentioned objectives, features and advantages of the present application, the present application is described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that the embodiments of the present application and the features therein can be combined with each other in the absence of conflict.

[0015] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application pertains. The terms used herein in the specification of this application are for the purpose of describing embodiments in some embodiments only and are not intended to limit this application.

[0016] It should be noted that, in this application, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A alone, A and B together, and B alone, where A and B can be singular or plural. The terms "first," "second," "third," "fourth," and so on (if any) in the specification, claims, and drawings of this application are used to distinguish similar objects, not to describe a specific order or precedence.

[0017] In the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a concrete manner. The following embodiments and features in the embodiments may be combined with each other unless there is a conflict.

[0018] In some embodiments, in intelligent voice interaction scenarios, users seek a highly natural and flexible conversational experience, which requires smart devices to not only accurately understand user commands but also respond in a manner that meets user expectations. Speech synthesis technology is a key link in achieving this goal, and its performance directly affects the quality of the user experience. However, while currently widely used speech synthesis models have made significant progress in sound quality and naturalness, they still face certain limitations in regulating speech speed.

[0019] In most intelligent voice interaction systems, the speech synthesis module typically generates speech based on preset speech rate parameters. This preset speech rate may be set based on general user preferences or considerations for specific application scenarios. However, user speech rate preferences vary. Some prefer fast-paced speech, believing that this improves efficiency, while others prefer slower speech rates, believing that this makes it easier to understand and absorb information. Therefore, a fixed speech rate setting is difficult to meet the personalized needs of all users.

[0020] When smart devices convert text into speech and output it to users based on a preset speaking speed, if the user thinks the speaking speed is inappropriate, the system faces a challenge: how to dynamically adjust the speaking speed to meet the user's immediate needs without retraining or adjusting the entire speech synthesis model. Traditional speech synthesis methods often lack this flexibility.

[0021] In order to solve the above problems, an embodiment of the present application provides a data processing method, which can integrate instruction texts containing speech rate indication information into the training process when training a speech synthesis model, as an important reference for model learning. Through this method, the model can learn to parse and respond to the speech rate instructions in these instruction texts, thereby having the ability to generate speech at the corresponding speed according to the instructions. In the application stage, the model can adjust the speech rate of its output speech in real time according to the user's speech rate instructions, achieving a high degree of flexibility in speech rate adjustment. In this way, the user's experience satisfaction in the process of intelligent voice interaction can be significantly improved, making the interaction more natural and smooth, in line with the user's personalized expectations.

[0022] The data processing method provided in the embodiments of the present application can be executed by a computer device, specifically by a processor of the computer device. The computer device may include a terminal device or a server. The terminal device may be a mobile phone, a tablet computer, a desktop computer, a portable notebook, etc. The server may be an independent server, a server cluster consisting of multiple servers, or a cloud server capable of cloud computing.

[0023] The data processing method provided in the embodiments of the present application can be divided into two parts. The first part includes pre-training the model so that the model can process text data and voice data simultaneously and synthesize corresponding voice based on the text data. The second part includes further fine-tuning the pre-trained model so that the model can generate voice data that is consistent with the content of the text data and has a speech speed consistent with the speech speed in the instruction text based on the speech speed indication information in the instruction text. Next, the structure of the model and the pre-training process will be introduced.

[0024] In some embodiments, to enable the model to simultaneously process and analyze both text and audio data, thereby achieving accurate text-based speech synthesis, the present application designs a bimodal data processing model specifically for processing text and audio. This model uses the relevant large text model as its foundational framework and makes targeted improvements to it, transforming the text encoding module into a bimodal data encoding module.

[0025] In one example, a large language model (Large Language Model Meta AI, LLaMA) is used as the basic framework. Since the LLaMA model is designed to process unimodal data (i.e., only text data), the text encoding module (such as the text embedding module embed_tokens) built into the LLaMA model can only be used for text embedding. However, the goal of the embodiment of the present application is to process bimodal data containing two modalities, text and speech. Therefore, the network of the LLaMA model is adaptively modified. The specific improvement measures are as follows: remove the original text embedding module in the LLaMA network, because the text embedding module is only applicable to unimodal text data processing, which does not meet the requirements of the embodiment of the present application for bimodal data processing. In order to replace the original text embedding module, a new bimodal embedding layer TAE (Text-Audio Embedding) is introduced. The TAE layer can be used to embed text and audio data into the same vector space for subsequent network processing.

[0026] In one example, the following embeddings are defined in the TAE layer: Sequence start position embedding (sosembedding): Located at position 0, it marks the beginning of the sequence. Speaker embedding: Located at position 1, it is extracted using a pre-trained speaker recognition model and mapped to the same TAE dimension, such as speaker features, through a linear layer. Text token embedding: Starting at position 2, it represents text tokens in the BPE (Byte Pair Encoding) vocabulary, such as text segmentation. To adapt to the sos and speaker embeddings in the data sequence, the BPE vocabulary is shifted two positions to the right, so the text token index starts at 2 and goes up to 51868. Task ID embedding: Located at position 51869, it is used to distinguish different tasks or datasets. Speech token embedding: Starting at position 51870, it represents speech tokens extracted from the audio data through the S3 Tokenizer, such as discrete features of speech. The speech token code table size is 4096, so the index range of speech tokens is from 51870 to 55966. Speech token eos embedding: located at position 55967, used to mark the end of the speech sequence.

[0027] Based on the above embodiments, the improved TAE+LLaMA model can not only receive and process text and audio data simultaneously, but also provide support for realizing speech synthesis based on text content.

[0028] refer to Figure 1 As shown, this is a flowchart of the pre-training method of the model provided in one embodiment of the present application. According to different requirements, the order of the steps in the flowchart can be changed, and some can be omitted.

[0029] S101, encode the second sample data using the model to obtain a second encoding vector.

[0030] In some embodiments, when constructing the second sample data, speech data from multiple speakers may be obtained, and the third speech data may be generated based on the speech data from the multiple speakers. The multiple speakers may have different characteristics such as gender, age, accent, and speaking speed to ensure data diversity. Thus, using the third speech data to train the model allows the model to learn the speech characteristics of different speakers, improving the model's generalization ability.

[0031] In some embodiments, each piece of voice data is preprocessed, such as by adjusting it to the same sampling rate, performing denoising, and length segmentation, to obtain third voice data. If the sampling rates corresponding to the third voice data are inconsistent, the sampling rate of the third voice data can be adjusted after obtaining the third voice data so that the sampling rates of all adjusted third voice data are consistent, thereby helping to ensure data consistency and compatibility. In one example, based on the data format of the third voice data (e.g., WAV, MP3, FLAC, etc.), the sampling rate can be adjusted using a sampling rate adjustment method for the corresponding format, such as, but not limited to, the FFmpeg command line tool, audio editing software, or the Python programming language's pydub or librosa libraries.

[0032] In some embodiments, the voice text corresponding to the third voice data may be obtained as the second voice text. The voice text corresponding to the third voice data may be determined by a voice recognition method and used as the second voice text.

[0033] In some embodiments, speaker characteristics of the speaker corresponding to the third voice data can be determined. In one example, the voice text corresponding to the third voice data can be obtained as the second voice text using a voice recognition algorithm. In another example, the speaker corresponding to the third voice data can be determined using an open-source user recognition model based on voice recognition, and an embedding vector corresponding to the third voice data can be obtained as the speaker characteristics of the speaker.

[0034] In some embodiments, second sample data can be constructed based on the third voice data, the second voice text, and speaker characteristics. Thus, each piece of second sample data includes the third voice data, the text data corresponding to the third voice data, and the speaker characteristics. During pre-training, the model can learn how to generate voice data that is consistent with the second text data and sounds like it was spoken by the speaker corresponding to the third voice data.

[0035] In some embodiments, the second sample data can be further processed to convert the data into a format understandable by the model, and the processed second sample data can be used as pre-training data for the model. In one example, the third speech data can be discretized to obtain second discrete features corresponding to the third speech data. For example, the second speech text can be segmented to obtain second text segmentations corresponding to the second speech text. The second sample data is constructed based on the speaker characteristics, the second text segmentations, and the second discrete features.

[0036] In one example, the third voice data can be discretized using a preset voice discretization method. For example, the third voice data is discretized using the vector quantization variational autoencoder VQ-VAE algorithm of S3Tokenizer, and the continuous voice signal in the third voice data is compressed into a series of discrete, representative feature vectors to obtain a second discrete feature. In one example, the second voice text can be segmented using a preset word segmentation tool. For example, the second voice text is segmented using the Whisper Tokenizer tool, which has the ability to process texts in multiple languages ​​and can decompose text data into smaller units, such as second text segmentation.

[0037] In one example, in order to separate different types of data, a separator can be set in the second sample data. For example, the second sample data includes a start symbol, a speaker feature, a second text segmentation, a first separator, a second discrete feature, and a terminator arranged in sequence. The first separator is used to separate the second text segmentation and the second discrete feature. Figure 2 As shown in the figure, the second sample data can be expressed as [sos, speaker_embedding, text token, task id, speechtoken]. In this figure, sos represents the start symbol, speaker_embedding represents the speaker feature, text token represents the second text segmentation, task id represents the first separator, and speech token represents the second discrete feature.

[0038] Based on the above embodiments, by converting speech data into discrete features, decomposing text data into word segments, and combining these elements with speaker features to construct sample data, the model can be allowed to utilize text information and speaker information when processing speech data, thereby improving the accuracy and generalization ability of the model's speech synthesis.

[0039] In some embodiments, when encoding the second sample data using the model, the third speech data, the second speech text, and the speaker features can be embedded into the same vector space based on the model's embedding layer, thereby obtaining an embedding vector corresponding to the second sample data as the second encoding vector. The model's embedding layer is an improved bimodal embedding layer (TAE), which can embed bimodal data consisting of text and audio data into the same vector space, thereby obtaining the first encoding vector.

[0040] refer to Figure 3As shown, it is an example diagram of the data distribution of the TAE layer provided by an embodiment of the present application. In one example, before the start of pre-training, the sos embedding is randomly initialized and recorded as e1; the task ID embedding is randomly initialized and recorded as e2. Among them, the dimensions of e1 and e2 are consistent with the TAE dimensions. The embedding at position 0 is replaced by e1, the embedding at position 1 is replaced by the speaker feature, and the embedding at position 51869 is replaced by e2. For text data, the index of the text segmentation is increased by 2 on this basis to adapt to the processing of TAE. For audio data, the index of the second discrete feature is increased by 51780 (i.e., 51870-4096, considering that the voice token starts at 51870 and the code table size is 4096) on this basis, also to adapt to the processing of TAE.

[0041] In this way, the second sample data is input into the TAE layer, and the TAE layer can learn how to extract useful features from each type of data in the second sample data and map these features into a common vector space. In this vector space, similar data points from speech, speech text, and speaker features are mapped to positions close to each other. This mapping relationship enables the model to understand and process bimodal data and achieve cross-modal information matching. In one example, if the second sample data includes speaker features, second text segmentation, and second discrete features arranged in sequence, the second encoding vector may include data features of speaker features arranged in sequence, data features corresponding to the second text segmentation, and data features corresponding to the second discrete features.

[0042] Based on the above embodiment, the second encoding vector not only contains the characteristic information of the three types of data, namely, speaker features, second text segmentation, and second discrete features, but also implies the association or mapping relationship between them, which can provide a basis for the subsequent generation of speech content that is consistent with the second speech text and sounds like the speech uttered by the speaker corresponding to the third speech data.

[0043] S102: Based on the second coding vector, use the model to predict the third speech data to obtain a second predicted speech.

[0044] In some embodiments, the model extracts data features corresponding to the second speech text and speaker features from the second encoded vector, and uses these features to predict discrete features of the speech, and generates predicted speech based on the predicted discrete features of the speech through speech synthesis technology.

[0045] In some embodiments, when determining the third data feature corresponding to the second speech text in the second coding vector, the data feature corresponding to the second speech text can be extracted from the second coding vector as the third data feature based on the position of the data feature corresponding to the second speech text in the second coding vector. The third data feature may include phonemes, syllables, intonation or other linguistic information related to speech generation in the second speech text. In some embodiments, when determining the fourth data feature corresponding to the speaker feature in the second coding vector, the data feature corresponding to the speaker feature can be extracted from the second coding vector as the fourth data feature based on the position of the data feature corresponding to the speaker feature in the second coding vector. The fourth data feature may include speaker information such as the speaker's age and accent. In some embodiments, a model is used to predict discrete features of speech based on the third and fourth data features. Discrete features of speech can be predicted based on the third and fourth data features. These features may be quantized representations of the speech signal, such as Mel-frequency cepstral coefficients, linear prediction coefficients, or other types of speech features. In one example, the model may use a forward propagation computational method to predict discrete features of speech based on the third and fourth data features.

[0046] In one example, reference Figure 2 As shown, the input data of the model, for example, the second sample data can be expressed as [sos, speaker_embedding, text token, task id, speech token], and the output data of the model can be expressed as [speaker_embedding, text token, task id, speech token, eos], where eos represents the terminator. The training goal of the model includes making the predicted speech discrete features as consistent as possible with the second discrete features.

[0047] In some embodiments, based on the discrete features of the speech, a speech synthesis technology is used to generate the second predicted speech. The speech synthesis technology can convert the discrete features into a continuous speech signal, thereby generating an audible predicted speech.

[0048] In one example, a Flow model can be used to generate a Mel spectrum based on discrete speech features. The Flow model can capture the complex distribution of speech signals, thereby generating high-quality Mel spectra. The generated Mel spectrum can then be input into a HiFi-GAN model, a generative adversarial network that generates speech waveforms based on the Mel spectrum, thereby generating realistic and natural first-pass predicted speech.

[0049] Based on the above embodiments, the feature prediction ability of the machine learning model, the signal conversion ability of the speech synthesis technology, and the high-quality speech generation ability of the Flow model and HiFi-GAN can be combined to realize text-to-speech conversion based on the second voice text and speaker characteristics. By continuously optimizing the model and algorithm, the naturalness and realism of the model speech synthesis can be improved to meet the needs of various application scenarios.

[0050] S103, determining a second loss value of the model based on the third speech data, the second predicted speech and a preset second loss function.

[0051] In some embodiments, a second loss function can be used to calculate the difference between the second predicted speech and the actual third speech data to obtain a second loss value. For example, the second loss function can include, but is not limited to, a cross-entropy function. A larger second loss value indicates a worse performance of the model in predicting the third speech data.

[0052] S104: Pre-train the model according to the second loss value.

[0053] In some embodiments, if the second loss value is greater than a preset second threshold, a back-propagation algorithm can be used to optimize the model and update the model parameters of the model, such as the connection weights between neurons, so that the second loss value of the model gradually decreases until the second loss value of the model is less than or equal to the second threshold.

[0054] The model pre-training method provided in the embodiment of the present application pre-trains the model by introducing second sample data containing speech data, corresponding speech text and speaker characteristics, so that the model can learn and generate speech that is consistent with the text content and the speaker characteristics, thereby improving the accuracy and flexibility of the model in generating speech.

[0055] The above embodiment introduces the pre-training method of the model. Next, a method for further training and fine-tuning the model based on the pre-trained model will be introduced.

[0056] refer to Figure 4 As shown, it is a flowchart of the data processing method provided by an embodiment of the present application. According to different requirements, the order of the steps in the flowchart can be changed, and some can be omitted.

[0057] S401: Encode first sample data using a model to obtain a first encoding vector.

[0058] In some embodiments, the model can be further trained using the first sample data based on the pre-trained model. The first sample data includes the first voice data, the first voice text corresponding to the first voice data, and the instruction text, and the instruction text is used to indicate the speaking speed of the first voice data. In this way, the model can be made to learn to generate a voice with a speaking speed corresponding to the instruction text based on the instruction text in the first sample data. In one example, the method for constructing the first sample data can refer to Figure 5 The flowchart shown.

[0059] In some embodiments, when encoding the first sample data using the model, the model's embedding layer can be used to embed the first voice data, the first voice text, and the instruction text into the same vector space, thereby obtaining an embedding vector corresponding to the first sample data as the first encoding vector. The model's embedding layer is an improved bimodal embedding layer (Text-Audio Embedding, TAE), which can embed bimodal data consisting of text and audio data into the same vector space, thereby obtaining the first encoding vector.

[0060] In one example, the TAE layer can learn how to extract useful features from the three types of data in the first sample data and map these features into a common vector space. In this vector space, similar data points from speech, speech text, and instruction text are mapped to positions close to each other. This mapping relationship enables the model to understand and process bimodal data and achieve cross-modal information matching. In one example, if the first sample data includes instruction segmentation words, a first separator, a first text segmentation word, a second separator, and a first discrete feature arranged in sequence, the first encoding vector may include data features of the instruction segmentation words arranged in sequence, data features corresponding to the first text segmentation word, and data features corresponding to the first discrete feature.

[0061] Based on the above embodiment, the first encoding vector not only contains the characteristic information of the three types of data, namely, the first voice data, the first voice text and the instruction text, but also implies the association or mapping relationship between them, which can provide a basis for the subsequent generation of voice content consistent with the first voice text and the speaking speed corresponding to the instruction text.

[0062] S402: Based on the first coding vector, use a model to predict the first speech data to obtain a first predicted speech.

[0063] In some embodiments, the model extracts data features corresponding to the first speech text and the instruction text from the first encoding vector, and uses these features to predict discrete features of the speech, and generates predicted speech based on the predicted discrete features of the speech using speech synthesis technology. In one example, the method for predicting the first speech data can refer to Figure 7The embodiment shown.

[0064] S403: Determine a first loss value of the model based on the first speech data, the first predicted speech, and a preset first loss function.

[0065] In some embodiments, a first loss function may be used to calculate the difference between the first predicted speech and the actual first speech data to obtain a first loss value. For example, the first loss function may include, but is not limited to, a cross-entropy function. A larger first loss value indicates a worse performance of the model in predicting the first speech data.

[0066] S404: Train the model according to the first loss value.

[0067] In some embodiments, if the first loss value is greater than a preset first threshold, a back-propagation algorithm can be used to optimize the model and update the model parameters of the model, such as the connection weights between neurons, so that the first loss value of the model gradually decreases until the first loss value of the model is less than or equal to the first threshold.

[0068] The data processing method provided in the embodiment of the present application further trains the pre-trained model by introducing first sample data including voice data, corresponding voice text and speaking speed instruction text, so that the model can learn and generate voice that matches the speaking speed specified by the instruction text, thereby improving the flexibility of the model in generating voice of corresponding speaking speed.

[0069] In some embodiments, as Figure 5 As shown, the method for constructing the first sample data may include the following steps. According to different requirements, the order of the steps in the flowchart can be changed, and some steps can be omitted.

[0070] S501: Perform speech rate conversion on second voice data to obtain first voice data.

[0071] In some embodiments, a smaller amount of first sample data can be used to fine-tune the model, wherein speech data of a single speaker can be obtained as the second speech data. For example, 10 hours of speech data of a single speaker can be obtained, which may include multiple speech data items, and each speech data item can be used as the second speech data item.

[0072] In some embodiments, the second voice data may be voice data obtained using the same sampling rate, for example, the sampling rate is 16kHz. If the sampling rates corresponding to the second voice data are inconsistent, the sampling rate of the second voice data may be adjusted after obtaining the second voice data so that the sampling rates of all adjusted second voice data are consistent, thereby helping to ensure data consistency and compatibility. In one example, based on the data format of the second voice data (such as WAV, MP3, FLAC, etc.), the sampling rate adjustment method of the corresponding format may be used to adjust the sampling rate, such as but not limited to the FFmpeg command line tool, audio editing software, pydub or librosa library of the Python programming language, and the like.

[0073] In some embodiments, the second voice data can be subjected to speech speed conversion to obtain corresponding speech data of different speech speeds as the first voice data. In one example, the second voice data can be adjusted to 0.8, 0.9, 1.0, 1.1, and 1.2 times the original speech speed to obtain 5 speech data of different speech speeds, and each voice data can be used as a first voice data. Among them, FFmpeg commands can be used to generate voice files of different speech speeds. In one example, a speech speed label can be set for each first voice data. For example, if the speech speed of the first voice data is 0.8 times the speech speed of the second voice data, a speech speed label "_speed_0.8" can be added. In one example, a unique identifier ID can also be set for each first voice data, and a storage path wav_path corresponding to each first voice data can be determined to indicate the storage location of the first voice data.

[0074] S502: Determine the first voice text corresponding to the first voice data based on the voice text corresponding to the second voice data.

[0075] In some embodiments, a speech text corresponding to the second speech data may be determined by a speech recognition method, and the speech text may be used as the speech text of each first speech data corresponding to the second speech data, hereinafter referred to as the first speech text.

[0076] S503: Construct first sample data according to the first voice data, the first voice text and the instruction text.

[0077] In some embodiments, instruction generation technology can be used to generate corresponding diversified instruction texts based on the speech rate label of the first voice data, such as multiple texts with the same meaning but different expressions. Among them, the instruction generation technology can be implemented based on an open source large language model, which has the ability to understand and generate natural language text. A task prompt can be input into the large model, which contains key information about the required instructions, such as the speech rate label "0.8 times the speech rate". The large language model generates multiple sentences with the same meaning but different expressions based on the input task prompt. These texts are all intended to convey a core instruction: adjust the speech rate to a specified ratio.

[0078] In one example, if the speech speed label is "0.8 times the speech speed", the task prompt can be "Generate multiple sentences with the same meaning but different expressions to guide the adjustment of the speech speed to 80% or 0.8 times the original speed. The instruction text should be clear, concise, and easy to understand. The instruction text can include a detailed description of the speech speed adjustment, a direct command, or a suggestive expression." In one example, the instruction text generated based on the above task prompt can be expressed as follows: "(1) Set the speech speed to 80%. (2) Reduce the speech speed to 80% of the original speed. (3) Adjust the speech speed to 80%. (4) Slow down the speech speed to 0.8 times the original speed. (5) Adjust the speech speed to 80% of the original speed. (6) Reduce the speech speed to 80% of the original speed. (7) Reduce the speech speed to 80% of the original speed. (8) Adjust the speech speed to 0.8 times the original speed. (9) Reduce the speech speed to 80% of the original speed. (10) Slow down the speech speed to 80% of the original speed."

[0079] Based on the above embodiment, instruction texts can be used to imitate various instruction statements that users may issue in actual applications to adjust the speaking speed. These instruction texts can not only cover direct commands and suggestive expressions that users may use, but also contain diverse expressions in different contexts and styles, increasing the richness and diversity of the data. In this way, training the model based on instruction texts can enable the model to learn multiple types of instruction texts, thereby enhancing its ability to understand and generate speech at corresponding speeds and improving the robustness of the model. Specifically, the model can analyze key information in these instruction texts, such as speech speed labels, adjustment directions (speeding up or slowing down), etc., and adjust the speaking speed of the speech it generates accordingly. Such a training strategy helps to improve the flexibility and adaptability of the model in actual applications, enabling it to more accurately respond to users' speech speed adjustment needs.

[0080] In some embodiments, first sample data can be constructed based on the first voice data, the first voice text, and the instruction text. In addition, the first sample data can also include an ID corresponding to the first voice data. For example, the first sample data can include the following CSV file content format: id, instruction_text, text, wav_path. Wherein instruction_text is the instruction text, text represents the first voice text, wav_path represents the storage path corresponding to the first voice data, and the speaking speed of the first voice data is consistent with the speaking speed required by the instruction text.

[0081] In one example, the first sample data may include: "1_1 (ID)", "wav_path (the first voice data with a speaking speed of 0.8 times)", "Excuse me, are you Ms. Li (first voice text)", and "Set the speaking speed to 80% (instruction text)". In another example, the first sample data may include: "1_2 (ID)", "wav_path (the first voice data with a speaking speed of 0.8 times)", "Excuse me, are you Ms. Li (first voice text)", and "Speaking speed reduced to 80% of the original speed (instruction text)". In yet another example, the first sample data may include: "1_3 (ID)", "wav_path (the first voice data with a speaking speed of 1.2 times)", "Excuse me, are you Ms. Li (first voice text)", and "Set the speaking speed to 1.2 times (instruction text)".

[0082] Based on the above embodiment, each piece of first sample data includes a piece of first speech data, text data corresponding to the first speech data, and instruction text data that explicitly indicates the speech rate corresponding to the first speech data. Training the model based on the first sample data enables the model to comprehensively learn the relationship between speech rate, speech content, and instructions, thereby more effectively understanding and generating speech that meets the specified speech rate requirements, greatly improving the model's training effectiveness and practical application capabilities.

[0083] In some embodiments, the first sample data may be further processed to convert the data into a format understandable by the model, and the processed first sample data may be used as model training data. In some embodiments, the first speech data may be discretized to obtain first discrete features corresponding to the first speech data; the first speech text may be segmented to obtain first text segmentations corresponding to the first speech text; the instruction text may be segmented to obtain instruction segmentations corresponding to the instruction text; and the first sample data may be constructed based on the instruction segmentations, the first text segmentations, and the first discrete features.

[0084] In one example, a preset speech discretization method can be used to discretize the first speech data. For example, the first speech data is discretized using the vector quantization variational autoencoder VQ-VAE algorithm of S3Tokenizer, and the continuous speech signal in the first speech data is compressed into a series of discrete, representative feature vectors to obtain a first discrete feature. In one example, a preset word segmentation tool can be used to segment the first speech text and the instruction text. For example, the Whisper Tokenizer tool is used to segment the text (such as the first speech text and the instruction text). The tool has the ability to process texts in multiple languages ​​and can decompose text data into smaller units, such as word segments.

[0085] In one example, in order to separate different types of data, a separator can also be set in the first sample data. For example, the first sample data includes an instruction segmentation word, a second separator, a first text segmentation word, a third separator, and a first discrete feature arranged in sequence. The second separator is used to separate the instruction segmentation word from the first text segmentation word, and the third separator is used to separate the first text segmentation word from the first discrete feature. Figure 6 As shown, the first sample data can be expressed as [sos, instructiontext token, <endofprompt>, text token, task id, speech token]. Among them, sos represents the start symbol, instruction text token represents the instruction segmentation, <endofprompt>Represents the second separator, text token represents the first text segmentation, ask id represents the third separator, and speech token represents the first discrete feature.

[0086] Based on the above embodiments, by converting speech data into discrete features, decomposing text data into word segments, and combining these elements to construct sample data, the model can be allowed to utilize text information when processing speech data, thereby improving the accuracy and generalization ability of the model's speech synthesis.

[0087] In some embodiments, as Figure 7 As shown, the method for predicting the use of the first voice data may include the following steps. According to different requirements, the order of the steps in the flowchart may be changed, and some steps may be omitted.

[0088] S701, determining a first data feature corresponding to a first speech text in a first encoding vector.

[0089] In some embodiments, the data feature corresponding to the first speech text can be extracted from the first encoding vector as the first data feature based on the position of the data feature corresponding to the first speech text in the first encoding vector. The first data feature may include phonemes, syllables, intonation or other linguistic information related to speech generation in the first speech text. S702: Determine a second data feature corresponding to the instruction text in the first encoding vector.

[0090] In some embodiments, the data feature corresponding to the instruction text can be extracted from the first encoding vector as the second data feature based on the position of the data feature corresponding to the instruction text in the first encoding vector. The second data feature may include instruction information for controlling the speech generation method, such as speaking speed.

[0091] S703: Based on the first data feature and the second data feature, obtain speech discrete features by using a model prediction.

[0092] In some embodiments, the model can predict discrete features of speech based on the first data feature and the second data feature. These features may be quantized representations of the speech signal, such as Mel-frequency cepstral coefficients, linear prediction coefficients, or other types of speech features. In one example, the model can use a forward propagation computational method to predict discrete features of speech based on the first data feature and the second data feature. The training objective of the model includes making the predicted discrete features of speech as consistent as possible with the first discrete features.

[0093] S704: synthesize the discrete features of the speech using speech synthesis technology and output synthesized speech data.

[0094] In some embodiments, synthesis parameters are set based on the discrete features of the speech and the selected speech synthesis technology, for example, parameters such as the fundamental frequency, duration, and timbre of the speech. Based on the discrete features of the speech and the synthesis parameters, the speech synthesis technology is used to generate a first predicted speech. The speech synthesis technology can convert the discrete features into a continuous speech signal based on the synthesis parameters, thereby generating an audible predicted speech.

[0095] In one example, a Flow model can be used to generate a Mel spectrum based on discrete speech features. The Flow model can capture the complex distribution of speech signals, thereby generating high-quality Mel spectra. The generated Mel spectrum can then be input into a HiFi-GAN model, a generative adversarial network that generates speech waveforms based on the Mel spectrum, thereby generating realistic and natural first-pass predicted speech.

[0096] Based on the above embodiments, the feature prediction ability of the machine learning model, the signal conversion ability of the speech synthesis technology, and the high-quality speech generation ability of the Flow model and HiFi-GAN can be combined to realize text-to-speech conversion based on the characteristics of the first voice text and the instruction text. By continuously optimizing the model and algorithm, the naturalness and realism of the model speech synthesis can be improved to meet the needs of various application scenarios.

[0097] The above examples describe how to train a model. The trained model can be used as a speech synthesis model to perform speech synthesis tasks. Figure 8 As shown, it is a flowchart of the speech synthesis method provided by an embodiment of the present application. According to different requirements, the order of the steps in the flowchart can be changed, and some can be omitted.

[0098] S801: Recognize the user's input voice and obtain recognized text.

[0099] In some embodiments, during the interaction between the intelligent voice interaction system and the user, the system recognizes the user's input speech to obtain recognized text, and then determines the user's intent based on the recognized text. The system can pre-store the response text corresponding to each user intent as the output text, convert the output text into speech and output it to the user, thereby meeting the user's needs.

[0100] S802: If the recognized text contains a preset keyword, determine the instruction text corresponding to the preset keyword.

[0101] In some embodiments, a user may be dissatisfied with the speech speed of the system and may issue a command to adjust the speech speed of the intelligent voice interaction system. To avoid missing such user-issued commands, the text can be identified to determine whether it contains preset keywords. For example, preset keywords may include, but are not limited to, one or more keywords or combinations of keywords such as "speech speed," "speed," "too fast," "speak slower," "inaudible," and "one point and one time."

[0102] In some embodiments, when the text input by the user is recognized as containing preset keywords, it can be determined that the user needs to adjust the system voice speed. To meet the user's needs, a fuzzy matching algorithm can be used to match the recognized preset keywords with a pre-stored instruction text library. The instruction text library contains a variety of texts with different expressions but the same meaning, each of which corresponds to a specific speed setting.

[0103] In one example, referring to the relevant embodiment in step S503, the instruction text library is pre-constructed and contains a variety of different expressions that users may use to express their demand for speech speed adjustment. For example, for the speech speed setting of "speak slower", the instruction text library may contain a variety of expressions such as "slow down the speech speed", "slow down the speech speed", and "reduce the speech speed to". In this way, after recognizing that the text input by the user contains preset keywords, the fuzzy matching algorithm is used to search the instruction text library for the expression that best matches the user input text. Since the fuzzy matching algorithm can handle the similarities and differences between texts, even if the user's input is not completely consistent with a specific expression in the instruction text library, the system can find the closest match as the instruction text corresponding to the preset keyword.

[0104] Based on the above embodiment, the user's speech speed adjustment requirements can be accurately understood based on the fuzzy matching algorithm and the pre-built instruction text library.

[0105] S803: Generate speech discrete features using a speech generation model based on the instruction text and the text to be output.

[0106] In some embodiments, the relevant steps in the model application reasoning process are similar to the steps in the data processing method. The text to be output can be used to replace the first speech text used in the forward reasoning of the model training process, so that discrete speech features can be generated by the speech generation model based on the instruction text and the text to be output.

[0107] In some embodiments, based on the instruction text and the text to be output, discrete speech features are generated using a speech generation model, including: using the embedding layer of the speech generation model to embed the text to be output and the instruction text into the same vector space to obtain an encoding vector, for example, refer to the description of the relevant embodiment of step S401. Determining the first data feature corresponding to the text to be output and the second data feature corresponding to the instruction text in the encoding vector, for example, refer to the description of the relevant embodiment of steps S701-S702. Based on the first data feature and the second data feature, the speech generation model is used to predict and obtain discrete speech features, for example, refer to the description of the relevant embodiment of step S703.

[0108] Based on the above embodiment, the text content corresponding to the discrete speech features generated by the trained speech generation model is consistent with the content of the text to be output, and the corresponding speaking speed is consistent with the corresponding speaking speed in the instruction text.

[0109] S804: synthesize the discrete features of the speech using speech synthesis technology and output synthesized speech data.

[0110] In some embodiments, the specific implementation of step S804 can refer to step S704 and will not be described again.

[0111] The speech synthesis method provided in the embodiments of the present application can directly generate speech with the desired content and speed based on a speech synthesis model. The speech synthesis model is derived based on the aforementioned data processing method and can therefore be used to flexibly adjust the output speech speed in real time, accurately responding to user demands for speed adjustment and thus improving the user experience.

[0112] like Figure 9 , is a functional block diagram of a data processing device provided by an embodiment of the present application. The data processing device 8 includes an encoding unit 91, a prediction unit 92, a determination unit 93, and a training unit 94. The module / unit referred to in this application refers to a type of data processing unit that can be processed by a processor (e.g. Figure 10 The processor 1101 shown in FIG. 1 is obtained and is capable of performing a series of computer-readable instruction segments that are stored in a memory (eg, Figure 10 1102).

[0113] The encoding unit 91 is used to encode the first sample data using the model to obtain a first encoding vector, wherein the first sample data includes first speech data; the prediction unit 92 is used to predict the first speech data using the model based on the first encoding vector to obtain a first predicted speech; the determination unit 93 is used to determine the first loss value of the model based on the first speech data, the first predicted speech and a preset first loss function; the training unit 94 is used to train the model according to the first loss value.

[0114] Figure 10 Schematic diagram of the structure of a computer device for implementing a data processing method provided in an embodiment of the present application. Figure 10 The computer device 10 shown is used to execute the methods in the above-mentioned method embodiments.

[0115] The computer device 10 includes at least one processor 1101 , a memory 1102 , and at least one network interface 1103 .

[0116] The processor 1101 is, for example, a general-purpose central processing unit (CPU), a network processor (NP), a graphics processing unit (GPU), a neural-network processing unit (NPU), a data processing unit (DPU), a microprocessor, or one or more integrated circuits for implementing the solution of the present application. For example, the processor 1101 includes an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The PLD is, for example, a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0117] The memory 1102 may be, for example, a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, a random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, an optical disc storage (including a compact disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but is not limited thereto. Optionally, the memory 1102 exists independently and is connected to the processor 1101 via the internal connection 1104. Alternatively, the memory 1102 and the processor 1101 may be integrated together.

[0118] The network interface 1103 uses any transceiver-like device for communicating with other devices or communication networks. For example, the network interface 1103 includes at least one of a wired network interface and a wireless network interface. For example, the wired network interface is an Ethernet interface. For example, the Ethernet interface is an optical interface, an electrical interface, or a combination thereof. For example, the wireless network interface is a wireless local area network (WLAN) interface, a cellular network interface, or a combination thereof.

[0119] In some embodiments, the processor 1101 includes one or more CPUs, such as Figure 10 CPU0 and CPU1 are shown in the figure.

[0120] In some embodiments, the computer device 10 optionally includes multiple processors, such as Figure 10 1 and 1105. Each of these processors is, for example, a single-CPU or a multi-CPU. A processor herein optionally refers to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions). In some embodiments, computer device 10 further includes internal connections 1104. Processor 1101, memory 1102, and at least one network interface 1103 are connected via internal connections 1104. Internal connections 1104 include pathways that transmit information between these components. Optionally, internal connections 1104 are boards or buses. Optionally, internal connections 1104 are divided into address buses, data buses, control buses, and the like.

[0121] In some embodiments, the computer device 10 further includes an input / output interface 1106 . The input / output interface 1106 is connected to the internal connection 1104 .

[0122] Optionally, the processor 1101 implements the method in the above embodiment by reading the program code 910 stored in the memory 1102, or the processor 1101 implements the method in the above embodiment by internally stored program code. In the case where the processor 1101 implements the method in the above embodiment by reading the program code 910 stored in the memory 1102, the memory 1102 stores the program code that implements the method provided in the embodiment of the present application.

[0123] For more details on how the processor 1101 implements the above functions, please refer to the descriptions in the previous method embodiments, which will not be repeated here.

[0124] This embodiment further provides a computer storage medium, in which computer instructions are stored. When the computer instructions are executed on a computer device, the computer device executes the above-mentioned related method steps to implement the method in the above-mentioned embodiment.

[0125] This embodiment further provides a computer program product, including a computer program / instruction. When the computer program product is run on a computer device, the computer device is caused to execute the above-mentioned related steps to implement the methods in the above-mentioned method embodiments.

[0126] This embodiment also provides a chip system, which is applied to a computer device. The chip system includes one or more processors, which are used to call computer instructions to enable the computer device to input a first training data set into the chip system and execute the methods in the above-mentioned method embodiments.

[0127] In addition, an embodiment of the present application also provides a device, which can specifically be a chip, component or module, and the device may include a connected processor and memory; wherein the memory is used to store computer-executable instructions, and when the device is running, the processor can execute the computer-executable instructions stored in the memory to enable the chip to execute the methods in the above-mentioned method embodiments.

[0128] Among them, the computer device, computer storage medium, computer program product or chip provided in this embodiment is used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method provided above, and will not be repeated here.

[0129] Through the description of the above implementation methods, technical personnel in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0130] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0131] The units described as separate components may or may not be physically separate, and the components shown as units may be one physical unit or multiple physical units, that is, they may be located in one place or distributed in multiple places. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0132] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0133] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for enabling a device (which can be a single-chip microcomputer, chip, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program code.

[0134] The above is only a specific embodiment of the present application, but the scope of protection of this application is not limited to this. Any changes or substitutions within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.< / endofprompt> < / endofprompt>

Claims

1. A data processing method, characterized in that: The method comprises: Encoding first sample data using the model to obtain a first encoding vector, wherein the first sample data includes first speech data; Based on the first encoding vector, using the model to predict the first speech data to obtain a first predicted speech; Determining a first loss value of the model based on the first speech data, the first predicted speech, and a preset first loss function; The model is trained according to the first loss value.

2. The data processing method according to claim 1, wherein: The method further includes constructing the first sample data, including: performing speech speed conversion on the second voice data to obtain the first voice data; Determining a first voice text corresponding to the first voice data based on the voice text corresponding to the second voice data; The first sample data is constructed according to the first voice data, the first voice text, and an instruction text, wherein the instruction text is used to indicate a speaking speed of the first voice data.

3. The data processing method according to claim 2, wherein: The constructing the first sample data according to the first voice data, the first voice text, and the instruction text includes: performing discretization processing on the first speech data to obtain first discrete features corresponding to the first speech data; Performing word segmentation processing on the first speech text to obtain first text segmentation corresponding to the first speech text; Performing word segmentation processing on the instruction text to obtain instruction word segmentations corresponding to the instruction text; The first sample data is constructed according to the instruction segmentation, the first text segmentation and the first discrete feature.

4. The data processing method according to claim 1, wherein: The first sample data also includes a first voice text and an instruction text corresponding to the first voice data. The encoding of the first sample data using the model to obtain a first encoding vector includes: The first speech data, the first speech text, and the instruction text are embedded into the same vector space by using the embedding layer of the model, and an embedding vector corresponding to the first sample data is obtained as the first encoding vector.

5. The data processing method according to claim 1, wherein: The first sample data also includes a first voice text and an instruction text corresponding to the first voice data. The method of predicting the first voice data using the model based on the first encoding vector to obtain a first predicted voice includes: Determining a first data feature corresponding to the first speech text in the first encoding vector; determining a second data feature corresponding to the instruction text in the first encoding vector; Based on the first data feature and the second data feature, using the model to predict and obtain discrete speech features; The speech discrete features are synthesized using speech synthesis technology to output synthesized speech data.

6. The data processing method according to claim 1, wherein: Before encoding the first sample data using the model, the method further includes: Encoding second sample data using the model to obtain a second encoding vector, wherein the second sample data includes third speech data, a second speech text corresponding to the third speech data, and speaker features; Based on the second encoding vector, using the model to predict the third speech data to obtain a second predicted speech; Determining a second loss value of the model based on the third speech data, the second predicted speech, and a preset second loss function; The model is pre-trained according to the second loss value.

7. The data processing method according to claim 6, wherein: The method further includes constructing the second sample data, including: Acquiring voice data of a plurality of speakers, and obtaining third voice data based on the voice data of the plurality of speakers; Acquire a speech text corresponding to the third speech data as the second speech text, and determine a speaker feature of a speaker corresponding to the third speech data; The second sample data is constructed according to the third voice data, the second voice text and the speaker characteristics.

8. The data processing method according to claim 7, wherein: The constructing the second sample data according to the third voice data, the second voice text and the speaker feature includes: performing discretization processing on the third voice data to obtain a second discrete feature corresponding to the third voice data; Performing word segmentation processing on the second voice text to obtain second text segmentation corresponding to the second voice text; The second sample data is constructed according to the speaker feature, the second text segmentation and the second discrete feature.

9. A speech synthesis method, characterized in that: The method comprises: Recognize the user's input voice and obtain the recognized text; If the recognized text contains a preset keyword, determining an instruction text corresponding to the preset keyword, wherein the instruction text is used to indicate the user's desired speaking speed; Based on the instruction text and the text to be output, generating speech discrete features using a speech generation model; The speech discrete features are synthesized using speech synthesis technology to output synthesized speech data.

10. The speech synthesis method according to claim 9, characterized in that: The generating of speech discrete features by using a speech generation model based on the instruction text and the text to be output includes: Using the embedding layer of the speech generation model, embed the to-be-output text and the instruction text into the same vector space to obtain an encoding vector; Determining a first data feature corresponding to the to-be-output text and a second data feature corresponding to the instruction text in the encoding vector; Based on the first data feature and the second data feature, the speech discrete feature is predicted using the speech generation model.

11. A computer device, characterized in that: The computer device comprises a memory and a processor: Wherein, the memory is used to store program instructions; The processor is used to read and execute the program instructions stored in the memory. When the program instructions are executed by the processor, the computer device implements the data processing method as described in any one of claims 1 to 8, or implements the speech synthesis method as described in any one of claims 9 to 10.