Voice generation model training method and device, equipment and medium
By introducing conditional distribution simulation of semantic features and acoustic feature modules into the speech generation model, combined with iterative training and loss function adjustment, the problem of low training efficiency of non-autoregressive models is solved, and efficient and accurate speech generation is achieved.
Patent Information
- Application Number
- CN202510499283.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-08-08
AI Technical Summary
The speech generation system using non-autoregressive models in the prior art has low efficiency and accuracy during training, making it difficult to meet the needs of efficient speech synthesis in financial and medical scenarios.
The semantic feature module and acoustic feature module in the preset training model are used to recognize semantic feature and acoustic feature through the conditional flow matching model, and speech generation is performed in combination with the acoustic decoding module, and the initial parameters are iteratively adjusted until the predicted loss value reaches the convergence condition.
It improves the efficiency and accuracy of speech generation, improves the performance of the model, and ensures high-quality output of the speech generation model.
Smart Images

Figure CN120452416A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech synthesis technology, and in particular to a speech generation model training method, device, equipment and medium. Background Art
[0002] Speech generation technology, which converts text into speech, has broad applications in a variety of fields, including the internet, finance, healthcare, and education. In financial and healthcare settings, human interaction is often required to answer user questions. However, due to the complexity and diversity of financial and healthcare operations, the high volume of simple tasks, such as consulting, can reduce staff efficiency and quality. Intelligent conversational methods based on speech synthesis can significantly reduce labor costs and improve customer service quality by controlling the synthesized speech. Therefore, speech synthesis technology plays an important supporting role in these two scenarios. Existing systems using non-autoregressive models (such as diffusion models) often require hundreds or even thousands of denoising steps during training to achieve high-quality speech generation, significantly limiting efficiency and accuracy. Summary of the Invention
[0003] Embodiments of the present invention provide a speech generation model training method, apparatus, device, and medium to solve the problem of low efficiency and accuracy in training systems using non-autoregressive models in the prior art.
[0004] A speech generation model training method, comprising: Acquire a sample data set, the sample data set including at least one sample text and its corresponding sample speech; Obtaining a preset training model, and performing semantic feature recognition on the sample text and the sample speech using a semantic feature module in the preset training model to obtain sample semantic features corresponding to each of the sample texts; Performing acoustic feature recognition on the sample speech and the sample semantic features through the acoustic feature module in the preset training model to obtain sample acoustic features corresponding to each of the sample texts; the semantic feature module and the acoustic feature module adopt a conditional flow matching model; Performing speech generation on the sample acoustic features through the acoustic decoding module in the preset training model to obtain predicted generated speech corresponding to each sample text; Determining a prediction loss value of the preset training model based on the predicted generated speech and the sample speech corresponding to the same sample text; When the predicted loss value does not reach the preset convergence condition, the initial parameters in the preset training model are iteratively updated until the predicted loss value reaches the convergence condition, and the preset training model after convergence is recorded as the speech generation model.
[0005] A speech generation model training device, comprising: A sample data acquisition module, configured to acquire a sample data set, wherein the sample data set includes at least one sample text and its corresponding sample speech; A semantic feature recognition module is used to obtain a preset training model, and perform semantic feature recognition on the sample text and the sample speech through the semantic feature module in the preset training model to obtain sample semantic features corresponding to each of the sample texts; An acoustic feature recognition module, configured to perform acoustic feature recognition on the sample speech and the sample semantic features using the acoustic feature module in the preset training model to obtain sample acoustic features corresponding to each of the sample texts; the semantic feature module and the acoustic feature module adopt a conditional flow matching model; A predicted speech generation module, configured to generate speech based on the acoustic features of the samples using the acoustic decoding module in the preset training model, to obtain predicted generated speech corresponding to each of the sample texts; A model loss prediction module, configured to determine a prediction loss value of the preset training model based on the predicted generated speech and the sample speech corresponding to the same sample text; The model convergence module is used to iteratively update the initial parameters in the preset training model when the predicted loss value does not reach the preset convergence condition, until the predicted loss value reaches the convergence condition, and record the preset training model after convergence as the speech generation model.
[0006] A speech generation method, comprising: Obtaining a text to be processed, and performing text encoding on the text to be processed to obtain target text features; Calling a speech generation model, wherein the speech generation model is trained using the speech generation model training method described above; The target text features are subjected to speech generation by the speech generation model to obtain target generated speech.
[0007] A speech generating device, comprising: A target text feature module is used to obtain the text to be processed and perform text encoding on the text to be processed to obtain target text features; A speech generation model module is used to call a speech generation model, wherein the speech generation model is trained by the speech generation model training method as described above; The target speech generation module is used to generate speech based on the target text features through the speech generation model to obtain target generated speech.
[0008] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the processor implements the above-mentioned speech generation model training method or the above-mentioned speech generation method.
[0009] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the above-mentioned speech generation model training method or the above-mentioned speech generation method.
[0010] The present invention provides a speech generation model training method, apparatus, equipment and medium. The method realizes the simulation of the conditional distribution of semantic features and acoustic features by using a semantic feature module and an acoustic feature module in a preset training model, thereby improving the quality of the generated speech, improving the efficiency of speech generation during training, and further improving the accuracy of the model and improving the performance of the model. Furthermore, the preset training model is iteratively trained using a large amount of sample training data, and the overall loss value of the preset training model is calculated by comparing the loss function, thereby realizing the determination of the predicted loss value of the preset training model. The initial parameters of the preset training model are adjusted according to the predicted loss value until the model converges, thereby realizing the training of the speech generation model and ensuring that the speech generation model has a high accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0012] Figure 1 2. It is a schematic diagram of an application environment of a speech generation model training method according to an embodiment of the present invention; Figure 2 is a flow chart of a method for training a speech generation model in one embodiment of the present invention; Figure 3 is a flow chart of a speech generation method according to an embodiment of the present invention; Figure 4 This is a functional block diagram of a speech generation model training device according to one embodiment of the present invention; Figure 5 is a principle block diagram of a speech generating device according to one embodiment of the present invention; Figure 6is a schematic diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0013] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0014] The speech generation model training method provided by the embodiment of the present invention can be applied as follows: Figure 1 Specifically, the speech generation model training method is applied in a speech generation model training device, which includes Figure 1 The client and server shown communicate with each other through a network to solve the problem in the prior art that local stylization of images cannot achieve the expected effect. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The client, also known as the user end, refers to a program that corresponds to the server and provides classification services to customers. The client can be installed on, but is not limited to, various computers, laptops, smartphones, tablets, and portable wearable devices.
[0015] In one embodiment, if Figure 2 As shown, a speech generation model training method is provided, which is applied in Figure 1 The server in the example is used as an example, and the steps are as follows: S10: Acquire a sample data set, where the sample data set includes at least one sample text and its corresponding sample speech.
[0016] Understandably, sample text refers to historical text data, such as medical route data, medical records, or answers to various consultation questions, such as responses to credit card application procedures. This sample text can also be obtained through speech conversion. Each sample text corresponds to a sample speech, and the sample speech refers to the audio data corresponding to the sample text. Sample training data and sample labels can be collected from different databases or pre-prepared and sent from the client to the database. A sample dataset is then constructed using all the sample text and all the sample speech.
[0017] S20: Obtain a preset training model, and perform semantic feature recognition on the sample text and the sample speech through the semantic feature module in the preset training model to obtain sample semantic features corresponding to each sample text.
[0018] Understandably, a preset training model refers to a pre-set model whose structure is derived through inference. Sample semantic features refer to the semantic features extracted from each sample text and sample speech. Semantic features refer to characteristic attributes that can represent the semantic information contained in a text or language segment.
[0019] Specifically, a preset training model is obtained and all sample text and speech samples are input into the preset training model. Semantic feature recognition is then performed on the sample text and speech samples using the semantic feature module within the preset training model. Specifically, the sample text is first encoded using a text encoder to obtain text features. Semantic features are then extracted from the sample speech using a semantic encoder to obtain masked semantic features and complete semantic features. The semantic extraction unit within the semantic feature module is then trained using the text features, masked semantic features, and complete semantic features, enabling the semantic extraction unit to extract the semantic features of the sample text from the text features. By performing semantic extraction on all sample texts, sample semantic features corresponding to each sample text are obtained. For example, in a financial scenario, to inquire about the credit card application process, the corresponding answer text is first found based on the user's question. This is then encoded using a text encoder, and then feature extracted using the semantic feature module. This yields the semantic features corresponding to the answer text. For example, in a medical scenario, to inquire about department routes, the corresponding answer text is first found based on the user's question. This is then encoded using a text encoder, and then feature extracted using the semantic feature module. This yields the semantic features corresponding to the answer text.
[0020] S30, performing acoustic feature recognition on the sample speech and the sample semantic features through the acoustic feature module in the preset training model to obtain sample acoustic features corresponding to each of the sample texts; the semantic feature module and the acoustic feature module adopt a conditional flow matching model.
[0021] It is understandable that the semantic feature module and the acoustic feature module adopt a conditional flow matching model. The semantic feature module and the acoustic feature module adopt a conditional flow matching model.
[0022] Specifically, acoustic feature recognition is performed on the sample speech and sample semantic features through the acoustic feature module in the preset training model, that is, the acoustic feature extraction is performed on the sample speech through the acoustic encoder in the acoustic feature module, so as to obtain masked acoustic features and complete acoustic features. Then, the acoustic extraction unit is trained with the sample semantic features, masked acoustic features and complete acoustic features, so that the acoustic extraction unit can extract acoustic features from semantic features, that is, the acoustic extraction unit automatically learns the mapping relationship between semantic features and acoustic features, and the sample acoustic features can be obtained. In this way, acoustic feature extraction is performed on all sample semantic features, so as to obtain sample acoustic features corresponding to each sample text.
[0023] S40, performing speech generation on the sample acoustic features through the acoustic decoding module in the preset training model to obtain predicted generated speech corresponding to each of the sample texts.
[0024] Understandably, predicted generated speech refers to audio information predicted and generated by a preset training model based on sample text.
[0025] Specifically, the acoustic features of the samples are used to generate speech through the acoustic decoding module in the preset training model. That is, the acoustic features corresponding to each sample text are first converted into the corresponding linear spectrum graph through the acoustic decoding module, and then the linear spectrum graph is converted into waveform audio through the voice coder, so that the predicted generated speech corresponding to each sample text can be obtained.
[0026] S50, determining the prediction loss value of the preset training model according to the predicted generated speech and the sample speech corresponding to the same sample text.
[0027] It can be understood that the prediction loss value is generated during the process of predicting sample data by the preset training model.
[0028] Specifically, after obtaining the predicted generated speech, all predicted generated speech corresponding to the sample texts are arranged in the order of the sample texts in the sample text set, and then the predicted generated speech associated with the sample texts is compared with the sample speech of the sample texts with the same sequence; that is, according to the sorting of the sample texts, the sample speech corresponding to the first sample text is compared with the predicted generated speech corresponding to the first sample text, and the loss value between the sample speech and the predicted generated speech is determined by the loss function, and then the sample speech corresponding to the second sample text is compared with the predicted generated speech corresponding to the second sample text, until all sample speech and the predicted generated speech are compared, and the prediction loss value of the preset training model can be obtained.
[0029] S60, when the predicted loss value does not reach the preset convergence condition, iteratively update the initial parameters in the preset training model until the predicted loss value reaches the convergence condition, and record the preset training model after convergence as a speech generation model.
[0030] Understandably, the convergence condition can be that the predicted loss value is less than a set threshold, or the training can be stopped when the predicted loss value is very small and will not decrease after 50,000 calculations.
[0031] Specifically, after obtaining the predicted loss value, when the predicted loss value does not reach the preset convergence condition, the initial parameters of the preset training model are adjusted according to the predicted loss value, and all sample texts and sample speech are re-input into the preset training model with the adjusted initial parameters, and the preset training model with the adjusted initial parameters is iteratively trained to obtain the predicted loss value corresponding to the preset training model with the adjusted initial parameters. Furthermore, when the predicted loss value does not reach the preset convergence condition, the initial parameters of the preset training model are adjusted again according to the predicted loss value, so that the predicted loss value of the preset training model with the adjusted initial parameters reaches the preset convergence condition. In this way, the accuracy of the preset training model becomes higher and higher, and the prediction result continuously approaches the correct result, until the predicted loss value of the preset training model reaches the preset convergence condition, and the preset training model after convergence is determined as the speech generation model.
[0032] In an embodiment of the present invention, a speech generation model training method realizes the simulation of the conditional distribution of semantic features and acoustic features through the semantic feature module and acoustic feature module in the preset training model, thereby improving the quality of the generated speech, improving the efficiency of speech generation during training, and further improving the accuracy of the model and improving the performance of the model. Furthermore, the preset training model is iteratively trained through a large amount of sample training data, and the overall loss value of the preset training model is calculated by comparing the loss function, thereby realizing the determination of the predicted loss value of the preset training model. The initial parameters of the preset training model are adjusted according to the predicted loss value until the model converges, thereby realizing the training of the speech generation model and ensuring that the speech generation model has a high accuracy.
[0033] In one embodiment, in step S20, semantic feature recognition is performed on the sample text and the sample speech by the semantic feature module in the preset training model to obtain sample semantic features corresponding to each of the sample texts, including: S201 , performing semantic encoding on the sample speech by using a semantic encoder in the semantic feature module to obtain a complete semantic feature and a masked semantic feature.
[0034] S202: Perform text encoding on the sample text by using the text encoder in the semantic feature module to obtain text encoding features.
[0035] S203 , performing semantic feature recognition on the complete semantic feature, the masked semantic feature, and the text encoding feature through the semantic extraction unit in the semantic feature module to obtain a sample semantic feature corresponding to the sample data.
[0036] In other words, complete semantic features are a set of features that fully and accurately reflect the text content expressed in language. Masked semantic features are features obtained by masking or hiding certain parts of the text content. Text encoding features are the various characteristics or attributes that appear after converting text into a digital vector form that computers can understand and process.
[0037] Specifically, after obtaining the sample text and sample speech, the semantic encoder in the semantic feature module performs a first semantic encoding on the sample speech to obtain a complete semantic feature. Then, the semantic encoder in the semantic feature module performs a second semantic encoding on the sample speech to obtain a masked semantic feature. Furthermore, the text encoder in the semantic feature module performs text encoding on the sample text, that is, the sample text is segmented, and each segmentation result is encoded, and the position encoding of each segmentation result is added, so that the text encoding feature corresponding to the sample text can be obtained. For example, the convolution layer can be used to extract local features in the sample text, such as capturing some phrases or local pattern features in the sample text. The sample text is scanned by convolution kernels of different sizes to extract features of different scales. For example, a small convolution kernel may capture features at the word level, and a large convolution kernel can capture features at the sentence or paragraph level.
[0038] Next, the semantic feature extraction unit in the semantic feature module performs semantic feature recognition on the complete semantic features, masked semantic features, and text encoding features, that is, the semantic extraction unit is trained through the complete semantic features, masked semantic features, and text encoding features, so that the semantic extraction unit has the ability to extract semantic features and the semantic extraction unit has the ability to complete, thereby obtaining sample semantic features corresponding to the sample data. For example, in a medical scenario, semantic feature extraction is performed on the case text, that is, the case text is first encoded by a text encoder, and then the encoding features are semantically extracted by the semantic extraction unit to obtain semantic features. In a financial scenario, when applying for a bank card, the business to be handled is selected, and then the corresponding operation process text is found. The semantic feature extraction unit performs semantic feature extraction on the operation process text to obtain semantic features.
[0039] In this embodiment, the semantic encoder is used to encode complete semantic features and masked semantic features, thereby completing missing semantic features. The text encoder is used to obtain text encoding features. The semantic extraction unit is used to extract semantic features, thereby improving the accuracy of semantic feature extraction and the quality of the generated speech.
[0040] In one embodiment, in step S203, semantic encoding is performed on the sample speech by the semantic encoder in the semantic feature module to obtain complete semantic features and masked semantic features, including: S2031 , performing a first semantic extraction on the sample speech by a first semantic extraction unit in the semantic encoder to obtain complete semantic features.
[0041] S2032: Perform a second semantic extraction on the sample speech by a second semantic extraction unit in the semantic encoder to obtain a masked semantic feature.
[0042] Specifically, the first semantic extraction unit in the semantic encoder performs a first semantic extraction on the sample speech, that is, the sample speech is convolved through a multi-layer convolutional network to obtain convolution features. At the same time, Mel-spectrum feature extraction is performed on the sample speech features, and feature clustering is performed using a K-clustering algorithm to obtain cluster features. The complete semantic features are then determined based on the convolution features and cluster features. Furthermore, the sample speech is convolved through a convolutional network, and the convolution features are masked using a dynamic random mask to obtain mask features. The mask features are then semantically extracted using a transformer module. At the same time, Mel-spectrum feature extraction is performed on the sample speech, and feature clustering is performed using a K-clustering algorithm. Then, based on the extracted mask features and cluster features, masked semantic features can be obtained.
[0043] In this embodiment, the first semantic extraction unit extracts complete semantic features from the sample speech, while the second semantic extraction unit extracts masked semantic features from the sample speech, thereby completing missing semantic features and improving the accuracy of semantic feature extraction.
[0044] In one embodiment, in step S30, acoustic feature recognition is performed on the sample speech and the semantic features by the acoustic feature module in the preset training model to obtain sample acoustic features corresponding to each sample text, including: S301 , acoustically encoding the sample speech through the acoustic encoder in the acoustic feature module to obtain complete acoustic features and masked acoustic features.
[0045] S302: Performing acoustic feature recognition on the complete acoustic feature, the masked acoustic feature, and the semantic feature through the acoustic extraction unit in the acoustic feature module to obtain a sample acoustic feature corresponding to the sample text.
[0046] It can be understood that a complete acoustic feature refers to a set of feature parameters that can fully and accurately describe a speech or other sound signal. A masked acoustic feature refers to a feature obtained by masking part of the acoustic feature.
[0047] Specifically, after obtaining the semantic features, the sample speech is acoustically encoded by the acoustic encoder in the acoustic feature module, that is, the sample speech is acoustically encoded for the first time by the acoustic encoder, thereby obtaining a complete acoustic feature. Similarly, the sample speech is acoustically encoded for the second time by the acoustic encoder, thereby obtaining a masked acoustic feature. Further, the acoustic extraction unit in the acoustic feature module performs acoustic feature recognition on the complete acoustic features, masked acoustic features and semantic features, that is, the acoustic extraction unit is trained by the complete acoustic features, masked acoustic features and semantic features, so that the acoustic extraction unit learns the mapping relationship of extracting acoustic features from semantic features, and the acoustic extraction unit is specifically complemented by the mask, so that the sample acoustic features can be obtained. For example, in the financial field, when a customer consults about a financial product, if the customer speaks fast and has an urgent tone, it may mean that he has a strong interest in the product or is anxious to get a reply. The customer service system can give priority to handling the problems of such customers. In the medical field, for patients with speech disorders caused by stroke, brain injury, and other causes, the effectiveness of rehabilitation treatment can be assessed by monitoring changes in the acoustic characteristics of the patient's voice, such as pitch and duration. For example, improvements in acoustic characteristics such as pitch stability and pronunciation accuracy before and after treatment can provide a direct indicator of the effectiveness of rehabilitation treatment.
[0048] In this embodiment, the acoustic encoder is used to obtain complete acoustic features and masked acoustic features, thereby completing missing acoustic features. The acoustic extraction unit is used to extract acoustic features, thereby improving the accuracy of acoustic feature extraction and the quality of generated speech.
[0049] In one embodiment, in step S302, the acoustic encoder in the acoustic feature module acoustically encodes the sample speech to obtain complete acoustic features and masked acoustic features, including: S3021: Perform a first acoustic extraction on the sample speech by using a first acoustic extraction unit in the acoustic encoder to obtain complete acoustic features.
[0050] S3022: Perform a second acoustic extraction on the sample speech by a second acoustic extraction unit in the acoustic encoder to obtain a masked acoustic feature.
[0051] Specifically, the first acoustic extraction unit in the acoustic encoder performs an initial acoustic extraction on the sample speech. This involves segmenting the sample speech into fixed-length frames and applying a window function to each frame. This window function smoothes the frame boundaries. Time-frequency conversion is then performed on the windowed audio frames to obtain information about the energy distribution of the audio at different frequencies. Next, a deep convolutional network in the encoder extracts acoustic features from the time-frequency representation. This involves performing sliding convolutions on the time-frequency feature map with kernels of varying sizes to extract local features. Normalization layers (such as batch normalization) accelerate network training and improve model stability. Activation functions (such as ReLU) introduce nonlinearity, enabling the network to learn more complex feature representations. The network gradually reduces the spatial dimensionality of the feature map while increasing the number of channels, thereby compressing the audio signal into a lower-dimensional feature space and generating complete acoustic features. Furthermore, the second acoustic extraction unit in the acoustic encoder performs a second acoustic extraction on the sample speech. Specifically, the sample speech is first segmented into frames of fixed length, and a window function is applied to each frame of the audio signal. The window function can make the frame boundaries smoother. The windowed audio frames are subjected to time-frequency conversion to obtain the energy distribution information of the audio at different frequencies. Then, the deep convolutional network in the encoder is used to extract acoustic features from the time-frequency representation. When extracting the acoustic features, the time-frequency representation is masked using a dynamic random mask to obtain masked acoustic features.
[0052] In this embodiment, the first acoustic extraction unit extracts complete acoustic features from the sample speech, while the second acoustic extraction unit extracts masked acoustic features from the sample speech, thereby completing missing acoustic features and improving the accuracy of acoustic feature extraction.
[0053] In one embodiment, if Figure 3 As shown, a speech generation method includes: S11, obtaining a text to be processed, and performing text encoding on the text to be processed to obtain target text features.
[0054] S12, calling a speech generation model, wherein the speech generation model is trained using the above-mentioned speech generation model training method.
[0055] S13, performing speech generation on the target text features through the speech generation model to obtain target generated speech.
[0056] Understandably, the "unprocessed text" refers to the text to be converted into speech. For example, in medical scenarios, information about patient locations, such as departments, examination rooms, and pharmacies, can be converted into speech, or operating instructions and warnings for medical equipment can be converted into voice prompts. In financial scenarios, when answering questions, responses can be converted into voice playback, or intelligent customer service can provide real-time voice guidance to customers, converting the text of the operational process into clear voice prompts. The target generated speech refers to the speech corresponding to the unprocessed text, for example, "Based on your investment preferences, I recommend a stock fund with high return potential..."
[0057] Specifically, a text to be processed is obtained, which can be generated text or preset text. The text to be processed is then encoded using a text encoder. Specifically, the text to be processed is first segmented, and then the segmentation results are vector-encoded to obtain target text features. Next, a speech generation model is invoked and the target text features are input into the speech generation model. The semantic feature module within the speech generation model performs semantic feature recognition on the target text features. This is done by using the semantic feature recognition capabilities learned during training to identify and extract semantic features from the target text features, thereby obtaining target semantic features. The acoustic feature module within the speech generation model then performs acoustic feature recognition on the target semantic features. This is done by using the acoustic feature recognition capabilities learned during training to extract acoustic features from the target semantic features, thereby obtaining target acoustic features. Finally, the target acoustic features are converted using a decoder to obtain the target generated speech. The target acoustic features can be Mel-frequency cepstral coefficients or linear prediction cepstral coefficients. The target semantic features refer to characteristic attributes that can represent the semantic information contained in a text or language segment.
[0058] In this embodiment, text to speech conversion is performed on the text to be processed through the speech generation model, thereby achieving text-to-speech conversion and improving the accuracy of the conversion, thereby achieving the quality of the target generated speech and improving the generation efficiency of the target generated speech.
[0059] It should be understood that the order of the steps in the above embodiments does not necessarily imply a specific order of execution. The order of execution of each process is determined by its function and inherent logic and does not constitute any limitation on the implementation of the embodiments of the present invention. Any software tools or components not owned by our company that appear in the embodiments of the present invention are for illustrative purposes only and do not represent actual use.
[0060] In one embodiment, a speech generating device is provided, which corresponds to the speech generating method in the above embodiment. Figure 5As shown, the speech generation device includes a target text feature module 11, a speech generation model module 12 and a target generated speech module 13. The functional modules are described in detail as follows: The target text feature module 11 is used to obtain the text to be processed and perform text encoding on the text to be processed to obtain target text features; The speech generation model module 12 is used to call the speech generation model, wherein the speech generation model is trained by the above-mentioned speech generation model training method; The target speech generation module 13 is configured to generate speech based on the target text features through the speech generation model to obtain target generated speech.
[0061] In one embodiment, a speech generation model training device is provided, which corresponds to the speech generation model training method in the above embodiment. Figure 4 As shown, the speech generation model training device includes a sample data acquisition module 10, a semantic feature recognition module 20, an acoustic feature recognition module 30, a predicted speech generation module 40, a model loss prediction module 50 and a model convergence module 60. The functional modules are described in detail as follows: The sample data acquisition module 10 is used to acquire a sample data set, wherein the sample data set includes at least one sample text and its corresponding sample speech; The semantic feature recognition module 20 is used to obtain a preset training model, and perform semantic feature recognition on the sample text and the sample speech through the semantic feature module in the preset training model to obtain sample semantic features corresponding to each sample text; An acoustic feature recognition module 30 is configured to perform acoustic feature recognition on the sample speech and the sample semantic features using the acoustic feature module in the preset training model to obtain sample acoustic features corresponding to each sample text; the semantic feature module and the acoustic feature module adopt a conditional flow matching model; A predicted speech generation module 40 is configured to generate speech based on the acoustic features of the sample using the acoustic decoding module in the preset training model to obtain predicted generated speech corresponding to each of the sample texts; A model loss prediction module 50 is configured to determine a prediction loss value of the preset training model based on the predicted generated speech and the sample speech corresponding to the same sample text; The model convergence module 60 is used to iteratively update the initial parameters in the preset training model when the predicted loss value does not reach the preset convergence condition, until the predicted loss value reaches the convergence condition, and record the preset training model after convergence as the speech generation model.
[0062] In one embodiment, the semantic feature recognition module 20 includes: A speech encoding unit, configured to perform semantic encoding on the sample speech through a semantic encoder in the semantic feature module to obtain complete semantic features and masked semantic features; A text encoding unit, configured to perform text encoding on the sample text using a text encoder in the semantic feature module to obtain text encoding features; The semantic feature extraction unit is used to perform semantic feature recognition on the complete semantic feature, the mask semantic feature and the text encoding feature through the semantic extraction unit in the semantic feature module to obtain a sample semantic feature corresponding to the sample data.
[0063] In one embodiment, the semantic feature extraction unit includes: A complete semantic feature subunit, configured to perform a first semantic extraction on the sample speech through the first semantic extraction unit in the semantic encoder to obtain a complete semantic feature; The masked semantic feature subunit is used to perform a second semantic extraction on the sample speech through the second semantic extraction unit in the semantic encoder to obtain a masked semantic feature.
[0064] In one embodiment, the acoustic feature recognition module 30 includes: an acoustic encoding unit, configured to acoustically encode the sample speech through the acoustic encoder in the acoustic feature module to obtain complete acoustic features and masked acoustic features; An acoustic feature extraction unit is used to perform acoustic feature recognition on the complete acoustic feature, the masked acoustic feature and the semantic feature through the acoustic extraction unit in the acoustic feature module to obtain a sample acoustic feature corresponding to the sample text.
[0065] In one embodiment, the acoustic feature extraction includes: A complete acoustic feature subunit, configured to perform a first acoustic extraction on the sample speech through the first acoustic extraction unit in the acoustic encoder to obtain a complete acoustic feature; The masked acoustic feature subunit is configured to perform a second acoustic extraction on the sample speech through the second acoustic extraction unit in the acoustic encoder to obtain a masked acoustic feature.
[0066] For the specific definition of the speech generation model training device, please refer to the definition of the speech generation model training method above, which will not be repeated here. The various modules in the above-mentioned speech generation model training device can be implemented in whole or in part by software, hardware and their combination. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0067] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 6 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the data used in the speech generation model training method in the above embodiment. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, the above speech generation model training method or the above speech generation method is implemented.
[0068] In one embodiment, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned speech generation model training method or the above-mentioned speech generation method when executing the computer program.
[0069] In one embodiment, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the computer program implements the above-mentioned speech generation model training method, or implements the above-mentioned speech generation method.
[0070] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0071] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0072] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the scope of protection of the present invention.
Claims
1. A speech generation model training method, characterized in that: include: Acquire a sample data set, the sample data set including at least one sample text and its corresponding sample speech; Obtaining a preset training model, and performing semantic feature recognition on the sample text and the sample speech using a semantic feature module in the preset training model to obtain sample semantic features corresponding to each of the sample texts; Performing acoustic feature recognition on the sample speech and the sample semantic features through the acoustic feature module in the preset training model to obtain sample acoustic features corresponding to each of the sample texts; The semantic feature module and the acoustic feature module adopt a conditional flow matching model; Performing speech generation on the sample acoustic features through the acoustic decoding module in the preset training model to obtain predicted generated speech corresponding to each of the sample texts; Determining a prediction loss value of the preset training model based on the predicted generated speech and the sample speech corresponding to the same sample text; When the predicted loss value does not reach the preset convergence condition, the initial parameters in the preset training model are iteratively updated until the predicted loss value reaches the convergence condition, and the preset training model after convergence is recorded as the speech generation model.
2. The speech generation model training method according to claim 1, wherein The semantic feature recognition of the sample text and the sample speech by the semantic feature module in the preset training model to obtain the sample semantic features corresponding to each of the sample texts includes: Performing semantic encoding on the sample speech by a semantic encoder in the semantic feature module to obtain a complete semantic feature and a masked semantic feature; Performing text encoding on the sample text by using the text encoder in the semantic feature module to obtain text encoding features; The semantic extraction unit in the semantic feature module performs semantic feature recognition on the complete semantic feature, the mask semantic feature and the text encoding feature to obtain a sample semantic feature corresponding to the sample data.
3. The speech generation model training method according to claim 2, wherein: The semantic encoding of the sample speech by the semantic encoder in the semantic feature module to obtain complete semantic features and masked semantic features includes: Performing a first semantic extraction on the sample speech by a first semantic extraction unit in the semantic encoder to obtain a complete semantic feature; The second semantic extraction unit in the semantic encoder performs a second semantic extraction on the sample speech to obtain a masked semantic feature.
4. The speech generation model training method according to claim 1, wherein: The acoustic feature recognition is performed on the sample speech and the semantic feature by the acoustic feature module in the preset training model to obtain the sample acoustic features corresponding to each of the sample texts, including: Acoustically encoding the sample speech by an acoustic encoder in the acoustic feature module to obtain complete acoustic features and masked acoustic features; The acoustic extraction unit in the acoustic feature module performs acoustic feature recognition on the complete acoustic feature, the mask acoustic feature and the semantic feature to obtain a sample acoustic feature corresponding to the sample text.
5. The speech generation model training method according to claim 4, wherein: The acoustic encoder in the acoustic feature module acoustically encodes the sample speech to obtain complete acoustic features and masked acoustic features, including: Performing a first acoustic extraction on the sample speech by a first acoustic extraction unit in the acoustic encoder to obtain complete acoustic features; The second acoustic extraction unit in the acoustic encoder performs a second acoustic extraction on the sample speech to obtain a masked acoustic feature.
6. A speech generation method, characterized in that: include: Obtaining a text to be processed, and performing text encoding on the text to be processed to obtain target text features; Calling a speech generation model, wherein the speech generation model is trained by the speech generation model training method according to any one of claims 1 to 5; The target text features are subjected to speech generation by the speech generation model to obtain target generated speech.
7. A speech generation model training device, characterized in that: include: A sample data acquisition module, configured to acquire a sample data set, wherein the sample data set includes at least one sample text and its corresponding sample speech; A semantic feature recognition module is used to obtain a preset training model, and perform semantic feature recognition on the sample text and the sample speech through the semantic feature module in the preset training model to obtain sample semantic features corresponding to each of the sample texts; An acoustic feature recognition module, configured to perform acoustic feature recognition on the sample speech and the sample semantic features using the acoustic feature module in the preset training model to obtain sample acoustic features corresponding to each of the sample texts; The semantic feature module and the acoustic feature module adopt a conditional flow matching model; A predicted speech generation module, configured to generate speech based on the acoustic features of the samples using the acoustic decoding module in the preset training model, to obtain predicted generated speech corresponding to each of the sample texts; A model loss prediction module, configured to determine a prediction loss value of the preset training model based on the predicted generated speech and the sample speech corresponding to the same sample text; The model convergence module is used to iteratively update the initial parameters in the preset training model when the predicted loss value does not reach the preset convergence condition, until the predicted loss value reaches the convergence condition, and record the preset training model after convergence as the speech generation model.
8. A speech generating device, characterized in that: include: A target text feature module is used to obtain the text to be processed and perform text encoding on the text to be processed to obtain target text features; A speech generation model module, configured to call a speech generation model, wherein the speech generation model is trained by the speech generation model training method according to any one of claims 1 to 5; The target speech generation module is used to generate speech based on the target text features through the speech generation model to obtain target generated speech.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, it implements the speech generation model training method as described in any one of claims 1 to 5, or implements the speech generation method as described in claim 6.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it implements the speech generation model training method as described in any one of claims 1 to 5, or implements the speech generation method as described in claim 6.