Speech generation method and device based on conditional flow matching model and related components
By employing a variable step size training and inference method in the conditional flow matching model, the problem of inaccurate speech generation caused by fixed step size is solved, achieving more efficient and higher quality speech generation.
Patent Information
- Application Number
- CN202510222737.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-02-26
AI Technical Summary
Existing conditional flow matching models do not produce accurate and reliable inference results during speech generation, resulting in generally poor speech quality. This is mainly due to information loss or wasted computational resources caused by fixed step sizes.
A variable step size is used to infer the conditional flow matching model. The model is trained by acquiring training text and noise data of the generated speech, and the speech generation model is constructed. The variable step size is then used for inference and prediction.
It improves the quality and efficiency of speech generation, ensuring that the model avoids wasting resources while acquiring key information, and the generated speech is more natural and accurate.
Smart Images

Figure CN119993115B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and in particular relates to a speech generation method and device based on a conditional flow matching model and related components. BACKGROUND
[0002] With the in-depth study of generative models, speech generation models based on different schemes are also being continuously proposed, including diffusion models based on stochastic differential equations and conditional flow matching (CFM) models based on ordinary differential equations. These models have shown very outstanding effects in the field of picture generation, and have gradually been applied to the field of speech generation, and have also shown good effects in speech generation quality. For example, intelligent diagnosis and treatment in medical scenarios, remote consultation, and some financial business processing platforms in financial scenarios.
[0003] Among them, the conditional flow matching model has higher inference efficiency than the diffusion model. However, since the conditional flow matching model was proposed relatively recently, the current field mostly focuses on improving the model framework, mainly in the model training stage, and there are relatively few improvements in the inference stage, which leads to the fact that the inference result based on the conditional flow matching model is not accurate and reliable enough, and further leads to the fact that the speech quality generated by the conditional flow matching model is generally poor. SUMMARY
[0004] The embodiments of the present application provide a speech generation method and device based on a conditional flow matching model, a computer device and a storage medium, aiming to improve the speech generation quality.
[0005] In a first aspect, the embodiments of the present application provide a speech generation method based on a conditional flow matching model, comprising:
[0006] obtaining training text of generated speech, and taking the generated speech as training speech;
[0007] sampling noise data, inputting the training text, training speech and noise data into a conditional flow matching model for training;
[0008] using a variable step length to infer the conditional flow matching model to obtain a vector field about noise and speech;
[0009] constructing a speech generation model based on the conditional flow matching model and the vector field, and using the speech generation model to infer and predict target text of speech to be generated.
[0010] In a second aspect, the embodiments of the present application provide a speech generation device based on a conditional flow matching model, comprising:
[0011] a text acquisition unit configured to acquire training text of generated speech and take the generated speech as training speech;
[0012] a data input unit configured to sample noise data and input the training text, the training speech and the noise data into the conditional flow matching model for training;
[0013] a model inference unit configured to infer the conditional flow matching model by using a variable step length to obtain a vector field related to noise and speech;
[0014] a prediction unit configured to construct a speech generation model based on the conditional flow matching model and the vector field and infer and predict target text to be generated speech by using the speech generation model.
[0015] In a third aspect, an embodiment of the present application provides a computer device, including a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor implements the speech generation method based on the conditional flow matching model according to the first aspect when executing the computer program.
[0016] In a fourth aspect, an embodiment of the present application provides a computer readable storage medium, and the computer readable storage medium stores a computer program, and the computer program is executable on a processor to implement the speech generation method based on the conditional flow matching model according to the first aspect.
[0017] The embodiment of the present application provides a speech generation method, device, computer device and storage medium based on a conditional flow matching model, and the method includes the following steps: acquiring training text of generated speech and taking the generated speech as training speech; sampling noise data and inputting the training text, the training speech and the noise data into the conditional flow matching model for training; inferring the conditional flow matching model by using a variable step length to obtain a vector field related to noise and speech; and constructing a speech generation model based on the conditional flow matching model and the vector field, and inferring and predicting target text to be generated speech by using the speech generation model. The embodiment of the present application trains the conditional flow matching model by using the text of the generated speech, constructs the speech generation model, and infers and solves the conditional flow matching model by using a variable step length in the training process of the conditional flow matching model, so that the speech can be quickly generated and the speech generation quality can be improved. BRIEF DESCRIPTION OF DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed to be used in the embodiments description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and for those skilled in the art, other drawings can be obtained without creative labor.
[0019] Figure 1 An application environment schematic diagram of the speech generation method based on the conditional flow matching model provided by the embodiments of the present application is shown in the figure.
[0020] Figure 2 A flow schematic diagram of the speech generation method based on the conditional flow matching model provided by the embodiments of the present application is shown in the figure.
[0021] Figure 3 A sub-flow schematic diagram of the speech generation method based on the conditional flow matching model provided by the embodiments of the present application is shown in the figure.
[0022] Figure 4 Another flow schematic diagram of the speech generation method based on the conditional flow matching model provided by the embodiments of the present application is shown in the figure.
[0023] Figure 5 A principle architecture diagram of the speech generation method based on the conditional flow matching model provided by the embodiments of the present application is shown in the figure.
[0024] Figure 6 A schematic block diagram of the speech generation device based on the conditional flow matching model provided by the embodiments of the present application is shown in the figure.
[0025] Figure 7 A sub-schematic block diagram of the speech generation device based on the conditional flow matching model provided by the embodiments of the present application is shown in the figure.
[0026] Figure 8 Another sub-schematic block diagram of the speech generation device based on the conditional flow matching model provided by the embodiments of the present application is shown in the figure.
[0027] Figure 9 A schematic block diagram of a computer device provided by the embodiments of the present application is shown in the figure. DETAILED DESCRIPTION
[0028] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the present application.
[0029] It should be understood that the terms "comprises" and "comprising," when used in this specification and accompanying claims, indicate the presence of the described features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0030] It should also be understood that the terms used in the specification and the appended claims are intended to describe particular embodiments and do not intend to limit the present application. As used in the specification and the appended claims, the singular forms "a," "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise.
[0031] It should further be understood that the term "and / or" used in the specification and the appended claims means one or more of the associated listed items as well as all possible combinations of the items and includes the combinations.
[0032] The speech generation method based on the conditional flow matching model provided by the embodiments of the present application can be applied in the application environment of Figure 1 , wherein the client communicates with the server through the network. The server can obtain the training text of the generated speech, and take the generated speech as the training speech; sample the noise data, input the training text, the training speech and the noise data into the conditional flow matching model for training; use the variable step to infer the conditional flow matching model to obtain the vector field about the noise and the speech; construct the speech generation model based on the conditional flow matching model and the vector field, and use the speech generation model to infer and predict the target text to be generated. The embodiments of the present application use the text of the generated speech to train the conditional flow matching model, thereby constructing the speech generation model, and in the training process of the conditional flow matching model, the variable step is used to infer and solve it, so that not only the speech can be quickly generated, but also the speech generation quality can be improved. The present application will be described in detail through specific embodiments.
[0033] Please refer to Figure 2 , the speech generation method based on the conditional flow matching model provided by the embodiments of the present application specifically includes steps S101-S104.
[0034] Step S101, obtaining the training text of the generated speech, and taking the generated speech as the training speech;
[0035] Step S102, sampling the noise data, inputting the training text, the training speech and the noise data into the conditional flow matching model for training;
[0036] Step S103, reasoning the conditional flow matching model with variable step length to obtain a vector field about noise and speech;
[0037] Step S104, constructing a speech generation model based on the conditional flow matching model and the vector field, and reasoning and predicting the target text to be generated speech by using the speech generation model.
[0038] In this embodiment, first, the training text corresponding to the successfully generated speech is obtained, then the noise data is collected, and the training text and the speech obtained are input into the conditional flow matching model, and then the conditional flow matching model is reasoned with variable step length to obtain the corresponding reasoning prediction result, thereby constructing the speech generation model. The specified target text can be reasoned and predicted by the speech generation model.
[0039] This embodiment uses the text of the generated speech to train the conditional flow matching model to construct the speech generation model, and in the training process of the conditional flow matching model, variable step length is used to reason and solve it, so it can not only quickly generate speech, but also improve the quality of speech generation.
[0040] It should be noted that the prior art uses fixed step length to reason the conditional flow matching model, which may result in too large or too small step length in a certain step of reasoning. When the step length is too large, some key information may be skipped when the input data (such as text) is divided and input into the conditional flow matching model. For example, in the text-to-speech conversion task, if the text is divided according to a larger step length, a complete semantic unit (such as a word group or a clause) may be divided. In this way, the model may not be able to completely obtain semantic information during reasoning, resulting in deviation in semantic expression of the generated speech. For example, a complex technical document is converted into speech, the document may contain professional terms and complex sentence structures. If the step length is too large, the key technical terms and their modifiers may be separated, so that the model cannot accurately understand the relationship between the words, and thus incorrectly emphasizes or pronounces these words when generating speech. In addition, for input data containing rich details, a larger step length will cause the model to ignore these details. In the field of speech generation, the details of the audio include tone, intonation, and connected reading. If the step length is too large, the model may not be able to accurately process these details during reasoning, and the generated speech may sound stiff. For example, when imitating the speech style of a certain character, some unique details of the character's speech (such as slight tremolo, specific intonation changes, etc.) may be ignored due to the large step length. The model may only be able to roughly generate the outline of the speech, but cannot accurately reproduce these subtle speech features.
[0041] If the step size is too small, it will significantly increase the number of steps that need to be processed during inference. This means that the model needs more computational resources and time to complete the inference. For example, when processing a long text-to-speech task, if the step size is set too small, the model may need to perform complex calculations and matching for each word or even each syllable, which will greatly increase the amount of computation. For resource-limited devices such as mobile devices, this situation may cause the program to run slowly or even crash. Moreover, too many calculation steps may also cause the model to overfit some local information, affecting the grasp of the overall content. In addition, too small step size will make the model pay too much attention to local input data, which may result in the generated speech lacking coherence. The model may overemphasize each small input unit, ignoring the overall connection between them. For example, when generating a continuous narrative voice, a too small step size may make the model accurate in pronouncing each word, but at the sentence level, the tone, speed, and other aspects may not correctly reflect the semantic and emotional meaning of the sentence, making it sound unnatural.
[0042] This embodiment adopts variable step size for inference, so that each step of inference has a suitable step size. This not only allows the model to effectively utilize input information while avoiding unnecessary calculations. It can ensure that the model fully captures important information such as semantic and speech features, without wasting resources by processing too many details. For example, when converting a news article into speech, a suitable step size can allow the model to reason according to the structure of sentences or phrases, accurately understanding the content of the article while efficiently generating natural and fluent speech. It also helps the model better capture the details and overall style of the speech. The model can adjust the prosody, rhythm, and timbre of the speech based on the input units divided by the step size. For example, when imitating the speech of a certain dialect, a suitable step size can allow the model to properly handle the special pronunciation, tone changes, and other details in the dialect, generating high-quality speech that meets the characteristics of the dialect.
[0043] For example, in a medical scenario, when obtaining medical-related documents as training text, sample corresponding noise data, and input the medical-related documents and corresponding training speech, noise data into the conditional flow matching model for training. Since the medical-related documents contain many medical-related professional terms and complex medical sentences, such as disease names, symptoms and signs, and treatment methods, etc., the variable step size is used to infer the conditional flow matching model with input medical-related documents, so that each step of inference has a suitable step size. A speech generation model suitable for a medical scenario is generated, so that the speech generation model can be used to infer and predict the medical text to be generated.
[0044] For example, in a financial scenario, when a financial-related document is obtained as training text, corresponding noise data is sampled, and the financial-related document and corresponding training speech and noise data are input into a conditional flow matching model for training. Since the financial-related document contains many professional terms related to finance and complex financial sentences, such as capital market, primary market, and financial tools, etc., a variable step is used to infer the conditional flow matching model input with the financial-related document, so that each step of inference has a suitable step. Thus, a speech generation model suitable for the financial scenario is generated, so that the speech generation model can be used to infer and predict the financial text to be generated.
[0045] In an embodiment, the sampling of noise data, the input of the training text, training speech, and noise data into the conditional flow matching model for training comprises:
[0046] The training text is encoded, and the text duration is obtained according to the encoding result.
[0047] A time is randomly sampled as a time step, and the time step, the training text, training speech, noise data, and text duration are input into the conditional flow matching model for training.
[0048] Specifically, the training expression of the conditional flow matching model is:
[0049] X t =tX1+(1-(1-σ min )t)X0;
[0050]
[0051] Wherein, X t represents the training result of the conditional flow matching model, X1 represents the training speech, X0 represents the noise data, t represents the time step, and σ min represents a coefficient.
[0052] In combination with Figure 5 , first introduce the general framework of the conditional flow matching (CFM) speech generation model. When constructing a speech generation model based on a conditional flow matching model, a common differential equation When t = 0, X t obeys a Gaussian distribution, i.e., noise, and when t = 1, it is the desired data distribution, which can be understood as the Mel spectrum distribution of the speech to be generated.
[0053] During the model training, the text (for example, medical related text) is first encoded, and after the encoding, the duration prediction is performed to ensure that the length of the generated voice can correspond to the length of the text. Here, the information obtained through the text is denoted as μ, which is input into the conditional flow matching model, and a time t is randomly sampled in [0, 1], and a Gaussian noise (corresponding to X(t=0)) is sampled, which is input into the conditional flow matching model together with the voice corresponding to the text (corresponding to X(t=1)) for training.
[0054] With the above conditions, the training of the model can be started. Here, the path from X(t=0) to X(t=1) is selected as the optimal transmission condition vector field, and the expression is as follows:
[0055]
[0056] In an embodiment, as shown in Figure 3 , the sampling noise data, the training text, the training voice, and the noise data are input into the conditional flow matching model for training, and further include steps S201-S203.
[0057] In step S201, a velocity field v t (X t |μ) of the training result X t is set, where μ represents an input parameter of the conditional flow matching model.
[0058] In step S202, the conditional flow matching model is target trained according to the following formula:
[0059]
[0060] Where minL represents the target training result of the conditional flow matching model, and represents the expected distribution based on the noise data.
[0061] In step S203, a common differential equation
[0062] In this embodiment, a velocity field v t (X t |μ) is trained by means of the conditional flow matching model, which represents the speed and direction of the noise data distribution to the real data distribution, and through which the above can be approximated, that is, the training target is:
[0063]
[0064] After the model training is completed, a common differential equation
[0065] In an embodiment, the inference of the conditional flow matching model with variable step length is used to obtain a vector field about noise and speech, comprising:
[0066] The ordinary differential equation is solved in an explicit Euler format to construct an inference expression; the inference expression is as follows:
[0067]
[0068] wherein X i+1 represents the inference result of the i+1 step, X i represents the inference result of the i step, t i+1 represents the time step of the i+1 step, t i represents the time step of the i step, represents the training result of the i step, represents the velocity field of the i step;
[0069] The inference expression is solved with variable step length to obtain the result of the ordinary differential equation.
[0070] In this embodiment, for any t in [0, 1], a v t (X t |μ) can be calculated by the conditional flow matching model, so that the ordinary differential equation can be numerically solved. Specifically, it is assumed that the inference process is divided into N steps, i.e., the process is X0→X1→…→X N-1 →X N =X1, and the corresponding time nodes are 0=t0→t1→…→t N-1 →t N =1. This embodiment uses explicit Euler format to numerically solve the ordinary differential equation, and the calculation expression is from t=0 to t=1 in turn.
[0071] Specifically, as shown in the figure, the solving of the inference expression with variable step length to obtain the result of the ordinary differential equation comprises steps S301-S303. Figure 4
[0072] Step S301, obtaining the speech signal amplitude square of the training result of the conditional flow matching model;
[0073] Step S302, mapping the energy of the training result based on the speech signal amplitude square;
[0074] Step S303, calculating the step length of the current step based on the energy of each step, and solving the ordinary differential equation of the current step through the step length of the current step.
[0075] Here, the expression for the variable step size is:
[0076] h = a / (e) i -e i-1 );
[0077] Where h represents the variable step size, a represents a constant, and e i Let e represent the energy at step i. i-1 This represents the energy at step i-1.
[0078] Most existing methods use a fixed step size for solving the problem, i.e., h = t1 - t0 = t2 - t1 = ... = t N-1 -t N-2 =t N -t N-1 This embodiment employs a variable step size. After the i-th step calculation is completed, for the obtained X... i Calculate energy e i At the same time, combined with the energy e from the previous step i-1 Therefore, the step size for the next step can be determined from this, that is, the step size expression is: Here, 'a' is a constant that can be determined based on the actual situation. Simply put, the step size of the next step is inversely proportional to the difference in energy gained in the previous two steps; the greater the energy change, the smaller the step size, and vice versa.
[0079] This embodiment calculates the energy value at each time point to compare where the recursive process fluctuates significantly. Therefore, near these time points, it's necessary to improve computational accuracy by reducing the step size to minimize error. When the energy is stable, the step size can be appropriately increased to accelerate inference without affecting the quality of the final generated speech. Compared to a fixed-step-size scheme, a variable-step-size scheme achieves better generation quality when selecting the same number of inference steps N. If the same initial step size is chosen, the variable-step-size scheme can select a larger step size after energy stabilization to accelerate inference. In particular, besides its application in speech generation, variable-step-size can also be applied to any other model based on conditional flow matching or ordinary differential equations to improve inference efficiency and output accuracy.
[0080] In practical application scenarios, after obtaining the training text, the training text is preprocessed, for example, removing redundant spaces, punctuation marks (if special processing is required), special characters, etc. If it is a text in other formats (such as document formats with text information), the text content needs to be extracted. If the text contains semantic labels and other information, these information also need to be parsed for better use of the conditional flow matching model for subsequent processing. Here, if the model needs voice-related features as auxiliary input (for example, referring to existing voice styles, etc.), the corresponding acoustic features need to be extracted from the voice sample library, such as fundamental frequency, intensity, timbre, and other acoustic parameters. These parameters can be obtained through signal processing techniques, for example, using short-time Fourier transform (STFT) to obtain the spectral features of the voice, or through linear predictive coding (LPC) to extract the vocal tract-related features. Thereafter, a pre-trained conditional flow matching model is loaded, and the model parameters can include the weights, biases, etc. of the neural network, which define the structure and function of the model. At the same time, the parameter initialization method related to variable step reasoning in the model is determined, for example, setting the initial step size, hyperparameters related to the step size adjustment strategy, etc.
[0081] Then variable step size is used for conditional flow matching model inference. Specifically, first, inference parameters are initialized, and an initial inference step size is determined. This step size can be set according to empirical values or statistical information in the model training process. For example, in some sequence generation tasks, the initial step size can be related to the average length of the sequence. Second, rules for step size adjustment are set. This can be determined based on performance indicators of the model (such as quality assessment indicators of generated speech), features of the data (such as complexity of the text), or number of iterations, etc. For example, when the quality of the generated speech segment is low, the step size is appropriately reduced to process more finely; when the text content is relatively simple and the generation quality is stable, the step size is appropriately increased to improve inference speed. Third, the inference loop is started, and the selected initial step size is used to start processing the text content (such as text). For text input, the text is divided into multiple segments according to the step size (for example, if the text is a sentence, the number of words determined by the initial step size is used to divide the segments). Each segment is input into the conditional flow matching model. The model generates the corresponding speech features or the prediction of the speech segment according to the input text segment and the previously accumulated context information (such as the text content corresponding to the previously generated speech part). In addition, during the inference process, it can be evaluated whether the step size needs to be adjusted according to the step size adjustment rule. For example, if the model has a large error when generating the current speech segment (which can be measured by speech quality assessment indicators such as speech clarity, naturalness, etc.), the step size is reduced according to the predetermined strategy, so that the text segment is processed more finely in the next round of inference. Continue this inference loop until the entire text content (such as text) is processed. Also, during the inference process, the intermediate generated speech segments or the intermediate state of the model can be saved. For example, when the generated speech needs to be partially modified or adjusted, these intermediate results can provide a convenient starting point. At the same time, the saving of the intermediate state is also helpful for the recovery of the inference process when unexpected situations (such as program interruption) occur.
[0082] Subsequently, in the speech generation stage, according to the speech features or speech segment predictions obtained by model inference, speech synthesis technology is used to convert them into actual speech. If the model directly outputs speech waveform parameters, it can be converted into playable speech signals through digital-to-analog conversion and other methods. If the model outputs the acoustic features of the speech (such as mel spectrum, etc.), a vocoder needs to be used to synthesize the speech. In the speech synthesis process, attention should be paid to the transition between speech segments to ensure the coherence of the speech. This can be achieved by smoothing the speech parameters (such as interpolation in the frequency domain) or using specialized speech splicing techniques. Further, post-processing of the generated speech is performed to improve the quality of the speech. For example, noise reduction processing is performed to remove noise or unnatural audio components that may be introduced during the synthesis process. Volume normalization can also be performed to keep the volume of the speech relatively stable throughout the audio text. In addition, according to the specific application scenario, the speech can be operated at different speeds and pitches to meet the different needs of users.
[0083] For example, in a medical scenario, medical-related text is obtained and preprocessed to extract key information such as disease description and treatment recommendations. Subsequently, according to the semantic content of the text, a suitable speech style and tone are selected. In the medical scenario, a formal, serious and easy-to-understand speech style may need to be selected to ensure that the patient or medical personnel can accurately understand the generated content. During model inference, a variable step strategy is used to adjust the step size according to the complexity and semantic changes of the text content. For medical text passages containing a large number of professional terms and complex descriptions, the model can use a smaller step size to ensure that each detail is accurately understood and generated. In the description of simple or repetitive medical text parts, the model can use a larger step size to improve inference speed and efficiency. In addition, to further improve the quality of the generated speech, some post-processing techniques can be introduced during the inference process. For example, speech synthesis post-processing algorithms can be used to smooth the generated speech waveform, reduce noise and distortion. Speech enhancement techniques can also be used to improve the clarity and intelligibility of the speech, making it more suitable for use in actual medical scenarios.
[0084] In summary, by applying a variable step strategy to the speech generation method based on the conditional flow matching model, the inference efficiency and output accuracy of the model in different application scenarios can be significantly improved. In particular, in the medical, financial and other professional fields, this method can generate more accurate, natural and easy-to-understand speech content, providing users with a better user experience.
[0085] It is also important to ensure that the input text is in the correct format for the model. For example, the text should be encoded correctly, such as using UTF-8, to avoid any issues with character encoding. Additionally, for text that contains special characters, abbreviations, slang, or other content, the model may require special preprocessing or handling strategies. For example, if the input includes additional features such as audio features as references, these features need to be normalized. For example, if the audio spectrum feature is used, the spectrum amplitude range may vary for different audio texts, and it needs to be normalized to a standard range (such as [0, 1] or [-1, 1]), so that the model can more stably calculate the weight of these features during inference. In addition, it is important to ensure that the model parameters are correctly loaded and updated. For example, before inference begins, it is important to ensure that the model parameters are complete and correctly loaded. If the model parameters are updated during inference (for example, using online learning or adaptive strategies), it is important to strictly follow the established update rules. These rules are usually determined during model training and the update process needs to consider the impact on the current inference task. For example, the update step size should not be too large, so as not to destroy the knowledge structure that the model has learned.
[0086] Preferably, during inference, the existing context information should be fully utilized. For example, when generating speech, the speech features and semantic information of the previous word or sentence can help better generate the speech of the next part. This can be achieved by storing the previous information in the model's hidden state or other specialized memory structures. Context information also needs to be updated in a timely manner. When new input data (such as new text segments) is processed, the model should be able to integrate this new information into the existing context and update the context representation for subsequent inference. Otherwise, there may be information lag or error accumulation. Of course, conditional flow matching models may face uncertainty during inference, such as for some ambiguous text content (with multiple possible speech expressions) or noisy audio features. At this time, the model needs to have a certain mechanism to estimate this uncertainty, such as using a probability distribution to represent multiple possible speech generation results.
[0087] In another practical application scenario, when constructing a speech generation model according to a conditional flow matching model and a vector field, a module can be designed to fuse the conditional flow matching model and the speech vector field. This module can be a multi-layer perceptron (MLP) or a more complex neural network structure, such as a combination of a convolutional neural network (CNN) and a recurrent neural network (RNN). The purpose is to effectively fuse the text-related features (such as semantic information, grammatical structure, etc.) output by the conditional flow matching model and the dynamic features (such as pitch variation, timbre variation, etc.) contained in the vector field. For example, in a text-to-speech conversion task, the conditional flow matching model may output speech feature predictions corresponding to each word in the text, such as phoneme sequences and prosody information. The vector field can provide dynamic change information of the actual speech signal in the time domain or frequency domain. The fusion module splices these information and generates a comprehensive speech feature representation through a series of linear and nonlinear transformations.
[0088] After the fusion module, a speech generator module is constructed, which is responsible for converting the fused speech features into actual speech waveforms. For example, using vocoder technology. The vocoder can be divided into parameter-based vocoder and waveform-based vocoder. The parameter-based vocoder (such as Mel Linear Predictive Coding-MLPC) first converts the fused speech features into a set of acoustic parameters (such as linear prediction coefficients, fundamental frequency, etc.), and then synthesizes speech waveforms through these parameters. The waveform-based vocoder (such as WaveNet) directly generates speech waveforms from the fused speech features, which usually utilizes the powerful fitting ability of deep neural networks to learn the generation rules of speech waveforms.
[0089] Figure 6 A schematic block diagram of a speech generation device 500 based on a conditional flow matching model is provided for an embodiment of the application, and the device 500 comprises:
[0090] A text acquisition unit 501 is configured to acquire training text of generated speech, and take the generated speech thereof as training speech.
[0091] A data input unit 502 is configured to sample noise data, and input the training text, the training speech, and the noise data into the conditional flow matching model for training.
[0092] A model inference unit 503 is configured to infer the conditional flow matching model by using a variable step length to obtain a vector field related to noise and speech.
[0093] An inference prediction unit 504 is configured to construct a speech generation model based on the conditional flow matching model and the vector field, and infer and predict a target text to be generated speech by using the speech generation model.
[0094] In this embodiment, first, training text corresponding to successfully generated speech is obtained, then noise data is collected, and the training text and speech obtained are input into the conditional flow matching model, and then the variable step is used to infer the conditional flow matching model to obtain the corresponding inference prediction result, thereby constructing a speech generation model. The speech generation model can be used to infer and predict the specified target text.
[0095] In this embodiment, the text of the generated speech is used to train the conditional flow matching model to construct a speech generation model. In the training process of the conditional flow matching model, the variable step is used to infer and solve it. In this way, not only can the speech be quickly generated, but also the quality of the generated speech can be improved.
[0096] It should be noted that the prior art uses a fixed step to infer the conditional flow matching model, which may result in a step that is too large or too small in a certain step of inference. When the step is too large, some key information may be skipped when the input data (such as text) is divided and input into the conditional flow matching model. For example, in a text-to-speech conversion task, if the text is divided according to a large step, a complete semantic unit (such as a word group or clause) may be divided. In this way, the model may not be able to completely obtain semantic information during inference, resulting in a deviation in the semantic expression of the generated speech. For example, a complex technical document is converted into speech. The document may contain professional terms and complex sentence structures. If the step is too large, the key technical terms and their modifiers may be separated, so that the model cannot accurately understand the relationship between the words, and thus incorrectly emphasizes or pronounces these words when generating speech. In addition, for input data containing rich details, a larger step size will cause the model to ignore these details. In the field of speech generation, the details of the audio include tone, intonation, and connected reading. If the step is too large, the model may not be able to accurately process these details during inference, and the generated speech may sound harsh. For example, when imitating the speech style of a certain character, some unique details of the character's speech (such as slight tremolo, specific intonation changes, etc.) may be ignored due to the large step. The model may only be able to roughly generate the outline of the speech, but may not be able to accurately reproduce these subtle speech features.
[0097] If the step size is too small, it will significantly increase the number of steps that need to be processed during inference. This means that the model needs more computational resources and time to complete the inference. For example, when processing a long text-to-speech task, if the step size is set too small, the model may need to perform complex calculations and matching for each word or even each syllable, which will greatly increase the amount of computation. For resource-limited devices such as mobile devices, this situation may cause the program to run slowly or even crash. Moreover, too many calculation steps may also cause the model to overfit some local information, affecting the grasp of the overall content. In addition, too small step size will make the model pay too much attention to local input data, which may result in the generated speech lacking coherence. The model may overemphasize each small input unit, ignoring the overall connection between them. For example, when generating a continuous narrative voice, a too small step size may make the model accurate in pronouncing each word, but at the sentence level, the tone, speed, and other aspects may not correctly reflect the semantic and emotional meaning of the sentence, making it sound unnatural.
[0098] The present embodiment adopts variable step size for inference, so that each step of inference has a suitable step size. This not only allows the model to effectively utilize input information while avoiding unnecessary calculations. It can ensure that the model fully obtains important information such as semantic and speech features, while not wasting resources by processing too many details. For example, when converting a news article into speech, a suitable step size can allow the model to reason according to the structure of sentences or phrases, accurately understanding the content of the article while efficiently generating natural and fluent speech. It also helps the model better capture the details and overall style of the speech. The model can adjust the prosody, rhythm, and timbre of the speech based on the input units divided by the step size. For example, when imitating the speech of a certain dialect, a suitable step size can allow the model to properly handle the special pronunciation, tone changes, and other details in the dialect, thereby generating high-quality speech that meets the characteristics of the dialect.
[0099] For example, in a medical scenario, when obtaining medical-related documents as training text, sample corresponding noise data, and input the medical-related documents and corresponding training speech and noise data into the conditional flow matching model for training. Since the medical-related documents contain many medical-related professional terms and complex medical sentences, such as disease names, symptoms and signs, and treatment methods, etc., the conditional flow matching model with input of medical-related documents is inferred using variable step size, so that each step of inference has a suitable step size. A speech generation model suitable for a medical scenario is generated, so that the speech generation model can be used to infer and predict the medical text to be generated.
[0100] For example, in a financial scenario, when a financial-related document is obtained as training text, corresponding noise data is sampled, and the financial-related document and corresponding training speech and noise data are input into the conditional flow matching model for training. Since the financial-related document contains many professional terms related to finance and complex financial sentences, such as capital market, primary market, and financial tools, etc., a variable step length is used to infer the conditional flow matching model input with the financial-related document, so that each step of inference has a suitable step length. Thus, a speech generation model suitable for the financial scenario is generated, so that the speech generation model can be used to infer and predict the financial text to be generated.
[0101] In an embodiment, the data input unit 502 comprises:
[0102] a text encoding unit configured to encode the training text and obtain a text duration according to the encoding result;
[0103] a sampling input unit configured to randomly sample a time as a time step and input the time step, the training text, training speech, noise data, and text duration into the conditional flow matching model for training.
[0104] In an embodiment, the training expression of the conditional flow matching model is:
[0105]
[0106]
[0107] wherein X t represents the training result of the conditional flow matching model, X1 represents the training speech, X0 represents the noise data, t represents the time step, σ min represents a coefficient.
[0108] In combination with Figure 5 , first introduce the general framework of the conditional flow matching (CFM) speech generation model. When constructing a speech generation model based on the conditional flow matching model, a common differential equation When t = 0, X t obeys a Gaussian distribution, i.e., noise, and when t = 1, it is the desired data distribution, which can be understood as the Mel spectrum distribution of the speech to be generated.
[0109] During model training, the text (e.g., medical related text) is first encoded, and after the encoding is completed, the duration prediction is performed to ensure that the length of the generated voice finally corresponds to the length of the text. Here, the information obtained through the text can be denoted as μ, which is input into the conditional flow matching model, and a time t is randomly sampled in [0, 1], and a Gaussian noise (corresponding to X(t=0)) is sampled, and the voice corresponding to the text (corresponding to X(t=1)) is input into the conditional flow matching model for training.
[0110] With the above conditions, the training of the model can be started. Here, the path from X(t=0) to X(t=1) is selected as the optimal transmission condition vector field, and the expression is as follows:
[0111]
[0112] In an embodiment, as shown in Figure 7 , the data input unit 502 further includes:
[0113] A velocity field setting unit 601 is configured to set a velocity field v t (X t | μ) about the training result X t , wherein μ represents an input parameter of the conditional flow matching model.
[0114] A target training unit 602 is configured to perform target training on the conditional flow matching model according to the following formula:
[0115]
[0116] Wherein minL represents the target training result of the conditional flow matching model, represents the expected distribution based on the noise data;
[0117] An equation obtaining unit 603 is configured to obtain an ordinary differential equation
[0118] In this embodiment, a velocity field v t (X t | μ) is trained by means of the conditional flow matching model, which represents the speed and direction of the noise data distribution to the real data distribution, and through which the above can be approximated, that is, the training target is:
[0119]
[0120] After the model training is completed, an ordinary differential equation
[0121] In an embodiment, the model inference unit 503 comprises:
[0122] an equation solving unit configured to solve the ordinary differential equation in an explicit Euler format to construct an inference expression, the inference expression being as follows:
[0123]
[0124] wherein X i+1 represents the inference result of the i+1th step, X i represents the inference result of the ith step, t i+1 represents the time step of the i+1th step, t i represents the time step of the ith step, represents the training result of the ith step, represents the velocity field of the ith step;
[0125] a variable inference unit configured to solve the inference expression in a variable step length to obtain the result of the ordinary differential equation.
[0126] In the embodiment, for any t in [0, 1], a v t (X t | μ) can be calculated by the conditional flow matching model, so that the ordinary differential equation can be solved numerically. Specifically, it is assumed that the inference process is divided into N steps, i.e., the process is X0→X1→…→X N-1 →X N = X1, and the corresponding time nodes are 0 = t0→t1→…→t N-1 →t N = 1. The explicit Euler format is used to solve the ordinary differential equation in the embodiment, and the calculation expression is which is recursively calculated from t = 0 to t = 1.
[0127] In an embodiment, as shown in Figure 8 the variable inference unit comprises:
[0128] a signal acquisition unit 701 configured to acquire a speech signal amplitude square of a training result of the conditional flow matching model;
[0129] an energy mapping unit 702 configured to map the speech signal amplitude square to obtain an energy of the training result;
[0130] a step length calculation unit 703 configured to calculate a step length of a current step based on the energy of each step, and solve the ordinary differential equation of the current step by using the step length of the current step.
[0131] In an embodiment, the expression of the variable step length is as follows:
[0132] h = a / (e) i -e i-1 );
[0133] Where h represents the variable step size, a represents a constant, and e i Let e represent the energy at step i. i-1 This represents the energy at step i-1.
[0134] Most existing methods use a fixed step size for solving the problem, i.e., h = t1 - t0 = t2 - t1 = ... = t N-1 -t N-2 =t N -t N-1 This embodiment employs a variable step size. After the i-th step calculation is completed, for the obtained X... i Calculate energy e i At the same time, combined with the energy e from the previous step i-1 Therefore, the step size for the next step can be determined from this, that is, the step size expression is: Here, 'a' is a constant that can be determined based on the actual situation. Simply put, the step size of the next step is inversely proportional to the difference in energy gained in the previous two steps; the greater the energy change, the smaller the step size, and vice versa.
[0135] This embodiment calculates the energy value at each time point to compare where the recursive process fluctuates significantly. Therefore, near these time points, it's necessary to improve computational accuracy by reducing the step size to minimize error. When the energy is stable, the step size can be appropriately increased to accelerate inference without affecting the quality of the final generated speech. Compared to a fixed-step-size scheme, a variable-step-size scheme achieves better generation quality when selecting the same number of inference steps N. If the same initial step size is chosen, the variable-step-size scheme can select a larger step size after energy stabilization to accelerate inference. In particular, besides its application in speech generation, variable-step-size can also be applied to any other model based on conditional flow matching or ordinary differential equations to improve inference efficiency and output accuracy.
[0136] In practical application scenarios, after obtaining the training text, the training text is preprocessed, for example, removing redundant spaces, punctuation marks (if special processing is required), special characters, etc. If it is a text in other formats (such as document formats with text information), the text content needs to be extracted. If the text contains semantic labels and other information, these information also need to be parsed for better use of the conditional flow matching model for subsequent processing. Here, if the model needs voice-related features as auxiliary input (for example, referring to existing voice styles, etc.), the corresponding acoustic features need to be extracted from the voice sample library, such as fundamental frequency, intensity, timbre, and other acoustic parameters. These parameters can be obtained through signal processing techniques, for example, using short-time Fourier transform (STFT) to obtain the spectral features of the voice, or through linear predictive coding (LPC) to extract the vocal tract-related features. Thereafter, a pre-trained conditional flow matching model is loaded, and the model parameters can include the weights, biases, etc. of the neural network, which define the structure and function of the model. At the same time, the parameter initialization method related to variable step reasoning in the model is determined, for example, setting the initial step size, hyperparameters related to the step size adjustment strategy, etc.
[0137] Then variable step size is used for conditional flow matching model inference. Specifically, first, inference parameters are initialized, and an initial inference step size is determined. This step size can be set according to empirical values or statistical information in the model training process. For example, in some sequence generation tasks, the initial step size can be related to the average length of the sequence. Second, rules for step size adjustment are set. This can be determined based on performance indicators of the model (such as quality evaluation indicators of generated speech), features of the data (such as complexity of the text), or number of iterations, etc. For example, when the quality of the generated speech segment is low, the step size is appropriately reduced to process more finely; when the text content is relatively simple and the generation quality is stable, the step size is appropriately increased to improve inference speed. Third, the inference loop is started, and the selected initial step size is used to start processing the text content (such as text). For text input, the text is divided into multiple segments according to the step size (for example, if the text is a sentence, the number of words determined by the initial step size is used to divide the segments). Each segment is input into the conditional flow matching model. The model generates the corresponding speech features or the prediction of the speech segment according to the input text segment and the previously accumulated context information (such as the text content corresponding to the previously generated speech part). In addition, during the inference process, it can be evaluated whether the step size needs to be adjusted according to the step size adjustment rule. For example, if the model has a large error when generating the current speech segment (which can be measured by speech quality evaluation indicators such as speech clarity, naturalness, etc.), the step size is reduced according to the predetermined strategy, so that the text segment is processed more finely in the next round of inference. Continue this inference loop until the entire text content (such as text) is processed. Also, during the inference process, the intermediate generated speech segments or the intermediate state of the model can be saved. For example, when the generated speech needs to be partially modified or adjusted, these intermediate results can provide a convenient starting point. At the same time, the saving of the intermediate state is also helpful for the recovery of the inference process when unexpected situations (such as program interruption) occur.
[0138] Subsequently, in the speech generation stage, according to the speech features or speech segment predictions obtained by model inference, speech synthesis technology is used to convert them into actual speech. If the model directly outputs speech waveform parameters, it can be converted into playable speech signals through digital-to-analog conversion and other methods. If the model outputs the acoustic features of the speech (such as mel spectrum, etc.), a vocoder needs to be used to synthesize the speech. In the speech synthesis process, attention should be paid to the transition between speech segments to ensure the coherence of the speech. This can be achieved by smoothing the speech parameters (such as interpolation in the frequency domain) or using specialized speech splicing techniques. Further, post-processing of the generated speech is performed to improve the quality of the speech. For example, noise reduction processing is performed to remove noise or unnatural audio components that may be introduced during the synthesis process. Volume normalization can also be performed to keep the volume of the speech relatively stable throughout the audio text. In addition, according to the specific application scenario, the speech can be operated at different speeds and pitches to meet the different needs of users.
[0139] For example, in a medical scenario, medical-related text is obtained and preprocessed to extract key information such as disease description and treatment recommendations. Subsequently, according to the semantic content of the text, a suitable speech style and tone are selected. In the medical scenario, a formal, serious and easy-to-understand speech style may need to be selected to ensure that the patient or medical personnel can accurately understand the generated content. During model inference, a variable step strategy is used to adjust the step size according to the complexity and semantic changes of the text content. For medical text passages containing a large number of professional terms and complex descriptions, the model can use a smaller step size to ensure that each detail is accurately understood and generated. In the description of simple or repetitive medical text parts, the model can use a larger step size to improve inference speed and efficiency. In addition, to further improve the quality of the generated speech, some post-processing techniques can be introduced during the inference process. For example, speech synthesis post-processing algorithms can be used to smooth the generated speech waveform, reduce noise and distortion. Speech enhancement techniques can also be used to improve the clarity and intelligibility of the speech, making it more suitable for use in actual medical scenarios.
[0140] In summary, by applying a variable step strategy to the speech generation method based on the conditional flow matching model, the inference efficiency and output accuracy of the model in different application scenarios can be significantly improved. In particular, in the medical, financial and other professional fields, this method can generate more accurate, natural and easy-to-understand speech content, providing users with a better user experience.
[0141] It is also important to ensure that the input text is in the correct format for the model. For example, the text should be encoded correctly, such as using UTF-8, to avoid any issues with character encoding. Additionally, for text that contains special characters, abbreviations, slang, or other content, the model may require special preprocessing or handling strategies. For example, if the input includes additional features such as audio features as references, these features need to be normalized. For example, if the audio spectrum feature is used, the spectrum amplitude range may vary for different audio texts, and it needs to be normalized to a standard range (such as [0, 1] or [-1, 1]), so that the model can more stably calculate the weight of these features during inference. Also, make sure the model parameters are correctly loaded and updated, for example, before inference starts, make sure the model parameters are complete and correctly loaded. If the parameters are updated during inference (for example, using online learning or adaptive strategies), strict update rules should be followed. These rules are usually determined during model training and the update process needs to consider the impact on the current inference task. For example, the update step cannot be too large, otherwise it may destroy the knowledge structure that the model has learned.
[0142] Preferably, during inference, make full use of existing context information. For example, when generating speech, the speech features and semantic information of the previous word or sentence can help better generate the speech of the next part. This can be achieved by storing the previous information in the model's hidden state or other specialized memory structures. Context information also needs to be updated in a timely manner. When new input data (such as new text fragments) is processed, the model should be able to integrate this new information into the existing context and update the context representation for subsequent inference. Otherwise, there may be information lag or error accumulation. Of course, conditional flow matching models may face uncertainty during inference, such as for some ambiguous text content (with multiple possible speech expressions) or noisy audio features. At this time, the model needs to have a certain mechanism to estimate this uncertainty, such as by using a probability distribution to represent multiple possible speech generation results.
[0143] In another practical application scenario, when constructing the speech generation model according to the conditional flow matching model and the vector field, a module can be designed to fuse the conditional flow matching model and the speech vector field. This module can be a multi-layer perception (MLP) or a more complex neural network structure, such as a combination of convolutional neural network (CNN) and recurrent neural network (RNN). The purpose is to effectively fuse the text-related features (such as semantic information, grammatical structure, etc.) output by the conditional flow matching model and the speech dynamic features (such as pitch change, timbre change, etc.) contained in the vector field. For example, in the text-to-speech conversion task, the conditional flow matching model may output the speech feature prediction corresponding to each word in the text, such as the phoneme sequence and prosody information. The vector field can provide dynamic change information of the actual speech signal in the time domain or frequency domain. The fusion module splices these information and generates a comprehensive speech feature representation through a series of linear and nonlinear transformations.
[0144] After the fusion module, a speech generator module is constructed, which is responsible for converting the fused speech features into actual speech waveforms. For example, using vocoder technology. The vocoder can be divided into parameter-based vocoder and waveform-based vocoder. The parameter-based vocoder (such as Mel Linear Predictive Coding-MLPC) first converts the fused speech features into a set of acoustic parameters (such as linear prediction coefficients, fundamental frequency, etc.), and then synthesizes speech waveforms through these parameters. The waveform-based vocoder (such as WaveNet) directly generates speech waveforms from the fused speech features, which usually utilizes the powerful fitting ability of deep neural networks to learn the generation rule of speech waveforms.
[0145] The embodiment of the present application also provides a computer device, which can include a memory and a processor, the memory has a computer program, and the processor can realize the steps provided by the above-mentioned embodiments when calling the computer program in the memory. Of course, the computer device can also include various network interfaces, power supplies and other components.
[0146] Please refer to Figure 9 , Figure 9 is a schematic block diagram of a computer device provided by the embodiment of the present application. The computer device is a device with wireless communication and wired communication.
[0147] The computer device includes a processor 802, a memory and a network interface 805 connected through a system bus 801, wherein the memory can include a non-volatile storage medium 803 and an internal memory 804.
[0148] The non-volatile storage medium 803 can store an operating system 8031 and a computer program 8032. The computer program 8032, when executed, can cause the processor 802 to perform the speech generation method based on the conditional flow matching model.
[0149] The processor 802 is configured to provide computing and control capabilities to support the operation of the entire computer device.
[0150] The non-volatile storage medium 803 provides an environment for the computer program 8032 to run. The computer program 8032, when executed by the processor 802, can cause the processor 802 to perform the speech generation method based on the conditional flow matching model.
[0151] The network interface 805 is configured to perform network communication with other devices. Those skilled in the art can understand that the network interface 805 can be configured to perform network communication with other devices through the network according to the structure shown in the figure, and details are not described herein. Figure 9 It should be understood that the structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0152] The processor 802 is configured to run the computer program 8032 stored in the memory to implement any embodiment of the speech generation method based on the conditional flow matching model.
[0153] It should be understood that in the embodiments of the present application, the processor 802 can be a central processing unit (CPU), and the processor 802 can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.
[0154] In one embodiment, a computer device is provided, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the following steps when executing the computer program:
[0155] Obtaining the training text of the generated speech, and taking the generated speech as the training speech;
[0156] Sample noise data, input the training text, training speech and noise data into a conditional flow matching model for training;
[0157] Infer the conditional flow matching model by using a variable step length to obtain a vector field about noise and speech;
[0158] Construct a speech generation model based on the conditional flow matching model and the vector field, and infer and predict a target text to be generated speech by using the speech generation model.
[0159] The embodiment of the present application also provides a computer readable storage medium, which has a computer program stored thereon, and the computer program can realize the steps provided by the above embodiment when being executed. The storage medium can include a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk and various program code storage media.
[0160] In one embodiment, a computer readable storage medium is provided, which has a computer program stored thereon, and the computer program realizes the following steps when being executed by a processor:
[0161] Obtain training text of generated speech, and take the generated speech as training speech;
[0162] Sample noise data, input the training text, training speech and noise data into a conditional flow matching model for training;
[0163] Infer the conditional flow matching model by using a variable step length to obtain a vector field about noise and speech;
[0164] Construct a speech generation model based on the conditional flow matching model and the vector field, and infer and predict a target text to be generated speech by using the speech generation model.
[0165] It should be noted that the functions or steps that can be realized by the computer readable storage medium or the computer device described above can be referred to the related description of the server side and the client side in the foregoing method embodiment, and the description will not be repeated here.
[0166] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when executed, can include the processes of the above-mentioned embodiment methods. Any reference to memory, storage, database or other medium used in the embodiments provided by the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0167] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of functional units and modules is exemplified, and in actual application, the above-mentioned functions can be completed by different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the above-described functions.
[0168] The above-mentioned embodiments are only used to illustrate the technical solutions of the present application, but not limit it. Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features. The modification or replacement does not make the essence of the corresponding technical solution deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.
Claims
1. A speech generation method based on a conditional flow matching model, characterized by, The method comprises the following steps: acquire training text of generated speech, and take the generated speech as training speech; sample noise data, input the training text, training speech and noise data into a conditional flow matching model for training; infer the conditional flow matching model by using a variable step length to obtain a vector field related to noise and speech; construct a speech generation model based on the conditional flow matching model and the vector field, and infer and predict target text to be generated speech by using the speech generation model; the step of sampling noise data, inputting the training text, training speech and noise data into a conditional flow matching model for training comprises: encode the training text, and acquire text duration according to the encoding result; randomly sample a time as a time step, and input the time step, the training text, training speech, noise data and text duration into the conditional flow matching model for training; the step of inferring the conditional flow matching model by using a variable step length to obtain a vector field related to noise and speech comprises: solve the ordinary differential equation by using an explicit Euler format to construct an inference expression; the inference expression is as follows: ; where X i+1 represents the inference result of the i+1th step, X i represents the inference result of the ith step, t i+1 represents the time step of the i+1th step, t i represents the time step of the ith step, represents the training result of the ith step, represents the velocity field of the ith step; solve the inference expression by using a variable step length to obtain the result of the ordinary differential equation; the expression of the variable step length is as follows: h = a / (e i - e i-1 ); where h represents the variable step size, a represents a constant, e i represents the energy of the i-th step, e i-1 represents the energy of the i-1-th step.
2. The voice generation method based on a conditional flow matching model according to claim 1, characterized in that, the training expression of the conditional flow matching model is as follows: ; ; wherein X t represents the training result of the conditional flow matching model, X1 represents the training voice, X0 represents the noise data, t represents the time step, represents a coefficient. 3.The voice generation method based on the conditional flow matching model according to claim 2, wherein, the step of sampling noise data, inputting the training text, training speech and noise data into a conditional flow matching model for training further comprises: Setting a result X of training t a velocity field wherein μ denotes an input parameter of the conditional flow matching model; perform target training on the conditional flow matching model according to the following formula: ; wherein minL represents a target training result of the conditional flow matching model, denotes the expected distribution based on the noise data; a target training based on the conditional flow matching model to obtain an ordinary differential equation .
4. The voice generation method based on a conditional flow matching model according to claim 3, characterized in that, the step of solving the inference expression by using a variable step length to obtain the result of the ordinary differential equation comprises: acquire the square of the amplitude of the speech signal of the training result of the conditional flow matching model; map the square of the amplitude of the speech signal to obtain the energy of the training result; calculate the step length of the current step based on the energy of each step, and solve the ordinary differential equation of the current step by using the step length of the current step.
5. A speech generation apparatus based on a conditional flow matching model, characterized by, The method comprises the following steps: a text acquisition unit is configured to acquire training text of generated speech, and take the generated speech as training speech; a data input unit is configured to sample noise data, input the training text, training speech and noise data into a conditional flow matching model for training; a model inference unit is configured to infer the conditional flow matching model by using a variable step length to obtain a vector field related to noise and speech; an inference prediction unit is configured to construct a speech generation model based on the conditional flow matching model and the vector field, and infer and predict target text to be generated speech by using the speech generation model; the data input unit comprises: a text encoding unit is configured to encode the training text, and acquire text duration according to the encoding result; a sampling input unit is configured to randomly sample a time as a time step, and input the time step, the training text, training speech, noise data and text duration into the conditional flow matching model for training; the model inference unit comprises: An equation solving unit is configured to solve the ordinary differential equation in an explicit Euler format to construct an inference expression as follows: ; wherein X i+1 represents the inference result of the i+1th step, X i represents the inference result of the ith step, t i+1 represents the time step of the i+1th step, t i represents the time step of the ith step, represents the training result of the ith step, represents the velocity field of the ith step; A variable inference unit is configured to solve the inference expression in a variable step length to obtain a result of the ordinary differential equation. An expression of the variable step length is as follows: h = a / (e i - e i-1 ); where h denotes the variable step size, a denotes a constant, e i denotes the energy of the i-th step, e i-1 denotes the energy of the i-1-th step.
6. A computer device, comprising: The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the speech generation method based on the conditional flow matching model according to any one of claims 1 to 4.
7. A computer readable storage medium characterized in that, The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the speech generation method based on the conditional flow matching model according to any one of claims 1 to 4.
Citation Information
Patent Citations
Voice generation method and device based on artificial intelligence, computer equipment and medium
CN119360818A
Speech synthesis method and device with conditional matching stream, equipment and medium
CN119517005A