Voice generation method and device based on conditional flow matching model, and related components
By using the training text and noise data of generated speech in the conditional flow matching model for training and using variable step size for inference, the speech generation model is constructed, which solves the problem of insufficient accurate and reliable speech generation results and improves speech quality.
Patent Information
- Application Number
- CN202510222737.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-26
AI Technical Summary
The speech generation results based on the conditional flow matching model are not accurate and reliable enough, resulting in a relatively average voice quality.
By obtaining the training text of the generated speech, sampling noise data, training its input conditional flow matching model, and inference is performed using variable steps to build a speech generation model and inference prediction of the target text.
It improves the quality of speech generation, realizes the rapid generation of speech, and improves the accuracy and reliability of speech.
Smart Images

Figure CN119993115A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a speech generation method, device and related components based on a conditional stream matching model. Background Art
[0002] With the deepening of generative model research, speech generation models based on different schemes have been continuously proposed, including diffusion models based on stochastic differential equations and conditional flow matching (CFM) models based on ordinary differential equations. These models first showed outstanding results in the field of image generation, and were gradually applied to the field of speech generation, showing good results in speech generation quality. For example, intelligent diagnosis and treatment in medical scenarios, remote consultation, and some financial business processing platforms in financial scenarios.
[0003] Among them, the conditional stream matching model has higher reasoning efficiency than the diffusion model. However, since the conditional stream matching model was proposed not long ago, most of the focus in the field is on improving the model framework, mainly concentrated in the model training stage, and there are relatively few improvements in the reasoning stage. This has led to the current reasoning results based on the conditional stream matching model being not accurate and reliable enough, which in turn leads to the general quality of the speech generated by the conditional stream matching model. Summary of the invention
[0004] The embodiments of the present invention provide a speech generation method, apparatus, computer equipment and storage medium based on a conditional stream matching model, aiming to improve the quality of speech generation.
[0005] In a first aspect, an embodiment of the present invention provides a method for generating speech based on a conditional stream matching model, comprising:
[0006] Obtain training text of the generated speech and use the generated speech as the training speech;
[0007] Sampling noise data, and inputting the training text, training speech, and noise data into a conditional stream matching model for training;
[0008] Inferring the conditional stream matching model using a variable step size to obtain a vector field of noise and speech;
[0009] A speech generation model is constructed based on the conditional flow matching model and the vector field, and the speech generation model is used to perform inference and prediction on a target text of the speech to be generated.
[0010] In a second aspect, an embodiment of the present invention provides a speech generation device based on a conditional stream matching model, comprising:
[0011] A text acquisition unit, used to acquire a training text of the generated speech and use the generated speech as the training speech;
[0012] A data input unit, used for sampling noise data, and inputting the training text, training speech and noise data into a conditional stream matching model for training;
[0013] A model inference unit, used for inferring the conditional stream matching model using a variable step size to obtain a vector field about noise and speech;
[0014] An inference prediction unit is used to construct a speech generation model based on the conditional flow matching model and the vector field, and use the speech generation model to perform inference prediction on the target text of the speech to be generated.
[0015] In a third aspect, an embodiment of the present invention provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the speech generation method based on the conditional stream matching model as described in the first aspect is implemented.
[0016] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the speech generation method based on the conditional flow matching model as described in the first aspect is implemented.
[0017] The embodiment of the present invention provides a speech generation method, device, computer equipment and storage medium based on a conditional flow matching model, the method comprising: obtaining a training text of generated speech, and using the generated speech as training speech; sampling noise data, inputting the training text, training speech and noise data into the conditional flow matching model for training; reasoning the conditional flow matching model using a variable step length to obtain a vector field about noise and speech; constructing a speech generation model based on the conditional flow matching model and the vector field, and using the speech generation model to infer and predict the target text of the speech to be generated. The embodiment of the present invention uses the text of the generated speech to train the conditional flow matching model, thereby constructing a speech generation model, and in the process of training the conditional flow matching model, a variable step length is used to reason and solve it, so that not only can speech be generated quickly, but also the quality of speech generation can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying any creative work.
[0019] Figure 1 A schematic diagram of an application environment of a speech generation method based on a conditional stream matching model provided by an embodiment of the present invention;
[0020] Figure 2 A flow chart of a method for generating speech based on a conditional stream matching model provided by an embodiment of the present invention;
[0021] Figure 3 A schematic diagram of a sub-process of a method for generating speech based on a conditional stream matching model provided by an embodiment of the present invention;
[0022] Figure 4 Another schematic diagram of a flow chart of a method for generating speech based on a conditional stream matching model provided by an embodiment of the present invention;
[0023] Figure 5 A schematic diagram of a speech generation method based on a conditional stream matching model provided by an embodiment of the present invention;
[0024] Figure 6 A schematic block diagram of a speech generation device based on a conditional stream matching model provided by an embodiment of the present invention;
[0025] Figure 7 A sub-schematic block diagram of a speech generation device based on a conditional stream matching model provided by an embodiment of the present invention;
[0026] Figure 8 Another sub-schematic block diagram of a speech generation device based on a conditional stream matching model provided by an embodiment of the present invention;
[0027] Fig. 9 A schematic block diagram of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0028] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0029] It should be understood that when used in this specification and the appended claims, the terms "include" and "comprises" indicate the presence of described features, integers, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof.
[0030] It should also be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.
[0031] It should be further understood that the term "and / or" used in the present description and the appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0032] The speech generation method based on the conditional stream matching model provided in the embodiment of the present invention can be applied in the following aspects: Figure 1 In an application environment, the client communicates with the server through a network. The server can obtain the training text of the generated speech, and use the generated speech as the training speech; sample noise data, and input the training text, training speech and noise data into the conditional stream matching model for training; use a variable step size to infer the conditional stream matching model to obtain a vector field about noise and speech; construct a speech generation model based on the conditional stream matching model and the vector field, and use the speech generation model to infer and predict the target text of the generated speech. The embodiment of the present invention uses the text of the generated speech to train the conditional stream matching model to construct a speech generation model, and in the training process of the conditional stream matching model, a variable step size is used to infer and solve it, so that not only can the speech be generated quickly, but also the quality of speech generation can be improved. The present invention is described in detail below through specific embodiments.
[0033] See below Figure 2 , a speech generation method based on a conditional stream matching model provided by an embodiment of the present invention specifically includes: steps S101 to S104.
[0034] Step S101, obtaining a training text of a generated speech, and using the generated speech as the training speech;
[0035] Step S102, sampling noise data, inputting the training text, training speech and noise data into a conditional stream matching model for training;
[0036] Step S103, using a variable step size to infer the conditional stream matching model to obtain a vector field of noise and speech;
[0037] Step S104: constructing a speech generation model based on the conditional flow matching model and the vector field, and using the speech generation model to perform inference and prediction on a target text of the speech to be generated.
[0038] In this embodiment, a training text is first obtained, which corresponds to a speech that has been successfully generated. Noise data is then collected and input into a conditional stream matching model together with the obtained training text and speech. The conditional stream matching model is then inferred using a variable step size to obtain corresponding inference prediction results, thereby constructing a speech generation model, through which the specified target text can be inferred and predicted.
[0039] This embodiment uses the text of generated speech to train the conditional stream matching model to construct a speech generation model. During the training process of the conditional stream matching model, a variable step size is used to infer and solve it. This not only enables rapid speech generation, but also improves the quality of speech generation.
[0040] It should be noted that the prior art uses a fixed step size to reason the conditional flow matching model, which may result in the step size being too large or too small in a certain step of reasoning. When the step size is too large, some key information may be skipped when the input data (such as text) is divided and input into the conditional flow matching model. For example, in the text-to-speech conversion task, if the text is segmented according to a larger step size, a complete semantic unit (such as a phrase or a clause) may be segmented. In this way, the model may not be able to fully obtain semantic information during the reasoning process, resulting in deviations in the semantic expression of the generated speech. Take the conversion of a complex technical document to speech as an example. The document may contain professional terms and complex sentence structures. If the step size is too large, key technical words and their modifiers may be separated, so that the model cannot accurately understand the relationship between the words, thereby erroneously emphasizing or pronouncing these words during speech generation. In addition, for input data containing rich details, a larger step size will cause the model to ignore these details. In the field of speech generation, the details of the audio include timbre, intonation, and connected reading. If the step size is too large, the model may not be able to accurately process these details during inference, and the generated speech will sound stiff. For example, when imitating the voice style of a specific person, some unique details of the person's voice (such as slight vibrato, specific intonation changes, etc.) may be ignored due to the large step size. The model may only be able to roughly generate the outline of the speech, but cannot accurately reproduce these subtle speech features.
[0041] If the step size is too small, the number of steps that need to be processed during the inference process will increase significantly. This means that the model needs more computing resources and time to complete the inference. For example, when processing a long text-to-speech task, if the step size is set too small, the model may need to perform complex calculations and matches for each word or even each syllable, which will greatly increase the amount of calculation. For devices with limited resources (such as mobile devices), this situation may cause the program to run slowly or even crash. Moreover, too many calculation steps may also cause the model to overfit some local information, thereby affecting the grasp of the overall content. In addition, if the step size is too small, the model will pay too much attention to local input data, which may cause the generated speech to lack coherence. The model may overemphasize each small input unit and ignore the overall connection between them. For example, when generating a continuous narrative speech, a step size that is too small may make the model pronounce each word accurately, but at the sentence level, intonation, speaking speed, etc. may not correctly reflect the semantics and emotions of the sentence, and it will sound unnatural.
[0042] This embodiment uses variable step length for reasoning, so that each step of reasoning can have a suitable step length, which not only allows the model to effectively utilize the input information while avoiding unnecessary calculations. It can ensure that the model fully obtains important information such as semantics and speech features, and at the same time does not waste resources due to processing too many details. For example, when converting news articles into speech, a suitable step length can enable the model to reason according to the structure of sentences or phrases, which can not only accurately understand the content of the article, but also efficiently generate natural and fluent speech. It also helps the model to better capture the details and overall style of the speech. The model can reasonably adjust the rhythm, rhythm and timbre of the speech according to the input units divided by the step length. For example, when imitating the speech of a specific dialect, a suitable step length can enable the model to properly handle details such as special pronunciation and intonation changes in the dialect, thereby generating high-quality speech that conforms to the characteristics of the dialect.
[0043] For example, in a medical scenario, when obtaining medical-related documents as training texts, the corresponding noise data is sampled, and the medical-related documents and the corresponding training speech and noise data are input into the conditional stream matching model for training. Since medical-related documents contain a lot of medical professional terms and more complex medical sentences, such as disease names, symptoms and signs, and treatment methods, etc., a variable step length is used to infer the conditional stream matching model with medical-related documents as input, so that each step of reasoning can have a suitable step length. In this way, a speech generation model suitable for medical scenarios is generated, so that the speech generation model can be used to infer and predict the medical text to be generated.
[0044] For another example, in a financial scenario, when obtaining financial-related documents as training texts, the corresponding noise data is sampled, and the financial-related documents and the corresponding training speech and noise data are input into the conditional stream matching model for training. Since financial-related documents contain a lot of financial professional terms and more complex financial statements, such as capital markets, primary markets, and financial instruments, etc., a variable step length is used to infer the conditional stream matching model with financial-related documents as input, so that each step of reasoning can have a suitable step length. In this way, a speech generation model suitable for financial scenarios is generated, so that the speech generation model can be used to infer and predict the financial text to be generated.
[0045] In one embodiment, the sampling noise data, inputting the training text, training speech and noise data into the conditional stream matching model for training, comprises:
[0046] Encode the training text, and obtain the text duration according to the encoding result;
[0047] A time is randomly sampled as a time step, and the time step is input into the conditional stream matching model together with the training text, training speech, noise data and text duration for training.
[0048] Specifically, the training expression of the conditional flow matching model is:
[0049] X t =tX1+(1-(1-σ min )t)X0;
[0050]
[0051] Among them, X t represents the training result of the conditional stream matching model, X1 represents the training speech, X0 represents the noise data, t represents the time step, σ min Represents the coefficient.
[0052] Combination Figure 5 First, let's introduce the general framework of the speech generation model based on conditional flow matching (CFM). When building a speech generation model based on the conditional flow matching model, we first define an ordinary differential equation When t = 0, X t It obeys Gaussian distribution, that is, it is noise. When t=1, it is the desired data distribution, which can be understood as the Mel spectrum distribution of the speech that we want to generate.
[0053] When training the model, the text (such as medical-related text) is first encoded, and then the duration is predicted to ensure that the length of the final generated speech corresponds to the length of the text. Here, the information obtained from the text can be recorded as μ, which is passed into the conditional stream matching model. At the same time, a time t is randomly sampled in [0,1], and a Gaussian noise (corresponding to X(t=0)) is sampled, and is passed into the conditional stream matching model together with the speech corresponding to the text (corresponding to X(t=1)) for training.
[0054] With the above conditions, we can start training the model. Here we choose the path of data flowing from X(t=0) to X(t=1) as the optimal transmission condition vector field, and the expression is as follows:
[0055]
[0056] In one embodiment, if Figure 3 As shown, the sampling noise data, inputting the training text, training speech and noise data into the conditional stream matching model for training, also includes: steps S201 to S203.
[0057] Step S201: Set the training result X t The velocity field v t (X t |μ), where μ represents the input parameter of the conditional flow matching model;
[0058] Step S202: Target training is performed on the conditional flow matching model according to the following formula:
[0059]
[0060] Wherein, minL represents the target training result of the conditional flow matching model, represents the expected distribution based on noisy data;
[0061] Step S203: Obtaining an ordinary differential equation based on the target training of the conditional flow matching model
[0062] In this embodiment, a velocity field v is trained by using the conditional flow matching model. t (X t |μ), the velocity field represents the speed magnitude and direction of the noise data distribution to the real data distribution. Through this velocity field, we can approximate the above That is, the training goal is:
[0063]
[0064] After the model training is completed, an ordinary differential equation can be obtained
[0065] In one embodiment, the inference of the conditional stream matching model using a variable step size to obtain a vector field about noise and speech includes:
[0066] The ordinary differential equation is solved by using an explicit Euler format to construct an inference expression; the inference expression is as follows:
[0067]
[0068] Among them, X i+1 represents the inference result of step i+1, X i represents the inference result of step i, t i+1 represents the time step of the i+1th step, t i represents the time step of the i-th step, represents the training result of the i-th step, represents the velocity field at step i;
[0069] The inference expression is solved with a variable step size to obtain a result of the ordinary differential equation.
[0070] In this embodiment, for any t in [0,1], a v can be calculated by the conditional flow matching model. t (X t |μ), so that this ordinary differential equation can be numerically solved. Specifically, assuming that the reasoning process is divided into N steps, that is, the process is X0→X1→…→X N-1 →X N =X1, the corresponding time node is 0=t0→t1→…→t N-1 →t N =1. This embodiment uses the explicit Euler format to numerically solve the ordinary differential equation, and the calculation expression is: Recursively from t=0 to t=1.
[0071] Specifically, Figure 4 As shown, the method of solving the inference expression with a variable step size to obtain the result of the ordinary differential equation includes: steps S301 to S303.
[0072] Step S301, obtaining the square of the speech signal amplitude of the training result of the conditional stream matching model;
[0073] Step S302, obtaining the energy of the training result based on the square mapping of the speech signal amplitude;
[0074] Step S303: Calculate the step length of the current step based on the energy of each step, and solve the ordinary differential equation of the current step through the step length of the current step.
[0075] Here, the expression of the variable step size is:
[0076] h=a / (e i -e i-1 );
[0077] Wherein, h represents the variable step size, a represents a constant, and e i represents the energy of the i-th step, e i-1 represents the energy of the i-1th step.
[0078] Most existing methods adopt a fixed step size for solving, that is, h = t1-t0 = t2-t1 = ... = t N-1 -t N-2 =t N -t N-1 This embodiment adopts a variable step size. After the i-th step calculation is completed, the obtained X i Calculate the energy e i , while combining the energy e from the previous step i-1 , then the step size of the next step can be determined from this, that is, the step size expression is Here a is a constant that can be determined according to the actual situation. In simple terms, the step size of the next step is inversely proportional to the difference between the energy obtained in the previous two steps. The greater the energy change, the smaller the step size, and the smaller the energy change, the larger the step size.
[0079] By calculating the energy value of each time node, this embodiment can compare which time nodes the recursive process changes greatly near, so it is necessary to improve the calculation accuracy near these time nodes, that is, reduce the step size and reduce the error; and when the energy is stable, the step size can be appropriately increased to speed up the reasoning speed without affecting the quality of the final generated speech. Compared with the fixed step size scheme, when the same number of reasoning steps N is selected, the variable step size scheme can obtain better generation quality; if the same initial step size is selected, the variable step size scheme can select a larger step size to accelerate reasoning after the energy is stable. In particular, in addition to applying variable step sizes in the field of speech generation, variable step sizes can also be applied in any other model based on conditional flow matching or ordinary differential equations to improve the reasoning efficiency and output accuracy of the model.
[0080] In actual application scenarios, after obtaining the training text, the training text is preprocessed, for example, removing redundant spaces, punctuation marks (if special processing is required), special characters, etc. If it is a text in other formats (such as a document format with text information), the text content needs to be extracted. If the text contains information such as semantic tags, this information also needs to be parsed so that it can be better processed using the conditional stream matching model in the future. Here, if the model requires speech-related features as auxiliary input (for example, referring to existing voice styles, etc.), it is necessary to extract corresponding acoustic features from the voice sample library, such as acoustic parameters such as fundamental frequency, sound intensity, and timbre. These parameters can be obtained through signal processing technology, for example, using short-time Fourier transform (STFT) to obtain the spectral features of speech, or extracting vocal tract-related features through linear predictive coding (LPC). Thereafter, the pre-trained conditional stream matching model is loaded, and the model parameters may include weights and biases of the neural network, which define the structure and function of the model. At the same time, it is necessary to determine the parameter initialization method related to variable step size inference in the model, for example, setting the initial step size, hyperparameters related to the step size adjustment strategy, etc.
[0081] Then, the conditional stream matching model is inferred using a variable step size. Specifically, the inference parameters are first initialized to determine the initial inference step size. The step size can be set based on empirical values or statistical information during the model training process. For example, in some sequence generation tasks, the initial step size may be related to the average length of the sequence. Secondly, the step size adjustment rule is set. This can be determined based on factors such as the performance indicators of the model (such as the quality evaluation indicators of the generated speech), the characteristics of the data (such as the complexity of the text), or the number of iterations. For example, when the quality of the generated speech segment is low, the step size is appropriately reduced for more refined processing; when the text content is relatively simple and the generation quality is stable, the step size is appropriately increased to increase the inference speed. Again, the inference loop is started to process the text content (such as text) with the selected initial step size. For text input, the text is divided into multiple segments according to the step size (for example, if the text is a sentence, the segments are divided according to the number of words determined by the initial step size). Each segment is input into the conditional stream matching model. The model generates the corresponding speech features or predictions of the speech segment based on the input text segment and the previously accumulated contextual information (such as the text content corresponding to the previously generated speech part). In addition, during the reasoning process, it is possible to evaluate whether the step size needs to be adjusted based on the step size adjustment rule. For example, if the model has a large error in generating the current speech segment (which can be measured by speech quality evaluation indicators, such as speech clarity, naturalness, etc.), the step size is reduced according to a predetermined strategy so that the text segment can be processed more finely in the next round of reasoning. This reasoning cycle continues until the entire text content (such as text) is processed. In addition, during the reasoning process, the intermediate generated speech segments or the intermediate states of the model can be saved. For example, when the generated speech needs to be partially modified or adjusted, these intermediate results can provide a convenient starting point. At the same time, the preservation of the intermediate state also helps to resume the reasoning process in the event of an unexpected situation (such as a program interruption).
[0082] Subsequently, in the speech generation stage, speech synthesis technology is used to convert the speech features or speech segment predictions obtained by model inference into actual speech. If the model directly outputs speech waveform parameters, it can be converted into a playable speech signal by means of digital-to-analog conversion. If the model outputs the acoustic features of speech (such as Mel spectrum, etc.), a vocoder is required to synthesize speech. In the speech synthesis process, attention should be paid to the transition between speech segments to ensure the coherence of speech. This can be achieved by smoothing speech parameters (such as interpolation in the spectral domain) or using special speech splicing technology. Further, the generated speech is post-processed to improve the speech quality. For example, noise reduction processing is performed to remove noise or unnatural audio components that may be introduced during the synthesis process. Volume normalization can also be performed to keep the volume of the speech relatively stable throughout the audio text. In addition, according to the specific application scenario, the speech can also be speed-changed, pitch-changed, etc. to meet the different needs of users.
[0083] For example, in a medical scenario, medical-related text is obtained and preprocessed to extract key information, such as disease descriptions, treatment recommendations, etc. Subsequently, the appropriate voice style and tone are selected based on the semantic content of the text. In a medical scenario, a formal, serious, and easy-to-understand voice style may need to be selected to ensure that patients or medical personnel can accurately understand the generated content. During the model reasoning process, a variable step size strategy is used to adjust the step size according to the complexity and semantic changes of the text content. For medical text paragraphs containing a large number of professional terms and complex descriptions, the model can use a smaller step size to ensure that every detail can be accurately understood and generated. For medical text parts with simpler or repetitive descriptions, the model can use a larger step size to improve the reasoning speed and efficiency. In addition, in order to further improve the quality of the generated speech, some post-processing techniques can be introduced in the reasoning process. For example, a speech synthesis post-processing algorithm can be used to smooth the generated speech waveform and reduce noise and distortion. Speech enhancement technology can also be used to improve the clarity and intelligibility of speech, making it more suitable for use in actual medical scenarios.
[0084] In general, by applying the variable step size strategy to the speech generation method based on the conditional stream matching model, the model's reasoning efficiency and output accuracy in different application scenarios can be significantly improved. Especially in professional fields such as medicine and finance, this method can generate more accurate, natural and easy-to-understand speech content, providing users with a better user experience.
[0085] It should also be noted that when using the conditional stream matching model to perform text-to-speech conversion generation, it is necessary to ensure that the format of the input text meets the requirements of the model. For example, the encoding method of the text (such as UTF-8, etc.) should be correct to avoid garbled characters. At the same time, for content containing special characters, abbreviations, slang, etc., the model may require special preprocessing or corresponding processing strategies. For example, feature normalization, that is, when the input contains other auxiliary features (such as audio features as references), these features need to be normalized. Taking audio spectrum features as an example, the spectrum amplitude range of different audio texts may be different, and it needs to be normalized to a standard range (such as [0,1] or [-1,1]), so that the model can calculate the weights of these features more stably during the inference process. In addition, it is necessary to ensure that the model parameters are correctly loaded and updated. For example, before the inference starts, it is necessary to ensure that the model parameters are complete and correctly loaded. If the parameter update is involved in the inference process (for example, using online learning or adaptive strategies), it must be strictly followed according to the established update rules. These rules are usually determined during the model training phase, and the update process needs to take into account the impact on the current inference task. For example, the update step size cannot be too large, so as not to destroy the knowledge structure that the model has learned.
[0086] Preferably, in the reasoning process, make full use of the existing context information. For example, when generating speech, the speech features and semantic information of the previous word or sentence can help better generate the speech of the next part. This can be achieved by storing the previous information in the hidden state of the model or other special memory structures. Context information also needs to be updated in time. When new input data (such as new text fragments) is processed, the model should be able to integrate this new information into the existing context and update the context representation for subsequent reasoning. Otherwise, information lag or error accumulation may occur. Of course, the conditional stream matching model may face uncertainty during the reasoning process, for example, for some ambiguous text content (with multiple possible speech expressions) or audio features interfered by noise. At this time, the model needs to have a certain mechanism to estimate this uncertainty, such as representing multiple possible speech generation results through probability distribution.
[0087] In another practical application scenario, when a speech generation model is constructed based on the conditional stream matching model and the vector field, a module can be first designed to fuse the conditional stream matching model and the speech vector field. This module can be a multi-layer perceptron (MLP) or a more complex neural network structure, such as a combination of a convolutional neural network (CNN) and a recurrent neural network (RNN). Its purpose is to effectively fuse the text-related features (such as semantic information, grammatical structure, etc.) output by the conditional stream matching model and the speech dynamic features (such as pitch change, timbre change, etc.) contained in the vector field. For example, in the text-to-speech conversion task, the conditional stream matching model may output the speech feature prediction corresponding to each word in the text, such as phoneme sequence and prosody information. The vector field can provide dynamic change information of the actual speech signal in the time domain or frequency domain. The fusion module splices this information and generates a comprehensive speech feature representation through a series of linear and nonlinear transformations.
[0088] After the fusion module, a speech generator module is constructed, which is responsible for converting the fused speech features into actual speech waveforms. For example, vocoder technology is used. Vocoders can be divided into parameter-based vocoders and waveform-based vocoders. Parameter-based vocoders (such as Mel Linear Prediction Coding - MLPC) first convert the fused speech features into a set of acoustic parameters (such as linear prediction coefficients, fundamental frequency, etc.), and then synthesize speech waveforms through these parameters. Waveform-based vocoders (such as WaveNet) directly generate speech waveforms from the fused speech features. They usually use the powerful fitting ability of deep neural networks to learn the generation rules of speech waveforms.
[0089] Figure 6 A schematic block diagram of a speech generation device 500 based on a conditional stream matching model provided by an embodiment of the present invention, the device 500 includes:
[0090] A text acquisition unit 501 is used to acquire a training text of the generated speech and use the generated speech as the training speech;
[0091] A data input unit 502 is used to sample noise data and input the training text, training speech and noise data into a conditional stream matching model for training;
[0092] A model inference unit 503, configured to use a variable step size to infer the conditional stream matching model to obtain a vector field about noise and speech;
[0093] The inference prediction unit 504 is used to construct a speech generation model based on the conditional flow matching model and the vector field, and use the speech generation model to perform inference prediction on the target text of the speech to be generated.
[0094] In this embodiment, a training text is first obtained, which corresponds to a speech that has been successfully generated. Noise data is then collected and input into a conditional stream matching model together with the obtained training text and speech. The conditional stream matching model is then inferred using a variable step size to obtain corresponding inference prediction results, thereby constructing a speech generation model, through which the specified target text can be inferred and predicted.
[0095] This embodiment uses the text of generated speech to train the conditional stream matching model to construct a speech generation model. During the training process of the conditional stream matching model, a variable step size is used to infer and solve it. This not only enables rapid speech generation, but also improves the quality of speech generation.
[0096] It should be noted that the prior art uses a fixed step size to reason the conditional flow matching model, which may result in the step size being too large or too small in a certain step of reasoning. When the step size is too large, some key information may be skipped when the input data (such as text) is divided and input into the conditional flow matching model. For example, in the text-to-speech conversion task, if the text is segmented according to a larger step size, a complete semantic unit (such as a phrase or a clause) may be segmented. In this way, the model may not be able to fully obtain semantic information during the reasoning process, resulting in deviations in the semantic expression of the generated speech. Take the conversion of a complex technical document to speech as an example. The document may contain professional terms and complex sentence structures. If the step size is too large, key technical words and their modifiers may be separated, so that the model cannot accurately understand the relationship between the words, thereby erroneously emphasizing or pronouncing these words during speech generation. In addition, for input data containing rich details, a larger step size will cause the model to ignore these details. In the field of speech generation, the details of the audio include timbre, intonation, and connected reading. If the step size is too large, the model may not be able to accurately process these details during inference, and the generated speech will sound stiff. For example, when imitating the voice style of a specific person, some unique details of the person's voice (such as slight vibrato, specific intonation changes, etc.) may be ignored due to the large step size. The model may only be able to roughly generate the outline of the speech, but cannot accurately reproduce these subtle speech features.
[0097] If the step size is too small, the number of steps that need to be processed during the inference process will increase significantly. This means that the model needs more computing resources and time to complete the inference. For example, when processing a long text-to-speech task, if the step size is set too small, the model may need to perform complex calculations and matches for each word or even each syllable, which will greatly increase the amount of calculation. For devices with limited resources (such as mobile devices), this situation may cause the program to run slowly or even crash. Moreover, too many calculation steps may also cause the model to overfit some local information, thereby affecting the grasp of the overall content. In addition, if the step size is too small, the model will pay too much attention to local input data, which may cause the generated speech to lack coherence. The model may overemphasize each small input unit and ignore the overall connection between them. For example, when generating a continuous narrative speech, a step size that is too small may make the model pronounce each word accurately, but at the sentence level, intonation, speaking speed, etc. may not correctly reflect the semantics and emotions of the sentence, and it will sound unnatural.
[0098] This embodiment uses variable step length for reasoning, so that each step of reasoning can have a suitable step length, which not only allows the model to effectively utilize the input information while avoiding unnecessary calculations. It can ensure that the model fully obtains important information such as semantics and speech features, and at the same time does not waste resources due to processing too many details. For example, when converting news articles into speech, a suitable step length can enable the model to reason according to the structure of sentences or phrases, which can not only accurately understand the content of the article, but also efficiently generate natural and fluent speech. It also helps the model to better capture the details and overall style of the speech. The model can reasonably adjust the rhythm, rhythm and timbre of the speech according to the input units divided by the step length. For example, when imitating the speech of a specific dialect, a suitable step length can enable the model to properly handle details such as special pronunciation and intonation changes in the dialect, thereby generating high-quality speech that conforms to the characteristics of the dialect.
[0099] For example, in a medical scenario, when obtaining medical-related documents as training texts, the corresponding noise data is sampled, and the medical-related documents and the corresponding training speech and noise data are input into the conditional stream matching model for training. Since medical-related documents contain a lot of medical professional terms and more complex medical sentences, such as disease names, symptoms and signs, and treatment methods, etc., a variable step length is used to infer the conditional stream matching model with medical-related documents as input, so that each step of reasoning can have a suitable step length. In this way, a speech generation model suitable for medical scenarios is generated, so that the speech generation model can be used to infer and predict the medical text to be generated.
[0100] For another example, in a financial scenario, when obtaining financial-related documents as training texts, the corresponding noise data is sampled, and the financial-related documents and the corresponding training speech and noise data are input into the conditional stream matching model for training. Since financial-related documents contain a lot of financial professional terms and more complex financial statements, such as capital markets, primary markets, and financial instruments, etc., a variable step length is used to infer the conditional stream matching model with financial-related documents as input, so that each step of reasoning can have a suitable step length. In this way, a speech generation model suitable for financial scenarios is generated, so that the speech generation model can be used to infer and predict the financial text to be generated.
[0101] In one embodiment, the data input unit 502 includes:
[0102] A text encoding unit, used to encode the training text and obtain the text duration according to the encoding result;
[0103] The sampling input unit is used to randomly sample a time as a time step, and input the time step together with the training text, training speech, noise data and text duration into the conditional stream matching model for training.
[0104] In one embodiment, the training expression of the conditional flow matching model is:
[0105]
[0106]
[0107] Among them, X t represents the training result of the conditional stream matching model, X1 represents the training speech, X0 represents the noise data, t represents the time step, σ min Represents the coefficient.
[0108] Combination Figure 5 First, let's introduce the general framework of the speech generation model based on conditional flow matching (CFM). When building a speech generation model based on the conditional flow matching model, we first define an ordinary differential equation When t = 0, X t It obeys Gaussian distribution, that is, it is noise. When t=1, it is the desired data distribution, which can be understood as the Mel spectrum distribution of the speech that we want to generate.
[0109] When training the model, the text (such as medical-related text) is first encoded, and then the duration is predicted to ensure that the length of the final generated speech corresponds to the length of the text. Here, the information obtained from the text can be recorded as μ, which is passed into the conditional stream matching model. At the same time, a time t is randomly sampled in [0,1], and a Gaussian noise (corresponding to X(t=0)) is sampled, and is passed into the conditional stream matching model together with the speech corresponding to the text (corresponding to X(t=1)) for training.
[0110] With the above conditions, we can start training the model. Here we choose the path of data flowing from X(t=0) to X(t=1) as the optimal transmission condition vector field, and the expression is as follows:
[0111]
[0112] In one embodiment, if Figure 7 As shown, the data input unit 502 also includes:
[0113] The velocity field setting unit 601 is used to set the velocity field of the training result X. t The velocity field v t (X t |μ), where μ represents the input parameter of the conditional flow matching model;
[0114] The target training unit 602 is used to perform target training on the conditional flow matching model according to the following formula:
[0115]
[0116] Wherein, minL represents the target training result of the conditional flow matching model, represents the expected distribution based on noisy data;
[0117] The equation acquisition unit 603 is used to obtain the ordinary differential equation based on the target training of the conditional flow matching model.
[0118] In this embodiment, a velocity field v is trained by using the conditional flow matching model. t (X t |μ), the velocity field represents the speed magnitude and direction of the noise data distribution to the real data distribution. Through this velocity field, we can approximate the above That is, the training goal is:
[0119]
[0120] After the model training is completed, an ordinary differential equation can be obtained
[0121] In one embodiment, the model reasoning unit 503 includes:
[0122] The equation solving unit is used to solve the ordinary differential equation in an explicit Euler format to construct an inference expression; the inference expression is as follows:
[0123]
[0124] Among them, X i+1 represents the inference result of step i+1, X i represents the inference result of step i, t i+1 represents the time step of the i+1th step, t i represents the time step of the i-th step, represents the training result of the i-th step, represents the velocity field at step i;
[0125] A variable reasoning unit is used to solve the reasoning expression with a variable step size to obtain a result of the ordinary differential equation.
[0126] In this embodiment, for any t in [0,1], a v can be calculated by the conditional flow matching model. t (X t |μ), so that this ordinary differential equation can be numerically solved. Specifically, assuming that the reasoning process is divided into N steps, that is, the process is X0→X1→…→X N-1 →X N =X1, the corresponding time node is 0=t0→t1→…→t N-1 →t N =1. This embodiment uses the explicit Euler format to numerically solve the ordinary differential equation, and the calculation expression is: Recursively from t=0 to t=1.
[0127] In one embodiment, if Figure 8 As shown, the variable reasoning unit includes:
[0128] A signal acquisition unit 701 is used to acquire the square of the speech signal amplitude of the training result of the conditional stream matching model;
[0129] An energy mapping unit 702, configured to obtain the energy of the training result based on the square mapping of the speech signal amplitude;
[0130] The step length calculation unit 703 is used to calculate the step length of the current step based on the energy of each step, and solve the ordinary differential equation of the current step through the step length of the current step.
[0131] In one embodiment, the expression of the variable step size is:
[0132] h=a / (e i -e i-1 );
[0133] Wherein, h represents the variable step size, a represents a constant, and e i represents the energy of the i-th step, e i-1 represents the energy of the i-1th step.
[0134] Most existing methods adopt a fixed step size for solving, that is, h = t1-t0 = t2-t1 = ... = t N-1 -t N-2 =t N -t N-1 This embodiment adopts a variable step size. After the i-th step calculation is completed, the obtained X i Calculate the energy e i , while combining the energy e from the previous step i-1 , then the step size of the next step can be determined from this, that is, the step size expression is Here a is a constant that can be determined according to the actual situation. In simple terms, the step size of the next step is inversely proportional to the difference between the energy obtained in the previous two steps. The greater the energy change, the smaller the step size, and the smaller the energy change, the larger the step size.
[0135] By calculating the energy value of each time node, this embodiment can compare which time nodes the recursive process changes greatly near, so it is necessary to improve the calculation accuracy near these time nodes, that is, reduce the step size and reduce the error; and when the energy is stable, the step size can be appropriately increased to speed up the reasoning speed without affecting the quality of the final generated speech. Compared with the fixed step size scheme, when the same number of reasoning steps N is selected, the variable step size scheme can obtain better generation quality; if the same initial step size is selected, the variable step size scheme can select a larger step size to accelerate reasoning after the energy is stable. In particular, in addition to applying variable step sizes in the field of speech generation, variable step sizes can also be applied in any other model based on conditional flow matching or ordinary differential equations to improve the reasoning efficiency and output accuracy of the model.
[0136] In actual application scenarios, after obtaining the training text, the training text is preprocessed, for example, removing redundant spaces, punctuation marks (if special processing is required), special characters, etc. If it is a text in other formats (such as a document format with text information), the text content needs to be extracted. If the text contains information such as semantic tags, this information also needs to be parsed so that it can be better processed using the conditional stream matching model in the future. Here, if the model requires speech-related features as auxiliary input (for example, referring to existing voice styles, etc.), it is necessary to extract corresponding acoustic features from the voice sample library, such as acoustic parameters such as fundamental frequency, sound intensity, and timbre. These parameters can be obtained through signal processing technology, for example, using short-time Fourier transform (STFT) to obtain the spectral features of speech, or extracting vocal tract-related features through linear predictive coding (LPC). Thereafter, the pre-trained conditional stream matching model is loaded, and the model parameters may include weights and biases of the neural network, which define the structure and function of the model. At the same time, it is necessary to determine the parameter initialization method related to variable step size inference in the model, for example, setting the initial step size, hyperparameters related to the step size adjustment strategy, etc.
[0137] Then, the conditional stream matching model is inferred using a variable step size. Specifically, the inference parameters are first initialized to determine the initial inference step size. The step size can be set based on empirical values or statistical information during the model training process. For example, in some sequence generation tasks, the initial step size may be related to the average length of the sequence. Secondly, the step size adjustment rule is set. This can be determined based on factors such as the performance indicators of the model (such as the quality evaluation indicators of the generated speech), the characteristics of the data (such as the complexity of the text), or the number of iterations. For example, when the quality of the generated speech segment is low, the step size is appropriately reduced for more refined processing; when the text content is relatively simple and the generation quality is stable, the step size is appropriately increased to increase the inference speed. Again, the inference loop is started to process the text content (such as text) with the selected initial step size. For text input, the text is divided into multiple segments according to the step size (for example, if the text is a sentence, the segments are divided according to the number of words determined by the initial step size). Each segment is input into the conditional stream matching model. The model generates the corresponding speech features or predictions of the speech segment based on the input text segment and the previously accumulated contextual information (such as the text content corresponding to the previously generated speech part). In addition, during the reasoning process, it is possible to evaluate whether the step size needs to be adjusted based on the step size adjustment rule. For example, if the model has a large error in generating the current speech segment (which can be measured by speech quality evaluation indicators, such as speech clarity, naturalness, etc.), the step size is reduced according to a predetermined strategy so that the text segment can be processed more finely in the next round of reasoning. This reasoning cycle continues until the entire text content (such as text) is processed. In addition, during the reasoning process, the intermediate generated speech segments or the intermediate states of the model can be saved. For example, when the generated speech needs to be partially modified or adjusted, these intermediate results can provide a convenient starting point. At the same time, the preservation of the intermediate state also helps to resume the reasoning process in the event of an unexpected situation (such as a program interruption).
[0138] Subsequently, in the speech generation stage, speech synthesis technology is used to convert the speech features or speech segment predictions obtained by model inference into actual speech. If the model directly outputs speech waveform parameters, it can be converted into a playable speech signal by means of digital-to-analog conversion. If the model outputs the acoustic features of speech (such as Mel spectrum, etc.), a vocoder is required to synthesize speech. In the speech synthesis process, attention should be paid to the transition between speech segments to ensure the coherence of speech. This can be achieved by smoothing speech parameters (such as interpolation in the spectral domain) or using special speech splicing technology. Further, the generated speech is post-processed to improve the speech quality. For example, noise reduction processing is performed to remove noise or unnatural audio components that may be introduced during the synthesis process. Volume normalization can also be performed to keep the volume of the speech relatively stable throughout the audio text. In addition, according to the specific application scenario, the speech can also be speed-changed, pitch-changed, etc. to meet the different needs of users.
[0139] For example, in a medical scenario, medical-related text is obtained and preprocessed to extract key information, such as disease descriptions, treatment recommendations, etc. Subsequently, the appropriate voice style and tone are selected based on the semantic content of the text. In a medical scenario, a formal, serious, and easy-to-understand voice style may need to be selected to ensure that patients or medical personnel can accurately understand the generated content. During the model reasoning process, a variable step size strategy is used to adjust the step size according to the complexity and semantic changes of the text content. For medical text paragraphs containing a large number of professional terms and complex descriptions, the model can use a smaller step size to ensure that every detail can be accurately understood and generated. For medical text parts with simpler or repetitive descriptions, the model can use a larger step size to improve the reasoning speed and efficiency. In addition, in order to further improve the quality of the generated speech, some post-processing techniques can be introduced in the reasoning process. For example, a speech synthesis post-processing algorithm can be used to smooth the generated speech waveform and reduce noise and distortion. Speech enhancement technology can also be used to improve the clarity and intelligibility of speech, making it more suitable for use in actual medical scenarios.
[0140] In general, by applying the variable step size strategy to the speech generation method based on the conditional stream matching model, the model's reasoning efficiency and output accuracy in different application scenarios can be significantly improved. Especially in professional fields such as medicine and finance, this method can generate more accurate, natural and easy-to-understand speech content, providing users with a better user experience.
[0141] It should also be noted that when using the conditional stream matching model to perform text-to-speech conversion generation, it is necessary to ensure that the format of the input text meets the requirements of the model. For example, the encoding method of the text (such as UTF-8, etc.) should be correct to avoid garbled characters. At the same time, for content containing special characters, abbreviations, slang, etc., the model may require special preprocessing or corresponding processing strategies. For example, feature normalization, that is, when the input contains other auxiliary features (such as audio features as references), these features need to be normalized. Taking audio spectrum features as an example, the spectrum amplitude range of different audio texts may be different, and it needs to be normalized to a standard range (such as [0,1] or [-1,1]), so that the model can calculate the weights of these features more stably during the inference process. In addition, it is necessary to ensure that the model parameters are correctly loaded and updated. For example, before the inference starts, it is necessary to ensure that the model parameters are complete and correctly loaded. If the parameter update is involved in the inference process (for example, using online learning or adaptive strategies), it must be strictly followed according to the established update rules. These rules are usually determined during the model training phase, and the update process needs to take into account the impact on the current inference task. For example, the update step size cannot be too large, so as not to destroy the knowledge structure that the model has learned.
[0142] Preferably, in the reasoning process, make full use of the existing context information. For example, when generating speech, the speech features and semantic information of the previous word or sentence can help better generate the speech of the next part. This can be achieved by storing the previous information in the hidden state of the model or other special memory structures. Context information also needs to be updated in time. When new input data (such as new text fragments) is processed, the model should be able to integrate this new information into the existing context and update the context representation for subsequent reasoning. Otherwise, information lag or error accumulation may occur. Of course, the conditional stream matching model may face uncertainty during the reasoning process, for example, for some ambiguous text content (with multiple possible speech expressions) or audio features interfered by noise. At this time, the model needs to have a certain mechanism to estimate this uncertainty, such as representing multiple possible speech generation results through probability distribution.
[0143] In another practical application scenario, when a speech generation model is constructed based on the conditional stream matching model and the vector field, a module can be first designed to fuse the conditional stream matching model and the speech vector field. This module can be a multi-layer perceptron (MLP) or a more complex neural network structure, such as a combination of a convolutional neural network (CNN) and a recurrent neural network (RNN). Its purpose is to effectively fuse the text-related features (such as semantic information, grammatical structure, etc.) output by the conditional stream matching model and the speech dynamic features (such as pitch change, timbre change, etc.) contained in the vector field. For example, in the text-to-speech conversion task, the conditional stream matching model may output the speech feature prediction corresponding to each word in the text, such as phoneme sequence and prosody information. The vector field can provide dynamic change information of the actual speech signal in the time domain or frequency domain. The fusion module splices this information and generates a comprehensive speech feature representation through a series of linear and nonlinear transformations.
[0144] After the fusion module, a speech generator module is constructed, which is responsible for converting the fused speech features into actual speech waveforms. For example, vocoder technology is used. Vocoders can be divided into parameter-based vocoders and waveform-based vocoders. Parameter-based vocoders (such as Mel Linear Prediction Coding - MLPC) first convert the fused speech features into a set of acoustic parameters (such as linear prediction coefficients, fundamental frequency, etc.), and then synthesize speech waveforms through these parameters. Waveform-based vocoders (such as WaveNet) directly generate speech waveforms from the fused speech features. They usually use the powerful fitting ability of deep neural networks to learn the generation rules of speech waveforms.
[0145] The embodiment of the present invention also provides a computer device, which may include a memory and a processor, wherein a computer program is stored in the memory, and when the processor calls the computer program in the memory, the steps provided in the above embodiment may be implemented. Of course, the computer device may also include various network interfaces, power supplies and other components.
[0146] See also Fig. 9 , Fig. 9 The present invention provides a schematic block diagram of a computer device provided in an embodiment of the present invention. The computer device is a device with wireless communication and wired communication.
[0147] The computer device includes a processor 802 , a memory, and a network interface 805 connected via a system bus 801 , wherein the memory may include a non-volatile storage medium 803 and an internal memory 804 .
[0148] The non-volatile storage medium 803 may store an operating system 8031 and a computer program 8032. When the computer program 8032 is executed, the processor 802 may execute a speech generation method based on a conditional stream matching model.
[0149] The processor 802 is used to provide computing and control capabilities to support the operation of the entire computer device.
[0150] The internal memory 804 provides an environment for the operation of the computer program 8032 in the non-volatile storage medium 803. When the computer program 8032 is executed by the processor 802, the processor 802 can execute a speech generation method based on a conditional stream matching model.
[0151] The network interface 805 is used to communicate with other devices over the network. Fig. 9 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present invention, and does not constitute a limitation on the computer device to which the solution of the present invention is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0152] The processor 802 is configured to run a computer program 8032 stored in the memory to implement any embodiment of the above-mentioned method for generating speech based on the conditional stream matching model.
[0153] It should be understood that in the embodiment of the present invention, the processor 802 may be a central processing unit (CPU), and the processor 802 may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0154] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the following steps when executing the computer program:
[0155] Obtain training text of the generated speech and use the generated speech as the training speech;
[0156] Sampling noise data, and inputting the training text, training speech, and noise data into a conditional stream matching model for training;
[0157] Inferring the conditional stream matching model using a variable step size to obtain a vector field of noise and speech;
[0158] A speech generation model is constructed based on the conditional flow matching model and the vector field, and the speech generation model is used to perform inference and prediction on a target text of the speech to be generated.
[0159] The embodiment of the present invention further provides a computer-readable storage medium on which a computer program is stored, and when the computer program is executed, the steps provided in the above embodiment can be implemented. The storage medium may include: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program codes.
[0160] In one embodiment, a computer readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:
[0161] Obtain training text of the generated speech and use the generated speech as the training speech;
[0162] Sampling noise data, and inputting the training text, training speech, and noise data into a conditional stream matching model for training;
[0163] Inferring the conditional stream matching model using a variable step size to obtain a vector field of noise and speech;
[0164] A speech generation model is constructed based on the conditional flow matching model and the vector field, and the speech generation model is used to perform inference and prediction on a target text of the speech to be generated.
[0165] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can refer to the relevant descriptions on the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.
[0166] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0167] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0168] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. A speech generation method based on a conditional stream matching model, characterized in that: include: Obtain training text of the generated speech and use the generated speech as the training speech; Sampling noise data, and inputting the training text, training speech, and noise data into a conditional stream matching model for training; Inferring the conditional stream matching model using a variable step size to obtain a vector field of noise and speech; A speech generation model is constructed based on the conditional flow matching model and the vector field, and the speech generation model is used to perform inference and prediction on a target text of the speech to be generated.
2. The speech generation method based on the conditional stream matching model according to claim 1, characterized in that: The sampling noise data, inputting the training text, training speech and noise data into the conditional stream matching model for training, comprises: Encode the training text, and obtain the text duration according to the encoding result; A time is randomly sampled as a time step, and the time step is input into the conditional stream matching model together with the training text, training speech, noise data and text duration for training.
3. The speech generation method based on the conditional stream matching model according to claim 2, characterized in that: The training expression of the conditional flow matching model is: X t =tX1+(1-(1-σ min )t)X0; Among them, X t represents the training result of the conditional stream matching model, X1 represents the training speech, X0 represents the noise data, t represents the time step, σ min Represents the coefficient.
4. The speech generation method based on the conditional stream matching model according to claim 3 is characterized in that: The sampling noise data, inputting the training text, training speech and noise data into the conditional stream matching model for training, further comprising: Set about training result X t The velocity field v t (X t |μ), where μ represents the input parameter of the conditional flow matching model; The conditional flow matching model is trained according to the following formula: Wherein, minL represents the target training result of the conditional flow matching model, represents the expected distribution based on noisy data; Based on the target training of the conditional flow matching model, the ordinary differential equation is obtained 5. The speech generation method based on the conditional stream matching model according to claim 4, characterized in that: The method of using a variable step size to infer the conditional stream matching model to obtain a vector field of noise and speech includes: The ordinary differential equation is solved by using an explicit Euler format to construct an inference expression; the inference expression is as follows: Among them, X i+1 represents the inference result of step i+1, X i represents the inference result of step i, t i+1 represents the time step of the i+1th step, t i represents the time step of the i-th step, represents the training result of the i-th step, represents the velocity field at step i; The inference expression is solved with a variable step size to obtain a result of the ordinary differential equation.
6. The speech generation method based on the conditional stream matching model according to claim 5, characterized in that: The adopting a variable step size to solve the inference expression to obtain the result of the ordinary differential equation includes: Obtaining the square of the speech signal amplitude of the training result of the conditional stream matching model; Obtaining energy of the training result based on the square mapping of the speech signal amplitude; The step length of the current step is obtained based on the energy calculation of each step, and the ordinary differential equation of the current step is solved by the step length of the current step.
7. The method for generating speech based on a conditional stream matching model according to claim 6, characterized in that: The expression of the variable step size is: h=a / (e i -e i-1 ); Wherein, h represents the variable step size, a represents a constant, and e i represents the energy of the i-th step, e i-1 represents the energy of the i-1th step.
8. A speech generation device based on a conditional stream matching model, characterized in that: include: A text acquisition unit, used to acquire a training text of the generated speech and use the generated speech as the training speech; A data input unit, used for sampling noise data, and inputting the training text, training speech and noise data into a conditional stream matching model for training; A model inference unit, used for inferring the conditional stream matching model using a variable step size to obtain a vector field about noise and speech; An inference prediction unit is used to construct a speech generation model based on the conditional flow matching model and the vector field, and use the speech generation model to perform inference prediction on the target text of the speech to be generated.
9. A computer device, characterized in that: It comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the speech generation method based on the conditional stream matching model as described in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for generating speech based on a conditional flow matching model according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Voice generation method and device based on artificial intelligence, computer equipment and medium
CN119360818A
Speech synthesis method and device with conditional matching stream, equipment and medium
CN119517005A
Signal-adaptive Remixing of Separated Audio Sources
US20230395079A1
Cited By
Training method, reasoning method and related device of stream matching generative model
CN120597951A
Motor control parameter data generation method based on conditional flow matching
CN121098201A
Speech generation method and apparatus based on conditional flow matching model, and related component
WO2026179344A1