Model training method and device

By introducing an intermediate model into the variational autoencoder and optimizing the model parameters of the encoder and decoder, the generated feature encoding satisfies the non-standard normal distribution, which solves the problems of long training time and low generated audio quality and achieves more efficient and higher-quality audio generation.

CN120596918AActive Publication Date: 2025-09-05SHANGHAI XIYU JIZHI TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510690121.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-09-05
Estimated Expiration
2045-05-27

AI Technical Summary

Technical Problem

In the existing technology, variational autoencoders have problems in training large language models, such as long training time, many training steps, and low generated audio quality, which makes it difficult to improve the quality of synthesized audio.

Method used

By introducing an intermediate model into the variational autoencoder and optimizing the model parameters of the encoder and decoder, the generated feature encoding satisfies the preset method of non-standard normal distribution, and the intermediate model is added to improve the encoding and decoding effect.

Benefits of technology

The training efficiency and quality of generated content are improved, the generated audio has more details and higher stability, the constraints on the standard normal distribution are reduced, and the difficulty of model training is simplified.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120596918A_ABST
    Figure CN120596918A_ABST
Patent Text Reader

Abstract

The invention discloses a model training method and device. The method comprises the following steps: acquiring a first training file, inputting the first training file into an encoding model in a variational auto-encoder, and encoding the first training file through the encoding model to generate a first feature code; generating a second feature code based on the first feature code and a first model or a first function in the intermediate model; wherein the probability distribution of the second feature codes meets a first preset distribution mode; generating a first generation file corresponding to the first training file based on the second feature code, a second model or a second function in the intermediate model and a decoding model in the variational auto-encoder; and comparing the first training file with the first generation file, and adjusting model parameters of the variational auto-encoder and / or the intermediate model based on a comparison result. According to the scheme, the encoding and decoding effects of the target file can be improved by optimizing the variational auto-encoder model and adding the intermediate model, so that the quality of the generated target content is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of model training technology, and in particular to a model training method and device. Background Art

[0002] In the field of large language model technology, in order to enable the large language model to learn training data, the training data needs to be encoded. How to encode is not only directly related to the learning and training effect and training efficiency of the large model, but also to the quality of the generated content.

[0003] For example, in the field of speech synthesis technology, in order to enable a large language model to learn and train speech data, the audio data needs to be compressed and encoded, and then the compressed and encoded audio data is input into the large language model for training. Related technologies generally use the encoder in the variational auto-encoder (VAE) to downsample the audio file several times. For example, for a 44.1KHz audio file, every 4410 sampling points is compressed into one code, which is equivalent to compressing every 0.1 second of audio into one audio feature code. The audio feature code is then input into the large language model for training, and the large language model outputs the predicted audio feature code. The decoder in the variational autoencoder (VAE) then upsamples the predicted audio feature code several times, for example, decoding one predicted audio feature code to generate an audio waveform of 4410 sampling points.

[0004] The encoder and decoder in a variational autoencoder (VAE) are two reversible models trained simultaneously. While VAEs have a simple structure, achieving good training results requires a long training time and multiple training steps. Furthermore, the probability distribution in the VAE model uses a standard normal distribution to fit real-world audio features. As a result, the audio quality generated by the trained decoder is generally poor, with a low upper limit, making it difficult to improve the quality of synthesized audio. Summary of the Invention

[0005] The present invention provides a model training method and device, which can improve the encoding and decoding effects of target files by optimizing the variational autoencoder model and adding intermediate models, thereby improving the quality of generated target content.

[0006] According to one aspect of the present invention, a model training method is provided, comprising:

[0007] Obtaining a first training file, inputting the first training file into an encoding model in a variational autoencoder, encoding the first training file using the encoding model to generate a first feature code; wherein a probability distribution of the first feature code satisfies a non-first preset distribution mode;

[0008] Generate a second feature code based on the first feature code and the first model or the first function in the intermediate model; wherein the probability distribution of the second feature code satisfies the first preset distribution mode;

[0009] Generate a first generated file corresponding to the first training file based on the second feature code, the second model or the second function in the intermediate model, and the decoding model in the variational autoencoder; wherein the second model or the second function in the intermediate model performs a process opposite to that of the first model or the first function;

[0010] The first training file and the first generated file are compared, and model parameters of the variational autoencoder and / or the intermediate model are adjusted based on the comparison result.

[0011] According to another aspect of the present invention, there is provided a model training device, comprising:

[0012] A first feature code generation module is configured to obtain a first training file, input the first training file into an encoding model in a variational autoencoder, encode the first training file using the encoding model, and generate a first feature code; wherein the probability distribution of the first feature code satisfies a non-first preset distribution mode;

[0013] a second feature code generating module, configured to generate a second feature code based on the first feature code and the first model or the first function in the intermediate model; wherein the probability distribution of the second feature code satisfies the first preset distribution mode;

[0014] a first generated file generating module, configured to generate a first generated file corresponding to the first training file based on the second feature code, the second model or the second function in the intermediate model, and the decoding model in the variational autoencoder; wherein the second model or the second function in the intermediate model performs a process opposite to that of the first model or the first function;

[0015] A model parameter adjustment module is used to compare the first training file and the first generated file, and adjust the model parameters of the variational autoencoder and / or the intermediate model based on the comparison result.

[0016] According to another aspect of the present invention, an electronic device is provided, comprising:

[0017] at least one processor; and

[0018] a memory communicatively connected to the at least one processor; wherein,

[0019] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the model training method described in any embodiment of the present invention.

[0020] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the model training method described in any embodiment of the present invention when executed.

[0021] The model training scheme of the embodiment of the present invention obtains a first training file, inputs the first training file into the encoding model in the variational autoencoder, encodes the first training file through the encoding model to generate a first feature code; wherein the probability distribution of the first feature code satisfies a non-first preset distribution mode; based on the first feature code and the first model or first function in the intermediate model, a second feature code is generated; wherein the probability distribution of the second feature code satisfies the first preset distribution mode; based on the second feature code, the second model or second function in the intermediate model and the decoding model in the variational autoencoder, a first generated file corresponding to the first training file is generated; wherein the second model or second function in the intermediate model performs the opposite process of the first model or first function; the first training file and the first generated file are compared, and the model parameters of the variational autoencoder and / or the intermediate model are adjusted based on the comparison result. Through the technical solution provided by the embodiment of the present invention, by optimizing the variational autoencoder model and adding the intermediate model, the encoding and decoding effects of the target file can be improved, thereby improving the quality of the generated target content. At the same time, compared with the existing technology, there is no need to constrain the file feature encoding to the standard normal distribution or other first preset distribution methods, but only needs to be constrained to the non-standard normal distribution (that is, other distribution methods other than the standard normal distribution) or other non-first preset distribution methods, which greatly reduces the difficulty of model training and improves the training efficiency; at the same time, since the improved variational autoencoder encoding has fewer constraints, the generated target content can be closer to the real data. At the same time, adding an intermediate model can make the feature information contained in the generated feature encoding richer, the generated target content has more details, and the quality and stability of the target content generation are greatly improved.

[0022] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0024] Figure 1 A flowchart of a model training method provided by an embodiment of the present invention;

[0025] Figure 2 A schematic structural diagram of a model training device provided by an embodiment of the present invention;

[0026] Figure 3 A schematic diagram of the structure of an electronic device for implementing the model training method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0027] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0028] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0029] Figure 1 This is a flow chart of a model training method provided by an embodiment of the present invention. This embodiment is applicable to the case of training a model. The method can be executed by a model training device. The model training device can be implemented in the form of hardware and / or software. The model training device can be configured in an electronic device. Figure 1 As shown, the method includes:

[0030] S110. Obtain a first training file, input the first training file into a coding model in a variational autoencoder, encode the first training file through the coding model, and generate a first feature code; wherein the probability distribution of the first feature code satisfies a non-first preset distribution mode.

[0031] In an embodiment of the present invention, a first training file is obtained, wherein the first training file includes at least one of text, audio, video, and picture types. For example, the first training file is an audio training file, a picture training file, or a video training file. For another example, the first training file is a combination file of an audio training file and a text training file. For another example, the first training file is a combination file of an audio training file, a video training file, and a text training file. It should be noted that the embodiment of the present invention does not limit the file type of the first training file. In addition, the first training file can be 100 or 1000, and the embodiment of the present invention does not limit the number of first training files.

[0032] The first training file is input into the encoding model of the variational autoencoder, and the first training file is encoded by the encoding model of the variational autoencoder to generate a first feature code; wherein the probability distribution of the first feature code satisfies a non-first preset distribution mode, that is, the probability distribution of the first feature code satisfies other preset distribution modes or any distribution mode other than the first preset distribution mode. Optionally, the probability distribution of the first feature code satisfies a second preset distribution mode, and the second preset distribution mode has fewer constraints than the first preset distribution mode. It can be understood that the second preset distribution mode is a non-first preset distribution mode, and the second preset distribution mode has fewer constraints than the first preset distribution mode. For example, the first preset distribution mode can be a standard normal distribution, that is, a standard Gaussian distribution, and the second preset distribution mode can be a more complex and diverse normal distribution, that is, a Gaussian distribution, or other predefined distribution modes. Optionally, the first preset distribution mode is a standard normal distribution, and the second preset distribution mode is a normal distribution. It can be understood that the second preset distribution method is not the first preset distribution method because the second preset distribution method has fewer constraints than the first preset distribution method, and it does not mean that the second preset distribution method has no intersection with the first preset distribution method. Continuing with the example, the first preset distribution method is the standard normal distribution, and the second preset distribution method is the normal distribution. It can be seen that the standard normal distribution of the first preset distribution method is a special case of the second preset distribution method-normal distribution with a mean of 0 and a standard deviation of 1.

[0033] When downsampling a file, the encoding model in a traditional variational autoencoder constrains the probability distribution of the output feature encoding to conform to a standard normal distribution (or standard Gaussian distribution). That is, the encoding model uses a standard normal distribution function or a standard Gaussian distribution function to fit the probability distribution of the feature encoding corresponding to the file and calculates the loss value until it converges to a preset range. In an embodiment of the present invention, when downsampling a first training file based on the encoding model in the variational autoencoder, the probability distribution of the output first feature encoding is constrained to conform to a distribution other than the first preset distribution.

[0034] S120. Generate a second feature code based on the first feature code and the first model or the first function in the intermediate model; wherein the probability distribution of the second feature code satisfies the first preset distribution mode.

[0035] In an embodiment of the present invention, a first feature code is input into a first model or a first function in an intermediate model to obtain a second feature code output by the first model or the first function, wherein the probability distribution of the second feature code satisfies a first preset distribution mode. It is understandable that the first feature code whose probability distribution satisfies a distribution mode other than the first preset distribution mode is converted into a second feature code whose probability distribution satisfies the first preset distribution mode by the first model or the first function in the intermediate model. For example, the first model or the first function in the intermediate model is a model or function for converting a feature code whose probability distribution satisfies a normal distribution or other predefined distribution mode into a feature code whose probability distribution satisfies a standard normal distribution.

[0036] S130. Based on the second feature code, the second model or second function in the intermediate model and the decoding model in the variational autoencoder, generate a first generated file corresponding to the first training file; wherein the second model or second function in the intermediate model performs a process opposite to that of the first model or first function.

[0037] Exemplarily, the second feature code is input into the second model or second function in the intermediate model, so that the second model or second function performs the opposite process of the first model or first function on the second feature code, obtains the feature code output by the second model or second function (such as the target feature code), and inputs the feature code output by the second model or second function (i.e., the target feature code) into the decoding model in the variational autoencoder, so that the decoding model in the variational autoencoder decodes the target feature code to generate a first generated file corresponding to the first training file. The first generated file is a prediction file obtained by the variational autoencoder and the intermediate model. It can be understood that the second feature code whose probability distribution satisfies the first preset distribution mode is converted into the target feature code whose probability distribution satisfies the non-first preset distribution mode by the second model or second function in the intermediate model. For example, the second model or second function in the intermediate model is a model or function for converting a feature code whose probability distribution satisfies the standard normal distribution into a feature code whose probability distribution satisfies a more complex and diverse normal distribution or other predefined distribution mode. The output code of the second model or second function in the intermediate model is the input code of the decoding model in the variational autoencoder. The second model or second function in the intermediate model can make the probability distribution of the feature encoding of the decoding model in the input variational autoencoder satisfy a more complex and diverse normal distribution or a feature encoding of other predefined distribution modes, so that the file features contained in the feature encoding of the input decoding model are richer. Exemplarily, if the first training file is an audio training file, the audio feature encoding whose probability distribution satisfies the first preset distribution mode is converted into a target audio feature encoding that satisfies a more complex and diverse normal distribution or other predefined distribution mode through the second model or second function in the intermediate model, so that the target audio feature encoding input to the decoder can contain richer audio features, so that the audio waveform generated by the decoder after decoding has more detailed information and higher quality.

[0038] Optionally, the second model or second function can be a Flow Matching model. The Flow Matching model is a reversible probability density transformation model that gradually transforms a simple distribution into a complex target distribution through a series of reversible transformation functions. This process can be regarded as an iterative process of a series of variable replacements, and each replacement follows the variable transformation principle of the probability density function. The Flow Matching model can achieve accurate mapping from simple distribution to complex distribution.

[0039] S140: Compare the first training file and the first generated file, and adjust model parameters of the variational autoencoder and / or the intermediate model based on the comparison result.

[0040] In an embodiment of the present invention, the first training file is compared with the first generated file, and the model parameters of the variational autoencoder (including the encoder and / or decoder) and / or the intermediate model (including the first model or the first function and / or the second model or the second function) are adjusted based on the comparison results. Through S110-S140, the loss value is continuously calculated based on the comparison results of the first training file and the first generated file, and S110-S140 are continuously iterated and looped until the loss value converges to a preset range, thereby completing the training of the variational autoencoder and the intermediate model.

[0041] The model training method of the embodiment of the present invention obtains a first training file, inputs the first training file into the encoding model in the variational autoencoder, encodes the first training file through the encoding model, and generates a first feature code; wherein the probability distribution of the first feature code satisfies a non-first preset distribution mode; based on the first feature code and the first model or first function in the intermediate model, generates a second feature code; wherein the probability distribution of the second feature code satisfies the first preset distribution mode; based on the second feature code, the second model or second function in the intermediate model and the decoding model in the variational autoencoder, generates a first generated file corresponding to the first training file; wherein the second model or second function in the intermediate model performs the opposite process of the first model or first function; compares the first training file and the first generated file, and adjusts the model parameters of the variational autoencoder and / or the intermediate model based on the comparison result. Through the technical solution provided by the embodiment of the present invention, by optimizing the variational autoencoder model and adding the intermediate model, the encoding and decoding effects of the target file can be improved, thereby improving the quality of the generated target content. At the same time, compared with the existing technology, there is no need to constrain the file feature encoding to the standard normal distribution or other first preset distribution methods, but only needs to be constrained to the non-standard normal distribution (that is, other distribution methods other than the standard normal distribution) or other non-first preset distribution methods, which greatly reduces the difficulty of model training and improves the training efficiency; at the same time, since the improved variational autoencoder encoding has fewer constraints, the generated target content can be closer to the real data. At the same time, adding an intermediate model can make the feature information contained in the generated feature encoding richer, the generated target content has more details, and the quality and stability of the target content generation are greatly improved.

[0042] In some embodiments, after the variational autoencoder and the intermediate model are trained through the above embodiments, it also includes: inputting the first generative code output by the first language model into the trained second model or second function, and outputting the second generative code; wherein, the first language model is a large language model that has completed pre-training; inputting the second generative code into the trained decoding model, and outputting the target generated content; wherein, the second model or second function and the decoding model are trained by the model training method provided by the above embodiments.

[0043] In an embodiment of the present invention, a file to be synthesized is obtained and input into a first language model to obtain a first generated code output by the first language model, wherein the first language model is a pre-trained large language model. The first generated code output by the first language model is then input into a second model or a second function in an intermediate model trained using the model training method provided in the above embodiment to obtain a second generated code output by the second model or the second function. The second generated code is input into a decoding model in a variational autoencoder trained using the model training method provided in the above embodiment, so that the decoding model decodes the second generated code to generate target generated content. It is understood that the probability distribution of the output code of the large language model generally satisfies a standard normal distribution. The first generated code output by the first language model is input into a second model or a second function, and the second model or the second function converts the first generated code, whose probability distribution satisfies a first preset distribution such as a standard normal distribution, into a second generated code, whose probability distribution satisfies a more complex and diverse normal distribution or other predefined distribution. This allows the feature code input into the decoder to contain richer feature information, thereby allowing the target generated content generated after decoding by the decoding model to have more detailed information and higher quality.

[0044] In an embodiment of the present invention, speech synthesis is performed using a first speech model as an example for explanation. When performing speech synthesis, the encoding model in the improved variational autoencoder that has completed training can be used to extract audio feature encoding from pre-trained speech data, and input into a large language model for pre-training, so that the probability distribution of the predicted audio feature encoding output by the large language model conforms to the normal distribution constraint in the improved variational autoencoder or other predefined distribution methods, so that the decoding model in the improved variational autoencoder that has completed training can be used to correctly decode and generate the corresponding audio waveform, but the cost of retraining the large language model is very high and time-consuming. Therefore, the optimization and improvement process of the speech synthesis process is as follows: the predicted audio feature encoding (that is, the first generated encoding) output by the pre-trained large language model is input into the trained second model or second function, and the probability distribution of the audio feature encoding (that is, the second generated encoding) converted by the second model or second function satisfies a more complex and diverse normal distribution or other predefined distribution methods, and then is decoded into audio by the decoding model in the improved variational autoencoder that has completed training.

[0045] Optionally, the first generated code output by the first language model is streamed into the second model or second function. Optionally, the second generated code is streamed into the trained decoding model to output the target generated content. Exemplarily, taking speech synthesis using the first speech model as an example, the text content to be synthesized is streamed into the trained first language model, outputting the first generated code, and then the first generated code is streamed into the trained second model or second function. For example, every 100 text contents are obtained as a streaming input package and input into the first language model to obtain the first generated code. The first generated code corresponding to the 100 texts is then streamed into the second model or second function without waiting for the generated codes corresponding to all 1000 text contents to be obtained and then input into the second model or second function. The second model or second function sequentially outputs the second generated code corresponding to the streaming input package. The second generated code is streamed into the trained decoding model to output the target generated content. Through the above method, the time interval for obtaining the target generated content can be greatly shortened, and the generation efficiency of the target generated content can be improved.

[0046] Optionally, the first generated code output by the first language model is a probability distribution of a hidden layer of the first language model. It is understandable that the first generated code output by the first language model is a probability distribution of a hidden layer in the first language model, not the existing last token, and the last token is sampled based on the probability distribution of the hidden layer.

[0047] In some embodiments, a second training file is obtained, and the second training file is input into the trained encoding model to generate a third feature code; wherein the third feature code satisfies a non-first preset distribution mode; based on the third feature code, a second language model is input to generate a third generated code; wherein the second language model is a large language model that needs to be pre-trained; the third generated code is input into the trained decoding model to generate a second generated file; the second training file and the second generated file are compared, and the model parameters of the second language model are adjusted based on the comparison result.

[0048] Obtain a second training file, wherein the second training file may include at least one of text, audio, video, and picture types. For example, the second training file is an audio training file, a picture training file, or a video training file. For another example, the second training file is a combination file of an audio training file and a text training file. For another example, the second training file is a combination file of an audio training file, a video training file, and a text training file. It should be noted that the embodiment of the present invention does not limit the file type of the second training file. In addition, the second training file can be 100 or 1000, and the embodiment of the present invention does not limit the number of second training files.

[0049] The second training file is input into the encoding model of the variational autoencoder trained by the model training method provided in the above embodiment, and the second training file is encoded by the encoding model to generate a third feature code; wherein the probability distribution of the third feature code satisfies a non-first preset distribution mode, that is, the probability distribution of the third feature code satisfies other preset distribution modes or arbitrary distribution modes other than the first preset distribution mode. The third feature code is input into the second language model that needs to be pre-trained (that is, the untrained large language model) to generate a third generated code, and the third generated code output by the second language model is input into the decoding model trained by the model training method provided in the above embodiment to generate a second generated file. It can be understood that the third generated code output by the second language model is a code whose probability distribution satisfies a non-first preset distribution mode, so that when the third generated code is input into the trained decoding model, it is effectively guaranteed that the decoding model can correctly decode the third generated code. The second training file and the second generated file are compared, and the model parameters of the second language model are adjusted based on the comparison results. The loss value is continuously calculated based on the comparison results of the second training file and the second generated file, and the above steps are continuously iterated and looped until the loss value of the second language model converges to a preset range, thereby completing the training of the second language model.

[0050] In some embodiments, the fourth generated code output by the trained second language model is input into a trained decoding model to output target generated content. This arrangement has the advantage of eliminating the need for the second model or second function in the intermediate model to convert the output code of the second language model. The probability distribution of the output code of the second language model can directly satisfy a non-first predetermined distribution mode, thereby effectively ensuring that the trained decoding model can correctly decode the output code of the second language model when decoding the output code of the second language model.

[0051] Obtain a file to be synthesized, and input the file to be synthesized into a trained second language model, obtain a fourth generated code output by the trained second language model, wherein the probability distribution of the fourth generated code satisfies a non-first preset distribution mode. Then, input the fourth generated code output by the trained second language model into the decoding model of the variational autoencoder trained by the model training method provided in the above embodiment, so that the decoding model decodes the fourth generated code to generate target generated content. It is understandable that the trained second language model can make its output code directly a fourth generated code whose probability distribution satisfies a more complex and diverse normal distribution or other predefined distribution mode, so that the code input into the decoder can contain richer feature information, so that the target generated content generated after decoding by the decoding model has more detailed information and higher quality.

[0052] In an embodiment of the present invention, speech synthesis using a trained second speech model is used as an example for explanation. When performing speech synthesis, the encoding model in the trained improved variational autoencoder can be used to extract audio feature encoding of the text data to be synthesized. Then, the audio feature encoding output by the encoding model is input into the trained second language model. The fourth generated encoding whose probability distribution directly satisfies the normal distribution constraint or other predefined distribution method is input into the trained decoding model, thereby using the decoding model in the trained improved variational autoencoder to correctly decode and generate the corresponding audio waveform.

[0053] Optionally, the fourth generated code output by the second language model is input into the decoding model in a streaming manner. Exemplarily, taking speech synthesis through the first speech model as an example, the text content to be synthesized is input into the trained second language model in a streaming manner, and the fourth generated code is output. The fourth generated code is then input into the trained decoding model in a streaming manner. For example, every 100 text contents are input as a streaming input package into the second language model to obtain the fourth generated code, and then the fourth generated code corresponding to the 100 characters is streamed into the decoding model without waiting for the generated codes corresponding to all 1,000 text contents to be obtained and then all input into the decoding model, and the decoding model sequentially outputs the fourth generated code corresponding to the streaming input package. Through the above method, the time interval for obtaining the target generated content can be greatly shortened, and the generation efficiency of the target generated content can be improved.

[0054] Optionally, the fourth generated code output by the second language model is a probability distribution of a hidden layer of the second language model. It is understandable that the fourth generated code output by the second language model is a probability distribution of a hidden layer in the second language model, not the existing last token, and the last token is sampled based on the hidden layer probability distribution.

[0055] Figure 2 This is a schematic diagram of the structure of a model training device provided by an embodiment of the present invention. Figure 2 As shown, the device includes:

[0056] A first feature code generating module 210 is configured to obtain a first training file, input the first training file into an encoding model in a variational autoencoder, encode the first training file using the encoding model, and generate a first feature code; wherein the probability distribution of the first feature code satisfies a non-first preset distribution mode;

[0057] A second feature code generating module 220 is configured to generate a second feature code based on the first feature code and the first model or the first function in the intermediate model; wherein the probability distribution of the second feature code satisfies the first preset distribution mode;

[0058] A first generated file generating module 230 is configured to generate a first generated file corresponding to the first training file based on the second feature code, the second model or the second function in the intermediate model, and the decoding model in the variational autoencoder; wherein the second model or the second function in the intermediate model performs a process opposite to that of the first model or the first function;

[0059] The model parameter adjustment module 240 is used to compare the first training file and the first generated file, and adjust the model parameters of the variational autoencoder and / or the intermediate model based on the comparison result.

[0060] Optionally, also include:

[0061] A second generated code output module, configured to input the first generated code output by the first language model into the trained second model or second function, and output a second generated code; wherein the first language model is a pre-trained large language model;

[0062] The first target generated content output module is used to input the second generated code into the trained decoding model and output the target generated content.

[0063] Optionally, also include:

[0064] A third feature code generation module is configured to obtain a second training file, input the second training file into the trained coding model, and generate a third feature code; wherein the third feature code satisfies a distribution method other than the first preset distribution method;

[0065] A third generative code generation module, configured to generate a third generative code based on the third feature code input into a second language model; wherein the second language model is a large language model that needs to be pre-trained;

[0066] A second generated file generating module is used to input the third generated code into the trained decoding model to generate a second generated file;

[0067] The second language model training module is used to compare the second training file with the second generated file and adjust the model parameters of the second language model based on the comparison result.

[0068] Optionally, also include:

[0069] The first target generated content output module is configured to input the fourth generated code output by the trained second language model into the trained decoding model to output target generated content.

[0070] Optionally, the probability distribution of the first feature code satisfies a second preset distribution mode, and the second preset distribution mode has fewer constraints than the first preset distribution mode.

[0071] Optionally, the first preset distribution mode is a standard normal distribution, and the second preset distribution mode is a normal distribution.

[0072] Optionally, the first training file includes at least one of text, audio, video, and picture types.

[0073] Optionally, the second model or second function includes a Flow Matching model.

[0074] Optionally, the output encoding of the language model is input into the second model or second function, or the decoding model in a streaming manner.

[0075] Optionally, the generation of the language model output is encoded as a probability distribution of a hidden layer of the language model.

[0076] The model training device provided in the embodiment of the present invention can execute the model training method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0077] Figure 3 A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0078] like Figure 3 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0079] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0080] The processor 11 may be any general-purpose and / or specialized processing component with processing and computing capabilities. Examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors for running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the model training method.

[0081] In particular, according to an embodiment of the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product that includes a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication unit 19, or installed from the storage unit 18, or installed from the ROM 12. When the computer program is executed by the processor 11, the above-mentioned functions defined in the method of the embodiment of the present invention are performed.

[0082] In some embodiments, the model training method may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the model training method described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to perform the model training method in any other appropriate manner (e.g., by means of firmware).

[0083] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0084] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0085] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0086] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0087] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0088] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.

[0089] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.

[0090] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.

Claims

1. A model training method, characterized in that: include: Obtaining a first training file, inputting the first training file into an encoding model in a variational autoencoder, encoding the first training file using the encoding model to generate a first feature code; wherein a probability distribution of the first feature code satisfies a non-first preset distribution mode; Generate a second feature code based on the first feature code and the first model or the first function in the intermediate model; wherein the probability distribution of the second feature code satisfies the first preset distribution mode; Generate a first generated file corresponding to the first training file based on the second feature code, the second model or the second function in the intermediate model, and the decoding model in the variational autoencoder; wherein the second model or the second function in the intermediate model performs a process opposite to that of the first model or the first function; The first training file and the first generated file are compared, and model parameters of the variational autoencoder and / or the intermediate model are adjusted based on the comparison result.

2. The method according to claim 1, characterized in that Also includes: Inputting the first generated code output by the first language model into the trained second model or second function, and outputting a second generated code; wherein the first language model is a large language model that has completed pre-training; The second generated code is input into the trained decoding model to output the target generated content.

3. The method according to claim 1, characterized in that Also includes: Obtaining a second training file, inputting the second training file into the trained encoding model to generate a third feature code; wherein the third feature code satisfies a distribution method other than the first preset distribution method; Inputting a second language model based on the third feature code to generate a third generated code; wherein the second language model is a large language model that needs to be pre-trained; Inputting the third generated code into the trained decoding model to generate a second generated file; The second training file and the second generated file are compared, and model parameters of the second language model are adjusted based on the comparison result.

4. The method according to claim 3, characterized in that Also includes: The fourth generated code output by the trained second language model is input into the trained decoding model to output the target generated content.

5. The method according to claim 1, wherein The probability distribution of the first feature code satisfies a second preset distribution mode, and the second preset distribution mode has fewer constraints than the first preset distribution mode.

6. The method according to claim 5, characterized in that The first preset distribution mode is a standard normal distribution, and the second preset distribution mode is a normal distribution.

7. The method according to claim 1, characterized in that The first training file includes at least one of text, audio, video, and picture types.

8. The method according to claim 1, characterized in that The second model or second function comprises a FlowMatching model.

9. The method according to claim 2 or 4, characterized in that The output encoding of the language model is input to the second model or second function, or the decoding model in a streaming manner.

10. The method according to claim 2 or 4, characterized in that The generation of the language model output is encoded as a probability distribution of the hidden layer of the language model.

11. A model training device, characterized in that: include: A first feature code generation module is configured to obtain a first training file, input the first training file into an encoding model in a variational autoencoder, encode the first training file using the encoding model, and generate a first feature code; wherein the probability distribution of the first feature code satisfies a non-first preset distribution mode; a second feature code generating module, configured to generate a second feature code based on the first feature code and the first model or the first function in the intermediate model; wherein the probability distribution of the second feature code satisfies the first preset distribution mode; a first generated file generating module, configured to generate a first generated file corresponding to the first training file based on the second feature code, the second model or the second function in the intermediate model, and the decoding model in the variational autoencoder; wherein the second model or the second function in the intermediate model performs a process opposite to that of the first model or the first function; A model parameter adjustment module is used to compare the first training file and the first generated file, and adjust the model parameters of the variational autoencoder and / or the intermediate model based on the comparison result.

Citation Information

Patent Citations

  • Interactive knowledge defining and processing method, system and device and readable medium

    CN112861515A

  • Training potentially score-based generative model

    CN115526222A

  • Model training method, audio generation method, computer equipment and storage medium

    CN118098268A

  • Large language model training and reasoning method and device

    CN119312863A

  • Voice generation method and device based on artificial intelligence, computer equipment and medium

    CN119360818A