Model training method, data generation method, image generation method, device, storage medium and program product

By training a second generative model using knowledge distillation and guided scale feature fusion, the problem of high inference cost in generative models is solved, achieving efficient generation while maintaining generation quality.

CN120996111APending Publication Date: 2025-11-21HANGZHOU ALIBABA INT INTERNET IND CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511021880.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-23
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing generative models have shortcomings in terms of generation efficiency, especially when using CFG technology, which requires multiple forward inferences, resulting in high inference costs and affecting generation efficiency.

Method used

We employ a knowledge distillation approach and a feature fusion method based on guided scales. The first generative model is used as the teacher model to train the second generative model. By reducing the number of inferences through feature fusion, only one forward inference is required. Combined with knowledge distillation, the knowledge of the first generative model is transferred to the second generative model.

Benefits of technology

It reduces inference costs, improves generation efficiency, and at the same time ensures generation quality and model performance, thus balancing generation quality and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120996111A_ABST
    Figure CN120996111A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a model training method, a data generation method, an image generation method, equipment, a storage medium and a program product. Generating first prediction data based on the conditional sample features and generating second prediction data based on the unconditional sample features by using a first generation model, and fusing the first prediction data and the second prediction data according to a sample guide scale to obtain first target prediction data; fusing the conditional sample features and the unconditional sample features according to a sample guide scale to obtain sample fusion features; generating second target prediction data based on the sample fusion features by using a second generation model; training a second generation model according to difference information between the second target prediction data and the first target prediction data; and the second generation model generates the target fusion feature according to the condition feature and the unconditional feature determined by the target condition data and the target guide scale, and generates the target data based on the target fusion feature, thereby ensuring the generation quality and the generation efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The embodiment of the present application relates to the technical field of artificial intelligence, and particularly relates to a model training method, a data generation method, an image generation method, a computing device, a computer readable storage medium and a computer program product. BACKGROUND

[0002] The generation model refers to a generative artificial intelligence model for generating target data such as text, images or audio based on conditional data such as text, images or audio. At present, CFG (Classifier-Free Guidance) is the main technology for ensuring high-quality generation of the generation model. The generation model using the CFG technology performs forward reasoning (Forward Propagation) on conditional data and unconditioned data respectively, fuses different reasoning results to obtain the final generated target data, and controls the fusion strength of the conditional data and the unconditioned data through a guide scale, so as to realize accurate control of the generated data.

[0003] As can be seen from the above description, the generation model using the CFG technology needs to perform multiple forward reasoning, which will obviously affect the generation efficiency of the generation model. SUMMARY

[0004] The embodiment of the present application provides a model training method, a data generation method, an image generation method, a computing device, a computer readable storage medium and a computer program product to solve the technical problem of low generation efficiency in the prior art.

[0005] In a first aspect, the embodiment of the present application provides a model training method, comprising:

[0006] determining target training data; the target training data comprises conditional sample data;

[0007] determining conditional sample features and unconditioned sample features according to the conditional sample data;

[0008] generating first predicted data based on the conditional sample features and second predicted data based on the unconditioned sample features by using a first generation model, and fusing the first predicted data and the second predicted data according to a sample guide scale to obtain first target predicted data;

[0009] fusing the conditional sample features and the unconditioned sample features according to the sample guide scale to obtain sample fusion features;

[0010] generating second target predicted data based on the sample fusion features by using a second generation model;

[0011] According to difference information of the second target prediction data and the first target prediction data, the second generation model is trained; the first generation model is used as a teacher model of the second generation model; and the second generation model is used for generating target fusion features according to conditional features and unconditional features determined according to target condition data and target guide scales, and generating target data based on the target fusion features.

[0012] In a second aspect, the embodiments of the present application provide a data generation method, comprising:

[0013] Obtaining target condition data and target guide scales;

[0014] According to the target condition data, conditional features and unconditional features are determined;

[0015] According to the target guide scales, the conditional features and the unconditional features are fused to obtain target fusion features;

[0016] Target data is generated based on the target fusion features by using a second generation model; wherein the second generation model takes a first generation model as a teacher model, and is trained according to difference information of first target prediction data and second target prediction data; the first target prediction data is obtained by fusing first prediction data and second prediction data according to sample guide scales; the first prediction data is generated by the first generation model based on conditional sample features determined by conditional sample data in training data; the second prediction data is generated by the first generation model based on unconditional sample features; the second target prediction data is generated by the second generation model based on sample fusion features; and the sample fusion features are obtained by fusing the conditional sample features and the unconditional sample features according to the sample guide scales.

[0017] In a third aspect, the embodiments of the present application provide a model training method, comprising:

[0018] Obtaining target training data; the target training data comprises conditional sample data and image sample data corresponding to the conditional sample data;

[0019] Determining a time step, sample guide scales and sample noise data corresponding to the target training data;

[0020] According to the conditional sample data, conditional sample features and unconditional sample features are determined;

[0021] According to the sample noise data, the image sample data and the time step, noise image data is generated;

[0022] generating first predicted data based on the conditional sample feature, the time step and the noisy image data, and generating second predicted data based on the unconditioned sample feature, the time step and the noisy image data, and fusing the first predicted data and the second predicted data according to a sample guidance scale to obtain first noisy predicted data;

[0023] fusing the conditional sample feature and the unconditioned sample feature according to the sample guidance scale to obtain sample fused feature;

[0024] generating second noisy predicted data based on the sample fused feature, the time step and the noisy image data by using a second image generation model;

[0025] training the second image generation model according to difference information between the second noisy predicted data and the first noisy predicted data; the second image generation model is used to generate target fused feature according to conditional feature and unconditioned feature determined by target conditional data and target guidance scale, and generate target image based on the target fused feature, time sequence and target noisy data.

[0026] In a fourth aspect, an image generation method is provided, including:

[0027] obtaining target conditional data and target guidance scale;

[0028] determining conditional feature and unconditioned feature according to the target conditional data;

[0029] randomly generating target noisy data and determining time sequence feature;

[0030] fusing the conditional feature and the unconditioned feature according to the target guidance scale to obtain target fused feature;

[0031] generate a target image based on the target fusion feature, the time sequence feature, and target noise data, wherein the second image generation model takes the first image generation model as a teacher model, is trained according to difference information between first noise prediction data and second noise prediction data, and is obtained; the first noise prediction data is obtained by fusing first prediction data and second prediction data according to a sample guidance scale; the first prediction data is generated by the first image generation model based on conditional sample feature, a time step, and noise image data in training data; the second prediction data is generated by the first image generation model based on unconditional sample feature, the time step, and the noise image data; the second noise prediction data is generated by the second generation model based on sample fusion feature, the time step, and the noise image data; and the sample fusion feature is obtained by fusing the conditional sample feature and the unconditional sample feature according to the sample guidance scale.

[0032] In a fifth aspect, an embodiment of the present application provides a computing device, including a processing component and a storage component.

[0033] The storage component stores a computer program; the computer program is used to be called and executed by the processing component to implement the model training method in the first aspect, the data generation method in the second aspect, the model training method in the third aspect, or the image generation method in the fourth aspect.

[0034] In a sixth aspect, an embodiment of the present application provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processing component to implement the model training method in the first aspect, the data generation method in the second aspect, the model training method in the third aspect, or the image generation method in the fourth aspect.

[0035] In a seventh aspect, an embodiment of the present application provides a computer program product, including a computer program or instructions, and the computer program or instructions are executed by a processing component to implement the model training method in the first aspect, the data generation method in the second aspect, the model training method in the third aspect, or the image generation method in the fourth aspect.

[0036] In the embodiment of the present application, target training data is determined; the target training data includes conditional sample data; conditional sample features and unconditional sample features are determined according to the conditional sample data; a first prediction data is generated based on the conditional sample features and a second prediction data is generated based on the unconditional sample features by using a first generation model, and the first prediction data and the second prediction data are fused according to a sample guidance scale to obtain first target prediction data; the conditional sample features and the unconditional sample features are fused according to the sample guidance scale to obtain sample fusion features; a second target prediction data is generated based on the sample fusion features by using a second generation model; the second generation model is trained according to difference information between the second target prediction data and the first target prediction data; the first generation model serves as a teacher model of the second generation model; and the second generation model is used to generate target fusion features according to conditional features and unconditional features determined based on target conditional data and a target guidance scale, and generate target data based on the target fusion features.

[0037] By using the knowledge distillation method and the feature fusion method based on the guidance scale for model training, the inference times of the second generation model obtained by training are reduced, the inference cost is reduced, the generation efficiency is improved, the knowledge of the first generation model can be refined into the second generation model through the knowledge distillation, and the generation quality of the second generation model can be ensured, and the generation quality and the inference cost are considered in the embodiment of the present application, so that the model performance is ensured.

[0038] These aspects or other aspects of the present application will be more apparent in the following description of the embodiments. BRIEF DESCRIPTION OF DRAWINGS

[0039] The accompanying drawings, which are included to provide a further understanding of the present application, constitute a part of the present application and illustrate embodiments of the present application and its description, which serve to explain the present application, and do not constitute improper limitations on the present application. In the drawings:

[0040] Figure 1 A flowchart of one embodiment of a model training method provided by the present application is shown;

[0041] Figure 2 A flowchart of one embodiment of a data generation method provided by the present application is shown;

[0042] Figure 3 A flowchart of another embodiment of a model training method provided by the present application is shown;

[0043] Figure 4 A model training schematic diagram of an actual application of the embodiment of the present application is shown;

[0044] Figure 5A flowchart of one embodiment of an image generation method provided by the present application is shown.

[0045] Figure 6 A system application schematic diagram in one practical application of an embodiment of the present application is shown.

[0046] Figure 7 A structural schematic diagram of one embodiment of a model training apparatus provided by the present application is shown.

[0047] Figure 8 A structural schematic diagram of one embodiment of a data generation apparatus provided by the present application is shown.

[0048] Figure 9 A structural schematic diagram of one embodiment of a computing device provided by the present application is shown. DETAILED DESCRIPTION

[0049] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be described below in detail with specific embodiments of the present application and corresponding drawings. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.

[0050] It should be noted that, in the case where the embodiments of the present application involve user information, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the embodiments of the present application are all information and data authorized by the user or authorized by all parties, and the collection, use, and processing of related data need to comply with relevant laws, regulations, and standards of relevant countries and regions, and provide corresponding operation portals for the user to choose authorization or rejection. In addition, the various models (including but not limited to language models or large models) involved in the present application are in compliance with relevant laws and standard regulations.

[0051] In addition, it should be noted that, in the case where the embodiments of the present application involve user interaction operations or trigger operations, the user interaction operations or trigger operations involved in the embodiments of the present application include but are not limited to: touch operations, gesture operations, voice operations, head movement operations, eye movement operations, and various modes of interaction operations; wherein the touch operations include but are not limited to: click operations, double-click operations, long-press operations, sliding operations, pinch operations, or mouse hovering operations, etc. The sliding operations include but are not limited to: straight-line sliding, curve sliding, etc.

[0052] The embodiments of the present application can be applied to the training and application scenarios of the generation model.

[0053] In an embodiment of the present application, the generation model is a generative artificial intelligence model, such as a diffusion model, an autoregressive model, or a flow-based generation model, and the generation model can generate corresponding target data based on conditional data, such as text, images, videos, and / or audio, and the target data can be text, images, videos, and / or audio, so as to realize text-to-image, image-to-image, text-to-video, or image-to-video, and the like. In an actual application, the generation model can be an image generation model using a diffusion model architecture, and the image generation model can generate target images based on conditional data such as text and / or images.

[0054] Taking an image generation model implemented using a diffusion model architecture as an example, it can be known from the description of the background art that, in order to improve the generation quality, the CFG (Classifier-Free Guidance) technology is usually used for model training and model application. However, the CFG technology can cause the model to perform multiple forward inferences, and especially for a diffusion model using a complex sampling strategy, the inference cost can be greatly increased, which seriously affects the generation efficiency of the generation model.

[0055] Therefore, how to reduce the inference cost and improve the generation efficiency while ensuring high-quality generation effect has become a technical problem to be solved.

[0056] In order to improve the generation efficiency and ensure the generation quality, the inventors have proposed the technical solution of the present application after a series of researches, and the basic idea is as follows: first, a first image generation model using the CFG technology and performing at least twice forward inference can be determined, the first generation model is used as a teacher model, a second generation model needing actual application is used as a student model, the knowledge distillation method is used for model training of the second generation model, and the guidance scale is injected into the feature input stage to control the conditional features and the unconditional features to obtain the fusion features. Since the fusion features generate target data, the second generation model only needs to perform one forward inference, and does not need to perform forward inference for the unconditional features and the conditional features respectively, which can greatly reduce the inference cost and improve the generation efficiency. In addition, the knowledge distillation can extract the knowledge of the first generation model into the second generation model, so as to ensure the generation quality of the second generation model. The generation model obtained by the knowledge distillation method and the feature fusion method based on the guidance scale in the embodiments of the present application can take into account the generation quality and the generation efficiency, and improve the model performance and the generation effect.

[0057] With reference to the drawings and embodiments of the present application, the technical solutions in the embodiments of the present application will be described clearly and completely. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the scope of protection of the present application.

[0058] Figure 1 A flowchart of one embodiment of a model training method provided in the embodiments of the present application. The embodiments of the present application explain the technical solutions of the present application from the perspective of model training. The technical solutions of the embodiments of the present application can be executed by a server. The server can be implemented as a distributed server cluster composed of multiple servers, or as a single server. The server can be a server of a distributed system, or a server combined with a blockchain. The server can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content distribution networks (CDN), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology, etc. The present application does not limit this.

[0059] The method can include the following steps:

[0060] 101: Determine target training data.

[0061] The target training data includes conditional sample data.

[0062] The target training data can be obtained from a training data set, and the training data set can be an open source data set. Since the model training operation is usually iteratively executed until the training requirements are met, each iteration of the training operation can use multiple training data, and one iteration of the training operation can be considered as one training batch. The multiple training data corresponding to the current training batch, i.e., the current iteration of the training operation, can be used as the target training data.

[0063] The conditional sample data can include any data type such as text data, image data, audio data, and / or video data, which can be different according to the actual task type of the generation model.

[0064] 102: Determine conditional sample features and unconditional sample features according to the conditional sample data.

[0065] The sample unconditional data corresponding to the conditional sample data can be null data or all-zero data, etc., which is not limited by the present application.

[0066] The conditional sample data can be encoded as conditional sample features, and the unconditional sample data can be encoded as unconditional sample features.

[0067] The encoding of the conditional sample data and the unconditional sample data can be implemented by an encoding algorithm or a statistical algorithm, and the application does not limit this. The encoding can be performed according to different types of data.

[0068] Of course, the feature generation model can also be used to generate the conditional sample features corresponding to the conditional sample data and the unconditional sample features corresponding to the unconditional sample data. The feature generation model can use a pre-trained machine learning model, such as a text encoder when the conditional sample data is text data. The feature generation model can be pre-trained according to the sample data and the corresponding sample features. At this time, the step 102 can be implemented as follows: using the feature generation model, generating the conditional sample features corresponding to the conditional sample data; determining the unconditional sample data corresponding to the conditional sample data; using the feature generation model, generating the unconditional sample features corresponding to the unconditional sample data.

[0069] 103: using the first generation model, generating the first prediction data based on the conditional sample features and the second prediction data based on the unconditional sample features, and fusing the first prediction data and the second prediction data according to the sample guidance scale to obtain the first target prediction data.

[0070] The first generation model can be pre-trained as a teacher model. The first generation model can be trained based on a traditional method, and can perform two forward inferences for the conditional data and the unconditional data to ensure the generation quality of the final generated result.

[0071] In the embodiment of the application, the first generation model can be used to perform forward inference on the conditional sample features and the unconditional sample features, respectively, to obtain the first prediction data and the second prediction data. Then, the first target prediction data can be obtained by fusing the first prediction data and the second prediction data according to the sample guidance scale.

[0072] The sample guidance scale can be pre-set. In addition, in order to enable the to-be-trained model (the second generation model described below) to adapt to different guidance intensities and ensure the generation quality, the sample guidance scale can be randomly selected within a predetermined scale range, so that the model iterative training process can cover different guidance scales. Alternatively, the sample guidance scale can be randomly generated for the current training batch, or the sample guidance scale can be generated for the target training data, and the application does not limit this.

[0073] Optionally, the first generation model can also be trained according to a sample guide scale randomly selected in a predetermined scale range, for example, the first generation model can randomly select a sample guide scale for each training batch or each training data, and train the first generation model according to the sample guide scale and based on the training data or the corresponding training data in the corresponding training batch. The training data of the first generation model can include conditional data and training labels corresponding to the conditional data. The first generation model can perform forward inference based on the conditional sample features corresponding to the conditional sample data and the unconditional sample features to obtain two inference data, and then fuse the two inference data according to the sample guide scale to obtain prediction data, and then train the first generation model based on the difference information between the prediction data and the training labels.

[0074] The fusion manner can be linear combination of the first prediction data and the second prediction data according to the sample guide scale, and the sample guide scale can be used as a scaling parameter to linearly combine the first prediction data and the second prediction data. For example, in an implementation, the difference data between the first prediction data and the second prediction data can be determined, and the sample guide scale can be used to control the control strength of the first prediction data and the second prediction data on the final result. For example, the sample guide scale can be used as a scaling parameter of the difference data, and the first target prediction data can be obtained by multiplying the sample guide scale and the difference data. The corresponding fusion formula can be Y=X1+ω*(X1-X2), where Y represents the first target prediction data, X1 represents the first prediction data, X2 represents the second prediction data, and ω represents the sample guide scale. Of course, other linear combination manners can also be used, such as Y=X1+ω*X2, and the application is not limited thereto.

[0075] 104: Fuse the conditional sample features and the unconditional sample features according to the sample guide scale to obtain sample fusion features.

[0076] The fusion manner of the conditional sample features and the unconditional sample features can be the same as the fusion manner of the first prediction data and the second prediction data, and the conditional sample features and the unconditional sample features can be linearly combined according to the sample guide scale.

[0077] In addition, in the case of feature fusion, in order to adapt to the model data standard, unify the data form, etc., the sample guide scale can be encoded as a sample scale feature, and the conditional sample features and the unconditional sample features can be fused based on the sample scale feature.

[0078] In an implementation manner, the step 104 can be implemented as: mapping the sample guide scale into sample scale features; determining feature difference data between the conditional sample features and the unconditional sample features; taking the sample scale features as scaling parameters of the feature difference data to linearly combine the conditional sample features and the feature difference data to obtain sample fusion features. The corresponding fusion formula can be, for example:

[0079] c can represent the conditional sample data, and F(c) is the conditional sample feature; represents the unconditional sample data, is the unconditional sample feature; and g(ψ(ω)) represents the sample scale feature. Of course, the feature fusion can also be in other linear combination manners, such as and the like, and the present application is not limited thereto.

[0080] In actual application, the sample guide scale can be first subjected to sine-cosine coding, and then the coding result is mapped into the sample scale feature by using a feature conversion model, which is pre-trained by using, for example, a multi-layer perception structure, for example, can be trained based on the sine-cosine coded sample vector and the corresponding sample feature.

[0081] 105: generating second target prediction data based on the sample fusion features by using the second generation model.

[0082] The second generation model is a student model, that is, a to-be-trained model. In the embodiments of the present application, before the conditional sample features and the unconditional sample features are input into the second generation model, the feature processing can be first performed, the sample fusion features are obtained by fusing according to the sample guide scale, so that the second generation model can only perform forward inference based on the sample fusion features, and does not need to perform forward inference on the conditional sample features and the unconditional sample features respectively, thereby reducing the number of inferences and reducing the inference cost. The second target prediction data obtained by the second generation model based on the sample fusion features is the final generation result of the second generation model.

[0083] The first generation model and the second generation model can be models with the same model architecture and the same number of model parameters, so as to ensure the model effect of the second generation model. Of course, considering the deployment cost and the like, the second generation model can also be a model with the same architecture as the first generation model but with a smaller number of model parameters, and the present application does not limit the same.

[0084] 106: training the second generation model according to the difference information between the second target prediction data and the first target prediction data.

[0085] Optionally, a first loss between the second target prediction data and the first target prediction data can be calculated, and the second generation model can be trained based on only the first loss, so as to reduce training calculation amount, reduce resource consumption, and the like.

[0086] Optionally, the target training data can further include a training label corresponding to the conditional sample data, a second loss between the second target prediction data and the training label can be calculated, and the second generation model can be trained based on a joint loss corresponding to the first loss and the second loss, the joint loss can be obtained by weighted calculation of the first loss and the second loss, and the like, which is not limited in the present application. The first loss and the second loss can be calculated by using the same or different loss functions, and the like, which can be determined according to actual effect, which is not limited in the present application.

[0087] The second generation model can be used as a final model for actual application, in the model application stage, the second generation model can be used to generate target fusion features according to the conditional features and the unconditional features determined by the target conditional data and the target guide scale, and generate target data based on the target fusion features.

[0088] In the case of pre-setting the sample guide scale, the target guide scale can be the sample guide scale, in order to improve the generality of the model, in the case of randomly selecting the sample guide scale from a predetermined scale range, the target guide scale is also located in the predetermined scale range.

[0089] The target conditional data is the to-be-processed conditional data, and specific implementation in the model application stage will be introduced in the corresponding embodiments below.

[0090] In the present embodiment, the model is trained by the knowledge distillation method and the feature fusion method based on the guide scale, so that the inference times of the second generation model obtained by training are reduced, the inference cost is reduced, the generation efficiency is improved, the knowledge of the first generation model can be refined into the second generation model by the knowledge distillation, and the generation quality of the second generation model can be ensured, the present embodiment considers both the generation quality and the inference cost, so as to ensure the model performance.

[0091] The feature fusion manner can be obtained by linear combination calculation, so that no additional convolution layer or attention sub-layer and other model network layers are needed to realize, and the calculation efficiency is high. Since the model does not need to introduce additional network structure or hyperparameters, it can ensure that the first generation model and the second generation model are consistent in architecture and deployment. The second generation model can use the same model architecture as the first generation model, and combined with knowledge distillation, it can ensure the same generation quality as the first generation model without introducing additional deployment costs. Moreover, the feature fusion manner makes the model have low gradient noise, so that the number of training iterations can also be reduced, that is, a second generation model that meets the training requirements can be obtained. In the model application stage, only one forward inference is needed, so that the generation efficiency of the second generation model is much higher than that of the first generation model.

[0092] In some embodiments, the determining the target training data can include: taking the plurality of training data of the current training batch as the target training data respectively;

[0093] The training of the second image generation model according to the difference information between the second target prediction data and the first target prediction data can include:

[0094] According to the difference information between the second target prediction data and the first target prediction data corresponding to the plurality of training data respectively, the joint difference information corresponding to the plurality of training data is determined; and the second generation model is trained according to the joint difference information.

[0095] In the case where the difference information is represented by the first loss, the joint difference information is the comprehensive loss. The comprehensive loss can be, for example, the average loss of the first loss of the plurality of training data, and the model training is realized by adjusting the model parameters to minimize the joint loss.

[0096] As described in the foregoing, the sample guide scale corresponding to the plurality of training data of the current training batch can be randomly selected from the predetermined scale range. Of course, the sample guide scale can also be randomly selected from the predetermined scale range for the current training batch. The plurality of training data of the same training batch can correspond to the same sample guide scale.

[0097] Figure 2 The flowchart of one embodiment of the data generation method provided by the embodiments of the present application introduces the technical solutions of the present application from the perspective of model application. The technical solutions in the embodiments can be executed by a server. Of course, in the case where a model is deployed in a client, the technical solutions can also be executed by the client.

[0098] The client can be a browser, an APP (application), or a web application such as an H5 (HyperText Markup Language 5) application, or a light application (also known as a small program, a lightweight application), or a cloud application, etc. The client can be deployed in an electronic device, run in dependence on the device or some app in the device, etc. The electronic device may, for example, have a display screen and support information browsing, etc., and may be, for example, a personal mobile terminal such as a mobile phone, a tablet computer, a personal computer, a desktop computer, a smart speaker, a smart watch, etc. Various other types of applications can also be configured in the electronic device, such as human-computer dialogue type applications, model training type applications, text processing type applications, web browser applications, shopping type applications, search type applications, instant messaging tools, email clients, social platform software, etc. The electronic device can refer to a device used by a user and having functions such as computing, networking, and communication required by the user, and may, for example, be a mobile phone, a tablet computer, a personal computer, a wearable device, etc. The electronic device can generally include at least one processing component and at least one storage component. The electronic device can also include a network card chip, an IO (input / output) bus, an audio / video component, etc. The present application does not limit this. Optionally, according to the implementation form of the electronic device, some peripheral devices such as a keyboard, a mouse, an input pen, a printer, etc. can also be included, and the present application does not limit this.

[0099] The method can include the following steps:

[0100] 201: Obtain target condition data and target guide scale.

[0101] The target condition data can be provided by a user, and the client can determine the target condition data in response to a user input operation. The client can display data prompt information to prompt the user to provide the target condition data.

[0102] The target guide scale can be pre-set by a technician when deploying the model, or it can also be provided by a user. The client can determine the target guide scale in response to a user input operation, etc.

[0103] The client can display scale prompt information, etc. The scale prompt information can include the predetermined scale range to prompt the user to select the target guide scale from the predetermined scale range.

[0104] In the case where the service end executes the technical solution of the present embodiment, the service end can obtain the target condition data and the target guide scale provided by the client.

[0105] 202: Determine condition features and unconditional features according to the target condition data.

[0106] The target condition data can encode the condition feature; and the unconditional data corresponding to the target condition data, such as empty data or zero data, can encode the unconditional feature. The condition feature can be generated based on the target condition data by using the feature generation model, and the unconditional feature can be generated based on the unconditional data.

[0107] 203: Obtain the target fusion feature by fusing the condition feature and the unconditional feature according to the target guide scale.

[0108] The feature fusion manner can be the same as the fusion manner of the condition sample feature and the unconditional sample feature in the model training process, such as linear combination. Alternatively, the target guide scale can be mapped to a target scale feature, the feature difference data between the condition feature and the unconditional feature can be determined, and the target guide scale is used as a scaling parameter of the feature difference data to fuse the condition feature and the unconditional feature to obtain the target fusion feature. For example, the target scale feature is multiplied by the feature difference data and then superimposed with the condition feature to obtain the target fusion feature.

[0109] 204: Generate the target data based on the target fusion feature by using the second generation model.

[0110] The second generation model takes the first generation model as a teacher model and is trained according to the difference information between the first target prediction data and the second target prediction data; the first target prediction data is obtained by fusing the first prediction data and the second prediction data according to the sample guide scale; the first prediction data is generated by the first generation model based on the condition sample feature determined by the condition sample data in the training data, and the second prediction data is generated by the first generation model based on the unconditional sample feature; the second target prediction data is generated by the second generation model based on the sample fusion feature; and the sample fusion feature is obtained by fusing the condition sample feature and the unconditional sample feature according to the sample guide scale.

[0111] The training manner of the second generation model can be detailed in the embodiments described in Figure 1 and will not be repeated here.

[0112] In this embodiment, the second generation model only needs to perform one forward inference by the feature fusion manner, that is, the target data can be generated, the generation efficiency is improved, and the second generation model is trained based on the first generation model which performs two forward inferences in a knowledge distillation manner, so that the knowledge and skills of the first generation model can be learned, thereby ensuring the generation quality.

[0113] In an actual application, the generation model applicable to the embodiments of the present application can be specifically an image generation model, and the image generation model can generate image data based on condition data. In an actual application, the image generation model can be a diffusion model.

[0114] A diffusion model is a generative artificial intelligence technology based on deep learning, which generates high-quality data by simulating diffusion phenomena in nature. For an image generation model implemented based on a diffusion model, noise data and time information need to be introduced, and the image generation process is a process of iteratively denoising according to the time information under the guidance of the conditional data to gradually reconstruct the image. Denoising means sampling. In the case of using a complex sampling strategy, two forward inferences will greatly increase the inference cost and affect the generation efficiency. The technical solution of the present application can solve this technical problem.

[0115] The technical solution of the present application will be introduced below taking an image generation model implemented based on a diffusion model as an example, Figure 3 A flowchart of another embodiment of a model training method provided by the present application, the technical solution of the present embodiment can be executed by a server, and the method can include the following steps.

[0116] 301: Obtain target training data.

[0117] The target training data can include conditional sample data and image sample data corresponding to the conditional sample data.

[0118] The image sample data is also the image data without any noise corresponding to the conditional sample data.

[0119] Taking the conditional sample data as text data as an example, the training data can be obtained from an open-source image-text alignment data set, and the resolution and text length can be unified without manual annotation.

[0120] Since the model training operation is usually iteratively executed until the training requirement is met, each iteration training operation can use multiple training data, and one iteration training operation can be considered as one training batch. The multiple training data corresponding to the current training batch, i.e., the current iteration training operation, can be used as the target training data.

[0121] 302: Determine the time step, sample guide scale and sample noise data corresponding to the target training data.

[0122] Since the image generation model implemented based on the diffusion model needs to introduce noise data and time information, in addition to determining the sample guide scale, the time step and the sample noise data also need to be determined in the present embodiment.

[0123] The image generation model is iteratively denoised in the model application stage, so the denoising process of multiple time steps is executed, and the multiple time steps constitute a time step sequence such as 0-1000 time steps. When training the model, in order to simplify the training process, the time sequence can not be introduced, and only one time step can be processed.

[0124] Optionally, the time step can be obtained by random selection from a predetermined numerical interval. The predetermined numerical interval may, for example, be U[0, T], representing uniform distribution over the interval [0, T], and in actual applications, T is usually normalized to the value 1. The time step t can be randomly sampled from the predetermined numerical interval by uniform distribution. Optionally, the time step t can be randomly sampled for the current training batch, or for the target training data, and the like, which are not limited in the present application.

[0125] The sample noise data can be randomly generated, for example, can be randomly sampled from a standard normal distribution N(0, 1). Optionally, the sample noise data can be randomly generated for the current training batch, or for the target training data, and the like, which are not limited in the present application.

[0126] The sample guide scale can be pre-set, and in addition, optionally, the sample guide scale can be randomly selected within a predetermined scale range, so that the model iterative training process can cover different guide scales.

[0127] Optionally, the sample guide scale can be randomly generated for the current training batch, or for the target training data, and the like, which are not limited in the present application.

[0128] 303: Determine the conditional sample feature and the unconditional sample feature according to the conditional sample data.

[0129] The feature generation process of step 303 is the same as that of step 102, which will not be repeated here.

[0130] 304: Generate noise image data corresponding to the time step according to the sample noise data and the image sample data.

[0131] By adding sample noise data to the image sample data, noise image data corresponding to the current time step can be obtained. The image generation model is used to denoise the noise image data of the current time step to predict the real noise data added, so that the image generation model can be trained according to the difference information between the predicted noise data and the real noise data. The image generation model obtained by training can obtain the target image from a target noise data by step-by-step denoising.

[0132] The noise adding mode can be implemented by a traditional method, which is not limited in the present application.

[0133] 305: generating, by using the first image generation model, first prediction data based on the conditional sample feature, the time step, and the noise image data, and second prediction data based on the unconditional sample feature, the time step, and the noise image data, and fusing the first prediction data and the second prediction data according to the sample guidance scale to obtain first noise prediction data.

[0134] In this embodiment, the first image generation model can be pre-trained as a teacher model, and can be trained based on a traditional method. The first image generation model can perform two forward inferences for conditional data and unconditional data to ensure the generation quality of the final generation result.

[0135] The first prediction data and the second prediction data generated by the first image generation model can be fused according to the sample guidance scale to obtain the final model generation result, i.e., the first noise prediction data.

[0136] As described above, the sample guidance scale can be randomly selected from a predetermined scale range. Alternatively, the first image generation model can be trained according to the randomly selected sample guidance scale in the predetermined scale range. For example, the first image generation model can randomly select a sample guidance scale for each training batch or each training data, and train the first image generation model according to the sample guidance scale based on the training data in the corresponding training batch or the corresponding training data. The training data of the first generation model can include conditional data and corresponding noise-free image data of the conditional data. The noise-free image data can be combined with random noise data to obtain noise image data. The first image generation model can combine the noise image time and the time step, and perform forward inference based on the conditional sample feature corresponding to the conditional sample data and the unconditional sample feature, respectively, to obtain two inference data. The two inference data are fused according to the sample guidance scale to obtain noise prediction data. Then, the first image generation model is trained based on the difference information between the noise prediction data and the noise image data.

[0137] The fusion manner can be linear combination of the first prediction data and the second prediction data according to the sample guidance scale, the sample guidance scale can be used as a scaling parameter to linearly combine the first prediction data and the second prediction data, and the like. The fusion manner can be specifically referred to in step 103, which will not be repeated here.

[0138] 306: fusing the conditional sample feature and the unconditional sample feature according to the sample guidance scale to obtain sample fusion feature.

[0139] The specific implementation manner of the sample fusion feature can be referred to in step 104, which will not be repeated here.

[0140] 307: generating, by using the second image generation model, second noise prediction data based on the sample fusion feature, the time step, and the noise image data.

[0141] The second image generation model is a student model of the first image generation model, that is, a to-be-trained model. Before inference, the second image generation model can first perform feature processing to obtain the sample fusion feature. Thus, the second image generation model can perform a forward inference based on the sample fusion feature, the time step, and the noise image data, without separately performing forward inferences on the conditional sample feature and the unconditional sample feature, thereby reducing the inference cost. The second noise prediction data is the final generation result of the second image generation model.

[0142] To adapt to model data standards, unify data formats, etc., the time step can also be mapped to a time feature. In actual application, the time step can be first encoded by using a sine-cosine coding method, and then the encoded result is mapped to a time feature by using a feature conversion model, which is pre-trained by using, for example, a multi-layer perception structure, for example, can be trained based on a sine-cosine coding sample vector and a corresponding sample feature.

[0143] The sample guide scale can be directly converted to a corresponding sample scale feature by using the feature conversion model, thereby reducing the deployment cost without increasing the network result or hyperparameters of the model, and ensuring the consistency of the student model and the teacher model in architecture and the generation quality.

[0144] The first image generation model and the second generation model can be diffusion models with the same model architecture and the same number of model parameters, such as a Diffusion Transformer model based on a flow matching framework of an invertible ordinary differential equation and a Transformer architecture.

[0145] 308: training the second image generation model according to difference information between the second noise prediction data and the first noise prediction data.

[0146] The difference information can be used to calculate a first loss by using a first loss function, for example, a mean square error loss function, etc. Of course, the present application is not limited thereto.

[0147] Optionally, a second loss between the second noise prediction data and the noise image data can also be calculated, so that the second image generation model can be trained based on a joint loss corresponding to the first loss and the second loss.

[0148] The second image generation model can be used as a final model in actual application. In the model application stage, the second image generation model is used to generate target fusion features according to the conditional features and unconditional features determined according to the target condition data and the target guide scale, and generate target images based on the target fusion features, the time sequence, and the target noise data.

[0149] In the case that the sample guide scale is randomly selected from the predetermined scale range, the target guide scale is also located in the predetermined scale range.

[0150] The technical scheme of the embodiment of the present application guarantees the generation quality and reduces the inference cost, and improves the generation efficiency. The feature fusion manner is realized by linear combination, which can further improve the generation efficiency and reduce the deployment cost.

[0151] In some embodiments, the determination of the target training data can include: taking each of the plurality of training data in the current training batch as the target training data.

[0152] The training of the second image generation model according to the difference information between the second noise prediction data and the first noise prediction data can include: calculating a first loss of the second noise prediction data and the first noise prediction data corresponding to each training data; determining a comprehensive loss according to the first loss corresponding to each of the plurality of training data; and training the second generation model according to the comprehensive loss.

[0153] The comprehensive loss can be an average loss of the first loss corresponding to the plurality of training data, and the model training is realized by adjusting the model parameters to minimize the joint loss.

[0154] For ease of understanding, taking the condition data as text data and the image generation model as an example of text-to-image, the technical scheme of the present application is introduced in combination with the model structure diagram shown in Figure 4 .

[0155] For each training data in each training batch, it is assumed that the unconditional sample data corresponding to the condition sample data c in the training data is, for example, empty data “”, and a certain condition sample data is, for example, text data of “Cyberpunk style car”.

[0156] The condition sample data and the empty data can be respectively used to generate corresponding condition sample features F(c) and unconditional sample features

[0157] In addition, for each training batch or each training data, a time step t can be randomly sampled from a uniform distribution U[0, 1], representing sampling a t from a uniform distribution between 0 and 1; a sample noise data ε can be randomly generated from a standard normal distribution N(0, 1), representing sampling a sample noise data ε from a standard normal distribution; and a sample guide scale ω can be randomly selected from [ω min , ω max ], representing randomly sampling a sample guide scale from a value range between ω min and ω max .

[0158] Then, the sample noise data is added to the sample image data x0 in the training data to obtain noise image data x t , and the noise addition manner can be, for example, x t = (1-t)x0+t*ε, of course, the present application is not limited thereto.

[0159] Wherein, the time step t and the sample guide scale ω can be encoded as a time step feature g(ψ(t)) and a sample scale feature g(ψ(ω)), such as after being encoded by a cosine, and being mapped by a feature conversion model.

[0160] As the first image generation model 402 of the teacher model, based on F(c), g(ψ(t)) and x t , two forward inferences are respectively performed to obtain first prediction data ∈ T (x t, t, c) and second prediction data Then, the first prediction data and the second prediction data can be linearly combined according to the sample guide scale to obtain target noise prediction data ∈ ~ T (x t , ω, c), that is:

[0161]

[0162] As the second image generation model 403 of the student model, first, the unconditional feature and the conditional feature are fused based on the sample guide scale to obtain a sample fusion feature z, for example Then, the time step feature can be fused to obtain The first student model can obtain second noise prediction data ∈ S (xt, z ~ ) based on z~and xt. Since the unconditional sample feature has been fused into z~, the first student model performs a forward inference.

[0163] Finally, the first loss of the first noise prediction data and the second noise prediction data can be continued, such as the mean square error loss L2 = ||∈ S (xt, z ~ )-∈ ~ T (x t , ω, c)| 2 .

[0164] For each training batch, the average loss of the first loss of multiple training data can be calculated, so that the model parameters of the second image generation model can be adjusted based on the average loss by back propagation until the training requirements are met. Since the feature fusion method does not need to introduce additional network structure, it has low gradient noise, can reduce the number of training iterations, improve training efficiency, and improve the generation efficiency of the first image generation model obtained by training.

[0165] The model structures of the first image generation model and the second image generation model can be consistent, so that knowledge distillation can be better performed to improve the generation quality of the second image generation model. In the case of using a diffusion model, the first image generation model uses a first sampling strategy, and the second image generation model uses a second sampling strategy. ~ T (x t , t, c) can be obtained through multiple derivations, and each derivation needs to pass through twice forward reasoning, which has relatively large reasoning cost. As a student model, the second image generation model can use a second sampling strategy, which only needs one forward reasoning for each derivation, reducing the reasoning cost. The second sampling strategy can select a sampling strategy with lower complexity than the first sampling strategy, such as the Euler sampling strategy, and the reasoning cost of the second image generation model can be further reduced.

[0166] In the embodiments of the present application, the model is trained by the knowledge distillation method and the feature fusion method based on the guide scale, so that the reasoning times of the second image generation model obtained by training are reduced, the reasoning cost is reduced, the generation efficiency is improved, and the knowledge distillation can extract the knowledge of the first image generation model into the second image generation model, which can ensure the generation quality of the second image generation model, and does not need to introduce any network structure in the second image generation model, so that the generation efficiency is higher, the model architecture of the first image generation model can be used, the deployment is simple, no additional deployment cost is introduced, and the training speed is faster.

[0167] Figure 5 The flowchart of one embodiment of the image generation method provided in the embodiments of the present application is introduced, and the technical solution in the embodiments of the present application can be executed by the server. Of course, in the case of deploying the model in the client, the client can also execute it.

[0168] The method can include the following steps:

[0169] 501: Obtain target condition data and target guidance scale.

[0170] 502: According to the target condition data, determine the condition feature and the unconditional feature.

[0171] 503: According to the target guidance scale, fuse the condition feature and the unconditional feature to obtain the target fusion feature.

[0172] The operations of steps 501-503 can refer to the operations of steps 201-202, which will not be repeated here.

[0173] 504: Randomly generate target noise data and determine time sequence feature.

[0174] The target noise data can be randomly sampled from a standard normal distribution N(0, 1).

[0175] Since the image generation stage is an iterative denoising process of target noise data under the guidance of condition data according to time information, a time sequence feature needs to be introduced. The time sequence can be uniformly sampled from a uniform distribution U[0, 1] to form multiple time steps, and each time step is converted into a time step feature to form a time sequence feature.

[0176] 505: Use the second image generation model to generate a target image based on the target fusion feature, the time sequence feature, and the target noise data.

[0177] The second image generation model only needs to perform a forward inference, and can iteratively denoise the target noise data according to the time information of the time sequence feature under the guidance of the target fusion feature, thereby gradually reconstructing and obtaining the target image.

[0178] The second image generation model takes the first image generation model as a teacher model and is trained according to the difference information between the first noise prediction data and the second noise prediction data; the first noise prediction data is obtained by fusing the first prediction data and the second prediction data according to the sample guidance scale; the first prediction data is generated by the first image generation model based on the condition sample feature, the time step, and the noise image data in the training data; the second prediction data is generated by the first image generation model based on the unconditional sample feature, the time step, and the noise image data; the second noise prediction data is generated by the second generation model based on the sample fusion feature, the time step, and the noise image data; and the sample fusion feature is obtained by fusing the condition sample feature and the unconditional sample feature according to the sample guidance scale.

[0179] The training manner of the second image generation model can be referred to Figure 3 The details are not repeated here.

[0180] In this embodiment, the second image generation model only needs to perform one forward inference to generate the target data through the feature fusion manner, which improves the generation efficiency. The second image generation model is trained based on the first image generation model performing two forward inferences in a knowledge distillation manner, which can learn the knowledge and skills of the first image generation model, thereby ensuring the generation quality.

[0181] The detailed implementation and beneficial effects of each step in the method of this embodiment have been described in detail in the foregoing embodiments, and will not be described in detail here.

[0182] In the model application stage, the technical solution of the embodiment of the present application can be applied to the system architecture as shown in Figure 6 The client 601 can provide a model application interface to perceive user input operations, thereby determining target condition data, such as text description information “orange cat, windowsill, sunlight”. The model application interface can provide scale selection prompt information to perceive user selection operations, thereby determining a target guide scale.

[0183] The client 601 sends the target condition data and the target guide scale to the server 602.

[0184] The server 602 can call the second image generation model to perform text-to-image operation. First, the feature generation model 401 is used to generate the condition features corresponding to the target condition data and the unconditional features corresponding to the unconditional data of the target condition data. Then, according to the target guide scale, the condition features and the unconditional features are fused to obtain target fusion features. The target fusion features, the time sequence features, and the target noise data can be input into the second image generation model 404, so as to generate the target image by gradually iterating and denoising the target noise data.

[0185] The server 602 can send the target image to the client 601, and the client 601 can display the target image on the model application interface, so that the user can obtain the target image.

[0186] It should be noted that the technical solution of the embodiment of the present application is applicable to a network virtual environment. The user described is generally a “virtual user”. A real user can register a user account in the server through a registration method to obtain a user identity in the network environment. The same user account can be logged into the server through different types of user terminals, so that the server can identify the same user.

[0187] The interaction between the service end and the user can be realized based on the user account, and the corresponding data received or sent by the service end to the user is also realized based on the user account, and in fact, the corresponding data is received or sent by the user end corresponding to the user account to the service end. In addition, communication and the like can also be realized between users through user accounts. Among them, the user can refer to an individual, or an institution such as an enterprise, and the present application does not make specific limitations thereto.

[0188] It should be noted that in some of the processes described in the above embodiments and the accompanying drawings, a plurality of operations appearing in a specific order are included, but it should be clearly understood that these operations can be executed in the order appearing in this text or in parallel, and the serial numbers of the operations such as 101, 102, etc. are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes can include more or fewer operations, and these operations can be executed in sequence or in parallel. It should be noted that the "first", "second" and the like described herein are used to distinguish different messages, devices, modules, etc., and do not represent the order of precedence, nor do "first" and "second" represent different types.

[0189] Figure 7 An embodiment of a model training device provided by the present application is shown in the structural schematic diagram. The device can include:

[0190] The data determination module 701 is configured to determine target training data. The target training data includes conditional sample data.

[0191] The first feature generation module 702 is configured to determine conditional sample features and unconditional sample features according to the conditional sample data.

[0192] The first prediction module 703 is configured to generate first prediction data based on the conditional sample features and second prediction data based on the unconditional sample features by using a first generation model, and fuse the first prediction data and the second prediction data according to a sample guide scale to obtain first target prediction data.

[0193] The first feature fusion module 704 is configured to fuse the conditional sample features and the unconditional sample features according to the sample guide scale to obtain sample fusion features.

[0194] The second prediction module 705 is configured to generate second target prediction data based on the sample fusion features by using a second generation model.

[0195] The model training module 706 is configured to train the second generation model according to the difference information between the second target prediction data and the first target prediction data; the first generation model is used as a teacher model of the second generation model; the second generation model is configured to generate target fusion features according to the conditional features and the unconditional features determined according to the target condition data and the target guide scale, and generate the target data based on the target fusion features.

[0196] In some embodiments, the apparatus can further include:

[0197] The scale selection module is configured to randomly select the sample guide scale from a predetermined scale range; and the target guide scale is located in the predetermined scale range.

[0198] In some embodiments, the data determination module can be specifically configured to take the plurality of training data in the current training batch as the target training data respectively;

[0199] The model training module can be specifically configured to determine joint difference information corresponding to the plurality of training data according to difference information between second target prediction data and first target prediction data corresponding to the plurality of training data respectively; and train the second generation model according to the joint difference information.

[0200] In some embodiments, the scale selection module can be specifically configured to randomly select the sample guide scale corresponding to the plurality of training data from a predetermined scale range; and the target guide scale is located in the predetermined scale range.

[0201] In some embodiments, the first feature fusion module can be specifically configured to map the sample guide scale to a sample scale feature; determine feature difference data between the conditional sample feature and the unconditional sample feature; and take the sample size feature as a scaling parameter of the feature difference data to linearly combine the conditional sample feature and the feature difference data to obtain the sample fusion feature.

[0202] In some embodiments, the first feature generation module can be specifically configured to generate the conditional sample feature corresponding to the conditional sample data by using the feature generation model; determine the unconditional sample data corresponding to the conditional sample data; and generate the unconditional sample feature corresponding to the unconditional sample data by using the feature generation model.

[0203] In some embodiments, the first generation model can be a first image generation model, and the second generation model can be a second image generation model; and the target training data can include the conditional sample data and image sample data.

[0204] The apparatus can further include:

[0205] The first data generation module is configured to determine a time step and sample noise data corresponding to the target training data.

[0206] The noise generation model is configured to generate noise image data corresponding to a time step according to sample noise data and image sample data;

[0207] The first prediction module can be specifically configured to generate first prediction data based on conditional sample features, a time step, and noise image data by using a first image generation model, generate second prediction data based on unconditional sample features, a time step, and noise image data, and fuse the first prediction data and the second prediction data according to a sample guidance scale to obtain first noise prediction data.

[0208] The second prediction module can be specifically configured to generate second noise prediction data based on sample fusion features, a time step, and noise image data by using a second image generation model.

[0209] The model training module can be specifically configured to train the second image generation model according to difference information between the second noise prediction data and the first noise prediction data; and the second image generation model is configured to generate target fusion features based on conditional features and unconditional features determined by target condition data and a target guidance scale, and generate a target image based on the target fusion features, a time sequence, and target noise data.

[0210] In some embodiments, the first data generation module can randomly select a time step corresponding to target training data from a predetermined numerical interval, and randomly generate sample noise data corresponding to the target training data.

[0211] In some embodiments, the model training module can be specifically configured to calculate a first loss of the second noise prediction data and the first noise prediction data corresponding to each training data, determine a comprehensive loss according to the first losses of the multiple training data, and train the second generation model according to the comprehensive loss.

[0212] In some embodiments, the first feature generation module can be specifically configured to generate conditional sample features corresponding to conditional sample data by using a feature generation model, determine empty data corresponding to the conditional sample data, and generate unconditional sample features corresponding to the empty data by using the feature generation model.

[0213] Figure 7 The model training device can perform Figure 1 or Figure 3 The implementation principle and technical effects of the model training method according to the embodiments are not repeated. The specific operation of each module and unit of the model training device in the above embodiments has been described in detail in the embodiments related to the method, and will not be described in detail here.

[0214] Figure 8A structural schematic diagram of one embodiment of a data generation apparatus provided in the present application can include:

[0215] The data acquisition module 801 is configured to acquire target condition data and target guide dimensions.

[0216] The second feature generation module 802 is configured to determine condition features and unconditional features according to the target condition data.

[0217] The second feature fusion module 803 is configured to fuse the condition features and the unconditional features according to the target guide dimensions to obtain target fusion features.

[0218] The data processing module 804 is configured to generate target data based on the target fusion features by using a second generation model, wherein the second generation model takes the first generation model as a teacher model and is trained according to difference information between first target prediction data and second target prediction data; the first target prediction data is obtained by fusing first prediction data and second prediction data according to sample guide dimensions; the first prediction data is generated by the first generation model based on condition sample features determined by the first generation model based on condition sample data in training data; the second prediction data is generated by the first generation model based on unconditional sample features; the second target prediction data is generated by the second generation model based on sample fusion features; and the sample fusion features are obtained by fusing the condition sample features and the unconditional sample features according to the sample guide dimensions.

[0219] In some embodiments, the apparatus can further include:

[0220] The second data generation module is configured to randomly generate target noise data and determine time sequence features.

[0221] The data acquisition module can be specifically configured to acquire target condition data and target guide dimensions.

[0222] The data processing module can be specifically configured to generate a target image based on the target fusion feature, the time sequence feature, and the target noise signal data by using a second image generation model; the second image generation model takes the first image generation model as a teacher model and is trained according to difference information between first noise prediction data and second noise prediction data; the first noise prediction data is obtained by fusing the first prediction data and the second prediction data according to a sample guide scale; the first prediction data is generated based on a conditional sample feature, a time step, and noise image data in the training data by the first image generation model, and the second prediction data is generated based on an unconditional sample feature, a time step, and noise image data by the first image generation model; the second noise prediction data is generated based on a sample fusion feature, a time step, and noise image data by the second generation model; and the sample fusion feature is obtained by fusing the conditional sample feature and the unconditional sample feature according to the sample guide scale.

[0223] Figure 8 The data generation apparatus can perform Figure 2 the data generation method of the embodiments shown in the Figure 5 the image generation method of the embodiments shown in the The implementation principles and technical effects of the data generation method of the embodiments shown in the

[0224] Figure 9 An embodiment of a computing device according to the present disclosure is shown in a structural schematic diagram. As shown in the Figure 9 In practice, the computing device can include a storage component 901 and a processing component 902.

[0225] The storage component 901 is configured to store computer programs and can be configured to store other various data to support operations on the computing device. Examples of these data include instructions of any application or method for operating on the computing device, data structures, contact data, phonebook data, messages, pictures, videos, and the like.

[0226] The processing component 902 is coupled to the storage component 901 and is configured to execute the computer programs in the storage component 901 for implementing the model training method as shown in the Figure 1 the data generation method as shown in the Figure 2 the data generation method as shown in the Figure 3 the data generation method as shown in the Figure 5 the image generation method as shown in the

[0227] Further, the computing device can further include a communication component 903, a display component 904, an input component 905, and a power supply component 906. Figure 9As shown, the computing device can further include a communication component 903, a display component 904, a power supply component 905, an audio component 906, and other components. Figure 9 Some components are only shown schematically and do not mean that the computing device only includes Figure 9 the components shown. In addition, Figure 9 The components in the dashed box are optional components, not mandatory components, and depend on the product form of the computing device. The computing device of the present embodiment can be implemented as a terminal device such as a desktop computer, a notebook computer, a smart phone, or an IOT (Internet of Things) device, or as a server device such as a general server, a cloud server, or a server array. If the computing device of the present embodiment is implemented as a terminal device such as a desktop computer, a notebook computer, or a smart phone, it can include Figure 9 the components in the dashed box; if the computing device of the present embodiment is implemented as a server device such as a general server, a cloud server, or a server array, it can not include Figure 9 the components in the dashed box.

[0228] The processing component described above includes one or more processors to execute computer instructions to complete all or part of the steps of the methods described above. Of course, the processing component can also be one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components, for executing the methods described above.

[0229] The storage component described above can be implemented by any type of volatile or non-volatile storage devices or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.

[0230] The communication component is configured to facilitate wired or wireless communication between the device on which the communication component is installed and other devices. The device on which the communication component is installed can access a wireless network based on a communication standard, such as a radio communication technology, a wireless local area network (WLAN) technology, a Bluetooth (BT) technology, a near field communication (NFC) technology, and / or a global positioning system (GPS) technology. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast related information from an external broadcast management system via a broadcast channel.

[0231] The display component can include a screen, which can include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive an input signal from a user. The touch panel includes one or more touch sensors to sense a touch, a slide and a gesture on the touch panel. The touch sensor can not only sense a boundary of a touching or a sliding action, but also detect duration and pressure related to the touching or sliding action.

[0232] The power supply component provides power to various components of the device on which the power supply component is installed. The power supply component can include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power for the device on which the power supply component is installed.

[0233] The audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC) that is configured to receive an external audio signal when the device on which the audio component is installed is in a particular mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in a memory or transmitted via the communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals.

[0234] Accordingly, the embodiments of the present application also provide a computer readable storage medium storing a computer program, when the computer program is executed by a processor, the processor is enabled to implement each step in the above-mentioned method embodiments. The computer readable storage medium includes volatile or non-volatile or a combination thereof, and can be removable or non-removable. Examples of the computer readable storage medium include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital video disc (DVD) or other optical storage, magnetic cassette, magnetic tape disk storage or other magnetic storage device, or any other non-transmission medium

[0235] Accordingly, the embodiments of the present application also provide a computer program product, the computer program product includes a computer program or instructions, when the computer program or instructions are executed by a processor, the processor is enabled to implement each step in the above-mentioned method embodiments. It should be understood that each process or a combination of multiple processes in the above-mentioned method flow can be implemented by the computer program or instructions. In addition, these computer programs or instructions can be applied to the processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device, so that the processor of the general-purpose computer, the special-purpose computer, the embedded processor or other programmable data processing device can be implemented as a device for implementing the corresponding functions in the above-mentioned method embodiments.

[0236] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-mentioned system, device and unit can refer to the corresponding process in the above-mentioned method embodiments, which will not be described here.

[0237] It should also be noted that the terms "comprising", "comprises", "including", "includes" or any other variation thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises a... " does not, without more constraints, exclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0238] Finally, it should be noted that only the preferred embodiments of the application have been described, and that all modifications and changes, which come within the spirit of the application, are therefore intended to be protected.

Claims

1. A model training method, characterized in that, include: Determine the target training data; The target training data includes conditional sample data; Based on the conditional sample data, determine the features of the conditional samples and the features of the unconditional samples; Using the first generative model, first prediction data is generated based on the conditional sample features, and second prediction data is generated based on the unconditional sample features. The first prediction data and the second prediction data are then fused according to the sample guidance scale to obtain the first target prediction data. According to the sample guidance scale, the conditional sample features and the unconditional sample features are fused to obtain sample fusion features; Using the second generative model, second target prediction data is generated based on the sample fusion features; The second generative model is trained based on the difference between the second target prediction data and the first target prediction data; the first generative model serves as the teacher model for the second generative model; the second generative model is used to generate target fusion features based on the conditional and unconditional features determined by the target conditional data and the target guidance scale, and to generate target data based on the target fusion features.

2. The method according to claim 1, characterized in that, Also includes: The sample guidance scale is randomly selected from a predetermined scale range; The target guidance scale is within the predetermined scale range.

3. The method according to claim 1, characterized in that, The target training data includes: Use multiple training data points from the current training batch as target training data respectively; The step of training the second image generation model based on the difference information between the second target prediction data and the first target prediction data includes: Based on the difference information between the second target prediction data and the first target prediction data corresponding to the plurality of training data, the joint difference information corresponding to the plurality of training data is determined. The second generative model is trained based on the joint difference information.

4. The method according to claim 3, characterized in that, Also includes: Randomly select sample guidance scales corresponding to the multiple training data within a predetermined scale range; The target guidance scale is within the predetermined scale range.

5. The method according to claim 1, characterized in that, The step of fusing the unconditional sample features into the conditional sample features according to the sample guidance scale to obtain sample fusion features includes: The sample guiding scale is mapped to sample scale features; Determine the feature difference data between the conditional sample features and the unconditional sample features; The sample size feature is used as a scaling parameter for the feature difference data to linearly combine the conditional sample feature with the feature difference data to obtain the sample fusion feature.

6. The method according to claim 1, characterized in that, The step of determining the conditional sample features and unconditional sample features based on the conditional sample data includes: Using a feature generation model, generate conditional sample features corresponding to the conditional sample data; Determine the unconditional sample data corresponding to the conditional sample data; Using the feature generation model, unconditional sample features corresponding to the unconditional sample data are generated.

7. A data generation method, characterized in that, include: Obtain target condition data and target guidance scale; Based on the target condition data, determine the conditional features and the unconditional features; According to the target guidance scale, the conditional features and the unconditional features are fused to obtain the target fusion features; Target data is generated using a second generative model based on the target fusion features. The second generative model uses the first generative model as a teacher model and is trained based on the difference information between the first and second target prediction data. The first target prediction data is obtained by fusing the first and second prediction data according to a sample guidance scale. The first prediction data is generated by the first generative model based on conditional sample features determined by the first generative model from conditional sample data in the training data. The second prediction data is generated by the first generative model based on unconditional sample features. The second target prediction data is generated by the second generative model based on sample fusion features. The sample fusion features are obtained by fusing the conditional sample features and the unconditional sample features according to the sample guidance scale.

8. A model training method, characterized in that, include: Obtain target training data; The target training data includes conditional sample data and image sample data corresponding to the conditional sample data; Determine the time step, sample guidance scale, and sample noise data corresponding to the target training data; Based on the conditional sample data, determine the features of the conditional samples and the features of the unconditional samples; Based on the sample noise data and the image sample data, generate the noise image data corresponding to the time step; Using a first image generation model, first prediction data is generated based on the conditional sample features, the time step, and the noisy image data, and second prediction data is generated based on the unconditional sample features, the time step, and the noisy image data. The first prediction data and the second prediction data are then fused according to the sample-guided scale to obtain the first noise prediction data. According to the sample guidance scale, the conditional sample features and the unconditional sample features are fused to obtain sample fusion features; Using a second image generation model, second noise prediction data is generated based on the sample fusion features, the time step, and the noisy image data; The second image generation model is trained based on the difference between the second noise prediction data and the first noise prediction data. The second image generation model is used to generate target fusion features based on the conditional features and unconditional features determined by the target conditional data and the target guiding scale, and to generate a target image based on the target fusion features, time series, and target noise data.

9. The method according to claim 8, characterized in that, The determination of the time step, sample guidance scale, and sample noise corresponding to the target training data target includes: Randomly select the time step corresponding to the target training data from a predetermined numerical range; Randomly generate sample noise data corresponding to the target training data; The sample guidance scale corresponding to the target training data is randomly selected from a predetermined scale range; the target guidance scale is located within the predetermined scale range.

10. The method according to claim 8, characterized in that, The target training data includes: Use multiple training data points from the current training batch as target training data respectively; The step of training the second image generation model based on the difference information between the second noise prediction data and the first noise prediction data includes: Calculate the first loss between the second noise prediction data and the first noise prediction data corresponding to each training data; Based on the first loss corresponding to each of the multiple training data, determine the comprehensive loss; The second generative model is trained based on the comprehensive loss.

11. The method according to claim 8, characterized in that, The step of determining the conditional sample features and unconditional sample features based on the conditional sample data includes: Using a feature generation model, generate conditional sample features corresponding to the conditional sample data; Determine the empty data corresponding to the conditional sample data; Using the feature generation model, unconditional sample features corresponding to the empty data are generated.

12. An image generation method, characterized in that, include: Acquire target condition data and target guidance scale; Based on the target condition data, determine the conditional features and the unconditional features; Randomly generate target noise data and determine time series characteristics; According to the target guidance scale, the conditional features and the unconditional features are fused to obtain the target fusion features; A target image is generated using a second image generation model based on the target fusion features, the time series features, and the target noise information data; wherein, the second image generation model uses the first image generation model as a teacher model and is trained based on the difference information between the first noise prediction data and the second noise prediction data. The first noise prediction data is obtained by fusing the first prediction data and the second prediction data according to the sample guidance scale; the first prediction data is generated by the first image generation model based on the conditional sample features, time step, and noisy image data determined by the conditional sample data in the training data; the second prediction data is generated by the first image generation model based on the unconditional sample features, the time step, and the noisy image data; the second noise prediction data is generated by the second generation model based on the sample fusion features, the time step, and the noisy image data; the sample fusion features are obtained by fusing the conditional sample features and the unconditional sample features according to the sample guidance scale.

13. A computing device, characterized in that, This includes processing components and storage components; The storage component stores a computer program; the computer program is invoked and executed by the processing component to implement the model training method as described in any one of claims 1 to 6, the data generation method as described in claim 7, the model training method as described in any one of claims 8 to 11, or the image generation method as described in claim 12.

14. A computer-readable storage medium, characterized in that, It stores a computer program, which, when executed by a processing component, implements the model training method as described in any one of claims 1 to 6, the data generation method as described in claim 7, the model training method as described in any one of claims 8 to 11, or the image generation method as described in claim 12.

15. A computer program product, characterized in that, It includes a computer program or instructions that, when executed by a processing component, implement the model training method as described in any one of claims 1 to 6, the data generation method as described in claim 7, the model training method as described in any one of claims 8 to 11, or the image generation method as described in claim 12.

Citation Information

Cited By

  • Image parallel reasoning method

    CN121168678A