Diffusion model optimization method and device based on semantic assistance and meta transfer learning, and medium
By embedding semantic information in the diffusion model and using a metatransfer learning framework, the diffusion process is optimized, and the problem of low efficiency in the generation of data related to input semantics of the diffusion model is solved, and rapid adaptation to new tasks and efficient generation is achieved.
Patent Information
- Application Number
- CN202510466194.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-18
AI Technical Summary
The diffusion model is inefficient and difficult to quickly adapt to new tasks when generating data related to input semantics, requiring a large amount of computing resources and data.
Embed semantic information through the cross attention mechanism, combines the metatransfer learning framework and DDIM optimization diffusion process, and uses the BERT model to extract text semantic features and run the diffusion process in the latent space.
The sampling efficiency and generation speed of the diffusion model are significantly improved, the semantic expression ability is enhanced, and the model can quickly adapt to new tasks.
Smart Images

Figure CN120338057A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning model optimization, and specifically to a diffusion model optimization method, device and medium based on semantic assistance and meta-transfer learning. Background Art
[0002] Diffusion models are a type of generative model that generates high-quality samples by simulating the inverse diffusion process of data from noise to real data. Although diffusion models have performed well in areas such as image generation and speech synthesis, their training process usually requires a lot of computing resources and data, and their generalization ability on different tasks is limited.
[0003] At the same time, traditional diffusion models have problems such as low sampling efficiency and difficulty in integrating semantic information. In addition, diffusion models require a large amount of data and computing resources when processing different tasks, and are difficult to quickly adapt to new tasks.
[0004] Therefore, how to enable the diffusion model to better understand and generate data related to the input semantics, while improving the sampling efficiency and enabling the diffusion model to quickly adapt to new tasks is a technical problem that needs to be solved urgently. Summary of the invention
[0005] The technical task of the present invention is to provide a diffusion model optimization method, device and medium based on semantic assistance and meta-transfer learning to solve the problem of how to enable the diffusion model to better understand and generate data related to the input semantics, while improving the sampling efficiency so that the diffusion model can quickly adapt to new tasks.
[0006] The technical task of the present invention is achieved in the following way: a diffusion model optimization method based on the combination of semantic assistance and meta-transfer learning, the method is as follows:
[0007] The semantic information is embedded into the generation process of the diffusion model through the cross-attention mechanism, and the interaction between the processing of semantic information and the diffusion model is realized through an independent semantic auxiliary module;
[0008] The meta-transfer learning framework uses the MAML algorithm to optimize the initial parameters of the diffusion model and the semantic auxiliary module;
[0009] The optimized diffusion sampling method of DDIM is used to optimize the diffusion process, and the diffusion process is run in the latent space through LDM to improve the generation efficiency.
[0010] Preferably, the semantic information is as follows:
[0011] Text preprocessing: preprocess the input text;
[0012] Extract text semantic features through the BERT model: Input the text description into the BERT model, and extract semantic embedding vectors related to the text semantics through the encoder structure of the BERT model. The semantic embedding vectors can capture the meanings of the words in the text, the structure of the sentences, and the semantic information of the context relationships;
[0013] Initialize the noise: In the diffusion model, the generation process starts from the noise and gradually denoises to generate the final samples; among them, the noise is randomly sampled from the Gaussian distribution; for each input sample, a corresponding noise vector is generated;
[0014] Noise feature processing: Use the noise vector directly as the input of the diffusion model. In the generator of the diffusion model, the noise vector will be gradually denoised, and finally samples close to the real data, that is, noise features, will be generated;
[0015] Fuse the semantic embedding vectors and the noise features: In the generator of the diffusion model, fuse the semantic embedding vectors and the noise features to obtain semantic information.
[0016] Preferably, the fusion of the semantic embedding vectors and the noise features is as follows:
[0017] Concatenate the semantic embedding vectors and the noise features: Concatenate the semantic embedding vectors and the noise features together as the input of the generator of the diffusion model;
[0018] Fuse through the cross-attention mechanism: Use the cross-attention mechanism to dynamically fuse the semantic embedding vectors and the noise features, so that the generated samples can better reflect the input text information.
[0019] Preferably, the diffusion model is a feed-forward neural network. The input of the diffusion model is the original data and the time step, and the output of the diffusion model is the reconstructed data; the output of the diffusion model is passed to the semantic auxiliary module for calculating the semantic loss;
[0020] The semantic auxiliary module is an independent neural network. The input of the semantic auxiliary module is the output generated by the diffusion model. The output generated by the diffusion model is the semantic classification result. The semantic auxiliary module is used to map the output generated by the diffusion model to the semantic space and optimize the semantic relevance of the output generated by the diffusion model through the classification loss;
[0021] Preferably, the diffusion model and the semantic auxiliary model are jointly trained as follows:
[0022] Construct a joint training dataset: Construct data for multiple related tasks. The data for each task includes input data and corresponding semantic labels; among them, for the image generation task, the input data is noise and the semantic label is the image category; for the text generation task, the input data is noise and the semantic label is the text category;
[0023] Initialize the diffusion model and the semantic assistance module for joint training;
[0024] Process multiple task data simultaneously and optimize the loss functions of multiple tasks as follows:
[0025] Data sampling: Randomly sample data from the dataset of each task;
[0026] Forward propagation: Input the sampled data into the diffusion model to generate reconstructed data;
[0027] Semantic assistance module: Input the generated reconstructed data into the semantic assistance module to calculate the semantic classification loss;
[0028] Calculate the total loss: Weightedly sum the reconstruction loss of the diffusion model and the classification loss of the semantic assistance module to obtain the total loss; The formula is as follows:
[0029] diffusion_loss = MSE(xrecon, x);
[0030] semantic_loss = CrossEntropy(semantic_module(xrecon), labels);
[0031] total_loss = diffusion_loss + λ · semantic_loss;
[0032] Among them, λ represents the weight of the classification loss, which is used to balance the reconstruction loss and the semantic loss;
[0033] Backward propagation: Perform backward propagation through the total loss to update the initial parameters of the diffusion model and the semantic assistance model.
[0034] Preferably, the meta-transfer learning framework uses the MAML algorithm to optimize the initial parameters of the diffusion model and the semantic assistance module as follows:
[0035] ①Construct few-shot tasks: Randomly sample a few-shot task from the dataset of each task, and each task contains a support set and a query set;
[0036] ②Support set training: Rapidly adapt the model on the support set, and update and optimize the parameters of the diffusion model and the semantic assistance module by calculating the loss on the support set, so that the optimized diffusion model and the semantic assistance module adapt to the current task;
[0037] ③Query set evaluation: Evaluate the performance of the model on the query set, calculate the loss on the query set, and use the loss on the query set to update and optimize the initial parameters of the diffusion model and the semantic assistance module;
[0038] ④ Meta-update: By repeating steps ① to ③ on multiple few-shot tasks, optimize the initial parameters of the model so that the diffusion model and the semantic assistance module can converge quickly on new tasks.
[0039] Preferably, the text preprocessing is as follows:
[0040] Tokenization: Split the text into word or word unit;
[0041] Encoding: Convert the tokenization result into the corresponding embedding vector;
[0042] During the text preprocessing, the input text is input_text, and the output of the BERT model is outputs, and the format is as follows:
[0043] inputs = tokenizer(input_text) outputs = BERT(inputs) text_embeddings = outputs.last_hidden_state[:, 0, :];
[0044] Among them, tokenizer is the BERT tokenizer, which converts the text into the input format that the model can process;
[0045] BERT is the pre-trained BERT model, which is used to extract the semantic features of the text; text_embeddings is the extracted semantic embedding vector, using the embedding vector of the [CLS] token of BERT;
[0046] Fusing the semantic embedding vector and the noise feature means concatenating the semantic embedding vector text_embeddings and the noise feature noise together as the input combined_input of the diffusion module generator, and the format is as follows:
[0047] combined_input = [noise, text_embeddings].
[0048] Preferably, using the DDIM optimization diffusion sampling method to optimize the diffusion process means that DDIM reduces the sampling steps by introducing latent variables and implicit update steps, and the formula is as follows:
[0049]
[0050] Among them, a t represents the scaling factor at time step t; z t represents the latent variable, which is used to generate the sample of the next time step;
[0051] Running the diffusion process in the latent space through LDM means that LDM maps the data to a low-dimensional latent space and performs the diffusion process in the latent space. The formula is as follows:
[0052] z = Encoder(x);
[0053] zrecon = Diffusion(z);
[0054] xrecon = Decoder(zrecon);
[0055] Among them, Encoder represents the encoder that maps the input data to a low-dimensional latent space; Diffusion represents the diffusion model running in the latent space; Decoder represents the decoder that decodes the latent variable back to the original data space.
[0056] An electronic device includes: a memory and at least one processor;
[0057] Among them, a computer program is stored on the memory;
[0058] The at least one processor executes the computer program stored in the memory, so that the at least one processor executes the diffusion model optimization method based on the combination of semantic assistance and meta-transfer learning as described above.
[0059] A computer-readable storage medium stores a computer program, and the computer program can be executed by a processor to implement the diffusion model optimization method based on the combination of semantic assistance and meta-transfer learning as described above.
[0060] The diffusion model optimization method, device and medium based on semantic assistance and meta-transfer learning of the present invention have the following advantages:
[0061] (1) By optimizing the diffusion process and the meta-learning strategy, the present invention significantly reduces the sampling steps, speeds up the generation speed, and improves the sampling efficiency;
[0062] (2) Through the semantic assistance module, the present invention enables the diffusion model to better understand and generate data related to the input semantics, enhancing the semantic expression;
[0063] (3) Through the meta-transfer learning framework, the present invention enables the diffusion model to quickly adapt to new tasks, reduces the demand for task data volume, and enables quick task adaptation. Description of the Drawings
[0064] The present invention will be further described below with reference to the drawings.
[0065] Att Figure 1 is a flowchart of the diffusion model optimization method based on semantic assistance and meta-transfer learning. Detailed implementation manners
[0066] The optimization method, device and medium of the diffusion model based on semantic assistance and meta-transfer learning of the present invention will be described in detail below with reference to the accompanying drawings of the specification and specific embodiments.
[0067] Embodiment 1:
[0068] As shown in the attached Figure 1 This embodiment provides an optimization method for a diffusion model based on semantic assistance and meta-transfer learning, and the method is as follows:
[0069] S1. Embed semantic information into the generation process of the diffusion model through a cross-attention mechanism, and implement the processing of semantic information and the interaction with the diffusion model through an independent semantic assistance module;
[0070] S2. The meta-transfer learning framework uses the MAML algorithm to optimize the initial parameters of the diffusion model and the semantic assistance module;
[0071] S3. Use the optimized diffusion sampling method of DDIM to optimize the diffusion process, and run the diffusion process in the latent space through LDM to improve the generation efficiency.
[0072] Among them, a cross-attention mechanism is added to the generator of the diffusion model. The cross-attention mechanism is a technology that enables the model to better integrate different modal information (such as text and images); specifically, when implemented, the extracted semantic feature vector is used as a query (Query) or key (Key) and input into the cross-attention mechanism, and at the same time, the noise feature in the generation process of the diffusion model is used as the value (Value) and input. The cross-attention mechanism fuses the semantic information with the noise feature by calculating the similarity between the query, key and value (usually using the dot product attention mechanism). In this way, the diffusion model can use semantic information to guide the generation process during the generation process, making the generated samples more conform to the input semantic description.
[0073] The semantic information in step S1 of this embodiment is specifically as follows:
[0074] S101. Text preprocessing: Preprocess the input text;
[0075] S102. Extract text semantic features through the BERT model: Input the text description into the BERT model, and extract the semantic embedding vector related to the text semantics through the encoder structure of the BERT model. The semantic embedding vector can capture the semantic information of the meaning of the words in the text, the structure of the sentence, and the context relationship.
[0076] S103. Initialize noise: In the diffusion model, the generation process starts from noise and gradually denoises to generate the final samples. Among them, the noise is randomly sampled from a Gaussian distribution. For each input sample, a corresponding noise vector is generated.
[0077] S104. Noise feature processing: The noise vector is directly used as the input of the diffusion model. In the generator of the diffusion model, the noise vector will be gradually denoised, and finally samples close to the real data, that is, noise features, will be generated.
[0078] S105. Fuse the semantic embedding vector and the noise feature: In the generator of the diffusion model, the semantic embedding vector and the noise feature are fused to obtain semantic information.
[0079] The specific process of fusing the semantic embedding vector and the noise feature in step S105 of this embodiment is as follows:
[0080] S10501. Concatenate the semantic embedding vector and the noise feature: The semantic embedding vector and the noise feature are concatenated together as the input of the generator of the diffusion model.
[0081] S10502. Fuse through the cross-attention mechanism: The cross-attention mechanism is used to dynamically fuse the semantic embedding vector and the noise feature, so that the generated samples can better reflect the input text information.
[0082] The diffusion model in this embodiment is a feed-forward neural network. The input of the diffusion model is the original data and the time step, and the output of the diffusion model is the reconstructed data. The output of the diffusion model is passed to the semantic auxiliary module for calculating the semantic loss.
[0083] The semantic auxiliary module in this embodiment is an independent neural network. The input of the semantic auxiliary module is the output generated by the diffusion model. The output generated by the diffusion model is the semantic classification result. The semantic auxiliary module is used to map the output generated by the diffusion model to the semantic space and optimize the semantic relevance of the output generated by the diffusion model through the classification loss.
[0084] The diffusion model and the semantic auxiliary model in this embodiment are jointly trained as follows:
[0085] (1) Construct a joint training dataset: Construct data for multiple related tasks. The data for each task includes input data and the corresponding semantic label. Among them, for the image generation task, the input data is noise and the semantic label is the image category. For the text generation task, the input data is noise and the semantic label is the text category.
[0086] (2) Initialize the diffusion model and the semantic auxiliary module for joint training.
[0087] (3) Process multiple task data simultaneously and optimize the loss functions of multiple tasks, specifically as follows:
[0088] ① Data sampling: Randomly sample data from the datasets of each task;
[0089] ② Forward propagation: Input the sampled data into the diffusion model to generate reconstructed data;
[0090] ③ Semantic assistance module: Input the generated reconstructed data into the semantic assistance module to calculate the semantic classification loss;
[0091] ④ Calculate the total loss: Weightedly sum the reconstruction loss of the diffusion model and the classification loss of the semantic assistance module to obtain the total loss; The formula is as follows:
[0092] diffusion_loss = MSE(xrecon, x);
[0093] semantic_loss = CrossEntropy(semantic_module(xrecon), labels);
[0094] total_loss = diffusion_loss + λ · semantic_loss;
[0095] Among them, λ represents the weight of the classification loss, which is used to balance the reconstruction loss and the semantic loss;
[0096] ⑤ Backward propagation: Perform backward propagation through the total loss to update the initial parameters of the diffusion model and the semantic assistance model.
[0097] In this embodiment, the meta-transfer learning framework in step S2 uses the MAML algorithm to optimize the initial parameters of the diffusion model and the semantic assistance module, specifically as follows:
[0098] ① Construct few-shot tasks: Randomly sample a few-shot task from the datasets of each task, and each task contains a support set and a query set;
[0099] ② Support set training: Rapidly adapt the model on the support set, and update and optimize the parameters of the diffusion model and the semantic assistance module by calculating the loss on the support set, so that the optimized diffusion model and the semantic assistance module adapt to the current task;
[0100] ③ Query set evaluation: Evaluate the performance of the model on the query set, calculate the loss on the query set, and use the loss on the query set to update and optimize the initial parameters of the diffusion model and the semantic assistance module;
[0101] ④Meta-update: By repeating steps ① to ③ on multiple few-shot tasks, the initial parameters of the model are optimized so that the diffusion model and the semantic assistance module can converge quickly on new tasks.
[0102] The text preprocessing in step S101 of this embodiment is specifically as follows:
[0103] S10101. Word segmentation: The text is segmented into word or word units;
[0104] S10102. Encoding: The word segmentation result is converted into a corresponding embedding vector;
[0105] During the text preprocessing, the input text is input_text, and the output of the BERT model is outputs, with the format as follows:
[0106] inputs = tokenizer(input_text) outputs = BERT(inputs) text_embeddings = outputs.last_hidden_state[:, 0, :];
[0107] Among them, tokenizer is the BERT tokenizer, which converts the text into the input format that the model can process;
[0108] BERT is the pre-trained BERT model, which is used to extract the semantic features of the text; text_embeddings is the extracted semantic embedding vector, using the embedding vector of the [CLS] token of BERT;
[0109] Fusing the semantic embedding vector and the noise feature means splicing the semantic embedding vector text_embeddings and the noise feature noise together as the input combined_input of the diffusion module generator, with the format as follows:
[0110] combined_input = [noise, text_embeddings].
[0111] In this embodiment, the diffusion process is optimized by using the DDIM optimized diffusion sampling method in step S3, which means that DDIM reduces the sampling steps by introducing latent variables and implicit update steps, and the formula is as follows:
[0112]
[0113] where, a t represents the scaling factor at time step t; z t represents the latent variable, which is used to generate the sample of the next time step.
[0114] The operation of the diffusion process in the latent space through LDM in step S3 of this embodiment means that LDM maps the data to a low-dimensional latent space and performs the diffusion process in the latent space. The formula is as follows:
[0115] z = Encoder(x);
[0116] zrecon = Diffusion(z);
[0117] xrecon = Decoder(zrecon);
[0118] Among them, Encoder represents the encoder that maps the input data to a low-dimensional latent space; Diffusion represents the diffusion model operating in the latent space; Decoder represents the decoder that decodes the latent variable back to the original data space.
[0119] DDIM (Denoising Diffusion Implicit Model) is as follows:
[0120] (1) Reducing the sampling steps: Traditional diffusion models (such as DDPM, Denoising Diffusion Probability Model) need to gradually denoise during the generation process, usually requiring a large number of sampling steps (such as 1000 steps). DDIM significantly reduces the sampling steps by introducing implicit variables and implicit update steps, thereby improving the generation efficiency.
[0121] (2) Improving the generation quality: DDIM further improves the quality of the generated samples by optimizing the noise schedule in the diffusion process. Optimization of the noise schedule: DDIM allows for a more flexible noise schedule strategy. By adjusting the distribution and step size of the noise, higher-quality samples can be generated. For example, DDIM can use a non-linear noise schedule, such that the noise decreases faster in the early steps and slower in the later steps, thus better balancing the generation quality and the denoising speed.
[0122] LDM (Latent Diffusion Model) is as follows:
[0123] (1) Diffusion process in the latent space: The core idea of LDM is to map the data to a low-dimensional latent space, perform the diffusion process in the latent space, and then decode the generated latent variable back to the original data space. This method significantly improves the generation efficiency and generation quality.
[0124] Introduction of the latent space: LDM first uses an encoder (such as an autoencoder or variational autoencoder) to map the input data to a low-dimensional latent space. In the latent space, the representation of the data is more compact, and the denoising process is also more efficient.
[0125] For example, for image generation tasks, LDM can map a high-resolution image to a low-dimensional latent vector and then perform the diffusion process in the latent space.
[0126] Diffusion process in the latent space: In the latent space, LDM uses a diffusion process similar to DDPM or DDIM, but operates on low-dimensional latent variables. Since the dimension of the latent space is low, the denoising process is more efficient and the generation speed is faster. After generation is completed, the decoder is used to decode the latent variables back to the original data space to obtain the final generated sample.
[0127] (2) Multimodal generation ability: LDM not only performs well in single-modal tasks (such as image generation), but also can handle multimodal tasks (such as text-image generation).
[0128] Multimodal generation: LDM can guide the generation process by introducing multimodal conditional information (such as text descriptions). For example, in text-image generation tasks, LDM can encode the text description as a latent variable and fuse it with the latent variable of the image in the latent space to generate an image related to the text description.
[0129] This multimodal generation ability makes LDM more flexible and adaptable in complex generation tasks.
[0130] (3) Multi-task learning and transfer: By jointly training on multiple related tasks, the model learns general semantic features and generation strategies, and transfers this knowledge to new tasks through a meta-learning mechanism.
[0131] Example 2:
[0132] This embodiment also provides an electronic device, including: a memory and a processor;
[0133] Wherein, the memory stores computer-executable instructions;
[0134] The processor executes the computer-executable instructions stored in the memory, so that the processor executes the diffusion model optimization method based on semantic assistance and meta-transfer learning in any embodiment of the present invention.
[0135] The processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0136] The memory can be used to store computer programs and / or modules. By running or executing the computer programs and / or modules stored in the memory and calling the data stored in the memory, the processor can implement various functions of the electronic device. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function, etc.; the data storage area can store data created according to the use of the terminal, etc. In addition, the memory can also include high-speed random access memory, and can also include non-volatile memory, such as hard disks, memory, plug-in hard disks, smart media cards (SMC), secure digital (SD) cards, flash memory cards, at least one magnetic disk storage period, flash memory devices, or other volatile solid-state storage devices.
[0137] Embodiment 3:
[0138] This embodiment also provides a computer-readable storage medium, which stores multiple instructions. The instructions are loaded by the processor to enable the processor to execute the diffusion model optimization method based on semantic assistance and meta-transfer learning in any embodiment of the present invention. Specifically, a system or device equipped with a storage medium can be provided. On this storage medium, software program codes for implementing the functions of any one of the above embodiments are stored, and the computer (or CPU or MPU) of the system or device is made to read and execute the program codes stored in the storage medium.
[0139] In this case, the program code read from the storage medium itself can implement the functions of any one of the above embodiments. Therefore, the program code and the storage medium storing the program code constitute a part of the present invention.
[0140] Embodiments of the storage medium for providing program codes include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Optionally, the program code can be downloaded from a server computer via a communication network.
[0141] In addition, it should be clear that not only can the actual operations be completed in part or in whole by executing the program code read by the computer, but also by the operating system etc. operating on the computer based on the instructions of the program code, so as to implement the functions of any one of the above embodiments.
[0142] In addition, it can be understood that the program code read from the storage medium is written into the memory provided in the expansion board inserted into the computer or into the memory provided in the expansion unit connected to the computer, and then based on the instructions of the program code, the CPU or the like installed on the expansion board or the expansion unit is made to execute some or all of the actual operations, thereby implementing the functions of any one of the above embodiments.
[0143] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An optimization method for a diffusion model based on the combination of semantic assistance and meta-transfer learning, characterized in that, The method is as follows: Embed semantic information into the generation process of the diffusion model through the cross-attention mechanism, and implement the processing of semantic information and the interaction with the diffusion model through an independent semantic assistance module; The meta-transfer learning framework uses the MAML algorithm to optimize the initial parameters of the diffusion model and the semantic assistance module; Adopt the optimized diffusion sampling method of DDIM to optimize the diffusion process, and run the diffusion process in the latent space through LDM.
2. The diffusion model optimization method based on the combination of semantic assistance and meta-transfer learning according to claim 1, wherein The semantic information is as follows: Text preprocessing: Preprocess the input text; Extract text semantic features through the BERT model: Input the text description into the BERT model, and extract semantic embedding vectors related to the text semantics through the encoder structure of the BERT model. The semantic embedding vectors can capture the semantic information of the meanings of words, sentence structures, and context relationships in the text; Initialize noise: In the diffusion model, the generation process starts from noise and gradually denoises to generate the final sample; among them, the noise is randomly sampled from a Gaussian distribution; for each input sample, a corresponding noise vector is generated; Noise feature processing: Directly use the noise vector as the input of the diffusion model. In the generator of the diffusion model, the noise vector will be gradually denoised, and finally a sample close to the real data, that is, the noise feature, will be generated. Fuse the semantic embedding vector and the noise feature: In the generator of the diffusion model, fuse the semantic embedding vector and the noise feature to obtain semantic information.
3. The optimization method of the diffusion model based on the combination of semantic assistance and meta-transfer learning according to claim 2, wherein The fusion of the semantic embedding vector and the noise feature is as follows: Concatenate the semantic embedding vector and the noise feature: Concatenate the semantic embedding vector and the noise feature together as the input of the diffusion model generator; Fuse through the cross-attention mechanism: Use the cross-attention mechanism to dynamically fuse the semantic embedding vector and the noise feature, so that the generated sample can better reflect the input text information.
4. The diffusion model optimization method based on the combination of semantic assistance and meta-transfer learning according to claim 2, characterized in that, The diffusion model is a feed-forward neural network. The input of the diffusion model is the original data and the time step, and the output of the diffusion model is the reconstructed data; the output of the diffusion model is passed to the semantic assistance module for calculating the semantic loss; The semantic assistance module is an independent neural network. The input of the semantic assistance module is the output generated by the diffusion model. The output generated by the diffusion model is the semantic classification result. The semantic assistance module is used to map the output generated by the diffusion model to the semantic space and optimize the semantic relevance of the output generated by the diffusion model through the classification loss.
5. The diffusion model optimization method based on the combination of semantic assistance and meta-transfer learning according to claim 4, characterized in that, The diffusion model and the semantic assistance model are jointly trained as follows: Construct a joint training dataset: Construct data for multiple related tasks. The data for each task includes input data and corresponding semantic labels; among them, for the image generation task, the input data is noise and the semantic label is the image category; for the text generation task, the input data is noise and the semantic label is the text category; Initialize the diffusion model and the semantic assistance module for joint training; Process multiple task data simultaneously and optimize the loss functions of multiple tasks as follows: Data sampling: Randomly sample data from the dataset of each task; Forward propagation: Input the sampled data into the diffusion model to generate reconstructed data; Semantic assistance module: Input the generated reconstructed data into the semantic assistance module to calculate the semantic classification loss; Calculate the total loss: Weighted sum the reconstruction loss of the diffusion model and the classification loss of the semantic assistance module to obtain the total loss; The formula is as follows: diffusion_loss = MSE(xrecon, x); semantic_loss = CrossEntropy(semantic_module(xrecon), labels); total_loss = diffusion_loss + λ · semantic_loss; Among them, λ represents the weight of the classification loss, which is used to balance the reconstruction loss and the semantic loss; Backpropagation: Perform backpropagation through the total loss to update the initial parameters of the diffusion model and the semantic assistance model.
6. The diffusion model optimization method based on the combination of semantic assistance and meta-transfer learning according to claim 1, characterized in that, The meta-transfer learning framework uses the MAML algorithm to optimize the initial parameters of the diffusion model and the semantic assistance module as follows: ①Construct few-shot tasks: Randomly sample a few-shot task from the dataset of each task. Each task contains a support set and a query set; ②Support set training: Quickly adapt the model on the support set. By calculating the loss on the support set and updating and optimizing the parameters of the diffusion model and the semantic assistance module, make the optimized diffusion model and semantic assistance module adapt to the current task; ③Query set evaluation: Evaluate the performance of the model on the query set, calculate the loss on the query set, and use the loss on the query set to update and optimize the initial parameters of the diffusion model and the semantic assistance module; ④Meta-update: By repeating steps ① to ③ on multiple few-shot tasks, optimize the initial parameters of the model, so that the diffusion model and the semantic assistance module can quickly converge on new tasks.
7. The optimization method of the diffusion model based on the combination of semantic assistance and meta-transfer learning according to claim 2, wherein The text preprocessing is as follows: Tokenization: Split the text into word or word unit; Encoding: Convert the tokenization result into the corresponding embedding vector; During the text preprocessing, the input text is input_text, and the output of the BERT model is outputs. The format is as follows: inputs = tokenizer(input_text) outputs = BERT(inputs) text_embeddings = outputs.last_hidden_state[:, 0, :]; Among them, tokenizer is the BERT tokenizer, which converts the text into the input format that the model can process; BERT is the pre-trained BERT model, which is used to extract the semantic features of the text; text_embeddings is the extracted semantic embedding vector, using the embedding vector of the [CLS] token of BERT; Fusing the semantic embedding vector and the noise feature means concatenating the semantic embedding vector text_embeddings and the noise feature noise together as the input combined_input of the generator of the diffusion module. The format is as follows: combined_input = [noise, text_embeddings].
8. The diffusion model optimization method based on the combination of semantic assistance and meta-transfer learning according to claim 1, characterized in that Optimizing the diffusion process using the DDIM optimized diffusion sampling method means that DDIM reduces the sampling steps by introducing latent variables and implicit update steps. The formula is as follows: where a t represents the scaling factor at time step t; z t represents the latent variable used to generate samples for the next time step; Running the diffusion process in the latent space through LDM means that LDM maps the data to a low-dimensional latent space and performs the diffusion process in the latent space. The formula is as follows: z = Encoder(x); zrecon = Diffusion(z); xrecon = Decoder(zrecon); Among them, Encoder represents the encoder that maps the input data to a low-dimensional latent space; Diffusion represents the diffusion model running in the latent space; Decoder represents the decoder that decodes the latent variable back to the original data space.
9. An electronic device, characterized in that, Including: A memory and at least one processor; Among them, a computer program is stored on the memory; The at least one processor executes the computer program stored in the memory, so that the at least one processor executes the diffusion model optimization method based on the combination of semantic assistance and meta-transfer learning as described in any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and the computer program can be executed by a processor to implement the diffusion model optimization method based on the combination of semantic assistance and meta-transfer learning as described in any one of claims 1 to 8.