Image generation method, device and equipment based on reinforcement learning and diffusion model

By introducing reinforcement learning and dynamic word generation technology into the diffusion model, combined with U-Net network and gated unit, the problem that diffusion model is difficult to generate target images with rich detailed information is solved, and higher image generation reliability and stability are achieved.

CN119478108BActive Publication Date: 2025-05-06XIANGJIANG LAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510059322.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-06
Estimated Expiration
2045-01-15

AI Technical Summary

Technical Problem

The existing diffusion model is difficult to generate target images with rich details, resulting in insufficient reliability of generated images.

Method used

The image generation method based on reinforcement learning and diffusion model is adopted, and the image generation is improved by obtaining preset text description information, input image and edge information, feature extraction and dynamic word generation are performed, combined with the U-Net network and gating unit, the initial image is generated, and the diffusion model is trained to minimize the loss value and maximize the reward value to improve the reliability of image generation.

Benefits of technology

The reliability and stability of the target image generated by the diffusion model is improved, and the detailed information capture capability of image generation is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119478108B_ABST
    Figure CN119478108B_ABST
Patent Text Reader

Abstract

The present invention relates to the fields of artificial intelligence technology and image technology, and discloses an image generation method, device and equipment based on reinforcement learning and diffusion model, the method comprising: obtaining preset text description information, preset input image and edge information of the preset input image; performing feature extraction on the preset text description information to obtain a first text embedding vector, performing feature extraction on the preset input image to obtain an image embedding vector, performing feature extraction on the edge information of the preset input image to obtain an edge information embedding vector; determining a second text embedding vector according to the first text embedding vector and the image embedding vector; processing the second text embedding vector through a diffusion model to obtain an initial image, determining a generated image according to the initial image and the edge information embedding vector; and obtaining a current target image output by the trained diffusion model. The present invention is conducive to improving the reliability of the current target image generated by the diffusion model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of artificial intelligence technology and image technology, and in particular to an image generation method, device and equipment based on reinforcement learning and diffusion model. Background Art

[0002] In recent years, deep learning technology has achieved remarkable results in the field of generative networks. Among them, the diffusion model, as an emerging generative network modeling method, has demonstrated excellent performance. Compared with traditional generative networks, the diffusion model performs better in generation quality and stability.

[0003] However, the existing diffusion model cannot learn detailed information, resulting in the current target image generated by the diffusion model being unreliable. Therefore, how to improve the reliability of the current target image generated by the diffusion model is an urgent problem to be solved. Summary of the invention

[0004] The present invention provides an image generation method, device, computer equipment and storage medium based on reinforcement learning and diffusion model to solve the technical problem of how to improve the reliability of the current target image generated by the diffusion model.

[0005] In a first aspect, a method for generating an image based on reinforcement learning and a diffusion model is provided, comprising:

[0006] Acquire preset text description information, a preset input image, and edge information of the preset input image;

[0007] Performing feature extraction on the preset text description information to obtain a first text embedding vector, performing feature extraction on the preset input image to obtain an image embedding vector, and performing feature extraction on edge information of the preset input image to obtain an edge information embedding vector;

[0008] Acquire a dynamic word according to the first text embedding vector and the image embedding vector, and determine a second text embedding vector based on the dynamic word and a predefined determination method;

[0009] Processing the second text embedding vector by a diffusion model to obtain an initial image, and determining a generated image according to the initial image and the edge information embedding vector;

[0010] With the goal of minimizing the loss value between the generated image and the preset target image and maximizing the reward value in reinforcement learning, obtaining the trained diffusion model;

[0011] The current text description information, the current input image and the edge information of the current input image are obtained, the current text description information, the current input image and the edge information of the current input image are input into the trained diffusion model, and the current target image output by the trained diffusion model is obtained.

[0012] Further, the performing feature extraction on the preset text description information to obtain a first text embedding vector, performing feature extraction on the preset input image to obtain an image embedding vector, and performing feature extraction on edge information of the preset input image to obtain an edge information embedding vector includes:

[0013] Perform feature extraction on the preset text description information through a text encoder of the CLIP model to obtain an embedding vector corresponding to the preset text description information, and select the embedding vector corresponding to the preset text description information as a first text embedding vector;

[0014] The image encoder of the CLIP model is used to perform feature extraction on the preset input image to obtain an embedding vector corresponding to the preset input image, and the embedding vector corresponding to the preset input image is selected as the image embedding vector; the image encoder of the CLIP model is used to perform feature extraction on the edge information to obtain an embedding vector corresponding to the edge information, and the embedding vector corresponding to the edge information is selected as the edge information embedding vector.

[0015] Further, acquiring a dynamic word according to the first text embedding vector and the image embedding vector, and determining a second text embedding vector based on the dynamic word and a predefined determination method includes:

[0016] Inputting the first text embedding vector and the image embedding vector into a preset adaptive neural fuzzy inference system, obtaining a dynamic word output by the adaptive neural fuzzy inference system, adding the dynamic word to the preset text description information, and obtaining the modified preset text description information;

[0017] The text encoder of the CLIP model is used to perform feature extraction on the modified preset text description information to obtain an embedding vector corresponding to the modified preset text description information, and the embedding vector corresponding to the modified preset text description information is selected as the second text embedding vector.

[0018] Furthermore, the step of processing the second text embedding vector by a diffusion model to obtain an initial image, and determining a generated image according to the initial image and the edge information embedding vector, comprises:

[0019] Inputting the second text embedding vector into a U-Net network in a diffusion model, and processing the second text embedding vector through the U-Net network to obtain an initial image;

[0020] The initial image and the edge information embedding vector are input into a gating unit, and the initial image and the edge information embedding vector are fused through a gating mechanism in the gating unit to obtain an enhanced initial image, and the enhanced initial image is selected as a generated image output by the diffusion model.

[0021] Furthermore, the method of obtaining the trained diffusion model with the goal of minimizing the loss value between the generated image and the preset target image and maximizing the reward value in reinforcement learning includes:

[0022] Obtaining a preset target image corresponding to the preset text description information, and calculating a current loss value between the generated image and the preset target image using a preset loss function;

[0023] Obtain a reward function corresponding to the diffusion model, optimize the parameters of the diffusion model using a reinforcement learning algorithm, obtain a reward value of the reward function during the optimization of the parameters, subtract the reward value from the current loss value to generate a total loss value, train the diffusion model based on the total loss value, stop training the diffusion model when the total loss value converges, and obtain the trained diffusion model.

[0024] Furthermore, the step of obtaining the current text description information, the current input image, and the edge information of the current input image, inputting the current text description information, the current input image, and the edge information of the current input image into the trained diffusion model, and obtaining the current target image output by the trained diffusion model includes:

[0025] Obtaining current text description information and a current input image, and obtaining edge information of the current input image through a Canny edge detection algorithm;

[0026] The current text description information, the current input image and the edge information of the current input image are input into the trained diffusion model to obtain the current target image output by the trained diffusion model.

[0027] Furthermore, after obtaining the current text description information, the current input image and the edge information of the current input image, inputting the current text description information, the current input image and the edge information of the current input image into the trained diffusion model, and obtaining the current target image output by the trained diffusion model, the image generation method includes:

[0028] A display window is created, and the current target image is displayed through the display window.

[0029] Furthermore, after obtaining the current text description information, the current input image and the edge information of the current input image, inputting the current text description information, the current input image and the edge information of the current input image into the trained diffusion model, and obtaining the current target image output by the trained diffusion model, the image generation method includes:

[0030] A push instruction is obtained, and the push instruction is executed to push the current target image to the target system.

[0031] In a second aspect, an image generation device based on reinforcement learning and diffusion model is provided, comprising:

[0032] A first acquisition module, used to acquire preset text description information, a preset input image and edge information of the preset input image;

[0033] An extraction module, configured to perform feature extraction on the preset text description information to obtain a first text embedding vector, perform feature extraction on the preset input image to obtain an image embedding vector, and perform feature extraction on edge information of the preset input image to obtain an edge information embedding vector;

[0034] A second acquisition module is used to acquire a dynamic word according to the first text embedding vector and the image embedding vector, and determine a second text embedding vector based on the dynamic word and a predefined determination method;

[0035] A first generating module, configured to process the second text embedding vector through a diffusion model to obtain an initial image, and determine a generated image according to the initial image and the edge information embedding vector;

[0036] A training module, used to obtain the trained diffusion model with the goal of minimizing the loss value between the generated image and the preset target image and maximizing the reward value in reinforcement learning;

[0037] The second generation module is used to obtain the current text description information, the current input image and the edge information of the current input image, input the current text description information, the current input image and the edge information of the current input image into the trained diffusion model, and obtain the current target image output by the trained diffusion model.

[0038] In a third aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned image generation method when executing the computer program.

[0039] In a fourth aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned image generation method are implemented.

[0040] The present application provides an image generation method, device, computer equipment and storage medium based on reinforcement learning and diffusion model, which obtain preset text description information, preset input image and edge information of the preset input image; perform feature extraction on the preset text description information to obtain a first text embedding vector, perform feature extraction on the preset input image to obtain an image embedding vector, perform feature extraction on the edge information of the preset input image to obtain an edge information embedding vector; obtain dynamic words according to the first text embedding vector and the image embedding vector, and determine a second text embedding vector based on the dynamic words and a predefined determination method; process the second text embedding vector through a diffusion model to obtain an initial image, and determine a generated image according to the initial image and the edge information embedding vector; obtain the trained diffusion model with the goal of minimizing the loss value between the generated image and the preset target image and maximizing the reward value in reinforcement learning; obtain the current text description information, the current input The trained diffusion model obtains the edge information of the current input image and the current input image, inputs the current text description information, the current input image and the edge information of the current input image into the trained diffusion model, and obtains the current target image output by the trained diffusion model. The beneficial effects are in two aspects. On the one hand, the current text description information, the current input image and the edge information of the current input image are obtained, and the current text description information, the current input image and the edge information of the current input image are input into the trained diffusion model, and the current target image output by the trained diffusion model is obtained. Since the edge information of the current input image can provide the detail information of the current input image, the trained diffusion model can capture the detail information of the current input image, which is beneficial to improve the reliability of the current target image output by the trained diffusion model. On the other hand, since the trained diffusion model will not be affected by human intervention, it is beneficial to improve the stability of the acquired current target image. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative labor.

[0042] Figure 1 is a schematic diagram of an application environment of an image generating method according to an embodiment of the present invention;

[0043] Figure 2 A schematic diagram of a flow chart of an image generating method provided by an embodiment of the present invention;

[0044] Figure 3 yes Figure 2 A schematic flow chart of a specific implementation of step S23;

[0045] Figure 4 yes Figure 2 A schematic flow chart of a specific implementation of step S25;

[0046] Figure 5 yes Figure 2 A schematic flow chart of a specific implementation of step S26;

[0047] Figure 6 is a structural schematic diagram of an image generating device in one embodiment of the present invention;

[0048] Figure 7 It is a structural diagram of a computer device in one embodiment of the present invention. DETAILED DESCRIPTION

[0049] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0050] See also Figure 1 , Figure 1 FIG. 1 is a schematic diagram of an application environment of an image generation method according to an embodiment of the present invention. The image generation method provided by the embodiment of the present invention can be applied in the following embodiments: Figure 1 In an application environment, a client communicates with a server through a network.

[0051] The server obtains preset text description information, a preset input image and edge information of the preset input image through the client;

[0052] Performing feature extraction on the preset text description information to obtain a first text embedding vector, performing feature extraction on the preset input image to obtain an image embedding vector, and performing feature extraction on edge information of the preset input image to obtain an edge information embedding vector;

[0053] Acquire a dynamic word according to the first text embedding vector and the image embedding vector, and determine a second text embedding vector based on the dynamic word and a predefined determination method;

[0054] Processing the second text embedding vector by a diffusion model to obtain an initial image, and determining a generated image according to the initial image and the edge information embedding vector;

[0055] With the goal of minimizing the loss value between the generated image and the preset target image and maximizing the reward value in reinforcement learning, obtaining the trained diffusion model;

[0056] The current text description information, the current input image and the edge information of the current input image are obtained, the current text description information, the current input image and the edge information of the current input image are input into the trained diffusion model, and the current target image output by the trained diffusion model is obtained.

[0057] In the scheme implemented by the above-mentioned image generation method, device, equipment and medium, the beneficial effects are in two aspects. On the one hand, the current text description information, the current input image and the edge information of the current input image are obtained, and the current text description information, the current input image and the edge information of the current input image are input into the trained diffusion model to obtain the current target image output by the trained diffusion model. Since the edge information of the current input image can provide the detail information of the current input image, the trained diffusion model can capture the detail information of the current input image, which is beneficial to improve the reliability of the current target image output by the trained diffusion model. On the other hand, since the trained diffusion model will not be affected by human intervention, it is beneficial to improve the stability of the acquired current target image.

[0058] The device running the client is referred to as a client device.

[0059] The device running the server is referred to as the server device.

[0060] Among them, client devices include but are not limited to smartphones, personal computers, Internet of Vehicles terminals, tablets and portable wearable devices.

[0061] The server device may be implemented by an independent server or a server cluster composed of multiple servers. The present invention is described in detail below through specific embodiments. Figure 2 , Figure 2 A flowchart of an image generation method provided by an embodiment of the present invention includes the following steps:

[0062] S21, obtaining preset text description information, a preset input image, and edge information of the preset input image;

[0063] Exemplarily, obtaining preset text description information, a preset input image, and edge information of the preset input image includes:

[0064] Get preset text description information and preset input image;

[0065] The edge information of the preset input image is obtained by using the Canny edge detection algorithm.

[0066] Among them, the Chinese name of the Canny edge detection algorithm is: Canny edge detection algorithm, and the English name of the Canny edge detection algorithm is Canny Edge Detection.

[0067] Among them, the Canny edge detection algorithm can accurately detect edge information in an image through multiple steps such as Gaussian filtering, gradient calculation, non-maximum suppression, dual threshold detection, edge tracking and connection, and has a certain robustness to noise. Its detection results are usually presented in the form of a binary image with clear and continuous edges, providing strong support for subsequent image analysis and processing.

[0068] S22, performing feature extraction on the preset text description information to obtain a first text embedding vector, performing feature extraction on the preset input image to obtain an image embedding vector, and performing feature extraction on edge information of the preset input image to obtain an edge information embedding vector;

[0069] The step of extracting features from the preset text description information to obtain a first text embedding vector, extracting features from the preset input image to obtain an image embedding vector, and extracting features from edge information of the preset input image to obtain an edge information embedding vector includes:

[0070] Perform feature extraction on the preset text description information through a text encoder of the CLIP model to obtain an embedding vector corresponding to the preset text description information, and select the embedding vector corresponding to the preset text description information as a first text embedding vector;

[0071] Performing feature extraction on the preset input image through the image encoder of the CLIP model to obtain an embedding vector corresponding to the preset input image, and selecting the embedding vector corresponding to the preset input image as the image embedding vector;

[0072] The image encoder of the CLIP model is used to extract features of the edge information to obtain an embedding vector corresponding to the edge information, and the embedding vector corresponding to the edge information is selected as the edge information embedding vector.

[0073] The Chinese name of the CLIP model is: Contrastive Language-Image Pre-training Model, and the English name of the CLIP model is: Contrastive Language-Image Pre-training. The CLIP model is a multimodal pre-training model based on contrastive learning. The CLIP model processes image and text data simultaneously and learns the matching relationship between images and text, thereby achieving cross-modal understanding and generation capabilities.

[0074] S23, acquiring a dynamic word according to the first text embedding vector and the image embedding vector, and determining a second text embedding vector based on the dynamic word and a predefined determination method;

[0075] Among them, dynamic words refer to those words that change with time, environment, context and other factors.

[0076] Among them, in the image generation process, dynamic words can be regarded as a key input information, which can provide the diffusion model with clues about the changes in image content, action or scene. These clues can guide the diffusion model to output generated images with dynamic features that conform to specific situations or action descriptions.

[0077] For the sake of illustration, an example is given below:

[0078] For example, when generating sports scenes, using dynamic words such as "running" and "jumping" can guide the diffusion model to output generated images with dynamic and coherent movements.

[0079] For example, when generating natural scenery, using dynamic words such as "sunrise" and "sunset" to describe time changes can guide the diffusion model to output generated images with light and shadow effects.

[0080] S24, processing the second text embedding vector by a diffusion model to obtain an initial image, and determining a generated image according to the initial image and the edge information embedding vector;

[0081] The step of processing the second text embedding vector by a diffusion model to obtain an initial image, and determining a generated image according to the initial image and the edge information embedding vector, includes:

[0082] Inputting the second text embedding vector into a U-Net network in a diffusion model, and processing the second text embedding vector through the U-Net network to obtain an initial image;

[0083] The initial image and the edge information embedding vector are input into a gating unit, and the initial image and the edge information embedding vector are fused through a gating mechanism in the gating unit to obtain an enhanced initial image, and the enhanced initial image is selected as a generated image output by the diffusion model.

[0084] Among them, the gating mechanism in the gating unit is a technology specifically used to control the flow of information in neural networks. The core idea of ​​the gating mechanism is to control the flow of information in the neural network by introducing gates. These gates are usually calculated based on the current input and the previous hidden state to decide which information to keep and which information to discard.

[0085] Among them, the diffusion model generates images by adding noise to the image and learning the denoising process. The U-Net network realizes efficient image generation and denoising functions in the diffusion model.

[0086] Among them, the Chinese name of the U-Net network is: U-type convolutional neural network.

[0087] Among them, the U-Net network is a fully convolutional network with a symmetrical encoder-decoder structure. The encoder extracts image features by gradually downsampling, while the decoder restores the features to the same size as the input image by upsampling. This structure enables the U-Net network to effectively capture and reconstruct image information.

[0088] The step of inputting the second text embedding vector into a U-Net network in a diffusion model and processing the second text embedding vector through the U-Net network to obtain an initial image includes:

[0089] A training time is obtained. When the training time is greater than a time threshold, the second text embedding vector is input into a U-Net network in a diffusion model, and the second text embedding vector is processed by the U-Net network to obtain an initial image.

[0090] The timing of using the first text embedding vector and the second text embedding vector is dynamically adjusted by using a time threshold. For ease of explanation, the details are as follows:

[0091] Among them, when the training time is not greater than the time threshold, it means that the generation process of the diffusion model is in the early stage. At this time, the first text embedding vector is input into the U-Net network in the diffusion model, and the first text embedding vector is processed by the U-Net network. The U-Net network of the diffusion model will use the first text embedding vector to guide the generation of the initial image to ensure that the content of the initial image meets the requirements of the first text embedding vector.

[0092] Among them, when the training time is greater than the time threshold, it means that the generation process of the diffusion model is in the later stage. At this time, the second text embedding vector is input into the U-Net network in the diffusion model, and the second text embedding vector is processed by the U-Net network. The U-Net network of the diffusion model will use the second text embedding vector to guide the generation of the initial image to ensure that the content of the initial image meets the requirements of the second text embedding vector.

[0093] Among them, based on the time threshold, the diffusion model realizes dynamic switching between the first text embedding vector and the second text embedding vector in the process of generating images.

[0094] S25, acquiring the trained diffusion model with the goal of minimizing the loss value between the generated image and the preset target image and maximizing the reward value in reinforcement learning;

[0095] The diffusion model is trained with the goal of minimizing the loss value between the generated image and the preset target image and maximizing the reward value in reinforcement learning, and the trained diffusion model is obtained.

[0096] The preset input image and the preset target image are different images.

[0097] For the sake of illustration, an example is given below:

[0098] For example, the preset input images are image one and image two, and the preset target image is image three.

[0099] Image 1, image 2, and image 3 are different images.

[0100] S26, obtaining current text description information, current input image and edge information of the current input image, inputting the current text description information, current input image and edge information of the current input image into the trained diffusion model, and obtaining the current target image output by the trained diffusion model.

[0101] Among them, the current input image and the current target image are different images.

[0102] For example, the current input images are image 4 and image 5, and the current target image is image 6.

[0103] Image 4, Image 5, and Image 6 are different images.

[0104] For the sake of illustration, an example is given below:

[0105] The current text description is: A cat is sitting on a chair. There are two current input images, one of a cat and one of a chair.

[0106] The current text description information, the cat image, the chair image, the edge information of the cat image, and the edge information of the chair image are input into the trained diffusion model to obtain the current target image output by the trained diffusion model. At this time, the current target image is an image of a cat sitting on a chair.

[0107] Wherein, after obtaining the current text description information, the current input image and the edge information of the current input image, inputting the current text description information, the current input image and the edge information of the current input image into the trained diffusion model, and obtaining the current target image output by the trained diffusion model, the image generation method includes:

[0108] Step A: creating a display window, and displaying the current target image through the display window.

[0109] The current target image is displayed through the display window, and the user can view the current target image through the display window.

[0110] Wherein, after obtaining the current text description information, the current input image and the edge information of the current input image, inputting the current text description information, the current input image and the edge information of the current input image into the trained diffusion model, and obtaining the current target image output by the trained diffusion model, the image generation method includes:

[0111] Step B, obtaining a push instruction, executing the push instruction, and pushing the current target image to the target system.

[0112] The current target image is pushed to the target system, and the user can view the current target image through the target system, which is helpful for the user to determine the application effect of the trained diffusion model, so as to perform targeted optimization.

[0113] Step A may be executed before or after step B, or simultaneously with step B, without limitation.

[0114] In the embodiment of the present invention, the beneficial effects lie in two aspects. On the one hand, the current text description information, the current input image and the edge information of the current input image are obtained, and the current text description information, the current input image and the edge information of the current input image are input into the trained diffusion model to obtain the current target image output by the trained diffusion model. Since the edge information of the current input image can provide the detail information of the current input image, the trained diffusion model can capture the detail information of the current input image, which is beneficial to improve the reliability of the current target image output by the trained diffusion model. On the other hand, since the trained diffusion model will not be affected by human intervention, it is beneficial to improve the stability of the acquired current target image.

[0115] See also Figure 3 , Figure 3 yes Figure 2 A specific implementation flow diagram of step S23 is described in detail as follows:

[0116] S31, inputting the first text embedding vector and the image embedding vector into a preset adaptive neural fuzzy inference system, obtaining a dynamic word output by the adaptive neural fuzzy inference system, adding the dynamic word to the preset text description information, and obtaining the modified preset text description information;

[0117] Among them, the adaptive neuro-fuzzy inference system is an intelligent system that organically combines neural networks with fuzzy reasoning. The adaptive neuro-fuzzy inference system integrates the learning mechanism of neural networks and the language reasoning ability of fuzzy systems, which not only brings out the advantages of both, but also makes up for their respective shortcomings.

[0118] S32, performing feature extraction on the modified preset text description information through the text encoder of the CLIP model to obtain an embedding vector corresponding to the modified preset text description information, and selecting the embedding vector corresponding to the modified preset text description information as the second text embedding vector.

[0119] In an embodiment of the present invention, since the modified preset text description information has dynamic words, the modified preset text description information has higher accuracy, and the embedding vector corresponding to the modified preset text description information is selected as the second text embedding vector, which can reduce ambiguity and uncertainty in the image generation process, thereby improving the quality of the generated image.

[0120] See also Figure 4 , Figure 4 yes Figure 2 A specific implementation flow diagram of step S25 is described in detail as follows:

[0121] S41, obtaining a preset target image corresponding to the preset text description information, and calculating a current loss value between the generated image and the preset target image through a preset loss function;

[0122] The loss function includes an expected loss function, a cross entropy loss function, or a combination thereof.

[0123] Preferably, the loss function adopts the expected loss function.

[0124] Among them, the expected loss function focuses on the average loss over the entire data distribution. This makes the expected loss function more advantageous in evaluating the generalization ability of the diffusion model.

[0125] S42, obtaining a reward function corresponding to the diffusion model, optimizing the parameters of the diffusion model using a reinforcement learning algorithm, obtaining a reward value of the reward function during the optimization of the parameters, subtracting the reward value from the current loss value to generate a total loss value, training the diffusion model based on the total loss value, and stopping training the diffusion model when the total loss value converges to obtain the trained diffusion model.

[0126] When the total loss value converges, the training of the diffusion model is stopped, and the trained diffusion model is obtained, including:

[0127] When the total loss value is less than a preset loss value, the training of the diffusion model is stopped, and the trained diffusion model is obtained.

[0128] Wherein, the reward function is:

[0129] ;

[0130] is the reward value, is the first weight parameter, is the second weight parameter, is the third weight parameter, is the consistency reward value, is the quality reward value, reward value for diversity;

[0131] Optionally, the first weight parameter and the second weight parameter are positive numbers, and the third weight parameter is a negative number.

[0132] Among them, the similarity between the generated image and the preset target image is selected as the consistency reward value. The higher the similarity, the greater the consistency reward value, and the smaller the similarity, the smaller the consistency reward value.

[0133] Among them, the reward value corresponding to the FID value is selected as the quality reward value. The smaller the FID value, the larger the quality reward value, and the larger the FID value, the smaller the quality reward value.

[0134] Among them, the Chinese name of FID can be: Fréchet initial distance, and the English name of FID is: Fréchet Inception Distance. The FID value is sensitive to the details and texture changes of the image.

[0135] Among them, the smaller the FID value is, the closer the feature distribution of the generated image is to the feature distribution of the preset target image, which means that the quality of the generated image is higher, and therefore, the greater the quality reward value.

[0136] Among them, the larger the FID value is, the less the feature distribution of the generated image is close to the feature distribution of the preset target image, which means that the quality of the generated image is worse, and therefore, the quality reward value is lower.

[0137] Among them, a color histogram of each of the generated images is obtained, and the difference values ​​of the color histograms between different generated images are calculated, the difference values ​​are added to generate a sum, the sum is divided by the total number of the generated images to obtain an average value, and the average value is selected as the diversity reward value.

[0138] The diversity reward value is based on the diversity of color and texture, and is used to measure the diversity of generated images. A larger value indicates that different generated images have greater differences in color and texture, that is, the image set composed of generated images has a higher diversity. For ease of explanation, an example is given below:

[0139] ;

[0140] in, represents the diversity reward value; Represents the total number of generated images.

[0141] represents the color histogram of the i-th generated image, Represents the color histogram of the j-th generated image;

[0142] represents the difference between the color histogram of the i-th generated image and the color histogram of the j-th generated image; represents the total number of pixels in the kth color interval in the color histogram of the i-th generated image;

[0143] represents the total number of pixels in the kth color interval in the color histogram of the jth generated image;

[0144] is the minimization function, express , The minimum value in .

[0145] in, The bigger, the and The more dissimilar the distribution of the number of pixels is, the greater the difference is, that is, the higher the diversity of the image set composed of different generated images. The smaller, the and The more similar the distribution of the number of pixels is, the smaller the difference is, that is, the lower the diversity of the image set composed of different generated images.

[0146] Among them, a diffusion model that can output different generated images usually means that the generated images have high flexibility and innovation. This diversity may come from the wide coverage of the input data by the generated images and the effective exploration of the latent space. When the generated images can generate diverse generated images, the diffusion model can capture the complexity and diversity in the data, thereby outputting more realistic and richer generated images.

[0147] In the embodiment of the present invention, the trained diffusion model is obtained, and the trained diffusion model is reloaded and used when needed without retraining, thereby reducing the training cost of the diffusion model.

[0148] See also Figure 5 , Figure 5 yes Figure 2 A specific implementation flow diagram of step S26 is described in detail as follows:

[0149] S51, obtaining current text description information and a current input image, and obtaining edge information of the current input image by using a Canny edge detection algorithm;

[0150] S52, inputting the current text description information, the current input image and the edge information of the current input image into the trained diffusion model, and acquiring the current target image output by the trained diffusion model.

[0151] In an embodiment of the present invention, the current text description information, the current input image and the edge information of the current input image are input into the trained diffusion model to obtain the current target image output by the trained diffusion model. Since the edge information of the current input image can provide detail information of the current input image, the trained diffusion model can capture the detail information of the current input image, which is beneficial to improving the reliability of the current target image output by the trained diffusion model.

[0152] See also Figure 6 , Figure 6 is a schematic diagram of a structure of an image generating device in one embodiment of the present invention. Figure 6 As shown, the image generation device includes a first acquisition module 101, an extraction module 102, a second acquisition module 103, a first generation module 104, a training module 105, and a second generation module 106. The functional modules are described in detail as follows:

[0153] A first acquisition module 101 is used to acquire preset text description information, a preset input image and edge information of the preset input image;

[0154] An extraction module 102 is used to perform feature extraction on the preset text description information to obtain a first text embedding vector, perform feature extraction on the preset input image to obtain an image embedding vector, and perform feature extraction on edge information of the preset input image to obtain an edge information embedding vector;

[0155] A second acquisition module 103 is used to acquire a dynamic word according to the first text embedding vector and the image embedding vector, and determine a second text embedding vector based on the dynamic word and a predefined determination method;

[0156] A first generating module 104 is used to process the second text embedding vector through a diffusion model to obtain an initial image, and determine a generated image according to the initial image and the edge information embedding vector;

[0157] A training module 105, configured to obtain the diffusion model after training with the goal of minimizing the loss value between the generated image and the preset target image and maximizing the reward value in reinforcement learning;

[0158] The second generation module 106 is used to obtain the current text description information, the current input image and the edge information of the current input image, input the current text description information, the current input image and the edge information of the current input image into the trained diffusion model, and obtain the current target image output by the trained diffusion model.

[0159] In the embodiment of the present invention, the beneficial effects lie in two aspects. On the one hand, the current text description information, the current input image and the edge information of the current input image are obtained, and the current text description information, the current input image and the edge information of the current input image are input into the trained diffusion model to obtain the current target image output by the trained diffusion model. Since the edge information of the current input image can provide the detail information of the current input image, the trained diffusion model can capture the detail information of the current input image, which is beneficial to improve the reliability of the current target image output by the trained diffusion model. On the other hand, since the trained diffusion model will not be affected by human intervention, it is beneficial to improve the stability of the acquired current target image.

[0160] For the specific definition of the image generating device, please refer to the definition of the image generating method above, which will not be repeated here.

[0161] Each module in the above-mentioned image generation device can be implemented in whole or in part by software, hardware or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute operations corresponding to each module above.

[0162] See also Figure 7 , Figure 7 1 is another structural diagram of a computer device in one embodiment of the present invention. In one embodiment, a computer device is provided. The computer device is a server device or a client device. The internal structure diagram thereof can be as shown in FIG. Figure 7 As shown. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external database. When the computer program is executed by the processor, the functions or steps of an image generation method based on reinforcement learning and diffusion model can be realized.

[0163] In one embodiment, a computer device is provided, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor.

[0164] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can refer to the relevant description of the aforementioned method embodiment. In order to avoid repetition, they will not be described one by one here.

[0165] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a graphics processing unit (GPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components.

[0166] The above description and the accompanying drawings fully illustrate the embodiments of the present disclosure so that those skilled in the art can practice them. Other embodiments may include structural, logical, electrical, process and other changes. The embodiments represent only possible changes. Unless explicitly required, individual components and functions are optional, and the order of operations may vary. Parts and subsamples of some embodiments may be included in or replace parts and subsamples of other embodiments. Moreover, the words used in this application are only used to describe the embodiments and are not used to limit the claims. As used in the description of the embodiments and the claims, unless the context clearly indicates otherwise, the singular forms of "a" and "an" are intended to also include plural forms. Similarly, the term "and / or" as used in this application refers to any and all possible combinations of one or more associated listings. In addition, in this article, each embodiment may focus on the differences from other embodiments, and the same and similar parts between the embodiments may refer to each other. For the methods, products, etc. disclosed in the embodiments, if they correspond to the method part disclosed in the embodiments, then the relevant parts can refer to the description of the method part.

[0167] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software may depend on the specific application and design constraints of the technical solution. Technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the embodiments of the present disclosure. Technicians can clearly understand that for the convenience and simplicity of description, the specific working process of the above-described system, device and unit can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0168] In the embodiments disclosed herein, the disclosed methods and products (including but not limited to devices, equipment, etc.) can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of units can be only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some sub-samples can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between each other shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms. The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the units may be selected according to actual needs to implement this embodiment. In addition, each functional unit in the embodiment of the present disclosure may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit.

[0169] The flowchart and block diagram in the accompanying drawings show the possible architecture, function and operation of the system, method and computer program product according to the embodiment of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of the code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. In some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, which can depend on the functions involved. In the description corresponding to the flowchart and the block diagram in the accompanying drawings, the operations or steps corresponding to different boxes can also occur in a different order from the order disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, which can depend on the functions involved. Each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified functions or actions, or may be implemented by a combination of dedicated hardware and computer instructions.

Claims

1. An image generation method based on reinforcement learning and diffusion model, characterized in that: include: Acquire preset text description information, a preset input image, and edge information of the preset input image; Performing feature extraction on the preset text description information to obtain a first text embedding vector, performing feature extraction on the preset input image to obtain an image embedding vector, and performing feature extraction on edge information of the preset input image to obtain an edge information embedding vector; Acquire a dynamic word according to the first text embedding vector and the image embedding vector, and determine a second text embedding vector based on the dynamic word and a predefined determination method; Processing the second text embedding vector by a diffusion model to obtain an initial image, and determining a generated image according to the initial image and the edge information embedding vector; With the goal of minimizing the loss value between the generated image and the preset target image and maximizing the reward value in reinforcement learning, obtaining the trained diffusion model; Acquire current text description information, current input image and edge information of the current input image, input the current text description information, current input image and edge information of the current input image into the trained diffusion model, and acquire a current target image output by the trained diffusion model; The acquiring a dynamic word according to the first text embedding vector and the image embedding vector, and determining a second text embedding vector based on the dynamic word and a predefined determination method, comprises: Inputting the first text embedding vector and the image embedding vector into a preset adaptive neural fuzzy inference system, obtaining a dynamic word output by the adaptive neural fuzzy inference system, adding the dynamic word to the preset text description information, and obtaining the modified preset text description information; The text encoder of the CLIP model is used to perform feature extraction on the modified preset text description information to obtain an embedding vector corresponding to the modified preset text description information, and the embedding vector corresponding to the modified preset text description information is selected as the second text embedding vector.

2. The image generation method according to claim 1, characterized in that: The step of extracting features from the preset text description information to obtain a first text embedding vector, extracting features from the preset input image to obtain an image embedding vector, and extracting features from edge information of the preset input image to obtain an edge information embedding vector comprises: Perform feature extraction on the preset text description information through a text encoder of the CLIP model to obtain an embedding vector corresponding to the preset text description information, and select the embedding vector corresponding to the preset text description information as a first text embedding vector; Performing feature extraction on the preset input image through the image encoder of the CLIP model to obtain an embedding vector corresponding to the preset input image, and selecting the embedding vector corresponding to the preset input image as the image embedding vector; The image encoder of the CLIP model is used to extract features of the edge information to obtain an embedding vector corresponding to the edge information, and the embedding vector corresponding to the edge information is selected as the edge information embedding vector.

3. The image generation method according to claim 1, characterized in that: The step of processing the second text embedding vector by using a diffusion model to obtain an initial image, and determining a generated image according to the initial image and the edge information embedding vector, comprises: Inputting the second text embedding vector into a U-Net network in a diffusion model, and processing the second text embedding vector through the U-Net network to obtain an initial image; The initial image and the edge information embedding vector are input into a gating unit, and the initial image and the edge information embedding vector are fused through a gating mechanism in the gating unit to obtain an enhanced initial image, and the enhanced initial image is selected as a generated image output by the diffusion model.

4. The image generation method according to claim 1, characterized in that: The step of obtaining the trained diffusion model with the goal of minimizing the loss value between the generated image and the preset target image and maximizing the reward value in the reinforcement learning includes: Obtaining a preset target image corresponding to the preset text description information, and calculating a current loss value between the generated image and the preset target image using a preset loss function; Obtain a reward function corresponding to the diffusion model, optimize the parameters of the diffusion model using a reinforcement learning algorithm, obtain a reward value of the reward function during the optimization of the parameters, subtract the reward value from the current loss value to generate a total loss value, train the diffusion model based on the total loss value, stop training the diffusion model when the total loss value converges, and obtain the trained diffusion model.

5. The image generation method according to claim 1, characterized in that: The step of obtaining the current text description information, the current input image, and the edge information of the current input image, inputting the current text description information, the current input image, and the edge information of the current input image into the trained diffusion model, and obtaining the current target image output by the trained diffusion model includes: Obtaining current text description information and a current input image, and obtaining edge information of the current input image through a Canny edge detection algorithm; The current text description information, the current input image and the edge information of the current input image are input into the trained diffusion model to obtain the current target image output by the trained diffusion model.

6. The image generation method according to claim 1, characterized in that: After obtaining the current text description information, the current input image and the edge information of the current input image, inputting the current text description information, the current input image and the edge information of the current input image into the trained diffusion model, and obtaining the current target image output by the trained diffusion model, the image generation method includes: A display window is created, and the current target image is displayed through the display window.

7. The image generation method according to claim 1, characterized in that: After obtaining the current text description information, the current input image and the edge information of the current input image, inputting the current text description information, the current input image and the edge information of the current input image into the trained diffusion model, and obtaining the current target image output by the trained diffusion model, the image generation method includes: A push instruction is obtained, and the push instruction is executed to push the current target image to the target system.

8. An image generation device based on reinforcement learning and diffusion model, characterized in that: include: A first acquisition module, used to acquire preset text description information, a preset input image and edge information of the preset input image; An extraction module, configured to perform feature extraction on the preset text description information to obtain a first text embedding vector, perform feature extraction on the preset input image to obtain an image embedding vector, and perform feature extraction on edge information of the preset input image to obtain an edge information embedding vector; A second acquisition module is used to acquire a dynamic word according to the first text embedding vector and the image embedding vector, and determine a second text embedding vector based on the dynamic word and a predefined determination method; A first generating module, configured to process the second text embedding vector through a diffusion model to obtain an initial image, and determine a generated image according to the initial image and the edge information embedding vector; A training module, used to obtain the trained diffusion model with the goal of minimizing the loss value between the generated image and the preset target image and maximizing the reward value in reinforcement learning; A second generation module is used to obtain current text description information, a current input image and edge information of the current input image, input the current text description information, the current input image and the edge information of the current input image into the trained diffusion model, and obtain a current target image output by the trained diffusion model; The acquiring a dynamic word according to the first text embedding vector and the image embedding vector, and determining a second text embedding vector based on the dynamic word and a predefined determination method, comprises: Inputting the first text embedding vector and the image embedding vector into a preset adaptive neural fuzzy inference system, obtaining a dynamic word output by the adaptive neural fuzzy inference system, adding the dynamic word to the preset text description information, and obtaining the modified preset text description information; The text encoder of the CLIP model is used to perform feature extraction on the modified preset text description information to obtain an embedding vector corresponding to the modified preset text description information, and the embedding vector corresponding to the modified preset text description information is selected as the second text embedding vector.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the image generating method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Novel image generation method jointly driven by text and semantic segmentation map

    CN117557683A

  • Image defogging method and device based on diffusion model, equipment and medium

    CN118521496A