Image generation method and electronic equipment
By decomposing it into a single instance task and using the global integrated information generation method, the problems of attribute obfuscation and instance loss in L2I technology are solved, and precise control and efficient generation of multi-instance image generation are realized.
Patent Information
- Application Number
- CN202410147641.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-31
- Publication Date
- 2025-08-01
AI Technical Summary
The existing L2I-based image generation technology is prone to the problems of attribute confusion and instance loss when processing multi-instance image generation, especially when there are many attributes such as number, size, type, color, material, and style of instances, the generated images do not meet user expectations.
The image generation method of division-conquer-combination is used to disassemble it into multiple single instance tasks, and the image features of each instance are generated one by one, and the instance position and attributes are accurately controlled through global integration of information, and the target image is finally generated.
Accurate control of the attributes and positions of each instance in the target image is realized. The instance positions in the generated image match the positions specified by the user, the success rate is higher than the prior art and the generation efficiency is higher.
Smart Images

Figure CN120411298A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of image processing, and in particular, to an image generation method and an electronic device. Background Art
[0002] In recent years, the technology of generating high-quality images through text has made great progress. Among them, a relatively key technology is Layout-to-Image Generation (L2I). Among them, the L2I task requires the user to provide, in addition to the text description related to the content of the generated image, the position distribution layout information of each object (or called object / instance) in the image; then, the L2I algorithm can generate an image that conforms to the object layout distribution according to the information provided by the user.
[0003] However, when the number, size, type, color, material, style and other attributes of the instances are relatively large, the images generated by the existing L2I-based image generation solutions have problems such as attribute confusion and instance loss. For example, the text input by the user is "a blue cat and a green dog", but the existing technology may generate an image of "two blue cats", or an image of "a blue cat and a blue dog", or an image of "a blue cat and a black and white dog", and so on. Summary of the Invention
[0004] In view of this, the present application provides an image generation method and an electronic device. The image generation method can solve the problems of attribute confusion and instance loss in the generated target image.
[0005] In a first aspect, the embodiments of the present application provide an image generation method, which includes: First, obtain a plurality of information groups; wherein, each information group includes the attribute information and position information of an instance; then, generate the image feature of an instance according to the attribute information of an instance included in each of the plurality of information groups; subsequently, integrate the image features of the plurality of instances according to the position information of the plurality of instances included in the plurality of information groups to generate a target image.
[0006] That is to say, the present application adopts a "divide-conquer-combine" scheme to implement image generation. Among them, "divide" can refer to "decomposing and partitioning a single-instance task", that is, decomposing the description information of each instance (including the attribute information and position information of the instance), that is, decomposing multiple information groups, and a single information group can correspond to a single-instance task. "Conquer" can refer to "generating and processing each single-instance task one by one", that is, executing the single-instance task multiple times, and each execution of the single-instance task can generate the image features of one instance. It should be noted that the present application can execute multiple single-instance tasks in parallel. "Combine" can refer to "integrating the generation results of each instance", that is, integrating the image features of multiple instances to generate the target image. In this way, a complex multi-instance can be decomposed into multiple relatively simple single-instance generation tasks; and in the process of executing the single-instance generation task, since each information group only includes the attribute information of the single instance itself, there is no problem of attribute leakage; furthermore, the present application can ensure that each instance can be generated according to the attributes specified by the user; thus, in the process of generating the target image, precise control of the attributes of each instance in the target image can be achieved.
[0007] In addition, in the global integration process, the position information of multiple instances is utilized, so the present application can also achieve precise control of the positions of each instance in the target image.
[0008] It should be noted that the multiple information groups can be input by the user or obtained by parsing the information input by the user; the embodiments of the present application do not limit this. When the multiple information groups are input by the user, "divide" is realized by the user to disassemble the multiple information groups.
[0009] Exemplarily, the instance can be any object; for example, a person, an animal, a plant, a hat, a stool, a head, limbs, a tail, eyes, etc.; the present application does not limit this.
[0010] Exemplarily, the attribute information of the instance can include the information used to describe the attributes of the instance, where the instance attributes can include but are not limited to: size, type (such as name), color, material, style, etc., and the present application does not limit this.
[0011] Exemplarily, the number of information groups can be represented by n (n is a positive integer); that is to say, the present application can generate the image features of n instances: the image features of instance 1, the image features of instance 2,..., the image features of instance n.
[0012] Exemplarily, a neural network can be used to process the attribute information of one instance included in one information group and output the image features of one instance.
[0013] Exemplarily, "integration" may refer to connecting scattered things to each other to form a whole; "integration" in this application can be understood as connecting the image features of multiple instances to each other to obtain the overall features including n instances; thereafter, a target image can be generated according to the overall features including n instances. Exemplarily, a neural network can be used to process the overall features including n instances to obtain the target image.
[0014] Exemplarily, the target image generated in this application includes these n instances, the positions of the n instances in the target image correspond one-to-one to the position information in the n information groups, and the attributes of the n instances in the target image correspond one-to-one to the attribute information in the n information groups.
[0015] It should be noted that when only one information group is obtained, the image generation method of this application can also be used to generate the target image.
[0016] It should be noted that "multiple information groups" may refer to two information groups, or more than two information groups.
[0017] According to the first aspect, multiple mask graphs are generated according to the position information of multiple instances; wherein, the multiple mask graphs correspond one-to-one to multiple pixel points in the target image; the size of each mask graph is the same as the size of the target image, and the mask graph corresponding to the first pixel point is the mask graph corresponding to the instance to which the first pixel point belongs, and the first pixel point is any pixel point in the target image; global integration information is generated according to the multiple mask graphs; integrating the image features of multiple instances according to the position information of the multiple instances included in the multiple information groups to generate a target image, including: integrating the image features of multiple instances according to the position information of multiple instances and the global integration information to generate a target image.
[0018] That is to say, in the global integration process, global integration information is newly introduced; in this way, the information for guiding the integration of the image features of multiple instances is richer, and therefore, more accurate position control of multiple instances in the target image can be achieved.
[0019] In addition, in the process of generating global integration information according to the multiple mask graphs, the probability that each pixel point is any instance or the background can be fine-tuned, that is, the positions of each instance in the target image are adjusted; furthermore, the positions of each instance in the target image can be made to match the position information of each instance more, and the accuracy of the positions of multiple instances in the target image can be further improved.
[0020] According to the first aspect, or any implementation of the above first aspect, obtain the global description information of the target image; generate global image features based on the global description information; integrate the image features of multiple instances included in multiple information groups according to the position information of the multiple instances to generate a target image, including: integrating the image features of multiple instances according to the position information of the multiple instances and the global image features to generate a target image.
[0021] Among them, the global description information of the target image may include at least one of foreground description information or background description information. Among them, the foreground description information may include at least one of the attribute information of all instances included in the target image, the positional relationship between each instance, and the action relationship between each instance. In this way, a target image matching the globally specified description information can be generated.
[0022] Exemplarily, the source of the global description information of the target image may include but is not limited to: the user's current input, the user's historical input, or pre-settings of the electronic device, etc., and this application does not limit this. This application takes the global description information of the target image as the user's current input as an example for illustration.
[0023] According to the first aspect, or any implementation of the above first aspect, generate multiple mask maps according to the position information of multiple instances; among them, the multiple mask maps correspond one-to-one with multiple pixel points in the target image; the size of each mask map is the same as the size of the target image, and the mask map corresponding to the first pixel point is the mask map corresponding to the instance to which the first pixel point belongs, and the first pixel point is any pixel point in the target image; generate global integration information according to the multiple mask maps; integrate the image features of multiple instances according to the position information of multiple instances included in multiple information groups and the global image features to generate a target image, including: integrating the image features of multiple instances according to the position information of multiple instances, the global integration information, and the global image features to generate a target image.
[0024] According to the first aspect, or any implementation of the above first aspect, obtain multiple information groups, including: receiving multiple information groups input by the user. In this way, the user can input information groups, and there is no need for the terminal device to disassemble single instances, which can improve the efficiency of the terminal device in generating target images.
[0025] According to the first aspect, or any implementation of the above first aspect, obtain multiple information groups, including: performing text parsing on the global description information of the target image to obtain the attribute information of multiple instances; performing instance position prediction on the global description information of the target image to obtain the position information of multiple instances.
[0026] That is to say, the input interface has diversity, and users can choose to input multiple information groups or input global description information according to their needs, which can better meet the needs of users.
[0027] According to the first aspect, or any implementation manner of the above first aspect, generate an image feature of an instance according to the attribute information of an instance included in each information group among multiple information groups, including: dividing the attribute information of multiple instances into multiple paths and inputting them into an instance processing network, and generating features by the instance processing network according to the attribute information of an instance to obtain an image feature of an instance; wherein, the attribute information of an instance is an input path of the instance processing network.
[0028] Exemplarily, the instance processing network can be implemented by a neural network and can also be referred to as an instance processing module.
[0029] According to the first aspect, or any implementation manner of the above first aspect, integrate the image features of multiple instances according to the position information of multiple instances included in multiple information groups to generate a target image, including: inputting the position information of multiple instances and the image features of multiple instances into a global integration network to obtain a target image.
[0030] Exemplarily, the global integration network can be implemented by a neural network and can also be referred to as a global integration module.
[0031] According to the first aspect, or any implementation manner of the above first aspect, generate an image feature of an instance according to the attribute information of an instance included in each information group among multiple information groups, including: in each generation process of the diffusion network performing T1 generation processes, using the first intermediate image generated in the previous generation as one path and the attribute information of multiple instances as multiple paths to input into the diffusion network, and performing multiple feature generations by the instance processing network in the diffusion network to obtain image features of multiple instances; wherein, the first intermediate image generated in the previous generation and the attribute information of an instance are used for one feature generation, and the attribute information of an instance is an input path; integrate the image features of multiple instances according to the position information of multiple instances included in multiple information groups to generate a target image, including: in each generation process of the diffusion network performing T1 generation processes, inputting the position information of multiple instances and the image features of multiple instances into the global integration network in the diffusion network to obtain a first intermediate feature; inputting the first intermediate feature into other networks in the diffusion network to obtain a first intermediate image; wherein, the global integration network is after the instance processing network, and the other networks are after the global integration network; determine the target image according to the first intermediate image obtained by the diffusion network performing the T1-th generation process.
[0032] According to the first aspect, or any implementation of the above first aspect, determining a target image based on a first intermediate image obtained by performing a T1-th generation process according to a diffusion network includes: in each generation process of the subsequent T2-th generation process after the diffusion network, taking the second intermediate image generated last time as one path and taking the attribute information of multiple instances as another path and inputting them into the diffusion network, and the instance processing network generates features according to the second intermediate image generated last time and the attribute information of multiple instances to obtain a second intermediate feature; inputting the second intermediate feature into another network to obtain a second intermediate image; taking the second intermediate image obtained by performing the T-th generation process of the diffusion network as the target image; where T is the sum of T1 and T2.
[0033] Since in the subsequent stages of the chain generation process of the diffusion network, details of each object in the target image, attributes, and light, shadow, style, tone, etc. of the entire image can be generated; in this way, without changing or basically not changing the image content, the interaction relationships of each object in the finally obtained target image are natural, without a sense of splicing and fragmentation, and the tone and light and shadow are harmonious.
[0034] According to the first aspect, or any implementation of the above first aspect, determining a target image based on a first intermediate image obtained by performing a T1-th generation process according to a diffusion network includes: taking the first intermediate image obtained by performing the T1-th generation process of the diffusion network as the target image.
[0035] According to the first aspect, or any implementation of the above first aspect, the attribute information includes at least one of the following: an image or text; the position information includes at least one of the following: an image or text.
[0036] For example, the attribute information is the text "a blue cat".
[0037] For example, the attribute information is an image of a blue cat.
[0038] For example, the attribute information is a pose map of a cat.
[0039] Exemplarily, the global description information can be text or an image, and the present application does not limit this.
[0040] For example, the position information is the coordinates of the four vertices of the minimum bounding box of the instance.
[0041] For example, the position information is a mask image generated based on the contour information of an instance. For example, the pixel values of the pixel points located within the contour box of the instance can be set to 0, and the pixel values of the pixel points located outside the contour box of the instance can be set to 1 to obtain the mask image as the position information. Alternatively, the pixel values of the pixel points located within the contour box of the instance can be set to 1, and the pixel values of the pixel points located outside the contour box of the instance can be set to 0 to obtain the mask image as the position information.
[0042] According to the first aspect, or any implementation manner of the above first aspect, the position information includes at least one of the following: the contour information of the instance and the information of the bounding box of the instance.
[0043] It should be noted that the image generation method of the present application can generate multiple instances with action relationships; the image generation method of the present application can also generate instances of different styles; the image generation method of the present application can also generate instances of different textures; the image generation method of the present application can also generate instances of different shapes; the image generation method of the present application can also generate instances of different quantities; the image generation method of the present application can also generate instances of different colors.
[0044] In addition, through testing, the success rate of the present application in generating target images containing a specified number is higher than that of the prior art; and compared with some prior arts, the present application takes less time to generate target images, and the success rate of generating target images containing a specified number is high. Moreover, the positions of the instances in the target images generated by the present application are more matched with the positions specified by the user; and the quantitative indicators of the target images generated by the present application are all better than those of the prior art.
[0045] In a second aspect, an image generation device is provided in an embodiment of the present application. The device includes:
[0046] An acquisition module, configured to acquire a plurality of information groups; wherein each information group includes the attribute information and position information of an instance;
[0047] A first generation module, configured to generate the image features of an instance according to the attribute information of an instance included in each information group among the plurality of information groups;
[0048] An integration module, configured to integrate the image features of the plurality of instances according to the position information of the plurality of instances included in the plurality of information groups to generate a target image.
[0049] According to the second aspect, the image generation device further includes:
[0050] A second generation module, configured to generate multiple mask maps according to the position information of multiple instances; wherein, the multiple mask maps correspond one by one to multiple pixel points in the target image; the size of each mask map is the same as the size of the target image, and the mask map corresponding to the first pixel point is the mask map corresponding to the instance to which the first pixel point belongs, and the first pixel point is any pixel point in the target image; and generate global integration information according to the multiple mask maps;
[0051] An integration module, specifically configured to integrate the image features of multiple instances according to the position information of multiple instances and the global integration information to generate a target image.
[0052] According to the second aspect, or any one of the implementation manners of the second aspect above, the acquisition module is further configured to acquire the global description information of the target image;
[0053] The image generation device further includes:
[0054] A third generation module, configured to generate global image features according to the global description information;
[0055] An integration module, specifically configured to integrate the image features of multiple instances according to the position information of multiple instances and the global image features to generate a target image.
[0056] According to the second aspect, or any one of the implementation manners of the second aspect above, the image generation device further includes:
[0057] A second generation module, configured to generate multiple mask maps according to the position information of multiple instances; wherein, the multiple mask maps correspond one by one to multiple pixel points in the target image; the size of each mask map is the same as the size of the target image, and the mask map corresponding to the first pixel point is the mask map corresponding to the instance to which the first pixel point belongs, and the first pixel point is any pixel point in the target image; and generate global integration information according to the multiple mask maps;
[0058] An integration module, specifically configured to integrate the image features of multiple instances according to the position information of multiple instances, the global integration information and the global image features to generate a target image.
[0059] According to the second aspect, or any one of the implementation manners of the second aspect above, the acquisition module is specifically configured to receive multiple information groups input by a user.
[0060] According to the second aspect, or any one of the implementation manners of the second aspect above, the acquisition module is specifically configured to perform text parsing according to the global description information of the target image to obtain the attribute information of multiple instances; perform instance position prediction according to the global description information of the target image to obtain the position information of multiple instances.
[0061] According to the second aspect, or any implementation manner of the above second aspect, the first generation module is specifically configured to divide the attribute information of multiple instances into multiple paths and input them into the instance processing network. The instance processing network generates features based on the attribute information of one instance to obtain the image features of one instance. Among them, the attribute information of one instance is one path of input of the instance processing network.
[0062] According to the second aspect, or any implementation manner of the above second aspect, the integration module is specifically configured to input the position information of multiple instances and the image features of multiple instances into the global integration network to obtain the target image.
[0063] According to the second aspect, or any implementation manner of the above second aspect,
[0064] The first generation module is specifically configured to, in each generation process of the diffusion network performing T1 generation processes, use the first intermediate image generated in the previous generation as one path and the attribute information of multiple instances as multiple paths and input them into the diffusion network. The instance processing network in the diffusion network performs multiple feature generations to obtain the image features of multiple instances. Among them, the first intermediate image generated in the previous generation and the attribute information of one instance are used for one feature generation, and the attribute information of one instance is one path of input.
[0065] The integration module is specifically configured to, in each generation process of the diffusion network performing T1 generation processes, input the position information of multiple instances and the image features of multiple instances into the global integration network in the diffusion network to obtain the first intermediate feature; input the first intermediate feature into other networks in the diffusion network to obtain the first intermediate image. Among them, the global integration network is after the instance processing network, and the other networks are after the global integration network; determine the target image according to the first intermediate image obtained by the diffusion network performing the T1th generation process.
[0066] According to the second aspect, or any implementation manner of the above second aspect, the integration module is specifically configured to, in each generation process of the diffusion network performing the subsequent T2 generation processes, use the second intermediate image generated in the previous generation as one path and the attribute information of multiple instances as one path and input them into the diffusion network. The instance processing network generates features based on the second intermediate image generated in the previous generation and the attribute information of multiple instances to obtain the second intermediate feature; input the second intermediate feature into other networks to obtain the second intermediate image; use the second intermediate image obtained by the diffusion network performing the Tth generation process as the target image. Among them, T is the sum of T1 and T2.
[0067] According to a second aspect, or any implementation manner of the above second aspect, the integration module is specifically configured to use the first intermediate image obtained by performing the T1-th generation process on the diffusion network as the target image.
[0068] According to a second aspect, or any implementation manner of the above second aspect, the attribute information includes at least one of the following: image or text; the location information includes at least one of the following: image or text.
[0069] According to a second aspect, or any implementation manner of the above second aspect, the location information includes at least one of the following: the contour information of the instance and the information of the bounding box of the instance.
[0070] The second aspect and any implementation manner of the second aspect respectively correspond to the first aspect and any implementation manner of the first aspect. For the technical effects corresponding to the second aspect and any implementation manner of the second aspect, reference can be made to the technical effects corresponding to the first aspect and any implementation manner of the first aspect above, which will not be elaborated here.
[0071] In a third aspect, an embodiment of the present application provides an electronic device, including: a memory and a processor, the memory is coupled to the processor; the memory stores program instructions, and when the program instructions are executed by the processor, the electronic device is enabled to execute the method in the first aspect or any possible implementation manner of the first aspect.
[0072] Exemplarily, the electronic device can be a terminal device or a server, and the embodiments of the present application do not limit this.
[0073] The third aspect and any implementation manner of the third aspect respectively correspond to the first aspect and any implementation manner of the first aspect. For the technical effects corresponding to the third aspect and any implementation manner of the third aspect, reference can be made to the technical effects corresponding to the first aspect and any implementation manner of the first aspect above, which will not be elaborated here.
[0074] In a fourth aspect, an embodiment of the present application provides a chip, including one or more interface circuits and one or more processors; the one or more processors receive or send data through the one or more interface circuits, and when the one or more processors execute computer instructions, the steps of the method in the first aspect or any possible implementation manner of the first aspect are executed.
[0075] The fourth aspect and any implementation manner of the fourth aspect respectively correspond to the first aspect and any implementation manner of the first aspect. For the technical effects corresponding to the fourth aspect and any implementation manner of the fourth aspect, reference can be made to the technical effects corresponding to the first aspect and any implementation manner of the first aspect above, which will not be elaborated here.
[0076] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, which, when running on a computer or a processor, causes the computer or the processor to execute the method in the first aspect or any possible implementation manner of the first aspect.
[0077] The fifth aspect and any implementation manner of the fifth aspect respectively correspond to the first aspect and any implementation manner of the first aspect. For the technical effects corresponding to the fifth aspect and any implementation manner of the fifth aspect, reference may be made to the technical effects corresponding to the first aspect and any implementation manner of the first aspect above, which will not be elaborated herein.
[0078] In a sixth aspect, an embodiment of the present application provides a computer program product including computer instructions, which, when executed by a computer or a processor, cause the computer or the processor to execute the method in the first aspect or any possible implementation manner of the first aspect.
[0079] The sixth aspect and any implementation manner of the sixth aspect respectively correspond to the first aspect and any implementation manner of the first aspect. For the technical effects corresponding to the sixth aspect and any implementation manner of the sixth aspect, reference may be made to the technical effects corresponding to the first aspect and any implementation manner of the first aspect above, which will not be elaborated herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0080] Figure 1 Schematic diagram of the framework of the exemplary image generation system 100;
[0081] Figure 2A Schematic diagram of the exemplary image generation process;
[0082] Figure 2B Schematic diagram of the exemplary image generation process;
[0083] Figure 3A Schematic diagram of the exemplary mobile phone interface;
[0084] Figure 3B Schematic diagram of the position information of the exemplary instance;
[0085] Figure 3C Schematic diagram of the exemplary mobile phone interface;
[0086] Figure 3D Schematic diagram of the exemplary mobile phone interface;
[0087] Figure 4A Schematic diagram of the structure of the exemplary diffusion network;
[0088] Figure 4B Schematic diagram of the training process of the exemplary diffusion network;
[0089] Figure 5A Schematic diagram of the exemplary image generation process 500;
[0090] Figure 5B Schematic diagram of the processing processes of the exemplary instance processing module and the global integration module;
[0091] Figure 5C Schematic diagram of the processing processes of the exemplary instance processing module and the global integration module;
[0092] Figure 6A Schematic diagram of the exemplary image generation process 600;
[0093] Figure 6B Schematic diagram of the processing processes of the exemplary instance processing module and the global integration module;
[0094] Figure 6C Schematic diagram of the processing processes of the exemplary instance processing module and the global integration module;
[0095] Figure 7 Schematic diagram of the exemplary image generation process 700;
[0096] Figure 8A Schematic diagram of the exemplary image generation process 800;
[0097] Figure 8B Schematic diagram of the exemplary mask map generation process;
[0098] Figure 8C Schematic diagram of the exemplary global integration information generation process;
[0099] Figure 8D Exemplarily shows a schematic diagram of the image generation process;
[0100] Figure 9A Schematic diagram of the exemplary target image;
[0101] Figure 9B Schematic diagram of the exemplary target image;
[0102] Figure 9C Schematic diagram of the exemplary target image;
[0103] Figure 9D Schematic diagram of the exemplary target image;
[0104] Figure 10A Schematic diagram of the exemplary target image;
[0105] Figure 10B Schematic diagram of the target image shown by way of example;
[0106] Figure 10C Schematic diagram of the target image shown by way of example;
[0107] Figure 10D Schematic diagram of the target image shown by way of example;
[0108] Figure 10E Schematic diagram of the target image shown by way of example;
[0109] Figure 10F Schematic diagram of the target image shown by way of example;
[0110] Figure 11 Schematic diagram of the image generation device 1100 shown by way of example;
[0111] Figure 12 Schematic diagram of the structure of the device shown by way of example. Detailed implementation manners
[0112] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts fall within the protection scope of the present application.
[0113] The term "and / or" in this document is only used to describe the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone.
[0114] The terms "first", "second", etc. in the description and claims of the embodiments of the present application are used to distinguish different objects, rather than to describe the specific order of the objects. For example, the first target object and the second target object are used to distinguish different target objects, rather than to describe the specific order of the target objects.
[0115] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly, using words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.
[0116] In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality" refers to two or more. For example, a plurality of processing units refers to two or more processing units; a plurality of systems refers to two or more systems.
[0117] Figure 1 It is a schematic diagram of the framework of the exemplary image generation system 100. Figure 1 The illustrated image generation system 100 includes a terminal device and a server 110.
[0118] Exemplarily, the terminal device includes, but is not limited to: a mobile phone 121, a laptop computer (personal computer) 122, a tablet computer 123, and wearable devices, etc., and the present application does not limit this.
[0119] Exemplarily, the specific implementation form of the server 110 can be a cloud server, a physical (independent) server, a cluster server, etc., and the present application does not limit this.
[0120] Exemplarily, the user can use the terminal device to generate the image required by himself / herself (which can be called the target image later) in the image generation application (Application, APP) / tool / web page. Exemplarily, the user can input the description information of the instance and / or the global description information of the target image in the image generation APP / tool / web page; then, the image generation operation can be executed. The terminal device can respond to the operation behavior of the user and send the description information of the instance and / or the global description information of the target image to the server 110 through the network. Subsequently, the server 110 can generate the target image according to the description information of the instance and / or the global description information of the target image; then, send the target image to the terminal device through the network. The terminal device can display the target image in the interface of the image generation APP / tool / web page.
[0121] Exemplarily, the description information of the instance can include the attribute information of the instance and the location information of the instance. The global description information of the target image can include at least one of foreground description information or background description information. Among them, the foreground description information can include at least one of the attribute information of all instances included in the target image, the positional relationship between each instance, and the action relationship between each instance.
[0122] Refer to Figure 1, Exemplarily, the user uses an image generation APP in the mobile phone 121 to generate a target image. Exemplarily, the user can input the attribute information of instance 1 (i.e., an image of a cat) and the location information of instance 1 (i.e., location 1); and the attribute information of instance 2 (i.e., the text "a pumpkin carriage") and the location information of instance 2 (i.e., location 2). Then, the user can perform an image generation operation; the mobile phone 121 can respond to the user's operation behavior and send to the server 110 via the network: an image of a cat and location 1, and the text "a pumpkin carriage" and location 2. Subsequently, the server 110 can generate a target image 11 containing a cat and a pumpkin carriage according to the information sent by the mobile phone 121; then, send the target image 11 to the mobile phone 121 via the network. The image generation APP in the mobile phone 121 can display the target image 11.
[0123] Refer to Figure 1 , Exemplarily, the user uses an image generation web page in the personal computer 122 to generate a target image. Exemplarily, the user can input the global description information of the target image (i.e., the text "4 boats and a bench"). Then, the user can perform an image generation operation; the personal computer 122 can respond to the user's operation behavior and send to the server 110 via the network: the text "4 boats and a bench". Subsequently, the server 110 can generate a target image 12 containing 4 boats and a bench according to the information sent by the personal computer 122; then, send the target image 12 to the personal computer 122 via the network. The image generation web page in the personal computer 122 can display the target image 12.
[0124] Refer to Figure 1 , Exemplarily, the user uses an image generation web page in the tablet computer 123 to generate a target image. Exemplarily, the user can input the global description information of the target image (i.e., the text "a red rabbit wearing a white wig"), the attribute information of instance 1 (the text "a red rabbit") and the location information of instance 1 (i.e., location 3), and the attribute information of instance 2 (i.e., the text "a white wig") and the location information of instance 2 (i.e., location 4). Then, the user can perform an image generation operation; the tablet computer 123 can respond to the user's operation behavior and send to the server 110 via the network: the text "a red rabbit wearing a white wig", the text "a red rabbit" and location 3, and the text "a white wig" and location 4. Subsequently, the server 110 can generate a target image 13 containing a red rabbit wearing a white wig according to the information sent by the tablet computer 123; then, send the target image 13 to the tablet computer 123 via the network. The image generation web page in the tablet computer 123 can display the target image 13.
[0125] It should be understood that the user can also use other terminal devices to generate a target image through interaction with the server, and the present application places no restrictions thereon.
[0126] It should be noted that Figure 1 The figure shows the image generation method involved in the present application executed by the server 110; it should be understood that the image can also be generated locally by the terminal device, that is, the image generation method involved in the present application is executed by the terminal device. In addition, the image generation method involved in the present application can also be executed jointly by the terminal device and the server. For example, after the terminal device preprocesses the description information of the instance input by the user and / or the global description information of the target image, it is sent to the server; then, the server generates the target image based on the preprocessed information and returns the generated target image to the terminal device for display. That is to say, the present application places no restrictions on the electronic device that executes the image generation method involved in the present application.
[0127] The following describes the image generation process of the present application.
[0128] Figure 2A It is a schematic diagram of an exemplary image generation process.
[0129] S201, obtain multiple information groups; where each information group includes the attribute information and location information of an instance.
[0130] In one possible way, the user can independently input the description information of each instance in the image generation APP. In this way, the description information of the instance can include the attribute information of the instance and the location information of the instance. Exemplarily, the attribute information and location information of an instance can form an information group; that is to say, the user can independently input multiple information groups, or in other words, the user can input the description information of any instance separately from the description information of other instances.
[0131] Exemplarily, the instance can be any object; for example, a person, an animal, a plant, a hat, a stool, a head, limbs, a tail, eyes, etc.; the present application places no restrictions thereon.
[0132] Exemplarily, the attribute information of the instance can include the information for describing the attributes of the instance, where the instance attributes can include but are not limited to: size, type (such as name), color, material, style, etc., and the present application places no restrictions thereon. Exemplarily, the attribute information of the instance can be text, or an image, etc., and the present application places no restrictions thereon.
[0133] Exemplarily, the location information of the instance can include but is not limited to: the bounding box information of the instance, the contour information of the instance, etc., and the present application places no restrictions thereon.
[0134] Exemplarily, the position information of the instance can be text (e.g., coordinates) or an image, and the present application does not limit this.
[0135] Figure 3A It is a schematic diagram of a mobile phone interface shown exemplarily.
[0136] Referring to Figure 3A , exemplarily, the image generation main interface 301 may include but is not limited to: an attribute editing box 11, a position editing box 12, an attribute editing box 21, a position editing box 22, image options 13, 14, 23, 24, and a confirmation option 302, etc., and the present application does not limit this.
[0137] Method (1):
[0138] Referring to Figure 3A , exemplarily, the user can input the attribute information of instance 1 in the attribute editing box 11 and the position information of instance 1 in the position editing box 12. For example, input "a blue cat" in the attribute editing box 11 and input the position information "(x11, y11), (x12, y12), (x13, y13), (x14, y14)" in the position editing box x12. Among them, (x11, y11), (x12, y12), (x13, y13), (x14, y14) are the coordinates of the four vertices (A1, A2, A3, and A4) of the minimum bounding box of the cat, as Figure 3B (1) shown.
[0139] Continuing to refer to Figure 3A , exemplarily, the user can input the attribute information of instance 2 in the attribute editing box 21 and the position information of instance 2 in the position editing box 22. For example, input "a green dog" in the attribute editing box 21 and input the position information "(x21, y21), (x22, y22), (x23, y23), (x24, y24)" in the position editing box 22.
[0140] Method (2):
[0141] Continuing to refer to Figure 3A , exemplarily, the user can click on the image option 13, and the mobile phone can respond to the user's operation and display the album main interface; the user can select an image of instance 1 (such as an image of a cat). And can input the position information "(x11, y11), (x12, y12), (x13, y13), (x14, y14)" of instance 1 in the position editing box 12. In this case, the user can input the attribute information of instance 1 in the attribute editing box 11 or not input the attribute information of instance 1, and the present application does not limit this.
[0142] Continue to refer to Figure 3A Exemplarily, the user can click on the image option 23, and the mobile phone can display the main interface of the photo album in response to the user's operation; the user can select an image of Instance 2 (such as an image of a dog). And the location information of Instance 2, i.e., "(x21, y21), (x22, y22), (x23, y23), (x24, y24)", can be entered in the location edit box 22. In this case, the user can enter the attribute information of Instance 2 in the attribute edit box 21, or can not enter the attribute information of Instance 2, and this application does not limit this.
[0143] Method (3):
[0144] Continue to refer to Figure 3A Exemplarily, the user can enter the attribute information of Instance 1 (such as "a blue cat") in the attribute edit box 11. And can click on the image option 14, the mobile phone can display the main interface of the photo album in response to the user's operation; the user can select a mask image of Instance 1 (which can be used to describe the contour information of Instance 1); can refer to Figure 3B as shown in Figure 3B In (2), a mask image of a cat is shown; the pixel values of the pixel points within the contour box of the cat in the mask image of the cat are 0, and the pixel values of the pixel points outside the contour box of the cat are 1. It should be understood that the pixel values of the pixel points within the contour box of the instance in the mask image can be 1, and the pixel values of the pixel points outside the contour box of the instance in the mask image can be 0; this embodiment of the application does not limit this. In this case, the user can enter the location information of Instance 1 in the location edit box 12, or can not enter the location information of Instance 1, and this application does not limit this.
[0145] Continue to refer to Figure 3A Exemplarily, the user can enter the attribute information of Instance 2 (such as "a green dog") in the attribute edit box 21. And can click on the image option 24, the mobile phone can display the main interface of the photo album in response to the user's operation; the user can select a mask image of Instance 2 (which can be used to describe the location information of the contour of Instance 2); for example, a mask image of a dog. In this case, the user can enter the location information of Instance 2 in the location edit box 22, or can not enter the location information of Instance 2, and this application does not limit this.
[0146] Method (4):
[0147] Continue to refer to Figure 3A, exemplarily, the user can click on image option 13, and the mobile phone can display the album main interface in response to the user's operation; the user can select an image of Instance 1 (such as an image of a cat). And click on image option 14, and the mobile phone can display the album main interface in response to the user's operation; the user can select a mask image of Instance 1 (such as a mask image of a cat).
[0148] Continue to refer to Figure 3A , exemplarily, the user can click on image option 23, and the mobile phone can display the album main interface in response to the user's operation; the user can select an image of Instance 2 (such as an image of a dog). And click on image option 24, and the mobile phone can display the album main interface in response to the user's operation; the user can select a mask image of Instance 2 (such as a mask image of a dog).
[0149] It should be noted that after the user clicks on image option 13 or image option 23, the mobile phone can display the album main interface in response to the user's operation, and the user can select a pose image of an instance.
[0150] It should be noted that when the user needs to generate a target image containing a green cat and a blue cat, the user also needs to input the attribute information and position information of 2 instances. Among them, the attribute information of Instance 1 is "a green cat", and the position information of Instance 1 is "(x31,y31), (x32,y32), (x33,y33), (x34,y34)". The attribute information of Instance 2 is "a blue cat", and the position information of Instance 2 is "(x41,y41), (x42,y42), (x43,y43), (x44,y44).
[0151] It should also be noted that when the user needs to generate a target image containing a cat with one blue eye and one green eye, the user needs to input the attribute information and position information of 3 instances. Among them, the attribute information of Instance 1 is "a blue eye", and the position information of Instance 1 is "(x51,y51), (x52,y52), (x53,y53), (x54,y54)". The attribute information of Instance 2 is "a green eye", and the position information of Instance 2 is "(x61,y41), (x62,y62), (x63,y63), (x64,y64). The attribute information of Instance 3 is "a cat", and the position information of Instance 3 is "(x71,y71), (x72,y72), (x73,y73), (x74,y74).
[0152] It should also be noted that when the user needs to generate a target image including a red rabbit wearing a white hat, the user needs to input the attribute information and position information of 2 instances. Among them, the attribute information of instance 1 is "a white hat", and the position information of instance 1 is "(x81, y81), (x82, y82), (x83, y83), (x84, y84)". The attribute information of instance 2 is "a red rabbit", and the position information of instance 2 is "(x91, y41), (x92, y92), (x93, y93), (x94, y94).
[0153] After the user completes the input of the attribute information and position information of all instances, the user can click the confirmation option 302; the mobile phone can respond to the user's operation information, and the mobile phone can obtain multiple information groups, and then execute S202-S203.
[0154] In a possible way, the user can input the global description information of the target image in the image generation APP. Exemplarily, the global description information of the target image can be text or image, and the present application does not limit this.
[0155] Method (1):
[0156] Exemplarily, the global description information input by the user is text, and the global description information includes foreground description information, and the foreground description information includes the attribute information of multiple instances.
[0157] Referring to Figure 3C (1), exemplarily, the image generation main interface 301 may include but is not limited to: an edit box 303, an image option 304, a confirmation option 302, etc., and the present application does not limit this. For example, the user can input "a blue dog and a green dog" in the edit box 303.
[0158] Method (2): The global description information input by the user is text, the global description information includes foreground description information, the foreground description information includes the attribute information of multiple instances, and the positional relationship of each instance.
[0159] Referring to Figure 3C (2), exemplarily, the image generation main interface 301 may include but is not limited to: an edit box 303, an image option 304, a confirmation option 302, etc., and the present application does not limit this. For example, the user can input "a blue dog is on the left of a green dog" in the edit box 303.
[0160] Method (3): The global description information input by the user is text, the global description information includes foreground description information, the foreground description information includes the attribute information of multiple instances, and the action relationship of each instance.
[0161] Reference Figure 3C (3), exemplarily, the main image generation interface 301 may include but is not limited to: an edit box 303, image options 304, a confirmation option 302, etc., and the present application does not limit this. For example, the user can enter "A player is playing football" in the edit box 303.
[0162] Method (4): The global description information input by the user is text, and the global description information includes foreground description information and background description information. The foreground description information includes the attribute information of multiple instances and the positional relationship of each instance.
[0163] Reference Figure 3C (4), exemplarily, the main image generation interface 301 may include but is not limited to: an edit box 303, image options 304, a confirmation option 302, etc., and the present application does not limit this. For example, the user can enter "A blue dog is on the left of a green dog on the street" in the edit box 303.
[0164] It should be understood that other global description information can also be input, and the embodiments of the present application do not limit this.
[0165] After the user completes the input of the global description information of the target image, the user can click the confirmation option 302. The mobile phone can respond to the user's operation information and execute the image generation method involved in the present application, that is, execute S201 to S203. One way to implement S201 can be to parse the global description information, extract the attribute information of each instance; and predict the instance positions according to the global description information to obtain the position information of each instance; then, the attribute information and position information of one instance can be used to form an information group; in this way, the terminal device can obtain multiple information groups.
[0166] Method (5): The global description information input by the user is an image.
[0167] Reference Figure 3C (1) to 3C(4), exemplarily, the user can click the image option 304, and the mobile phone can respond to the user's operation and display the main album interface; the user can select an image (the background of the target image is expected to be similar to the background of this image, and the foreground of the target image is expected to be similar to the foreground of this image). Then, the user can click the confirmation option 302. The mobile phone can respond to the user's operation information and execute the image generation method involved in the present application, that is, execute S201 to S203. One way to implement S201 can be to detect this image, extract the background description information, the attribute information and position information of each instance, and the description text such as the positional relationship and action relationship between each instance. Then, the attribute information and position information of one instance can be used to form an information group; in this way, the terminal device can obtain multiple information groups.
[0168] In one possible way, the user can input the global description information of the target image and independently input the description information of each instance in the target image in the image generation APP.
[0169] Referring to Figure 3D , exemplarily, the main image generation interface 301 may include but is not limited to: the attribute editing box 11, the position editing box 12, the attribute editing box 21, the position editing box 22, the image option 13, the image option 14, the image option 23, the image option 24, the global description information editing box 31, the image option 32, and the confirmation option 302, etc. The present application does not limit this.
[0170] For example, input "a blue cat" in the attribute editing box 11, and input the position information "(x11,y11), (x12,y12), (x13,y13), (x14,y14)" in the position editing box 12. Input "a green dog" in the attribute editing box 21, and input the position information "(x21,y21), (x22,y22), (x23,y23), (x24,y24)" in the position editing box 22. Input "a blue cat is on the left of a green dog on the street" in the global description information editing box 31.
[0171] It should be understood that Figure 3D the forms of the attribute information and position information of each instance input by the user in Figure 3A and Figure 3C , and the form of the global description information, can refer to the above description for
[0172] It should be understood that the present application does not limit the form of the main image generation interface, nor the form of the information input by the user.
[0173] Exemplarily, the number of information groups can be represented by n (n is an integer greater than 1), that is, information group 1, information group 2,..., information group n can be obtained. Among them, information group 1 includes the attribute information and position information of instance 1, information group 2 includes the attribute information and position information of instance 2,..., and information group n includes the attribute information and position information of instance n.
[0174] S202, generate the image feature of one instance according to the attribute information of one instance included in each of the multiple information groups.
[0175] Exemplarily, the present application can generate the image features of an instance by using the attribute information of an instance included in an information group; thus, according to the description information of multiple information groups, the image features of multiple instances can be generated. That is to say, S202 can generate the image features of n instances: the image features of instance 1, the image features of instance 2,..., the image features of instance n.
[0176] Exemplarily, a neural network can be used to process the attribute information of an instance included in an information group and output the image features of an instance.
[0177] S203. Integrate the image features of multiple instances according to the location information of multiple instances included in multiple information groups to generate a target image.
[0178] Exemplarily, after obtaining the image features of each instance in multiple instances, the image features of multiple instances can be integrated to obtain the image features of the target image; then, the image features of the target image can be processed to generate the target image.
[0179] Exemplarily, according to the location information of multiple instances, the image features of multiple instances can be integrated so that each instance in the generated target image is at the position corresponding to its location information.
[0180] Exemplarily, a neural network can be used to process the location information of n instances and the image features of n instances to obtain the target image.
[0181] Exemplarily, the target image generated by S203 includes these n instances, and the positions of the n instances correspond one-to-one with the location information in the n information groups, and the attributes of the n instances in the target image correspond one-to-one with the attribute information in the n information groups.
[0182] Figure 2B For exemplary illustration, a schematic diagram of the image generation process is shown.
[0183] Refer to Figure 2B , exemplarily, the user inputs the global description information of the target image "A blue cat is on the left of a green dog". Among them, according to the global description information, the position layout and attribute information of instance 1 (cat) and instance 2 (dog) can be determined. Then, instance division can be performed to obtain the attribute information of the cat "A blue cat" and the attribute information of the dog "A green dog".
[0184] Refer to Figure 2B, the attribute information of a cat, "a blue cat", can be processed to output the image features of the cat (i.e., Blue Cat’s Feature). And the attribute information of a dog, "a green dog", can be processed to output the image features of the dog (i.e., Green Dog’s Feature).
[0185] Referring to Figure 2B , exemplarily, according to the position layout of the cat and the dog, the image features of the cat and the dog can be integrated to obtain the image features of the target image; afterwards, the image features of the target image can be processed to obtain the target image. For example, the target image can be Figure 2B Image 1 or Image 2 in
[0186] In summary, the present application adopts a "divide-conquer-combine" scheme to implement image generation. Among them, "divide" can correspond to S201, which can refer to "decomposing and partitioning single-instance tasks", that is, decomposing the description information of each instance, that is, decomposing multiple information groups, and a single information group can correspond to a single-instance task. It should be noted that when multiple information groups are input by the user, that is, the user realizes the decomposition of multiple information groups. "Conquer" can correspond to S202, which can refer to "generating and processing each single-instance task one by one", that is, executing the single-instance task multiple times, and each execution of the single-instance task can generate the image features of an instance. It should be noted that S202 can execute multiple single-instance tasks in parallel. "Combine" can correspond to S203, which can refer to "integrating the generation results of each instance", that is, integrating the image features of multiple instances to generate the target image. In this way, complex multi-instances can be decomposed into multiple relatively simple single-instance generation tasks; and during the execution of the single-instance generation task, since each information group only includes the attribute information of the single instance itself, there is no problem of attribute leakage; furthermore, the present application can ensure that each instance can be generated according to the attributes specified by the user; thus, during the process of generating the target image, precise control of the attributes of each instance in the target image can be achieved.
[0187] In addition, during the global integration process, the position information of multiple instances is utilized, so the present application can also achieve precise control of the positions of each instance in the target image.
[0188] It should be noted that when only one information group is obtained, the image generation method of the present application can also be adopted to generate the target image.
[0189] Exemplarily, the present application can adopt a diffusion network to implement S202 and S203. Furthermore, the structure of the diffusion network and the training process of the diffusion network can be described first. The following takes the diffusion network implemented by U-NET as an example for illustration.
[0190] Figure 4ASchematic structural diagram of the diffusion network shown exemplarily.
[0191] Referring to Figure 4A (1), exemplarily, U-NET (U-shaped structure) may include multiple blocks. The left half and the right half of U-NET may each include g (g is a positive integer) blocks. Among them, each block may include two parts: the first part and the second part (as shown by the two white rectangles in each block in Figure 4A (1)); both the first part and the second part may include one or more network layers (such as convolutional layers, activation layers, etc.).
[0192] Exemplarily, the input Y of U-NET is a noisy image and a time node, and the output X may be the noise in the predicted noisy image.
[0193] In one possible way, an instance processing module and a global integration module may be inserted into each block in U-NET, as shown by the black rectangles in Block_1 to Block_g in Figure 4A (1). In this case, the diffusion network may include U-NET, an instance processing module, and a global integration module.
[0194] Exemplarily, the instance processing module may include one or more cross-attention modules. Optionally, the instance processing module may further include an enhanced attention module. Among them, the instance processing module may be implemented by a neural network, so the instance processing module may also be called an instance processing network.
[0195] Exemplarily, the global integration module may be an instance of a neural network, so the global integration module may also be called a global integration network. For example, the global integration module may include 2 self-attention modules and a softmax function.
[0196] Referring to Figure 4A (2), exemplarily, the instance processing module may include m + 1 inputs; where m is an integer greater than or equal to n. Among them, the input R0 may be the output of the first part of the block where the instance processing module is located (which may also be called the output of the previous network layer of the instance processing module). Any one of the inputs R1 to Rm may be used to input the attribute information of an instance included in an information group.
[0197] Exemplarily, the instance processing module may include m outputs. Any one of the outputs U1 to Um can be used to output the image features of an instance. Among them, the output U1 is the output obtained by the instance processing module processing the input R0 and the input R1; the output U2 is the output obtained by the instance processing module processing the input R0 and the input R2,..., and the output Um is the output obtained by the instance processing module processing the input R0 and the input Rm. It should be noted that the instance processing module can perform m-way processing simultaneously, and these m-way processes are independent of each other.
[0198] Exemplarily, the global integration module may include m + 1 inputs. Among them, the m outputs of the instance processing module are the m inputs of the global integration module: the output U1 of the instance processing module serves as the input U1 of the global integration module; the output U2 of the instance processing module serves as the input U2 of the global integration module;...; the output Um of the instance processing module serves as the input Um of the global integration module. In addition, the location information of all instances can be used as the input E0 of the global integration module. The global integration module may include one output, and this output can be used as the input of the second part of the Block where the global integration module is located (which can also be referred to as the input of the subsequent network layer of the global integration module). It should be understood that the processing performed by the global integration module is to integrate multiple inputs into one output.
[0199] In one possible way, an instance processing module, a global integration information generation module, and a global integration module can be inserted into each Block in the U-NET, as shown by the black rectangles in Block_1 to Block_g in Figure 4A (1). In this case, the diffusion network may include the U-NET, the instance processing module, the global integration information generation module, and the global integration module.
[0200] Exemplarily, the global integration information generation module may include one or more self-attention modules. Among them, the global integration information generation module can be implemented using a neural network, so the global integration information generation module can be referred to as the global integration information generation network.
[0201] Exemplarily, Figure 4A (3) The instance processing module is similar to the instance processing module in Figure 4A (2), and will not be elaborated here.
[0202] Referring to Figure 4A(3) The global integration information generation module may include two inputs: input R0 and input H0. Among them, input R0 may be the output of the first part of the Block where the global integration information generation module is located (which may also be referred to as the output of the previous network layer of the global integration information generation module). Input H0 may be the mask map corresponding to each pixel point in the target image. Among them, the mask map corresponding to each pixel point in the target image is generated based on the position information of all instances in the target image, which will be specifically described later. The global integration information generation module includes one output, that is, output D0 (D0 is the integrated information).
[0203] Exemplarily, the global integration module may include m + 2 inputs. Among them, the m outputs of the instance processing module are the m inputs of the global integration module; the 1 output D0 of the global integration information generation module is the 1 input E0 of the global integration module; and, the position information of all instances is used as the input E0 of the global integration module. The global integration module may include one output, which is used as the input of the second part of the Block where the global integration module is located.
[0204] Figure 4B It is a schematic diagram of the training process of the diffusion network shown exemplarily.
[0205] Referring to Figure 4B , exemplarily, the training process of the diffusion network may include a forward process (or referred to as a diffusion process, as shown in Figure 4B (1)) and a reverse process (or referred to as a generation process, as shown in Figure 4B (2)).
[0206] Exemplarily, the training process of the diffusion network may be as shown in S41~S46 below:
[0207] S41, Obtain the training image and the global description information of the training image.
[0208] Exemplarily, multiple text-image data pairs may be obtained. Each text-image data pair may include a training image and the global description information of the training image.
[0209] S42, Parse out multiple information groups from the global description information of the training image.
[0210] Exemplarily, each training image may include one or more instances. From the global description information of the training image, the attribute information of each instance contained in the training image can be extracted; and the position information of each instance contained in the training image (which can be called the bounding box) can be detected from the training image. Among them, the attribute information and position information of each instance contained in the training image can form an information group; in this way, for each training image, multiple information groups can be determined.
[0211] S43. Sample a noisy step t from time node 0 to time node T, and gradually add noise to the training image to obtain a noisy image / noisy latent space feature.
[0212] Refer to Figure 4B (1). Exemplarily, Z0 is the training image, and the corresponding time node is 0. Z0 can be called the state at time node 0.
[0213] Exemplarily, between time node 0 and time node 1, random Gaussian noise can be sampled to obtain a preset noise ε ∼ N(0, 1) (N(0, 1) is the Gaussian distribution, and ε ∼ N(0, 1) means the preset noise conforms to the Gaussian distribution). Then, based on this preset noise, Z0 is noise-added to obtain the state Z1 at time node 1; among them, the specific noise-adding method can refer to the description in the prior art and will not be elaborated here.
[0214] Exemplarily, between time node 1 and time node 2, random Gaussian noise can be sampled to obtain a preset noise ε ∼ N(0, 1); then, based on this preset noise, Z1 is noise-added to obtain the state Z2 at time node 2; among them, the specific noise-adding method can refer to the description in the prior art and will not be elaborated here.
[0215] And so on, after performing noise addition t times, the state Z at time node t is obtained t . Among them, Z t can be a noisy image / noisy latent space feature.
[0216] S44. Input the noisy image / noisy latent space feature, the global description information of the training image, multiple information groups, and the noisy step into the diffusion network, and the diffusion network performs t generation processes to obtain the predicted noise ε' of the noisy image / latent space feature at time node 0.
[0217] Exemplarily, between time node t and time node t - 1, the noisy image / noisy latent space feature at time node t, the global description information of the training image, multiple information groups, and t can be input into the diffusion network to obtain the state X at time node t - 1 t-1 , and X t-1The predicted noise in
[0218] Among them, the attribute information in each of the n information groups is an input to the instance processing module, that is, the attribute information in the n information groups corresponds to the inputs R1 to Rn of the instance processing module. The global description information of the training image is the input R(n + 1) of the instance processing module.
[0219] Exemplarily, between time node t - 1 and time node t - 2, the state X at time node t - 1 t-1 , the global description information of the training image, multiple information groups, and t - 1 are input into the diffusion network to obtain the state X at time node t - 2 t-2 , and the predicted noise in X t-2 The predicted noise in
[0220] By analogy, after the diffusion network performs t generation processes, the predicted noise ε' of the noisy image / latent space feature at time node 0 can be obtained.
[0221] S45, Calculate the loss value according to the predicted noise and the true noise.
[0222] Exemplarily, calculate the L1 distance and / or L2 distance between the predicted noise ε' of the noisy image / latent space feature at time node 0 and the true noise sampled between time node 1 and time node 0. Determine the loss value according to the L1 distance and / or L2 distance.
[0223] S46, Perform backpropagation according to the loss value and adjust the network parameters of the diffusion network.
[0224] By analogy, loop through S42 - S46 until the preset conditions are met. Among them, the preset conditions can be set according to requirements. For example, the number of training times reaches the threshold; or, the loss value is less than or equal to the loss threshold; etc. This application does not limit this.
[0225] After that, the trained diffusion network can be used to implement image generation; specifically, it can refer to the following description.
[0226] Figure 5A It is a schematic diagram of the image generation process 500 shown exemplarily. Figure 5A In the embodiment of , the diffusion network adopted includes a U - NET, an instance processing module, and a global integration module. The main image generation interface corresponding to the image generation process 500 can be as Figure 3A shown.
[0227] S501, Receive multiple information groups input by the user.
[0228] Exemplarily, the number of information groups can be represented by n.
[0229] For example, when n = 2, information group 1 includes: "a blue cat" and "(x11, y11), (x12, y12), (x13, y13), (x14, y14)"; information group 2 includes: "a green dog" and "(x21, y21), (x22, y22), (x23, y23), (x24, y24)".
[0230] S502: Input multiple information groups into the diffusion network, and let the diffusion network perform T1 generation processes to obtain a first intermediate image.
[0231] Exemplarily, the diffusion network in S502 includes a U-NET, an instance processing module, and a global integration module.
[0232] Exemplarily, assume that the total number of times the diffusion network performs the generation process is T (T is a positive integer); that is, during the process from time node T to time node 0, the diffusion network performs T generation processes. Among them, between time node T and time node T - 1, the diffusion network performs the first generation process; between time node T - 1 and time node T - 2, the diffusion network performs the second generation process; and so on.
[0233] Exemplarily, T1 is a positive integer less than or equal to T. When T1 = T, then S5 is not required to be executed; when T1 is less than T, then S503 can be executed. This application takes the case where T1 is less than T as an example for illustration.
[0234] [[ID=It is described below by taking the processing processes of the instance processing module and the global integration module embedded in Block_1 of the U-NET during the i-th generation process of the diffusion network as an example. Among them, i is a positive integer less than or equal to T1.
[0235] S5021: Take the first intermediate image generated in the previous time as one path and the attribute information of multiple instances as multiple paths and input them into the diffusion network. Let the instance processing module in the diffusion network perform multiple feature generations to obtain image features of multiple instances; among them, the first intermediate image generated in the previous time and the attribute information of one instance are used for one feature generation, and the attribute information of one instance is one path of input.
[0236] Exemplarily, the first intermediate image (subsequently referred to as the (i - 1)-th first intermediate image) generated by the diffusion network during the (i - 1)-th generation process and the time node can be taken as input Y and input into the diffusion network; that is, the (i - 1)-th first intermediate image and the time node are input into the first part of Block_1. After being processed by the first part of Block_1, feature 1 is output. It should be noted that when i is equal to 1, the first intermediate image generated by the (i - 1)-th generation process can refer to a preset noisy image.
[0237] After that, the feature 1 is used as the input R0 of the instance processing module in Block_1 and input into the instance processing module. In addition, the attribute information in the n information groups can also be used as the n-channel inputs of the instance processing module, that is, the attribute information of the first information group is used as the input R1 of the instance processing module, the attribute information of the second information group is used as the input R2 of the instance processing module,..., and the attribute information of the nth information group is used as the input Rn of the instance processing module.
[0238] That is to say, in the image generation process 500, the instance processing module includes n + 1-channel inputs.
[0239] Figure 5B It is a schematic diagram of the processing procedures of the exemplary instance processing module and the global integration module.
[0240] Refer to Figure 5B , for example, the instance processing module can generate features based on the feature 1 output by the first part of Block_1 and the attribute information of instance 1 to obtain the image features of instance 1; generate features based on the feature 1 output by the first part of Block_1 and the attribute information of instance 2 to obtain the image features of instance 2;...; generate features based on the feature 1 output by the first part of Block_1 and the attribute information of instance n to obtain the image features of instance n.
[0241] It can also be understood that the instance processing module performs n-channel processing in parallel. The first-channel processing is to generate features based on the feature 1 output by the first part of Block_1 and the attribute information of instance 1, the second-channel processing is to generate features based on the feature 1 output by the first part of Block_1 and the attribute information of instance 2,..., and the nth-channel processing is to generate features based on the feature 1 output by the first part of Block_1 and the attribute information of instance n.
[0242] It should be noted that Figure 5B shows that the n instance processing modules are essentially the same instance processing module, Figure 5B and the schematic way of
[0243] Figure 5C is to illustrate that the instance processing module processes the attribute information of each instance independently.
[0244] In Among them, n = 2. The attribute information of the instances in information group 1 is "a blue cat", and the attribute information of the instances in information group 2 is "a green dog". The instance processing module can generate features based on feature 1 and "a blue cat" and output the image features of the cat; and can also generate features based on feature 1 and "a green dog" and output the image features of the dog.
[0245] S5022, Input the position information of multiple instances and the image features of multiple instances into the global integration module in the diffusion network to obtain the first intermediate feature.
[0246] Exemplarily, the image features of n instances can be used as the n-way input of the global integration module. That is, the image feature of instance 1 is used as the input U1 of the global integration module, the image feature of instance 2 is used as the input U2 of the global integration module,..., and the image feature of instance n is used as the input Un of the global integration module. And the position information of n instances is used as one-way input (i.e., input E0) of the global integration module. Then, the global integration module can process these n + 1-way inputs to obtain the first intermediate feature. Specifically, the global integration module can integrate the n-way inputs from input U1 to input Un according to input E0 to obtain the first intermediate feature.
[0247] Continue to refer to , Exemplarily, assume that the position information in information group 1 is "(x11, y11), (x12, y12), (x13, y13), (x14, y14)", and the position information in information group 2 is "(x21, y21), (x22, y22), (x23, y23), (x24, y24)". Then, (x11, y11), (x12, y12), (x13, y13), (x14, y14), (x21, y21), (x22, y22), (x23, y23), (x24, y24) can be used as input E0 and input into the global integration module.
[0248] S5023, Input the first intermediate feature into other modules in the diffusion network to obtain the third intermediate feature.
[0249] Exemplarily, other modules can include the second part of the Block in U-NET.
[0250] Exemplarily, the global integration module can input the first intermediate feature into the second part of Block_1, and the second part of Block_1 processes it and outputs the third intermediate feature to the 2nd Block.
[0251] After that, the 2nd block to the 2g-th block can be processed sequentially in the same way as Block_1 to complete the i-th generation process of the diffusion network, and the i-th first intermediate image is obtained. That is to say, the i-th first intermediate image can refer to the output of the second part of the last block of the U-NET (or the output of the diffusion network when performing the i-th generation process) during the process of the diffusion network performing the i-th generation process.
[0252] By analogy, after the diffusion network performs the generation process T1 times according to the above process, the T1-th first intermediate image can be obtained.
[0253] It should be noted that when T1 equals T, S503 does not need to be executed. At this time, the T1-th first intermediate image can be used as the target image. When T1 is less than T, S503 can be executed.
[0254] S503: Input the first intermediate image into the diffusion network, and the diffusion network performs the generation process T2 times to obtain the target image.
[0255] It should be noted that the diffusion network in S503 includes a U-NET and an instance processing module.
[0256] Exemplarily, S503 may include S5031 to S5033:
[0257] The following takes the processing process of the instance processing module embedded in Block_1 of the U-NET during the process of the diffusion network performing the (T1 + j)-th generation process as an example for explanation. Among them, j is a positive integer less than or equal to T2, T2 is a positive integer, and T1 + T2 = T.
[0258] S5031: Take the second intermediate image generated in the previous time as one path and the attribute information of multiple instances as another path and input them into the diffusion network. The instance processing module generates features according to the second intermediate image generated in the previous time and the attribute information of multiple instances to obtain the second intermediate features.
[0259] Exemplarily, the second intermediate image generated by the diffusion network when performing the (T1 + j - 1)-th generation process (subsequently referred to as the (T1 + j - 1)-th second intermediate image) and the time node can be used as the input Y and input into the diffusion network; that is to say, the (T1 + j - 1)-th second intermediate image and the time node are input into the first part of Block_1. After being processed by the first part of Block_1, the feature 2 is output. It should be noted that when j equals 1, the second intermediate image generated by the (T1 + j - 1)-th generation process is the T-th first intermediate image generated by S502.
[0260] After that, feature 2 is used as the input R0 of the instance processing module in Block_1 and input into the instance processing module; in addition, the attribute information in the n information groups can also be used as an input path of the instance processing module, namely input R1, and input into the instance processing module. Then, the instance processing module can perform a feature generation based on feature 2 and the attribute information in the n information groups to obtain a second intermediate feature.
[0261] S5032: Input the second intermediate feature into other modules to obtain a second intermediate image.
[0262] Exemplarily, other modules may include the second part of the Block in U-NET.
[0263] Exemplarily, the instance processing module can input the second intermediate feature into the second part of Block_1, which is processed by the second part of Block_1, and output a fourth intermediate feature to the second Block.
[0264] After that, the second Block to the 2g-th Block can be processed in sequence according to the manner of Block_1 to complete the (T1 + j)-th generation process of the diffusion network, and obtain the (T1 + j)-th second intermediate image. That is to say, the (T1 + j)-th second intermediate image may refer to the output of the second part of the last Block of U-NET (or, the output of the diffusion network when performing the (T1 + j)-th generation process) during the process of the diffusion network performing the (T1 + j)-th generation process.
[0265] And so on, when the diffusion network performs the generation process T2 times according to the above process, the T2-th second intermediate image can be obtained.
[0266] S5033: Use the second intermediate image obtained by the diffusion network performing the T-th generation process as the target image.
[0267] Exemplarily, that is to use the T2-th second intermediate image as the target image.
[0268] It should be understood that S503 is an optional step.
[0269] Since in the subsequent stage of the chain generation process of the diffusion network, details of each object in the target image, attributes, and light, style, tone, etc. of the whole image can be generated; in this way, without changing or basically not changing the image content, the interaction relationship between each object in the finally obtained target image is natural, without a sense of splicing and splitting, and the tone and light are harmonious.
[0270] It is a schematic diagram of the exemplary image generation process 600. In the embodiments, the diffusion network adopted includes a U-NET / instance processing module and a global integration module. The main image generation interface corresponding to the image generation process 600 may be as follows as shown.
[0271] S601, Receive the global description information of the target image input by the user.
[0272] Exemplarily, the user may input the global description information of the target image in the main image generation interface 301 in .
[0273] S602, Based on the global description information of the target image input by the user, determine multiple information groups.
[0274] Exemplarily, a text parser may be called to perform text parsing according to the global description information to obtain the attribute information of each instance; and a position predictor may be called to perform instance position prediction according to the global description information to obtain the position information of each instance.
[0275] S603, Input the multiple information groups into the diffusion network, and the diffusion network performs T1 generation processes to obtain a first intermediate image.
[0276] Exemplarily, T1 is a positive integer less than or equal to T. When T1 = T, then S604 may not need to be executed; when T1 is less than T, then S604 may be executed. This application is described by taking T1 less than T as an example.
[0277] The following takes the processing processes of the instance processing module and the global integration module embedded in Block_1 of the U-NET during the i-th generation process executed by the diffusion network as an example for description. Wherein, i is a positive integer less than or equal to T1.
[0278] Exemplarily, the implementation manner of S603 may refer to the above S5021~S5023, which will not be elaborated here.
[0279] In a possible manner, S603 may include: Input the multiple information groups and the global description information into the diffusion network, and the diffusion network performs T1 generation processes to obtain an intermediate image; it may refer to the following: S6031~S6034:
[0280] S6031, Use the first intermediate image generated last time as one path and the attribute information of multiple instances as multiple paths to input into the diffusion network, and the instance processing module in the diffusion network performs multiple feature generations to obtain the image features of multiple instances; wherein, the first intermediate image generated last time and the attribute information of one instance are used for one feature generation, and the attribute information of one instance is used as one path of input.
[0281] Exemplarily, S6031 may refer to the description of S5021 above and will not be elaborated here.
[0282] S6032, input the global description information into the instance processing module for one-time feature generation to obtain global image features.
[0283] Exemplarily, the global description information of the target image may be used as the input R(n + 1) of the instance processing module and input into the instance processing module.
[0284] That is to say, in the image generation process 600, the instance processing module includes n + 2 inputs.
[0285] It is a schematic diagram of the processing procedures of the exemplarily shown instance processing module and global integration module.
[0286] Refer to , exemplarily, the instance processing module may perform feature generation based on the feature 1 output from the first part of Block_1 and the attribute information of instance 1 to obtain the image feature of instance 1; perform feature generation based on the feature 1 output from the first part of Block_1 and the attribute information of instance 2 to obtain the image feature of instance 2;......; perform feature generation based on the feature 1 output from the first part of Block_1 and the attribute information of instance n to obtain the image feature of instance n; perform feature generation based on the feature 1 output from the first part of Block_1 and the global description information of the target image to obtain global image features.
[0287] It can also be understood that the instance processing module performs n + 1 parallel processes. The first process is to perform feature generation based on the feature 1 output from the first part of Block_1 and the attribute information of instance 1, the second process is to perform feature generation based on the feature 1 output from the first part of Block_1 and the attribute information of instance 2,......, the nth process is to perform feature generation based on the feature 1 output from the first part of Block_1 and the attribute information of instance n; the (n + 1)th process is to perform feature generation based on the feature 1 output from the first part of Block_1 and the global description information of the target image.
[0288] It should be noted that shows that the n + 1 instance processing modules are essentially the same instance processing module, and the schematic manner is to illustrate that the instance processing module independently processes the attribute information of each instance and the global description information of the target image.
[0289] It is a schematic diagram of the processing procedures of the exemplarily shown instance processing module and global integration module.
[0290] In it, the global description information of the target image is "a blue cat is on the left of a green dog". n = 2. The attribute information of the instance in information group 1 is "a blue cat", and the attribute information of the instance in information group 2 is "a green dog". The instance processing module can generate features based on feature 1 and "a blue cat" to output the image features of the cat; and can generate features based on feature 1 and "a green dog" to output the image features of the dog; generate global image features based on feature 1 and "a blue cat is on the left of a green dog".
[0291] S6033, input the position information of multiple instances, the image features of multiple instances, and the global image features into the global integration module in the diffusion network to obtain the first intermediate feature.
[0292] Exemplarily, compared with S5022, the global integration module in S6033 has one more input, namely the global image feature; the global image feature can be used as the input U(n + 1) of the global integration module. Then, the global integration module can integrate the n inputs from input U1 to input Un according to the position information of multiple instances and the global image feature to obtain the first intermediate feature. For details, please refer to the description of S5022 and will not be elaborated here.
[0293] S6034, input the first intermediate feature into other modules in the diffusion network to obtain the third intermediate feature.
[0294] Exemplarily, S6034 can refer to the description of S5023 above and will not be elaborated here.
[0295] S604, input the first intermediate image into the diffusion network, and the diffusion network performs T2 generation processes to obtain the target image.
[0296] Exemplarily, S604 can refer to the description of S503 above and will not be elaborated here.
[0297] It should be noted that when the global description information of the target image does not include background description information, background prediction can be performed according to the global description information to obtain the background description information of the target object.
[0298] In this way, a target image matching the globally specified description information by the user can be generated.
[0299] It is a schematic diagram of the exemplary image generation process 700. In the embodiment of, the diffusion network adopted includes an instance processing module and a global integration module. The main image generation interface corresponding to the image generation process 700 can be as shown.
[0300] S701 receives multiple information groups input by a user.
[0301] Exemplarily, the number of information groups can be represented by n.
[0302] For example, n = 2. Information group 1 includes: "a blue cat" and "(x11, y11), (x12, y12), (x13, y13), (x14, y14)"; Information group 2 includes: "a green dog" and "(x21, y21), (x22, y22), (x23, y23), (x24, y24)".
[0303] S702 receives global description information of a target image input by a user.
[0304] For example, the global description information of the target image is "a blue cat is on the left of a green dog".
[0305] S703 inputs the multiple information groups and the global description information into a diffusion network, and the diffusion network performs T1 generation processes to obtain a first intermediate image.
[0306] Exemplarily, T1 is a positive integer less than or equal to T. When T1 = T, then S704 can be skipped; when T1 is less than T, then S704 can be executed. This application takes T1 being less than T as an example for illustration.
[0307] The following takes the processing procedures of an instance processing module and a global integration module embedded in Block_1 of U-NET during the i-th generation process performed by the diffusion network as an example for illustration. Wherein, i is a positive integer less than or equal to T1.
[0308] Exemplarily, S703 may include S7031 to S7034:
[0309] S7031 inputs the first intermediate image generated last time as one path and the attribute information of multiple instances as multiple paths into the diffusion network, and the instance processing module in the diffusion network performs multiple feature generations to obtain image features of multiple instances; wherein, the first intermediate image generated last time and the attribute information of one instance are used for one feature generation, and the attribute information of one instance is one path of input.
[0310] S7032 inputs the global description information into the instance processing module for one feature generation to obtain global image features.
[0311] S7033 inputs the position information of multiple instances, the image features of multiple instances, and the global image features into the global integration module in the diffusion network to obtain a first intermediate feature.
[0312] S7034. Input the first intermediate feature into other modules in the diffusion network to obtain a third intermediate feature.
[0313] Exemplarily, S7033 to S7034 may refer to the description of S6033 to S6034 above, and will not be elaborated here.
[0314] S704. Input the first intermediate image into the diffusion network, and the diffusion network performs T2 generation processes to obtain a target image.
[0315] Exemplarily, S704 may refer to the description of S503 above, and will not be elaborated here.
[0316] Exemplarily, global integration information of the target image (which may also be referred to as a global integration template, and is an image feature) may also be generated according to the position information of multiple instances; the global integration information of the target image can be used to guide how to generate the target image using the image features of multiple instances.
[0317] It is a schematic diagram of an exemplary image generation process 800. It is shown based on <> the following. In the embodiment of [], the diffusion network adopted includes a U-NET, an instance processing module, a global integration module, and a global integration information generation module. The main image generation interface corresponding to the image generation process 800 may be as shown.
[0318] S801. Receive multiple information groups input by the user.
[0319] Exemplarily, the number of information groups may be represented by n. [[ID=]]
[0320] For example, n = 2. Information group 1 includes: "a blue cat" and "(x11, y11), (x12, y12), (x13, y13), (x14, y14)"; information group 2 includes: "a green dog" and "(x21, y21), (x22, y22), (x23, y23), (x24, y24)".
[0321] S802. Receive the global description information of the target image input by the user.
[0322] For example, the global description information of the target image is "a blue cat is on the left of a green dog".
[0323] S803. Input the multiple information groups and the global description information into the diffusion network, and the diffusion network performs T1 generation processes to obtain a first intermediate image.
[0324] Exemplarily, in S803, the diffusion network includes a U-NET, an instance processing module, a global integration information generation module, and a global integration module.
[0325] Exemplarily, T1 is a positive integer less than or equal to T. When T1 = T, then S804 may not need to be executed; when T1 is less than T, then S804 may be executed. This application takes the case where T1 is less than T as an example for illustration.
[0326] The following takes the processing procedures of the instance processing module, the global integration information generation module, and the global integration module embedded in Block_1 of the U-NET during the i-th generation processing of the diffusion network as an example for illustration. Wherein, i is a positive integer less than or equal to T1.
[0327] Exemplarily, S803 may include S8031 to S8036:
[0328] S8031, taking the first intermediate image generated last time as one path and taking the attribute information of multiple instances as multiple paths and inputting them into the diffusion network, and performing multiple feature generations by the instance processing module in the diffusion network to obtain the image features of multiple instances; wherein, the first intermediate image generated last time and the attribute information of one instance are used for one feature generation, and the attribute information of one instance is one path of input.
[0329] S8032, inputting the global description information into the instance processing module for one feature generation to obtain the global image features.
[0330] Exemplarily, S8031 to S8032 may refer to the descriptions of S6031 to S6032 above, and will not be elaborated here.
[0331] S8033, generating multiple mask maps according to the position information of multiple instances; wherein, the multiple mask maps correspond one by one to multiple pixel points in the target image; the size of each mask map is the same as the size of the target image, and the mask map corresponding to the first pixel point is the mask map corresponding to the instance to which the first pixel point belongs, and the first pixel point is any pixel point in the target image.
[0332] The following takes n = 2 as an example to illustrate the process of generating the mask map corresponding to each pixel point in the target image.
[0333] It is a schematic diagram of the mask map generation process shown exemplarily.
[0334] Exemplarily, the size of the layout map in is the same as the size of the target image, and the pixel points in the layout map correspond one by one to the pixel points in the target image; The layout diagram includes the position layouts of two instances. Among them, each square in the layout diagram can represent a pixel. The white pixel represents the pixel located in the background of the target image, the gray pixel represents the pixel located within the bounding box of Instance 1, and the black pixel represents the pixel located within the bounding box of Instance 2.
[0335] Exemplarily, for a pixel in the layout diagram, a mask diagram corresponding to the instance (or background) to which the pixel belongs (that is, the mask diagram corresponding to the pixel) can be generated and stored. Assuming that the height of the layout diagram is H and the width is W, then for H * W pixels, H * W mask diagrams can be generated, with one mask diagram corresponding to one pixel.
[0336] It should be noted that when the position information of the instance is the bounding box information of the instance, the mask diagram corresponding to the instance to which pixel P1 belongs can be generated according to the bounding box of the instance to which pixel P1 belongs. Specifically, the pixel values of the pixels in the layout diagram that are within the bounding box of the instance to which pixel P1 belongs can be set to 0; the pixel values of the pixels in the layout diagram that are outside the bounding box of the instance to which pixel P1 belongs can be set to 1, to obtain the mask diagram corresponding to the instance to which pixel P1 belongs. Or the pixel values of the pixels in the layout diagram that are within the bounding box of the instance to which pixel P1 belongs can be set to 1; the pixel values of the pixels in the layout diagram that are outside the bounding box of the instance to which pixel P1 belongs can be set to 0, to obtain the mask diagram corresponding to the instance to which pixel P1 belongs.
[0337] It should be noted that when the position information of the instance is the contour information of the instance, the mask diagram corresponding to the instance to which pixel P2 belongs can be generated according to the contour information of the instance to which pixel P2 belongs. Specifically, the pixel values of the pixels in the layout diagram that are within the contour box of the instance to which pixel P2 belongs can be set to 0; the pixel values of the pixels in the layout diagram that are outside the contour box of the instance to which pixel P2 belongs can be set to 1, to obtain the mask diagram corresponding to the instance to which pixel P belongs. Or the pixel values of the pixels in the layout diagram that are within the contour box of the instance to which pixel P2 belongs can be set to 1; the pixel values of the pixels in the layout diagram that are outside the contour box of the instance to which pixel P2 belongs can be set to 0, to obtain the mask diagram corresponding to the instance to which pixel P belongs.
[0338] Exemplarily, for pixel P3 belonging to the background, the pixel values of the pixels in the layout diagram that are within the bounding box (or contour) of all instances can be set to 0, and the pixel values of the pixels in the layout diagram that are outside the bounding box (or contour) of all instances can be set to 1, to obtain the mask diagram corresponding to the background.
[0339] It should be understood that multiple pixel points belonging to the same instance correspond to the same mask image, or multiple pixel points belonging to the background correspond to the same mask image. Among them, the size of the mask image is the same as the size of the layout image.
[0340] Referring to , for white pixel points, a mask image corresponding to the background can be generated as shown in mask 1. For a black pixel point, a mask image corresponding to instance 2 can be generated as shown in mask 2. For a gray pixel point, a mask image corresponding to instance 1 can be generated as shown in mask 4. Among them, for the two black pixel points located within the bounding box of instance 2, the corresponding mask images are mask 2 and mask 3 respectively, and mask 2 and mask 3 are the same.
[0341] S8034, generate global integration information according to multiple mask images.
[0342] Exemplarily, the feature 1 output by the first part of the Block and the mask image corresponding to each pixel point in the target image can be input into the global integration information generation module. Among them, feature 1 can be used as the input R0 of the global integration information generation module, and the mask image corresponding to each pixel point in the target image can be used as the input H0 of the global integration information generation module.
[0343] Next, the global integration information generation module can process the feature 1 and the mask image corresponding to each pixel point in the target image, and output the global integration information.
[0344] It is a schematic diagram of the global integration information generation process shown exemplarily. It shows a possible processing process for the global integration information generation module to generate global integration information. Among them, the global integration information generation module can include a feature transformation module and a self-attention module.
[0345] Referring to , Exemplarily, feature 1 can be first input into the feature transformation module, and the feature transformation module performs three types of feature transformations (such as linear transformation) on feature 1 to obtain feature map Q, feature map K, and feature map V. Among them, the height of feature map Q, feature map K, and feature map V is H, and the width is W. Then, feature map Q and feature map K are input into the self-attention module, and the self-attention module processes feature map Q and feature map K to generate H*W attention maps; among them, the height of each attention map is H, and the width is W. After that, H*W mask maps are multiplied pointwise with H*W attention maps to obtain H*W feature maps R (the height of feature map R is H, and the width is W); then, H*W feature maps R are multiplied (i.e., matrix multiplication) with feature map V to obtain the global integration information (shading template); among them, the global integration information is a feature map with a height of H and a width of W.
[0346] S8035, Input the position information of multiple instances, the image features of multiple instances, the global integration information, and the global image features into the global integration module in the diffusion network to obtain the first intermediate feature.
[0347] Exemplarily, compared with S6033, the global integration module in S8035 has one more input, namely the global integration information; the global integration information can be used as the input U(n+2) of the global integration module. After that, the global integration module can integrate the n inputs from input U1 to input Un according to the position information of multiple instances, the global image features, and the global integration information to obtain the first intermediate feature. For details, please refer to the description of S6033 and will not be elaborated here.
[0348] S8036, Input the first intermediate feature into other modules in the diffusion network to obtain the third intermediate feature.
[0349] Exemplarily, S8036 can refer to the description of S5023 above and will not be elaborated here.
[0350] S804, Input the first intermediate image into the diffusion network, and the diffusion network performs T2 generation processes to obtain the target image.
[0351] Exemplarily, S804 can refer to the description of S503 above and will not be elaborated here.
[0352] To exemplarily show the schematic diagram of the image generation process.
[0353] Refer to , Exemplarily, the user inputs the global description information of the target image "A blue cat is on the left of a green dog". Among them, according to the global description information, the position layout and attribute information of Instance 1 (the cat) and Instance 2 (the dog) can be determined. Then, instance partitioning can be performed to obtain the attribute information of the cat "A blue cat" and the attribute information of the dog "A green dog".
[0354] Referring to , the instance processing module may include a cross-attention module and an enhancement attention module. The attribute information of the cat "A blue cat" can be input into the cross-attention module (CrossAttention) and the enhancement attention module (EnhancementAttention) respectively, and Feature 1 (i.e., the output of the first part of the Block) can be input into the cross-attention module and the enhancement attention module respectively. Then, the cross-attention module can process Feature 1 and the attribute information of the cat "A blue cat" to output the first feature of the cat; and the enhancement attention module can process Feature 1 and the attribute information of the cat "A blue cat" to output the second feature of the cat. Then, the first feature of the cat and the second feature of the cat are added together to obtain the image feature of the cat.
[0355] Similarly, the attribute information of the dog "A green dog" can be input into the cross-attention module (CrossAttention) and the enhancement attention module (EnhancementAttention) respectively, and Feature 1 (i.e., the output of the first part of the Block) can be input into the cross-attention module and the enhancement attention module respectively. Then, the cross-attention module can process Feature 1 and the attribute information of the dog "A green dog" to output the first feature of the dog; and the enhancement attention module can process Feature 1 and the attribute information of the dog "A green dog" to output the second feature of the dog. Then, the first feature of the dog and the second feature of the dog are added together to obtain the image feature of the dog.
[0356] Referring to , Feature 1 and the global description information can be input into the cross-attention module of the instance processing module to obtain the image feature (which can also be called the background feature) of the target image.
[0357] Referring to , Feature 1 and the mask map corresponding to each pixel point in the target image can be input into the global integrated information generation module (which may include a layout attention module) to obtain the global integrated information.
[0358] Referring to , the position information of the cat, the position information of the dog, the global integration information, the image features of the target image, the image features of the cat, and the image features of the dog can be input into the global integration module to obtain the global image features. After that, by processing the global image features, a target image as shown in Target Image 1 or Target Image 2 can be obtained.
[0359] Compared with the above Image Generation Process 200, Image Generation Process 400, Image Generation Process 500, Image Generation Process 600, and Image Generation Process 700, in the global integration process of Image Generation Process 800, global integration information is newly introduced; that is to say, the information used by Image Generation Process 800 to guide the integration of image features of multiple instances is richer. Therefore, Image Generation Process 800 can achieve more accurate position control of multiple instances in the target image.
[0360] In addition, the attention map represents the probability that each pixel point in the target image is any instance or the background; since the H*W mask maps are generated based on the position information of each instance, multiplying the H*W mask maps by the H*W attention maps can fine-tune the probability that each pixel point in the H*W attention maps is any instance or the background, that is, adjust the positions of each instance in the target image; thereby enabling the positions of each instance in the target image to match the position information of each instance more, and further improving the accuracy of the positions of multiple instances in the target image.
[0361] Schematic diagram of the target image shown for illustration.
[0362] Exemplarily, the global description information of the target image input by the user is "There is a fantasy world in the crystal ball on the table". According to this global description information, text parsing is performed to obtain the attribute information of Instance 1 "a crystal ball", the attribute information of Instance 2 "a white moon", the attribute information of Instance 3 "a Christmas tree", the attribute information of Instance 4 "a red house", the attribute information of Instance 5 "snowy ground", and the attribute information of Instance 6 "a table". And according to this global description information, instance position prediction is performed to obtain the position information of the above 6 instances; among them, according to the position information of these 6 instances, the position layout of the bounding boxes of these 6 instances (which can also be called the position layout of these 6 instances) is as (11) shown. The target image generated according to this global description information by using the image generation method of the present application is as shown in 9A(12).
[0363] Exemplarily, the global description information of the target image input by the user is "a cute little cat and a pumpkin carriage". According to this global description information, text parsing is performed to obtain the attribute information of instance 1 as "a cute little cat", the attribute information of instance 2 as "a pumpkin carriage", the attribute information of instance 3 as "a green donut wheel", the attribute information of instance 4 as "a blue donut wheel", and the attribute information of instance 5 as "hole". And according to this global description information, instance position prediction is performed to obtain the position information of the above 5 instances; among them, according to the position information of these 5 instances, the position layout of the bounding boxes of these 5 instances (which can also be called the position layout of these 5 instances) is as (21) shown. The target image generated according to this global description information by using the image generation method of the present application is as shown in 9A(22).
[0364] Exemplarily, the global description information of the target image input by the user is "a squirrel artist". According to this global description information, text parsing is performed to obtain the attribute information of instance 1 as "a beret", the attribute information of instance 2 as "a squirrel head", and the attribute information of instance 3 as "a coat". And according to this global description information, instance position prediction is performed to obtain the position information of the above 3 instances; among them, according to the position information of these 3 instances, the position layout of the bounding boxes of these 3 instances (which can also be called the position layout of these 3 instances) is as (31) shown. The target image generated according to this global description information by using the image generation method of the present application is as shown in 9A(32).
[0365] It is a schematic diagram of the target image shown exemplarily.
[0366] Exemplarily, the global description information of the target image input by the user is "a corgi wearing a Christmas hat sitting on a cake, with a teddy bear standing beside". According to this global description information, text parsing is performed to obtain the attribute information of instance 1 as "a Christmas hat", the attribute information of instance 2 as "a blurred Christmas tree", the attribute information of instance 3 as "a cute corgi", the attribute information of instance 4 as "a cake", and the attribute information of instance 5 as "a blurred blue teddy bear". And according to this global description information, instance position prediction is performed to obtain the position information of the above 5 instances; among them, according to the position information of these 5 instances, the position layout of the bounding boxes of these 4 instances (which can also be called the position layout of these 5 instances) is as (11) shown. The target image generated according to this global description information by using the image generation method of the present application is as shown in 9B(12).
[0367] Exemplarily, the global description information of the target image input by the user is "on a green table, a blue plant grows in a yellow flowerpot, next to a red landline and a white cup". According to this global description information, text parsing is performed to obtain the attribute information of instance 1 as "a blue plant", the attribute information of instance 2 as "a yellow flowerpot", the attribute information of instance 3 as "a red landline", the attribute information of instance 4 as "a white cup", and the attribute information of instance 5 as "a green table". And according to this global description information, instance position prediction is performed to obtain the position information of the above 5 instances; among them, according to the position information of these 5 instances, the position layout of the bounding boxes of these 5 instances (which can also be called the position layout of these 5 instances) is as shown in (21). The target image generated according to this global description information by the image generation method of the present application is shown in 9B(22).
[0368] Exemplarily, the global description information of the target image input by the user is "four boats". According to this global description information, text parsing is performed to obtain the attribute information of instance 1 as "a blue boat", the attribute information of instance 2 as "a red boat", the attribute information of instance 3 as "a green boat", the attribute information of instance 4 as "a black boat", the attribute information of instance 5 as "a yellow bench", and the attribute information of instance 6 as "a beach". And according to this global description information, instance position prediction is performed to obtain the position information of the above 6 instances; among them, according to the position information of these 6 instances, the position layout of the bounding boxes of these 6 instances (which can also be called the position layout of these 6 instances) is as shown in (31). The target image generated according to this global description information by the image generation method of the present application is shown in 9B(32).
[0369] It is a schematic diagram of the target image shown exemplarily.
[0370] Exemplarily, the global description information of the target image input by the user is "two teddy bears holding hands". According to this global description information, text parsing is performed to obtain the attribute information of instance 1 as "a yellow teddy bear", the attribute information of instance 2 as "a pink dress", the attribute information of instance 3 as "a red teddy bear", the attribute information of instance 4 as "a blue tie", and the attribute information of instance 5 as "a black piece of clothing". And according to this global description information, instance position prediction is performed to obtain the position information of the above 5 instances; among them, according to the position information of these 5 instances, the position layout of the bounding boxes of these 5 instances (which can also be called the position layout of these 5 instances) is as (11) as shown. The target image generated according to the global description information by the image generation method of the present application is as shown in 9C(12).
[0371] Exemplarily, the global description information of the target image input by the user is "A blue rabbit wearing a yellow wig is typing on a green keyboard". Performing text parsing according to the global description information, the attribute information of instance 1 is "a yellow wig", the attribute information of instance 2 is "a blue rabbit", the attribute information of instance 3 is "a green keyboard", and the attribute information of instance 4 is "a table". And performing instance position prediction according to the global description information, the position information of the above 4 instances is obtained; among them, according to the position information of these 4 instances, the position layout of the bounding boxes of these 4 instances (which can also be called the position layout of these 4 instances) is as (21) as shown. The target image generated according to the global description information by the image generation method of the present application is as shown in 9C(22).
[0372] Exemplarily, the global description information of the target image input by the user is "A squirrel wearing a spacesuit is holding an apple". Performing text parsing according to the global description information, the attribute information of instance 1 is "a squirrel", the attribute information of instance 2 is "a spacesuit", the attribute information of instance 3 is "a glass helmet", the attribute information of instance 4 is "an apple", and the attribute information of instance 5 is "a pair of gloves". And performing instance position prediction according to the global description information, the position information of the above 5 instances is obtained; among them, according to the position information of these 5 instances, the position layout of the bounding boxes of these 5 instances (which can also be called the position layout of these 5 instances) is as (31) as shown. The target image generated according to the global description information by the image generation method of the present application is as shown in 9C(32).
[0373] It is a schematic diagram of the target image shown exemplarily.
[0374] Exemplarily, the global description information of the target image input by the user is "A corgi is looking out of a sunny window". Performing text parsing according to the global description information, the attribute information of instance 1 is "a window with a sunny view", the attribute information of instance 2 is "a corgi", and the attribute information of instance 3 is "a dog leash"; in addition, the determined background description information is "a dark background". And performing instance position prediction according to the global description information, the position information of the above 3 instances is obtained; among them, according to the position information of these 3 instances, the position layout of the bounding boxes of these 3 instances (which can also be called the position layout of these 3 instances) is as (11) as shown. In addition, (11) It also includes the position layout of the background. The target image generated according to this global description information by the image generation method of the present application is as shown in 9D(12).
[0375] Exemplarily, the global description information of the target image input by the user is "A blue lion is reading a red book on the beach". According to this global description information, text parsing is performed to obtain the attribute information of instance 1 as "a coconut tree", the attribute information of instance 2 as "a blue lion", the attribute information of instance 3 as "a sandy beach", the attribute information of instance 4 as "a clear sky", and the attribute information of instance 5 as "a red book". And according to this global description information, instance position prediction is performed to obtain the position information of the above 5 instances; among them, according to the position information of these 5 instances, the position layout of the bounding boxes of these 5 instances (which can also be called the position layout of these 5 instances) is as (21) shown. The target image generated according to this global description information by the image generation method of the present application is as shown in 9D(22).
[0376] Exemplarily, the global description information of the target image input by the user is "A corgi is sitting on the beach looking at the sunset sky". According to this global description information, text parsing is performed to obtain the attribute information of instance 1 as "a lake", the attribute information of instance 2 as "a sunset sky", the attribute information of instance 3 as "a corgi", and the attribute information of instance 4 as "a sandy beach". And according to this global description information, instance position prediction is performed to obtain the position information of the above 4 instances; among them, according to the position information of these 4 instances, the position layout of the bounding boxes of these 4 instances (which can also be called the position layout of these 4 instances) is as (31) shown. The target image generated according to this global description information by the image generation method of the present application is as shown in 9D(32).
[0377] It is a schematic diagram of the target image shown for illustration. The target image of (1) includes an orange cat, a blue cat, a green cat, a yellow sofa, and a red cake. The target image of (2) includes a green cake, a yellow cake, a blue cake, a red cake, and an orange cake. It can be seen that the image generation method of the present application can generate instances of different colors.
[0378] It is a schematic diagram of the target image shown for illustration. The target image of (1) includes 3 dogs and 2 cats. The target image of (2) includes 5 birds. It can be seen that the image generation method of the present application can generate instances of different quantities.
[0379] is a schematic diagram of an exemplary target image. The target images of (1) include triangular pyramids, cuboids, spheres, and cylinders. The target image in (2) includes a triangular cake, a round cake, and a square cake. It can be seen that the image generation method of the present application can generate instances of different shapes.
[0380] is a schematic diagram of an exemplary target image. The target image of (1) includes vases made of five materials: wooden vase, glass vase, woven vase, stone vase and metal; it can be seen that the image generation method of the present application can generate instances of different materials. The target image of (2) includes five towels: a pure blue towel, a striped towel, a spotted towel, a checkered towel, and a printed towel; it can be seen that the image generation method of the present application can generate instances of different textures.
[0381] is a schematic diagram of an exemplary target image. The target images in (1) include cyberpunk-style cities, realistic corgis, and cartoon cars. The target images in (2) include a watercolor elephant, a Van Gogh starry sky, a cartoon cat, and a realistic lake. It can be seen that the image generation method of the present application can generate instances of different styles.
[0382] is a schematic diagram of an exemplary target image. The target image includes a player wearing two shoes of different colors, a football and a blue chest; it can be seen that the image generation method of the present application can generate multiple instances with action relationships.
[0383] For example, Tables 1 to 3 show the comparison results of the effects of the present application and the prior art:
[0384] Table 1
[0385]
[0386] In Table 1, when the number of instances is two, the instance generation success rate of the prior art 1 is 6.87, the instance generation success rate of the prior art 2 is 20.47, the instance generation success rate of the prior art 3 is 24.61, the instance generation success rate of the prior art 4 is 24.88, the instance generation success rate of the prior art 5 is 42.30, and the instance generation success rate of this application is 67.70. When the number of instances is 3, 4, 5, and 6 respectively, the instance generation success rates of the prior art 1 to the prior art 5 and the instance generation success rate of this application can be referred to as shown in Table 1, which will not be elaborated here. By comparison, the instance generation success rate of this application is higher than that of the prior art.
[0387] Exemplarily, one way to determine whether an instance is successfully generated can be as follows: for an instance in the target image (subsequently referred to as instance 1), detect the position of instance 1 in the target image (subsequently referred to as position 1); then, calculate the intersection over union (IoU) between position 1 and the position of instance 1 specified by the user (subsequently referred to as position 0); when the IoU between position 1 and position 0 is greater than 0.5, calculate the IoU between the region where the attributes of instance 1 in the target image are the same as the attributes of instance 1 specified by the user (subsequently referred to as region 1) and the region corresponding to instance 1 in the target image (subsequently referred to as region 2, that is, the region in the target image that only contains instance 1). When the IoU between region 1 and region 2 is greater than 0.2, it is determined that instance 1 in the target image is successfully generated; otherwise, it is determined that instance 1 in the target image is failed to be generated. And so on, determine whether other instances in the target image are successfully generated.
[0388] In Table 1, the time taken by the prior art 1 to generate the target image is 9.18 s, the time taken by the prior art 2 to generate the target image is 19.92 s, the time taken by the prior art 3 to generate the target image is 44.17 s, the time taken by the prior art 4 to generate the target image is 25.15 s, the time taken by the prior art 5 to generate the target image is 22.00 s, and the time taken by the prior art 4 to generate the target image is 15.16 s. For the prior art 2 to the prior art 5, this application takes less time to generate the target image and has a higher instance generation success rate.
[0389] Table 2
[0390]
[0391] Among them, the intersection over union ratio in Table 2 refers to the intersection over union ratio of region 1 and region 2 mentioned above.
[0392] In Table 2, when the number of instances is two, the intersection over union (IoU) of the prior art 1 is 18.92, the IoU of the prior art 2 is 29.34, the IoU of the prior art 3 is 32.64, the IoU of the prior art 4 is 29.41, the IoU of the prior art 5 is 37.58, and the IoU of this application is 59.39. When the number of instances is 3, 4, 5, and 6 respectively, the IoUs of the prior art 1 to the prior art 5 and the IoU of this application can be referred to as shown in Table 2, which will not be elaborated here. It can be seen that the IoU of this application is higher than that of the prior art; that is to say, the positions of the instances in the target image generated by this application are more matched with the positions of the instances specified by the user.
[0393] Table 3
[0394]
[0395] Exemplarily, the success rate in Table 3 is the success rate of generating the target image; among them, for each target image, if all the instances in the target image are successfully generated, it can be determined that the target image is successfully generated.
[0396] Exemplarily, the intersection over union (IoU) in Table 3 refers to the IoU of the positions of all the instances in the target image and the positions of all the instances specified by the user.
[0397] Exemplarily, AP (average precision) 50 refers to the proportion of target images in which the IoU of the positions of all the instances and the positions of all the instances specified by the user is greater than 0.5 among multiple target images.
[0398] Exemplarily, AP75 refers to the proportion of target images in which the IoU of the positions of all the instances and the positions of all the instances specified by the user is greater than 0.75 among multiple target images.
[0399] Exemplarily, AP refers to the average value of AP50, AP55, AP60, AP65, AP70, AP75, AP80, AP85, AP90, and AP95.
[0400] Exemplarily, CLIP (Contrastive Language-Image Pretraining) is to determine the similarity between the target image and the global description information by using CLIP.
[0401] Exemplarily, assume that the target image includes 2 instances (instance 1 and instance 2). The region of instance 1 can be extracted from the target image, and CLIP is used to determine the similarity between the region of instance 1 and the attributes of instance 1 specified by the user, obtaining similarity 1. And the region of instance 2 can be extracted from the target image, and CLIP is used to determine the similarity between the region of instance 2 and the attributes of instance 2 specified by the user, obtaining similarity 2. Among them, the mean of similarity 1 and similarity 2 is Local CLIP.
[0402] Exemplarily, the full name of the FID metric is Fréchet Inception Distance, and the Chinese full name is Fréchet initial distance. This is a metric for evaluating the performance of image generation models (such as GANs, generative adversarial networks), which measures the quality and diversity of generated images by comparing the distribution differences between generated images and real images in a specific space. 6K means that the generated target images are 6000 in number.
[0403] Exemplarily, the real images in Table 3 can refer to the images in the dataset. The position (i.e., the position of the instance specified by the user) and attributes (i.e., the attributes of the instance specified by the user) of each instance can be obtained by manually annotating the images in the dataset. Then, a detection tool can be used to detect the position and attributes of each instance in the real image; according to the position and attributes of each instance obtained by the detection tool's detection of the real image and the position and attributes of the instance specified by the user, it can be determined whether each instance is successfully generated (reference can be made to the above description and will not be elaborated here); subsequently, according to whether the instance is successfully generated, it can be determined whether the real image is successfully generated. After that, the intersection over union, AP, AP50, and AP75 can also be calculated based on the position of the instance specified by the user and the position of each instance obtained by the detection tool's detection of the real image.
[0404] It should be understood that the success rate, intersection over union, AP, AP50, and AP75 of the target images generated by the prior art 1 to the prior art 6, as well as the target images generated by the present application, can be determined in the same way as determining the success rate, intersection over union, AP, AP50, and AP75 of the real image; this will not be elaborated here.
[0405] Referring to Table 3, the quantitative metrics of the target images generated by the present application: success rate, intersection over union, AP, AP50, and AP75; as well as CLIP, Local CLIP, and FID-6K, are all better than the prior art.
[0406] Schematic diagram of the exemplary image generation device 1100. The image generation device 1100 can be used to execute the methods of the foregoing embodiments. Therefore, the beneficial effects it can achieve can be referred to the beneficial effects in the corresponding methods provided above, and will not be elaborated here.
[0407] An acquisition module 1101, configured to acquire a plurality of information groups; wherein each information group includes attribute information and location information of an instance.
[0408] A first generation module 1102, configured to generate image features of an instance according to the attribute information of an instance included in each of the plurality of information groups.
[0409] An integration module 1103, configured to integrate the image features of a plurality of instances according to the location information of the plurality of instances included in the plurality of information groups to generate a target image.
[0410] Exemplarily, the image generation device 1100 further includes:
[0411] A second generation module, configured to generate a plurality of mask maps according to the location information of the plurality of instances; wherein the plurality of mask maps correspond one-to-one to a plurality of pixel points in the target image; the size of each mask map is the same as the size of the target image, the mask map corresponding to the first pixel point is the mask map corresponding to the instance to which the first pixel point belongs, and the first pixel point is any pixel point in the target image; and generate global integration information according to the plurality of mask maps.
[0412] The integration module 1103 is specifically configured to integrate the image features of the plurality of instances according to the location information of the plurality of instances and the global integration information to generate a target image.
[0413] Exemplarily, the acquisition module 1101 is further configured to acquire global description information of the target image.
[0414] The image generation device 1100 further includes:
[0415] A third generation module, configured to generate global image features according to the global description information.
[0416] The integration module 1103 is specifically configured to integrate the image features of the plurality of instances according to the location information of the plurality of instances and the global image features to generate a target image.
[0417] Exemplarily, the image generation device 1100 further includes:
[0418] A second generation module for generating multiple mask images based on the position information of multiple instances; wherein, the multiple mask images correspond one-to-one with multiple pixel points in the target image; the size of each mask image is the same as the size of the target image, the mask image corresponding to the first pixel point is the mask image corresponding to the instance to which the first pixel point belongs, and the first pixel point is any pixel point in the target image; and generating global integration information based on the multiple mask images.
[0419] An integration module 1103, specifically for integrating the image features of multiple instances based on the position information of multiple instances, the global integration information, and the global image features to generate a target image.
[0420] Exemplarily, an acquisition module 1101 is specifically configured to receive multiple information groups input by a user.
[0421] Exemplarily, an acquisition module 1101 is specifically configured to perform text parsing on the global description information of the target image to obtain the attribute information of multiple instances; and perform instance position prediction on the global description information of the target image to obtain the position information of multiple instances.
[0422] Exemplarily, a first generation module 1102 is specifically configured to divide the attribute information of multiple instances into multiple paths and input them into an instance processing network, and the instance processing network generates features based on the attribute information of one instance to obtain the image features of one instance.
[0423] Wherein, the attribute information of one instance is an input path of the instance processing network.
[0424] Exemplarily, an integration module 1103 is specifically configured to input the position information of multiple instances and the image features of multiple instances into a global integration network to obtain a target image.
[0425] Exemplarily, a first generation module 1102 is specifically configured to, in each generation process of the diffusion network performing T1 generation processes, use the first intermediate image generated in the previous generation as one path and the attribute information of multiple instances as multiple paths and input them into the diffusion network, and the instance processing network in the diffusion network performs multiple feature generations to obtain the image features of multiple instances; wherein, the first intermediate image generated in the previous generation and the attribute information of one instance are used for one feature generation, and the attribute information of one instance is an input path.
[0426] The integration module 1103 is specifically configured to, during each generation process of the T1-th generation process in the diffusion network, input the position information of multiple instances and the image features of multiple instances into the global integration network in the diffusion network to obtain a first intermediate feature; input the first intermediate feature into other networks in the diffusion network to obtain a first intermediate image; wherein the global integration network is after the instance processing network, and the other networks are after the global integration network; determine the target image according to the first intermediate image obtained by the diffusion network in the T1-th generation process.
[0427] Exemplarily, the integration module 1103 is specifically configured to, during each generation process of the subsequent T2-th generation process in the diffusion network, input the second intermediate image generated in the previous time as one path and the attribute information of multiple instances as one path into the diffusion network, and the instance processing network generates features according to the second intermediate image generated in the previous time and the attribute information of multiple instances to obtain a second intermediate feature; input the second intermediate feature into other networks to obtain a second intermediate image; use the second intermediate image obtained by the diffusion network in the T-th generation process as the target image; wherein T is the sum of T1 and T2.
[0428] Exemplarily, the integration module 1103 is specifically configured to use the first intermediate image obtained by the diffusion network in the T1-th generation process as the target image.
[0429] Exemplarily, the attribute information includes at least one of the following: image or text;
[0430] The position information includes at least one of the following: image or text.
[0431] Exemplarily, the position information includes at least one of the following: the contour information of the instance and the information of the bounding box of the instance.
[0432] In one example FIG. shows a schematic block diagram of a device 1200 according to an embodiment of the present application. The device 1200 may include: a processor 1201 and a transceiver / transceiver pin 1202. Optionally, a memory 1203 is further included.
[0433] Each component of the device 1200 is coupled together through a bus 1204. The bus 1204 includes, in addition to a data bus, a power bus, a control bus, and a status signal bus. However, for the sake of clarity, all various buses are referred to as the bus 1204 in the figure.
[0434] Optionally, the memory 1203 may be used to store the instructions in the foregoing method embodiments. The processor 1201 may be configured to execute the instructions in the memory 1203, control the receiving pin to receive signals, and control the sending pin to send signals.
[0435] The device 1200 may be an electronic device or a chip of an electronic device in the above method embodiments.
[0436] Among them, all relevant contents of each step involved in the above method embodiments can be cited in the function descriptions of the corresponding functional modules, and will not be elaborated here.
[0437] The embodiment of the present application further provides a chip, including one or more interface circuits and one or more processors; the one or more processors receive or send data through the one or more interface circuits, and when the one or more processors execute computer instructions, the relevant method steps described above are executed to implement the steps of the method in the above embodiments. Among them, the interface circuit is a transceiver / transceiver pin 1202.
[0438] The present embodiment further provides a computer-readable storage medium, in which computer instructions are stored. When the computer instructions run on an electronic device, the electronic device is caused to execute the relevant method steps to implement the method in the above embodiments.
[0439] The present embodiment further provides a computer program product, which includes computer instructions. When the computer instructions are executed by a computer or a processor, the computer is caused to execute the above relevant steps to implement the method in the above embodiments.
[0440] In addition, the embodiment of the present application further provides a device, which may specifically be a chip, a component or a module. The device may include a processor and a memory connected to each other; among them, the memory is used to store computer execution instructions. When the device runs, the processor may execute the computer execution instructions stored in the memory, so that the chip executes the methods in the above method embodiments.
[0441] Among them, the electronic device, computer-readable storage medium, computer program product or chip provided in the present embodiment are all used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method provided above, and will not be elaborated here.
[0442] Through the description of the above embodiments, those skilled in the art can understand that for the convenience and brevity of description, only the above division of each functional module is used as an example for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.
[0443] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.
[0444] The units described as separate components may or may not be physically separated. The components displayed as units may be one physical unit or multiple physical units, that is, they can be located in one place, or they can be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0445] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0446] Any content in each embodiment of the present application, as well as any content in the same embodiment, can be freely combined. Any combination of the above content is within the scope of the present application.
[0447] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods in the various embodiments of the present application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read only memory (ROM), random access memory (RAM), magnetic disks or optical discs and other various media that can store program codes.
[0448] The steps of the method or algorithm described in connection with the disclosed content of the embodiments of the present application may be implemented in a hardware manner or by a processor executing software instructions. The software instructions may be composed of corresponding software modules, and the software modules may be stored in a random access memory (RAM), flash memory, read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disk, removable hard disk, compact disc read-only memory (CD-ROM), or any other form of storage medium well-known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium may also be a component of the processor. The processor and the storage medium may be located in an ASIC.
[0449] Those skilled in the art should be able to realize that in one or more of the above examples, the functions described in the embodiments of the present application can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium. The computer-readable medium includes a computer-readable storage medium and a communication medium, where the communication medium includes any medium that facilitates the transmission of a computer program from one place to another. The storage medium may be any available medium accessible by a general-purpose or special-purpose computer.
[0450] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.
Claims
1. An image generation method, characterized in that, The method includes: Obtaining a plurality of information groups; wherein each information group includes the attribute information and location information of an instance; Generating the image features of an instance according to the attribute information of an instance included in each of the plurality of information groups; Integrating the image features of multiple instances according to the location information of the multiple instances included in the plurality of information groups to generate a target image.
2. The method according to claim 1, wherein The method further includes: Generating multiple mask maps according to the location information of the multiple instances; wherein the multiple mask maps correspond one-to-one to multiple pixel points in the target image; the size of each mask map is the same as the size of the target image, and the mask map corresponding to the first pixel point is the mask map corresponding to the instance to which the first pixel point belongs, and the first pixel point is any pixel point in the target image; Generating global integration information according to the multiple mask maps; The integrating the image features of multiple instances according to the location information of the multiple instances included in the plurality of information groups to generate a target image includes: Integrating the image features of the multiple instances according to the location information of the multiple instances and the global integration information to generate the target image.
3. The method according to claim 1, characterized in that, The method further includes: Obtaining the global description information of the target image; Generating the global image features according to the global description information; The integrating the image features of multiple instances according to the location information of the multiple instances included in the plurality of information groups to generate a target image includes: Integrating the image features of the multiple instances according to the location information of the multiple instances and the global image features to generate the target image.
4. The method according to claim 3, wherein The method further includes: Generating multiple mask maps according to the location information of the multiple instances; wherein the multiple mask maps correspond one-to-one to multiple pixel points in the target image; the size of each mask map is the same as the size of the target image, and the mask map corresponding to the first pixel point is the mask map corresponding to the instance to which the first pixel point belongs, and the first pixel point is any pixel point in the target image; Generating global integration information according to the multiple mask maps; The integrating the image features of the multiple instances according to the location information of the multiple instances and the global image features to generate the target image includes: Integrating the image features of the multiple instances according to the location information of the multiple instances, the global integration information and the global image features to generate the target image.
5. The method according to any one of claims 1 to 4, characterized in that, The obtaining a plurality of information groups includes: Receiving a plurality of information groups input by a user.
6. The method according to any one of claims 1 to 4, characterized in that The obtaining a plurality of information groups includes: Performing text parsing according to the global description information of the target image to obtain the attribute information of the multiple instances; Performing instance location prediction according to the global description information of the target image to obtain the location information of the multiple instances.
7. The method according to any one of claims 1 to 6, characterized in that, The generating the image features of an instance according to the attribute information of an instance included in each of the plurality of information groups includes: The attribute information of the multiple instances is divided into multiple paths and input into an instance processing network. The instance processing network generates features based on the attribute information of one instance to obtain the image features of one instance. Among them, the attribute information of one instance is one path of input to the instance processing network.
8. The method according to any one of claims 1 to 7, characterized in that, The integrating the image features of multiple instances according to the position information of the multiple instances included in the multiple information groups to generate a target image includes: Inputting the position information of the multiple instances and the image features of the multiple instances into a global integration network to obtain the target image.
9. The method according to any one of claims 1 to 6, wherein The generating the image features of one instance according to the attribute information of one instance included in each of the multiple information groups includes: In each generation process of the diffusion network performing T1 generation processes, using the first intermediate image generated in the previous generation as one path and using the attribute information of the multiple instances as multiple paths and inputting them into the diffusion network. The instance processing network in the diffusion network performs multiple feature generations to obtain the image features of the multiple instances. Among them, the first intermediate image generated in the previous generation and the attribute information of one instance are used for one feature generation, and the attribute information of one instance is one path of input. The integrating the image features of multiple instances according to the position information of the multiple instances included in the multiple information groups to generate a target image includes: In each generation process of the diffusion network performing T1 generation processes, inputting the position information of the multiple instances and the image features of the multiple instances into the global integration network in the diffusion network to obtain a first intermediate feature. Inputting the first intermediate feature into other networks in the diffusion network to obtain a first intermediate image. Among them, the global integration network is after the instance processing network, and the other networks are after the global integration network. Determining the target image according to the first intermediate image obtained by the diffusion network performing the T1-th generation process.
10. The method according to claim 11, characterized in that, The determining the target image according to the first intermediate image obtained by the diffusion network performing the T1-th generation process includes: In each generation process of the diffusion network performing the subsequent T2 generation processes, using the second intermediate image generated in the previous generation as one path and using the attribute information of the multiple instances as one path and inputting them into the diffusion network. The instance processing network generates features according to the second intermediate image generated in the previous generation and the attribute information of the multiple instances to obtain a second intermediate feature. Inputting the second intermediate feature into the other networks to obtain a second intermediate image. Using the second intermediate image obtained by the diffusion network performing the T-th generation process as the target image. Wherein, T is the sum of T1 and T2.
11. The method according to claim 10, wherein The determining the target image according to the first intermediate image obtained by the diffusion network performing the T1-th generation process includes: Using the first intermediate image obtained by the diffusion network performing the T1-th generation process as the target image.
12. The method according to any one of claims 1 to 11, wherein The attribute information includes at least one of the following: an image or text; The location information includes at least one of the following: an image or text.
13. The method according to any one of claims 1 to 12, characterized in that The location information includes at least one of the following: the contour information of the instance and the information of the bounding box of the instance.
14. An image generation device, characterized in that, The apparatus includes: An acquisition module, configured to acquire a plurality of information groups; wherein each information group includes the attribute information and location information of an instance; A first generation module, configured to generate an image feature of an instance according to the attribute information of an instance included in each information group among the plurality of information groups; An integration module, configured to integrate the image features of a plurality of instances according to the location information of the plurality of instances included in the plurality of information groups to generate a target image.
15. The device according to claim 14, characterized in that, The apparatus further includes: A second generation module, configured to generate a plurality of mask maps according to the location information of the plurality of instances; wherein the plurality of mask maps correspond one-to-one to a plurality of pixel points in the target image; the size of each mask map is the same as the size of the target image, and the mask map corresponding to the first pixel point is the mask map corresponding to the instance to which the first pixel point belongs, and the first pixel point is any pixel point in the target image; and generate global integration information according to the plurality of mask maps; The integration module is specifically configured to integrate the image features of the plurality of instances according to the location information of the plurality of instances and the global integration information to generate the target image.
16. The apparatus according to claim 14, characterized in that The acquisition module is further configured to acquire the global description information of the target image; The apparatus further includes: A third generation module, configured to generate the global image feature according to the global description information; The integration module is specifically configured to integrate the image features of the plurality of instances according to the location information of the plurality of instances and the global image feature to generate the target image.
17. The apparatus according to any one of claims 14 to 16, characterized in that The acquisition module is specifically configured to receive a plurality of information groups input by a user.
18. The apparatus according to any one of claims 14 to 16, characterized in that The acquisition module is specifically configured to perform text parsing on the global description information of the target image to obtain the attribute information of the plurality of instances; perform instance location prediction on the global description information of the target image to obtain the location information of the plurality of instances.
19. An electronic device, characterized in that, Includes: A memory and a processor, the memory is coupled to the processor; The memory stores program instructions, and when the program instructions are executed by the processor, the electronic device is caused to execute the method according to any one of claims 1 to 13.
20. A chip, characterized in that, Includes one or more interface circuits and one or more processors; the one or more processors receive or send data through the one or more interface circuits, and when the one or more processors execute computer instructions, the steps of the method according to any one of claims 1 to 13 are caused to be executed.
21. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program runs on a computer or a processor, the computer or the processor is caused to execute the method according to any one of claims 1 to 13.
22. A computer program product, characterized in that, The computer program product includes computer instructions, and when the computer instructions are executed by a computer or a processor, the steps of the method according to any one of claims 1 to 13 are caused to be executed.
Citation Information
Cited By
Image processing method and apparatus, device, computer-readable storage medium, and product
US20250356555A1