Image generation method and device, equipment, storage medium and computer program product
By setting up a control module in the target control layer of the image generation model, the target control layer is determined based on the influence of each control layer, and the problem of low computing resource allocation efficiency in the prior art is solved, and image generation efficiency and effect are improved.
Patent Information
- Application Number
- CN202510192113.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-10
AI Technical Summary
The existing image generation method has low efficiency in allocation of computing resources for the image generation model, resulting in low image generation efficiency and poor image effect.
In response to the image generation task, the task is input into the preset image generation model, and the control module is used to set it in the target control layer, and the target control layer is determined based on the impact of each control layer on the model output results, thereby realizing the effective allocation of computing resources.
Reduces computational redundancy and complexity, and improves image generation efficiency and image effect.
Smart Images

Figure CN120125692A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to an image generation method, apparatus, device, storage medium, and computer program product. Background Art
[0002] Currently, image generation models have greatly promoted the development of various technologies in the generation field with their excellent scalability and multi-modal alignment capabilities. However, existing image generation methods have the defects of low image generation efficiency and poor image quality due to the low efficiency of computing resource allocation for image generation models. Summary of the Invention
[0003] The main purpose of this application is to provide an image generation method, apparatus, device, storage medium, and computer program product, aiming to solve the technical problem that existing image generation methods have the defects of low image generation efficiency and poor image quality due to the low efficiency of computing resource allocation for image generation models.
[0004] To achieve the above object, this application provides an image generation method, which includes:
[0005] In response to an image generation task, input the image generation task into a preset image generation model;
[0006] Process the image generation task through the preset image generation model to generate a target image, where the preset image generation model includes at least one control module, the control module is set in a target control layer of the preset image generation model, and the target control layer is determined based on the influence of each control layer in the preset image generation model on the model output result.
[0007] Optionally, before responding to the image generation task and inputting the image generation task into the preset image generation model, it further includes:
[0008] Calculate the correlation scores of each control layer in the initial image generation model, where the correlation scores are used to measure the influence of each control layer on the model output result;
[0009] Select a target control layer from each control layer according to the correlation scores;
[0010] Set a control module in the target control layer to obtain the preset image generation model.
[0011] Optionally, the calculating the correlation scores of each control layer in the initial image generation model includes:
[0012] Skip each control layer in the initial image generation model in turn for model inference;
[0013] Obtain the model metrics during the model inference process, and generate the correlation scores of each control layer in the model based on the model metrics.
[0014] Optionally, the obtaining the model metrics during the model inference process and generating the correlation scores of each control layer in the model based on the model metrics includes:
[0015] Obtain the image quality loss and the control precision loss during the model inference process, where the image quality loss and the control precision loss are used to measure the impact of skipping each control module for model inference on the initial image generation model;
[0016] Calculate the correlation scores of each control layer in the model based on the image quality loss and the control precision loss.
[0017] Optionally, the selecting the target control layer from each control layer according to the correlation score includes:
[0018] Sort each control layer according to the correlation score;
[0019] Select a preset number of target control layers from each control layer according to the sorting result.
[0020] Optionally, the preset image generation model is a diffusion Transformer model, and the control module is a correlation-guided control module; the processing the image generation task through the preset image generation model to generate a target image includes:
[0021] Process the image generation task through the correlation-guided control module to generate a target image, where the correlation-guided control module is used to unify the token mixing operation and the channel mixing operation of the diffusion Transformer model into a target operation.
[0022] Optionally, the correlation-guided control module includes: a two-dimensional random mixing module; the processing the image generation task through the preset image generation model based on the target operation to generate a target image includes:
[0023] Convert the image generation task into a token sequence;
[0024] Process the token sequence through the two-dimensional random mixing module to generate a target image, where the two-dimensional random mixing module is used to perform information interaction across the token dimension and the channel dimension.
[0025] Optionally, the processing the token sequence through the two-dimensional random mixing module to generate a target image includes:
[0026] Randomly divide the token sequence in the channel dimension based on the correlation scores of each target control layer to form multiple channel groups;
[0027] Randomly divide the tokens within each channel group to form multiple token groups;
[0028] Perform attention calculation within the multiple channel groups and multiple token groups after random division;
[0029] After completing the attention calculation, reverse the positions of the tokens and channels to generate the target image.
[0030] Optionally, the step of, after completing the attention calculation, reversing the positions of the tokens and channels to generate the target image includes:
[0031] After completing the attention calculation, obtain the first random factor for randomly dividing the token sequence and the second random factor for randomly dividing the tokens;
[0032] Based on the first random factor, reverse the positions of the tokens, and based on the second random factor, reverse the positions of the channels to generate the target image.
[0033] Optionally, the step of converting the image generation task into a token sequence includes:
[0034] Preprocess the image generation task to obtain a processed image generation task;
[0035] Convert the processed image generation task into an embedding vector;
[0036] Sort the embedding vector according to the image generation task to obtain a token sequence.
[0037] Optionally, after generating the target image by processing the image generation task through the preset image generation model, the method further includes:
[0038] Obtain user information, and construct a user portrait based on the user information through a large language model;
[0039] Perform personalized adjustment on the target image based on the user portrait to obtain a personalized image.
[0040] In addition, to achieve the above object, the present application also proposes an image generation device, where the image generation device includes:
[0041] A task input module, configured to input the image generation task into a preset image generation model in response to the image generation task;
[0042] An image generation module, configured to process the image generation task through the preset image generation model to generate a target image, where the preset image generation model includes at least one control module, the control module is set in the target control layer of the preset image generation model, and the target control layer is determined based on the influence of each control layer in the preset image generation model on the model output result.
[0043] Optionally, the image generation device further includes:
[0044] A model construction module, configured to calculate the correlation scores of each control layer in the initial image generation model, where the correlation scores are used to measure the influence of each control layer on the model output result; select a target control layer from each control layer according to the correlation scores; set a control module in the target control layer to obtain the preset image generation model.
[0045] Optionally, the model construction module is further configured to sequentially skip each control layer in the initial image generation model for model inference; obtain the model metrics during the model inference process, and generate the correlation scores of each control layer in the model based on the model metrics.
[0046] Optionally, the model construction module is further configured to obtain the image quality loss and the control precision loss during the model inference process, where the image quality loss and the control precision loss are used to measure the influence of skipping each control module for model inference on the initial image generation model; calculate the correlation scores of each control layer in the model based on the image quality loss and the control precision loss.
[0047] Optionally, the model construction module is further configured to sort each control layer according to the correlation scores; select a preset number of target control layers from each control layer according to the sorting result.
[0048] Optionally, the preset image generation model is a diffusion Transformer model, and the control module is a correlation-guided control module; the image generation module is further configured to process the image generation task through the correlation-guided control module to generate a target image, where the correlation-guided control module is configured to unify the token mixing operation and the channel mixing operation of the diffusion Transformer model into a target operation.
[0049] In addition, to achieve the above object, the present application further provides an image generation device, the image generation device includes a memory, a processor, and an image generation program stored on the memory and executable on the processor, and the image generation program is configured to implement the image generation method as described above.
[0050] In addition, to achieve the above object, the present application also provides a storage medium, on which an image generation program is stored. When the image generation program is executed by a processor, the image generation method described above is implemented.
[0051] In addition, to achieve the above object, the present application also provides a computer program product, which includes an image generation program. When the image generation program is executed by a processor, the image generation method described above is implemented.
[0052] One or more technical solutions proposed by the present application have at least the following technical effects:
[0053] In the present application, in response to an image generation task, the image generation task is input into a preset image generation model, and the preset image generation model processes the image generation task to generate a target image. Among them, a control module is set in a target control layer of the preset image generation model, and the target control layer is determined based on the influence of each control layer in the preset image generation model on the model output result. Since the present application determines the target control layer based on the influence of each control layer in the image generation model on the model output result, and sets a control module in the target control layer for image generation, it is possible to effectively allocate computing resources, thereby reducing computational redundancy and complexity, and improving image generation efficiency and image effect. Description of the Drawings
[0054] The drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0055] To more clearly illustrate the technical solutions in the embodiments of the present application or in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0056] Figure 1 It is a schematic flowchart of the first embodiment of the image generation method of the present application;
[0057] Figure 2 It is a schematic flowchart of the second embodiment of the image generation method of the present application;
[0058] Figure 3 It is a schematic flowchart of the third embodiment of the image generation method of the present application;
[0059] Figure 4 It is a schematic diagram of the module structure of the image generation device in the embodiment of the present application;
[0060] Figure 5Schematic diagram of the device structure of the hardware operating environment involved in the image generation method in the embodiment of the present application.
[0061] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0062] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.
[0063] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.
[0064] At present, the diffusion transformer model has greatly promoted the development of various technologies in the field of generation with its excellent scalability and multimodal alignment capabilities. On this basis, methods such as PixArt-δ and OminiControl further explored the controllable text-to-image generation method based on the diffusion transformer, effectively promoting its implementation in practical application scenarios such as artificial intelligence (AI)-driven content creation and e-commerce shopping, and provided a new direction for the development of intelligent generation technology. However, these control generation methods for the diffusion transformer still face two major problems. First, these methods usually introduce a large number of additional parameters and computational overhead, which increases the burden of training and reasoning. PixArt-δ directly copies the first half of the network, which increases the number of parameters and computational complexity by 50%. Secondly, existing methods often ignore the changes in the correlation of control information in different layers of the network, resulting in inefficient allocation of computing resources. These two problems seriously hinder the deployment and application of controllable generation of diffusion transformers in practical scenarios.
[0065] Therefore, in order to overcome the above-mentioned defects, the present application provides a solution, which includes: in response to an image generation task, inputting the image generation task into a preset image generation model, processing the image generation task through the preset image generation model, and generating a target image, wherein the control module is set in a target control layer of the preset image generation model, and the target control layer is determined based on the influence of each control layer in the preset image generation model on the model output result; since the present application determines the target control layer based on the influence of each control layer in the image generation model on the model output result, and sets the control module in the target control layer for image generation, it is possible to achieve effective allocation of computing resources, thereby reducing computing redundancy and complexity, and improving image generation efficiency and image effect.
[0066] It should be noted that the execution subject of this embodiment can be an image generation device with data processing, network communication, and program running functions, such as a server, or other electronic devices that can achieve the same or similar functions. This embodiment does not limit this.
[0067] Based on this, an image generation method is provided in an embodiment of this application. Refer to Figure 1 , Figure 1 which is a schematic flowchart of the first embodiment of the image generation method of this application.
[0068] In the first embodiment, the image generation method includes:
[0069] Step S10: In response to an image generation task, input the image generation task into a preset image generation model.
[0070] It should be understood that an image generation task may refer to a task that requires a computer program to generate a corresponding image according to input information (such as text description, style requirements, etc.). For example, if a user hopes to generate a landscape painting based on a text description, then this text description will be input into a preset image generation model as an image generation task (such as "A desolate land covered with dirt and gravel, with a dead tree in the foreground, and warm sunlight shining on this desolate scene"). A preset image generation model may refer to a model that has been trained and set up for performing image generation tasks, and it can generate a target image according to the input task information. In a specific implementation, the preset image generation model may be a preset diffusion Transformer model, where the diffusion Transformer model may refer to a generative model composed of multiple stacked Transformer modules, which converts random noise into a target image through an iterative denoising process (such as PixArt-α, Flux, etc.).
[0071] In a specific implementation, when an image generation task is received, the image generation task is input into a preset image generation model that has been prepared.
[0072] Step S20: Process the image generation task through the preset image generation model to generate a target image, where the preset image generation model includes at least one control module, the control module is set in the target control layer of the preset image generation model, and the target control layer is determined based on the influence of each control layer in the preset image generation model on the model output result.
[0073] It can be understood that the control module can refer to a module in a preset image generation model that is used to introduce or adjust control information and can affect the output results of the model, such as the style, color, etc. of the image. The control layer can refer to all potential positions in the model where control modules can be inserted (such as the 27 Transformer layers of PixArt-α). The target control layer can refer to the position where the control module is actually deployed in the preset image generation model, and the selection of the target control layer is based on the degree of influence of the control layer on the output results of the model. For example, the 11 control layers with the greatest influence are selected as the target control layers to set the control module.
[0074] In a specific implementation, the preset image generation model processes and generates a target image step by step according to the input task information through its internal structure and algorithm. The preset image generation model contains at least one control module for adjusting the output results, and the control module is placed on the target control layer in the model. When setting the position of the control module, it is necessary to evaluate the degree of influence of each control layer in the model on the output results, and then select the position with the greatest influence as the target control layer.
[0075] Furthermore, in order to make the generated image more in line with the user's preferences and needs, after the step S20, it further includes: obtaining user information and constructing a user portrait based on the user information through a large language model; performing personalized adjustment on the target image based on the user portrait to obtain a personalized image. Among them, user information can refer to the data provided by the user for constructing the user portrait, including but not limited to the user's preferences, historical behaviors, age, gender, etc. The large language model can be an artificial intelligence model that can understand and generate natural language text and is used for natural language processing tasks. The user portrait can refer to a user feature description constructed based on user information and is used to understand the user's preferences and behavior patterns. The personalized image can refer to an image that meets the user's personalized needs after adjusting the target image according to the user portrait.
[0076] In a specific implementation, a user portrait is constructed using the information provided by the user through a large language model. According to the constructed user portrait, the generated target image is personalized adjusted to obtain a personalized image.
[0077] This embodiment determines the target control layer based on the influence of each control layer in the image generation model on the output results of the model, and sets a control module on the target control layer for image generation, thereby enabling effective allocation of computing resources, further reducing computational redundancy and complexity, and improving image generation efficiency and image quality.
[0078] Refer to Figure 2 , Figure 2 which is a schematic flowchart of the second embodiment of the image generation method of this application. Based on the above Figure 1Based on the first embodiment shown, a second embodiment of the image generation method of the present application is proposed.
[0079] In the second embodiment, before the step S10, the following steps are further included:
[0080] Step S01: Calculate the correlation scores of each control layer in the initial image generation model, where the correlation scores are used to measure the influence of each control layer on the model output result.
[0081] It should be understood that, in order to better measure the influence of each control layer on the model output result and achieve effective allocation of computing resources, in this embodiment, the target control layer is selected from each control layer according to the correlation scores of each control layer in the initial image generation model, and a control module is set in the target control layer to obtain a preset image generation model. Among them, the initial image generation model may refer to an image generation model that has not been optimized or does not have a specific control module added, and the initial image generation model is the basis for subsequent addition of control modules and correlation analysis. The control layer may refer to all potential positions in the model where a control module can be inserted (such as the 27 Transformer layers of PixArt-α). The correlation score may refer to a quantitative index that measures the influence degree of each control layer on the model output result and is used to determine which control layers are more important for the model output. The higher the correlation score, the more important the control layer is for the model output.
[0082] Furthermore, in order to accurately evaluate the influence of each control layer on the model output result, in this embodiment, the correlation scores of each control layer in the model are generated by sequentially skipping each control layer in the initial image generation model for model inference. The step S01 includes: sequentially skipping each control layer in the initial image generation model for model inference; obtaining the model metrics during the model inference process, and generating the correlation scores of each control layer in the model based on the model metrics. Among them, model inference may refer to the process of using the model to process input data to generate a prediction result or an image. The model metric may refer to a metric used to measure the performance of the model during inference, such as image quality, control accuracy, etc.
[0083] In a specific implementation, during the model inference process, each control layer in the initial model is removed (or "skipped") one by one, and the change of the model output is observed. After sequentially skipping the control layers for inference, the model metrics of the model output are collected, and then the model metrics are used to calculate the correlation score of each control layer.
[0084] Furthermore, to improve the accuracy of the relevance scores, in this embodiment, the relevance scores of each control layer in the model are calculated based on the image quality loss and the control precision loss. Obtaining the model metrics during the model inference process and generating the relevance scores of each control layer in the model based on the model metrics includes: obtaining the image quality loss and the control precision loss during the model inference process, where the image quality loss and the control precision loss are used to measure the impact of skipping each control module for model inference on the initial image generation model; calculating the relevance scores of each control layer in the model based on the image quality loss and the control precision loss. Among them, the image quality loss can be a quantitative indicator for measuring the quality of the generated image, and the image quality loss can be generated based on factors such as the clarity, color accuracy, and detail richness of the image. The control precision loss can be a quantitative indicator for measuring the accuracy of the model's response to external control instructions, and the control precision loss reflects the ability of the model to generate according to the expected control conditions during the generation process.
[0085] In a specific implementation, after skipping each control module for inference one by one, a preset image quality evaluation method and a control precision evaluation method are used to quantify the changes in the model output, and the image quality loss and the control precision loss during the model inference process are obtained. Combining the image quality loss and the control precision loss, a weighted sum or other preset statistical methods are used to calculate the relevance scores of each control layer.
[0086] Step S02: Select a target control layer from each of the control layers according to the relevance scores.
[0087] It can be understood that the target control layer can be a control layer selected based on the relevance scores and used to set control modules to maximize the model performance.
[0088] In a specific implementation, according to the calculated relevance scores, the control layers with relevance scores greater than a preset score are selected as the target control layers, where the preset score can be set in advance to ensure the effective allocation of computing resources under limited computing resources.
[0089] Furthermore, to select a target control layer that has a greater impact on the model output result, the step S02 includes: sorting each of the control layers according to the relevance scores; selecting a preset number of target control layers from each of the control layers according to the sorting result.
[0090] In a specific implementation, the calculated relevance scores are sorted from high to low to determine the importance order of the control layers. According to the sorting result and a preset quantity limit (such as "the first N positions"), the target control layers are selected from the control layers.
[0091] Step S03: Set a control module in the target control layer to obtain the preset image generation model.
[0092] It should be understood that the preset image generation model may refer to an image generation model that has been optimized and includes at least one control module set in the target control layer. After setting the control module in the target control layer, the image generation model needs to be trained to obtain the preset image generation model.
[0093] In a specific implementation, set a control module in the target control layer to obtain an image generation model to be trained; train the image generation model to be trained based on image generation task samples to obtain the preset image generation model.
[0094] The training steps may specifically be inputting the image generation task samples into the image generation model to be trained. The image generation model to be trained is a diffusion Transformer model, and the control module is a correlation-guided control module. The correlation-guided control module includes: a two-dimensional random mixing module; converting the image generation task samples into token sequence samples, and processing the token sequence samples through the two-dimensional random mixing module to generate sample images, where the two-dimensional random mixing module is used for information interaction across the token dimension and the channel dimension. Calculate the loss function based on the sample images, and train the image generation model to be trained based on the loss function to obtain the preset image generation model.
[0095] For ease of understanding, the following is an example, but it does not limit the present application. As an example, the method proposed in this embodiment first calculates the control relevance scores for the diffusion transformer. The calculation method is as follows. For a diffusion transformer with complete replication, during inference, different modules of the control branch are skipped in sequence, and the impact of removing this module on the final result is measured using image quality metrics and control metrics. According to the comprehensive ranking of the two metrics, the relevance scores from high to low can be obtained, and the top 11 control layers with higher relevance are selected to integrate the control module (PixArt-δ has 13 control modules).
[0096] Different from the direct replication of the main branch of the diffusion Transformer in related methods, in this embodiment, a Relevance-Guided Lightweight Control Block (RGLC) is introduced to further eliminate the internal redundancy of the control module. The core idea of the Relevance-Guided Lightweight Control Block is to unify the token mixing and channel mixing in the Transformer into a single operation to improve the computational efficiency. Specifically, in this embodiment, a Two Dimensions Shuffle Mixer (TDSM) is designed to perform information interaction and modeling across the token and channel dimensions, thereby significantly reducing the number of parameters and computational requirements of the replication block while maintaining the original controllable generation ability. According to the relevance score, the number of channel group partitions is assigned to the input token sequence, followed by random channel group partitioning, and then random token group partitioning and attention calculation are performed within different channel groups. After the above operations, the positions of the tokens and channels are restored in reverse according to the random factor used in the random partitioning. Relying on this design method, efficient attention calculation without information loss in the TDSM is achieved, thereby effectively reducing the computational resource consumption of the control branch.
[0097] In this embodiment, the target control layer is selected from each control layer according to the relevance scores of each control layer in the initial image generation model, and a control module is set in the target control layer to obtain a preset image generation model, so as to better measure the influence of each control layer on the model output result and achieve the effective allocation of computational resources.
[0098] Refer to Figure 3 , Figure 3 FIG.
[0099] In the third embodiment, the preset image generation model is a diffusion Transformer model, and the control module is a relevance-guided control module; the step S20 includes:
[0100] Step S201: Process the image generation task through the relevance-guided control module to generate a target image, where the relevance-guided control module is used to unify the token mixing operation and the channel mixing operation of the diffusion Transformer model into a target operation.
[0101] It should be understood that, in order to eliminate the internal redundancy of the control module, in this embodiment, the control module is set as a Relevance-Guided Lightweight Control Block (RGLC). The Relevance-Guided Lightweight Control Block is used to unify the token mixing operation and the channel mixing operation of the diffusion Transformer model into a target operation to improve the computational efficiency. Among them, the token mixing operation can refer to the operation of mixing or transforming the input tokens (or called tokens) in the diffusion Transformer model to enhance the model's ability to understand the input data. The channel mixing operation can refer to the operation of mixing or transforming the channels of the feature map in the neural network to extract more useful feature information. The target operation can refer to a single operation that unifies the token mixing operation and the channel mixing operation, aiming to improve the computational efficiency and performance of the model.
[0102] In a specific implementation, during the process of the model processing the image generation task, the Relevance-Guided Lightweight Control Block, through a special design, unifies the token mixing operation and the channel mixing operation in the diffusion Transformer model into a new target operation to improve the computational efficiency and performance of the model.
[0103] Furthermore, in order to improve the computational efficiency and maintain the controllable generation ability, the Relevance-Guided Lightweight Control Block includes: a two-dimensional random mixing module. The step S201 includes: converting the image generation task into a token sequence; processing the token sequence through the two-dimensional random mixing module to generate a target image, where the two-dimensional random mixing module is used to perform information interaction across the token dimension and the channel dimension. In a specific implementation, first, the image generation task (such as text description) is converted into a series of token sequences. During the process of the model processing the token sequence, the two-dimensional random mixing module performs information interaction and modeling across tokens and channels to generate a target image.
[0104] Furthermore, the processing the token sequence through the two-dimensional random mixing module to generate a target image includes: randomly dividing the token sequence in the channel dimension based on the relevance scores of each target control layer to form multiple channel groups; randomly dividing the tokens within each channel group to form multiple token groups; performing attention calculation within the randomly divided multiple channel groups and multiple token groups; after completing the attention calculation, reversely restoring the positions of the tokens and channels to generate a target image.
[0105] In a specific implementation, based on the correlation scores of each target control layer, the token sequence is randomly partitioned in the channel dimension to form multiple channel groups. The higher the correlation score, the fewer the number of channel groups assigned to the token sequence. In this way, there can be more tokens in one channel group, facilitating the analysis of the relationships between tokens. Within each channel group, the tokens are further randomly partitioned to form multiple token groups. In the randomly partitioned channel groups and token groups, attention calculation is performed to capture the dependencies between tokens. After the attention calculation is completed, according to the random factors used in the random partitioning, the positions of the tokens and channels are restored in reverse to generate the final target image.
[0106] It can be understood that after the attention calculation is completed, restoring the positions of the tokens and channels in reverse to generate the target image can be to obtain the first random factor for randomly partitioning the token sequence and the second random factor for randomly partitioning the tokens after the attention calculation is completed; based on the first random factor, the positions of the tokens are restored in reverse, and based on the second random factor, the positions of the channels are restored in reverse to generate the target image.
[0107] It can be understood that converting the image generation task into a token sequence can be to preprocess the image generation task to obtain the processed image generation task; convert the processed image generation task into an embedding vector; sort the embedding vector according to the image generation task to obtain the token sequence.
[0108] In this embodiment, the control module is set as a correlation-guided control module. The correlation-guided control module is used to unify the token mixing operation and the channel mixing operation of the diffusion Transformer model into a target operation to eliminate internal redundancy in the control module and improve computational efficiency.
[0109] It should be understood that through the correlation-guided position arrangement and the correlation-guided lightweight control module design in this application, the computational complexity required for the generation model is significantly reduced. Taking a representative model PixArt-α in the diffusion Transformer as an example. Compared with the traditional ControlNet implementation (PixArt-δ), this embodiment only introduces a very small amount of parameter and computational increase.
[0110] Moreover, the efficient control method of the diffusion Transformer based on the correlation score in this application can effectively complete controllable generation tasks under various conditions, while accurately ensuring the consistency of the text and control conditions. At the same time, this method can be compatible with the Lora fine-tuning of the main branch of the diffusion Transformer to achieve controllable generation in different styles.
[0111] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the image generation method of this application. Based on this technical concept, more forms of simple transformations are within the protection scope of this application.
[0112] This application also provides an image generation device. Please refer to Figure 4 , the image generation device includes:
[0113] A task input module 10, configured to input the image generation task into a preset image generation model in response to an image generation task;
[0114] An image generation module 20, configured to process the image generation task through the preset image generation model to generate a target image, where the preset image generation model includes at least one control module, the control module is disposed in a target control layer of the preset image generation model, and the target control layer is determined based on the influence of each control layer in the preset image generation model on the model output result.
[0115] The image generation device provided by this application adopts the image generation method in the above embodiment, and can solve the technical problem that the existing image generation method has low efficiency in allocating computing resources for the image generation model, resulting in low image generation efficiency and poor image effects. Compared with the prior art, the beneficial effects of the image generation device provided by this application are the same as those of the image generation method provided by the above embodiment, and other technical features in the image generation device are the same as those disclosed in the method of the above embodiment, and will not be elaborated here.
[0116] This application provides an image generation device. The image generation device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the image generation method in the first embodiment above.
[0117] Next, refer to Figure 5 , which shows a schematic structural diagram of an image generation device suitable for implementing the embodiments of this application. The image generation device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions, tablet computers), PMPs (Portable Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 5The illustrated image generation device is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the present application.
[0118] As Figure 5 shown, the image generation device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a ROM (Read Only Memory) 1002 or a program loaded from a storage device 1003 into a RAM (Random Access Memory) 1004. In the RAM 1004, various programs and data required for the operation of the image generation device are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, an LCD (Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the image generation device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows an image generation device having various systems, it should be understood that it is not required to implement or have all the shown systems. Instead, more or fewer systems may be implemented or had.
[0119] Specifically, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device, or installed from the storage device 1003, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above functions defined in the methods of the embodiments disclosed in the present application are executed.
[0120] The image generation device provided by this application adopts the image generation method in the above-mentioned embodiment, and can solve the technical problem that the existing image generation method has defects such as low image generation efficiency and poor image effect due to the low computational resource allocation efficiency of the image generation model. Compared with the prior art, the beneficial effects of the image generation device provided by this application are the same as those of the image generation method provided by the above-mentioned embodiment, and other technical features in this image generation device are the same as those disclosed in the method of the previous embodiment, which will not be elaborated here.
[0121] It should be understood that each part disclosed in this application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in a suitable manner in any one or more embodiments or examples.
[0122] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
[0123] This application provides a computer-readable storage medium with computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the image generation method in the above-mentioned embodiment.
[0124] The computer-readable storage medium provided by the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, RAM (Random Access Memory), ROM (Read Only Memory), EPROM (Erasable Programmable Read Only Memory), or flash memory, optical fibers, CD-ROM (CD-Read Only Memory), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.
[0125] The above computer-readable storage medium may be included in an image generation device; or it may exist independently without being assembled into the image generation device.
[0126] The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed by an image generation device, the image generation device is caused to execute the above image generation method.
[0127] Computer program code for performing the operations of the present application may be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a LAN (Local Area Network) or a WAN (Wide Area Network), or it may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0128] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0129] The modules described in the embodiments of the present application can be implemented in software or in hardware. In some cases, the name of the module does not constitute a limitation on the unit itself.
[0130] The readable storage medium provided by the present application is a computer-readable storage medium that stores computer-readable program instructions (i.e., computer programs) for executing the above-described image generation method, and can solve the technical problem that the existing image generation method has low efficiency in allocating computing resources for the image generation model, resulting in low image generation efficiency and poor image quality. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by the present application are the same as those of the image generation method provided in the above embodiments, and will not be elaborated here.
[0131] The present application also provides a computer program product, including a computer program that, when executed by a processor, implements the image generation method as described above.
[0132] The computer program product provided by the present application can solve the technical problem that the existing image generation method has low efficiency in allocating computing resources for the image generation model, resulting in low image generation efficiency and poor image quality. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as those of the image generation method provided in the above embodiments, and will not be elaborated here.
[0133] The above are only some embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structural transformation made using the description and drawings of the present application under the technical concept of the present application, or any direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.
[0134] This application discloses A1, an image generation method, which includes:
[0135] In response to an image generation task, input the image generation task into a preset image generation model;
[0136] Process the image generation task through the preset image generation model to generate a target image, where the preset image generation model includes at least one control module, the control module is set in the target control layer of the preset image generation model, and the target control layer is determined based on the influence of each control layer in the preset image generation model on the model output result.
[0137] A2. The image generation method as described in A1, before inputting the image generation task into the preset image generation model in response to the image generation task, further includes:
[0138] Calculate the correlation scores of each control layer in the initial image generation model, where the correlation scores are used to measure the influence of each control layer on the model output result;
[0139] Select a target control layer from each of the control layers according to the correlation scores;
[0140] Set a control module in the target control layer to obtain the preset image generation model.
[0141] A3. The image generation method as described in A2, the calculation of the correlation scores of each control layer in the initial image generation model includes:
[0142] Skip each control layer in the initial image generation model in turn for model inference;
[0143] Obtain the model metrics during the model inference process, and generate the correlation scores of each control layer in the model based on the model metrics.
[0144] A4. The image generation method as described in A3, the obtaining of the model metrics during the model inference process and generating the correlation scores of each control layer in the model based on the model metrics includes:
[0145] Obtain the image quality loss and control precision loss during the model inference process, where the image quality loss and the control precision loss are used to measure the influence of skipping each control module for model inference on the initial image generation model;
[0146] Calculate the correlation scores of each control layer in the model based on the image quality loss and the control precision loss.
[0147] A5. The image generation method as described in A2, wherein the step of selecting a target control layer from each of the control layers according to the correlation score includes:
[0148] Sorting each of the control layers according to the correlation score;
[0149] Selecting a preset number of target control layers from each of the control layers according to the sorting result.
[0150] A6. The image generation method as described in any one of A1 to A5, wherein the preset image generation model is a diffusion Transformer model, and the control module is a correlation-guided control module; the step of processing the image generation task through the preset image generation model to generate a target image includes:
[0151] Processing the image generation task through the correlation-guided control module to generate a target image, wherein the correlation-guided control module is used to unify the token mixing operation and the channel mixing operation of the diffusion Transformer model into a target operation.
[0152] A7. The image generation method as described in A6, wherein the correlation-guided control module includes: a two-dimensional random mixing module; the step of processing the image generation task through the preset image generation model based on the target operation to generate a target image includes:
[0153] Converting the image generation task into a token sequence;
[0154] Processing the token sequence through the two-dimensional random mixing module to generate a target image, wherein the two-dimensional random mixing module is used to perform information interaction across the token dimension and the channel dimension.
[0155] A8. The image generation method as described in A7, wherein the step of processing the token sequence through the two-dimensional random mixing module to generate a target image includes:
[0156] Randomly dividing the token sequence in the channel dimension based on the correlation scores of each target control layer to form a plurality of channel groups;
[0157] Randomly dividing tokens within each channel group to form a plurality of token groups;
[0158] Performing attention calculation within the randomly divided plurality of channel groups and plurality of token groups;
[0159] After completing the attention calculation, reversing the positions of the tokens and channels to generate a target image.
[0160] A9. The image generation method as described in A8, after completing the attention calculation, reversely restoring the positions of the tokens and channels to generate a target image, including:
[0161] After completing the attention calculation, obtaining a first random factor for randomly partitioning the token sequence and a second random factor for randomly partitioning the tokens;
[0162] Based on the first random factor, reversely restoring the positions of the tokens, and based on the second random factor, reversely restoring the positions of the channels to generate a target image.
[0163] A10. The image generation method as described in A7, converting the image generation task into a token sequence, including:
[0164] Preprocessing the image generation task to obtain a preprocessed image generation task;
[0165] Converting the preprocessed image generation task into an embedding vector;
[0166] Sorting the embedding vector according to the image generation task to obtain a token sequence.
[0167] A11. The image generation method as described in any one of A1 to A5, after processing the image generation task through the preset image generation model to generate a target image, further including:
[0168] Obtaining user information and constructing a user portrait based on the user information through a large language model;
[0169] Based on the user portrait, performing personalized adjustment on the target image to obtain a personalized image.
[0170] This application also discloses B12. An image generation device, the image generation device includes:
[0171] A task input module, configured to input the image generation task into a preset image generation model in response to an image generation task;
[0172] An image generation module, configured to process the image generation task through the preset image generation model to generate a target image, wherein the preset image generation model includes at least one control module, the control module is arranged in a target control layer of the preset image generation model, and the target control layer is determined based on the influence of each control layer in the preset image generation model on the model output result.
[0173] B13. The image generation device as described in B12, the image generation device further includes:
[0174] A model construction module, configured to calculate the correlation scores of each control layer in an initial image generation model, where the correlation scores are used to measure the influence of each control layer on the model output result; select a target control layer from each of the control layers according to the correlation scores; and set a control module in the target control layer to obtain the preset image generation model.
[0175] B14. The image generation device according to B13, wherein the model construction module is further configured to sequentially skip each control layer in the initial image generation model for model inference; obtain model metrics during the model inference, and generate the correlation scores of each control layer in the model based on the model metrics.
[0176] B15. The image generation device according to B14, wherein the model construction module is further configured to obtain an image quality loss and a control precision loss during the model inference, where the image quality loss and the control precision loss are used to measure the influence of skipping each control module for model inference on the initial image generation model; and calculate the correlation scores of each control layer in the model based on the image quality loss and the control precision loss.
[0177] B16. The image generation device according to B13, wherein the model construction module is further configured to sort each of the control layers according to the correlation scores; and select a preset number of target control layers from each of the control layers according to the sorting result.
[0178] B17. The image generation device according to any one of B12 to B16, wherein the preset image generation model is a diffusion Transformer model, and the control module is a correlation-guided control module; the image generation module is further configured to process the image generation task through the correlation-guided control module to generate a target image, where the correlation-guided control module is configured to unify the token mixing operation and the channel mixing operation of the diffusion Transformer model into a target operation.
[0179] The present application also discloses C18. An image generation device, comprising: a memory, a processor, and an image generation program stored on the memory and executable on the processor, where when the image generation program is executed by the processor, the image generation method described above is implemented.
[0180] The present application also discloses D19. A storage medium, on which an image generation program is stored, where when the image generation program is executed by a processor, the image generation method described above is implemented.
[0181] The present application also discloses E20, a computer program product, which includes an image generation program. When the image generation program is executed by a processor, it implements the image generation method as described above.
Claims
1. An image generation method, characterized in that: The image generation method comprises: In response to an image generation task, inputting the image generation task into a preset image generation model; The image generation task is processed by the preset image generation model to generate a target image, wherein the preset image generation model includes at least one control module, and the control module is arranged in a target control layer of the preset image generation model, and the target control layer is determined based on the influence of each control layer in the preset image generation model on the model output result.
2. The image generation method according to claim 1, characterized in that: In response to the image generation task, before inputting the image generation task into a preset image generation model, the method further includes: Calculating the correlation score of each control layer in the initial image generation model, wherein the correlation score is used to measure the influence of each control layer on the model output result; Selecting a target control layer from the control layers according to the correlation scores; A control module is set in the target control layer to obtain the preset image generation model.
3. The image generation method according to claim 2, characterized in that: The calculating the correlation scores of each control layer in the initial image generation model includes: Skipping each control layer in the initial image generation model in turn to perform model reasoning; The model indicators in the model reasoning process are obtained, and the correlation scores of each control layer in the model are generated based on the model indicators.
4. The image generation method according to claim 3, characterized in that: The obtaining of model indicators in the model reasoning process and generating correlation scores of each control layer in the model based on the model indicators include: Acquire image quality loss and control accuracy loss during model reasoning, wherein the image quality loss and the control accuracy loss are used to measure the impact of skipping the control modules to perform model reasoning on the initial image generation model; The correlation scores of the various control layers in the model are calculated based on the image quality loss and the control accuracy loss.
5. The image generation method according to claim 2, characterized in that: The selecting a target control layer from the various control layers according to the correlation scores includes: Sorting the control layers according to the correlation scores; A preset number of target control layers are selected from the control layers according to the sorting result.
6. The image generation method according to any one of claims 1 to 5, characterized in that: The preset image generation model is a diffusion Transformer model, and the control module is a correlation guidance control module; the image generation task is processed by the preset image generation model to generate a target image, including: The image generation task is processed by the relevance guidance control module to generate a target image, wherein the relevance guidance control module is used to unify the word-unit mixing operation and the channel mixing operation of the diffusion Transformer model into a target operation.
7. An image generating device, characterized in that: The image generating device comprises: A task input module, configured to input the image generation task into a preset image generation model in response to the image generation task; An image generation module is used to process the image generation task through the preset image generation model to generate a target image, wherein the preset image generation model includes at least one control module, and the control module is arranged in a target control layer of the preset image generation model, and the target control layer is determined based on the influence of each control layer in the preset image generation model on the model output result.
8. An image generating device, characterized in that: The image generating device comprises: a memory, a processor, and an image generating program stored in the memory and executable on the processor, wherein the image generating program implements the image generating method according to any one of claims 1 to 6 when executed by the processor.
9. A storage medium, characterized in that: The storage medium stores an image generation program, and when the image generation program is executed by the processor, the image generation method according to any one of claims 1 to 6 is implemented.
10. A computer program product, characterized in that The computer program product comprises an image generation program, and when the image generation program is executed by a processor, the image generation method according to any one of claims 1 to 6 is implemented.