Image generation method, system and equipment based on attribute control and medium
By extracting and decoupling semantic features in the diffusion model, screening and intervening in key features, the control inaccuracy problem in the diffusion model generation process is solved, and precise control of image generation and content consistency improvement is achieved.
Patent Information
- Application Number
- CN202510283453.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-07-04
AI Technical Summary
The existing diffusion models lack precise control of the generated results in the image generation process, resulting in randomness and inaccuracy of content, making it difficult to meet the application needs of precise content control.
By obtaining the image to be adjusted and its target attributes, semantic feature groups are extracted using the diffusion model and feature decoupling, sparse semantic features are screened out, feature intervention is performed based on gradient attribution, reconstructing and iteratively denoising to generate target images.
It improves the controllability and accuracy of image generation, ensures precise control of target features while retaining the original image features, and improves the reliability and consistency of generated content.
Smart Images

Figure CN120259459A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly relates to an image generation method, system, device and medium based on attribute control. Background Art
[0002] In recent years, diffusion models, as a cutting-edge technology in generative artificial intelligence, have shown great potential and broad application prospects in the fields of image, text and multimodal content generation. This generation technology can gradually generate high-quality and high-fidelity image content from random noise through probabilistic inference and denoising processes. However, despite the rapid development of the technology, diffusion models still face serious controllability challenges in the content generation process. The generation process of traditional diffusion models is essentially a highly random and uncertain process. The model often lacks precise control over the generation results during content synthesis, which makes the generated content and attributes show great randomness and inaccuracy. Especially in application scenarios that require precise content control, this randomness of the generated content and inaccuracy of control have become the key bottlenecks restricting the implementation of generative technologies. Therefore, there is a need to provide an image generation method, system, device and medium based on attribute control. Summary of the Invention
[0003] In view of the above-mentioned disadvantages of the prior art, the purpose of the present invention is to provide an image generation method, system, device and medium based on attribute control, which improves the problems of poor controllability and accuracy of the images generated by the existing diffusion models.
[0004] To achieve the above object and other related objects, the present invention provides an image generation method based on attribute control, including: obtaining an image to be adjusted and its target attribute; wherein, the target attribute is used to characterize the target that the image needs to be adjusted; inputting the image into a diffusion model, and extracting a semantic feature group of the image through the diffusion model; decoupling the features of the semantic feature group to obtain a sparse semantic feature group; screening at least one target semantic feature from the sparse semantic feature group according to the gradient attribution method; intervening in the screened target semantic features based on the target attribute to obtain the intervened target semantic features; reconstructing the semantic feature group according to the intervened target semantic features, and inputting the reconstructed semantic feature group into the diffusion model, and iteratively denoising the image based on the reconstructed semantic feature group to generate an intervened target image.
[0005] In one embodiment of the present invention, the sparse semantic feature group is obtained through the encoder module of the autoencoder. The decoupling of the semantic feature group to obtain the sparse semantic feature group includes: inputting the semantic feature into the encoder module of the autoencoder, mapping the semantic feature to a high-dimensional semantic space to generate a plurality of initial sparse semantic features; calculating the activation values of each initial sparse semantic feature; sorting all the activation values to form an activation value sequence; selecting a preset first number of initial sparse semantic features from the activation value sequence in the descending order of activation values as the final sparse semantic features to form a sparse semantic feature group.
[0006] In one embodiment of the present invention, the screening of at least one target semantic feature from the sparse semantic feature group according to the gradient attribution method includes: calculating the contribution degree of each sparse semantic feature to the target attribute based on the integrated gradient method; sorting the contribution degrees of all the sparse semantic features to form a contribution degree sequence; selecting a preset second number of sparse semantic features from the contribution degree sequence in the descending order of contribution degrees as the target semantic features.
[0007] In one embodiment of the present invention, the intervention on the screened target semantic feature based on the target attribute to obtain the intervened target semantic feature includes: judging the relationship between the target attribute and the screened target semantic feature: if the target attribute and the screened target semantic feature are positively correlated, amplifying the screened target semantic feature to obtain the intervened target semantic feature; if the target attribute and the screened target semantic feature are negatively correlated, suppressing the screened target semantic feature to obtain the intervened target semantic feature.
[0008] In one embodiment of the present invention, the intervened target semantic feature reconstructs the semantic feature group through the decoder module of the autoencoder. Reconstructing the semantic feature group according to the intervened target semantic feature and inputting the reconstructed semantic feature group into the diffusion model, and iteratively denoising the image based on the reconstructed semantic feature group to generate the intervened target image includes: inputting the intervened target semantic feature into the decoder module of the autoencoder to reconstruct the intervened target semantic feature back to the dimension of the original semantic feature group to obtain the reconstructed semantic feature group; inputting the reconstructed semantic feature group into the diffusion model and iteratively denoising the image based on the reconstructed semantic feature group to generate the intervened target image; wherein, the position where the reconstructed semantic feature group is input into the diffusion model is the same as the position where the original semantic feature group is extracted from the diffusion model.
[0009] In one embodiment of the present invention, the encoder is a k-sparse autoencoder.
[0010] In one embodiment of the present invention, the diffusion model is an unconditional diffusion model or a conditional diffusion model.
[0011] In one embodiment of the present invention, there is also provided an image generation system based on attribute control, the system comprising: an image acquisition module for acquiring an image to be adjusted and its target attribute; wherein the target attribute is used to characterize the target for which the image needs to be adjusted; a feature extraction module for inputting the image into a diffusion model and extracting a semantic feature group of the image through the diffusion model; a feature decoupling module for decoupling the semantic feature group to obtain a sparse semantic feature group; a target extraction module for screening at least one target semantic feature from the sparse semantic feature group according to the gradient attribution method; an intervention module for intervening on the screened target semantic feature based on the target attribute to obtain an intervened target semantic feature; and a reconstruction module for reconstructing the semantic feature group according to the intervened target semantic feature and inputting the reconstructed semantic feature group into the diffusion model, and iteratively denoising the image based on the reconstructed semantic feature group to generate an intervened target image.
[0012] In one embodiment of the present invention, there is also provided an electronic device, comprising: one or more processors; a storage device for storing one or more programs, which when executed by the one or more processors, cause the electronic device to implement the image generation method based on attribute control as described in any one of the above.
[0013] In one embodiment of the present invention, there is also provided a computer-readable storage medium, on which a computer program is stored, which when executed by a processor of a computer, causes the computer to execute the image generation method based on attribute control as described in any one of the above.
[0014] As described above, an image generation method, system, device and medium based on attribute control of the present invention have the following beneficial effects: decoupling the semantic feature group extracted from the diffusion model to generate a sparse semantic feature group, so as to map the hidden state in the diffusion model from the original latent space to a sparse semantic space, thereby improving the independence and controllability of features. Screening target semantic features from the sparse semantic feature group to ensure that subsequent adjustment interventions act on key features. Intervening and regulating these key features to precisely control the expression of target features, and generating an intervened target image through the diffusion model, ensuring that the adjustment is precisely regulated for the target attribute while retaining the original image features. The present invention improves the controllability and accuracy of the generated image. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 It is a schematic flowchart of an image generation method based on attribute control provided by an embodiment of the present invention;
[0016] Figure 2 Flow chart of feature decoupling steps provided by an embodiment of the present invention;
[0017] Figure 3 Flow chart of precise positioning of target content features provided by an embodiment of the present invention;
[0018] Figure 4 Flow chart of precise intervention of features provided by an embodiment of the present invention;
[0019] Figure 5 is an example test diagram provided by an embodiment of the present invention;
[0020] Figure 6 Shown is a structural block diagram of an image generation system based on attribute control provided by an embodiment of the present invention;
[0021] Figure 7 Shown is a schematic structural diagram of an electronic device according to an embodiment of the present invention. Detailed implementation manners
[0022] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0023] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Therefore, only the components related to the present invention are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The types, quantities, and proportions of the components in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.
[0024] In the following description, a large number of details are explored to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.
[0025] The present invention provides an attribute - based image generation method, which can deeply analyze and precisely control the generation process of the diffusion model. For the first time, the feature decoupling and gradient attribution methods are systematically applied to the generation control of the diffusion model. By revealing the internal generation decision mechanism of the model, identifying and adjusting the target attribute content features in the model, an accurate regulation technology at the feature level is established to generate content highly consistent with expectations. The present invention proposes an accurate feature intervention paradigm for the generation system. This method significantly improves the controllability and consistency of the generated content while retaining the original performance of the model through targeted regulation of key generation features. This not only provides an innovative technical path for multi - modal generative artificial intelligence systems but also opens up a new research direction for improving the accuracy and reliability of the generated content.
[0026] Please refer to Figure 1 , the attribute - based image generation method proposed by the present invention includes the following steps:
[0027] S1. Obtain the image to be adjusted and its target attribute; wherein, the target attribute is used to characterize the target for which the image needs to be adjusted.
[0028] The target attribute is used to characterize the features that the image needs to be adjusted in specific aspects, such as gender, age, expression, style, etc. The target attribute not only determines the direction of adjustment, for example, adjusting the young to the old or adjusting the non - smiling to smiling, but also provides a clear guiding basis for subsequent feature extraction and intervention. To ensure the accuracy of adjustment, this attribute can be user - preset label information or obtained by the model's automatic inference method. Through the target attribute, the diffusion model can be guided to gradually generate the target image that meets the expected adjustment during the iterative denoising process, so as to ensure that while retaining the main features of the original image, only the target features are accurately modified.
[0029] S2. Input the image into the diffusion model, and extract the semantic feature group of the image through the diffusion model.
[0030] Input the image to be adjusted into the diffusion model, and use the powerful feature extraction ability of the diffusion model to deeply analyze the image. Among them, the diffusion model is a U - Net model or an improved model based on U - Net, and the diffusion model can be an unconditional diffusion model or a conditional diffusion model. To further improve the control ability of the image content, a semantic feature group is extracted from the middle layer of the diffusion model, that is, a group of multi - dimensional semantic features, and each semantic feature is used to represent a representation of the image in the latent space. This kind of feature, as the hidden state generated by the diffusion model at the corresponding time step, carries the semantic information of the image. These features will be used as the basis for subsequent feature decoupling, target feature screening and intervention, ensuring that subsequent adjustment operations can accurately act on the target attribute without affecting other irrelevant content.
[0031] S3. Decouple the semantic feature group to obtain a sparse semantic feature group.
[0032] Decouple the extracted semantic feature group to more clearly separate the representation information of different semantic attributes. Since the semantic feature group of the diffusion model is high-dimensional and multiple features are mixed together, there are a large number of features that are invalid for the current target attribute, making it difficult to precisely control the target attribute directly through feature intervention. Therefore, a feature decoupling method is needed to decouple the original semantic feature group, mapping it from the hidden state space to a high-dimensional sparse semantic space. Thus, each semantic feature is decomposed into more independent and controllable sparse features, minimizing the mutual influence between different features, and obtaining a sparse semantic feature group.
[0033] In an embodiment of the present invention, the sparse semantic feature group is obtained through the encoder module of an autoencoder. The decoupling of the semantic feature group to obtain a sparse semantic feature group includes:
[0034] Input the semantic feature group into the encoder module of the autoencoder, map the semantic feature to a high-dimensional semantic space, and generate multiple initial sparse semantic features;
[0035] Calculate the activation values of each initial sparse semantic feature;
[0036] Sort all the activation values to form an activation value sequence;
[0037] Select a preset first number of initial sparse semantic features from the activation value sequence in descending order of activation values as the final sparse semantic features, forming a sparse semantic feature group.
[0038] Input the semantic feature group into the encoder module of a pre-trained autoencoder. This module uses learnable weights to map the input semantic feature group from the hidden state space of the diffusion model to a high-dimensional sparse semantic space, thereby obtaining multiple initial sparse semantic features and forming an initial sparse semantic feature group. These sparse semantic features represent decoupled semantic features, such as age, gender, style, lighting, etc. By calculating the activation values of each sparse semantic feature and using the TopK operation to select the most important k features, these features are called "fired features", and the remaining features are called "unfired features", thereby screening out representative features. Specifically, calculate the activation values of each initial sparse semantic feature as shown in formula (1):
[0039] s = φ(h) = TopK(W enc (h - b pre )) (1)
[0040] Among them, s is the selected sparse semantic feature group, and φ(h) is a feature transformation function used to transform the hidden state h of the diffusion model into the sparse semantic space. W enc and b pre are the weight matrix and bias term of the encoder module respectively, h is the hidden state of the diffusion model, that is, the semantic feature group extracted from the middle layer of the diffusion model, TopK is a sorting operation, that is, selecting the top k sparse semantic features with the highest activation values. By calculating the activation values of different initial sparse semantic features, the importance of different sparse semantic features can be measured. After the calculation is completed, all the activation values are sorted to form an activation value sequence, and the preset first quantity (such as k) of initial sparse semantic features are selected in descending order of the activation values as the final sparse semantic features, and these sparse semantic features form a sparse semantic feature group. This sparse semantic feature group retains the most critical semantic information while removing redundant or unimportant features, making the subsequent target feature screening and intervention more accurate, and laying a foundation for the controllability adjustment of the diffusion model. It can be understood that the dimension of the finally formed sparse semantic feature group is the same as that of the initial sparse semantic feature group. When obtaining the final sparse semantic feature group from the initial sparse semantic feature group, the features are not screened by reducing the number of features, but by adjusting the feature values, for example, assigning higher weights or enhancing signals to the activated features, and suppressing or setting to zero the unactivated features, so as to form the final sparse semantic feature group in the feature space of the same dimension.
[0041] S4. According to the gradient attribution method, at least one target semantic feature is screened out from the sparse semantic feature group.
[0042] Using the gradient attribution method, calculate the contribution degree of each sparse semantic feature to the target attribute. According to the magnitude of the contribution degree, all sparse semantic features are sorted, and at least one target semantic feature with the greatest influence on the target attribute is screened out. Through the above process, it can be ensured that the selected features can accurately reflect the change direction of the target attribute and are used for subsequent feature intervention. Exemplarily, when adjusting the gender, select the semantic features that affect facial features (such as beards or jaw contours), and when adjusting the age, select the features that affect wrinkles, skin smoothness, etc.
[0043] In an embodiment of the present invention, the screening of at least one target semantic feature from the sparse semantic feature group according to the gradient attribution method includes:
[0044] Based on the integrated gradient method, calculate the contribution degree of each sparse semantic feature to the target attribute;
[0045] Sort the contribution degrees of all sparse semantic features to form a contribution degree sequence;
[0046] Select a preset second number of sparse semantic features from the contribution degree sequence in descending order of contribution degree as the target semantic features.
[0047] As shown in formula (2), based on the integrated gradient method, calculate the contribution degree of each sparse semantic feature to the target attribute through the attribution evaluation function:
[0048]
[0049] Among them, S(s i ; x) is the contribution degree of the i-th target semantic feature s i to the target attribute with respect to the input x of the diffusion model, s′ i is the preset baseline feature, and F x (*) is the measurement of the target attribute with respect to the input x of the diffusion model for the input data *, which can be obtained by inputting the sparse semantic feature group generated by the encoder model of the autoencoder into the pre-trained classifier F, so as to obtain the measurement of the target attribute x for the input data. is the gradient of s i to the target attribute, and α is the integration variable. Specifically, input the semantic feature group generated by the diffusion model into the encoder module of the autoencoder, and the encoder calculates the contribution degree of this feature to the image target attribute (such as age, gender, style, etc.), which can accurately evaluate the role played by each sparse semantic feature in the image generation process. Sort the contribution degrees of all sparse semantic features to form a contribution degree sequence, where the higher the contribution degree, the greater the impact of the feature on the target attribute. To ensure the accuracy of feature intervention, select a preset second number (such as τ) of sparse semantic features from the contribution degree sequence in descending order of contribution degree as the target semantic features. This set of target semantic features has a decisive role in the target attribute during the generation process and will be used as the input for subsequent feature intervention. It can be understood that the baseline feature can be a zero vector or other values, and the specific value can be adaptively set based on actual needs and is not limited here.
[0050] S5. Intervene on the selected target semantic features based on the target attribute to obtain the intervened target semantic features.
[0051] In an embodiment of the present invention, the intervening on the selected target semantic features based on the target attribute to obtain the intervened target semantic features includes:
[0052] Judge the relationship between the target attribute and the selected target semantic features:
[0053] If the target attribute and the selected target semantic features are positively correlated, perform amplification processing on the selected target semantic features to obtain the intervened target semantic features;
[0054] If the target attribute is negatively correlated with the selected target semantic feature, the selected target semantic feature is suppressed to obtain the intervened target semantic feature.
[0055] Determine the relationship between the target attribute and the selected target semantic feature, that is, analyze the action direction of this semantic feature in the diffusion model and its influence trend on the target attribute. If the target attribute is positively correlated with the selected target semantic feature, that is, the enhancement of the target semantic feature will increase the expression intensity of the target attribute, then this feature is amplified, such as using scaling operations or linear gain adjustments to enhance the expression of this feature in the image and obtain the intervened target semantic feature. Conversely, if the target attribute is negatively correlated with the selected target semantic feature, that is, the enhancement of this semantic feature will suppress the target attribute, then this feature is suppressed, such as reducing the feature amplitude or performing an inhibitory offset to reduce the influence of this feature on image generation and ensure that the adjustment direction meets the requirements of the target attribute. The intervened target semantic feature is used to update the feature input of the diffusion model, so that the model gradually generates an image that meets the intervention goal during the denoising iteration process, while maintaining the stability of non-target features and ensuring the naturalness and consistency of the adjustment. Exemplarily, in the sparse semantic space, an amplification (Amplify) or suppression (Suppress) operation is performed on the i-th target feature s i to adjust the performance of the target attribute in the generated content, as shown in formula (3):
[0056]
[0057] where β is a preset weighting parameter. When scaling, when β is greater than 1, it means amplifying the target feature, and when β is less than 1, it means shrinking the target feature; when adding, when β is a positive number, it means amplifying or adding the target feature, and when β is a negative number, it means shrinking the target feature.
[0058] S6. Reconstruct the semantic feature group according to the intervened target semantic feature, input the reconstructed semantic feature group into the diffusion model, and perform iterative denoising on the image based on the reconstructed semantic feature group to generate the intervened target image.
[0059] Use the decoder module of the autoencoder to reconstruct the sparse semantic feature group, reconstruct the sparse semantic features back to the original latent space, and superimpose them back on the original latent state h to obtain the reconstructed semantic feature group To match the input format of the diffusion model. The reconstructed semantic feature group is input into the diffusion model as the guiding information for the next denoising iteration. Based on these adjusted features, the diffusion model gradually updates the image in multiple time steps during the denoising process, making it change towards the target attributes while maintaining the overall consistency and structural stability of the original image. After iterative denoising, the model generates the target image after intervention, and its visual features meet the requirements of the target attributes.
[0060] The target semantic features after intervention are used to reconstruct the semantic feature group by the decoder module of the autoencoder. Reconstruct the semantic feature group based on the target semantic features after intervention, and input the reconstructed semantic feature group into the diffusion model. Iteratively denoise the image based on the reconstructed semantic feature group to generate the target image after intervention, including:
[0061] Input the target semantic features after intervention into the decoder module of the autoencoder, and reconstruct the target semantic features after intervention back to the dimension of the original semantic feature group to obtain the reconstructed semantic feature group;
[0062] Input the reconstructed semantic feature group into the diffusion model, and iteratively denoise the image based on the reconstructed semantic feature group to generate the target image after intervention; wherein, the position where the reconstructed semantic feature group is input into the diffusion model is the same as the position where the original semantic feature group is extracted from the diffusion model.
[0063] Input the target semantic features after intervention into the decoder module of the autoencoder to map the target semantic features after intervention back to the dimension of the original semantic feature group to obtain the reconstructed semantic feature group. It can be understood that the reconstructed semantic feature group and the initial semantic feature group have the same dimension, thus maintaining the consistency of the feature space to ensure that the diffusion model can normally receive the modified feature information. Input the reconstructed semantic feature group into the diffusion model, and during the denoising iteration of the diffusion model, adjust the image content based on the modified feature information to make it gradually change towards the target attributes, and finally generate the target image after intervention. It should be noted that the position where the reconstructed semantic feature group is input into the diffusion model is the same as the position where the original semantic feature group is extracted from the diffusion model to ensure that the diffusion model can correctly apply the modified information during the processing without affecting the overall structure of the model or other irrelevant features. This processing process ensures the accuracy of the adjustment while maintaining the overall style of the original image, making the finally generated image meet the target modification requirements and be natural, coherent, and distortion-free.
[0064] It is understandable that the autoencoder and classifier of the present invention are pre-trained. The autoencoder can be a k-sparse autoencoder. The specific implementation method of training the k-sparse autoencoder is as follows: First, use an unconditional diffusion model or a conditional diffusion model to generate 35,000 pictures, and extract the hidden state h from the diffusion model as a training data set. Initialize the weights of the encoder module and the decoder module of the k-sparse autoencoder (k-SAE), and initialize the weight of the decoder module to the transpose of the encoder module weight to ensure the structural symmetry of the initial autoencoder. Use the reconstruction error as the loss function, and the training goal is to minimize the original hidden state h and the reconstructed hidden state h. The difference between them is shown in formula (4):
[0065]
[0066] in, is the original hidden state and the hidden state reconstructed by k-SAE through encoding-decoding The mean square error between them, h is the original hidden state of the diffusion model, that is, the semantic feature group generated by the diffusion model, is the reconstructed hidden state, that is, the reconstructed semantic feature group. As described in formula (5), the unexplained variance score (Fraction of Variance Unexplained, FVU) is used as the evaluation index to measure the reconstruction performance of the k-sparse autoencoder:
[0067]
[0068] Among them, var[h] is the variance of the hidden state h, and FVU is the variance of the original hidden state h and the reconstructed hidden state h. The smaller the FVU value, the better k-SAE can reconstruct the hidden state, that is, the model has a stronger feature extraction ability. Through the above training process, k-SAE can effectively decouple the hidden state of the diffusion model and generate sparse and highly controllable semantic features to support subsequent target feature screening and intervention.
[0069] Furthermore, in the present invention, the specific implementation method of training a lightweight classifier is as follows: for each category, 1000 images are generated using a target diffusion model, and a hidden state h is obtained as a training data set. The classification result of the image classifier is used as the label of the hidden state h data set. The training classifier is a linear layer, and the hidden state h is classified based on the classification result of the image classifier.
[0070] See also Figure 2, the present invention first uses a k-sparse autoencoder (k-SAE) to decouple the feature of the hidden state of the U-Net neural network architecture in the diffusion model from a complex multi-semantic space. This process involves mapping the hidden state in the U-Net from the original latent space to a sparse semantic space. Specifically, during the denoising process of the diffusion model, the image x t+1 generated at the (t + 1)-th time step is input into the U-Net of the diffusion model. The encoder on the left extracts multi-level feature information through multiple layers of 3×3 convolutions (Conv 3×3) and downsampling. The decoder on the right gradually restores the image details through upsampling and skip connections (Copy). During the denoising process of the U-Net, the hidden state h (i.e., the semantic feature group) generated at any intermediate layer is extracted. This hidden state is a multi-semantic neuron, and one feature can be encoded as multiple different semantic units simultaneously. The hidden state h of the diffusion model is input into the k-sparse autoencoder, and through the encoder module, h is mapped from the hidden state space to a high-dimensional sparse semantic space, obtaining a sparse semantic feature group. In the sparse semantic feature group, the black dots are the activated features, indicating the features that play a key role in the current image generation process, such as the features affecting gender or age, and the white dots are the unactivated features, indicating the features that are not significantly activated in the current image and have little impact on the current generation task. Through the decoder module of the k-sparse autoencoder, the features obtained by the above encoder module are reconstructed, thereby reconstructing the original hidden state space and inputting it to the original position of the diffusion model. After denoising, the target image is obtained.
[0071] As Figure 3 shown, the original semantic features are input into the encoder module of the k-sparse autoencoder, and the original semantic features are decoupled and mapped to a high-dimensional sparse semantic space, obtaining a sparse semantic feature group. Among them, the activated features in this feature group indicate the features that are highly activated in the current input image and play a key role in the generation result, and the unactivated features indicate the features that do not have an obvious impact on the generation process currently. The target features are screened out and activated through the gradient attribution method. The activated target features refer to specific target attributes (such as gender, age, race, etc.), which are in the activated state in the current state, and the unactivated target features indicate that they are not activated currently, that is, the features that have little impact on the current image. Exemplarily, S 24 represents the feature that controls the gender of image generation. Through this feature, the gender can be adjusted from female to male, or vice versa. S 18 and S 26 represent the features used to adjust the age, making the image change from young to old, or vice versa. S 30 represents the feature used to adjust the racial feature, for example, changing from Caucasian to African American, or vice versa.
[0072] The present invention realizes precise control of the generated content by intervening in the identified target features. For example Figure 4 As shown, the noisy image x t+1 in the diffusion process is input into the diffusion model to extract the hidden state, that is, the semantic feature group. It is input into the k-sparse autoencoder, and the encoder module of the k-sparse autoencoder maps the hidden state generated by the diffusion model to a high-dimensional sparse semantic space to separate different semantic features, obtaining a sparse semantic feature group. Among the sparse semantic feature groups output by the encoder, the target feature S 24 is selected for intervention and adjustment. Specifically, the performance of the target feature in the finally generated image can be affected by magnifying or suppressing the target feature. The original target feature in the sparse semantic feature group is updated with the intervened target feature, and the updated feature group is reconstructed through the decoder module, and the reconstructed feature is input into the diffusion model to continue denoising, thereby generating the target image.
[0073] The method of the present invention has a certain control effect on the target feature intervention of the image generated by the diffusion model. Figure 5a shows the control of the features related to the "gender" attribute, while Figure 5b shows the control of the features of the "age" attribute. In Figure 5a , the original images are all female. By applying the target feature intervention steps of the present invention, we can transform these images into images with male features. While keeping other features of the original images (such as smiling) unchanged, this transformation effectively adjusts the gender feature, thereby generating images with the desired gender features. In Figure 5b , the original images are all young. By intervening in the age feature, we can make these images show older age features, such as wrinkles and different hairstyles, while other features (such as gender) remain unchanged. This shows that the method of the present invention can precisely control and adjust specific age features in the image.
[0074] Please refer to Figure 6, the attribute-based image generation system 100 includes: an image acquisition module 110, a feature extraction module 120, a feature decoupling module 130, a target extraction module 140, an intervention module 150, and a reconstruction module 160. The image acquisition module 110 is configured to acquire an image to be adjusted and its target attribute; wherein, the target attribute is used to characterize the target for which the image needs to be adjusted. The feature extraction module 120 is configured to input the image into a diffusion model and extract a semantic feature group of the image from the diffusion model. The feature decoupling module 130 is configured to perform feature decoupling on the semantic feature group to obtain a sparse semantic feature group. The target extraction module 140 is configured to screen at least one target semantic feature from the sparse semantic feature group according to the gradient attribution method. The intervention module 150 is configured to intervene in the screened target semantic feature based on the target attribute to obtain an intervened target semantic feature. The reconstruction module 160 is configured to reconstruct the semantic feature group according to the intervened target semantic feature, input the reconstructed semantic feature group into the diffusion model, and perform iterative denoising on the image based on the reconstructed semantic feature group to generate an intervened target image.
[0075] For the specific limitations of the attribute-based image generation system, reference can be made to the limitations of the attribute-based image generation method described above, which will not be elaborated here. Each module in the above-mentioned attribute-based image generation system can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in or independent of the processor in the computer device in hardware format, or stored in the memory in the computer device in software format, so as to facilitate the processor to call the corresponding operations of each of the above modules.
[0076] It should be noted that, in order to highlight the innovative part of the present invention, modules not closely related to solving the technical problems proposed by the present invention are not introduced in this embodiment, but this does not mean that there are no other modules in this embodiment.
[0077] Please refer to Figure 7 , the electronic device 1 may include a memory 11, a processor 12, and a bus, and may further include a computer program stored in the memory 11 and executable on the processor 12, such as an attribute-based image generation program.
[0078] Among them, the memory 11 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the memory 11 may be an internal storage unit of the electronic device 1, such as the mobile hard disk of the electronic device 1. In some other embodiments, the memory 11 may also be an external storage device of the electronic device 1, such as a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device 1. Further, the memory 11 may also include both the internal storage unit and the external storage device of the electronic device 1. The memory 11 can be used not only to store application software and various types of data installed in the electronic device 1, such as codes for image generation based on attribute control, etc., but also to temporarily store data that has been output or will be output.
[0079] In some embodiments, the processor 12 may be composed of integrated circuits. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple integrated circuits with the same or different functions packaged together, including a combination of one or more Central Processing Units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips, etc. The processor 12 is the control core (Control Unit) of the electronic device 1, connecting various components of the entire electronic device 1 through various interfaces and circuits, and by running or executing programs or modules stored in the memory 11 (such as programs for image generation based on attribute control, etc.), and calling data stored in the memory 11, to execute various functions of the electronic device 1 and process data.
[0080] The processor 12 executes the operating system of the electronic device 1 and various installed application programs. The processor 12 executes the application programs to implement the steps in the above-mentioned image generation method based on attribute control.
[0081] Exemplarily, the computer program may be divided into one or more modules, and the one or more modules are stored in the memory 11 and executed by the processor 12 to complete this application. The one or more modules may be a series of computer program instruction segments capable of completing specific functions, and these instruction segments are used to describe the execution process of the computer program in the electronic device 1. For example, the computer program may be divided into an image acquisition module 110, a feature extraction module 120, a feature decoupling module 130, a target extraction module 140, an intervention module 150, and a reconstruction module 160.
[0082] The integrated unit implemented in the form of software functional modules described above can be stored in a computer-readable storage medium, which can be non-volatile or volatile. The above software functional modules are stored in a storage medium and include several instructions for causing a computer device (which can be a personal computer, a computer device, or a network device, etc.) or a processor to execute part of the functions of the image generation method based on attribute control described in various embodiments of the present application.
[0083] In summary, a method, system, device, and medium for image generation based on attribute control disclosed in the present invention have the ability to precisely control specific attributes (such as gender and age) during the generation process. Through steps of feature decoupling, target feature recognition, and target feature intervention, the present invention provides an innovative technical path, significantly improving the controllability and consistency of the generated content while maintaining the original performance of the model. This method provides an innovative technical path for the precise control of generative artificial intelligence systems and opens up a new research direction for improving the accuracy and reliability of the generated content. The present invention is applicable to diffusion generative models for unconditional and conditional generation, can process multiple attributes of the generated content simultaneously, and realizes precise adjustment of the generation distribution while maintaining the image generation quality. The systematic feature regulation method proposed in the present invention can directly and precisely adjust specific features in the model that affect the generation result, achieving fine-grained control of the generated content. The present invention is an innovative method that can deeply analyze and precisely control the generation process of the diffusion model. By revealing the internal generation decision mechanism of the model and establishing precise regulation techniques at the feature level, it can not only improve the controllability of the generated content but also open up new theoretical and practical paths for the research on the interpretability and reliability of generative artificial intelligence technologies. Therefore, the present invention effectively overcomes various shortcomings in the prior art and has high industrial utilization value.
[0084] The above embodiments are only illustrative of the principles and effects of the present invention and are not intended to limit the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes made by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed by the present invention should still be covered by the claims of the present invention.
Claims
1. An image generation method based on attribute control, characterized in that The method includes: Obtaining an image to be adjusted and its target attribute; wherein, the target attribute is used to characterize the target that the image needs to be adjusted; Inputting the image into a diffusion model, and extracting a semantic feature group of the image through the diffusion model; Performing feature decoupling on the semantic feature group to obtain a sparse semantic feature group; Screening at least one target semantic feature from the sparse semantic feature group according to the gradient attribution method; Intervening on the screened target semantic feature based on the target attribute to obtain an intervened target semantic feature; Reconstructing the semantic feature group according to the intervened target semantic feature, inputting the reconstructed semantic feature group into the diffusion model, and performing iterative denoising on the image based on the reconstructed semantic feature group to generate an intervened target image.
2. The image generation method based on attribute control according to claim 1, wherein The sparse semantic feature group is obtained through the encoder module of an autoencoder. The performing feature decoupling on the semantic feature group to obtain a sparse semantic feature group includes: Inputting the semantic feature into the encoder module of the autoencoder, mapping the semantic feature to a high-dimensional semantic space, and generating a plurality of initial sparse semantic features; Calculating the activation values of each initial sparse semantic feature; Sorting all the activation values to form an activation value sequence; Selecting a preset first number of initial sparse semantic features from the activation value sequence in the descending order of the activation values as the final sparse semantic features to form a sparse semantic feature group.
3. The image generation method based on attribute control according to claim 1, wherein The screening at least one target semantic feature from the sparse semantic feature group according to the gradient attribution method includes: Calculating the contribution degree of each sparse semantic feature to the target attribute based on the integrated gradient method; Sorting the contribution degrees of all the sparse semantic features to form a contribution degree sequence; Selecting a preset second number of sparse semantic features from the contribution degree sequence in the descending order of the contribution degrees as the target semantic features.
4. The method for generating an image based on attribute control according to claim 1, wherein The intervening on the screened target semantic feature based on the target attribute to obtain an intervened target semantic feature includes: Judging the relationship between the target attribute and the screened target semantic feature: If the target attribute and the screened target semantic feature are positively correlated, performing amplification processing on the screened target semantic feature to obtain an intervened target semantic feature; If the target attribute and the screened target semantic feature are negatively correlated, performing suppression processing on the screened target semantic feature to obtain an intervened target semantic feature.
5. The method for generating an image based on attribute control according to claim 1, wherein The intervened target semantic feature reconstructs the semantic feature group through the decoder module of the autoencoder. The reconstructing the semantic feature group according to the intervened target semantic feature, inputting the reconstructed semantic feature group into the diffusion model, and performing iterative denoising on the image based on the reconstructed semantic feature group to generate an intervened target image includes: Inputting the intervened target semantic feature into the decoder module of the autoencoder, reconstructing the intervened target semantic feature back to the dimension of the original semantic feature group to obtain a reconstructed semantic feature group; Input the reconstructed semantic feature group into the diffusion model, and iteratively denoise the image based on the reconstructed semantic feature group to generate the target image after intervention; wherein, the position where the reconstructed semantic feature group is input into the diffusion model is the same as the position where the original semantic feature group is extracted from the diffusion model.
6. The method for generating an image based on attribute control according to claim 2 or 5, characterized in that The encoder is a k-sparse autoencoder.
7. The method for generating an image based on attribute control according to claim 1, wherein The diffusion model is an unconditional diffusion model or a conditional diffusion model.
8. An image generation system based on attribute control, characterized in that, The system includes: An image acquisition module, configured to acquire an image to be adjusted and its target attribute; wherein, the target attribute is used to characterize the target for which the image needs to be adjusted. A feature extraction module, configured to input the image into the diffusion model and extract the semantic feature group of the image through the diffusion model. A feature decoupling module, configured to decouple the semantic feature group to obtain a sparse semantic feature group. A target extraction module, configured to screen at least one target semantic feature from the sparse semantic feature group according to the gradient attribution method. An intervention module, configured to intervene in the screened target semantic feature based on the target attribute to obtain the target semantic feature after intervention. A reconstruction module, configured to reconstruct the semantic feature group according to the target semantic feature after intervention, input the reconstructed semantic feature group into the diffusion model, and iteratively denoise the image based on the reconstructed semantic feature group to generate the target image after intervention.
9. An electronic device, characterized in that, The electronic device includes: One or more processors; A storage device, configured to store one or more programs, and when the one or more programs are executed by the one or more processors, enable the electronic device to implement the attribute control-based image generation method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and when the computer program is executed by the processor of the computer, enable the computer to execute the attribute control-based image generation method according to any one of claims 1 to 7.