Image generation and understanding unification method and system based on joint diffusion modeling
By combining diffusion modeling and an improved model architecture, the problem of unifying image generation and understanding tasks was solved, achieving efficient image generation and understanding tasks and improving the model's cross-domain alignment and modeling accuracy.
Patent Information
- Application Number
- CN202511205486.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-27
- Publication Date
- 2025-11-14
AI Technical Summary
Existing technologies have failed to effectively unify image generation and understanding tasks. Traditional methods cannot simultaneously complete generation and understanding in a single model, and the discriminative feature modeling capabilities of pre-trained diffusion models are limited, making them unable to effectively handle image segmentation, detection, and classification tasks.
We adopt a joint diffusion modeling approach, construct an initial joint diffusion model using the Flux architecture, and optimize and train it using an improved DINOv2 self-supervised visual encoder, Segmenter model, and DETR model. We introduce a random role assignment mechanism and a masked full attention mechanism to achieve cross-domain feature interaction between the image domain, label domain, and feature domain.
It improves the efficiency and accuracy of image generation and understanding tasks, enhances the aggregation of classification feature clusters and the details of segmentation boundaries, and strengthens the representation of small target features, outperforming existing unified models.
Smart Images

Figure CN120953442A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image generation and understanding technology, and in particular to a unified method and system for image generation and understanding based on joint diffusion modeling. Background Technology
[0002] Visual generation and understanding are typically viewed as two entirely different tasks, handled with different approaches—generation tasks usually rely on generative models, such as diffusion models, while understanding tasks are often transformed into classification or regression problems, solved by discriminative models. Simultaneously achieving generation and understanding within a single model is a valuable and meaningful task.
[0003] In the generation task, controllable generation methods generate corresponding images using interpretable labels such as object detection boxes, semantic segmentation maps, and human pose maps as control conditions, but lack the ability to derive labels from images. In the understanding task, traditional discriminative models undergo supervised training for individual tasks, failing to achieve uniformity. Methods that mine semantic information from pre-trained diffusion models do not inherently jointly model generation and understanding tasks, and the discriminative feature modeling capabilities in pre-trained diffusion models are limited, resulting in performance inferior to traditional discriminative models. Methods that transform understanding tasks into image label generation tasks are mostly used for dense prediction, but are unsuitable for understanding tasks such as segmentation, detection, and classification. For example, image segmentation requires outputting label maps (the semantic category of each pixel). Representing label maps using RGB images introduces uncertainty in color mapping, potentially leading to concentrated mapping to a portion of the RGB space, causing learning difficulties. Object detection requires explicit bounding box coordinates and category labels; these structures are sparse and non-image format, unable to be represented by pixel images. Image classification requires category labels, which cannot be modeled using RGB. Regarding joint modeling, recent methods have attempted to model the conditional distributions in both the image-to-label and label-to-image directions simultaneously, but they haven't modeled the joint distribution of these two directions. While some unified methods share a single set of model parameters across all tasks, achieving model-level uniformity, they still treat different tasks separately, failing to unify all image domains with the output domains of each task. Therefore, existing methods do not jointly model the generation and understanding tasks. Summary of the Invention
[0004] The purpose of this invention is to provide a unified method and system for image generation and understanding based on joint diffusion modeling, so as to improve the above-mentioned technical problems.
[0005] To achieve the above-mentioned objectives, the embodiments of the present invention provide the following technical solutions:
[0006] A unified method for image generation and understanding based on joint diffusion modeling, comprising:
[0007] The image domain, label domain, or feature domain of the task to be generated / understood is determined; an initial joint diffusion model is constructed using the Flux architecture; the task to be generated / understood includes a generation task, an understanding task, and a joint distribution task; the image domain includes the original image; the label domain includes depth map, normal map, albedo, edge map, and line drawing; the feature domain includes image segmentation features, object detection features, and image classification features.
[0008] Based on image classification model, image segmentation model and object detection model, the learning objective is determined, and the initial joint diffusion model is optimized and trained by combining encoder and random role assignment mechanism to obtain joint diffusion model;
[0009] Based on the task to be generated / understood, a joint diffusion model is used to output the task results, thus achieving the unification of image generation and understanding;
[0010] The image classification model uses an improved DINOv2 self-supervised visual encoder; the image segmentation model uses an improved Segmenter model; and the object detection model uses an improved DETR model.
[0011] A unified system for image generation and understanding based on joint diffusion modeling includes:
[0012] The data acquisition module is used to determine the image domain, label domain, or feature domain of the task to be generated / understood;
[0013] The diffusion model construction training module is used to determine the learning objective based on the image classification model, image segmentation model and object detection model, and to optimize and train the initial joint diffusion model by combining the encoder and random role assignment mechanism to obtain the joint diffusion model;
[0014] The generation / understanding task completion module is used to output task results based on the task to be generated / understood through a joint diffusion model, thereby achieving the unification of image generation and understanding.
[0015] The beneficial effects of this invention are as follows:
[0016] This invention unifies image generation and understanding tasks through joint diffusion modeling, eliminating the need to design separate models for generation and understanding tasks and improving efficiency. Improved DINOv2, Segmenter, and DETR models enhance the aggregation of classification feature clusters, segmentation boundary details, and the representation of small target features in detection. Random role assignment and masked full attention mechanisms flexibly handle multi-domain information, while domain-invariant positional encoding assists cross-domain alignment, improving modeling accuracy. Optimized training enables the model to simultaneously support joint generation, controlled generation, and image perception tasks, outperforming existing unified models and even surpassing proprietary models in tasks such as edge detection. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a flowchart of the method in Embodiment 1 of the present invention;
[0019] Figure 2 This is a structural diagram of the initial joint diffusion model in Embodiment 1 of the present invention;
[0020] Figure 3 This is a structural diagram of the improved DINOv2 self-supervised visual encoder in Embodiment 1 of the present invention;
[0021] Figure 4 This is a structural diagram of the improved Segmenter model in Embodiment 1 of the present invention;
[0022] Figure 5 This is a structural diagram of the improved DETR model in Embodiment 1 of the present invention;
[0023] Figure 6 This is a schematic diagram of the overall joint diffusion model in Embodiment 1 of the present invention;
[0024] Figure 7 This is a structural diagram of the random role allocation mechanism in Embodiment 1 of the present invention;
[0025] Figure 8 This is a system structure diagram in Embodiment 2 of the present invention;
[0026] Figure 9 This is a diagram illustrating the joint generation task in Embodiment 3 of the present invention;
[0027] Figure 10 This is a comparison diagram of the generation tasks of the method in Embodiment 4 of the present invention with those of other models;
[0028] Figure 11 This is a comparison diagram of the results of the domain-invariant position encoding in Embodiment 5 of the present invention. Detailed Implementation
[0029] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0030] Example 1:
[0031] Please see Figure 1 This embodiment provides a unified method for image generation and understanding based on joint diffusion modeling. Figure 1 The execution entity of the method shown can be a software and / or hardware device. The execution entity of this application can include, but is not limited to, at least one of the following: user equipment, network equipment, etc. User equipment can include, but is not limited to, computers, smartphones, personal digital assistants (PDAs), and the aforementioned electronic devices. Network equipment can include, but is not limited to, a single network server, a server group consisting of multiple network servers, or a cloud based on cloud computing consisting of a large number of computers or network servers. Cloud computing is a type of distributed computing, consisting of a super virtual computer composed of a group of loosely coupled computers. This embodiment does not impose any limitations on this.
[0032] A unified method for image generation and understanding based on joint diffusion modeling, comprising:
[0033] S1. Determine the image domain, label domain, or feature domain of the task to be generated / understood; construct an initial joint diffusion model using the Flux architecture; the task to be generated / understood includes a generation task, an understanding task, and a joint distribution task; the image domain includes the original image; the label domain includes depth map, normal map, albedo, edge map, and line drawing; the feature domain includes image segmentation features, object detection features, and image classification features;
[0034] Specifically, the joint generation task refers to the ability to simultaneously generate an image, all corresponding labels, and corresponding image features using a joint diffusion model. The generation task involves learning to generate an image given several labels and features using a joint diffusion model. The understanding task involves learning to predict multiple labels and image features simultaneously using a given image, i.e., generating both the label domain and the feature domain.
[0035] When the task to be generated / understood is a joint generation task, the input to the joint diffusion model is textual information, i.e., a description of the image to be generated. While generating the required image, the corresponding label and feature domains are extracted. When the task to be generated / understood is a generation task, the input to the joint diffusion model is multiple label domains and multiple feature domains, and the output is an image with multiple label domains determined. When the task to be generated / understood is a comprehension task, the input to the joint diffusion model is the image domain, and the corresponding output is the label and feature domains. If different tasks exist simultaneously, each task's corresponding domain is input to a different joint diffusion model, processed independently, and the required results are output separately.
[0036] In this embodiment, the Transformer model's backbone is based on the open-source FLUX model and improved for multi-domain joint modeling tasks. Unlike the original Flux model, which is mainly used for image generation tasks, this method extends it to include the image domain, label domain, and feature domain. To achieve multi-domain joint modeling, the Flux architecture is modified to handle token sequences from M domains and cross-domain feature interaction is performed through a unified masked full attention mechanism.
[0037] Therefore, as Figure 2 As shown, the initial joint diffusion model includes a Transformer model; the Transformer model includes a domain-invariant positional encoding layer and a Transformer module; the Transformer module includes multiple stacked Transformer layers;
[0038] The domain-invariant positional coding layer employs 2D sine and cosine coding. Specifically, for the same spatial location in different domains, the same 2D sine and cosine coding is added to the input token, thereby providing explicit positional information to the joint diffusion model.
[0039] Each Transformer layer employs a masked full attention mechanism, which improves modeling accuracy. To enable the joint diffusion model to ignore invalid domains with role type X, a masking mechanism is introduced on top of the standard full attention mechanism, adding log(m) to the tokens marked as invalid. j A penalty is applied, causing the attention weights to become zero after the softmax function. Therefore, the formula corresponding to the masked full attention mechanism is:
[0040]
[0041] Where softmax(·) denotes the maximum value function, O i Q represents the attention score corresponding to the i-th Transformer layer. i Kj V j Let d represent the i-th query vector, the j-th key vector, and the j-th value vector, respectively, and let d represent Q. i Dimensions Let m denote the transpose of the key vectors, ∑(·) denote the summation function, log(·) denote the logarithmic function to the natural constant, and m j This represents the j-th mask indicator, with a value range of [0, 1]. M and N represent the number of queries and the number of keys or values, respectively. When m j When the value is 0, the attention score region of the j-th token is negative infinity, which can completely mask the invalid region, and the corresponding output has no effect.
[0042] S2. Based on the image classification model, image segmentation model, and object detection model, the learning objective is determined, and the initial joint diffusion model is optimized and trained by combining the encoder and random role assignment mechanism to obtain the joint diffusion model. Among them, the image classification model adopts the improved DINOv2 self-supervised visual encoder; the image segmentation model adopts the improved Segmenter model; and the object detection model adopts the improved DETR model.
[0043] like Figure 3 As shown, the improved DINOv2 self-supervised visual encoder includes a cascaded geometry-aware patch embedding layer, a relative position Transformer layer, and a channel selection compression module; the relative position Transformer layer includes multiple stacked first Transformers, with a cluster-aware module inserted between every two first Transformers, and the rest of the structure is the same as the existing Transformer; the channel selection compression module includes a global average pooling layer, two MLP layers, and a weighted activation layer.
[0044] like Figure 4 As shown, the improved Segmenter model includes a cascaded MS-CAP module, a channel attention layer, and a segmentation encoder. The segmentation encoder includes a cascaded one-dimensional flattening layer, a position embedding layer, and multiple stacked second Transformers. An SREB module (semantic residual enhancement module) is inserted between every two second Transformers to introduce cross-layer feature connections and fuse the output of the i-th second Transformer with the output of the i+k-th second Transformer. The MS-CAP module includes a cascaded parallel convolutional layer, a channel attention layer, and a feature fusion layer.
[0045] like Figure 5As shown, the improved DETR model includes a cascaded dual-branch CNN backbone network, channel attention layers, a Transformer-based encoder, a Transformer-based decoder, and a prediction head. The dual-branch CNN backbone network includes a parallel global semantic CNN network and a target edge CNN network. The target edge CNN network includes a cascaded CNN network, local perceptual convolutional layers, and a boundary response enhancement module. The Transformer-based encoder includes multiple block layers, with a spatial guidance module inserted between every two block layers. The global semantic CNN network is the same as the CNN network.
[0046] The local perceptual convolutional layers include convolutional layers with 3×3 kernels and convolutional layers with 5×5 kernels. In the boundary response enhancement module, a boundary attention map (sigmoid activation) is first generated through 1×1 convolution, and then element-wise multiplied with the original features to achieve feature enhancement of the boundary region.
[0047] like Figure 6 As shown, the optimization training of the initial joint diffusion model to obtain the joint diffusion model includes:
[0048] S2-1. Obtain the training image domain and training label domain; based on the learning objective, process the training image domain through image classification model, image segmentation model and object detection model to obtain the classification task domain, segmentation domain and detection domain, and use them as training feature domains.
[0049] Since generative tasks typically rely on generative models, such as diffusion models, while understanding tasks are often transformed into classification or regression problems and solved by discriminative models, image classification models, image segmentation models, and object detection models are selected to obtain corresponding training feature domains for training the initial joint diffusion model. The feature domains generated by the initial joint diffusion model are more suitable for discriminative tasks (classification, detection, and segmentation). Subsequently, the generated feature domains are input into the decoder to achieve the discriminative task, thus achieving the purpose of image classification, image segmentation, and object detection.
[0050] When the learning objective is image classification, the original structure of the traditional DINOv2 self-supervised visual encoder suffers from insufficient inter-class separability in the intermediate layers, and the high feature dimension makes it difficult to model and compress diffusion models, hindering effective learning of diffusion models. Therefore, a cluster perception module and a channel selection compression module are introduced. Thus, the image classification model processing procedure is as follows:
[0051] S2-101. Obtain the training image domain and input it into the geometry-aware Patch embedding layer. Divide the training image domain by rotational equivariant convolution to obtain a fixed-size classification Patch and convert it into classification tokens.
[0052] S2-102. Extract the classification space features of each category token using the first Transformer in the relative position Transformer layer;
[0053] S2-103. The classification space features are aggregated by the cluster perception module, and the classification cluster perception features are obtained by combining the attention weight and the aggregation vector.
[0054] It should be noted that the original class tokens are used to semantically aggregate the categorization cluster perception features, guiding similar categorization cluster perception features to form high-density semantic clusters. Specifically, a set of class tokens is reserved for each first Transformer, that is, one categorization cluster perception feature is selected as the class token, and the remaining categorization cluster perception features are semantically aggregated. Taking the first first Transformer as an example, cluster aggregation Attention is performed on its input categorization cluster perception features, and the attention weight A between each categorization cluster perception feature and its corresponding class token is calculated. i The corresponding formula is:
[0055]
[0056] in, x represents the perceptual feature of the i-th classification cluster. i The corresponding query vector, This represents the transpose of the key vector corresponding to the class token.
[0057] Based on each attention weight, the corresponding aggregation vector is calculated and used as a residual to guide the clustering of class centers for each category cluster perceptual feature item, thus obtaining the category cluster perceptual features. The corresponding formula is:
[0058] x i '=x i +A i ·V cls ;
[0059] Where, x i 'Represents the perceptual features of classification clusters, V cls A represents the value vector corresponding to the class token. i ·V cls This represents an aggregate vector.
[0060] S2-104. Repeat the same operations for the first Transformer and the cluster awareness module until the output of the last first Transformer is obtained;
[0061] S2-105. The output of the last first Transformer is pooled and importance weights are calculated through the channel selection compression module. Based on the importance weights, the dimensionality of the pooled classification cluster perception features is reduced to obtain the classification task domain, which is the training feature domain.
[0062] The channel selection compression module pools the output of the last first Transformer and calculates its importance weights; based on the importance weights, the dimensionality of the pooled output of the last first Transformer is reduced.
[0063] Specifically, a channel selection compression module is used to compress the invalid redundant dimensions in the output features of the last high-dimensional first Transformer, suppressing invalid channel activations and compressing feature dimensions, thereby improving the learning efficiency of the diffusion model. The global average S over the channel dimensions is calculated using a global average pooling layer, with the corresponding formula being:
[0064] S=GAP(x'∈R 1 ×D;
[0065] Where R represents a constant, GAP(·) represents the pooling operation, which averages all elements of each feature map to obtain a single value, and D represents the feature dimension of the output of the last first Transformer.
[0066] The importance weight W is calculated using a cascaded two-layer MLP, and the corresponding formula is:
[0067] W=σ(MLP(S))∈R 1×D ;
[0068] Where σ(·) represents the activation function and MLP(·) represents the multilayer perceptron operation.
[0069] By selectively activating important channels based on importance weights and discarding channels with importance weights W, we obtain the dimensionality-reduced feature X, with the corresponding formula as follows:
[0070] X = x'·W.
[0071] When the learning objective is image segmentation, the image segmentation model processes the following steps:
[0072] S2-111. Obtain the training image domain and divide it using the MS-CAP module (multi-scale channel-aware embedding module), and extract patch tokens under different receptive fields in parallel; calculate the channel attention weights of different patch tokens through the channel attention layer; process each patch token based on the channel attention weights to obtain segmentation enhancement multi-scale features.
[0073] Specifically, the traditional Segmenter model has shortcomings in boundary detail modeling and the completeness of semantic information in intermediate layer features, which is not conducive to the effective learning of the diffusion model. Therefore, the MS-CAP module was chosen to replace the original flatten and project modules. In the MS-CAP module, the training image domain is processed through multiple parallel convolutional layers, and patch tokens under different receptive fields are extracted using convolutional kernels of different sizes. Small convolutional kernels (such as 3×3) have small receptive fields, focusing on capturing local details of the image and accurately extracting subtle textures; large convolutional kernels (such as 7×7) have large receptive fields, which can obtain information from a wider area and focus on the overall structure and semantics.
[0074] Each patch token is input into the channel attention layer, and the feature channels of each branch are analyzed. The feature channels that are sensitive to boundaries and textures are adaptively selected. By calculating the importance weight of each channel, the response of channels related to texture and semantic boundaries is enhanced, while irrelevant channels are suppressed, thereby highlighting the key feature information of the image. This helps the joint diffusion model to focus more on boundary and texture parts.
[0075] Finally, the multi-scale features enhanced by channel attention are concatenated, integrating the enhanced features from different scales to form a feature set containing rich multi-scale information. Next, the concatenated features are uniformly projected to the patch token embedding dimension required by the encoder in the first Transformer, making the features adaptable to the processing requirements of subsequent Transformers. Then, positional encoding is added to record the spatial location information of each feature in the original image, preserving the spatial order of the features and avoiding the loss of spatial structure information during processing.
[0076] S2-112. Input the segmentation enhancement multi-scale features into the segmentation encoder, extract the semantic residual fusion features and use them as the segmentation task domain to obtain the training feature domain.
[0077] In the segmentation encoder, segmentation enhancement multi-scale features are divided to obtain segmentation patches, which are then input into a one-dimensional flattening layer to flatten each segmentation patch into a one-dimensional vector. A positional embedding layer embeds each one-dimensional vector to obtain an embedding sequence. The first second Transformer performs feature encoding and context aggregation on the embedding sequence to obtain the first semantic residual fusion feature. The SREB module fuses the first semantic residual fusion feature with the outputs of the second Transformers spaced k times apart, and uses this as the input to the next second Transformer. This process of second Transformer and SREB module operations is repeated until the output of the last second Transformer is obtained, yielding the semantic residual fusion feature. In this implementation, K takes the value of 2 or 3.
[0078] It needs to be explained that the SREB module is used to fuse shallow and deep features; at the same time, it introduces cross-layer feature connections to improve the transmission of semantic information between layers, reduce the loss of intermediate semantic information, and enable the extracted features to capture local edges, semantic subjects, and global structural information respectively, thus making the supervision target for training the diffusion model a complete semantic feature set. The output F of the second Transformer in the i-th layer... i The output F of the second Transformer in the (i+k)th layer i+k The formula for fusion is:
[0079] F' i+k =F i+k +gate(F i );
[0080] Among them, F' i+k The fused semantic residual features are represented by `gate(·)`, and `gate(·)` represents a convolution operation with a 1×1 kernel and the Sigmoid activation function.
[0081] During the inference phase, a joint diffusion model is used to generate semantic segmentation features. These generated features are then input into a two-layer lightweight MaskTransformer decoder, which outputs a semantic mask probability map for each category. After softmax and category argmax operations, the final segmentation result map is obtained.
[0082] When the learning objective is object detection, the processing procedure of the object detection model is as follows:
[0083] S2-121. Obtain the training image domain; extract global semantic information of the training image domain through a global semantic CNN network; extract initial edge features of the training image domain through a CNN network; extract detailed textures of the initial edge features using local perceptual convolutional layers to obtain edge convolutional features; and dynamically extract edge enhancement features of the edge convolutional features using a boundary response enhancement module, combined with edge detection operators and boundary attention gating.
[0084] S2-122. By fusing edge enhancement features and global semantic information through the channel attention layer, global-edge fusion features are obtained.
[0085] S2-123. The global edge fusion features are encoded through the first block layer to obtain global-fusion encoding information; the edge response map of the training image domain is extracted using the spatial guidance module, and the edge response map and global-fusion encoding information are fused together by combining the mask attention mechanism and modulation to obtain global-fusion guidance information.
[0086] It needs to be explained that in the spatial guidance module, the global-fusion coding information is modulated based on the edge response map extracted from the training image domain to obtain dynamically weighted attention features, thereby enhancing the joint diffusion model's response capability to small targets and edge regions. Specifically, a low-resolution edge response map (E) is extracted from the input image and used as a location information guidance signal. The output features of the current block layer are fused with the guidance map E to obtain a mask attention value M. This mask attention value M is then used to modulate the output features F of the current block layer, as shown in the following formula:
[0087] F' = F⊙(1+M);
[0088] Obtain global-fusion guidance information F'.
[0089] The global-fusion guidance information is then fed into the next block layer. Within the block layer, each encoder can dynamically adjust the feature focus position, and the guidance signal comes from the original input image (training image domain) or a shallow branch, without the need for additional supervision.
[0090] S2-124. Repeat the operations of the block layer and the spatial guidance module until the output of the last block layer is obtained and used as the detection domain, that is, the training feature domain is obtained.
[0091] During the training phase, the output of the last block layer serves as the object detection domain, acting as a supervisory target for training the joint diffusion model. During the inference phase, the diffusion model can be used to generate detection features.
[0092] S2-2. Roles are assigned to the training feature domain, training image domain, and training label domain respectively through the encoder and random role assignment mechanism to determine the corresponding role types. Based on each role type, the domain-encoding information corresponding to the training feature domain, training image domain, and training label domain is processed to obtain the corresponding joint diffusion model input data.
[0093] Since existing methods cannot unify the generation and understanding tasks, a random role assignment mechanism is proposed. The process involves assigning roles to the training feature domain, training image domain, and training label domain through both the encoder and the random role assignment mechanism to determine the corresponding role types, including:
[0094] S2-2-1. Based on the training generation / understanding task, the training image domain, training label domain, and training feature domain are input into the encoder to obtain training domain-encoding information; at this time, the task to be generated / understood is the generation task and the understanding task. For example, when the task to be generated / understood is the generation task, the input to the encoder is the corresponding label domain and feature domain.
[0095] S2-2-2. The training domain-encoding information is processed through a random role assignment mechanism to assign roles, determine the role type and assignment method; if the role type is G, the assignment method is to add noise; if the role type is C, the assignment method is to keep it as is; if the role type is X, the assignment method is to set the data to zero.
[0096] Specifically, such as Figure 7 As shown, if the role is G (generated), the allocation method is to add noise, that is, to add noise using the standard training method of Rectified Flow, and predict the velocity field; if the role type is C (conditional), the allocation method is to keep it as is; if the role type is X (ignore), the allocation method is to zero out the data, that is, to set the data to zero and use a masking mechanism to shield its influence. Figure 2 In this context, the switcher represents the random role allocation mechanism.
[0097] S2-2-3. The training domain-encoding allocation information is processed by the allocation processing method to obtain the input data of the joint diffusion model.
[0098] Specifically, the formula corresponding to the random character allocation mechanism is:
[0099]
[0100] in, ρ represents the output data of the initial joint diffusion model obtained by processing domain-encoded information through a random role assignment mechanism. k This represents the assigned role, k represents the index, and τ represents the time step. Represents the domain-encoding allocation information, ξ k Indicates the noise term;
[0101] Through a random role assignment mechanism, different tasks are assigned. For example, if the image domain is C and other domains are G, it is understanding; if the image domain is G and a certain label domain is C, it is controlled generation; if all domains are G, it is joint generation. During training, multiple tasks are trained in a single model through the role assignment mechanism, thereby unifying generation and understanding.
[0102] S2-3. Obtain training text information; according to the training generation / understanding task, input the training text information and the input data of each joint diffusion model into the initial joint diffusion model to generate the corresponding training task results;
[0103] It should be noted that before the initial joint diffusion model, the input data needs to be partitioned to obtain multiple different training tokens. The processing of the initial joint diffusion model can be explained in three cases, depending on the training generation / understanding task.
[0104] When training the generation / understanding task is a joint generation task, the training token corresponding to the training text information is input into the joint diffusion model. The role of all domains is G, and the token is randomly initialized noise. Finally, a graph and the corresponding label domain and feature domain of the graph are generated (generated simultaneously in one inference, without any order).
[0105] When training to generate / understand is a generation task, the training tokens corresponding to the training label domain and training feature domain are input into the domain-invariant positional encoding layer. Positional information is introduced into the training tokens to obtain the corresponding training generation position tokens. Stacked Transformer layers capture the relationships between the training generation position tokens, generating shallow and deep features at different levels and fusing them to finally obtain the corresponding training image.
[0106] When training the generation / understanding task to be an understanding task, the process is the same as when training the generation / understanding task to be a generation task, except that the input data is the training images. It should be noted that in all three tasks, the conditional domain remains unchanged from the original token, while the domain to be generated is initialized with random noise.
[0107] S2-4. Based on the results of each training task
[0108]
[0109] Calculate the diffusion loss function Loss; where, Let ξ denote the expectation function, U(0,1) denote a random variable sampled from a uniform distribution [0,1], and ξ denote the expectation function. 0:K ξk Let N(0,1) represent noise terms with indices from 0 to K and with index k, respectively. Let N(0,1) represent random variables sampled from a standard normal distribution with mean 0 and variance 1. Let represent the target values in the training set from index 0 to K and the target value of the k-th domain, respectively; D represents the training set consisting of the training image domain, the training label domain, and the training feature domain; ∑(·) represents the summation function; ρ k Let G represent the role, ||·|| represent the norm, θ represent the parameters of the initial joint diffusion model, and τ represent the time step. v represents the target value corresponding to the fields with indices 0, 1, and K at time step τ, respectively. θ (·) indicates a speed prediction model;
[0110] The velocity prediction model, as described in this invention, predicts the "velocity of change" of data from the current state to the true state within a diffusion framework, gradually reconstructing the true data from noise. The loss function optimizes the model by minimizing the difference between the model's predicted velocity and the true velocity (noise - true value), thereby improving the generation quality.
[0111] S2-5. Iterate the initial joint diffusion model based on the diffusion loss function and adjust the weight parameters of the initial joint diffusion model to complete the optimization training of the initial joint diffusion model.
[0112] S3. Based on the task to be generated / understood, the joint diffusion model outputs the task results, thus achieving the unification of image generation and understanding.
[0113] Example 2:
[0114] like Figure 8 As shown, a unified system for image generation and understanding based on joint diffusion modeling includes:
[0115] The data acquisition module is used to determine the image domain, label domain, or feature domain of the task to be generated / understood;
[0116] The role assignment module is used to assign roles through an encoder and a random role assignment mechanism to obtain input data for joint diffusion modeling.
[0117] The diffusion model construction and training module is used to build an initial joint diffusion model in conjunction with the Flux architecture; by combining the image classification model, image segmentation model and object detection model, the learning objective is determined and the initial joint diffusion model is optimized and trained to obtain the joint diffusion model;
[0118] The generation / understanding task completion module inputs the generation / understanding task data into the joint diffusion model and outputs the task results, thus achieving the unification of image generation and understanding.
[0119] The diffusion model construction and training module includes:
[0120] The diffusion model construction submodule is used to build an initial joint diffusion model in conjunction with the Flux architecture;
[0121] The classification task domain acquisition submodule is used to acquire the training image domain and process it through the image classification model to obtain the classification task domain.
[0122] The domain acquisition submodule is used to acquire the training image domain and process it through the image segmentation model to obtain the segmentation task domain;
[0123] The detection domain acquisition submodule is used to acquire the training image domain and process it through the object detection model to obtain the detection domain.
[0124] The role assignment submodule is used to assign the classification task domain, segmentation task domain, and detection domain to the encoder and random role assignment mechanism respectively to obtain the corresponding role types; based on each role type, the domain-encoding information corresponding to the training feature domain, training image domain, and training label domain is processed to obtain the corresponding joint diffusion model input data.
[0125] The model optimization and training submodule is used to obtain the training label domain and training text information, and calculate the diffusion loss function based on the classification task domain, segmentation task domain, and detection domain; iterates the initial joint diffusion model based on the diffusion loss function and adjusts the weight parameters of the initial joint diffusion model.
[0126] It should be noted that the specific methods by which each module performs operations in the system described in the above embodiments have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0127] Example 3:
[0128] The joint generation task was processed using the method described in Example 1, and the results are as follows: Figure 9 As shown, given input text, the model can simultaneously generate images and their corresponding 7 labels. The generated images have a wide range of aspect ratios and are all around 1024 resolution.
[0129] Example 4:
[0130] The ControlNet, UniControl, EasyControl, OmniGen, PixWizard, and OneDiffusion models were selected as other models and compared with the joint diffusion model of this invention. As shown in Table 1, performance was compared for the generated label domain. Compared with other models, the joint diffusion model of this invention achieved better performance under all conditions. Performance metrics included LPIPS (Perceptual Image Similarity Points) and FID distance, where FID distance is a metric used to measure the distance between the generated image and the real image distribution.
[0131] Table 1
[0132]
[0133] The generation task was implemented using this method, the OmniGen model, the OneDiffusion model, and the PixWizard model, respectively, and the corresponding results are as follows: Figure 10 As shown, compared with the OmniGen model, OneDiffusion model, and PixWizard model, this method performs superiorly on the generation task, producing images that better match the input labels and features, far outperforming the OmniGen model, OneDiffusion model, and PixWizard model.
[0134] This method sequentially performs depth estimation, normal estimation, albedo estimation, and edge detection.
[0135] Table 2
[0136] method NYUv2 ScanNet DIODE OmniGen 9.2 10.1 30.6 PixWizard 7.0 7.9 25.4 OneDiffusion 8.9 9.7 25.2 This method 10.1 12.1 25.9 This method (integrated) 8.3 9.9 25.8
[0137] Table 3
[0138] method NYUv2 ScanNet iBims OmniGen 28.9 28.9 31.3 PixWizard 23.5 26.6 22.5 This method 21.1 24.3 20.1 This method (integration) 18.6 20.3 18.2
[0139] Table 4
[0140] method PSNR ScanNet OrdinalShading 15.6 0.37 Kocsisetal. 11.3 0.49 Careaga and Aksoy 15.7 0.36 RGB2X 20.6 0.18 This method 15.5 0.31 This method (integration) 16.5 0.33
[0141] Table 5
[0142] method ODS OIS HED 0.788 0.808 PiDiNet 0.807 0.823 OmniGen 0.767 0.781 PixWizard 0.605 0.633 OneDiffusion 0.682 0.691 This method 0.826 0.851
[0143] For depth estimation, the absolute mean relative error was tested on the NYUv2, ScanNet, and DIODE datasets using our method and its ensemble model, the OneDiffusion model, the OmniGen model, and the PixWizard model, respectively. As shown in Table 2, arranged in ascending order, our method and its ensemble model ranked fifth and second in absolute mean relative error on the NYUv2 dataset; fifth and third on the ScanNet dataset; and fourth and third on the DIODE dataset.
[0144] For normal vector estimation, the mean angle error was tested on the NYUv2 dataset, ScanNet dataset, and iBims dataset using our proposed method and its ensemble model, the OmniGen model, and the PixWizard model, respectively. As shown in Table 3, on the NYUv2 dataset, our proposed method and its ensemble model ranked third and fourth in mean angle error, respectively, in descending order. On the ScanNet dataset, the same ranking was achieved. On the iBims dataset, the same ranking was achieved. Compared to other models, our method exhibits more stable performance across various datasets.
[0145] For albedo estimation, PSNR and LPIPS were tested on the Hypersim dataset test set using our proposed method and its ensemble model, the Ordinal Shading model, the Kocsis et al. model, the Careaga and Aksoy model, and the RGB2X model, respectively. As shown in Table 4, in descending order, our proposed method and its ensemble model ranked third and second in PSNR, respectively; and in descending order, they ranked second and third in LPIPS, respectively. Therefore, our proposed method outperforms the Ordinal Shading model, the Kocsis et al. model, and the Careaga and Aksoy model.
[0146] For edge detection, the F-scores of our method, the HED model, the PiDiNet model, the OmniGen model, the PixWizard model, and the OneDiffusion model were tested on the BSDS500 dataset at the optimal dataset scale (ODS) and the optimal image scale (OIS). As shown in Table 5, our method achieved the highest F-scores at both the optimal dataset scale (ODS) and the optimal image scale (OIS), demonstrating that our method significantly outperforms the other models.
[0147] Example 5:
[0148] A comparison is made between the joint diffusion model of our method and the joint diffusion model that removes the domain-invariant position coding layer. For example... Figure 11 As shown, without domain-invariant positional coding, there is a significant misalignment between the image domain and the label domain. This problem is improved after adding positional coding, demonstrating the effectiveness of positional coding. Therefore, the proposed domain-invariant positional coding layer can achieve cross-domain spatial alignment in conjunction with the diffusion-assisted model.
[0149] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
[0150] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A unified method for image generation and understanding based on joint diffusion modeling, characterized in that, include: Determine the image domain, label domain, or feature domain for the task to be generated / understood; An initial joint diffusion model is constructed using the Flux architecture; the tasks to be generated / understood include generation, understanding, and joint distribution tasks; the image domain includes the original image; the label domain includes depth map, normal map, albedo, edge map, and line drawing; the feature domain includes image segmentation features, object detection features, and image classification features. Based on image classification model, image segmentation model and object detection model, the learning objective is determined, and the initial joint diffusion model is optimized and trained by combining encoder and random role assignment mechanism to obtain joint diffusion model; Based on the task to be generated / understood, a joint diffusion model is used to output the task results, thereby achieving the unification of image generation and understanding; The image classification model uses an improved DINOv2 self-supervised visual encoder; the image segmentation model uses an improved Segmenter model; and the object detection model uses an improved DETR model.
2. The unified method for image generation and understanding based on joint diffusion modeling according to claim 1, characterized in that, The initial joint diffusion model includes a Transformer model; the Transformer model includes a domain-invariant positional encoding layer and a Transformer module; the Transformer module includes multiple stacked Transformer layers; the domain-invariant positional encoding layer uses 2D sine and cosine encoding; each Transformer layer uses a masked full attention mechanism; the formula corresponding to the masked full attention mechanism is: Where softmax(·) denotes the maximum value function, O i Q represents the attention score corresponding to the i-th Transformer layer. i K j V j Let represent the query vector of the i-th token, the key vector of the j-th token, and the value vector of the j-th token, respectively, and d represent Q. i Dimensions Let m denote the transpose of the key vectors, ∑(·) denote the summation function, log(·) denote the logarithmic function to the natural constant, and m j The mask indicator represents the j-th token, and M and N represent the number of queries and the number of keys or values, respectively.
3. The unified method for image generation and understanding based on joint diffusion modeling according to claim 1, characterized in that, The improved DINOv2 self-supervised visual encoder includes a cascaded geometry-aware patch embedding layer, a relative position Transformer layer, and a channel selection compression module. The relative position Transformer layer consists of multiple stacked first Transformers, with a cluster-aware module inserted between every two first Transformers; The improved Segmenter model includes a cascaded MS-CAP module, a channel attention layer, and a segmentation encoder. The segmentation encoder includes a cascaded one-dimensional flattening layer, a position embedding layer, and multiple stacked second Transformers. An SREB module is inserted between every two second Transformers to introduce cross-layer feature connections and fuse the output of the i-th second Transformer with the output of the i+k-th second Transformer. The improved DETR model includes a cascaded dual-branch CNN backbone network, channel attention layers, a Transformer-based encoder, a Transformer-based decoder, and a prediction head. The dual-branch CNN backbone network includes a parallel global semantic CNN network and a target edge CNN network. The target edge CNN network includes a cascaded CNN network, local perceptual convolutional layers, and a boundary response enhancement module. The Transformer-based encoder includes multiple block layers, with a spatial guidance module inserted between every two block layers.
4. The unified image generation and understanding method based on joint diffusion modeling according to claim 3, characterized in that, The optimization training of the initial joint diffusion model to obtain the joint diffusion model includes: Obtain the training image domain and training label domain; based on the learning objective, process the training image domain through image classification model, image segmentation model and object detection model to obtain the classification task domain, segmentation domain and detection domain, all of which are used as training feature domains; Roles are assigned to the training feature domain, training image domain, and training label domain respectively through encoder and random role assignment mechanism to determine the corresponding role type; based on each role type, the domain-encoding information corresponding to the training feature domain, training image domain, and training label domain is processed to obtain the corresponding joint diffusion model input data; Obtain training text information; according to the training generation / understanding task, input the training text information and the input data of each joint diffusion model into the initial joint diffusion model to generate the corresponding training task results; Based on the results of each training task, combined with the formula: Calculate the diffusion loss function Loss; where, Let ξ denote the expectation function, U(0,1) denote a random variable sampled from a uniform distribution [0,1], and ξ denote the expectation function. 0:K ξ k Let N(0,1) represent noise terms with indices from 0 to K and with index k, respectively. Let N(0,1) represent random variables sampled from a standard normal distribution with mean 0 and variance 1. Let represent the target values in the training set from index 0 to K and the target value of the k-th domain, respectively; D represents the training set consisting of the training image domain, the training label domain, and the training feature domain; ∑(·) represents the summation function; ρ k Let G represent the role, ||·|| represent the norm, θ represent the parameters of the initial joint diffusion model, and τ represent the time step. v represents the target value corresponding to the fields with indices 0, 1, and K at time step τ, respectively. θ (·) indicates a speed prediction model; The initial joint diffusion model is iterated based on the diffusion loss function, and the weight parameters of the initial joint diffusion model are adjusted to complete the optimization training of the initial joint diffusion model.
5. The unified method for image generation and understanding based on joint diffusion modeling according to claim 4, characterized in that, The process of assigning roles to the training feature domain, training image domain, and training label domain through an encoder and a random role assignment mechanism, respectively, to determine the corresponding role types, includes: Based on the training generation / understanding task, the training image domain, training label domain, and training feature domain are input into the encoder to obtain training domain-encoding information; A random role assignment mechanism is used to process the training domain-encoding information for role assignment, determining the role type and assignment method. If the role type is G, the assignment method is to add noise; if the role type is C, the assignment method is to keep it as is; if the role type is X, the assignment method is to set the data to zero. The training domain-encoding allocation information is processed using an allocation processing method to obtain the input data for the joint diffusion model.
6. The unified method for image generation and understanding based on joint diffusion modeling according to claim 4, characterized in that, When the learning objective is image classification, the image classification model processes the following steps: The training image domain is obtained and input into the geometry-aware Patch embedding layer. The training image domain is divided by rotational equivariant convolution to obtain a fixed-size classification Patch and convert it into classification tokens. The classification space features of each category token are extracted using the first Transformer in the relative position Transformer layer; The cluster perception module performs cluster aggregation on the classification space features, and combines the attention weights and aggregation vectors to obtain the classification cluster perception features; Repeat the same operations for the first Transformer and the cluster awareness module until the output of the last first Transformer is obtained; The channel selection compression module pools the output of the last first Transformer and calculates its importance weights. Based on importance weights, the dimensionality of the pooled classification cluster perception features is reduced to obtain the classification task domain, which is the training feature domain.
7. The unified method for image generation and understanding based on joint diffusion modeling according to claim 4, characterized in that, When the learning objective is image segmentation, the image segmentation model processes the following steps: The training image domain is acquired and divided using the MS-CAP module, and patch tokens under different receptive fields are extracted in parallel. Channel attention weights for different patch tokens are calculated using a channel attention layer. Each patch token is processed based on its channel attention weights to obtain segmentation enhancement multi-scale features. The segmentation enhancement multi-scale features are input into the segmentation encoder, and the semantic residual fusion features are extracted and used as the segmentation task domain, thus obtaining the training feature domain. In the segmentation encoder, segmentation enhancement multi-scale features are divided to obtain segmentation patches, which are then input into a one-dimensional flattening layer to flatten each segmentation patch into a one-dimensional vector. Each one-dimensional vector is then embedded through a position embedding layer to obtain an embedding sequence. The first and second Transformers are used to encode features and aggregate context in the embedding sequence to obtain the first semantic residual fusion feature. The first semantic residual fusion feature is fused with the output of the second Transformer k times apart using the SREB module, and used as the input of the next second Transformer. The operation of the second Transformer and the SREB module is repeated until the output of the last second Transformer is obtained, thus obtaining the semantic residual fusion feature.
8. The unified method for image generation and understanding based on joint diffusion modeling according to claim 4, characterized in that, When the learning objective is object detection, the processing procedure of the object detection model is as follows: Obtain the training image domain; extract global semantic information of the training image domain through a global semantic CNN network; extract initial edge features of the training image domain through a CNN network; extract detailed textures of the initial edge features using local perceptual convolutional layers to obtain edge convolutional features; and dynamically extract edge enhancement features of the edge convolutional features using a boundary response enhancement module, combined with edge detection operators and boundary attention gating. By fusing edge enhancement features and global semantic information through a channel attention layer, global-edge fusion features are obtained. The global edge fusion features are encoded in the first block layer to obtain global-fusion encoding information; the edge response map of the training image domain is extracted using the spatial guidance module, and modulated by the mask attention mechanism to fuse the edge response map and global-fusion encoding information to obtain global-fusion guidance information. Repeat the operations of the block layer and the spatial guidance module until the output of the last block layer is obtained and used as the detection domain, that is, the training feature domain is obtained.
9. A unified system for image generation and understanding based on joint diffusion modeling, used to implement the unified method for image generation and understanding based on joint diffusion modeling as described in any one of claims 1 to 8, characterized in that, include: The data acquisition module is used to determine the image domain, label domain, or feature domain of the task to be generated / understood; The diffusion model construction training module is used to determine the learning objective based on the image classification model, image segmentation model and object detection model, and to optimize and train the initial joint diffusion model by combining the encoder and random role assignment mechanism to obtain the joint diffusion model; The generation / understanding task completion module is used to output task results based on the task to be generated / understood through a joint diffusion model, thereby achieving the unification of image generation and understanding.
10. A unified image generation and understanding system based on joint diffusion modeling according to claim 9, characterized in that, The diffusion model construction and training module includes: The diffusion model construction submodule is used to build an initial joint diffusion model in conjunction with the Flux architecture; The classification task domain acquisition submodule is used to acquire the training image domain and process it through the image classification model to obtain the classification task domain. The domain acquisition submodule is used to acquire the training image domain and process it through the image segmentation model to obtain the segmentation task domain; The detection domain acquisition submodule is used to acquire the training image domain and process it through the object detection model to obtain the detection domain. The role assignment submodule is used to assign the classification task domain, segmentation task domain, and detection domain to the encoder and random role assignment mechanism respectively to obtain the corresponding role types; based on each role type, the domain-encoding information corresponding to the training feature domain, training image domain, and training label domain is processed to obtain the corresponding joint diffusion model input data. The model optimization training submodule is used to obtain the training label domain and training text information; calculate the diffusion loss function according to the training generation / understanding task; and iterate the initial joint diffusion model based on the diffusion loss function and adjust the weight parameters of the initial joint diffusion model.
Citation Information
Cited By
Individualized health management system and method for child special case
CN121388001A