Method for training a visual language model using a diffusion model supervision

By combining expert cross-attention mechanism and dynamic routing mechanism, the problem of insufficient multi-level feature fusion in the training of visual language models by diffusion model is solved, realizing more efficient image reconstruction and supervision signal transmission, and improving the generation quality and training efficiency of visual language models.

CN120580446BActive Publication Date: 2025-11-21XIAMEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511094970.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-11-21
Estimated Expiration
2045-08-06

AI Technical Summary

Technical Problem

Existing diffusion models lack a flexible fusion mechanism for multi-level features during visual language model training, resulting in insufficient information fusion, inadequate feature selection and weighting, which affects the quality of image generation and the effective transmission of supervisory signals.

Method used

A hybrid expert cross-attention mechanism is adopted, which processes low, medium and high-level image features and text features through multiple adapters respectively. A dynamic routing mechanism is introduced to adaptively calculate the feature contribution ratio, thereby realizing the dynamic fusion and fine-grained control of multimodal features.

Benefits of technology

It significantly improves the image reconstruction quality and semantic fidelity of visual language models, enhances the expressiveness of supervision signals and training convergence speed, and improves multimodal understanding capabilities and downstream task performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120580446B_ABST
    Figure CN120580446B_ABST
Patent Text Reader

Abstract

The application discloses a method for training a visual language model by using a diffusion model. The method inputs an image and text, extracts multi-level image features through a visual encoder, combines text features output by a connector, and inputs the combined features into a diffusion model as conditional information. The core of the method is to propose a hybrid expert cross-attention mechanism, which respectively constructs independent attention branches for low, medium and high level image features and text features, and dynamically fuses the features through a gating routing mechanism. The fused features are used to guide the step-by-step reconstruction of the image, and the final image result is output. The visual encoder and the connector are optimized in the reverse direction through the perception loss compared with the original image, so as to enhance the perception ability and semantic expression ability of the visual language model.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of artificial intelligence, and particularly relates to an application method and structural design of a multi-modal learning and diffusion generation model in visual language model training, in particular to a method for supervising visual language model training using a diffusion model. BACKGROUND

[0002] Current diffusion models are mainly applied to text-to-image generation tasks. Representative methods such as Stable Diffusion, GLIDE and DALL·E2 input text features as conditions to guide the U-Net model to reconstruct images in the latent space. In recent years, in order to improve the control ability of style, edge and structure, control methods such as ControlNet and T2I-Adapter introduce lightweight adapters, but their control mechanisms are mostly based on single modal or single channel input, lacking the modeling of multi-modal and multi-level feature fusion. In addition, the cross-attention in existing methods usually adopts a unified structure to process different modal features, which is difficult to adapt to the differences in semantic expression and hierarchical structure between images and text.

[0003] Specifically, when used for visual language model supervision, existing diffusion models usually input features as unified conditions and inject them into the diffusion network through a single cross-attention mechanism. However, this approach has the following three main defects: (1) Different levels of features differ significantly in semantic granularity and structural information, and unified processing can easily lead to insufficient or redundant information fusion; (2) The fixed attention structure lacks dynamic modeling ability for the importance of different levels of features, limiting the effective transmission of supervision signals; (3) Current diffusion models lack flexible fusion mechanisms when integrating multiple sets of features, making it impossible to achieve fine-grained feature selection and weighting, resulting in insufficient feedback of generated images to visual encoders and affecting the feature optimization effect of visual language models.

[0004] Therefore, the existing technology cannot fully utilize the guidance of multi-level features on the diffusion reconstruction process. SUMMARY

[0005] The main purpose of the present application is to provide a method for supervising visual language model training using a diffusion model, which solves the problems existing in the prior art and improves the supervision ability of diffusion models in visual language model training.

[0006] In order to achieve the above purpose, the solution of the present application is:

[0007] A method for supervising visual language model training using a diffusion model, comprising the following steps:

[0008] Step S1. Extracting multi-level supervision features

[0009] input image and text; extract multiple levels of intermediate features from the visual encoder respectively represented as low-level features , mid-level features and high-level features , and extract supervised features representing language semantics from the output features of the connector ; wherein the low-level features refer to the features of the 8th layer, the mid-level features refer to the features of the 16th layer, and the high-level features refer to the features of the 24th layer;

[0010] Step S2. Multi-adapter mapping processing

[0011] Through multiple adapters, the four groups of extracted features are respectively mapped and normalized, transforming the original features into a unified conditional vector dimension, and entering the hybrid expert cross-attention mechanism of the diffusion model;

[0012] Step S3. Hybrid expert cross-attention injection mechanism

[0013] An independent cross-attention branch is established for each group of transformed features; each cross-attention layer in the diffusion model simultaneously receives independent processing results from low, mid, and high three different levels of image features and feature responses from the text modality; by learning a lightweight linear transformation matrix, the query vector of the current diffusion stage is mapped to a fusion weight vector, and finally the fused features are output;

[0014] Step S4. Image reconstruction and supervised optimization

[0015] The fused features in step S3 are input into the U-Net backbone network of the diffusion model, and the image is gradually restored at each time step to obtain the final output image; the final output image is compared with the original image, and a perceptual loss is constructed to build a reconstruction loss, which is directly affected by the visual encoder and the connector through gradient backpropagation, thereby improving the perceptual ability and semantic alignment effect of the visual language model, achieving the goal of optimizing the visual language model.

[0016] In the step S1, the input image is defined as , and the output features of the visual encoder at the layer are represented as:

[0017] ;

[0018] wherein represents the 8th layer; represents the 16th layer; represents the 24th layer, represents the visual encoder; ​

[0019] The projection process of the connector to the image features is denoted as:

[0020] ;

[0021] wherein, denotes the connector.

[0022] Preferably, each adapter in the step S2 is a light-weight feed-forward network, denoted as The mapping and normalization process is denoted as:

[0023] ;

[0024] wherein, denotes the four groups of features extracted in step S1, is one of ; denotes the transformed features.

[0025] The adapter is denoted as:

[0026] ;

[0027] wherein, denotes the layer normalization; denotes the non-linear activation function; denotes the linear layer weight matrix; denotes the bias term.

[0028] Preferably, in the step S3, the cross-attention branch is denoted as:

[0029] ;

[0030] wherein, denotes the attention mechanism; denotes the query vector of the current processing stage in the U-Net of the diffusion model; , denote the key-value pair of the th supervision feature, respectively, , denote the trainable parameters, respectively; denotes the Softmax function; denotes the transpose of ; denotes the feature dimension.

[0031] Preferably, in the step S3, the process of mapping the query vector to the fusion weight vector is denoted as:

[0032] ;

[0033] in, Represents the fusion weight vector; Represents a 4-dimensional vector; Indicates the pooling layer; Represents a linear transformation matrix;

[0034] The final output is represented as:

[0035] ;

[0036] in, Indicates the characteristics after fusion; This indicates the contribution ratio of low-level, mid-level, high-level, and textual features.

[0037] Preferably, in step S4, the diffusion model employs a reverse denoising process, expressed as:

[0038] ;

[0039] in, This indicates the final output image; This represents the noise component predicted by the denoising network. Indicates the first A noisy image of the step; Represents the original image; and respectively The proportion of original information retained and the intensity of added noise; Indicates Gaussian noise;

[0040] Reconstruction losses are expressed as:

[0041] ;

[0042] in, Indicates the losses incurred during reconstruction; Represents the square of the L2 norm; Indicates the first Perceptual features extracted from layers; This represents the set of feature layer indices used in perceptual loss.

[0043] After adopting the above technical solution, the present invention has the following technical effects:

[0044] (1) The mixed expert cross-attention mechanism proposed in the application can receive image features from different levels of visual encoders and text features output by the connector in the visual language model, respectively perform modal-specific cross-attention modeling, and realize dynamic fusion through a mixed expert routing mechanism, thereby significantly improving the reconstruction ability and feature supervision efficiency of the diffusion model in the visual language model (VLM) training framework. In the training process of the visual language model, the reconstruction quality and semantic fidelity of the image features can be significantly improved, filling the gap in the current technology in the fine control and supervision of multi-modal propagation.

[0045] (2) By designing independent attention paths for visual features at different levels and introducing a dynamic routing fusion strategy, the diffusion model can adaptively focus on the most discriminative semantic information, thereby more efficiently feeding back image supervision signals to the visual encoder and connector modules. Compared with the traditional unified attention mechanism, the application effectively alleviates the problems of feature redundancy and weak supervision, brings stronger supervision signal expression, higher training convergence speed and better image-text alignment effect, and further improves the multi-modal understanding ability and downstream task performance of the visual language model. BRIEF DESCRIPTION OF DRAWINGS

[0046] Figure 1 The method flowchart of the embodiment of the application.

[0047] Figure 2 The overall framework diagram of the embodiment of the application.

[0048] Figure 3 The mixed expert cross-attention mechanism diagram of the embodiment of the application.

[0049] Figure 4 The performance comparison results of the application and several mainstream visual language models on six widely used multi-modal evaluation benchmarks. DETAILED DESCRIPTION

[0050] In order to further explain the technical solutions of the application, the application will be described in detail below through specific embodiments.

[0051] The key to this invention lies in proposing a diffusion model with a Mixture-of-Experts (MoE) cross-attention mechanism, specifically designed to enhance the supervision capabilities of diffusion models in visual language model training. Specifically, independent cross-modal attention pathways are constructed for image features from different levels (low, middle, and high) of the visual encoder to fully leverage the complementary advantages of each layer's features in semantic granularity and spatial detail. To avoid information redundancy or representational bias caused by traditional fixed fusion methods, this invention introduces a Mixture-of-Experts (MoE) routing mechanism to dynamically fuse the outputs of the aforementioned attention pathways. This routing mechanism automatically calculates the importance weights of each layer's features based on the context query vector content of the current diffusion step, and uses this weighted aggregation of multi-source features to achieve adaptive selection of the optimal feature combination under different contexts, significantly improving the accuracy of contextual modeling and the quality of supervision signals during the diffusion reconstruction process. This structure significantly improves upon the problems of single supervision signal processing methods and insufficient fusion capabilities in existing technologies, solving the technical bottleneck of difficulty in uniformly modeling differences in semantic abstraction and detail expression among image features at different levels. Through this design, the diffusion model can adaptively select the optimal feature combination for image reconstruction based on the input context, thereby enhancing the gradient supervision efficiency of the visual language model, improving the quality of its visual feature representation, and ultimately achieving the training of a more stable and semantically richer visual language model.

[0052] refer to Figures 1 to 3 As shown, this invention discloses a method for supervising the training of a visual language model using a diffusion model, comprising the following steps:

[0053] Step S1. Extract multi-level supervised features

[0054] Input images and text; to achieve high-quality supervision of the visual language model, this invention first starts from the visual encoder. Extract intermediate features from multiple levels, which are then represented as low-level features. Mid-layer characteristics and high-level characteristics Furthermore, it extracts supervisory features representing language semantics from the output features of the connector. The above four sets of characteristics These will be used together as conditional inputs in the subsequent diffusion model. Here, low-level features refer to features from layer 8, mid-level features refer to features from layer 16, and high-level features refer to features from layer 24.

[0055] Specifically, in step S1 above, if the input image is Then the visual encoder in the th Layer output features is represented as:

[0056] ;

[0057] wherein, represents the 8th layer; represents the 16th layer; represents the 24th layer, represents a visual encoder;

[0058] And the projection process of the connector to the image feature is represented as:

[0059] ;

[0060] wherein, represents a connector.

[0061] Step S2. Multi-adapter mapping processing

[0062] Since different levels of features have different semantic granularity and distribution characteristics, in order to improve their compatibility in the diffusion model, the present application maps and normalizes the four groups of features extracted by the multi-adapter respectively, transforms the original features into a unified conditional vector dimension, and enters the mixed expert cross-attention mechanism of the diffusion model.

[0063] Specifically, each adapter in the above step S2 is a lightweight feedforward network, denoted as , and the mapping and normalization process is represented as:

[0064] ;

[0065] wherein, represents the four groups of features extracted in step S1, which is one of ; represents the transformed feature;

[0066] The adapter is generally composed of a linear layer, a normalization layer and an activation function, and is represented as:

[0067] ;

[0068] wherein, represents layer normalization; represents a nonlinear activation function; represents a linear layer weight matrix; represents a bias term.

[0069] Step S3. Mixed expert cross-attention injection mechanism

[0070] In the U-Net structure of the diffusion model, each layer contains a cross-attention module for processing the conditional input. In order to fully utilize the multi-level supervised features, the present application proposes a "Mixture-of-Experts (MoE) cross-attention injection mechanism" to establish an independent cross-attention branch for each group of transformed features (low, medium, high, and text).

[0071] Specifically, the above cross-attention branch is represented as:

[0072] ;

[0073] wherein, represents the attention mechanism; represents the query vector of the current processing stage in the U-Net of the diffusion model; , respectively represent the key-value pairs of the first level supervised feature, , respectively represent trainable parameters; represents the Softmax function; represents the transpose of ; represents the feature dimension.

[0074] In order to realize the automatic adjustment of the influence of each layer feature on image reconstruction according to the input content, the present application designs a dynamic routing mechanism based on input feature content adaptive weight calculation. Each cross-attention layer in the diffusion model simultaneously receives independent processing results from low, medium, and high image features of three different levels and feature responses from the text modality. In order to dynamically control the output contribution of these features, the present application introduces a Gating Router, which learns a lightweight linear transformation matrix to map the query vector of the current diffusion stage to a fusion weight vector, and finally outputs the fused features.

[0075] Specifically, the process of mapping the above query vector to the fusion weight vector is represented as:

[0076] ;

[0077] wherein, represents the fusion weight vector; represents a 4-dimensional vector; represents a pooling layer; represents a linear transformation matrix; wherein the Softmax function ensures that the output weight is normalized between the four cross-attention branches (low, medium, high, and text four-level features), reflecting "adaptive allocation".

[0078] The core of the dynamic routing mechanism is that different input images correspond to different visual semantic complexity, and low-level features may be more critical in detail structure reconstruction, while high-level features are more important in global semantic modeling. Therefore, the router automatically learns and assigns weights to different levels of features according to the context semantics of the query vector at each time step, realizes "input-related dynamic adjustment of feature contribution", and finally outputs the representation as:

[0079] ;

[0080] wherein, denotes the fused features; denotes the contribution proportion of low-level, middle-level, high-level and text features.

[0081] The above dynamic routing mechanism has the following advantages: it can automatically adjust the supervision focus according to the specific input (for example, structure-rich images prefer low-level, and semantic complex images prefer high-level), thereby realizing more accurate context modeling; at the same time, it avoids the problem of insufficient adaptability of fixed fusion strategy in different tasks or scenarios; finally, it improves the semantic fidelity and detail restoration quality of the diffusion image reconstruction, effectively improving the supervision efficiency of the visual encoder.

[0082] Step S4. Image reconstruction and supervision optimization

[0083] The fused features in step S3 are input into the U-Net backbone network of the diffusion model, and the image is gradually restored at each time step to obtain the final output image; the final output image is compared with the original image, and the reconstruction loss is constructed with the perception loss, and the reconstruction loss is directly affected on the visual encoder and the connector through gradient back propagation, thereby improving its perception ability and semantic alignment effect, achieving the goal of optimizing the visual language model.

[0084] Specifically, in the above step S4, the diffusion model uses a reverse denoising process, represented as:

[0085] ;

[0086] wherein, denotes the final output image; denotes the noise component predicted by the denoising network; denotes the noisy image at the th step; denotes the original image; and denote the proportion of preserving original information and the intensity of adding noise, respectively, which are functions changing with the step number; denotes Gaussian noise;

[0087] The reconstruction loss is represented as:

[0088]

[0089] wherein, represents the reconstruction loss; represents the square of the L2 norm (Euclidean norm); represents the perceptual feature extracted at the layer; represents the set of feature layer indices used in the perceptual loss.

[0090] Through the above scheme, the hybrid expert cross attention mechanism proposed by the application can receive image features from different levels of the visual encoder and text features output by the connector in the visual language model, respectively perform modal-specific cross attention modeling, and realize dynamic fusion through a hybrid expert routing mechanism, thereby significantly improving the reconstruction capability and feature supervision efficiency of the diffusion model in the visual language model (VLM) training framework. In the visual language model training process, the reconstruction quality and semantic fidelity of the image features can be significantly improved, filling the gap in the current technology in the fine control and supervision of multi-modal propagation.

[0091] Specifically, by designing independent attention paths for visual features at different levels and introducing a dynamic routing fusion strategy, the diffusion model can adaptively focus on the most discriminative semantic information, thereby more efficiently feeding back image supervision signals to the visual encoder and connector modules. Compared with the traditional unified attention mechanism, the application effectively alleviates the problems of feature redundancy and weakened supervision, brings stronger supervision signal expression, higher training convergence speed and better image-text alignment effect, and thus comprehensively improves the multi-modal understanding ability and downstream task performance of the visual language model.

[0092] The test data of the application are as follows:

[0093] Referring to Figure 4 It is shown that the diffusion supervised visual language model and a plurality of mainstream visual language models (VLMs) are compared in performance on six widely used multi-modal evaluation benchmarks. The diffusion supervised visual language model of the application exhibits comprehensive performance superior to existing SOTA methods. Notably, the diffusion supervised visual language model introduces a diffusion model to perform pixel-level supervision on the visual encoder and the connector during the training stage, effectively shortens the gradient propagation path, and improves the semantic integrity and structural fidelity of image feature expression. Without introducing additional inference overhead, the diffusion supervised visual language model achieves better performance, verifying the universality and efficiency of the diffusion supervised visual language model as a plug-in optimization scheme.

[0094] ​The above embodiments and drawings are not intended to limit the product shape and style of the present application, and any appropriate changes or modifications made by those skilled in the art should be considered as not departing from the scope of the patent of the present application.

Claims

1. A method for supervising the training of a visual language model using a diffusion model, characterized in that... The method comprises the following steps: Step S1. Extracting multi-level supervision features input image and text; extracting multiple levels of intermediate features from a visual encoder respectively denoted as low-level features , mid-level features , and high-level features , and extracting supervised features representing language semantics from the output features of the connector ; wherein the low-level features refer to features at the 8th layer, the mid-level features refer to features at the 16th layer, and the high-level features refer to features at the 24th layer. Step S2. Multi-adapter mapping processing Through the multi-adapter, the four groups of extracted features are respectively mapped and normalized, the original features are transformed into a unified conditional vector dimension, and then enter the mixed expert cross-attention mechanism of the diffusion model; Step S3. Mixed expert cross-attention injection mechanism An independent cross-attention branch is established for each group of transformed features; each cross-attention layer in the diffusion model simultaneously receives independent processing results from low, medium and high three different levels of image features and feature responses from the text mode; by learning a lightweight linear transformation matrix, the query vector of the current diffusion stage is mapped into a fusion weight vector, and finally the fused features are output; Step S4. Image reconstruction and supervised optimization The fused features in step S3 are input into the U-Net backbone network of the diffusion model, and the image is gradually restored at each time step to obtain the final output image; the final output image is compared with the original image, and a reconstruction loss is constructed using a perception loss, and the reconstruction loss is directly affected on the visual encoder and the connector through gradient back propagation, thereby improving the perception ability and semantic alignment effect of the visual encoder and the connector, and achieving the goal of optimizing the visual language model.

2. The method for training a visual language model supervised by a diffusion model according to claim 1, wherein: In the step S1, the input image is defined as , the output features of the visual encoder at the first layer are represented as: ​ ; wherein, represents the 8th layer; represents the 16th layer; represents the 24th layer, represents a visual encoder; The projection process of the connector on the image features is represented as: ; wherein represents a connector.

3. The method for training a visual language model supervised by a diffusion model according to claim 2, wherein: Each of the adapters in the step S2 is a light-weight feedforward network, denoted as The process of mapping and normalizing is denoted as: ; wherein, represents the four groups of features extracted in step S1, is one of the groups in the set represents the transformed features; The adapter is represented as: ; wherein, denotes a layer normalization; denotes a nonlinear activation function; denotes a linear layer weight matrix; denotes a bias term.

4. The method for training a visual language model supervised by a diffusion model according to claim 3, wherein: In the step S3, the cross-attention branch is represented as: ; wherein, denotes the attention mechanism; denotes the query vector of the current processing stage in the diffusion model’s U-Net; , denote the first key-value pair of the supervision feature, , denote the trainable parameters; denotes the Softmax function; denotes the transpose of ; denotes the feature dimension.

5. The method for training a visual language model supervised by a diffusion model according to claim 4, wherein: In step S3, the process of mapping the query vector into the fusion weight vector is represented as: ; wherein, denotes a fusion weight vector; denotes a 4-dimensional vector; denotes a pooling layer; denotes a linear transformation matrix; The final output is represented as: ; wherein, represents the fused features; represents the contribution proportion of low-level, middle-level, high-level, and text features.

6. The method for training a visual language model supervised by a diffusion model according to claim 5, wherein: In step S4, the diffusion model adopts a reverse denoising process, represented as: ; wherein, denotes the final output image; denotes the noise component predicted by the denoising network; denotes the noisy image of the step; denotes the original image; and denote the proportion of the original information preserved and the intensity of the added noise, respectively; denote the proportion of the original information preserved and the intensity of the added noise, respectively; denotes Gaussian noise; The reconstruction loss is represented as: ; wherein, denotes a reconstruction loss; denotes a square of an L2 norm; denotes a perceptual feature extracted at the layer; denotes a set of feature layer indices used in the perceptual loss.

Citation Information

Patent Citations

  • Automatic driving perception method and device based on multi-modal multi-scale fusion

    CN118072286A

  • Remote sensing image super-resolution method and product based on diffusion model and multi-modal large language model

    CN119722462A