An exchange-based multi-modal multi-scale transformation fusion method and system

Through the multimodal multi-scale transformation fusion method, combining the multi-head self-attention and multi-modal information exchange unit of image and text data, the information loss and heterogeneity problems in the integration of medical images and text data are solved, and more efficient cross-modal feature representation and diagnostic accuracy are achieved, which is suitable for medical image analysis.

CN119538188BActive Publication Date: 2025-07-11CHONGQING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411596655.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-11
Publication Date
2025-07-11
Estimated Expiration
2044-11-11

AI Technical Summary

Technical Problem

The existing multimodal fusion technology has problems of information loss and cross-modal relationship complexity in the integration of medical images and text data, especially in the heterogeneity processing between images and text data, and the existing multi-scale method mainly focuses on image modality, ignoring the multi-scale complementary learning of text data.

Method used

The multimodal multi-scale transformation fusion method based on exchange is adopted. Through a multimodal encoder, a decoder, a channel-based information exchange module and a multi-scale fusion module, combined with image and text data, information exchange and multi-scale feature fusion are used to perform information exchange and multi-scale feature fusion, thereby realizing complementary learning of cross-modal features.

Benefits of technology

Improve the feature learning ability of medical images and text data, enhance the convergence and diagnostic accuracy of the model, especially in the case of scarcity of labeled data, and achieve higher classification performance and interpretability, supporting clinician decision support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119538188B_ABST
    Figure CN119538188B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for multi-modal multi-scale transformation fusion based on exchange. The method includes: obtaining original image data and original text data and inputting them into the MMTF model to generate a fusion result. Among them, the MMTF model includes: a multi-modal encoder module, a decoder module, a channel-based information exchange module, and a multi-scale fusion module. The multi-modal encoder module includes a text encoder and a dual-branch image decoder; the decoder module decodes the embeddings generated by the encoder; the channel-based information exchange module exchanges information on the embeddings of different modalities on different channels; the multi-scale fusion module is used to fuse the cls tokens from one branch and the patch tokens from another branch according to the image features and text features on different branches. The method of the present invention can provide reliable decision support in various medical environments, improve diagnostic accuracy and reduce the workload of clinicians.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of machine learning, and particularly relates to a multi-modal multi-scale transformation fusion method and system based on swapping. Background Art

[0002] In recent years, data-driven artificial intelligence (AI) has developed rapidly, profoundly changing the field of multi-modal machine learning. Multi-modal fusion is a core challenge in this field, aiming to unify heterogeneous data into a single representation form while retaining the semantic features of each modality while effectively reducing the dimension. In the field of healthcare, patient data usually includes imaging (such as radiological scans) and text records (such as electronic health records or diagnostic reports). Especially text data contains detailed descriptions of the medical conditions observed by radiologists, thus generating complex and less transparent cross-modal relationships, making data integration more complex.

[0003] To address this challenge, research increasingly adopts deep multi-modal fusion techniques. Deep multi-modal fusion techniques are mainly divided into aggregation-based and alignment-based methods. Aggregation-based methods represent each modality through sub-networks and combine these representations using various operators. However, this method sacrifices the ability to make independent predictions for each modality, resulting in some loss of within-modal information. In contrast, alignment-based methods use alignment losses to maintain the consistency of multi-modal features, while retaining the independent outputs of multiple sub-networks and weighting the final predictions. However, alignment-based fusion only adds a regularization term to the original single-modal optimization objective, without achieving true cross-modal fusion or promoting the necessary inter-modal information exchange. In addition, some research has explored channel-based modality swapping methods to balance within-modal processing and inter-modal fusion. For example, by using the Batch Normalization (BN) scaling factor to measure the channel importance, channels with factors close to zero are replaced with the average channel of another modality. However, this channel swapping technique often ignores the heterogeneity between image and text data, making it difficult to effectively represent cross-modal data in a unified low-dimensional space. In addition, since pathological regions usually only occupy a small part of medical images, single-scale feature representations face limitations in capturing local lesion details.

[0004] As a solution, integrating multi-scale structures into the model has been proven to effectively improve the feature learning ability. The multi-scale network structure samples the input data at different granularities, which may result in regions of interest (ROIs) having different fine-grained and coarse-grained features. Fine-grained features retain more detailed input data, while coarse-grained features capture the overall trend. Most disease detection tasks consider multi-scale factors during image classification.

[0005] During the artificial intelligence training process in medical applications, integrating features of various data scales can enhance the convergence of deep neural networks. Fine-grained features capture detailed information, while coarse-grained features reveal the overall trend, both of which are crucial for optimizing model performance. This understanding has led to an increasing focus on combining multi-modal and multi-scale learning strategies. However, existing multi-scale methods in medical artificial intelligence mainly focus on combining image modalities, and there is limited exploration of multi-scale complementary learning between images and text reports. Summary of the Invention

[0006] In view of this, the purpose of the present invention is to provide an exchange-based multi-modal multi-scale transformation fusion method and system.

[0007] The purpose of the present invention is achieved through the following technical solutions:

[0008] An exchange-based multi-modal multi-scale transformation fusion method, comprising:

[0009] Obtain multi-modal data, namely, original image data and original text data corresponding to the original image data;

[0010] Input the original image data and the original text data into a multi-modal multi-scale transformation fusion model to generate a fusion result, wherein the multi-modal multi-scale transformation fusion model includes: a multi-modal encoder module, a decoder module, a channel-based information exchange module, and a multi-scale fusion module.

[0011] The multi-modal encoder module includes a text encoder and a dual-branch image decoder.

[0012] The text encoder is used to encode the input text data to generate text embeddings.

[0013] The dual-branch image encoder includes a small-branch image encoder and a large-branch image encoder, which are respectively used to extract fine-grained image embeddings and coarse-grained image embeddings of the input image data.

[0014] The decoder module contains two decoders that respectively decode the embeddings generated by the text encoder and the dual-branch image encoder, so as to perform two generation tasks of generating images from text and generating text from images, in order to regularize the embeddings generated by the text encoder and the dual-branch image decoder;

[0015] The channel-based information exchange module is used to exchange information of the embeddings of different modalities on different channels according to the attention scores of different channels, so as to generate image features and text features on different branches;

[0016] The multi-scale fusion module is used to fuse the cls tokens from one branch and the patch tokens from another branch according to the image features and text features on different branches.

[0017] The present invention also provides a multi-modal multi-scale transformation fusion system based on exchange, including:

[0018] The multi-modal data acquisition engine is used to acquire multi-modal data, that is, image data and text data corresponding to the image data;

[0019] The model training engine is used to train the multi-modal multi-scale transformation fusion model,

[0020] The fusion generation engine is used to generate a fusion result according to the acquired multi-modal data by using the trained multi-modal multi-scale transformation fusion model,

[0021] wherein, the multi-modal multi-scale transformation fusion model includes: a multi-modal encoder module, a decoder module, a channel-based information exchange module and a multi-scale fusion module,

[0022] The multi-modal encoder module includes a text encoder and a dual-branch image decoder,

[0023] The text encoder is used to encode the input text data to generate text embeddings;

[0024] The dual-branch image encoder includes a small-branch image encoder and a large-branch image encoder, which are respectively used to extract fine-grained image embeddings and coarse-grained image embeddings of the input image data;

[0025] The decoder module contains two decoders that respectively decode the embeddings generated by the text encoder and the dual-branch image encoder, so as to perform two generation tasks of generating images from text and generating text from images, in order to regularize the embeddings generated by the text encoder and the dual-branch image decoder;

[0026] The channel-based information exchange module is used to exchange information of the embeddings of different modalities on different channels according to the attention scores of different channels, so as to generate image features and text features on different branches;

[0027] A multi-scale fusion module for fusing the cls token from one branch and the patch token from another branch according to the image features and text features on different branches.

[0028] Furthermore, the information exchange of the embeddings of different modalities includes the information exchange between the coarse-grained image embedding and the text embedding and the information exchange between the fine-grained image embedding and the text embedding. The channel-based information exchange module also uses two hyperparameters η and μ to control the start layer and the end layer of the multi-modal information exchange respectively.

[0029] Furthermore, the channel-based information exchange module includes a multi-head self-attention unit and a multi-modal information exchange unit. The text embedding, the fine-grained image embedding, and the coarse-grained image embedding, after adding their respective cls tokens, are fed into the multi-head self-attention unit to obtain the attention scores of each token in the text embedding, the fine-grained image embedding, and the coarse-grained image embedding. The calculated attention scores are sent to the multi-modal information exchange unit through a feed-forward network with residual connections. The multi-modal information exchange unit uses multiple Transformer encoders with shared parameters to exchange multi-modal information of different scales in a complementary manner.

[0030] Furthermore, the multi-modal information exchange unit exchanges multi-modal information of different scales in a complementary manner, including: for the information exchange of any layer, for any channel, if the attention score of the current channel is lower than the set threshold, the embedding vector of this channel will be replaced by the average embedding of a preset percentage of tokens in another modality, otherwise, its embedding vector remains unchanged, so as to obtain an updated embedding matrix.

[0031] Furthermore, the channel-based information exchange module also includes an FFN and a CB unit. After the information exchange of the current layer is completed, the FFN generates the input embedding for the information exchange of the next layer according to the updated embedding matrix, and the CB is used to broadcast the context to each token.

[0032] Furthermore, fusing the cls token from one branch and the patch token from another branch specifically includes:

[0033] Performing a concatenation operation on the image cls token from the large branch and the text patch token from the small branch, and performing cross-attention with the image cls token of the large branch, and then performing a concatenation operation on the patch token in the result of the cross-attention with the image cls token of the large branch, so as to obtain the fused image feature of the large branch, that is, the coarse-grained fused image feature;

[0034] Perform a concatenation operation on the image cls token from the small branch and the text patch token from the large branch, and perform cross-attention with the image cls token of the small branch. Then, perform a concatenation operation on the patch token in the result of the cross-attention with the image cls token of the small branch, thereby obtaining the fused image feature of the small branch, that is, the fine-grained fused image feature;

[0035] Perform a concatenation operation on the text cls token from the large branch and the image patch token from the small branch, and perform cross-attention with the text cls token of the large branch. Then, perform a concatenation operation on the patch token in the result of the cross-attention with the text cls token of the large branch, thereby obtaining the fused text feature of the large branch, that is, the coarse-grained fused text feature;

[0036] Perform a concatenation operation on the text cls token from the small branch and the image patch token from the large branch, and perform cross-attention with the text cls token of the small branch. Then, perform a concatenation operation on the patch token in the result of the cross-attention with the text cls token of the small branch, thereby obtaining the fused text feature of the small branch, that is, the fine-grained fused text feature.

[0037] The beneficial effects of the present invention are as follows:

[0038] The present invention captures the correlation between patient imaging and medical reports through image and report generation tasks, jointly regularizes multi-modal embeddings, and aligns them into a shared latent space to achieve cross-modal interaction; the present invention constructs a multi-modal multi-scale transformation fusion model (i.e., the MMTF model), introduces multi-path Transformer, and effectively combines features from different modalities and scales through the cross-attention mechanism, thereby aggregating cross-modal correlations at multiple scales and realizing semantic knowledge exchange between images and text reports at different scales; extensive experiments on four public lung disease datasets show that the proposed MMTF model achieves state-of-the-art classification performance.

[0039] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following specification. Brief Description of the Drawings

[0040] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings, where:

[0041] Figure 1 is a schematic diagram of the MMTF model;

[0042] Figure 2 It is a schematic diagram of channel-based multi-modal information exchange;

[0043] Figure 3 It is a schematic diagram of large-branch multi-scale fusion;

[0044] Figure 4 It is the comparison result of the MMTF model and the baseline model with different proportions of labeled data in the target domain;

[0045] Figure 5 It is a comparison chart of the MMTF model and the SOTA medical image text pre-training model in NIH ChestX-ray and VinBigData chest X-ray anomaly detection;

[0046] Figure 6 It is a comparison chart of the MMTF model and SOTA multi-modal lung diagnosis in NIH ChestX-ray, VinBigData ChestX X-ray anomaly detection and Shenzhen tuberculosis dataset;

[0047] Figure 7 It is the visualization of randomly selected samples;

[0048] Figure 8 It is a t-SNE visual clustering result chart of the Shenzhen tuberculosis and pneumonia image data acquisition datasets caused by the novel coronavirus. Detailed implementation manners

[0049] The preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the preferred embodiments are only for illustrating the present invention rather than limiting the protection scope of the present invention.

[0050] The present invention proposes a multi-modal multi-scale transformation fusion system based on exchange, including:

[0051] A multi-modal data acquisition engine for acquiring multi-modal data, that is, image data and text data corresponding to the image data;

[0052] A model training engine for training the multi-modal multi-scale transformation fusion model;

[0053] A fusion generation engine for generating a fusion result according to the acquired multi-modal data by using the trained multi-modal multi-scale transformation fusion model.

[0054] The present invention also provides a multi-modal multi-scale transformation fusion method based on exchange, and the method includes:

[0055] Acquiring multi-modal data, that is, original image data and original text data corresponding to the original image data;

[0056] The original image data and the original text data are input into the multimodal multi-scale transformation fusion model to generate a fusion result.

[0057] Figure 1 The following is a schematic diagram of a multi-modal multi-scale transform fusion model (MMTF model) using lung disease as an example. Figure 1 As shown, the multi-modal multi-scale transformation fusion model mainly includes four functional modules, namely: a multi-modal encoder module (i.e. Figure 1 The components marked in ①), the decoder module (i.e. Figure 1 ② and ③), channel-based information exchange module (i.e. Figure 1 The component marked in (4) and the multi-scale fusion module (i.e. Figure 1 Components marked in ⑤).

[0058] The multimodal encoder module consists of a text encoder and a dual-branch image decoder.

[0059] The text encoder is used to encode the input text data (e.g., medical reports, etc.) and generate text embeddings.

[0060] A dual-branch image encoder is designed to obtain multi-scale features of an image (e.g., a medical radiograph), which may include a small-branch image encoder and a large-branch image encoder, which are respectively used to extract fine-grained image embedding and coarse-grained image embedding of the input image data.

[0061] Given a pair of inputs [x v ,x t ], where x v stands for radiograph, x t is the corresponding report. The text encoder is used to extract the semantic information of the report and obtain the text embedding T t , and obtain a wider embedding of the image (i.e., coarse-grained image embedding I L ) and smaller embeddings (i.e., fine-grained image embeddings I S ). These encoded features are input into subsequent components to learn multimodal and multi-scale representations. Finally, the learned multimodal fusion feature representation is used for multi-label classification tasks.

[0062] In some embodiments, a dual-branch image encoder is used to extract features from radiographs to obtain multi-scale image features. For example, a Resnet34 architecture can be used to extract coarse-grained features. That is, coarse-grained image embedding, while the Resnet152 architecture can be used to extract fine-grained features That is, coarse-grained image embedding. Where M and N are the number of tags and feature dimensions at each scale, respectively. The set of real numbers is denoted as \( \mathbb{R} \), and \( C \) represents the number of channels of the image, which is expressed by the formula:

[0063] I L = Renet34(x v ), (Formula 1)

[0064] I S = Resnet152(x v ), (Formula 2)

[0065] Reports usually consist of long paragraphs and require reasoning across multiple sentences. Therefore, the present invention uses a language model based on the self-attention mechanism to learn long-range semantic dependencies in reports. Different from using a dual-branch image encoder to obtain multi-scale image features, the original report has no additional processing. The cumbersome and complex processing of local text features not only ignores rare words but may also confuse synonyms in the medical field. In some embodiments, the MMTF model can adopt the BioClinicalBERT model as a text encoder to obtain clinically-aware text embeddings, which is expressed by the formula::

[0066] T t = BertEncoder(x t ), (Formula 3)

[0067] BertEncoder in Formula 3 is the BioClinicalBERT model.

[0068] In the medical field, due to the modality variability between radiographs and reports, the encoded multi-modal features are usually in different vector spaces. Therefore, for channel-based multi-modal information exchange, these embeddings need to be first pulled into the same space. To capture the correlation between reports and radiographs, the present invention designs two generation tasks (mutual generation between reports and radiographs). As shown in Components ② and ③, the generation tasks regularize the embeddings of the encoder through two decoders.

[0069] As Figure 1 shown, the decoder module contains two decoders, which respectively decode the embeddings generated by the text encoder and the dual-branch image encoder, so as to generate images according to the text and generate text according to the image for two generation tasks, in order to realize the regularization of the embeddings generated by the text encoder and the dual-branch image decoder.

[0070] In Component ②, the received text embeddings are processed by an MLP and then decoded by an image decoder to generate an image; in Component ③, the received large-branch image embeddings (i.e., coarse-grained image embeddings) are processed by an MLP and then decoded by a text decoder to generate text corresponding to different scales (or called "branches").

[0071] For the generation task of generating text from images (i.e., the generation task shown in Component ③), the NIC model can be used as a decoder to generate a corresponding report based on the input radiograph. Similarly, the image generation model PixelCNN can be used as a decoder to extract embeddings from the text encoder and generate the corresponding radiograph (such as the generation task shown in Component ②).

[0072] By utilizing these two generation tasks, the relationship between images and text can be strengthened and explored. By comparing the generated radiographs and reports with the original input radiographs and reports, two generation losses are generated, namely and wherein is the generation loss of generating text from images, is the generation loss of generating images from text. These two generation tasks are auxiliary tasks for the prediction task, Figure 1 in is the loss of the main task of multi-label classification (e.g., disease classification).

[0073] The channel-based information exchange module is used to exchange information of different modalities' embeddings on different channels according to the attention scores of different channels, so as to generate image features and text features at different scales (i.e., branches).

[0074] Exchanging information of different modalities' embeddings includes exchanging information between the coarse-grained image embedding and the text embedding, and exchanging information between the fine-grained image embedding and the text embedding.

[0075] As Figure 2 shown, the channel-based information exchange module includes a multi-head self-attention unit and a multi-modal information exchange unit (i.e., the "transformer based on information exchange" in Figure 1 ). The multi-modal information exchange unit employs multiple transformer encoders with shared parameters (i.e., Transformer encoders, or transformer encoders), which are used to learn to embed text and visual modalities at multiple scales. The Transformer encoder learns global context information from the input vectors in its shallow layer and performs channel-based information exchange between modalities, that is, exchanging multi-modal information at different scales in a complementary manner. Since the downstream task is to perform multi-label classification, a cls token (or called "token") for global feature aggregation is added at the beginning of the input sequence (i.e., text embedding and image embedding). After adding the respective cls tokens to the text embedding, the fine-grained image embedding, and the coarse-grained image embedding, they are fed into the multi-head self-attention unit to learn global context information and obtain the attention scores of each token (i.e., patch tokens) in the text embedding, the fine-grained image embedding, and the coarse-grained image embedding.

[0076] In the self-attention mechanism, the input is linearly mapped to obtain three matrices: a query matrix Q, a key matrix K, and a value matrix V. The output calculation is expressed as:

[0077]

[0078] where d k is the scaling factor, and Softmax represents the normalization function.

[0079] The channel-based information exchange module sets its shallow layer as a regular Transformer encoder layer, followed by multiple exchange layers for information exchange. When the multi-modal multi-scale fusion ends, the exchange process ends. To this end, the channel-based information exchange module also uses two hyperparameters η and μ to control the start layer and the end layer of the multi-modal information exchange respectively. When the multi-modal embedding enters the η layer, the information exchange starts and stops until the μ layer. In some embodiments, there are 6 exchange layers in total. Therefore, 6 information exchanges will be performed (as Figure 1 in component ② of

[0080] The "transformer based on information exchange" is a 6-layer structure). After the channel-based multi-modal information exchange ends, the exchanged embedding results are fed into the subsequent module for multi-modal multi-scale fusion.

[0081] For the information exchange of any layer (i.e., any layer of the exchange layer), for any channel, if the attention score of the current channel is lower than the set threshold, the embedding vector of this channel will be replaced by the average embedding of a preset percentage of tokens in another modality, otherwise, the embedding vector of this channel remains unchanged, thus obtaining an updated embedding matrix.

[0082]

[0083] where θ is the preset threshold, γ is the replacement rate of the information exchange, and are the intermediate embeddings of the coarse-grained image and the text corresponding to the coarse-grained image generated by the multi-head self-attention unit respectively; represents the attention score of the i-th row in , j represents the row of the intermediate embedding of the text ; n represents the total number of tokens on one channel in

[0084] The information exchange process between the fine-grained image and the text is the same as that between the coarse-grained image and the text, which will not be elaborated here.

[0085] The text embedding on one channel during the multi-modal information exchange process can be expressed as:

[0086]

[0087] Wherein, represents the corresponding coarse-grained text embedding generated by the information exchange between the coarse-grained image embedding and the text, and n represents the total number of tokens in on one channel.

[0088] The corresponding fine-grained text embedding generated by the information exchange between the fine-grained image embedding and the text can be expressed as The specific implementation method is the same as Formula 6, and only needs to replace in Formula 6 with the intermediate embedding of the fine- and coarse-grained images generated by the multi-head self-attention unit That's it.

[0089] The channel-based information exchange module also includes an FFN (i.e., Feed-Forward Network) and a CB (i.e., Context Broadcasting) unit. After the information exchange at the current layer is completed, the FFN generates the input embedding for the next layer of information exchange based on the updated embedding matrix, and the CB is used to broadcast the context to each token.

[0090] Figure 2 also magnifies the detailed structures of the FNN and the CB, and the specific details can be referred to Figure 2 , which will not be elaborated here.

[0091] As mentioned above, the radiograph multi-modal data contains extensive pathological information and different concerns. To provide sufficient feature extraction on each modality, the present invention also designs a multi-modal multi-scale fusion module (abbreviated as "multi-scale fusion module"). After being processed by the channel-based multi-modal information exchange module, coarse-grained coarse-grained fine-grained and fine-grained multi-modal features can be obtained. Subsequently, they need to be sent to the multi-scale fusion module for further representation learning.

[0092] The multi-modal multi-scale fusion module can be used to fuse the cls tokens from one branch and the patch tokens from another branch according to the image features and text features on different branches. As Figure 1As shown, the multi-modal multi-scale fusion module may include a multi-scale fusion transformer (or transducer), which can fuse image features and text features on different branches to generate coarse-grained fused image features, fine-grained fused image features, coarse-grained fused text features, and fine-grained fused image features. Each fusion operation generates the above four fused features, and the multi-scale fusion transformer has a total of four layers, that is, four fusion processes are performed.

[0093] Figure 3 Figure 4 is a schematic diagram of multi-scale fusion of the large branch (i.e., large scale or coarse-grained). As Figure 3 shown, the cls token of the large branch acts as a query token and interacts with the patch tokens of the small branch through the cross-attention mechanism. The smaller branch follows the same process but exchanges the patch tokens from the other branch.

[0094] The model proposed by the present invention can use the information exchange between the image cls token from the large branch and the text patch token from the small branch, and then return to the corresponding branch (i.e., the large branch). At the same time, the same operation is performed between the image cls token of the small branch and the text patch token of the large branch; it can also use the information exchange between the text cls token from the large branch and the image patch token from the small branch, and then return to the corresponding branch (i.e., the large branch). At the same time, the same operation is performed between the text cls token of the small branch and the image patch token of the large branch.

[0095] Since the image cls tokens at each scale have learned the abstract information in all patch tokens in their branches, interacting with the text patch tokens on other branches helps to learn information at different scales. After fusing with the text tokens from other branches, the image cls tokens interact with their patch tokens again in the subsequent transformer encoder. This operation helps to efficiently transfer the information learned from other branches to its patch tokens, thus enriching the multi-scale representation of the image patch tokens.

[0096] Perform a concatenation operation on the image cls token from the large branch and the text patch token from the small branch, and perform cross-attention with the image cls token of the large branch. Then, concatenate the patch tokens in the result of the cross-attention with the image cls token of the large branch to obtain the fused image feature of the large branch, that is, the coarse-grained fused image feature I′ L .

[0097] Concatenate the image cls token from the small branch and the text patch token from the large branch, perform cross-attention with the image cls token of the small branch, and then concatenate the patch token in the result of the cross-attention with the image cls token of the small branch to obtain the fused image feature of the small branch, that is, the fine-grained fused image feature I′ S ;

[0098] Concatenate the text cls token from the large branch and the image patch token from the small branch, perform cross-attention with the text cls token of the large branch, and then concatenate the patch token in the result of the cross-attention with the text cls token of the large branch to obtain the fused text feature of the large branch, that is, the coarse-grained fused text feature;

[0099] Concatenate the text cls token from the small branch and the image patch token from the large branch, perform cross-attention with the text cls token of the small branch, and then concatenate the patch token in the result of the cross-attention with the text cls token of the small branch to obtain the fused text feature of the small branch, that is, the fine-grained fused text feature.

[0100] The fused image features of different scales generated in the above multi-scale interaction process can be expressed by formulas 7 and 8 respectively as follows:

[0101]

[0102] Among them, the cls token with a subscript represents the cls token of the image or text feature generated after the information exchange is completed; the coarse-grained is represented by the superscript L, and the fine-grained is represented by the superscript S; the one with a patch subscript represents the corresponding patch token; is the concatenation operation, CA is the cross-attention, specifically as formula 9:

[0103]

[0104] where, W q ,W k ,W v are learnable matrix parameters.

[0105] It should be noted that the above formulas 7 and 8 only represent the process of generating fused image features of different scales in the multi-scale interaction process; since the process of generating fused text features of different scales in the multi-scale interaction process and its formula expression are similar, it is necessary to exchange I representing the image and T representing the text in formulas 7 and 8 t to achieve this.

[0106] After multi-scale information fusion, the output embedding of the multi-scale feature fusion module is finally connected to a fully connected network to obtain the final fused embedding matrix and subsequent multi-label classification predictions.

[0107] In summary, the overall loss of the overall optimization objective function during the MMTF training process can be expressed as:

[0108]

[0109] where α and β are hyperparameters used to measure the importance of the three tasks. (In the subsequent experiments, both α and β are set to 0.0001).

[0110] To evaluate the effectiveness and reliability of the MMTF model, it is pre-trained using the large-scale chest X-ray image and radiology report dataset MIMIC-CXR-JPG, and then fine-tuned on four public datasets. The experimental results show that the method of the present invention outperforms the previously state-of-the-art network models in terms of classification performance. In the following, an in-depth overview of the dataset, evaluation metrics, and experimental settings will be provided. Then, the experimental results of all datasets are presented, including a comprehensive performance analysis of the model of the present invention based on a series of ablation studies conducted on the VinBigData dataset.

[0111] The present invention uses the MIMIC-CXR-JPG dataset for pre-training. MIMIC-CXR-JPG is a publicly accessible dataset of chest X-ray photographs with attached X-ray reports. MIMIC-CXR-JPG contains more than 370,000 radiographs, corresponding to 227,835 radiology studies on 65,279 patients. Each radiology study is accompanied by a chest X-ray image, a free-text radiology report, and abnormal / disease labels obtained through manual-assisted intervention. Among them, the report is a summary by radiologists of their findings, including examinations, inductions, impressions, findings, techniques, and comparisons. In fact, only two main parts are retained in the report: findings and impressions. Findings are descriptions of the basic parts of the radiograph, while impressions summarize the most directly relevant findings.

[0112] To evaluate the model, fine-tuning and ablation experiments are also conducted on the model. The training set, validation set, and test are all divided using the official partition, with partition percentages of 70%, 10%, and 20% respectively. A detailed description of the fine-tuning dataset and implementation is provided below.

[0113] NIHChestX-ray14: NIHChestX-ray14 has 112,120 disease-labeled chest X-rays from 30,805 patients. This dataset includes 14 chest abnormalities: atelectasis, cardiomegaly, effusion, infiltration, mass, nodule, pneumonia, pneumothorax, consolidation, edema, emphysema, fibrosis, pleural thickening, and hernia. In the radiographs, all labels were obtained by mining relevant reports using natural language processing tools.

[0114] VinBigData: VinBigData provides 14 chest abnormalities, namely aortic enlargement, atelectasis, pneumothorax, pulmonary opacity, pleural thickening, interstitial lung disease, pulmonary fibrosis, calcification, pleural effusion, consolidation, cardiomegaly, other lesions, nodule / mass, and infiltration. The training set contains 15,000 individually labeled chest X-ray images, and the test set contains 3,000 chest X-ray images.

[0115] ShenzhenTuberculosis: Shenzhen Tuberculosis was collected through a collaboration between Guangzhou Medical University and the Third People's Hospital of Shenzhen. The chest X-rays mainly came from the outpatient department, including 662 frontal chest X-ray images. Among them, 326 were normal subjects, and the remaining 336 were cases with tuberculosis manifestations.

[0116] COVID-19 Image DataCollection: The pneumonia image dataset caused by the novel coronavirus has more than 900 cases of chest X-ray images of pneumonia. Experiments were conducted to distinguish between pneumonia cases caused by the novel coronavirus and those not caused by the novel coronavirus, as well as to distinguish between viral pneumonia and bacterial pneumonia cases.

[0117] Pre-training was performed on the MIMIC-CXR-JPG dataset. During the entire pre-training process, the batch size was set to 64, and the number of epochs was set to 5. The AdamW optimizer was used, and its initial learning rate was set to 5×10 -5 . A linear learning rate warm-up strategy was adopted to accelerate model convergence. The training was carried out on a single 4090 GPU in a PyTorch environment using the MIMIC-CXR-JPG dataset.

[0118] In the fine-tuning stage, the batch size was set to 32, and the number of epochs was set to 10. The AdamW optimizer was used, and the learning rate was set to 1×10 -5 , and the decay strategy was consistent with the pre-training process. Except for cropping the images to a uniform pixel size, no additional data augmentation was applied. The area under the ROC curve (AUROC) was used to evaluate the model performance to ensure the consistency of the metrics in the experiment. This setting allows for a direct transition from pre-training to fine-tuning, emphasizing the stability and simplicity of parameter handling.

[0119] According to the patterns used in the pre-training stage, state-of-the-art models are generally divided into two categories: medical image pre-training methods and medical image-text pre-training methods. Therefore, MMTF was compared and validated with existing state-of-the-art models on four fine-tuning datasets.

[0120] The medical image pre-training methods involved in the comparison include:

[0121] ModelGenesis: A self-supervised visual representation learning method that can perform various types of generative restoration based on the input radiographs;

[0122] Comparedto Learn: A self-supervised visual representation learning method that constructs homogeneous and heterogeneous data pairs by mixing images and features;

[0123] ImageNet Pre-training: A representative pre-training method for supervised learning on large-scale natural image datasets.

[0124] In addition, for better comparison, supervised ConvNet and Transformer-based were also used for comparative experiments.

[0125] The medical image-text pre-training methods involved in the comparison include:

[0126] MuSE: Performs multimodal fusion by channel swapping and is good at entity recognition and sentiment analysis tasks.

[0127] ConVIRT: Through bidirectional contrastive learning, jointly trains visual and text encoders with paired medical images and reports.

[0128] GLoRIA: Performs multimodal feature representation of medical images through global and local contrastive learning.

[0129] MedKLIP: Uses a report filter to extract medical entities and uses more complex modal fusion to aggregate features.

[0130] Embodiments of the present invention were compared with other SOTA image pre-training and image-text pre-training models in the same deployment environment, and ablation experiments were accordingly carried out. All image pre-training and image-text pre-training baselines were first pre-trained on MIMIC-CXR-JPG. In the fine-tuning stage, the transferability of the models was evaluated by fine-tuning the models using 1%, 10%, and 100% portions of four mature datasets (including NIH ChestX-ray, VinBigData Chest X-ray anomaly detection, Shenzhen tuberculosis, and the pneumonia image dataset caused by the novel coronavirus).

[0131] In Table 1, the classification results of the image pre-training methods for four fine-tuning datasets were examined, covering limited and full supervision. These results are consistent with the broader medical image analysis literature. The evaluation results for specific disease types are detailed in the appendix. Overall, this analysis emphasizes that MMTF consistently outperforms the baseline methods on datasets with different labeling ratios, especially in cases where annotated data is limited.

[0132] Table 1 Classification results of the image pre-training methods for four fine-tuning datasets

[0133]

[0134]

[0135] In Table 1, NIH, VBD, and SZ represent the NIH ChestX-ray, VinBigData ChestX-ray anomaly detection, and Shenzhen tuberculosis datasets. C-T1 (or written as CT1) and C-T2 (or written as CT2) represent two tasks in the collection of pneumonia image data caused by the novel coronavirus: distinguishing pneumonia cases caused by the novel coronavirus from non-novel coronavirus pneumonia cases (C-T1) and distinguishing viral pneumonia cases from bacterial pneumonia cases (C-T2). The evaluation metric is AUROC (i.e., Area Under the Receiver Operating Characteristic), and the best results are shown in bold.

[0136] Specifically, in the fine-tuning stage, compared with traditional pre-trained models, MMTF demonstrated excellent performance by achieving the highest AUROC, as Figure 4 shown. In Figure 4 , NIH ChestX-ray (a) and VinbBigData chest X-ray anomaly detection (b). AUROC indicates the macro-average of all diseases. It should be noted the percentage of labels used in the training data. As Figure 4 shown, on the NIH ChestX-ray dataset, MMTF improved by an average of 6% compared to pre-trained models such as Model Genesis and ImageNet-based models under limited supervision. In the context of medical image analysis, this result is particularly important because obtaining a large amount of expert annotations is challenging. Additionally, on the VinBigData dataset, MMTF achieved an AUROC of 90.9% with only 10% of the labeled data, exceeding models trained on fully annotated datasets. These findings illustrate the strong representation ability of MMTF, especially when dealing with small datasets.

[0137] Further analysis shows that MMTF performs excellently in tasks related to tuberculosis detection, outperforming the C2L model by 2% in AUROC. Even in simpler tasks such as the classification of pneumonia caused by the novel coronavirus, MMTF also exceeds the Transformer model by 1.4%, mainly due to its efficient channel-based multi-modal information exchange module. The most significant advantage of this model lies in distinguishing viral infections from bacterial infections, with its AUROC at least 4.8% higher than that of the baseline method. These results highlight MMTF's ability to transfer multi-modal representations learned across tasks, making it particularly valuable in clinical settings where labeled data is often scarce.

[0138] Table 2 shows the image-text pre-training classification results on four fine-tuning datasets, where the image-text pre-training baseline is based on the same backbone as MMTF. AUROC is the evaluation metric. The best results are shown in bold. Table 2 demonstrates the consistent advantage of MMTF over the baseline model.

[0139] Table 2 Comparison between the present invention and the image-text pre-training baseline

[0140]

[0141]

[0142] Figure 5 It is a comparison between the MMTF model proposed by the present invention and the SOTA medical image-text pre-training model in the abnormal detection of NIH ChestX-ray and VinBigData chest X-rays. The AUROC scores for the multi-label classification task on the NIH ChestX-ray dataset are shown in (a), and the AUROCs scores for multi-label classification on the VinBigData dataset are shown in (b). The optimal results are used as the upper limit for each category in the radar chart.

[0143] As Figure 5 shown, MMTF outperforms existing image-text pre-training models in almost all chest diseases. Under constrained supervision, this performance is particularly remarkable. For example, when only 1% of the labeled data is used on the NIH ChestX-ray dataset, MMTF's performance is significantly better than that of MuSE by at least 15%. This huge advantage can be attributed to MMTF's multi-scale representation learning, which enhances its generalization ability between tasks and datasets, especially in cases where labeled data is scarce.

[0144] Note that both MMTF and MuSE adopt channel-based multimodal information exchange. Although MuSE is effective in entity recognition and sentiment analysis tasks, it is not optimized for complex medical image analysis, where capturing fine-grained hierarchical features is crucial. MMTF extends the potential of channel-based fusion by incorporating multi-scale learning, making it perform better in medical tasks such as identifying consolidation, lung opacity, and pneumothorax. For example, the AUROC scores of MMTF in these cases are 97.3%, 95.7%, and 97.9% respectively, exceeding other models such as MedKLIP and ConVIRT.

[0145] In addition, MMTF shows greater advantages on the VinBigData dataset, performing well in differentiating viral and bacterial pneumonia (a 9% improvement over the baseline) and diagnosing pneumonia caused by the novel coronavirus (nearly 3% higher than MedKLIP). These findings are not only of practical significance but also statistically supported. The p-value analysis emphasizes that the performance improvement of MMTF is extremely unlikely to be accidental, with a p-value as low as 3.25×10 -4 , when using a dataset with 10% labeled data, demonstrating the robustness of the model under different regulatory levels. Even when limited to 1% of the labeled data, the benefits of MMTF are still statistically significant, confirming that its multimodal fusion and multi-scale learning methods provide specific improvements in medical image analysis.

[0146] In the embodiments, the proposed MMTF model is also compared with other state-of-the-art multi-modal pulmonary disease diagnosis methods, including MultiFusionNet, Late Fusion, and CheXMed. All three of these methods involve inputting multimodal data, including electronic medical records and chest X-rays. Notably, for fairness, in subsequent experiments, all of these models were first pre-trained on the MIMIC-CXR dataset and then fine-tuned on the downstream datasets.

[0147] Figure 6 are the comparison results between the proposed MMTF and SOTA multi-modal lung diagnosis on the NIH ChestX-ray, VinBigData ChestXX-ray anomaly detection, and Shenzhen tuberculosis datasets.

[0148] As Figure 6As shown, the MMTF model consistently outperforms other competing methods in terms of generalization and diagnostic accuracy, especially when dealing with scarce labeled data. For example, MultiFusionNet uses a multi-layer fusion technique to combine features from different convolutional layers, improving classification by leveraging different feature maps from deeper layers. However, MMTF achieved an accuracy of 76.4% on the NIH dataset with only 1% labeled data, significantly higher than 68.9% of MultiFusionNet, indicating that MMTF has stronger generalization ability under limited supervision. Similarly, Late Fusion integrates multimodal data such as clinical records and chest X-rays and has some advantages in a multimodal environment but performs poorly in more complex tasks. On the SZ and C-T2 datasets, the accuracies of MMTF were 97.8% and 92.9% respectively, superior to 92.8% and 70.2% of Late Fusion, highlighting MMTF's ability to capture hierarchical features crucial for precise medical image analysis. CheXMed combines clinical records with chest X-ray data and performs well in specific areas such as age-based pneumonia detection but is generally still lacking. On the NIH dataset with 1% labeled data, the score of CheXMed was only 66.8%, while that of MMTF reached 76.4%. MMTF's multi-scale fusion mechanism enables it to capture finer details of lung pathologies such as pneumothorax and pulmonary opacity, resulting in more accurate diagnoses on multiple datasets. Overall, MMTF's advanced multimodal fusion, multi-scale aggregation, and Transformer-based feature extraction enable it to more effectively handle the complexity of medical data and have higher robustness and diagnostic accuracy compared to its peers.

[0149] MMTF provides reliable decision support for clinicians by accurately identifying and localizing lung lesions. In addition to diagnostic accuracy, interpretability plays a crucial role in artificial intelligence-assisted diagnosis. For this purpose, interpretability analysis was conducted using the NIH ChestX-ray and VinBigData chest X-ray datasets to evaluate MMTF's ability to detect lesion areas.

[0150] To make the decision-making process of the model more transparent, the cross-attention maps of each transformation layer in the multi-scale fusion module were averaged and visualized.

[0151] Figure 7 These are the visualization results of randomly selected samples. Figure 7 Figure shows eight randomly sampled images from the NIH ChestX-ray dataset (a - d) and VinBigData (e - h). On the left side of the eight randomly sampled images, the original images are compared with the attention Figure 1Starting display. In each original image, the red box highlights the lesion area annotated by the radiologist. The attention map uses a color spectrum from red to blue overlaid on the original image, where red represents high-attention areas and blue represents low-attention areas.

[0152] Through Figure 7 It can be seen that MMTF effectively integrates information from different regions and is closely related to the pathological location determined by the radiologist. It is worth noting that the multi-scale fusion mechanism allows MMTF to detect lung lesions at various scales, ensuring accuracy even for smaller lesions (e.g., the first image in NIH ChestX-ray, Figure 7 a).

[0153] In addition, to evaluate the transfer learning ability of the model on a smaller dataset, t-distributed stochastic neighbor embedding (t-SNE) was applied and visualized on the Shenzhen tuberculosis dataset. Figure 8 Figures a-c in it show the t-SNE visual clustering results of the Shenzhen tuberculosis and the dataset of image data collection of pneumonia caused by the novel coronavirus. Figure 8 The clustering results in it show that MMTF can accurately capture the local and global structures of the data. Whether it is the relatively simple tuberculosis classification task or the more challenging task of differentiating pneumonia caused by the novel coronavirus from other types of pneumonia, MMTF demonstrates strong discriminative ability. This highlights the robustness of the model, especially in identifying viral and bacterial pneumonia cases - a key challenge in the real-world clinical setting.

[0154] Therefore, from the above experiments, it can be seen that MMTF not only performs excellently in terms of diagnostic accuracy but also improves interpretability, providing practical clinical value by offering more transparent and reliable diagnostic insights. This feature enhances its practicality for clinicians and increases trust in AI-driven medical solutions.

[0155] Thorough ablation experiments were conducted on MMTF by removing or replacing individual modules. Table 3 shows the quantitative results of the ablation experiments. In Table 3, Task1 represents the generation task of generating images according to the text; Task2 represents the generation task of generating text according to the image; L-encoder represents the large-branch image encoder; S-encoder represents the small-branch image encoder; Channel-based module represents the channel-based information exchange module. The checkmark in Table 3 indicates the presence of this component or unit or module, and the absence of a checkmark means that this component or unit or module has been removed. The results show that the proposed module can effectively solve the previous problem of poor multimodal representation.

[0156] Table 3 Ablation study of MMTF by removing or replacing individual modules

[0157]

[0158] First, the impact on prediction performance was examined by replacing two generation tasks in the inspection order (rows 0 and 1 in Table 3). When only image generation (i.e., Task1) or text generation (i.e., Task2) was retained in the model of the present invention, the overall performance of MMTF on the Vinbigdata dataset decreased by 1.28% and 1.09% respectively (compared with row 6). Notably, removing the generation tasks led to a 2.38% decrease in the overall model performance. This indicates that reducing the generation tasks used in the model improved the loss of fine-grained semantics caused by modality fusion.

[0159] In addition to the generation tasks within the model, the effect of the multi-scale feature fusion module was also evaluated. First, the multi-scale fusion module was removed to study the impact of the single-branch image encoder on the model performance. The results showed that in the single-scale case (rows 3 and 4 in Table 3), compared with the double-branch model (row 6 in Table 3), the model prediction performance decreased significantly (6.77% and 7.31%). This result implies that single-scale multimodal fusion cannot effectively represent the relationship between lesions and text descriptions in different image regions. Next, the channel-based information exchange module was removed to verify the effectiveness of channel-based multimodal information exchange. The results showed that multimodal information fusion directly between multiple scales after encoding was 2.84% lower in AUROC than the original model. This indicates that channel-based inter-modal information exchange is necessary before performing multi-scale information fusion.

[0160] The MMTF model proposed by the present invention has great potential in clinical applications, especially in environments with limited resources and scarce annotated data. By leveraging existing free text radiology reports, MMTF reduces the dependence on a large number of manual labels and promotes a more efficient diagnostic process. This is particularly important in real-world clinical settings where the availability of fully annotated datasets is often limited.

[0161] In addition, the strong performance of MMTF in various datasets (such as NIH ChestX-ray, VinBigData, and Shenzhen tuberculosis) highlights its adaptability to different clinical tasks, from common diseases such as tuberculosis to more complex diseases such as pneumonia caused by the novel coronavirus and pneumonia. This indicates that the model can provide reliable decision support in various medical environments, improving diagnostic accuracy and reducing the workload of clinicians.

[0162] In addition, MMTF integrates multi-scale feature fusion and auxiliary generation tasks, enabling clinicians to better visualize and understand the model's decision-making process, thereby improving interpretability. This function is highly consistent with the growing demand for transparent artificial intelligence systems in the healthcare field.

[0163] Finally, the model's ability to achieve high performance under limited supervision and its ability to capture fine-grained semantic details indicate that it can be applied to other medical imaging tasks, including segmentation and reconstruction, making it a valuable tool for a wide range of medical applications.

[0164] A novel Multimodal Multi-scale Transform Fusion (MMTF) method proposed in this invention can be used to analyze medical images and corresponding reports. MMTF integrates multi-modal, multi-stage, and multi-level fine-grained features to address the limitations of cross-modal fusion of heterogeneous medical data. Detailed fusion is achieved by exchanging token embeddings between image and text modalities, while a cross-attention-based multi-scale fusion module enables effective information exchange between branches. In addition, two generation tasks capture the correlation between text and image, narrowing the gap in cross-modal attention representations. Experimental results show that MMTF achieves state-of-the-art performance on several benchmark lung disease datasets, and interpretability analysis further confirms the model's ability to accurately capture lung abnormalities. Compared with existing methods, good performance is obtained on four public datasets, supporting the effectiveness of the model in diagnosing lung abnormalities.

[0165] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the present technical solution, and they should all be covered by the scope of the claims of the present invention.

Claims

1. A multi-modal multi-scale transformation fusion method based on exchange, characterized in that Including: Obtain multimodal data, that is, original image data and original text data corresponding to the original image data; Input the original image data and the original text data into a multimodal multi-scale transformation fusion model to generate a fusion result, where the multimodal multi-scale transformation fusion model includes: a multimodal encoder module, a decoder module, a channel-based information exchange module, and a multi-scale fusion module. The multimodal encoder module includes a text encoder and a dual-branch image encoder. The text encoder is used to encode the input text data to generate a text embedding. The dual-branch image encoder includes a small-branch image encoder and a large-branch image encoder, which are respectively used to extract fine-grained image embeddings and coarse-grained image embeddings of the input image data. The decoder module contains two decoders, which respectively decode the embeddings generated by the text encoder and the dual-branch image encoder, so as to perform two generation tasks of generating an image according to the text and generating text according to the image, in order to regularize the embeddings generated by the text encoder and the dual-branch image encoder. The channel-based information exchange module is used to perform information exchange on the embeddings of different modalities on different channels according to the attention scores of different channels, so as to generate image features and text features on different branches. The multi-scale fusion module is used to fuse the cls token from one branch and the patch token from another branch according to the image features and text features on different branches.

2. The multi-modal multi-scale transform fusion method based on exchange according to claim 1, wherein Performing information exchange on the embeddings of different modalities includes information exchange between the coarse-grained image embedding and the text embedding and information exchange between the fine-grained image embedding and the text embedding. The channel-based information exchange module also uses two hyperparameters η and μ to respectively control the start layer and the end layer of the multimodal information exchange.

3. The multi-modal multi-scale transformation fusion method based on exchange according to claim 1, wherein The channel-based information exchange module includes a multi-head self-attention unit and a multimodal information exchange unit. After the text embedding, the fine-grained image embedding, and the coarse-grained image embedding are added with their respective corresponding cls tokens, they are fed into the multi-head self-attention unit to obtain the attention scores of each token in the text embedding, the fine-grained image embedding, and the coarse-grained image embedding. The calculated attention scores are sent to the multimodal information exchange unit through a feed-forward network with residual connections. The multimodal information exchange unit adopts multiple Transformer encoders with shared parameters to exchange multimodal information at different scales in a complementary manner.

4. The multi-modal multi-scale transformation fusion method based on exchange according to claim 3, characterized in that The multimodal information exchange unit exchanges multimodal information at different scales in a complementary manner, including: for information exchange at any layer, for any channel, if the attention score of the current channel is lower than a set threshold, the embedding vector of this channel will be replaced by the average embedding of a preset percentage of tokens in another modality, otherwise, the embedding vector of this channel remains unchanged, so as to obtain an updated embedding matrix.

5. The multi-modal multi-scale transformation fusion method based on exchange according to claim 4, wherein The channel-based information exchange module also includes an FFN and a CB unit. After the information exchange at the current layer is completed, the FFN generates the input embedding for the next layer of information exchange according to the updated embedding matrix, and the CB is used to broadcast the context to each token.

6. The multi-modal multi-scale transformation fusion method based on exchange according to claim 4, characterized in that Fuse the cls token from one branch and the patch tokens from another branch, specifically including: Perform a concatenation operation on the image cls token from the large branch and the text patch tokens from the small branch, and perform cross-attention with the image cls token of the large branch. Then, concatenate the patch tokens in the result of the cross-attention with the image cls token of the large branch to obtain the fused image features of the large branch, that is, the coarse-grained fused image features; Perform a concatenation operation on the image cls token from the small branch and the text patch tokens from the large branch, and perform cross-attention with the image cls token of the small branch. Then, concatenate the patch tokens in the result of the cross-attention with the image cls token of the small branch to obtain the fused image features of the small branch, that is, the fine-grained fused image features; Perform a concatenation operation on the text cls token from the large branch and the image patch tokens from the small branch, and perform cross-attention with the text cls token of the large branch. Then, concatenate the patch tokens in the result of the cross-attention with the text cls token of the large branch to obtain the fused text features of the large branch, that is, the coarse-grained fused text features; Perform a concatenation operation on the text cls token from the small branch and the image patch tokens from the large branch, and perform cross-attention with the text cls token of the small branch. Then, concatenate the patch tokens in the result of the cross-attention with the text cls token of the small branch to obtain the fused text features of the small branch, that is, the fine-grained fused text features.

7. A multi-modal multi-scale transformation fusion system based on exchange, characterized in that Including: A multimodal data acquisition engine for acquiring multimodal data, that is, image data and text data corresponding to the image data; A model training engine for training the multimodal multi-scale transformation fusion model, A fusion generation engine for generating a fusion result according to the acquired multimodal data by using the trained multimodal multi-scale transformation fusion model, Among them, the multimodal multi-scale transformation fusion model includes: a multimodal encoder module, a decoder module, a channel-based information exchange module, and a multi-scale fusion module, The multimodal encoder module includes a text encoder and a two-branch image encoder, The text encoder is used to encode the input text data to generate text embeddings; The two-branch image encoder includes a small-branch image encoder and a large-branch image encoder, which are respectively used to extract the fine-grained image embeddings and coarse-grained image embeddings of the input image data; The decoder module contains two decoders, which respectively decode the embeddings generated by the text encoder and the two-branch image encoder, so as to perform two generation tasks of generating an image according to the text and generating text according to the image, in order to regularize the embeddings generated by the text encoder and the two-branch image encoder; The channel-based information exchange module is used to exchange information of the embeddings of different modalities on different channels according to the attention scores of different channels, so as to generate image features and text features on different branches; A multi-scale fusion module is used to fuse the cls tokens from one branch and the patch tokens from another branch based on the image features and text features on different branches.

8. The exchange-based multi-modal multi-scale transformation fusion system according to claim 7, characterized in that, The channel-based information exchange module includes a multi-head self-attention unit and a multi-modal information exchange unit. After text embeddings, fine-grained image embeddings, and coarse-grained image embeddings are added with their respective corresponding cls tokens, they are fed into the multi-head self-attention unit to obtain the attention scores of each token in the text embeddings, fine-grained image embeddings, and coarse-grained image embeddings. The calculated attention scores are sent to the multi-modal information exchange unit through a feed-forward network with residual connections. The multi-modal information exchange unit adopts multiple Transformer encoders with shared parameters to exchange multi-modal information at different scales in a complementary manner.

9. The exchange-based multi-modal multi-scale transformation fusion system according to claim 8, wherein The multi-modal information exchange unit exchanges multi-modal information at different scales in a complementary manner, including: for the information exchange of any layer, for any channel, if the attention score of the current channel is lower than the set threshold, the embedding vector of this channel will be replaced by the average embedding of a preset percentage of tokens in another modality, otherwise, the embedding vector of this channel remains unchanged.

10. The exchange-based multi-modal multi-scale transformation fusion system according to claim 9, wherein The multi-scale fusion module performs the following operations: Perform a concatenation operation on the image cls token from the large branch and the text patch token from the small branch, perform cross-attention with the image cls token from the large branch, and then perform a concatenation operation on the patch token in the result of the cross-attention with the image cls token from the large branch, so as to obtain the fused image feature of the large branch, that is, the coarse-grained fused image feature; Perform a concatenation operation on the image cls token from the small branch and the text patch token from the large branch, perform cross-attention with the image cls token from the small branch, and then perform a concatenation operation on the patch token in the result of the cross-attention with the image cls token from the small branch, so as to obtain the fused image feature of the small branch, that is, the fine-grained fused image feature; Perform a concatenation operation on the text cls token from the large branch and the image patch token from the small branch, perform cross-attention with the text cls token from the large branch, and then perform a concatenation operation on the patch token in the result of the cross-attention with the text cls token from the large branch, so as to obtain the fused text feature of the large branch, that is, the coarse-grained fused text feature; Perform a concatenation operation on the text cls token from the small branch and the image patch token from the large branch, perform cross-attention with the text cls token from the small branch, and then perform a concatenation operation on the patch token in the result of the cross-attention with the text cls token from the small branch, so as to obtain the fused text feature of the small branch, that is, the fine-grained fused text feature.

Citation Information

Patent Citations

  • Multimodal medical image fusion method based on global information fusion

    CN114565816A

  • Multi-spectral pedestrian detection method based on cross Transform fusion

    CN116580425A