A multimodal data fusion classification method based on large model and attention mechanism

By introducing the sliding window cross attention fusion module into the multimodal data fusion classification method, the problem of over-fusion of image features is solved, and the accuracy and efficiency of classification results are improved.

CN118823528BActive Publication Date: 2025-05-16JIANGNAN UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410791163.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-19
Publication Date
2025-05-16
Estimated Expiration
2044-06-19

AI Technical Summary

Technical Problem

The classification method based on multimodal data fusion in the prior art has problems with excessive fusion of image features of the target object, resulting in low classification efficiency and classification results accuracy.

Method used

The multimodal data fusion classification method based on large models and attention mechanisms is adopted to feature fusion of different image features through the sliding window cross attention fusion module, avoiding excessive fusion and fully fusing feature information between different images.

Benefits of technology

Improve the accuracy of classification results, reduce information redundancy and noise, and enable the model to better balance information between text modes and image modes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118823528B_ABST
    Figure CN118823528B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of computer technology, and relates to a multimodal data fusion classification method based on a large model and an attention mechanism; a first image feature vector and a second image feature vector are input into a sliding window cross-attention fusion module in a classification model, and a first target image feature vector and a second target image feature vector are output; the first target image feature vector, the second target image feature vector and the text feature vector are input into a heterogeneous data cross-attention fusion module in a classification model, and a target feature vector of a target object is output; the target feature vector of a target object is input into a fully connected layer in a classification model, and a classification result of the target object is output. The present application directly fuses different image features, which not only fuses the feature information between different images, but also avoids the risk of overfitting caused by excessive fusion, reduces information redundancy and noise, can better balance text modality and image modality, and improves the accuracy of classification results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a multimodal data fusion classification method, device and computer-readable storage medium based on a large model and an attention mechanism. Background Art

[0002] Classification methods based on deep learning are widely used in various fields, such as face recognition, disease classification, etc. However, in practical applications, relying solely on single modal information to classify target objects may fail to fully obtain the feature information of the target objects, resulting in inaccurate classification results. Therefore, existing target object classification methods often perform decision-level fusion or feature-level fusion of different modal information of the target objects to fully obtain the feature information of the target objects, thereby improving the accuracy of the classification results.

[0003] Among them, decision-level fusion refers to the fusion of classification results from different modalities. Common fusion methods include voting and weighted averaging. For example, in face recognition tasks, face recognition results from different modalities are voted on, and the result with the most votes is selected as the final recognition result. However, simply integrating the classification results of different modalities cannot truly explore the correlation between different modal information of the target object. Even when there is noise or error in the classification result of a single modality, it will introduce noise interference and affect the accuracy of the classification result.

[0004] Based on this, most classification methods currently use feature-level fusion. Feature-level fusion refers to the fusion of information from different data during the feature extraction process. A common fusion method first obtains text data and multiple image data of the target object, performs feature fusion on multiple image data in pairs to obtain fusion features containing information from two images, and finally fuses all fusion features again to obtain target image features of the target object, and fuses the target image features with text features to obtain feature representations containing information from different modalities, thereby classifying the target object. However, when the prior art performs feature fusion on any two image data, it is often necessary to perform multi-level feature extraction on the two image data, fuse the features of the two image data after each level of feature extraction, thereby obtaining multiple fused image features, and finally fuse the multiple fused image features again to obtain fusion features containing information from two images. This method has the problem of over-fusion, which not only leads to information redundancy, increases computational complexity, and reduces classification efficiency, but also causes the model to over-focus on the image features of the target object due to the over-fusion of image features, and cannot correctly balance the information of the text modality and the image modality, thereby reducing the accuracy of the classification results. Summary of the invention

[0005] To this end, the technical problem to be solved by the present invention is to overcome the problem that the classification method based on multimodal data fusion in the prior art over-fuses the image features of the target object, thereby resulting in low classification efficiency and accuracy of the classification results.

[0006] To solve the above technical problems, the present invention provides a multimodal data fusion classification method based on a large model and an attention mechanism, comprising:

[0007] Obtaining a first image feature vector, a second image feature vector, and a text feature vector of a target object;

[0008] Inputting the first image feature vector and the second image feature vector into a sliding window cross attention fusion module in a trained classification model, and outputting a first target image feature vector and a second target image feature vector, which specifically includes:

[0009] The first image feature vector and the second image feature vector are respectively inputted into a first normalization submodule, a window multi-head self-attention submodule and a first feedforward neural network submodule which are sequentially connected in series along the forward propagation direction, and a first intermediate image feature vector and a second intermediate image feature vector are outputted;

[0010] The query vector of the first intermediate image feature vector, the key vector and the eigenvalue vector of the second intermediate image feature vector are simultaneously input into the second normalization submodule, the sliding window multi-head cross attention submodule and the second feedforward neural network submodule which are sequentially connected in series along the forward propagation direction, and the first target image feature vector is output;

[0011] The query vector of the second intermediate image feature vector, the key vector and the eigenvalue vector of the first intermediate image feature vector are simultaneously input into the second normalization submodule, the sliding window multi-head cross attention submodule and the second feedforward neural network submodule which are sequentially connected in series along the forward propagation direction, and the second target image feature vector is output;

[0012] Inputting the first target image feature vector, the second target image feature vector and the text feature vector into a heterogeneous data cross-attention fusion module in a trained classification model, and outputting a target feature vector of a target object;

[0013] The target feature vector of the target object is input into the fully connected layer in the trained classification model, and the classification result of the target object is output.

[0014] Preferably, the first feedforward neural network submodule and the second feedforward neural network submodule both include a normalization unit and a nonlinear transformation unit connected in series.

[0015] Preferably, the first intermediate image feature vector is expressed as:

[0016]

[0017]

[0018] Among them, f′ A represents the first intermediate image feature vector, FFN1 represents the first feedforward neural network submodule, LN represents the normalization operation, WMSA represents the window multi-head self-attention submodule, and f A represents the first image feature vector;

[0019] The second intermediate image feature vector is expressed as:

[0020]

[0021]

[0022] Among them, f′ B represents the second intermediate image feature vector, f B represents the second image feature vector.

[0023] Preferably, the first target image feature vector is expressed as:

[0024] F A =FFN2(LN(f″) A ))+f″ A ,

[0025] f″ A =SWMCA(LN(Q A , K B , V B ))+f′ A ,

[0026] Among them, F A represents the first target image feature vector, FFN2 represents the second feedforward neural network submodule, LN represents the normalization operation, SWMCA represents the sliding window multi-head cross attention submodule, Q A The query vector K represents the first intermediate image feature vector. B The key vector representing the feature vector of the second intermediate image, V B The eigenvalue vector representing the eigenvector of the second intermediate image, f′ A represents the first intermediate image feature vector;

[0027] The second target image feature vector is expressed as:

[0028] F B =FFN2(LN(f″) B ))+f″ B ,

[0029] f″ B =SWMCA(LN(Q B ,K A ,V A ))+f′ B ,

[0030] Among them, F B represents the second target image feature vector, Q B The query vector K represents the feature vector of the second intermediate image. A The key vector representing the first intermediate image feature vector, V A The eigenvalue vector representing the first intermediate image eigenvector, f′ B Represents the second intermediate image feature vector.

[0031] Preferably, inputting the first target image feature vector, the second target image feature vector and the text feature vector into a heterogeneous data cross attention fusion module in a trained classification model, and outputting a target feature representation of the target object comprises:

[0032] Inputting the key vector of the first target image feature vector and the query vector of the text feature vector into a first matrix cross multiplication submodule for matrix cross multiplication, and outputting a first heterogeneous data fusion feature vector;

[0033] Normalizing the first heterogeneous data fusion feature vector, and inputting the normalized first heterogeneous data fusion feature vector and the eigenvalue vector of the first target image feature vector into a second matrix cross multiplication submodule for matrix cross multiplication, and outputting a second heterogeneous data fusion feature vector;

[0034] Inputting the query vector of the second heterogeneous data fusion feature vector and the text feature vector into a first residual connection submodule, and outputting a first target fusion feature vector;

[0035] Inputting the key vector of the second target image feature vector and the query vector of the text feature vector into a third matrix cross multiplication submodule for matrix cross multiplication, and outputting a third heterogeneous data fusion feature vector;

[0036] Normalizing the third heterogeneous data fusion feature vector, and inputting the normalized third heterogeneous data fusion feature vector and the eigenvalue vector of the second target image feature vector into a fourth matrix cross multiplication submodule for matrix cross multiplication, and outputting a fourth heterogeneous data fusion feature vector;

[0037] Inputting the query vector of the fourth heterogeneous data fusion feature vector and the text feature vector into a second residual connection submodule, and outputting a second target fusion feature vector;

[0038] The first target fusion feature vector and the second target fusion feature vector are concatenated to obtain a target feature vector of the target object.

[0039] Preferably, the first target fusion feature vector is expressed as:

[0040]

[0041]

[0042] Among them, Z A represents the first target fusion feature vector, f cls represents the text feature vector, g 1 (·) represents the projection function used for eigenvector alignment, represents the second heterogeneous data fusion feature vector, MCA represents the cross attention operation, LN represents the normalization operation, represents the first heterogeneous data fusion feature vector, F A represents the first target feature vector;

[0043] The second target fusion feature vector is expressed as:

[0044]

[0045]

[0046] Among them, Z B represents the second target fusion feature vector, represents the fourth heterogeneous data fusion feature vector, represents the third heterogeneous data fusion feature vector, F B represents the second target feature vector;

[0047] The target feature vector of the target object is expressed as:

[0048] F = concat(Z A , Z B ),

[0049] Wherein, F represents the target feature vector of the target object.

[0050] Preferably, obtaining the first image feature vector, the second image feature vector and the text feature vector of the target object includes:

[0051] Acquire first image data, second image data, and text data of a target object;

[0052] Inputting the first image data and the second image data into a visual feature extractor for feature extraction, and outputting a first image feature vector and a second image feature vector;

[0053] Inputting the text data into a text feature extractor for feature extraction, and outputting a text feature vector;

[0054] Among them, the visual feature extractor includes a patch embedding module, a first EfficientViT Block module, a first downsampling module, a second EfficientViT Block module, a second downsampling module and a third EfficientViT Block module which are connected in series in the forward propagation direction. The first EfficientViT Block module, the second EfficientViTBlock module and the third EfficientViT Block module each include a first deep convolution submodule, a third feedforward neural network submodule, a cascaded group attention submodule, a second deep convolution submodule and a fourth feedforward neural network submodule which are connected in series in the forward propagation direction.

[0055] Preferably, the first image feature vector is expressed as:

[0056] f A =EfficientViT(Embed(Modality A )),

[0057] Among them, f A represents the first image feature vector, EfficientViT represents the visual feature extractor, Embed represents embedding, Modality A represents first image data;

[0058] The second image feature vector is expressed as:

[0059] f B =EfficientViT(Embed(Modality B )),

[0060] Among them, f B Represents the second image feature vector, Modality B represents second image data;

[0061] The text feature vector is expressed as:

[0062] f cls =Pooling(BERT(Modality T )),

[0063] Among them, fcls represents a text feature vector, Pooling represents a pooling operation, BERT represents a text feature extractor, Modality T Represents text data.

[0064] The present invention also provides a multimodal data fusion classification device based on a large model and an attention mechanism, comprising:

[0065] A feature acquisition module, used to acquire a first image feature vector, a second image feature vector and a text feature vector of a target object;

[0066] The first feature fusion module is used to input the first image feature vector and the second image feature vector into the sliding window cross attention fusion module in the trained classification model, and output the first target image feature vector and the second target image feature vector, which specifically includes:

[0067] An intermediate image feature vector acquisition submodule, used to input the first image feature vector and the second image feature vector respectively into a first normalization submodule, a window multi-head self-attention submodule and a first feedforward neural network submodule which are sequentially connected in series along a forward propagation direction, and output a first intermediate image feature vector and a second intermediate image feature vector;

[0068] A first target image feature vector acquisition submodule is used to simultaneously input the query vector of the first intermediate image feature vector, the key vector and the eigenvalue vector of the second intermediate image feature vector to the second normalization submodule, the sliding window multi-head cross attention submodule and the second feedforward neural network submodule which are sequentially connected in series along the forward propagation direction, and output the first target image feature vector;

[0069] A second target image feature vector acquisition submodule is used to simultaneously input the query vector of the second intermediate image feature vector, the key vector and the eigenvalue vector of the first intermediate image feature vector to the second normalization submodule, the sliding window multi-head cross attention submodule and the second feedforward neural network submodule which are sequentially connected in series along the forward propagation direction, and output the second target image feature vector;

[0070] A second feature fusion module, configured to input the first target image feature vector, the second target image feature vector and the text feature vector into a heterogeneous data cross attention fusion module in a trained classification model, and output a target feature representation of a target object;

[0071] The classification module is used to input the target feature representation of the target object into the fully connected layer in the trained classification model and output the classification result of the target object.

[0072] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned multimodal data fusion classification method based on a large model and an attention mechanism are implemented.

[0073] The multimodal data fusion classification method based on a large model and an attention mechanism provided in the present application utilizes a window multi-head self-attention submodule to capture the local features and global features of an image feature vector, and then utilizes a feedforward neural network submodule to extract useful feature information from the output of the window multi-head self-attention submodule, thereby obtaining a first intermediate image feature vector and a second intermediate image feature vector; then, the query vector of an intermediate image feature vector and the key vector and eigenvalue vector of another intermediate image feature vector are input into a sliding window multi-head cross-attention submodule to focus on the correlation between different image features and capture the dependencies in different image features, and the feedforward neural network submodule is used again to perform feature extraction on the output of the sliding window multi-head cross-attention submodule, thereby outputting a first target image feature vector and a second target image feature vector fused with different image features; different from the multi-stage fusion method in the prior art, the sliding window cross-attention fusion module provided in the present application directly performs feature fusion on the extracted different image features, which not only fully integrates the feature information between different images, but also avoids the risk of overfitting caused by excessive fusion, reduces information redundancy and noise, and enables the model to better balance the information of the text modality and the image modality, thereby improving the accuracy of the classification results. BRIEF DESCRIPTION OF THE DRAWINGS

[0074] In order to make the content of the present invention more clearly understood, the present invention is further described in detail below according to specific embodiments of the present invention in conjunction with the accompanying drawings, wherein:

[0075] Figure 1 Flowchart of the multimodal data fusion classification method based on large model and attention mechanism provided for this application;

[0076] Figure 2 A schematic diagram of the structure of the visual feature extractor provided for this application;

[0077] Figure 3 A schematic diagram of the structure of the sliding window cross attention fusion module provided for this application;

[0078] Figure 4 Schematic diagram of the heterogeneous data cross-attention fusion module structure provided for this application;

[0079] Figure 5 Schematic diagram of the framework of the multimodal data fusion classification method based on a large model and attention mechanism provided for this application;

[0080] Figure 6 The radar chart is a comparison of the indicators of different methods on different classification tasks; Figure 6 (a) is a schematic diagram of the comparison results of different methods in the BMV classification task. Figure 6 (b) is a schematic diagram of the comparison results of indicators of different methods in the DAG classification task. Figure 6 (c) is a schematic diagram of the comparison results of different methods in the DIAG classification task. Figure 6 (d) is a schematic diagram of the comparison results of indicators of different methods in the PIG classification task. Figure 6 (e) is a schematic diagram of the comparison results of indicators of different methods in the PN classification task. Figure 6 (f) is a schematic diagram of the comparison results of indicators of different methods in RS classification tasks. Figure 6 (g) is a schematic diagram of the comparison results of indicators of different methods in the STR classification task. Figure 6 (h) is a schematic diagram of the comparison results of indicators of different methods in VS classification tasks;

[0081] Figure 7 Schematic diagram of the structure of the multimodal data fusion classification device based on a large model and attention mechanism provided in this application. DETAILED DESCRIPTION

[0082] The present invention is further described below in conjunction with the accompanying drawings and specific embodiments so that those skilled in the art can better understand the present invention and implement it, but the embodiments are not intended to limit the present invention.

[0083] See also Figure 1 , Figure 1 A flow chart of a multimodal data fusion classification method based on a large model and an attention mechanism provided in this application, the method specifically includes:

[0084] S10: Acquire a first image feature vector, a second image feature vector, and a text feature vector of the target object;

[0085] Specifically, the first image feature vector and the second image feature vector of the target object are image features obtained by extracting features from different image data of the target object; for example, if the classification task is skin disease classification, the first image feature vector is a feature vector obtained by extracting features from a dermoscopic image, and the second image feature vector is a feature vector obtained by extracting features from a clinical image;

[0086] S20: Inputting the first image feature vector and the second image feature vector into the sliding window cross attention fusion module in the trained classification model, and outputting the first target image feature vector and the second target image feature vector, which specifically includes:

[0087] S200: Input the first image feature vector and the second image feature vector into the first normalization submodule, the window multi-head self-attention submodule and the first feedforward neural network submodule connected in series in the forward propagation direction, and output the first intermediate image feature vector and the second intermediate image feature vector.

[0088] vector;

[0089] S201: inputting the query vector of the first intermediate image feature vector, the key vector and the eigenvalue vector of the second intermediate image feature vector simultaneously into the second normalization submodule, the sliding window multi-head cross attention submodule and the second feedforward neural network submodule which are sequentially connected in series along the forward propagation direction, and outputting the first target image feature vector;

[0090] S202: inputting the query vector of the second intermediate image feature vector, the key vector and the eigenvalue vector of the first intermediate image feature vector simultaneously into a second normalization submodule, a sliding window multi-head cross attention submodule and a second feedforward neural network submodule which are sequentially connected in series along the forward propagation direction, and outputting a second target image feature vector;

[0091] Specifically, the first feedforward neural network submodule and the second feedforward neural network submodule each include a normalization unit and a nonlinear transformation unit connected in series;

[0092] S30: inputting the first target image feature vector, the second target image feature vector and the text feature vector into a heterogeneous data cross attention fusion module in the trained classification model, and outputting a target feature vector of the target object;

[0093] S40: Input the target feature vector of the target object into the fully connected layer in the trained classification model, and output the classification result of the target object.

[0094] The multimodal data fusion classification method based on a large model and an attention mechanism provided in the present application designs a sliding window cross-attention module to fuse different image features. The local features and global features of the image feature vector are first captured by the window multi-head self-attention sub-module, and then the useful feature information is extracted from the output of the window multi-head self-attention sub-module by using the feedforward neural network sub-module to obtain the first intermediate image feature vector and the second intermediate image feature vector; then, the query vector of one intermediate image feature vector and the key vector and eigenvalue vector of another intermediate image feature vector are input into the sliding window multi-head cross-attention sub-module to focus on the correlation between different image features and capture the dependency relationship in different image features, and the feedforward neural network sub-module is used again. The output of the sliding window multi-head cross-attention submodule is subjected to feature extraction, thereby outputting a first target image feature vector and a second target image feature vector which fuse different image features. Different from the multi-stage fusion method in the prior art, the sliding window cross-attention module designed in the present application directly fuses the extracted different image features, uses the window multi-head self-attention submodule to capture the local features and global features of a single image, and then uses the sliding window multi-head cross-attention submodule to capture the correlation between different images, which fully fuses the feature information between different images and avoids the risk of overfitting caused by excessive fusion, reduces information redundancy and noise, and enables the model to better balance the information of text modality and image modality, thereby improving the accuracy of classification results.

[0095] Specifically, the specific implementation steps of step S10 include:

[0096] S100: Acquire first image data, second image data and text data of a target object;

[0097] S102: Inputting the first image data and the second image data into a visual feature extractor for feature extraction, and outputting a first image feature vector and a second image feature vector;

[0098] The first image feature vector is expressed as:

[0099] f A =EfficientViT(Embed(Modality A )),

[0100] Among them, f A represents the first image feature vector, EfficientViT represents the visual feature extractor, Embed represents embedding, Modality A represents first image data;

[0101] The second image feature vector is expressed as:

[0102] fB =EfficientViT(Embed(Modality B )),

[0103] Among them, f B Represents the second image feature vector, Modality B represents second image data;

[0104] Specifically, the visual feature extractor includes a patch embedding module, a first EfficientViT Block module, a first downsampling module, a second EfficientViT Block module, a second downsampling module and a third EfficientViT Block module which are sequentially connected in series along the forward propagation direction, and the first EfficientViT Block module, the second EfficientViTBlock module and the third EfficientViT Block module each include a first deep convolution submodule, a third feedforward neural network submodule, a cascaded group attention submodule, a second deep convolution submodule and a fourth feedforward neural network submodule which are sequentially connected in series along the forward propagation direction;

[0105] For example, Figure 2 The figure shows a schematic diagram of the structure of a visual feature extractor provided in an embodiment of the present application. After the image is input into the visual feature extractor, it is first patch-embedded in sequence, and the 16*16 patch is embedded into the token of the specified dimension, and then enters the EfficientViT Block. In the Efficient Block, it first passes through DWConv (deep convolution) and FFN (feedforward neural network) in sequence, and then enters the calculation of cascade group attention, and then passes through DWConv (deep convolution) and FFN (feedforward neural network) in sequence to obtain the output. After the first EfficientViT Block module, it passes through a downsampling module, and the number of tokens is reduced by 4 times (2 times subsampling of the resolution), and then the downsampled features are input into the second EfficientViT Block module again, and the image feature vector is obtained based on the output of the third EfficientViT Block module.

[0106] S103: Input the text data into a text feature extractor for feature extraction, and output a text feature vector; the text feature vector is represented as:

[0107] f cls =Pooling(BERT(Modality T )),

[0108] Among them, f clsrepresents a text feature vector, Pooling represents a pooling operation, BERT represents a text feature extractor, Modality T Represents text data.

[0109] Specifically, most existing technologies use convolutional neural networks (CNNs) to extract image modal information and use one-hot encoding to process text modal information. However, convolutional neural networks capture local information through a fixed-size window and do not have the ability to process medium- and long-distance dependencies. Therefore, the features they extract lack globality and are not conducive to the interaction of multimodal information. The one-hot encoding method will encode each word in the text information independently, which cannot express the similarity and correlation between text information and cannot give full play to the role of text modality.

[0110] Therefore, in the embodiments of the present application, a large-model visual feature extractor EfficientViT and a text feature extractor BERT are used for feature extraction. Both EfficientViT and BERT have the ability to capture medium- and long-term relationship dependencies based on the attention mechanism, and can extract more effective image features and text features.

[0111] For example, Figure 3 The figure shows a schematic diagram of a sliding window cross attention fusion module provided by the present application. For a given input feature vector f of size H*W*C, A and f B , the window division mechanism divides it into non-overlapping M*M local windows, and then calculates the standard self-attention for each local window. Specifically, the first intermediate image feature vector is expressed as:

[0112]

[0113]

[0114] Among them, f′ A represents the first intermediate image feature vector, FFN1 represents the first feedforward neural network submodule, LN represents the normalization operation, WMSA represents the window multi-head self-attention submodule, and f A represents the first image feature vector;

[0115] The second intermediate image feature vector is expressed as:

[0116]

[0117]

[0118] Among them, f′ B represents the second intermediate image feature vector, f Brepresents the second image feature vector;

[0119] Furthermore, the conventional window partitioning method will cause information disconnection at the window boundary, so in the second half of the model, the window partition is moved to generate a new window, and the new window crosses the window boundary to provide a connection to better obtain features. In this application, a sliding window multi-head cross attention submodule is designed to fuse information from different images, such as Figure 3 As shown in , for the query vector of the first intermediate image feature vector, the cross-modal information is fused by focusing on the key vector and the eigenvalue vector of the second intermediate image feature vector, and the information of the first image data is retained through the residual connection to achieve global contextual interaction of different image data;

[0120] Specifically, the first target image feature vector is expressed as:

[0121] F A =FFN2(LN(f″) A ))+f″ A ,

[0122] f″ A =SWMCA(LN(Q A , K B , V B ))+f′ A ,

[0123] Among them, F A represents the first target image feature vector, FFN2 represents the second feedforward neural network submodule, LN represents the normalization operation, SWMCA represents the sliding window multi-head cross attention submodule, Q A The query vector K represents the first intermediate image feature vector. B The key vector representing the feature vector of the second intermediate image, V B The eigenvalue vector representing the eigenvector of the second intermediate image, f′ A represents the first intermediate image feature vector;

[0124] The second target image feature vector is expressed as:

[0125] F B =FFN2(LN(f″) B ))+f″ B ,

[0126] f″ B =SWMCA(LN(Q B , K A , V A ))+f′ B ,

[0127] Among them, FB represents the second target image feature vector, Q B The query vector K represents the feature vector of the second intermediate image. A The key vector representing the first intermediate image feature vector, V A The eigenvalue vector representing the first intermediate image eigenvector, f′ B Represents the second intermediate image feature vector.

[0128] Optionally, in some embodiments of the present application, the first target image feature vector, the second target image feature vector and the text feature vector may be directly weighted and fused to obtain the target feature vector of the target object; however, considering that this fusion method cannot capture the relationship between text data and different image data, based on this, another fusion method of image feature vector and text feature vector is provided in the embodiments of the present application, such as Figure 4 As shown in the figure, the text features are first fused with the two image features to obtain the feature Z with three mode information. A and Z B , and then Z A and Z B Connect the input to a fully connected layer or classifier for final decision making;

[0129] Specifically, the specific implementation steps of step S30 include:

[0130] S500: Inputting the key vector of the first target image feature vector and the query vector of the text feature vector into a first matrix cross multiplication submodule for matrix cross multiplication, and outputting a first heterogeneous data fusion feature vector;

[0131] S501: normalizing the first heterogeneous data fusion feature vector, and inputting the normalized first heterogeneous data fusion feature vector and the eigenvalue vector of the first target image feature vector into a second matrix cross multiplication submodule for matrix cross multiplication, and outputting a second heterogeneous data fusion feature vector;

[0132] S502: Inputting the query vector of the second heterogeneous data fusion feature vector and the text feature vector into the first residual connection submodule, and outputting a first target fusion feature vector;

[0133] S503: Inputting the key vector of the second target image feature vector and the query vector of the text feature vector into a third matrix cross multiplication submodule for matrix cross multiplication, and outputting a third heterogeneous data fusion feature vector;

[0134] S504: normalizing the third heterogeneous data fusion feature vector, and inputting the normalized third heterogeneous data fusion feature vector and the eigenvalue vector of the second target image feature vector into a fourth matrix cross multiplication submodule for matrix cross multiplication, and outputting a fourth heterogeneous data fusion feature vector;

[0135] S505: Inputting the query vector of the fourth heterogeneous data fusion feature vector and the text feature vector into the second residual connection submodule, and outputting a second target fusion feature vector;

[0136] S506: Concatenate the first target fusion feature vector and the second target fusion feature vector to obtain a target feature vector of the target object.

[0137] Specifically, the first target fusion feature vector is expressed as:

[0138]

[0139]

[0140] Among them, Z A represents the first target fusion feature vector, f cls represents the text feature vector, g 1 (-) indicates the projection function used for eigenvector alignment, represents the second heterogeneous data fusion feature vector, MCA represents the cross attention operation, LN represents the normalization operation, represents the first heterogeneous data fusion feature vector, F A represents the first target feature vector;

[0141] The second target fusion feature vector is expressed as:

[0142]

[0143]

[0144] Among them, Z B represents the second target fusion feature vector, represents the fourth heterogeneous data fusion feature vector, represents the third heterogeneous data fusion feature vector, F B represents the second target feature vector;

[0145] The target feature vector of the target object is expressed as:

[0146] F = concat(Z A ,Z B ),

[0147] Wherein, F represents the target feature vector of the target object.

[0148] For example, Figure 5Shown is a schematic diagram of the framework of a multimodal data fusion classification method based on a large model and an attention mechanism provided in an embodiment of the present application, which is specifically composed of a visual feature extractor, a text feature extractor, a sliding window cross-attention fusion module for image modality fusion, and a heterogeneous data cross-attention fusion module for image-text fusion.

[0149] In order to verify the effectiveness of the above method, the embodiment of the present application also experiments the above method in the field of dermatology in the medical field. The experiment uses the public multimodal Derm7pt dataset, which contains 1011 multimodal instances, where each case includes three data types: dermoscopic images, clinical images and text information, and the text information includes gender, symptom location, shape, treatment method and diagnostic difficulty level. The dataset is divided into a training set, a validation set and a test set, where the training set has 413 cases, the validation set has 203 cases, and the test set has 395 cases. In addition, the dataset is a multi-task dataset, which requires the classification tasks of eight labels at the same time, namely diagnosis (DIAG), bluewhitish Veil (BMV), dots and globules (DAG), pigment (PIG), pigment network (PN), regression structures (RS), streak (STR), and vascular structures (VS).

[0150] In order to estimate the experimental effect of this method, this application uses five commonly used evaluation indicators for classification tasks, namely accuracy, precision, sensitivity, specificity and F1-score, among which accuracy is the most important indicator:

[0151] Accuracy indicates the percentage of correct predictions in the total samples, and its calculation formula is as follows:

[0152]

[0153] Among them, TP means that the prediction is positive and the actual is positive, and the prediction is correct; FP means that the prediction is positive and the actual is negative, and the prediction is wrong; FN means that the prediction is negative and the actual is positive, and the prediction is wrong; TN means that the prediction is negative and the actual is negative, and the prediction is correct;

[0154] Precision represents the probability of actually being a positive sample among all samples predicted to be positive. Precision refers to the prediction results. It represents the accuracy of the prediction of the positive sample results. Its expression is:

[0155]

[0156] Sensitivity refers to the original sample, which means the probability of being predicted as a positive sample in the actual positive sample. Its expression is:

[0157]

[0158] Specificity refers to the ability of the model to correctly identify negative samples, which is relative to sensitivity and is expressed as:

[0159]

[0160] In order to consider both precision and recall, a threshold needs to be selected to achieve the highest balance between the two. Therefore, the concept of F1-score is proposed, and its calculation formula is:

[0161]

[0162] Furthermore, the experimental environment is shown in Table 1:

[0163] Table 1

[0164]

[0165] The size of all input images is 224*224*3, and features are extracted using EfficientViT pre-trained on ImageNet 1K. In order to improve the generalization of the model, the embodiments of the present application first perform data enhancement processing on these images, including random horizontal or vertical flipping, translation, scaling, rotation, and adjustment of brightness and contrast. The text end uses the pre-trained BERT model for feature extraction, and the dimension after pooling is 768. During the training phase, the model is set to train 200 times from scratch with a batch size of 32. The training samples of each epoch are randomly shuffled. The optimizer uses the Adam optimizer with a learning rate of 0.0001 and a weight decay of 1e-4, and the learning rate decay uses a cosine decay strategy.

[0166] In addition, the embodiment of the present application also selected several most advanced methods currently used on the Derm7pt dataset for comparison with the method provided by the present application. The comparison methods specifically include: Inception-combine, EmbeddingNet, TripleNet, HcCNN, AMFAM, FusionM4Net, CAFNet and TFormer; Table 2 shows the accuracy of the various methods obtained in the embodiment of the present application on the classification tasks of eight labels. From the data in Table 2, it can be seen that the classification effect of the method provided by the present application on most classification tasks has reached the highest, and its average accuracy is also the highest, reaching 79.14%, which is 1.54% higher than the prior art;

[0167] Table 2

[0168] Methods DIAG PN BMV VS PIG STR DAG RS Avg. Inception-combine 74.2 70.9 87.1 79.7 66.1 74.2 60.0 77.2 73.7 EmbeddingNet 68.6 65.1 84.3 82.5 64.3 73.4 57.5 78.0 71.7 TripleNet 68.6 63.3 87.9 83.0 67.3 74.4 61.3 76.0 72.7 HcCNN 69.9 70.6 87.1 84.8 68.6 71.6 65.6 80.8 74.9 AMFAM 75.4 70.6 88.1 83.3 70.9 74.7 63.8 80.8 76.0 FusionM4Net 77.6 69.2 88.5 81.6 71.3 76.1 64.4 81.4 76.3 CAFNet 78.2 70.1 87.8 84.3 73.4 77.0 61.5 81.8 76.8 TFormer 79.49 74.34 86.67 83.03 70.29 76.71 66.85 82.11 77.60 Ours 77.21 77.97 89.87 84.05 76.71 78.73 67.85 80.76 79.14

[0169] In addition, in order to fully verify the superiority of the method provided by the present application, the present application embodiment also adopts four other evaluation indicators for evaluation. In order to more intuitively demonstrate the effect of the present method, the present application embodiment adopts a radar chart to present it, such as Figure 6 The following is a radar chart showing the comparison of indicators of different methods on different classification tasks, where: Figure 6 (a) is a schematic diagram showing the comparison results of SEN, SPE, PRE and F1-score indicators of Inception-combine, AMFAM, FusionM4Net, TFormer and the method provided in this application in the BMV classification task. Figure 6 (b) is a schematic diagram showing the comparison results of SEN, SPE, PRE and F1-score indicators of Inception-combine, AMFAM, FusionM4Net, TFormer and the method provided in this application in the DAG classification task. Figure 6 (c) is a schematic diagram showing the comparison results of SEN, SPE, PRE and F1-score indicators of Inception-combine, AMFAM, FusionM4Net, TFormer and the method provided in this application in the DIAG classification task. Figure 6 (d) is a schematic diagram showing the comparison results of SEN, SPE, PRE and F1-score indicators of Inception-combine, AMFAM, FusionM4Net, TFormer and the method provided in this application in the PIG classification task. Figure 6(e) is a schematic diagram showing the comparison results of SEN, SPE, PRE and F1-score indicators of Inception-combine, AMFAM, FusionM4Net, TFormer and the method provided in this application in the PN classification task. Figure 6 (f) is a schematic diagram showing the comparison results of SEN, SPE, PRE and F1-score indicators of Inception-combine, AMFAM, FusionM4Net, TFormer and the method provided in this application in the RS classification task. Figure 6 (g) is a schematic diagram showing the comparison results of SEN, SPE, PRE and F1-score indicators of Inception-combine, AMFAM, FusionM4Net, TFormer and the method provided in this application in the STR classification task. Figure 6 (h) is a schematic diagram of the comparison results of SEN, SPE, PRE and F1-score indicators of Inception-combine, AMFAM, FusionM4Net, TFormer and the method provided in this application in the VS classification task; it can be seen from the figure that the method provided in this application has achieved the best results in most indicators. F1-score, as the arithmetic mean of SEN and PRE, increased by 2.3%, which indicates that the method provided in this application has better robustness and ability to handle unbalanced classes. PRE increased by 3.89%, SEN decreased by 2.32%, and SPE only decreased by 0.89%.

[0170] In addition, in order to verify the effectiveness of the modules designed in this application, the embodiment of this application also provides an ablation experiment to analyze the various modules of the model. As shown in Table 3, C represents the clinical image modality, D represents the dermoscopic image modality, and T represents the text modality. As far as single images are concerned, the accuracy of dermoscopic images is better than that of clinical images. In multimodal experiments, it can be observed that the use of connections alone cannot improve the accuracy of classification results. For example, the combined use of clinical images and dermoscopic images (75.17%) is lower than that of using only dermoscopic images (77.57%). This shows that the advantage of the model depends on an effective feature fusion strategy. The method provided in this application promotes effective interaction between different modes by designing appropriate fusion modalities, and its overall accuracy can reach 79.14%;

[0171] Table 3

[0172]

[0173] Based on the above experimental results, the method provided in this application can be actually applied to the medical field, such as the diagnosis of skin diseases, because in actual clinical diagnosis, doctors often do not rely on information from only one modality for diagnosis, but need to conduct multiple examinations and integrate multiple modality information for joint diagnosis, which can greatly improve the accuracy of diagnosis and reduce the risk of misdiagnosis.

[0174] In some embodiments, the Spring Boot framework can also be used for encapsulation, and a callable interface can be provided through the Spring Boot framework to achieve quick calls to fine-tune and deploy large-scale pre-trained models in downstream tasks. Spring Boot is a new framework provided by the Pivotal team. It is designed to simplify the initial setup and development process of new Spring applications. The framework uses a specific method for configuration, so that developers no longer need to define boilerplate configurations, and can quickly adapt to scenarios for actual classification tasks. It requires lower development costs and lightweight project standards. Therefore, Spring Boot is very suitable.

[0175] When creating a project with IntelliJ IDEA, choose Spring Inttializr to directly win the SpringBoot framework. First, create a class to encapsulate the data preprocessing method and the best model into a function, then write the calling logic in the service layer, and write the interface in the controller layer for the front-end page to call the interface. In this way, the new data set can be fine-tuned and inferred through the interface.

[0176] In the specific implementation, a web platform system can be designed for human-computer interaction. Developers can train different data sets in the platform, and the loss and other information generated during the training process will also be displayed in the platform, which is convenient for technical personnel to conduct real-time observation and analysis and summary. Then the developer can save the weight of the best model locally or in the cloud for actual use. Each user has his own account. When the user enters the platform, he will choose the appropriate model according to his needs and pass the image data and text data to be predicted to the platform. The platform will call the interface corresponding to the model, predict the input information, and feedback the results to the user. The output can be a direct classification result or presented in the form of probability. All prediction results will be saved in the user's usage record for easy query. Even if the model is deployed on a device with low computing power, you can enjoy the significant effects and convenience brought by the large model.

[0177] Based on the multimodal data fusion classification method based on a large model and an attention mechanism provided in the above embodiment, the embodiment of the present application also provides a multimodal data fusion classification device based on a large model and an attention mechanism, such as Figure 7 As shown, the device specifically includes:

[0178] A feature acquisition module 10 is used to acquire a first image feature vector, a second image feature vector and a text feature vector of a target object;

[0179] The first feature fusion module 20 is used to input the first image feature vector and the second image feature vector into the sliding window cross attention fusion module in the trained classification model, and output the first target image feature vector and the second target image feature vector, which specifically includes:

[0180] An intermediate image feature vector acquisition submodule, used to input the first image feature vector and the second image feature vector respectively into a first normalization submodule, a window multi-head self-attention submodule and a first feedforward neural network submodule which are sequentially connected in series along a forward propagation direction, and output a first intermediate image feature vector and a second intermediate image feature vector;

[0181] A first target image feature vector acquisition submodule is used to simultaneously input the query vector of the first intermediate image feature vector, the key vector and the eigenvalue vector of the second intermediate image feature vector to the second normalization submodule, the sliding window multi-head cross attention submodule and the second feedforward neural network submodule which are sequentially connected in series along the forward propagation direction, and output the first target image feature vector;

[0182] A second target image feature vector acquisition submodule is used to simultaneously input the query vector of the second intermediate image feature vector, the key vector and the eigenvalue vector of the first intermediate image feature vector to the second normalization submodule, the sliding window multi-head cross attention submodule and the second feedforward neural network submodule which are sequentially connected in series along the forward propagation direction, and output the second target image feature vector;

[0183] A second feature fusion module 30 is used to input the first target image feature vector, the second target image feature vector and the text feature vector into a heterogeneous data cross attention fusion module in a trained classification model, and output a target feature representation of the target object;

[0184] The classification module 40 is used to input the target feature representation of the target object into the fully connected layer in the trained classification model, and output the classification result of the target object.

[0185] An embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned multimodal data fusion classification method based on a large model and an attention mechanism are implemented.

[0186] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.

[0187] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0188] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0189] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0190] Obviously, the above embodiments are merely examples for the purpose of clear explanation and are not intended to limit the implementation methods. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the implementation methods here. The obvious changes or modifications derived therefrom are still within the scope of protection of the present invention.

Claims

1. A multimodal data fusion classification method based on a large model and an attention mechanism, characterized in that: include: Acquiring a first image feature vector, a second image feature vector, and a text feature vector of a target object specifically includes: Acquire first image data, second image data, and text data of a target object; Inputting the first image data and the second image data into a visual feature extractor for feature extraction, and outputting a first image feature vector and a second image feature vector; Inputting the text data into a text feature extractor for feature extraction, and outputting a text feature vector; Inputting the first image feature vector and the second image feature vector into a sliding window cross attention fusion module in a trained classification model, and outputting a first target image feature vector and a second target image feature vector, which specifically includes: The first image feature vector and the second image feature vector are respectively inputted into a first normalization submodule, a window multi-head self-attention submodule and a first feedforward neural network submodule which are sequentially connected in series along the forward propagation direction, and a first intermediate image feature vector and a second intermediate image feature vector are outputted; The query vector of the first intermediate image feature vector, the key vector and the eigenvalue vector of the second intermediate image feature vector are simultaneously input into the second normalization submodule, the sliding window multi-head cross attention submodule and the second feedforward neural network submodule which are sequentially connected in series along the forward propagation direction, and the first target image feature vector is output; The query vector of the second intermediate image feature vector, the key vector and the eigenvalue vector of the first intermediate image feature vector are simultaneously input into the second normalization submodule, the sliding window multi-head cross attention submodule and the second feedforward neural network submodule which are sequentially connected in series along the forward propagation direction, and the second target image feature vector is output; Inputting the first target image feature vector, the second target image feature vector and the text feature vector into a heterogeneous data cross attention fusion module in a trained classification model, and outputting a target feature vector of a target object; The target feature vector of the target object is input into the fully connected layer in the trained classification model, and the classification result of the target object is output.

2. The multimodal data fusion classification method based on a large model and an attention mechanism according to claim 1 is characterized in that: The first feedforward neural network submodule and the second feedforward neural network submodule both include a normalization unit and a nonlinear transformation unit connected in series.

3. The multimodal data fusion classification method based on large model and attention mechanism according to claim 1 is characterized in that: The first intermediate image feature vector is expressed as: , , in, represents the first intermediate image feature vector, represents the first feedforward neural network submodule, represents the normalization operation, represents the window multi-head self-attention submodule, represents the first image feature vector; The second intermediate image feature vector is expressed as: , , in, represents the second intermediate image feature vector, represents the second image feature vector.

4. The multimodal data fusion classification method based on large model and attention mechanism according to claim 1 is characterized in that: The first target image feature vector is expressed as: , , in, represents the first target image feature vector, represents the second feedforward neural network submodule, represents the normalization operation, represents the sliding window multi-head cross attention submodule, The query vector representing the first intermediate image feature vector, The key vector representing the feature vector of the second intermediate image, The eigenvalue vector representing the eigenvalue vector of the second intermediate image, represents the first intermediate image feature vector; The second target image feature vector is expressed as: , , in, represents the second target image feature vector, The query vector representing the feature vector of the second intermediate image, The key vector representing the first intermediate image feature vector, represents the eigenvalue vector of the first intermediate image eigenvector, Represents the second intermediate image feature vector.

5. The multimodal data fusion classification method based on large model and attention mechanism according to claim 1 is characterized in that: Inputting the first target image feature vector, the second target image feature vector, and the text feature vector into a heterogeneous data cross attention fusion module in a trained classification model, and outputting a target feature representation of the target object comprises: Inputting the key vector of the first target image feature vector and the query vector of the text feature vector into a first matrix cross multiplication submodule for matrix cross multiplication, and outputting a first heterogeneous data fusion feature vector; Normalizing the first heterogeneous data fusion feature vector, and inputting the normalized first heterogeneous data fusion feature vector and the eigenvalue vector of the first target image feature vector into a second matrix cross multiplication submodule for matrix cross multiplication, and outputting a second heterogeneous data fusion feature vector; Inputting the query vector of the second heterogeneous data fusion feature vector and the text feature vector into a first residual connection submodule, and outputting a first target fusion feature vector; Inputting the key vector of the second target image feature vector and the query vector of the text feature vector into a third matrix cross multiplication submodule for matrix cross multiplication, and outputting a third heterogeneous data fusion feature vector; Normalizing the third heterogeneous data fusion feature vector, and inputting the normalized third heterogeneous data fusion feature vector and the eigenvalue vector of the second target image feature vector into a fourth matrix cross multiplication submodule for matrix cross multiplication, and outputting a fourth heterogeneous data fusion feature vector; Inputting the query vector of the fourth heterogeneous data fusion feature vector and the text feature vector into a second residual connection submodule, and outputting a second target fusion feature vector; The first target fusion feature vector and the second target fusion feature vector are concatenated to obtain a target feature vector of the target object.

6. The multimodal data fusion classification method based on large model and attention mechanism according to claim 5 is characterized in that: The first target fusion feature vector is expressed as: , , in, represents the first target fusion feature vector, represents the text feature vector, represents the projection function for eigenvector alignment, represents the second heterogeneous data fusion feature vector, represents the cross attention operation, represents the normalization operation, represents the first heterogeneous data fusion feature vector, represents the first target feature vector; The second target fusion feature vector is expressed as: , , in, represents the second target fusion feature vector, represents the fourth heterogeneous data fusion feature vector, represents the third heterogeneous data fusion feature vector, represents the second target feature vector; The target feature vector of the target object is expressed as: , in, The target feature vector representing the target object.

7. The multimodal data fusion classification method based on large model and attention mechanism according to claim 1 is characterized in that: The visual feature extractor includes a patch embedding module, a first EfficientViTBlock module, a first downsampling module, a second EfficientViT Block module, a second downsampling module and a third EfficientViT Block module which are connected in series in a forward propagation direction. The first EfficientViT Block module, the second EfficientViT Block module and the third EfficientViT Block module each include a first deep convolution submodule, a third feedforward neural network submodule, a cascaded group attention submodule, a second deep convolution submodule and a fourth feedforward neural network submodule which are connected in series in a forward propagation direction.

8. The multimodal data fusion classification method based on large model and attention mechanism according to claim 1 is characterized in that: The first image feature vector is expressed as: , in, represents the first image feature vector, represents the visual feature extractor, represents embedding, represents first image data; The second image feature vector is expressed as: , in, represents the second image feature vector, represents second image data; The text feature vector is expressed as: , in, represents the text feature vector, represents the pooling operation, represents a text feature extractor, Represents text data.

9. A multimodal data fusion classification device based on a large model and an attention mechanism, characterized in that: include: The feature acquisition module is used to acquire the first image feature vector, the second image feature vector and the text feature vector of the target object, which specifically includes: Acquire first image data, second image data, and text data of a target object; Inputting the first image data and the second image data into a visual feature extractor for feature extraction, and outputting a first image feature vector and a second image feature vector; Inputting the text data into a text feature extractor for feature extraction, and outputting a text feature vector; The first feature fusion module is used to input the first image feature vector and the second image feature vector into the sliding window cross attention fusion module in the trained classification model, and output the first target image feature vector and the second target image feature vector, which specifically includes: An intermediate image feature vector acquisition submodule, used to input the first image feature vector and the second image feature vector respectively into a first normalization submodule, a window multi-head self-attention submodule and a first feedforward neural network submodule which are sequentially connected in series along a forward propagation direction, and output a first intermediate image feature vector and a second intermediate image feature vector; A first target image feature vector acquisition submodule is used to simultaneously input the query vector of the first intermediate image feature vector, the key vector and the eigenvalue vector of the second intermediate image feature vector to the second normalization submodule, the sliding window multi-head cross attention submodule and the second feedforward neural network submodule which are sequentially connected in series along the forward propagation direction, and output the first target image feature vector; A second target image feature vector acquisition submodule is used to simultaneously input the query vector of the second intermediate image feature vector, the key vector and the eigenvalue vector of the first intermediate image feature vector to the second normalization submodule, the sliding window multi-head cross attention submodule and the second feedforward neural network submodule which are sequentially connected in series along the forward propagation direction, and output the second target image feature vector; A second feature fusion module, configured to input the first target image feature vector, the second target image feature vector and the text feature vector into a heterogeneous data cross attention fusion module in a trained classification model, and output a target feature representation of a target object; The classification module is used to input the target feature representation of the target object into the fully connected layer in the trained classification model and output the classification result of the target object.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the steps of the multimodal data fusion classification method based on a large model and an attention mechanism according to any one of claims 1 to 8 are implemented.