Image classification method based on bidirectional guide aggregation multi-mode prompt learning

By introducing a two-way guided aggregation of multimodal prompt learning method in the visual language model, the problem of insufficient performance of the visual language model in downstream tasks is solved, efficient interaction and flexible integration of multimodal information are achieved, and the generalization ability and classification performance of the model are improved.

CN120472211APending Publication Date: 2025-08-12GUILIN UNIV OF ELECTRONIC TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510545356.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-08-12

Smart Images

  • Figure CN120472211A_ABST
    Figure CN120472211A_ABST
Patent Text Reader

Abstract

The invention discloses a multimodal prompt learning image classification method based on bidirectional guide aggregation. According to the method, a bidirectional hierarchical interaction mechanism is innovatively constructed in a visual language model CLIP framework. Specifically, independent prompts and aggregation prompts are introduced into an image encoder and a text encoder respectively, and the aggregation prompts are generated through a guide prompt module and an aggregation prompt module: firstly, the independent prompts generate cross-modal guide prompts for another modal through a guide module; and the guide prompt is adaptively fused with the independent prompt of the previous layer through an attention mechanism, and finally, the independent prompt generated by each layer and the aggregation prompt are spliced and input into an encoder for learning. According to the method, deep integration of multi-modal information among different abstract hierarchies is realized, on the premise that pre-training knowledge is completely reserved, the recognition capability of the model for unseen categories can be remarkably improved only by a small number of samples, and the problem of poor generalization performance caused by insufficient modal interaction in a traditional method is effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical field

[0001] The present invention relates to the technical field of prompt learning, and more particularly to an image classification method based on bidirectional guided aggregated multimodal prompt learning. [Background Technology]

[0002] Visual language models such as CLIP (Contrastive Language-Image Pre-training) have demonstrated powerful capabilities in understanding and generating image descriptions. These models, pre-trained on large-scale image-text pairs, are able to capture the deep connections between images and text. While CLIP performs well on pre-training tasks, its direct application to specific downstream tasks can be challenging. Therefore, improving CLIP's capabilities on downstream tasks is a worthy research topic. Fine-tuning typically requires a large amount of labeled data to adjust the model's parameters, making it unsuitable for transfer to few-shot scenarios. Furthermore, fine-tuning can destroy the general knowledge learned by visual language models during pre-training, leading to catastrophic forgetting. Cued learning, an emerging technique, has been shown to effectively guide pre-trained models to perform better on specific tasks. Applying cue learning to CLIP models is expected to further improve their performance on few-shot image classification tasks.

[0003] To address the above issues, inspired by prompt learning in natural language processing, many researchers have proposed prompt learning methods to adapt visual language models. As a lightweight transfer learning paradigm, prompt learning introduces trainable prompt vectors to adapt to downstream tasks while freezing pre-trained parameters. This technology has been extended from the field of NLP to visual language models (VLMs). Early methods such as CoOp used static text prompts to improve task adaptability, but suffered from domain generalization bottlenecks; VPT added learnable prompts to the visual branch but lacked cross-modal collaboration capabilities. Subsequent research has gradually broken through the limitations of single modality; CoCoOp enhanced generalization through image-conditional dynamic prompts, while MaPLe first constructed a text-visual dual-branch prompt framework and used linear coupling functions to achieve one-way interaction between modalities. However, existing methods are still limited by shallow cross-modal fusion and fail to fully tap the multimodal collaboration potential of VLMs. [Summary of the invention]

[0004] The purpose of this invention is to address the problems of the prior art by proposing an image classification method based on bidirectional guided aggregated multimodal cue learning. This method inserts independent text cues and text aggregate cues into the text encoder, and independent image cues and image aggregate cues into the image encoder. Aggregate cues are generated by the guided cue module and the aggregated cue module, ensuring a more efficient and flexible integration of the two modes, thereby achieving alignment for specific downstream tasks. This method enables efficient interaction of multimodal information, conserves computing resources, and improves the model's generalization capabilities.

[0005] In order to solve the above problems, the technical solution of the present invention is: an image classification method based on bidirectional guided aggregate multimodal prompt learning, comprising the following steps:

[0006] Step 1: Construct the initial learnable text prompt:

[0007] 1.1 By constructing a framework for each category that combines human prior knowledge with learnable text prompts. The learnable text prompts consist of independent text prompts and image-guided cross-modal prompts.

[0008] 1.2 Independent text prompts are initialized through a normal distribution, and image-guided cross-modal prompts are generated by independent image prompts through a guided prompt module (BGPM). The guided prompt module (BGPM) is composed of multi-head self-attention and a multi-layer perceptron (MLP).

[0009] Step 2: Construct the initial learnable image prompts:

[0010] 2.1 Before being input into the image encoder, the image is divided into multiple image blocks. Each image block is mapped into an embedding vector. The resulting vector sequence is input into the image encoder together with learnable image cues. The learnable image cues consist of independent image cues and text-guided cross-modal cues.

[0011] 2.2 Independent image cues are initialized by normal distribution, and text-guided cross-modal cues are generated by independent text cues through the guided prompt module BGPM;

[0012] Step 3: Build the aggregation prompt:

[0013] 3.1 Aggregation Prompt Module APM includes a text aggregation prompt module and an image different source module. These two modules are responsible for integrating prompt information from different sources;

[0014] The text aggregation prompt module dynamically fuses the information from the previous layer prompts of the text encoder and the image guidance prompts to generate the aggregated prompts of the current layer as the input of the text encoder. It uses the attention mechanism to adaptively weight the prompt information from different sources and aggregate the guidance prompt information of the image branch and the prompt information of the previous layer.

[0015] 3.3 The image aggregation prompt module dynamically fuses the information from the previous layer prompts and the text guidance prompts of the image encoder to generate the aggregated prompts of the current layer as the input of the image encoder. It is also based on the attention mechanism and adaptively aggregates the guidance prompt information of the text branch and the prompt information of the previous layer;

[0016] Step 4 combines and optimizes the independent prompts generated by the normal distribution initialization method with the aggregated prompts to form the final input prompt. The optimized prompts are input into the pre-trained CLIP visual language model. Through its powerful cross-modal understanding capabilities, the semantic feature vectors of the text prompts of each category and the visual feature vectors of the input image are extracted respectively.

[0017] Step 5: In the joint embedding space, the learned contextual cue vector is used to guide the cross-modal similarity calculation between the image features and the parameterized text template. By measuring the prototype similarity between the image and each category of text, a normalized category probability distribution is generated, and finally a classification prediction result based on cross-modal matching is output;

[0018] In step 6, based on the extracted bimodal feature representation, the cross entropy loss between the text feature vector and the image feature vector is calculated as the classification loss function. The loss gradient is propagated back to the prompt parameter space through the backpropagation algorithm to achieve end-to-end learnable prompt parameter optimization.

[0019] The guidance prompt module (BGPM) in step 1 adopts a cross-modal interaction architecture based on multi-head self-attention and multi-layer perceptron (MLP). The generation process of image-guided cross-modal prompts is as follows: first, independent image prompts are mapped into query vectors Q, key vectors K, and value vectors V through three basic linear projection layers, and then the cross-modal attention weights are calculated through the multi-head self-attention mechanism. The output of each attention head is concatenated and then passed through an MLP module containing two fully connected layers and a ReLU activation function for feature transformation and enhancement. The final output is used as the image-guided cross-modal prompt. The calculation process is shown as follows:

[0020] P v→l =BGPM(P v )=MLP(MA(P v )) (1)

[0021] The MLP consists of two linear layers and a ReLU activation function in the middle of the two linear layers, and MA stands for multi-head self-attention.

[0022] Q=P l W q ,K=P l W k ,V=P l W v (2)

[0023] Where W q ,W k ,W k ∈R dv×dh is a linear projection layer;

[0024] Then, through the self-attention mechanism calculation, the calculation method of each self-attention mechanism head is as follows:

[0025]

[0026] where d H =d v / H is the dimensional feature output by each attention head.

[0027] Step 2: For the generation of text-guided cross-modal prompts, a cross-modal interaction architecture based on multi-head self-attention and multi-layer perceptron (MLP) is also used: the independent text prompts are respectively generated into a query vector Q, a key vector K, and a value vector V through three basic linear transformation layers; then the cross-modal attention weights are calculated through the multi-head self-attention mechanism, where the output of each attention head is concatenated and passed through an MLP module containing two fully connected layers and a ReLU activation function for feature transformation and enhancement. The final output is used as the text-guided cross-modal prompt. The calculation process is shown as follows:

[0028] P l→v =BGPM(P l )=MLP(MA(P l )) (4)

[0029] The MLP consists of two linear layers and a ReLU activation function in the middle of the two linear layers. MA stands for multi-head self-attention, and the calculation process is similar to that of image-guided cross-modal prompts.

[0030] In the process of constructing the aggregated prompt in step 3, the generation of the prompt in the i-th layer adopts a dynamic aggregation mechanism. The specific implementation method is: the guidance prompt of the current layer is used as the query vector Q, and the independent prompts of the i-1th layer are used as the key vector K and value vector V respectively. They are dynamically fused through the attention mechanism and then added to the guidance prompt by the α weight coefficient. The calculation process is shown as follows:

[0031]

[0032]

[0033] in and denote the text-guided cross-modal prompt and image-guided cross-modal prompt of the i-th layer, respectively. and They represent independent image prompts and independent text prompts at layer i-1, respectively, and α is a learnable weight coefficient.

[0034] Step 4 combines and optimizes the independent prompts generated by the normal distribution initialization method with the aggregated prompts to form the final input prompt. The optimized prompts are input into the pre-trained CLIP visual language model. Through its powerful cross-modal understanding capabilities, the semantic feature vectors of the text prompts of each category and the visual feature vectors of the input image are extracted respectively.

[0035] For the text encoder L, the input text embedding Independent text prompts and text aggregation prompts Splicing composition Input to the text encoder L i :

[0036]

[0037] Where [·,·,·] represents the concatenation operation. After K layers of prompt learning, the text features are further processed by TextProj(·) to obtain the final text representation z:

[0038]

[0039]

[0040] For the image encoder V, given an RGB input image, it is first divided into M non-overlapping image blocks, and then the image embedding is obtained by linear projection. Final and independent image prompts and image aggregation tips Splicing composition Input to the image encoder V i :

[0041]

[0042] Where [·,·,·] represents the concatenation operation. After K layers of prompt learning, the image features are further processed by ImageProj(·) to obtain the final image representation x:

[0043]

[0044] x=ImageProj(c k ),c k ∈R dv (12).

[0045] Steps 5 and 6 calculate the predicted probability based on the cosine similarity between the text representation z and the image representation x:

[0046]

[0047] where sim(·) represents the cosine similarity function, τ is the temperature parameter, L is the number of categories, and y and x represent the true label and input image.

[0048] The final training goal is to use the cross entropy loss function L ce Calculate the loss and propagate the loss gradient back to the prompt parameter space through the backpropagation algorithm to achieve end-to-end learnable prompt parameter optimization:

[0049]

[0050] Here, B represents the batch size.

Brief Description of the Drawings

[0051] Figure 1 This is a flowchart of an image classification method based on bidirectional guided aggregated multimodal prompt learning of the present invention.

[0052] Figure 2 This is a framework diagram of an image classification method based on bidirectional guided aggregated multimodal prompt learning in the present invention. [Specific implementation method]

[0053] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments, but the present invention is not limited thereto.

[0054] Example 1

[0055] According to an embodiment of the present invention, a method for image classification based on bidirectional guided aggregate multimodal prompt learning is provided. Figure 1 The figure shows a flow chart of the two-way guidance prompt learning method provided in this example, which includes the following steps:

[0056] Step 1: Construct the initial learnable text prompt:

[0057] 1.1 By constructing a framework for each category that combines human prior knowledge with learnable text prompts. The learnable text prompts consist of independent text prompts and image-guided cross-modal prompts.

[0058] 1.2 Independent text prompts are initialized through a normal distribution, and image-guided cross-modal prompts are generated by independent image prompts through a guided prompt module (BGPM). The guided prompt module (BGPM) is composed of multi-head self-attention and a multi-layer perceptron (MLP).

[0059] Step 2: Construct the initial learnable image prompts:

[0060] 2.1 Before being input into the image encoder, the image is divided into multiple image blocks. Each image block is mapped into an embedding vector. The resulting vector sequence is input into the image encoder together with learnable image cues. The learnable image cues consist of independent image cues and text-guided cross-modal cues.

[0061] 2.2 Independent image cues are initialized by normal distribution, and text-guided cross-modal cues are generated by independent text cues through the guided prompt module BGPM;

[0062] Step 3: Build the aggregation prompt:

[0063] 3.1 Aggregation Prompt Module APM includes a text aggregation prompt module and an image aggregation prompt module. These two modules are responsible for integrating prompt information from different sources;

[0064] 3.2 The text aggregation prompt module dynamically fuses the information from the previous layer prompts of the text encoder and the image guidance prompts to generate the aggregated prompts of the current layer as the input of the text encoder. It uses the attention mechanism to adaptively weight the prompt information from different sources, aggregating the guidance prompt information of the image branch and the prompt information of the previous layer;

[0065] 3.3 The image aggregation prompt module dynamically fuses the information from the previous layer prompts and the text guidance prompts of the image encoder to generate the aggregated prompts of the current layer as the input of the image encoder. It is also based on the attention mechanism and adaptively aggregates the guidance prompt information of the text branch and the prompt information of the previous layer;

[0066] Step 4 combines and optimizes the independent prompts generated by the normal distribution initialization method with the aggregated prompts to form the final input prompt. The optimized prompts are input into the pre-trained CLIP visual language model. Through its powerful cross-modal understanding capabilities, the semantic feature vectors of the text prompts of each category and the visual feature vectors of the input image are extracted respectively.

[0067] Step 5: In the joint embedding space, the learned contextual cue vector is used to guide the cross-modal similarity calculation between the image features and the parameterized text template. By measuring the prototype similarity between the image and each category of text, a normalized category probability distribution is generated, and finally a classification prediction result based on cross-modal matching is output;

[0068] In step 6, based on the extracted bimodal feature representation, the cross entropy loss between the text feature vector and the image feature vector is calculated as the classification loss function. The loss gradient is propagated back to the prompt parameter space through the backpropagation algorithm to achieve end-to-end learnable prompt parameter optimization.

[0069] The guidance prompt module (BGPM) in step 1 adopts a cross-modal interaction architecture based on multi-head self-attention and multi-layer perceptron (MLP). The generation process of image-guided cross-modal prompts is as follows: first, independent image prompts are mapped into query vectors Q, key vectors K, and value vectors V through three basic linear projection layers, and then the cross-modal attention weights are calculated through the multi-head self-attention mechanism. The output of each attention head is concatenated and then passed through an MLP module containing two fully connected layers and a ReLU activation function for feature transformation and enhancement. The final output is used as the image-guided cross-modal prompt. The calculation process is shown as follows:

[0070] P v→l =BGPM(P v )=MLP(MA(P v )) (1)

[0071] The MLP consists of two linear layers and a ReLU activation function in the middle of the two linear layers, and MA stands for multi-head self-attention.

[0072] Q=P l W q ,K=P l W k ,V=P l W v (2)

[0073] in is a linear projection layer;

[0074] Then, through the self-attention mechanism calculation, the calculation method of each self-attention mechanism head is as follows:

[0075]

[0076] where d H =d v / H is the dimensional feature output by each attention head.

[0077] Step 2: For the generation of text-guided cross-modal prompts, a cross-modal interaction architecture based on multi-head self-attention and multi-layer perceptron (MLP) is also used: the independent text prompts are respectively generated into a query vector Q, a key vector K, and a value vector V through three basic linear transformation layers; then the cross-modal attention weights are calculated through the multi-head self-attention mechanism, where the output of each attention head is concatenated and passed through an MLP module containing two fully connected layers and a ReLU activation function for feature transformation and enhancement. The final output is used as the text-guided cross-modal prompt. The calculation process is shown as follows:

[0078] P l→v =BGPM(P l )=MLP(MA(P l)) (4)

[0079] The MLP consists of two linear layers and a ReLu activation function in the middle of the two linear layers. MA stands for multi-head self-attention, and the calculation process is similar to that of image-guided cross-modal prompts.

[0080] In the process of constructing the aggregated prompt in step 3, the generation of the prompt in the i-th layer adopts a dynamic aggregation mechanism. The specific implementation method is: the guidance prompt of the current layer is used as the query vector Q, and the independent prompts of the i-1th layer are used as the key vector K and value vector V respectively. They are dynamically fused through the attention mechanism and then added to the guidance prompt by the α weight coefficient. The calculation process is shown as follows:

[0081]

[0082]

[0083] in and denote the text-guided cross-modal prompt and image-guided cross-modal prompt of the i-th layer, respectively. and They represent independent image prompts and independent text prompts at layer i-1, respectively, and α is a learnable weight coefficient.

[0084] Step 4 combines and optimizes the independent prompts generated by the normal distribution initialization method with the aggregated prompts to form the final input prompt. The optimized prompts are input into the pre-trained CLIP visual language model. Through its powerful cross-modal understanding capabilities, the semantic feature vectors of the text prompts of each category and the visual feature vectors of the input image are extracted respectively.

[0085] For the text encoder L, the input text embedding Independent text prompts and text aggregation prompts Splicing composition Input to the text encoder L i :

[0086]

[0087] Where [·,·,·] represents the concatenation operation. After K layers of prompt learning, the text features are further processed by TextProj(·) to obtain the final text representation z:

[0088]

[0089]

[0090] For the image encoder V, given an RGB input image, it is first divided into M non-overlapping image blocks, and then the image embedding is obtained by linear projection. Final and independent image prompts and image aggregation tips Splicing composition Input to the image encoder V i :

[0091]

[0092] Where [·,·,·] represents the concatenation operation. After K layers of prompt learning, the image features are further processed by ImageProj(·) to obtain the final image representation x:

[0093]

[0094] x=ImageProj(c k ),c k ∈R dv (12).

[0095] Steps 5 and 6 calculate the predicted probability based on the cosine similarity between the text representation z and the image representation x:

[0096]

[0097] where sim(·) represents the cosine similarity function, τ is the temperature parameter, L is the number of categories, and y and x represent the true label and input image.

[0098] The final training goal is to use the cross entropy loss function L ce Calculate the loss and propagate the loss gradient back to the prompt parameter space through the backpropagation algorithm to achieve end-to-end learnable prompt parameter optimization:

[0099]

[0100] Here, B represents the batch size.

[0101] Implementation 2

[0102] The present invention can be widely applied in multiple different downstream application scenarios. Taking the remote sensing image recognition task as an example, the dataset used in this task is EuroSAT.

[0103] refer to Figure 2As shown in the figure, a framework diagram of an image classification method based on bidirectional guided aggregate multimodal prompt learning is used to train the model using a single RTX3090 GPU, adopting a few-shot training strategy, taking 16 samples for each class, setting the epoch to 5, using a batch size of 4 and a learning rate of 0.0035, and training through the SGD optimizer. The pre-trained ViT-B / 16CLIP model is selected as the base model, where the dimensions d of the image and text representations are v and d t They are set to 768 and 512 respectively, and α is set to 0.75.

[0104] During the training phase, a text prompt template that integrates domain prior knowledge and learnable prompts is constructed for each category. The input image is divided into multiple image blocks and encoded into an embedding vector sequence. This sequence and the learnable image prompt tags are jointly input into the image encoder. The prompt guidance module (BGPM) and the aggregation prompt module (APM) are used to achieve bidirectional interaction and joint optimization of text and visual prompts, thereby enhancing the model's ability to capture features of specific categories.

[0105] The cross-entropy loss function is used to quantify the difference between the classification results and the true labels, and then all relevant parameters including text and image prompts are updated based on the calculated classification loss. After multiple rounds of iterative training, when the model reaches the predetermined performance standard, the final optimized prompt parameters can be used for prediction to generate classification results.

[0106] The classification performance comparison between the proposed method and the current advanced hint learning method on the EuroSAT dataset is shown in Table 1, where Base represents the classification accuracy of the base class, Novel represents the classification accuracy of the new class, and HM is the harmonic mean of the two.

[0107] Table 1

[0108] Comparative experiments on the EuroSAT dataset demonstrate that the proposed method comprehensively surpasses state-of-the-art cued learning techniques in classification performance. As shown in Table 1, the proposed method achieves a base class classification accuracy of 94.43%, a novel class classification accuracy of 82.73%, and a harmonic mean accuracy of 88.19%, significantly outperforming the comparative method in all three key metrics.

[0109] In summary, the method provided by this invention enables efficient interaction of multimodal information, conserves computing resources, and improves the generalization capability of the model, providing an efficient, flexible, and scalable solution for this field. Experimental data shows that this method significantly improves the model's performance in classification tasks while maintaining computational efficiency.

[0110] The above embodiments are only used to illustrate the technical principles and core features of the present invention. Those skilled in the art should understand that the present invention is not limited to the specific details shown in the aforementioned embodiments. In practical applications, based on the basic principles of the present invention, if the technical features are equivalently replaced, the implementation scheme is appropriately modified, etc., without departing from the essence of the present invention, the obtained technical solutions should be deemed to fall within the scope of protection of the present invention.

Claims

1. An image classification method based on bidirectional guided aggregate multimodal prompt learning, characterized in that: Including steps: Step 1: 1.1 By constructing a framework for each category that combines human prior knowledge with learnable text prompts. The learnable text prompts consist of independent text prompts and image-guided cross-modal prompts. 1.2 Independent text prompts are initialized through a normal distribution, and image-guided cross-modal prompts are generated by independent image prompts through a guided prompt module (BGPM). The guided prompt module (BGPM) is composed of multi-head self-attention and a multi-layer perceptron (MLP). Step 2: 2.1 Before being input into the image encoder, the image is divided into multiple image blocks. Each image block is mapped into an embedding vector. The resulting vector sequence is input into the image encoder together with learnable image cues. The learnable image cues consist of independent image cues and text-guided cross-modal cues. 2.2 Independent image cues are initialized by normal distribution, and text-guided cross-modal cues are generated by independent text cues through the guided prompt module BGPM; Step 3: 3.1 Aggregation Prompt Module APM includes a text aggregation prompt module and an image aggregation prompt module, which are responsible for integrating prompt information from different sources; 3.2 The text aggregation prompt module dynamically fuses the information from the previous layer prompts of the text encoder and the image guidance prompts to generate the aggregated prompts of the current layer as the input of the text encoder. It uses the attention mechanism to adaptively weight the prompt information from different sources, aggregating the guidance prompt information of the image branch and the prompt information of the previous layer; 3.3 The image aggregation prompt module dynamically fuses the information from the previous layer prompts and the text guidance prompts of the image encoder to generate the aggregated prompts of the current layer as the input of the image encoder. It is also based on the attention mechanism and adaptively aggregates the guidance prompt information of the text branch and the prompt information of the previous layer; Step 4 combines and optimizes the independent prompts generated by the normal distribution initialization method with the aggregated prompts to form the final input prompt. The optimized prompts are input into the pre-trained CLIP visual language model. Through its powerful cross-modal understanding capabilities, the semantic feature vectors of the text prompts of each category and the visual feature vectors of the input image are extracted respectively. Step 5: In the joint embedding space, the learned contextual cue vector is used to guide the cross-modal similarity calculation between the image features and the parameterized text template. By measuring the prototype similarity between the image and each category of text, a normalized category probability distribution is generated, and finally a classification prediction result based on cross-modal matching is output; In step 6, based on the extracted bimodal feature representation, the cross entropy loss between the text feature vector and the image feature vector is calculated as the classification loss function. The loss gradient is propagated back to the prompt parameter space through the backpropagation algorithm to achieve end-to-end learnable prompt parameter optimization.

2. The image classification method based on bidirectional guided aggregated multimodal prompt learning according to claim 1, characterized in that: The guidance prompt module (BGPM) in step 1 adopts a cross-modal interaction architecture based on multi-head self-attention and multi-layer perceptron (MLP). The generation process of image-guided cross-modal prompts is as follows: first, independent image prompts are mapped into query vectors Q, key vectors K, and value vectors V through three basic linear projection layers, and then the cross-modal attention weights are calculated through the multi-head self-attention mechanism. The output of each attention head is concatenated and then passed through an MLP module containing two fully connected layers and a ReLU activation function for feature transformation and enhancement. The final output is used as the image-guided cross-modal prompt. The calculation process is shown as follows: P v→l =BGPM(P v )=MLP(MA(P v )) (1) The MLP consists of two linear layers and a ReLU activation function between the two linear layers. MA stands for multi-head self-attention. Q=P l W q ,K=P l W k ,V=P l W v (2) in is a linear projection layer; Then, through the self-attention mechanism calculation, the calculation method of each self-attention mechanism head is as follows: where d H =d v / H is the dimensional feature output by each attention head.

3. The image classification method based on bidirectional guided aggregated multimodal prompt learning according to claim 1, characterized in that: Step 2: For the generation of text-guided cross-modal prompts, a cross-modal interaction architecture based on multi-head self-attention and multi-layer perceptron (MLP) is also used: the independent text prompts are respectively generated into a query vector Q, a key vector K, and a value vector V through three basic linear transformation layers; then the cross-modal attention weights are calculated through the multi-head self-attention mechanism, where the output of each attention head is concatenated and passed through an MLP module containing two fully connected layers and a ReLU activation function for feature transformation and enhancement. The final output is used as the text-guided cross-modal prompt. The calculation process is shown as follows: P l→v =BGPM(P l )=MLP(MA(P l )) (4) The MLP consists of two linear layers and a ReLU activation function in the middle of the two linear layers. MA stands for multi-head self-attention, and the calculation process is similar to that of image-guided cross-modal prompts.

4. The image classification method based on bidirectional guided aggregated multimodal prompt learning according to claim 1, characterized in that: In the process of constructing the aggregated prompt in step 3, the generation of the prompt in the i-th layer adopts a dynamic aggregation mechanism. The specific implementation method is: the guidance prompt of the current layer is used as the query vector Q, and the independent prompts of the i-1th layer are used as the key vector K and value vector V respectively. They are dynamically fused through the attention mechanism and then added to the guidance prompt by the α weight coefficient. The calculation process is shown as follows: in and denote the text-guided cross-modal prompt and image-guided cross-modal prompt of the i-th layer, respectively. and They represent the independent image prompt and independent text prompt of the i-1 layer respectively, and α is the weight coefficient.

5. The image classification method based on bidirectional guided aggregated multimodal prompt learning according to claim 1, characterized in that: Step 4 is the final input to the encoder prompt construction process, which uses the normal distribution initialization method to generate learnable independent prompts, which are combined and optimized with the aggregated prompts to form the final input prompts; For the text encoder L, the input text embedding Independent text prompts and text aggregation prompts Splicing composition Input to the text encoder L i : Where [·,·,·] represents the concatenation operation. After K layers of prompt learning, the text features are further processed by TextProj(·) to obtain the final text representation z: For the image encoder V, given an RGB input image, it is first divided into M non-overlapping image blocks, and then the image embedding is obtained by linear projection. Final and independent image prompts and image aggregation tips Splicing composition Input to the image encoder V i : Where [·,·,·] represents the concatenation operation. After K layers of prompt learning, the image features are further processed by ImageProj(·) to obtain the final image representation x: x=ImageProj(c k ),c k ∈R dv (12)。 6. The image classification method based on bidirectional guided aggregated multimodal prompt learning according to claim 1, characterized in that: Steps 5 and 6 calculate the predicted probability based on the cosine similarity between the text representation z and the image representation x: Where sim(·) represents the cosine similarity function, τ is the temperature parameter, L is the number of categories, y and x represent the true label and input image; The final training goal is to use the cross entropy loss function L ce Calculate the loss and propagate the loss gradient back to the prompt parameter space through the backpropagation algorithm to achieve end-to-end learnable prompt parameter optimization: Here, B represents the batch size.