Adapter-based infrared and visible light parameter efficient transfer learning method

Through the adapter-based efficient transfer learning method of infrared and visible light parameters, the model structure is simplified, the versatility and scalability of the model are improved, the computing resources and time costs are reduced, and the processing effects of infrared and visible light tasks are improved.

CN120671767APending Publication Date: 2025-09-19CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510730087.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing Transformer-based infrared and visible light methods lack universality, have inconsistent performance due to domain differences, consume large amounts of computing resources, and have high time costs, making them difficult to effectively apply in different tasks.

Method used

An adapter-based efficient transfer learning method for infrared and visible light parameters is adopted. Through the infrared modality embedding layer, the external modality prompt adapter, the visible light pre-trained visual model with the adapter inserted, and the target task decoder, the model structure is simplified, only a small amount of backbone network parameters are fine-tuned, and the visual representation is enhanced using the hybrid adapter and the internal prompt adapter.

Benefits of technology

It improves the versatility and scalability of the model, reduces computing resources and time costs, improves the processing effects of infrared and visible light tasks, and can detect and segment targets more accurately.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120671767A_ABST
    Figure CN120671767A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of computer vision, and relates to an adapter-based infrared and visible light parameter efficient transfer learning method, which comprises the following steps of: obtaining an infrared image and a visible light image which are paired, and inputting the infrared image and the visible light image into a trained target task prediction model to obtain a prediction result; the training process of the prediction model comprises the following steps: acquiring multiple groups of paired infrared images xir and visible light images xvis; inputting the image xir and the image xvis into an embedded layer of a corresponding mode to obtain an infrared feature token Zir and a visible light feature token Zvis; inputting the tokens Zir and Zvis into an external modal prompt adapter to obtain an external modal prompt PE; inputting the prompt PE and the Zvis token into the visible light pre-training visual model inserted into the adapter for coding to obtain coding features; inputting the coding features into a target task decoder to obtain a prediction result; updating parameters of the prediction model according to the prediction result until a trained prediction model is obtained; according to the method, a plurality of adapters are introduced to adapt the pre-trained visual model to various infrared and visible light downstream tasks, so that the universality and expandability of the model are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision technology, and in particular relates to an adapter-based efficient transfer learning method for infrared and visible light parameters. Background Art

[0002] In the field of computer vision, research on infrared and visible light tasks is of great significance and is widely used in security monitoring, autonomous driving, remote sensing, and other fields. These tasks aim to utilize the complementary information of infrared and visible light images to improve the perception and understanding of targets.

[0003] Early approaches to infrared and visible light tasks were mostly based on traditional image processing techniques, such as using handcrafted features and machine learning algorithms for image fusion and object detection. However, these methods were limited by the limitations of handcrafted features and struggled to effectively extract key features from infrared and visible light images in complex scenes, resulting in poor performance.

[0004] With the rapid development of deep learning, methods based on convolutional neural networks (CNNs) have gradually become mainstream. Researchers have designed various convolutional structures to attempt to learn features from infrared and visible light images. However, CNNs have difficulties processing cross-modal information, making it difficult to fully exploit the complementary relationship between infrared and visible light images. Furthermore, they are computationally intensive and inefficient when dealing with large amounts of data.

[0005] In recent years, the Transformer architecture has achieved remarkable results in computer vision. Its self-attention mechanism effectively captures long-range dependencies, providing new insights for processing multimodal information. Several Transformer-based methods have begun to be applied to infrared and visible light tasks, improving performance to a certain extent by fine-tuning pre-trained vision models.

[0006] However, existing Transformer-based infrared and visible light methods still have many problems. On the one hand, they usually use complex fusion networks, which are often designed for specific tasks and lack versatility, making it difficult to achieve good results in different infrared and visible light tasks. On the other hand, due to the lack of infrared pre-training models, existing methods often use visual encoders trained on visible light data as the backbone network of the infrared branch, which also leads to problems such as inconsistency due to domain differences. In addition, due to the small scale of commonly used infrared and visible light datasets, when fine-tuning the pre-trained visual model, it is easy to destroy its pre-trained knowledge space, resulting in performance degradation. Finally, the process of fully fine-tuning the pre-trained visual model is resource-intensive and time-consuming, making it unsuitable for rapid deployment in practical applications.

[0007] In summary, with the current trend of utilizing powerful pre-trained vision models, how to simplify the utilization of infrared modalities and apply visible light pre-trained vision models to infrared and visible light tasks in a more efficient and universal way has become an urgent problem to be solved. Summary of the Invention

[0008] To address the above-mentioned problems in the prior art, the present invention adopts an adapter-based efficient transfer learning method for infrared and visible light parameters, comprising: obtaining paired infrared images and visible light images, inputting the paired infrared and visible light images into a trained target task prediction model to obtain a prediction result; the target task model includes: an infrared modality embedding layer, a visible light modality embedding layer, an external modality prompt adapter, a visible light pre-trained visual model inserted into the adapter, and a target task decoder;

[0009] The training process of the target task prediction model includes:

[0010] S1. Obtain an image dataset, where the image dataset includes: multiple sets of paired infrared images and visible light images; input the paired infrared images and visible light images into the embedding layer of the corresponding modality to obtain infrared feature tokens and visible light feature tokens;

[0011] S2. Inputting the infrared feature token and the visible light feature token into the external modal prompt adapter to obtain an external modal prompt;

[0012] S3, inputting the external modality prompt and the visible light feature token into the visible light pre-trained vision model of the adapter for encoding to obtain the encoded features;

[0013] S4. Input the encoded features into the target task decoder to obtain the prediction results;

[0014] S5. Calculate the loss function value based on the prediction results, and update the parameters of the target task prediction model based on the loss function value. When the loss function value is minimized, the trained target task prediction model is obtained.

[0015] Beneficial effects:

[0016] 1. Based on the powerful representation and expression capabilities of the most advanced pre-trained visual models, the present invention innovatively introduces multiple adapters to effectively adapt the pre-trained visual models to a variety of infrared and visible light downstream tasks, including but not limited to semantic segmentation and target detection. This universal framework design enables the model to flexibly switch between different tasks without the need to redesign a complex network structure for each task, thereby improving the versatility and scalability of the model; 2. The present invention utilizes an external modality prompt adapter to encode infrared data into the external modality prompt, avoiding the introduction of additional visual encoders, solving the problem of performance degradation caused by domain offset, and greatly simplifying the model structure; 3. The present invention introduces a hybrid adapter to enable the pre-trained visible light visual model to be used for training images. Efficient fine-tuning of the visual model. When the visual model parameters are frozen, the hybrid adapter can effectively enhance the visual representation of the backbone network of the pre-trained visual model, thereby improving the processing effect of infrared and visible light tasks; 4. The present invention introduces an internal prompt adapter to convert external prompts into internal prompts, which effectively supplements the prior information for the backbone network representation of the visible light pre-trained visual model, so that the model can more accurately grasp the overall characteristics and detailed information of the target when detecting and segmenting the target, thereby effectively improving the processing effect of infrared and visible light tasks; 5. The present invention only fine-tunes a small amount of backbone network parameters. Compared with full fine-tuning, it greatly reduces the computing resources and time cost required for training, and significantly improves the practicality and adaptability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 A flowchart of an adapter-based efficient transfer learning method for infrared and visible light parameters provided by an embodiment of the present invention;

[0018] Figure 2 A flowchart of an adapter-based efficient transfer learning method for infrared and visible light parameters provided by an embodiment of the present invention;

[0019] Figure 3 A schematic diagram of an external modal prompt adapter provided by an embodiment of the present invention;

[0020] Figure 4 A schematic diagram of a hybrid adapter provided in an embodiment of the present invention;

[0021] Figure 5 A schematic diagram of an internal adapter provided by an embodiment of the present invention;

[0022] Figure 6 A schematic diagram of the structure of a mixing unit provided in an embodiment of the present invention;

[0023] Figure 7 A schematic diagram of data in a data set provided by an embodiment of the present invention;

[0024] Figure 8A visualization of the model prediction results provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0025] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0026] This paper proposes an efficient transfer learning method for infrared and visible light parameters based on adapters, such as Figure 1 、 Figure 2 As shown, the method includes: obtaining a paired infrared image and a visible light image, inputting the paired infrared image and visible light image into a trained target task prediction model, and obtaining a prediction result; the target task model includes: an infrared modality embedding layer, a visible light modality embedding layer, an external modality prompt adapter, a visible light pre-trained vision model inserted into the adapter, and a target task decoder;

[0027] The training process of the target task prediction model includes:

[0028] S1. Obtain an image dataset, which includes multiple sets of paired infrared images and visible light images and their annotations; input the paired infrared images and visible light images into the embedding layer of the corresponding modality and map them to the same dimension to obtain infrared feature tokens and visible light feature tokens;

[0029] Infrared images reflect the distribution of thermal radiation in a scene, while visible light images reflect the color information of the scene (RGB channels). The image dataset also includes real labels specific to downstream tasks for model training.

[0030] S2. Inputting the infrared feature token and the visible light feature token into the external modal prompt adapter to obtain an external modal prompt;

[0031] like Figure 3 As shown, the external modal prompt adapter includes: an infrared modal feature down-mapping layer, an attention enhancement unit, a visible light modal feature down-mapping layer, a mixing unit, and a dual-modal feature up-mapping layer; the external modal prompt adapter processes the infrared feature token and the visible light feature token including:

[0032] S21, inputting the infrared feature token and the visible light feature token into the feature down-mapping layer of the corresponding modality respectively to obtain infrared mapping features and visible light mapping features;

[0033] S22, inputting the visible light mapping feature into the mixing unit to obtain a mixed feature;

[0034] S23, inputting the infrared mapping feature into the attention enhancement unit to obtain an enhanced feature;

[0035] S24. Element-wise addition is performed on the hybrid features and the enhanced features, and the result of the addition is input into the bimodal feature upper mapping layer to obtain an external modal prompt that fuses the bimodal information.

[0036] The above steps S21-S24 are the data processing process of the external modality prompt adapter proposed by the present invention, which introduces an asymmetric processing structure for infrared and visible light data to generate prompts. The feature down-mapping layer and up-mapping layer of the external modality prompt adapter map the features to lower dimensions to reduce the number of parameters; the hybrid processing unit uses depth convolution and 1×1 convolution aggregation to effectively capture the rich features of the mapped visible light features and reduce the information loss problem caused by down-mapping; the attention enhancement unit SimAM captures the structural information of the infrared data; the output elements of the hybrid unit and the attention enhancement unit are added and mapped to the original dimension through the up-mapping layer to generate the external modality prompt, avoiding the introduction of additional visual encoders, thereby greatly simplifying the model structure and achieving efficient interaction and feature fusion between modalities.

[0037] S3. Insert the external modality prompt and visible light feature token input into the adapter’s visible light pre-trained vision model for encoding to obtain the encoded features.

[0038] The adapter-inserted visible light pre-trained vision model includes L layers of encoders, and each layer of encoder is inserted with two types of adapters, namely, a hybrid adapter and a prompt adapter; preferably, the encoder is a Transformer encoder.

[0039] The visible light pre-trained vision model plugged into the adapter processes external modality cues and visible light feature tokens including:

[0040] S31. Add the visible light feature token and the external modal prompt element-wise, input the result of the element-wise addition and the external modal prompt into a first encoder to obtain a first encoding feature; and update the external modal prompt according to the first encoding feature to obtain a first external modal prompt.

[0041] S32. Input the first coding feature and the first external modal prompt into a second encoder to obtain a second coding feature; and update the first external modal prompt according to the second coding feature to obtain a second external modal prompt.

[0042] S33, inputting the previous coding feature and the previous external modal prompt into the current encoder to obtain a current coding feature; updating the previous external modal prompt according to the current coding feature to obtain a current external modal prompt;

[0043] The encoder includes: a multi-head self-attention module and a feedforward neural network; the inserted adapters include: a first hybrid adapter, a second hybrid adapter, a first internal prompt adapter, and a second internal prompt adapter; the current encoder processes the previous encoded feature and the previous external modal prompt including:

[0044] S331. Input the previous encoding feature into the multi-head self-attention module of the current encoder to obtain a self-attention feature;

[0045] S332. Input the self-attention feature into the first hybrid adapter inserted by the current encoder to obtain a first hybrid feature; wherein the hybrid adapter is used to optimize the feature;

[0046] like Figure 4 As shown, the hybrid adapter includes: a feature down-mapping layer, a hybrid unit, and a feature up-mapping layer; the first hybrid adapter inserted by the current encoder processes the self-attention feature by inputting the self-attention feature into the feature down-mapping layer, inputting the feature output from the feature down-mapping layer into the hybrid unit, and inputting the feature output from the hybrid unit into the feature up-mapping layer to obtain the first hybrid feature. The specific formula includes:

[0047]

[0048] in, and They are the down-mapping and up-mapping parameters of dimension d2, and r is the feature The dimension, superscript T indicates transposition, GeLu is the activation function, g h (·) is a mixed unit, is a mixed feature, is the self-attention feature.

[0049] The hybrid adapter follows the architecture of the standard adapter as a whole, but introduces additional hybrid units to further enhance the visual signal. When the visual model parameters are frozen, the hybrid adapter can effectively enhance the visual representation of the backbone network of the pre-trained visual model.

[0050] like Figure 6As shown, the hybrid unit (Hybrid Operation) includes: a partial convolution module (Partial Conv), a first 1×1 convolution layer and a second 1×1 convolution layer, and the partial convolution module includes a depth convolution module (Depth-wiseConv); the hybrid unit processes the input features (i.e., the features output by the feature mapping layer) including: reshaping the input features into a feature map structure to obtain an initial feature map; inputting the initial feature map into the partial convolution module to obtain a partial convolution feature map; performing a residual connection (i.e., element addition) on the initial feature map and the partial convolution feature map to obtain a first fused feature map; inputting the first fused feature map into the first 1×1 convolution layer to obtain a first convolution feature map; performing a residual connection (i.e., element addition) on the first convolution feature map and the first fused feature map to obtain a second fused feature map; inputting the second fused feature map into the second 1×1 convolution layer to obtain a second convolution feature map; and reorganizing the second convolution feature map into the structure of the input feature token to obtain the output feature.

[0051] Preferably, before performing residual connection on the first convolution feature map and the first fusion feature map, batch normalization BN is performed on the first convolution feature map, and the batch normalized feature map is processed using the Relu activation function.

[0052] The partial convolution module processes the initial feature map by dividing the initial feature map into two groups of features according to the channel. The number of channels of the two groups of features is C and C respectively. P and CC P ; Set the number of channels to C P The features of are input into the deep convolution module, and the features after deep convolution and the number of channels are CC P The features of the initial feature map are spliced ​​to obtain a partial convolution feature map; where C is the number of channels of the initial feature map, C P The dimension size after segmentation, the default is C / 4.

[0053] S333: Input the previous external modal prompt into the first internal prompt adapter inserted by the current encoder to obtain a first internal prompt feature; wherein the internal prompt adapter is used to convert the external prompt into an internal prompt;

[0054] like Figure 5 As shown, the internal prompt adapter includes: a feature down-mapping layer, an attention enhancement unit, and a feature up-mapping layer. The first internal prompt adapter inserted by the current encoder processes the previous external modality prompt by inputting the previous external modality prompt into the feature down-mapping layer, inputting the features output by the feature down-mapping layer into the attention enhancement unit, and inputting the features output by the attention enhancement unit into the feature up-mapping layer to obtain the first internal prompt feature. The specific formula is:

[0055]

[0056] in, and are the lower and upper mapping parameters of dimension d3, respectively, and r is the previous external modal prompt The dimension of GeLu is the activation function, g s (·) is the attention enhancement unit, It is the first internal prompt feature.

[0057] Theoretically, directly fusing external modal cues with internal features will lead to feature conflicts. Therefore, the present invention proposes an internal cue adapter to solve this problem. Similar to the hybrid adapter, the internal cue adapter still maps features to a lower dimension, and uses an attention enhancement unit to effectively capture prior information. It is then mapped to the original dimension through the upper mapping layer to generate internal cues. At each layer of the pre-trained visual model, the internal cue features will be fused with the output of the hybrid adapter, thereby effectively supplementing the prior information for the backbone network representation of the pre-trained visual model. This enables the model to more accurately grasp the overall features and detailed information of the target when detecting and segmenting the target, effectively improving the processing effect of infrared and visible light tasks.

[0058] Attention Enhancement Unit g s (·) does not contain any parameters, but enhances features by calculating channel-aware spatial attention weights applied to features. The specific formula is:

[0059]

[0060] in, is the input feature of the attention enhancement unit, is the output feature of the attention enhancement unit, λ is a hyperparameter, is the input feature The spatial attention weight of channel t, B is the input feature The weights of all channels The combined vector, ⊙ is the Hadamard product, δ 2 、 is the input feature The variance and mean of are calculated by the following formulas:

[0061]

[0062] Among them, H and W are input features The height and width of i are the input features The index of the element.

[0063] In this way, the attention enhancement unit derives the attention scaling value for the down-mapped features by calculating the correlation between neurons, thereby enhancing the early features in the external modality cues and supplementing the subsequent feature fusion.

[0064] S334, fusing the attention feature, the first mixed feature, and the first internal prompt feature to obtain a fused feature;

[0065] S335: Input the fused features into the feedforward neural network of the current encoder, and input the features output by the feedforward neural network into the second hybrid adapter inserted into the current encoder to obtain a second hybrid feature;

[0066] S336 , inputting the previous external modal prompt into the second internal prompt adapter inserted into the current encoder to obtain a second internal prompt feature;

[0067] S337: Fuse the feature output by the feedforward neural network, the second mixed feature, and the second internal prompt feature to obtain the current coding feature.

[0068] The previous external modal prompt is updated according to the current coding feature, that is, the current coding feature directly replaces the previous external modal prompt as the current external modal prompt.

[0069] S34, repeat step S33 until the encoding features output by the last layer encoder are obtained

[0070] S4. Encoding features Input the target task decoder to obtain the prediction result;

[0071] S5. Calculate the loss function value based on the prediction results, and update the parameters of the target task prediction model based on the loss function value. When the loss function value is minimized, the trained target task prediction model is obtained.

[0072] During the overall model training, the encoder portion of the original pre-trained vision model will be frozen, and only the infrared modality embedding layer, the external modality prompt adapter, the adapter inserted by the visible light pre-trained vision model, and the target task decoder will be trained. Therefore, the model training process can be modeled as:

[0073]

[0074] Among them, θ IV represents the optimization process of the prediction model, represents the image dataset for the target task, represents the loss function of the target task, θ is all the updateable parameters of the model, and φ is the target task decoder, which can encode the features Decoded into task-specific outputs, y is the true label of the target task, that is, the annotation map.

[0075] In one embodiment, Figure 7 As shown in the figure, when the target task is semantic segmentation, the image dataset for semantic segmentation includes: infrared images, visible light images, and annotation maps. The annotation maps are pixel-level labels, and each pixel corresponds to a semantic category (background, curve, building, car, pedestrian, etc.). The target task decoder outputs the predicted category of each pixel and uses cross entropy loss to train the model based on the predicted category and the actual category of each pixel. The visualization example of the model prediction result is as follows Figure 8 shown.

[0076] The above embodiments further illustrate the purpose, technical solutions and advantages of the present invention in detail. It should be understood that the above embodiments are only preferred implementation plans of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made to the present invention within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. An adapter-based efficient transfer learning method for infrared and visible light parameters, characterized in that: include: Obtain paired infrared images and visible light images, input the paired infrared images and visible light images into the trained target task prediction model to obtain the prediction results; The target task model includes: an infrared modality embedding layer, a visible light modality embedding layer, an external modality cue adapter, a visible light pre-trained vision model inserted into the adapter, and a target task decoder; The training process of the target task prediction model includes: S1. Obtain an image dataset, where the image dataset includes: multiple sets of paired infrared images and visible light images; input the paired infrared images and visible light images into the embedding layer of the corresponding modality to obtain infrared feature tokens and visible light feature tokens; S2. Inputting the infrared feature token and the visible light feature token into the external modal prompt adapter to obtain an external modal prompt; S3, inputting the external modality prompt and the visible light feature token into the visible light pre-trained vision model of the adapter for encoding to obtain the encoded features; S4. Input the encoded features into the target task decoder to obtain the prediction results; S5. Calculate the loss function value based on the prediction results, and update the parameters of the target task prediction model based on the loss function value. When the loss function value is minimized, the trained target task prediction model is obtained.

2. The adapter-based efficient transfer learning method for infrared and visible light parameters according to claim 1, characterized in that: The external modal prompt adapter includes: an infrared modal feature down-mapping layer, an attention enhancement unit, a visible light modal feature down-mapping layer, a mixing unit, and a bimodal feature up-mapping layer; the external modal prompt adapter processes infrared feature tokens and visible light feature tokens including: S21, inputting the infrared feature token and the visible light feature token into the feature down-mapping layer of the corresponding modality respectively to obtain infrared mapping features and visible light mapping features; S22, inputting the visible light mapping feature into the mixing unit to obtain a mixed feature; S23, inputting the infrared mapping feature into the attention enhancement unit to obtain an enhanced feature; S24. Add the mixed features and the enhanced features, and input the added result into the bimodal feature upper mapping layer to obtain the external modality prompt.

3. The adapter-based efficient transfer learning method for infrared and visible light parameters according to claim 1, characterized in that: The visible light pre-trained vision model inserted into the adapter includes multiple layers of encoders, each of which is inserted into the adapter; The visible light pre-trained vision model plugged into the adapter processes external modality cues and visible light feature tokens including: S31, adding the visible light feature token and the external modal prompt, inputting the added result and the external modal prompt into a first encoder to obtain a first encoding feature; updating the external modal prompt according to the first encoding feature to obtain a first external modal prompt; S32. Input the first coding feature and the first external modal prompt into a second encoder to obtain a second coding feature; and update the first external modal prompt according to the second coding feature to obtain a second external modal prompt. S33, inputting the previous coding feature and the previous external modal prompt into the current encoder to obtain a current coding feature; updating the previous external modal prompt according to the current coding feature to obtain a current external modal prompt; S34. Repeat step S33 until the encoding features output by the last layer encoder are obtained.

4. The adapter-based efficient transfer learning method for infrared and visible light parameters according to claim 3, characterized in that: The encoder includes: Multi-head self-attention module and feed-forward neural network; The inserted adapters include: a first hybrid adapter, a second hybrid adapter, a first internal prompt adapter, and a second internal prompt adapter; the current encoder processes the previous encoded feature and the previous external modality prompt including: S331. Input the previous encoding feature into the multi-head self-attention module of the current encoder to obtain a self-attention feature; S332. Input the self-attention feature into the first hybrid adapter inserted by the current encoder to obtain a first hybrid feature; S333, inputting the previous external modal prompt into the first internal prompt adapter inserted into the current encoder to obtain a first internal prompt feature; S334, fusing the attention feature, the first mixed feature, and the first internal prompt feature to obtain a fused feature; S335: Input the fused features into the feedforward neural network of the current encoder, and input the features output by the feedforward neural network into the second hybrid adapter inserted into the current encoder to obtain a second hybrid feature; S336 , inputting the previous external modal prompt into the second internal prompt adapter inserted into the current encoder to obtain a second internal prompt feature; S337. Fuse the feature output by the feedforward neural network, the second mixed feature, and the second internal prompt feature to obtain the current encoding feature.

5. The adapter-based efficient transfer learning method for infrared and visible light parameters according to claim 4, characterized in that: The hybrid adapter includes: a feature down-mapping layer, a hybrid unit and a feature up-mapping layer; the first hybrid adapter inserted by the current encoder processes the self-attention feature including: inputting the self-attention feature into the feature down-mapping layer, inputting the feature output from the feature down-mapping layer into the hybrid unit, and inputting the feature output from the hybrid unit into the feature up-mapping layer to obtain a first hybrid feature.

6. The adapter-based efficient transfer learning method for infrared and visible light parameters according to claim 5, characterized in that: The mixing unit includes: a partial convolution module, a first convolution layer and a second convolution layer, and the partial convolution module includes a depth convolution module; the mixing unit processes the features output by the feature mapping layer, including: reorganizing the features output by the feature mapping layer into a feature map structure to obtain an initial feature map; inputting the initial feature map into the partial convolution module to obtain a partial convolution feature map; performing a residual connection between the initial feature map and the partial convolution feature map to obtain a first fused feature map; inputting the first fused feature map into the first convolution layer to obtain a first convolution feature map; performing a residual connection between the first convolution feature map and the first fused feature map to obtain a second fused feature map; inputting the second fused feature map into the second convolution layer to obtain a second convolution feature map; reorganizing the second convolution feature map into the structure of the features output by the feature mapping layer to obtain an output feature.

7. The adapter-based efficient transfer learning method for infrared and visible light parameters according to claim 6, characterized in that: The partial convolution module processes the initial feature map by dividing the initial feature map into two groups of features according to the channel. The number of channels of the two groups of features is C and C respectively. P and CC P ; Set the number of channels to C P The features of are input into the deep convolution module, and the features after deep convolution and the number of channels are CC P The features of the initial feature map are spliced ​​to obtain a partial convolution feature map; where C is the number of channels of the initial feature map, C P is the channel dimension size after segmentation.

8. The adapter-based efficient transfer learning method for infrared and visible light parameters according to claim 4, characterized in that: The internal prompt adapter includes: a feature down-mapping layer, an attention enhancement unit and a feature up-mapping layer; the processing process of the first internal prompt adapter inserted by the current encoder on the previous external modal prompt includes: inputting the previous external modal prompt into the feature down-mapping layer, inputting the features output by the feature down-mapping layer into the attention enhancement unit, and inputting the features output by the attention enhancement unit into the feature up-mapping layer to obtain the first internal prompt feature.

9. The adapter-based efficient transfer learning method for infrared and visible light parameters according to claim 8, characterized in that: The attention enhancement unit includes processing the features output by the feature mapping layer including: in, is the feature output by the feature mapping layer, is the feature output by the attention enhancement unit, λ is a hyperparameter, Features The spatial attention weight of channel t, δ 2 、 Features The mean and variance of B are the characteristics The spatial attention weights of all channels The combined weight vector is ⊙, which is the Hadamard product.

10. The adapter-based efficient transfer learning method for infrared and visible light parameters according to claim 1, characterized in that: Updating the parameters of the prediction model according to the loss function value includes updating the parameters of the infrared modality embedding layer, the external modality prompt adapter, the adapter inserted by the visible light pre-trained vision model, and the target task decoder according to the loss function value.