An optical-infrared target detection method of cross-modal attention fusion mechanism
By using implicit neural representations and the Mamba architecture, the RGB-T object detection method is extended from the discrete space to the continuous function space, solving the bottlenecks of cross-modal alignment and interaction in traditional methods, achieving high robustness and generalization ability, and improving detection performance.
Patent Information
- Application Number
- CN202511430366.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-09
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2045-10-09
AI Technical Summary
Existing RGB-T target detection methods lack robustness and accuracy in low-light environments, and traditional discrete mesh fusion limits cross-modal alignment and interaction capabilities, resulting in blurred boundaries and limited scale adaptability.
By combining implicit neural representation with the Mamba architecture, cross-modal fusion is extended from discrete space to continuous function space. Continuous representation of features is achieved through implicit feature interaction units, and adaptive aggregation is performed using implicit attention mechanisms. Multi-scale feature alignment and global context interaction are achieved by combining implicit cross-scale fusion units.
It improves the robustness and generalization ability of object detection, achieves sub-pixel-level feature interaction and alignment, maintains the computational efficiency of the model, and enhances the detection performance under different lighting and weather conditions.
Smart Images

Figure CN120913023B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of image processing, and particularly relates to an optical-infrared target detection method based on a cross-modal attention fusion mechanism. BACKGROUND
[0002] Object detection is a basic task in the field of computer vision, and has been widely applied to the fields of automatic driving and intelligent monitoring. At present, most detection methods mainly rely on RGB images, which can provide rich visual features under good lighting conditions. However, its performance will decrease significantly in low-light environments, such as at night or in bad weather. In contrast, thermal infrared (TIR) sensors achieve illumination-invariant perception by capturing thermal radiation, making them robust in poor visibility conditions. However, TIR data usually has the problem of insufficient spatial details and resolution, and is difficult to meet the requirements when used alone. Based on this, the collaborative fusion of RGB and TIR modalities has become an important means to improve the performance of target detection. The high-resolution structural information provided by the RGB image and the robust thermal radiation features of the TIR image have high complementarity in the semantic level. Visible-infrared fusion (RGB-T) target detection aims to fully exploit the complementary characteristics of the two modalities, improve the robustness and accuracy of the detection system under various lighting conditions and weather environments, and meet the intelligent perception needs in all-day and multi-scene.
[0003] At present, most RGB-T target detection methods focus on deep feature fusion. In order to overcome the limitations of early fusion schemes based on splicing, researchers have recently explored learnable fusion strategies. Among them, the attention mechanism has become the dominant paradigm—from modeling global correlations to constructing iterative cross-attention transformers. Customized modules can also be used to improve fusion effects, including dedicated convolutional neural network architectures, segmentation aggregation, and difference perception designs. In addition, auxiliary tasks such as modality perception reconstruction are used to improve the discrimination ability. However, most methods are still based on discrete 2D grids, which not only limits their ability to model continuous spatial relationships, but also easily leads to problems such as boundary ambiguity, alignment errors, and limited scale adaptability.
[0004] In summary, in order to overcome the inherent limitations of traditional discrete grid fusion in cross-modal alignment and interaction, there is an urgent need to develop a new implicit feature cross-attention fusion framework for RGB-T target detection, which expands cross-modal fusion from discrete space to continuous function space to effectively solve the above problems. SUMMARY
[0005] The optical-infrared target detection method with the cross-modal attention fusion mechanism aims to provide an optical-infrared target detection method with a cross-modal attention fusion mechanism, which combines implicit neural representation with a Mamba architecture, expands cross-modal fusion from discrete space to continuous function space, solves lossless alignment of different scale features, improves detection robustness and generalization ability, and maintains the calculation efficiency of the model.
[0006] The present application is realized by the following technical solutions:
[0007] The optical-infrared target detection method with the cross-modal attention fusion mechanism comprises the following steps:
[0008] S1, input the images of two modalities of visible light and infrared into respective feature extraction backbone networks to obtain discrete multi-scale feature maps corresponding to each modality;
[0009] S2, input the discrete multi-scale feature maps into an implicit feature interaction unit, and convert the discrete features into a feature mapping function in a continuous space domain through an implicit neural representation module in the unit; the feature mapping function can output a feature vector corresponding to a given two-dimensional space coordinate, realizing continuous representation of the feature;
[0010] S3, in the continuous space domain, use an implicit attention mechanism to adaptively aggregate cross-modal features; use the features of the first modality as a query to query and weightedly aggregate the continuous feature mapping function of the second modality, to generate fusion features precisely aligned with the first modality, thereby completing sub-pixel level feature interaction and alignment;
[0011] S4, use an implicit cross-scale fusion unit to realize alignment of multi-scale fusion features, and then use an efficient state space model in the unit to model long-distance dependencies and interact global context information of features at different levels, realize scale restoration of features at different scales, and obtain multi-scale fusion features after cross-scale feature interaction;
[0012] S5, input the final multi-scale fusion features into a detection neck and head to decode and output position and category information of the target in the image.
[0013] Further, in step S1, a double-branch feature extraction structure is used, a pair of visible light images and infrared images that have been spatially registered are input into two parallel feature extraction backbone networks based on convolutional neural networks, and feature maps with decreasing spatial resolution and increasing semantic information are extracted from multiple different depth stages of the two backbone networks, to generate a set of discrete multi-scale feature maps for the visible light modality and the infrared modality respectively.
[0014] Further, step S2 specifically comprises the following steps:
[0015] Any discrete feature map obtained from step S1 is regarded as the sampling result of a continuous function defined by the following mapping relationship:
[0016]
[0017] wherein, denotes any query coordinate within the normalized two-dimensional continuous coordinate domain Ω, denotes the C-dimensional feature vector corresponding to the coordinate point, and represents the C-dimensional feature space; in order to evaluate the value of the continuous function at any query coordinate , a differentiable sampling operator is adopted to calculate by the following formula:
[0018]
[0019] wherein, is the original discrete feature map, is a differentiable sampling operator, which is specifically implemented by locating the four nearest neighbor grid points of the query coordinate in the discrete feature map and performing bilinear interpolation operation on the feature vectors of the four points; thus, the inherent discrete feature grid can be converted into a continuous and differentiable feature field.
[0020] Further, step S3 is specifically: defining the query in the traditional attention mechanism as a coordinate-based position encoding, replacing the traditional content-dependent query with a coordinate-based query mechanism, so that the attention process is directly affected by the condition of the absolute spatial position; given the input feature map , the total number of query points is , the query tensor is constructed by filling each spatial position with its corresponding normalized coordinate vector ; for any query point , the corresponding query tensor is the two-dimensional coordinate of the point, as shown in the formula:
[0021]
[0022] For each query position , the key is obtained by sampling from the continuous feature field of the guide feature at the same coordinate:
[0023]
[0024] wherein, is the discrete feature map of the guide modality;
[0025] Querying position The value comes from the continuous field of source modality, which is used to enhance:
[0026]
[0027] where, is the discrete feature map of source modality; the final preserves the original modality information, while the attention weight is used to adaptively adjust ;
[0028] The above-sampled and vectors are linearly projected and reshaped into groups, each with a dimension of ; the query coordinates are also expanded to adapt to multi-head calculation;
[0029] The matrix multiplication is used to replace the traditional dot product to calculate the interaction between content and position, and the position query is multiplied by the content to produce a 2D directional attention :
[0030]
[0031] Then, this scalar score is used to condition ; As an implicit function, the correlation score is decoded into a modulation vector :
[0032]
[0033] where, the vector provides a set of generic-specific weights; finally, the output is obtained by calculating the Hadamard product of these weights and the value vector , thus obtaining the enhanced feature vector :
[0034] .
[0035] Further, step S4 specifically comprises the following steps:
[0036] For any feature vector located at coordinates , first generate an enhanced representation by concatenating it with the normalized coordinates, and then process the vector through a multi-layer perceptron to output the converted features at the target scale :
[0037]
[0038] By conditionalizing the transformation in spatial coordinates, It learns a spatial adaptive function that enables intelligent feature sampling;
[0039] To effectively model the global dependencies of multi-scale features aligned across spatial dimensions, a sequence processor consisting of multiple stacked Mamba blocks is introduced. This process treats the concatenated multi-scale feature sequence as a whole and then applies positional encoding. The sequence is then fed into the Mamba block stack, as shown below:
[0040]
[0041]
[0042] in, This represents the feature information at each scale segmented from the output sequence with a global aggregation context. This represents feature tensors from different scales, which have been aligned in continuous space by a coordinate-aware feature scale transformation module.
[0043] The Mamba processor scans the entire feature sequence to capture long-range dependencies; then, a coordinate-aware feature scaling module is reapplied to restore each feature map to its original channel dimensions; finally, these globally enhanced features are coupled with the initial features via residual connections. integrated:
[0044] .
[0045] Compared with the prior art, the beneficial effects of the present invention are:
[0046] The detection method of this invention incorporates an implicit feature interaction unit and an implicit cross-scale fusion unit. The implicit feature interaction unit, by introducing implicit neural representations, elevates multimodal features from a traditional discrete pixel grid to a continuous function space, learning and constructing a continuous mapping function from spatial coordinates to feature vectors. This design aims to leverage a novel implicit attention mechanism to adaptively aggregate target modality feature information at any query coordinate point, thereby overcoming the bottlenecks of traditional methods in receptive field and alignment accuracy. The implicit cross-scale fusion unit is responsible for processing feature information at different levels. Based on the continuity advantage of implicit neural representations, it achieves lossless alignment of features at different scales, fully exploiting complementary information between scales to improve fusion performance. By combining implicit neural representations with the Mamba architecture, this invention extends cross-modal fusion from a discrete space to a continuous function space, improving detection robustness and generalization ability while maintaining the computational efficiency of the model. Attached Figure Description
[0047] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required to be used in the embodiments will be briefly introduced as follows. It should be understood that the following drawings only show some of the embodiments of the present application, and therefore should not be considered as a limitation to the scope, and other related drawings can also be obtained by those of ordinary skill in the art without creative effort based on these drawings.
[0048] Figure 1 Flow chart of the optical-infrared target detection method of the cross-modal attention fusion mechanism of the present application;
[0049] Figure 2 Network structure diagram of the optical-infrared target detection method of the cross-modal attention fusion mechanism of the present application;
[0050] Figure 3 Comparison diagram of fusion results of different visible-infrared image fusion detection algorithms on a visible-infrared image pair. DETAILED DESCRIPTION
[0051] The present application will be further described below in conjunction with the embodiments:
[0052] The present application will be further described below in conjunction with the embodiments:
[0053] It should be noted that: similar labels and letters represent similar items in the following drawings, therefore, once an item is defined in one drawing, it does not need to be further defined and explained in the subsequent drawings. Meanwhile, in the description of the present application, the terms "first", "second" and the like are only used for distinguishing description, and cannot be understood as indicating or implying relative importance.
[0054] As Figures 1-2As shown, the optical-infrared target detection method of the cross-modal attention fusion mechanism of the application comprises: converting input visible light and infrared images into multi-scale features containing different levels of information required for subsequent processing through parallel feature extraction networks; through an implicit feature interaction unit, the discrete features are lifted to a continuous function space, and a coordinate-driven attention mechanism is used to realize accurate alignment and deep fusion between visible light and infrared modalities; using an implicit cross-scale fusion unit, long-distance dependency modeling and global context information interaction are performed on features of different levels to obtain multi-scale fusion features after cross-scale feature interaction; the feature map obtained after cross-modal and cross-scale double fusion, which is highly information-rich, is input into a detection network, and finally high-precision target position and category information is output.
[0055] Specifically, the following steps are included:
[0056] S1, input the images of the two modalities of visible light and infrared into respective feature extraction backbone networks to obtain discrete multi-scale feature maps corresponding to each modality.
[0057] Specifically, a double-branch feature extraction structure is adopted, a pair of spatially registered visible light images and infrared images are input into two parallel feature extraction backbone networks based on convolutional neural networks, and feature maps with decreasing spatial resolution and increasing semantic information are extracted from multiple different depths of the two backbone networks, thereby generating a set of discrete multi-scale feature maps for the visible light modality and the infrared modality.
[0058] S2, input the discrete multi-scale feature maps into an implicit feature interaction unit, and through the implicit neural representation module in the unit, convert the discrete features into a feature mapping function in the continuous spatial domain; the feature mapping function can output the feature vector corresponding to the coordinate point according to any given two-dimensional spatial coordinate, thereby realizing continuous representation of the features.
[0059] Specifically, any discrete feature map obtained from step S1 is regarded as a sampling result of a potential and more fundamental continuous function, and the continuous function is defined by the following mapping relationship:
[0060]
[0061] wherein, represents any query coordinate in the normalized two-dimensional continuous coordinate domain Ω, represents the C-dimensional feature vector corresponding to the coordinate point, and represents a C-dimensional feature space; in order to evaluate the value of the continuous function at any query coordinate , a differentiable sampling operator is used to calculate it by the following formula:
[0062] wherein, is the original discrete feature map, is a differentiable sampling operator, which is specifically implemented by locating to the four nearest neighbor grid points of the query coordinate in the discrete feature map and performing bilinear interpolation operation on the feature vectors of the four points; thus, the inherent discrete feature grid can be converted into a continuous and differentiable feature field, thereby breaking the limitation that the model can only operate on fixed pixel points and endowing the model with the ability to accurately query and process features at any sub-pixel position, providing a basis for the subsequent continuous spatial attention mechanism.
[0063] S3, in the continuous spatial domain, using an implicit attention mechanism for adaptive aggregation of cross-modal features; specifically, the features of the first modality are taken as queries, and the query and weighted aggregation are performed on the continuous feature mapping function of the second modality to generate fused features that are precisely aligned with the first modality, thereby completing the feature interaction and alignment at the sub-pixel level.
[0064] Specifically, the coordinate-driven attention mechanism based on the implicit continuous field defines the continuous feature field of one modality as the guide modality and the other as the source modality, redefines the traditional query (Query), key (Key) and value (Value) modeling method, and explicitly incorporates spatial information as guide information, fundamentally changing the attention from pure content dependence to spatial awareness. This method uses such a mechanism to achieve mutual enhancement between visible light and thermal infrared image feature streams, enabling each modality to dynamically obtain complementary information provided by the other.
[0065] Specifically, the query (Quey) in the traditional attention mechanism is defined as a coordinate-based position encoding, and the coordinate-based query mechanism is used to replace the traditional content-dependent query, so that the attention process is directly affected by the condition of the absolute spatial position. Given the input feature map , the total number of query points is , by filling each spatial position with its corresponding normalized coordinate vector , a query tensor is constructed. For any query point , the corresponding query tensor is the two-dimensional coordinate of the point, as shown in the formula:
[0066]
[0067] This explicit spatial guidance is crucial for learning a continuous mapping from coordinates to feature modulation weights, thereby achieving robust alignment across modalities.
[0068] The mechanism treats the query (Key) in the traditional attention mechanism as an information vector of guiding features. For each query position , the key is sampled from the continuous feature field of guiding features at the same coordinates:
[0069] ,
[0070] where, is the discrete feature map of the guiding modality. Crucially, this formalization establishes a direct mapping between the spatial position and its corresponding semantic feature in the guiding modality. This provides an explicit, position-specific context that enables the implicit attention mechanism to learn how to best augment the features of the source modality at each point of the feature map.
[0071] Similarly, the value (Value) in the attention mechanism is treated as the raw information of the source features. The value at query position comes from the continuous feature field of the source modality to augment:
[0072]
[0073] where, is the discrete feature map of the source modality. The final preserves the original modality information, while the attention weights (computed by the interaction of and ) are used to adaptively adjust to selectively enhance complementary features while preserving the integrity of the original source content.
[0074] To implement multi-head attention, the above sampled and vectors are linearly projected and reshaped into groups, each with a dimension of . Similarly, the query coordinates are also expanded to adapt to multi-head computation.
[0075] Based on the modeling of the attention component implemented by design, a new type of modulation process is introduced, the core of which is the generation and application of attention weights, aiming to generate high-dimensional, content-aware modulation vectors to achieve fine-grained, channel-level feature adaptation. Matrix multiplication is used to replace the traditional dot product to calculate the interaction between content and position, and the product of the position query and the content reference produces a 2D directional attention :
[0076]
[0077] This scalar score is then used to condition . As an implicit function, this simple correlation score is decoded into a rich modulation vector :
[0078]
[0079] where the vector provides a set of channel-specific weights. Finally, the output is obtained by computing the Hadamard product (element-wise multiplication) of these weights with the value vector , resulting in the enhanced feature vector :
[0080]
[0081] This implicit attention modulation has two key advantages. First, it learns a continuous and universal fusion rule that is not restricted by a fixed grid, thus improving generalization ability and alignment accuracy. Second, by generating a modulation vector instead of a scalar, fine-grained, channel-wise feature recalibration is achieved. This enables the model to selectively amplify complementary information and suppress redundant features on each channel, resulting in a more discriminative and semantically consistent fusion representation.
[0082] The process is executed in a symmetric dual-flow manner, enabling bidirectional feature enhancement. This design promotes strong synergy: the enhanced features of one modality can serve as better guidance for the feature fusion of the other modality. This forms a virtuous cycle of mutual reinforcement, where the RGB and TIR feature representations gradually improve through interaction, ultimately producing a more discriminative and powerful fusion output.
[0083] S4, using an implicit cross-scale fusion unit to realize the alignment of multi-scale fusion features, then, through the efficient state space model in the unit, long-distance dependence modeling and global context information interaction of features at different levels are realized, finally, the scale restoration of each scale feature is realized, and the multi-scale fusion features after cross-scale feature interaction are obtained.
[0084] Specifically, first, in order to realize the lossless alignment of features at different scales, a coordinate-aware feature scale transformation module is introduced. Unlike spatial invariant methods such as interpolation, which often introduce artifacts and reduce feature quality, this method learns a continuous, position-dependent mapping. For any feature vector located at coordinate , first, an enhanced representation is generated by concatenating it with the normalized coordinates. This vector is then processed by a multi-layer perceptron , outputting the transformed feature at the target scale :
[0085]
[0086] By conditioning the transformation on spatial coordinates, a spatially adaptive function is learned, enabling intelligent sampling of features. This preserves fine details and object boundaries more effectively than traditional grid-based operations. The process generates a set of high-fidelity, information-rich features that are perfectly aligned in continuous space, laying the foundation for efficient global fusion.
[0087] To effectively model the global dependencies of multiscale features that are aligned across space, a sequential processor composed of multiple stacked Mamba blocks is introduced. The process treats the concatenated multiscale feature sequence as a whole and then applies position encoding . Subsequently, the sequence is input into the Mamba block stack. This is shown in the following equation:
[0088]
[0089]
[0090] where, represents each scale of feature information partitioned from the output sequence with global aggregation context, represents feature tensors from different scales that have been aligned in continuous space by the coordinate-aware feature scale transformation module.
[0091] With its selective state space mechanism and linear complexity, the Mamba processor can efficiently scan the entire feature sequence to capture long-range dependencies. This facilitates bidirectional information flow, enabling high-resolution details at the shallow layers to interact with high-level semantics at the deep layers. The mechanism is particularly good at preserving signals from small objects. Subsequently, the coordinate-aware feature scale transformation module is reapplied to restore each feature map to the original channel dimension. Finally, these globally enhanced features are integrated with the initial features through a residual connection:
[0092]
[0093] This residual design constrains the module to learn incremental enhancements, ensuring that the final output features preserve the original high-fidelity information while enriching the global and multiscale context through deep fusion. The process enhances key target signals and suppresses irrelevant background noise, providing a more accurate feature representation for the final detection task.
[0094] S5, input the final multiscale fusion features into the detection neck and head to decode and output the location and category information of the target in the image.
[0095] In order to prove the effect of the detection method of the present application, a cross-modal fusion detection experiment is carried out on the visible light thermal infrared image pairs obtained by the thermal imaging and visual camera installed on the vehicle, and the new fusion method at the present stage is compared, and the results are shown in Figure 3 The visualization comparison clearly reveals the significant superiority of the present application compared with other methods in complex real scenes. In the face of small targets far away that are difficult to detect due to dim light, other methods have obvious missed detection, while the present application successfully realizes the accurate capture of the target, showing excellent recall ability. In addition, for the main target that can be detected by all methods, the boundary box regression quality of the present application is also far superior to the opponent, and the detection box generated is more compact and more consistent with the real label. Table 1 shows that the present application has reached the current best level in many key indicators compared with the current advanced multi-modal fusion detection method. In summary, through multi-level innovative design, the present application has achieved comprehensive surpassing in recall rate, positioning accuracy and robustness of detection, and has proved the advanced perception ability and practical application value of the present application in all-weather and multi-scene.
[0096] Table 1 Quantitative evaluation of fusion experiment of the present application and comparative algorithm
[0097] Method People (%) Car (%) Bicycle (%) mAP50 (%) mAP75 (%) RGB only 60.8 73.9 37.2 57.3 17.6 Infrared only 82.9 82.8 50.8 72.2 33.4 GAFF 76.6 85.5 59.4 72.9 32.9 CFT 84.1 89.5 61.4 78.7 35.5 TarDAL 85.1 85.3 69.3 79.9 37.9 CSAA 83.2 86.7 68.6 79.4 37.2 ICAFusion 81.6 89.0 66.9 79.2 36.9 TFDet 85.3 90.8 68.4 81.5 41.9 FD2Net 85.3 89.9 73.2 82.9 42.5 JTMDet 87.8 91.5 71.5 83.6 42.8 The invention 87.2 91.2 74.4 84.3 45.0
[0098] Note that the design of the above digital prototype model, the parameter setting of each material, and the selection of the measurement plane are only the preferred embodiments of the present application and the technical principles applied. Those skilled in the art will understand that the present application is not limited to the specific embodiments described herein, and various obvious changes, readjustments and substitutions can be made by those skilled in the art without departing from the scope of the present application. Therefore, although the present application has been described in more detail through the above embodiments, the present application is not limited to the above embodiments, and can include more other equivalent embodiments without departing from the concept of the present application, and the scope of the present application is determined by the scope of the appended claims.
Claims
1. An optical-infrared target detection method of cross-modal attention fusion mechanism, characterized in that, The method comprises the following steps: S1, inputting images of two modalities of visible light and infrared into respective feature extraction backbone networks to obtain discrete multi-scale feature maps corresponding to each modality; S2, inputting the discrete multi-scale feature maps into an implicit feature interaction unit, and converting the discrete feature maps into a feature mapping function in a continuous space domain through an implicit neural representation module in the unit; the feature mapping function can output a feature vector corresponding to a given two-dimensional space coordinate, realizing continuous representation of the feature; S3, in the continuous space domain, using an implicit attention mechanism to adaptively aggregate cross-modal features; taking the features of the first modality as a query, querying and weightedly aggregating on the continuous feature mapping function of the second modality to generate fused features precisely aligned with the first modality, thereby completing sub-pixel level feature interaction and alignment; S4, using an implicit cross-scale fusion unit to realize alignment of multi-scale fused features, and then using an efficient state space model in the unit to model long-distance dependencies and interact global context information of features at different levels, realize scale restoration of features at different scales, and obtain multi-scale fused features after cross-scale feature interaction; S5, inputting the final multi-scale fused features into a detection neck and head to decode and output position and category information of the target in the image.
2. The optical-infrared target detection method of claim 1, wherein: Step S1, a double-branch feature extraction structure is adopted, and a pair of spatially registered visible light images and infrared images are input into two parallel feature extraction backbone networks based on convolutional neural networks; And from multiple different depth stages of the two backbone networks, features with decreasing spatial resolution and increasing semantic information are extracted, respectively, to generate a set of discrete multi-scale feature maps for the visible light modality and the infrared modality.
3. The optical-infrared target detection method of claim 1, wherein, Step S2, specifically comprising the following steps: Any discrete feature map obtained from step S1 is regarded as a sampling result of a continuous function, and the continuous function is defined by the following mapping relationship: , wherein, denotes an arbitrary query coordinate within the normalized two-dimensional continuous coordinate domain Ω, denotes the C-dimensional feature vector corresponding to this coordinate point, and represents the C-dimensional feature space; in order to evaluate the value of this continuous function at an arbitrary query coordinate a differentiable sampling operator is employed, which is calculated by the following formula: , wherein, is the original discrete feature map, is a differentiable sampling operator, which is implemented by locating to the four nearest-neighbor grid points of the query coordinate in the discrete feature map and performing a bilinear interpolation operation on the feature vectors of the four points; thereby, the inherent discrete feature grid is converted into a continuous and differentiable feature field.
4. The optical-infrared target detection method of claim 1, wherein, Step S3, specifically: define the query in the traditional attention mechanism as a coordinate-based position encoding, replace the traditional content-dependent query with a coordinate-based query mechanism, so that the attention process is directly affected by the condition of the absolute spatial position; given the input feature map , the total number of query points is , by filling each spatial position with its corresponding normalized coordinate vector , the query tensor is constructed; for any query point , the corresponding query tensor is the two-dimensional coordinates of the point, as shown in the formula: , For each query position The key is sampled from the continuous field of features of the guiding feature at the same coordinates: , wherein, is a discrete feature map of the guidance modality; Querying locations The value comes from the continuous field of features of the source modality, used to augment: , wherein, is a discrete feature map of the source modality; the final preserves the original modality information, while the attention weights are used to adaptively adjust ; The above sampling results in and Vectors are linearly projected and reshaped to groups, each group has a dimension of ; Query coordinates are also extended to adapt to multi-head computation; Matrix multiplication is used to replace the traditional dot product to calculate the interaction between content and position, position query is multiplied by content to produce a 2D directional attention : , The scalar score is then used to condition ; The relevance score is decoded as a modulation vector : , where the vector A set of weights specific to the feature is provided; ultimately, the output is obtained by computing the Hadamard product of these weights with the vector of values, resulting in the enhanced feature vector 。 5. The optical-infrared target detection method of claim 1, wherein, Step S4, specifically comprising the following steps: For any coordinate eigenvectors First, an enhanced representation is generated by concatenating it with normalized coordinates. This vector is then passed through a multilayer perceptron. Process and output the transformation features at the target scale. : , By conditioning the transformation on spatial coordinates, Learn a spatial adaptive function, which can intelligently sample features; To effectively model the global dependency of multi-scale features aligned across space, a sequential processor consisting of multiple stacked Mamba blocks is introduced; the process treats the concatenated multi-scale feature sequence as a whole, then applies position encoding ; subsequently, the sequence is input into the Mamba block stack as follows: , , wherein, represents each scale feature information segmented from the output sequence with a global aggregation context, represents feature tensors from different scales that have been aligned in continuous space by a coordinate-aware feature scale transformation module; The Mamba processor can scan the entire sequence of features to capture long-range dependencies; subsequently reapply the coordinate-aware feature scale transform module to restore each feature map to the original channel dimension; finally, these globally augmented features are concatenated with the initial features through a residual connection integration: 。
Citation Information
Patent Citations
Modal sharing information layered unwrapping fusion network for RGB-T target tracking
CN120339780A
RGB-t multispectral pedestrian detection method based on target aware fusion strategy
US20240331403A1