Infrared multi-source information fusion method and system based on hierarchical cross-modal attention
Patent Information
- Application Number
- CN202611194158.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-07
- Publication Date
- 2026-09-29
AI Technical Summary
然而,现有融合方法在应对红外与文本这类模态差异显著的信息源时,仍存在明显不足
[0026]实现细粒度语义对齐与全局语义理解相统一。通过局部级注意力建立红外图像各空间位置与文本描述间的细粒度关联,再经全局级注意力整合场景级语义,形成层次化、互补性的融合特征,显著提升了跨模态语义融合的完整性与准确性。
Smart Images

Figure CN122841906A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of multimodal information fusion and artificial intelligence technology, specifically relating to an infrared multi-source information fusion method and system based on hierarchical cross-modal attention. Background Technology
[0002] In intelligent sensing and decision-making systems, infrared imaging has become a key information source due to its all-weather capability. To obtain a deep semantic understanding of a scene, it is often necessary to fuse multi-source heterogeneous information such as infrared images, natural language descriptions (such as reconnaissance reports and operational instructions), and geographic information system data. However, existing fusion methods still have significant shortcomings when dealing with information sources with significant modal differences, such as infrared and text.
[0003] Traditional multimodal fusion methods often employ simple operations such as feature concatenation and weighted summation. While easy to implement, these methods cannot model complex nonlinear semantic relationships between modalities and treat all feature regions equally, making them prone to introducing noise in complex contexts and leading to a decrease in the discriminative power of the fused features. Subsequent attention-based methods can calculate global weights, but they are mostly single-level attention structures, lacking the ability to align local fine-grained semantics. They struggle to accurately map specific hotspots in infrared images to specific descriptions in text, failing to achieve pixel / region-level semantic grounding.
[0004] In existing technologies, visible light and infrared image fusion (such as methods based on generative adversarial networks or part alignment) mainly focus on visual feature compensation and alignment between modalities. However, their goals are mostly to generate enhanced images or achieve identity recognition, without addressing the "semantic gap" between infrared and text. Furthermore, some text-guided image fusion methods, while using text cues to modulate visual features, treat the text only as a static guiding signal and do not engage in bidirectional, hierarchical interactive attention fusion with the image. Moreover, the final output is still an image, rather than a unified cross-modal feature representation for high-level semantic tasks (such as scene understanding and decision reasoning).
[0005] Meanwhile, existing methods lack a flexible and scalable architecture when fusing multi-source information (such as infrared, text, and GIS), cannot dynamically assess the importance of information from different modalities and regions, and are difficult to adaptively fuse structured prior knowledge, resulting in poor semantic consistency and weak generalization ability in complex application scenarios.
[0006] Therefore, there is an urgent need for a method that can adaptively and multi-layeredly understand and fuse heterogeneous information such as infrared images and text, in order to extract more discriminative, robust and semantically consistent unified feature representations to support accurate high-level semantic understanding and intelligent decision-making. Summary of the Invention
[0007] To address the aforementioned technical problems, this invention provides an infrared multi-source information fusion method and system based on hierarchical cross-modal attention. By designing a two-level attention fusion mechanism from local to global, it achieves the unification of fine-grained semantic alignment and global scene understanding, with a clear structure and the ability to effectively utilize complementary information.
[0008] To achieve the above objectives, the present invention adopts the following technical solution:
[0009] Infrared multi-source information fusion methods based on hierarchical cross-modal attention include:
[0010] Receives an infrared image feature sequence from an infrared image encoder and a text feature sequence from a text encoder;
[0011] The feature of each spatial location in the infrared image feature sequence is used as a query vector and interacted with the text feature sequence through attention to obtain local semantic enhancement features.
[0012] The local semantic enhancement features are spatially aggregated to obtain a global visual feature vector; the global visual feature vector is then used as a query vector and interacted with the text feature sequence again to generate the final unified multimodal fusion feature vector.
[0013] Furthermore, the infrared image encoder is a visual transformer or a residual network, and the text encoder is a text encoder based on a transformer-based bidirectional encoder representation technology model or a contrastive language-image pre-trained model.
[0014] Furthermore, obtaining local semantic enhancement features includes: reshaping the infrared image feature sequence into a query matrix, using the text feature sequence as the key matrix and value matrix, calculating the association weight between the query matrix and the key matrix through scaling dot product attention, and weighted summing the association weight with the value matrix to obtain the local features of each spatial location after text semantic enhancement.
[0015] Furthermore, obtaining the global visual feature vector includes: performing global average pooling on the local semantic enhancement features in the spatial dimension, and using the pooling result as a global visual feature vector representing the overall scene.
[0016] Furthermore, generating the final unified multimodal fusion feature vector includes: using the global visual feature vector as the query matrix, using the text feature sequence as the key matrix and value matrix, calculating the association weight between the query matrix and the key matrix through scaling dot product attention, and weighting and summing the association weight with the value matrix to obtain the final feature vector that fuses global visual semantics and text semantics.
[0017] Furthermore, when geographic information system (GIS) data is available, the GIS data is encoded into a GIS feature sequence. In the attention interaction that generates the unified multimodal fusion feature vector, the text feature sequence is concatenated with the GIS feature sequence, and the concatenated sequence is used as the key matrix and value matrix to participate in attention calculation.
[0018] Furthermore, the unified multimodal fusion feature vector can be directly used for downstream classification tasks, retrieval tasks, or scene understanding tasks.
[0019] On the other hand, the present invention provides an infrared multi-source information fusion system based on hierarchical cross-modal attention, comprising:
[0020] The receiving module is used to receive infrared image feature sequences from an infrared image encoder and text feature sequences from a text encoder.
[0021] The first attention interaction module is used to take the features of each spatial location in the infrared image feature sequence as a query vector and perform attention interaction with the text feature sequence to obtain local semantic enhancement features.
[0022] The second attention interaction module is used to aggregate the spatial dimensions of the local semantic enhancement features to obtain a global visual feature vector; using the global visual feature vector as a query vector, it performs attention interaction with the text feature sequence again to generate the final unified multimodal fusion feature vector.
[0023] Thirdly, the present invention provides an electronic device, comprising: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the aforementioned method for fusion of infrared multi-source information based on hierarchical cross-modal attention.
[0024] Fourthly, the present invention provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, enable the processor to implement the aforementioned method for infrared multi-source information fusion based on hierarchical cross-modal attention.
[0025] The beneficial effects of this invention are as follows:
[0026] Achieving a unified approach between fine-grained semantic alignment and global semantic understanding. By establishing fine-grained associations between spatial locations in infrared images and textual descriptions through local-level attention, and then integrating scene-level semantics through global-level attention, hierarchical and complementary fusion features are formed, significantly improving the completeness and accuracy of cross-modal semantic fusion.
[0027] It supports adaptive deep fusion of infrared, text, and structured data. This invention is not only applicable to dual-modal fusion of infrared and text, but can also be flexibly extended to structured prior information such as GIS. It dynamically evaluates the importance of information from different modalities and regions through an attention mechanism, achieving intelligent weighting and noise suppression of multi-source information.
[0028] It provides robust feature representations for high-level semantic tasks. The generated unified multimodal fusion feature vectors have strong discriminativeness and semantic consistency, and can be directly used for downstream tasks such as classification, retrieval, and scene understanding, providing high-quality feature input for intelligent perception and decision-making systems.
[0029] The structure is clear and highly scalable. The hierarchical attention module design is clear and easy to integrate with other encoders or fusion modules. It is suitable for different modal combinations and complex application scenarios, and has good engineering applicability and technical extensibility. Attached Figure Description
[0030] Figure 1 This is a flowchart of the infrared multi-source information fusion method based on hierarchical cross-modal attention of the present invention;
[0031] Figure 2 This is a schematic diagram of the structure of the hierarchical cross-modal attention fusion module in an embodiment of the present invention;
[0032] Figure 3 This is a schematic diagram of the standard scaled dot product attention mechanism. Detailed Implementation
[0033] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0034] like Figure 1 As shown, this invention provides an infrared multi-source information fusion method based on hierarchical cross-modal attention, the method mainly comprising the following steps:
[0035] Step 1: Receive the infrared image feature sequence from the infrared image encoder and the text feature sequence from the text encoder;
[0036] Both the infrared image encoder and the text encoder are pre-trained feature extraction models. The infrared image encoder (such as Vision Transformer, ResNet, etc.) is responsible for extracting deep visual features with spatial structure and semantic information from the raw infrared image; the text encoder (such as BERT, CLIP text encoder, etc.) is responsible for converting natural language descriptions into text feature sequences with contextual semantics. This invention does not limit the specific network structure, number of layers, or training data of the above encoders; their output features only need to meet the agreed dimensional format to be used as input to this fusion method. This design allows this invention to flexibly access mature visual and language models in the industry, focusing on the hierarchical fusion of cross-modal features, and has good cutting-edge compatibility and engineering practicality.
[0037] Here, we assume that the features output by the infrared image encoder (such as VisionTransformer) are... H, W, and C represent the image dimensions (length, width, and number of channels), respectively (preferably...). The features output by a text encoder (such as BERT) are: L represents the text length (preferably...). ).
[0038] like Figure 2 As shown, the hierarchical cross-modal attention fusion module of this invention mainly consists of two connected parts: a local-level fusion unit and a global-level fusion unit. The local-level fusion unit takes the infrared image feature sequence extracted by the front-end encoder as input, uses its spatial location features as a query, and performs cross-modal attention interaction with the text feature sequence, outputting a local feature map enhanced with fine-grained text semantics. The global-level fusion unit receives this local feature map, first compresses it into a global visual vector through a feature aggregator (such as global average pooling), then uses this vector as a query to perform a secondary attention interaction with the text feature sequence, finally outputting a unified multimodal fusion feature vector.
[0039] like Figure 3 As shown, the aforementioned local and global cross-modal attention interactions are all based on Figure 3The standard scaled dot product attention mechanism is implemented as shown in the figure. Its input consists of three vectors: query (Q), key (K), and value (V), obtained from feature sequences mapped from different modalities. During computation, the dot product of Q and K is first calculated, and then scaled by dividing by the square root of the dimension to obtain the original attention score, stabilizing the gradient. Subsequently, the score is converted into probability distribution attention weights using a softmax function. Finally, the weights are multiplied by V and summed to output the weighted feature representation. The introduction of this mechanism enables the model to adaptively select and fuse the most relevant information from key-value pairs based on the query content, providing the mathematical foundation for fine-grained semantic alignment and dynamic information fusion. Therefore, this invention also includes:
[0040] Step 2, Local-level Cross-modal Attention Fusion: The features at each spatial location in the infrared image feature sequence are used as query vectors and interacted with the text feature sequence through attention to obtain local semantic enhancement features; the features output by the infrared image encoder are then... Remodeling into a query matrix Z is an intermediate parameter. ;make K is the key, and V is the value. Local-level fusion is performed using a single-layer Transformer decoder. The core of this process is calculating the association weight between each spatial location in the image feature sequence and all text tokens.
[0041] ,
[0042] This represents the attention mechanism, where softmax represents the normalized exponential function, and the superscript T indicates transpose. The output of this operation is... It can be reconstructed back to 14×14×768. Its physical meaning is that each spatial unit in the image actively searches for the most relevant semantic fragment in the text description (such as "high temperature" or "smoke exhaust"), generating local semantic enhancement features. .
[0043] Step 3, Global-level Cross-modal Attention Fusion: The local semantic enhancement features are spatially aggregated to obtain a global visual feature vector; using this global visual feature vector as the query vector, attention interaction is performed again with the text feature sequence to generate the final unified multimodal fusion feature vector. Specifically, this includes:
[0044] First, local semantic enhancement features Perform global average pooling to obtain the global visual vector. :
[0045] ,
[0046] in, express The Okay. At this time. .
[0047] Then, with For query matrix Features output by the text encoder For key value Perform global attention calculation:
[0048] ,
[0049] The physical meaning of this step is: from the integrated global visual perspective, re-evaluate the importance of each part of the entire text description, perform weighted synthesis, and finally obtain a unified multimodal fusion feature vector representing the entire scene. .
[0050] Without loss of generality, if GIS data exists, it should be encoded as a GIS feature. (Right now In global attention computation, the features output by the text encoder are... With GIS features splicing together a new context:
[0051] ,
[0052] In the formula, Concat represents the concatenation operation. The key of the concatenated feature is the value. ;
[0053] The fusion formula then becomes:
[0054] ,
[0055] This allows the fusion process to take into account both semantic description and spatial prior knowledge.
[0056] Final output or It is a compact and robust feature vector that integrates local detail correlations and global semantics, and can be directly used for downstream tasks such as classification and retrieval.
[0057] In summary, this invention achieves intelligent and adaptive fusion of multi-source information, such as infrared, through a hierarchical attention design that ranges from fine-grained to global, effectively improving perception capabilities in complex scenarios.
[0058] On the other hand, the present invention provides an infrared multi-source information fusion system based on hierarchical cross-modal attention, which includes various modules for implementing the various steps of the aforementioned method, specifically including:
[0059] The receiving module is used to receive infrared image feature sequences from an infrared image encoder and text feature sequences from a text encoder.
[0060] The first attention interaction module is used to take the features of each spatial location in the infrared image feature sequence as a query vector and perform attention interaction with the text feature sequence to obtain local semantic enhancement features.
[0061] The second attention interaction module is used to aggregate the spatial dimensions of the local semantic enhancement features to obtain a global visual feature vector; using the global visual feature vector as a query vector, it performs attention interaction with the text feature sequence again to generate the final unified multimodal fusion feature vector.
[0062] Thirdly, the present invention provides an electronic device, comprising: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the aforementioned method for fusion of infrared multi-source information based on hierarchical cross-modal attention.
[0063] Fourthly, the present invention provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, enable the processor to implement the aforementioned method for infrared multi-source information fusion based on hierarchical cross-modal attention.
[0064] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. An infrared multi-source information fusion method based on hierarchical cross-modal attention, characterized in that, include: Receives an infrared image feature sequence from an infrared image encoder and a text feature sequence from a text encoder; The feature of each spatial location in the infrared image feature sequence is used as a query vector and interacted with the text feature sequence through attention to obtain local semantic enhancement features. The local semantic enhancement features are spatially aggregated to obtain a global visual feature vector; the global visual feature vector is then used as a query vector and interacted with the text feature sequence again to generate the final unified multimodal fusion feature vector.
2. The infrared multi-source information fusion method based on hierarchical cross-modal attention according to claim 1, characterized in that, The infrared image encoder is a visual transformer or a residual network, and the text encoder is a text encoder based on a transformer-based bidirectional encoder representation technology model or a contrastive language-image pre-trained model.
3. The infrared multi-source information fusion method based on hierarchical cross-modal attention according to claim 1, characterized in that, The process of obtaining local semantic enhancement features includes: reshaping the infrared image feature sequence into a query matrix, using the text feature sequence as the key matrix and value matrix, calculating the association weight between the query matrix and the key matrix through scaling dot product attention, and weighted summing the association weight with the value matrix to obtain the local features of each spatial location after text semantic enhancement.
4. The infrared multi-source information fusion method based on hierarchical cross-modal attention according to claim 1, characterized in that, The process of obtaining the global visual feature vector includes: performing global average pooling on the local semantic enhancement features in the spatial dimension, and using the pooling result as the global visual feature vector representing the overall scene.
5. The infrared multi-source information fusion method based on hierarchical cross-modal attention according to claim 1, characterized in that, The process of generating the final unified multimodal fusion feature vector includes: using the global visual feature vector as the query matrix, using the text feature sequence as the key matrix and value matrix, calculating the association weight between the query matrix and the key matrix through scaling dot product attention, and weighted summing the association weight with the value matrix to obtain the final feature vector that fuses global visual semantics and text semantics.
6. The infrared multi-source information fusion method based on hierarchical cross-modal attention according to claim 1, characterized in that, When geographic information system (GIS) data is available, the GIS data is encoded into a GIS feature sequence. In the attention interaction that generates the unified multimodal fusion feature vector, the text feature sequence is concatenated with the GIS feature sequence, and the concatenated sequence is used as the key matrix and value matrix to participate in attention calculation.
7. The infrared multi-source information fusion method based on hierarchical cross-modal attention according to claim 1, characterized in that, The unified multimodal fusion feature vector is directly used in downstream classification tasks, retrieval tasks, or scene understanding tasks.
8. An infrared multi-source information fusion system based on hierarchical cross-modal attention, characterized in that, include: The receiving module is used to receive infrared image feature sequences from an infrared image encoder and text feature sequences from a text encoder. The first attention interaction module is used to take the features of each spatial location in the infrared image feature sequence as a query vector and perform attention interaction with the text feature sequence to obtain local semantic enhancement features. The second attention interaction module is used to aggregate the spatial dimensions of the local semantic enhancement features to obtain a global visual feature vector; using the global visual feature vector as a query vector, it performs attention interaction with the text feature sequence again to generate the final unified multimodal fusion feature vector.
9. An electronic device, characterized in that, include: One or more processors; Memory, used to store one or more programs; When one or more programs are executed by the one or more processors, the one or more processors implement the infrared multi-source information fusion method based on hierarchical cross-modal attention as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, It stores executable instructions that, when executed by a processor, enable the processor to implement the infrared multi-source information fusion method based on hierarchical cross-modal attention as described in any one of claims 1-7.