A multi-modal semantic generation method and system for infrared and SAR images

By constructing a structured semantic automatic generation framework for infrared and SAR images and fusing a visual encoder with a large language model, the problem of insufficient semantic description in multimodal generation of infrared and SAR images is solved, achieving fine-grained scene understanding and logical enhancement, and improving the model's understanding accuracy in complex scenes.

CN122392062APending Publication Date: 2026-07-14UNIV OF SCI & TECH BEIJING
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
UNIV OF SCI & TECH BEIJING
Filing Date
2026-04-21
Publication Date
2026-07-14

Smart Images

  • Figure CN122392062A_ABST
    Figure CN122392062A_ABST
Patent Text Reader

Abstract

The application provides a kind of multi-modal semantic generation method and system for infrared and SAR image, belong to computer vision field.Firstly, the application manually labels infrared and SAR image, obtains image bottom layer label and a small amount of semantic description information;Then, based on image bottom layer label and a small amount of semantic description information, a structured semantic automatic generation framework under the constraint of perception is constructed, by mapping low-level visual perception information to high-level fine-grained semantic space, the automatic hierarchical change from pixel feature to structured label is realized;Secondly, in the model architecture design, the fusion scheme of visual encoder and large language model is adopted, and the modal characteristics of infrared and SAR are coupled to a unified semantic representation space through a lightweight visual-language alignment mapping mechanism;Finally, under the guidance of modal prompt vector and semantic demand instruction, language decoding is carried out through large language model, and the comprehensive expression of spatial topology logic and overall scene attribute is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision, and in particular to a method and system for generating multimodal semantics for infrared and SAR images. Background Technology

[0002] For decades, image understanding and semantic description, as crucial supports for computer vision, have been research hotspots in applications such as remote sensing monitoring, target reconnaissance, and disaster assessment. From land cover classification in urban planning to target identification in military operations, infrared and synthetic aperture radar (SAR) images, with their all-weather, all-time detection advantages, have demonstrated immense practical value in scene cognition under complex environments. With the rapid development of Large Multimodal Models (LMMs), significant progress has been made in using visual-language alignment techniques to assist in image semantic generation. This advancement has introduced a new perspective to cross-modal semantic mapping. By encoding visual features into the semantic space of a pre-trained language model, the model can leverage the superior reasoning capabilities of large models to generate rich scene descriptions. Particularly in natural optical image processing, visual-language pre-training based on natural optical images has enabled the semantic generation of scene attribute information using large-scale image-text datasets.

[0003] However, existing multimodal generation methods still face serious challenges when applied to non-natural optical modalities such as infrared and SAR. Due to the unique physical imaging mechanisms of non-natural optical modalities such as infrared and SAR and their relatively insufficient semantic expression capabilities, existing methods for processing large multimodal models lack deep adaptation to the physical characteristics of specific modalities and are difficult to accurately characterize the complex spatial topological relationships between multiple targets in a scene. This results in the generated semantic information remaining at a shallow level of category stacking, lacking fine-grained structured semantic descriptions, which limits the performance breakthrough of large multimodal models in special remote sensing scene understanding tasks. Summary of the Invention

[0004] To address the problems in the existing technologies, this invention provides a method and system for multimodal semantic generation of infrared and SAR images. First, the acquired infrared and SAR images are manually annotated to obtain low-level image annotations and a small amount of semantic description information. Then, based on the low-level image annotations and the limited semantic description information, a structured semantic automatic generation framework under perceptual constraints is constructed. By mapping low-level visual perceptual information to a high-level fine-grained semantic space, an automatic hierarchical transformation from pixel features to structured annotations is achieved. Second, in terms of model architecture design, this invention adopts a fusion scheme of a visual encoder and a large language model. Through a lightweight visual-language alignment mapping mechanism, the modal features of infrared and SAR are coupled to a unified semantic representation space. Finally, through the large language model, guided by modal cue vectors and semantic requirement instructions, language decoding is performed, realizing the model's comprehensive expression of spatial topological logic and overall scene attributes. To achieve the above objectives, the technical solution is as follows: On the one hand, the present invention provides a multimodal semantic generation method for infrared and SAR images, the method comprising: S1. Use remote sensing sensors to acquire infrared and SAR images to obtain raw datasets of infrared and SAR images; S2. Based on the original infrared image dataset and the original SAR image dataset, a perception annotation library is obtained through manual annotation; S3. Based on the original infrared image dataset and the original SAR image dataset, train the visual encoder adaptation network through the perceptual annotation library to obtain a visual encoder with an adaptation network. S4. Based on the original infrared image dataset and the original SAR image dataset, the infrared image feature dataset and the SAR image feature dataset are obtained through inference by the visual encoder with the adaptive network; S5. Based on the perception annotation library, the infrared image feature dataset, the SAR image feature dataset, and the semantic requirement instructions, a structured semantic dataset is obtained through manual annotation and automatic generation of structured semantics; S6. Based on the infrared image feature dataset and the SAR image feature dataset, a dual-modal aggregated visual feature dataset is obtained through fusion and local aggregation processing; S7. Based on the bimodal aggregated visual feature dataset, a multimodal mapping network is trained under the supervision of the structured semantic dataset to obtain a cross-domain transformation network model; S8. Based on the bimodal aggregated visual feature dataset, the visual semantic vector dataset is obtained by transformation and alignment through the cross-domain transformation network model; S9. Based on the visual semantic vector dataset, the semantic requirement instruction and modal cue vector, the parameter efficiency module of the large language model is trained under the supervision of the structured semantic dataset to obtain a large language model with efficient parameter tuning. S10. Based on the visual semantic vector dataset, the semantic requirement instruction and modal cue vector, and through the reasoning of the efficient parameter-tuned large language model, fine-grained semantic text is obtained.

[0005] Optionally, in step S4, based on the original infrared image dataset and the original SAR image dataset, the infrared image feature dataset and the SAR image feature dataset are obtained through inference by the visual encoder with the adaptive network, including: S41. Based on the original infrared image dataset and the original SAR image dataset, the first-stage infrared image feature dataset and the first-stage SAR image feature dataset are obtained through the basic visual feature extraction of the visual encoder with the adaptive network. S42. Based on the first-stage infrared image feature dataset and the first-stage SAR image feature dataset, the second-stage infrared image feature dataset and the second-stage SAR image feature dataset are obtained through filtering and key weighted optimization. S43. Based on the second-stage infrared image feature dataset and the second-stage SAR image feature dataset, the third-stage infrared image feature dataset and the third-stage SAR image feature dataset are obtained through a mixture of local window attention and a small amount of global attention. S44. Based on the third-stage infrared image feature dataset and the third-stage SAR image feature dataset, spatial features are enhanced by two-dimensional rotational position coding to obtain the infrared image feature dataset and the SAR image feature dataset.

[0006] Optionally, in step S41, based on the original infrared image dataset and the original SAR image dataset, the first-stage infrared image feature dataset and the first-stage SAR image feature dataset are obtained through basic visual feature extraction by the visual encoder with the adaptive network, including: S411. Based on the original infrared image dataset and the original SAR image dataset, obtain the infrared image patch dataset and the SAR image patch dataset through image segmentation; S412. Based on the infrared image patch dataset and the SAR image patch dataset, the first-stage infrared image feature dataset and the first-stage SAR image feature dataset are obtained through the embedding layer mapping of the visual encoder with the adaptive network.

[0007] Optionally, in S5, based on the perception annotation library, the infrared image feature dataset, the SAR image feature dataset, and the semantic requirement instructions, a structured semantic dataset is obtained through manual annotation and automatic generation of structured semantics, including: S51. Based on this perceptual annotation library, an artificially structured semantic description text library is obtained through artificial natural language organization; S52. The perception annotation library, the infrared image feature dataset, the SAR image feature dataset, and the semantic requirement instructions are used to train a multimodal large model under the supervision of the artificially structured semantic description text library to obtain a multimodal structured semantic automatic generation model. S53. Based on the perception annotation library, the infrared image feature dataset, the SAR image feature dataset, and the semantic requirement instruction, the structured semantic dataset is obtained through reasoning using the multimodal structured semantic automatic generation model.

[0008] Optionally, in step S53, based on the perception annotation library, the infrared image feature dataset, the SAR image feature dataset, and the semantic requirement instruction, a structured semantic dataset is obtained through reasoning using the multimodal structured semantic automatic generation model, including: S531. Based on this perception annotation library, through statistical processing, the target layer semantics are obtained, including: target type set, target bounding box set, target location set, and target quantity set; S532. Based on the semantics of the target layer, the semantic vector of the relation layer is obtained by calculating the geometric relationship based on the set of target boxes, including: regional distribution features, clustering features and arrangement morphology features; S533. Based on the infrared image feature dataset, the SAR image feature dataset, the relational layer semantic vector, and the semantic requirement instruction, the structured semantic dataset is obtained through reasoning by the multimodal structured semantic automatic generation model.

[0009] Optionally, in step S6, based on the infrared image feature dataset and the SAR image feature dataset, a dual-modal aggregated visual feature dataset is obtained through fusion and local aggregation processing, including: S61. Based on the infrared image feature dataset and the SAR image feature dataset, a dual-modal fused visual feature dataset is obtained through transformation and fusion; S62. Based on the bimodal fusion visual feature dataset, a bimodal aggregated visual feature dataset is obtained through local aggregation processing.

[0010] Optionally, in S7, based on the bimodal aggregated visual feature dataset, a multimodal mapping network is trained under the supervision of the structured semantic dataset to obtain a cross-domain conversion network model, including: S71. Based on this structured semantic dataset, word-level semantic vectors are obtained through the word embedding mechanism of the language model; S72. Based on the bimodal aggregated visual feature dataset, a multimodal mapping network is trained under the supervision of word-level semantic vectors to obtain a cross-domain conversion network model.

[0011] On the other hand, the present invention provides a multimodal semantic generation system for infrared and SAR images. This system is applied to a multimodal semantic generation method for infrared and SAR images, and includes: The dual-modal image acquisition module is used to acquire infrared and SAR images using remote sensing sensors to obtain raw infrared and SAR image datasets. The perception annotation module is used to obtain a perception annotation library by manually annotating the original infrared image dataset and the original SAR image dataset. The visual encoder adaptation module is used to train the visual encoder adaptation network based on the original infrared image dataset and the original SAR image dataset through the perceptual annotation library to obtain a visual encoder with the adaptation network. The dual-modal visual feature acquisition module is used to obtain the infrared image feature dataset and the SAR image feature dataset by reasoning through the visual encoder with the adaptive network based on the original infrared image dataset and the original SAR image dataset. The automatic generation module for structured semantics is used to obtain a structured semantic dataset by combining manual annotation and automatic generation of structured semantics based on the perception annotation library, the infrared image feature dataset, the SAR image feature dataset and semantic requirement instructions. The dual-modal aggregated visual feature acquisition module is used to obtain a dual-modal aggregated visual feature dataset by fusing and local aggregation processing based on the infrared image feature dataset and the SAR image feature dataset. The cross-domain transformation network training module is used to train a multimodal mapping network under the supervision of the structured semantic dataset based on the bimodal aggregated visual feature dataset to obtain the cross-domain transformation network model. The visual semantic transformation module is used to transform and align the bimodal aggregated visual feature dataset through the cross-domain transformation network model to obtain a visual semantic vector dataset. The efficient parameter tuning module for large language models is used to train the parameters of a large language model under supervision using the visual semantic vector dataset, the semantic requirement instruction and modal cue vector, through the structured semantic dataset, and obtain an efficient parameter tuning module for large language models. The fine-grained semantic text generation module is used to obtain fine-grained semantic text based on the visual semantic vector dataset, the semantic requirement instruction and the modal cue vector through reasoning of the efficient parameter-tuned large language model.

[0012] Compared with the prior art, the technical solution of the present invention has at least the following beneficial effects: The above-mentioned solution addresses several key aspects. First, it automates the construction of a framework to transform low-level target information into high-level semantics, effectively solving the problems of scarce semantic data and high manual annotation costs in infrared and SAR images. Second, it optimizes the visual encoding mechanism for non-visible light imaging characteristics, enhancing the ability to depict structural contours and geometric shapes, and significantly improving the model's understanding accuracy in complex scenes. Third, it employs lightweight feature aggregation and efficient parameter fine-tuning strategies, significantly reducing model computational overhead and training costs while retaining the generalization ability of large models. Fourth, it achieves multi-level, structured semantic generation from "object-level" to "scene-level," resulting in more granular output content with stronger logic and interpretability. Attached Figure Description

[0013] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0014] Figure 1 This is a flowchart of an embodiment of the multimodal semantic generation method for infrared and SAR images of the present invention; Figure 2 This is a flowchart illustrating the dual-modal visual feature acquisition process of an embodiment of the multimodal semantic generation method for infrared and SAR images of the present invention. Figure 3 This is a flowchart of the bimodal basic visual feature extraction process of an embodiment of the multimodal semantic generation method for infrared and SAR images of the present invention; Figure 4 This is a flowchart illustrating the automatic generation of structured semantics, based on an embodiment of the multimodal semantic generation method for infrared and SAR images of the present invention. Figure 5 This is a flowchart illustrating the generation of high-level structured semantics from target-layer semantics, based on an embodiment of the multimodal semantics generation method for infrared and SAR images of the present invention. Figure 6 This is a flowchart illustrating the acquisition of bimodal aggregated visual features in an embodiment of the multimodal semantic generation method for infrared and SAR images of the present invention. Figure 7This is a flowchart of the cross-domain conversion network training process in an embodiment of the multimodal semantic generation method for infrared and SAR images of the present invention; Figure 8 This is a system block diagram of an embodiment of the multimodal semantic generation system for infrared and SAR images of the present invention. Detailed Implementation

[0015] The technical solution of the present invention will now be described with reference to the accompanying drawings.

[0016] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.

[0017] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0018] like Figure 1 The flowchart shown is an embodiment of the multimodal semantic generation method for infrared and SAR images of the present invention. The present invention provides a multimodal semantic generation method for infrared and SAR images, which is implemented by a multimodal semantic generation system for infrared and SAR images. The method includes: S1. Use remote sensing sensors to acquire infrared and SAR images to obtain raw datasets of infrared and SAR images.

[0019] S2. Based on the original infrared image dataset and the original SAR image dataset, a perception annotation library is obtained through manual annotation.

[0020] S3. Based on the original infrared image dataset and the original SAR image dataset, train the visual encoder adaptation network using the perceptual annotation library to obtain a visual encoder with the adaptation network.

[0021] For example, in this visual modality adaptation stage, a lightweight modality adaptation layer, such as an Adapter or LoRA structure, is introduced only on the input side of the visual encoder to learn the spectral structure and target morphology representation specific to infrared and SAR images. The goal of this stage is to gradually adapt the feature distribution output by the visual encoder to the statistical characteristics of infrared and SAR images, providing a stable input representation for subsequent cross-modal alignment. Supervision comes from existing perceptual annotation data, such as manually annotated bounding boxes and category information. These low-level visual labels provide a reliable structural reference for the model, enabling the newly added modality adaptation layer to learn the contour morphology, scattering structure, and texture statistical patterns specific to infrared and SAR images. The supervision in this stage is characterized by weak semantics and strong structural constraints; the core objective is to achieve modal alignment of visual feature distributions, rather than directly learning semantic representation capabilities.

[0022] S4. Based on the original infrared image dataset and the original SAR image dataset, the infrared image feature dataset and the SAR image feature dataset are obtained through inference by the visual encoder with the adaptive network; Specifically, such as Figure 2 The flowchart shown is a bimodal visual feature acquisition process of an embodiment of the multimodal semantic generation method for infrared and SAR images of the present invention. In step S4, based on the original infrared image dataset and the original SAR image dataset, the infrared image feature dataset and the SAR image feature dataset are obtained through inference by the visual encoder with an adaptive network, including: S41. Based on the original infrared image dataset and the original SAR image dataset, the first-stage infrared image feature dataset and the first-stage SAR image feature dataset are obtained through the basic visual feature extraction of the visual encoder with the adaptive network. S42. Based on the first-stage infrared image feature dataset and the first-stage SAR image feature dataset, the second-stage infrared image feature dataset and the second-stage SAR image feature dataset are obtained through filtering and key weighted optimization. S43. Based on the second-stage infrared image feature dataset and the second-stage SAR image feature dataset, the third-stage infrared image feature dataset and the third-stage SAR image feature dataset are obtained through a mixture of local window attention and a small amount of global attention. S44. Based on the third-stage infrared image feature dataset and the third-stage SAR image feature dataset, spatial features are enhanced by two-dimensional rotational position coding to obtain the infrared image feature dataset and the SAR image feature dataset; Furthermore, such as Figure 3The flowchart shown is a bimodal basic visual feature extraction process of an embodiment of the multimodal semantic generation method for infrared and SAR images of the present invention. In step S41, based on the original infrared image dataset and the original SAR image dataset, the basic visual feature extraction of the visual encoder with an adaptive network yields a first-stage infrared image feature dataset and a first-stage SAR image feature dataset, including: S411. Based on the original infrared image dataset and the original SAR image dataset, obtain the infrared image patch dataset and the SAR image patch dataset through image segmentation; S412. Based on the infrared image patch dataset and the SAR image patch dataset, the first-stage infrared image feature dataset and the first-stage SAR image feature dataset are obtained through the embedding layer mapping of the visual encoder with the adaptive network.

[0023] For example, in this visual representation stage, the model uses the Vision Transformer as the backbone for image feature extraction. First, the input infrared or SAR image is divided into fixed-size image patches and mapped to a series of visual token sequences. Then, filtering and weighted optimization are performed. Unlike traditional natural image encoding methods that rely on statistical features such as color and texture, this model pays more attention to visual information that is more stable in infrared and SAR imaging, such as structural contours, geometric shapes, boundary continuity, and intensity distribution, during training and adaptation. Heat source areas in infrared images are usually characterized by concentrated high brightness, while SAR images contain bright spots, linear structures, and speckle noise determined by scattering mechanisms. Therefore, the visual encoder emphasizes spatial structural relationships and shape expression capabilities when modeling features, thereby providing visual representations with physical meaning and scene consistency for subsequent semantic generation. Secondly, in order to control computational costs while ensuring high-resolution detail representation, the visual encoder internally adopts a mechanism that combines local window attention with a small amount of global attention. Most Transformer layers perform self-attention computation within a local window to effectively model neighborhood structure and local texture morphology, while a few layers introduce global attention to maintain the ability to model long-distance spatial relationships. This design retains the advantage of VisionTransformer in expressing global dependencies while avoiding the problem of excessive computational complexity under high-resolution input, making it particularly suitable for infrared and SAR scenes with sparse targets and a large background ratio. Finally, the visual encoder introduces two-dimensional rotational position encoding to enhance the ability to characterize spatial structure and geometric layout, enabling the model to better understand the relative positional relationships between targets and the spatial distribution patterns between regions.

[0024] S5. Based on the perception annotation library, the infrared image feature dataset, the SAR image feature dataset, and the semantic requirement instructions, a structured semantic dataset is obtained through manual annotation and automatic generation of structured semantics; Specifically, such as Figure 4 The flowchart shown is a structured semantic automatic generation flowchart of an embodiment of the multimodal semantic generation method for infrared and SAR images of the present invention. In step S5, based on the perceptual annotation library, the infrared image feature dataset, the SAR image feature dataset, and the semantic requirement instruction, a structured semantic dataset is obtained through manual annotation and automatic generation of structured semantics, including: S51. Based on this perceptual annotation library, an artificially structured semantic description text library is obtained through artificial natural language organization; S52. The perception annotation library, the infrared image feature dataset, the SAR image feature dataset, and the semantic requirement instructions are used to train a multimodal large model under the supervision of the artificially structured semantic description text library to obtain a multimodal structured semantic automatic generation model. S53. Based on the perception annotation library, the infrared image feature dataset, the SAR image feature dataset, and the semantic requirement instruction, the structured semantic dataset is obtained through reasoning using the multimodal structured semantic automatic generation model; Furthermore, such as Figure 5 The flowchart shown is from the present invention, an embodiment of the multimodal semantic generation method for infrared and SAR images, illustrating the generation of high-level structured semantics from target-layer semantics. In step S53, based on the perceptual annotation library, the infrared image feature dataset, the SAR image feature dataset, and the semantic requirement instruction, the structured semantic dataset is obtained through inference using the multimodal structured semantic automatic generation model, including: S531. Based on this perception annotation library, through statistical processing, the target layer semantics are obtained, including: target type set, target bounding box set, target location set, and target quantity set; S532. Based on the semantics of the target layer, the semantic vector of the relation layer is obtained by calculating the geometric relationship based on the set of target boxes, including: regional distribution features, clustering features and arrangement morphology features; S533. Based on the infrared image feature dataset, the SAR image feature dataset, the relational layer semantic vector, and the semantic requirement instruction, the structured semantic dataset is obtained through reasoning by the multimodal structured semantic automatic generation model.

[0025] For example, for each infrared or SAR image, a set of perceived labeled targets. Recorded as: (1) Where N represents the total number of target instances in the image, denotes the target class label corresponding to the \(i\)-th instance, and all existing are integrated into the target class vector \(\mathbf{c}\). denotes the corresponding target bounding box. Based on the instance set \(\mathcal{O}\), the semantic construction of the target layer is carried out, and the number of targets of each category is counted: (2) where \(k\) represents the category index, and \(1(\cdot)\) is an indicator function that takes 1 when the condition in the parentheses holds and 0 otherwise. By counting the number of instances of targets of each category, the target number distribution vector \(\mathbf{n}=(n_1,n_2,\cdots,n_K)\) can be obtained. , ,\(\cdots\), ). The target class vector \(\mathbf{c}\) and the target number distribution vector \(\mathbf{n}\) directly constitute the semantic of the target layer.

[0026] Based on the semantic of the target layer, this study further constructs the relational layer semantic vector \(\mathbf{r}\) that reflects the spatial organizational structure by using the geometric distribution information of the target bounding boxes. The relational layer semantics pays more attention to the spatial distribution and organization of the current targets, and mainly constructs the spatial distribution pattern of the targets through regional statistics. Specifically, the image is divided into regular grid regions, and the distribution ratios of different category targets in each region are counted, so as to obtain the distribution feature vector which mainly reflects whether the targets are concentrated on one side, evenly distributed, or show a local dense phenomenon.

[0027] Subsequently, the degree of aggregation or dispersion is statistically analyzed through the distances between targets. Let the average normalized distance between targets of the same category be: (3) where and represent the centers of each target bounding box normalized to the image coordinates, with \(i \lt j\). A smaller indicates that the target distribution is more aggregated, while a larger indicates that the target distribution is more dispersed. The distance statistics of all categories together constitute the aggregation degree feature . In addition, by analyzing the main direction of the overall target distribution, the arrangement pattern can be obtained. By simply performing directional statistics on all target position points, the main distribution direction and linearity degree can be obtained, which are used to distinguish spatial patterns such as linear arrangement, block distribution, or no obvious structure. This feature is denoted as . Finally, the relational layer semantic feature \(\mathbf{r}\) can be uniformly represented as: (4) Based on obtaining the target layer semantic vectors n and c and the relation layer semantic vector r, this study further introduces a multimodal large language model to achieve fine-grained generation of high-level natural language scene semantics from the target layer semantics and relation layer semantics: (5) Where v represents the visual features of the image. The structured semantic conditions are defined by c (target category set), n (target quantity statistics), r (spatial relationship features), and p (target location set). q is the semantic requirement instruction, which guides the model to generate text descriptions containing multi-granular semantic information, such as: "Please describe the target category and quantity based on the image content, and summarize the overall scene attributes." Under this instruction constraint, the multimodal large language model needs to combine visual evidence and structured semantic information to generate high-level scene semantic text consistent with the target category, quantity statistics, and spatial organization features, thereby achieving a controllable and interpretable semantic generation process.

[0028] S6. Based on the infrared image feature dataset and the SAR image feature dataset, a dual-modal aggregated visual feature dataset is obtained through fusion and local aggregation processing; Specifically, such as Figure 6 The flowchart shown is for obtaining bimodal aggregated visual features in an embodiment of the multimodal semantic generation method for infrared and SAR images of the present invention. In step S6, based on the infrared image feature dataset and the SAR image feature dataset, a bimodal aggregated visual feature dataset is obtained through fusion and local aggregation processing, including: S61. Based on the infrared image feature dataset and the SAR image feature dataset, a dual-modal fused visual feature dataset is obtained through transformation and fusion; S62. Based on the bimodal fusion visual feature dataset, a bimodal aggregated visual feature dataset is obtained through local aggregation processing.

[0029] S7. Based on the bimodal aggregated visual feature dataset, a multimodal mapping network is trained under the supervision of the structured semantic dataset to obtain a cross-domain transformation network model; Specifically, such as Figure 7 The flowchart shown is from an embodiment of the multimodal semantic generation method for infrared and SAR images of the present invention, illustrating the training of a cross-domain conversion network. In step S7, based on the bimodal aggregated visual feature dataset, the multimodal mapping network is trained under the supervision of the structured semantic dataset to obtain the cross-domain conversion network model, including: S71. Based on this structured semantic dataset, word-level semantic vectors are obtained through the word embedding mechanism of the language model; S72. Based on the bimodal aggregated visual feature dataset, a multimodal mapping network is trained under the supervision of word-level semantic vectors to obtain a cross-domain conversion network model.

[0030] For example, this stage belongs to the cross-modal semantic alignment stage, where the training focus shifts from visual statistical adaptation to visual-semantic mapping optimization: while keeping the visual backbone basically frozen, the visual-language alignment module is fine-tuned on a small scale to ensure that visual tokens can be stably mapped to the corresponding text space. The supervision signal for this stage is provided by a structured semantic dataset, which includes target category, target quantity, and fine-grained scene semantic information. Its goal is to enhance the conditional constraint ability of structured semantics on the multimodal mapping network, and establish a stable cross-modal semantic alignment foundation for the subsequent semantic generation stage.

[0031] S8. Based on the bimodal aggregated visual feature dataset, the visual semantic vector dataset is obtained by transformation and alignment through the cross-domain transformation network model.

[0032] S9. Based on the visual semantic vector dataset, the semantic requirement instruction and modal cue vector, the parameter efficiency module of the large language model is trained under the supervision of the structured semantic dataset to obtain the efficient parameter tuning large language model.

[0033] For example, in the semantic generation and modality modulation stage, after completing visual adaptation and semantic alignment, the final semantic generation optimization stage begins. This stage primarily trains the parameter-efficient modules on the large language model side, thereby enhancing the model's semantic generation capability with minimal parameter cost. Only the low-rank subspace of the attention matrix on the language model side is updated: (6) in, The original weights are frozen, and A and B are low-rank trainable matrices.

[0034] During training, standard text descriptions constructed by the multimodal structured semantic automatic generation model are used as training targets. The parameter efficiency module of the large language model is optimized through language generation loss to ensure that the generated semantics are consistent with the real perceived semantics in terms of target category, quantity and scene.

[0035] S10. Based on the visual semantic vector dataset, the semantic requirement instruction and modal cue vector, and through the reasoning of the efficient parameter-tuned large language model, fine-grained semantic text is obtained.

[0036] For example, to address the issues of different imaging mechanisms and significant differences in texture statistical properties between infrared and SAR, an explicit modality cue vector m is constructed during training to identify the image source modality (infrared, SAR, or hybrid modality), and together with the semantic generation requirement cue q, serves as the conditional input to the large language model.

[0037] Ultimately, the semantic generation process can be represented as: (7) Where I is the input image, m is the modal cue vector, and q is the semantic generation requirement cue. The modal cue vector m plays a conditional modulation role in the language decoding stage, enabling the model to form a modality-specific expression bias in the shared parameter space: when generating semantics from SAR images, more emphasis is placed on structural contours and scattering patterns, while infrared images focus more on thermal targets and temperature difference information. The mixed mode simultaneously highlights structural contours and scattering patterns as well as thermal targets and temperature difference information.

[0038] Furthermore, the semantic generation results have different emphases under the input constraints of modal cue vectors and semantic generation requirement cue vectors.

[0039] Question and Answer 1: User question: [SAR] Please describe the target category and quantity based on the image content; Model response: The image shows ground vehicle targets from a SAR perspective, totaling 6 vehicles.

[0040] Question and Answer 2: User Question: [TIR] Please describe the overall scene. Model Answer: This is a road traffic scene from a thermal infrared (TIR) ​​perspective. The image is taken from a high-altitude overhead view. There are no obvious buildings or vegetation with high temperatures in the surrounding area. It is a relatively open road or parking area. There are a total of 5 vehicles in the image, which are arranged in a line along one side of the road.

[0041] Fine-grained semantic information generation of multimodal images based on large models meets the usage requirements.

[0042] like Figure 8 The diagram shown is a system block diagram of an embodiment of the multimodal semantic generation system for infrared and SAR images according to the present invention. The present invention provides a multimodal semantic generation system for infrared and SAR images, which is applied to a multimodal semantic generation method for infrared and SAR images. The system includes: a dual-modal image acquisition module, a perceptual annotation module, a visual encoder adaptation module, a dual-modal visual feature acquisition module, a structured semantic automatic generation module, a dual-modal aggregated visual feature acquisition module, a cross-domain conversion network training module, a visual semantic conversion module, a large language model efficient parameter tuning module, and a fine-grained semantic text generation module. Specifically... The dual-modal image acquisition module is used to acquire infrared and SAR images using remote sensing sensors to obtain raw infrared and SAR image datasets. The perception annotation module is used to obtain a perception annotation library by manually annotating the original infrared image dataset and the original SAR image dataset. The visual encoder adaptation module is used to train the visual encoder adaptation network based on the original infrared image dataset and the original SAR image dataset through the perceptual annotation library to obtain a visual encoder with the adaptation network. The dual-modal visual feature acquisition module is used to obtain the infrared image feature dataset and the SAR image feature dataset by reasoning through the visual encoder with the adaptive network based on the original infrared image dataset and the original SAR image dataset. The automatic generation module for structured semantics is used to obtain a structured semantic dataset by combining manual annotation and automatic generation of structured semantics based on the perception annotation library, the infrared image feature dataset, the SAR image feature dataset and semantic requirement instructions. The dual-modal aggregated visual feature acquisition module is used to obtain a dual-modal aggregated visual feature dataset by fusing and local aggregation processing based on the infrared image feature dataset and the SAR image feature dataset. The cross-domain transformation network training module is used to train a multimodal mapping network under the supervision of the structured semantic dataset based on the bimodal aggregated visual feature dataset to obtain the cross-domain transformation network model. The visual semantic transformation module is used to transform and align the bimodal aggregated visual feature dataset through the cross-domain transformation network model to obtain a visual semantic vector dataset. The efficient parameter tuning module for large language models is used to train the parameters of a large language model under supervision using the visual semantic vector dataset, the semantic requirement instruction and modal cue vector, through the structured semantic dataset, and obtain an efficient parameter tuning module for large language models. The fine-grained semantic text generation module is used to obtain fine-grained semantic text based on the visual semantic vector dataset, the semantic requirement instruction and the modal cue vector through reasoning of the efficient parameter-tuned large language model.

[0043] This invention provides a method and system for multimodal semantic generation of infrared and SAR images, belonging to the field of computer vision. First, the invention manually annotates infrared and SAR images to obtain low-level image annotations and a small amount of semantic description information. Then, based on the low-level image annotations and the small amount of semantic description information, a structured semantic automatic generation framework under perceptual constraints is constructed. By mapping low-level visual perceptual information to a high-level fine-grained semantic space, an automatic hierarchical transformation from pixel features to structured annotations is achieved. Second, in terms of model architecture design, a fusion scheme of a visual encoder and a large language model is adopted. Through a lightweight visual-language alignment mapping mechanism, the modal features of infrared and SAR are coupled to a unified semantic representation space. Finally, guided by modal cue vectors and semantic requirement instructions, language decoding is performed through a large language model, realizing a comprehensive expression of spatial topological logic and overall scene attributes.

[0044] It is understood that the present invention has been described through the above embodiments and should not be construed as limiting the implementation and scope of the present invention. Those skilled in the art will recognize that various changes or equivalent substitutions can be made to these features and embodiments without departing from the spirit and scope of the present invention. Furthermore, under the teachings of the present invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application are within the protection scope of the present invention.

Claims

1. A method for generating multimodal semantics for infrared and SAR images, characterized in that, The method includes: S1. Use remote sensing sensors to acquire infrared and SAR images to obtain raw datasets of infrared and SAR images; S2. Based on the original infrared image dataset and the original SAR image dataset, a perception annotation library is obtained through manual annotation; S3. Based on the original infrared image dataset and the original SAR image dataset, train the visual encoder adaptation network through the perceptual annotation library to obtain a visual encoder with an adaptation network; S4. Based on the original infrared image dataset and the original SAR image dataset, the infrared image feature dataset and the SAR image feature dataset are obtained through inference by the visual encoder with the adaptive network; S5. Based on the perception annotation library, the infrared image feature dataset, the SAR image feature dataset, and the semantic requirement instructions, a structured semantic dataset is obtained through manual annotation and automatic generation of structured semantics; S6. Based on the infrared image feature dataset and the SAR image feature dataset, a dual-modal aggregated visual feature dataset is obtained through fusion and local aggregation processing; S7. Based on the bimodal aggregated visual feature dataset, a multimodal mapping network is trained under the supervision of the structured semantic dataset to obtain a cross-domain conversion network model; S8. Based on the bimodal aggregated visual feature dataset, the visual semantic vector dataset is obtained by transformation and alignment through the cross-domain transformation network model; S9. Based on the visual semantic vector dataset, the semantic requirement instructions and modal cue vectors, the parameter efficiency module of the large language model is trained under the supervision of the structured semantic dataset to obtain a large language model with efficient parameter tuning; S10. Based on the visual semantic vector dataset, the semantic requirement instructions, and the modal cue vectors, fine-grained semantic text is obtained through reasoning by the efficient parameter-tuned large language model.

2. The multimodal semantic generation method for infrared and SAR images according to claim 1, characterized in that, In step S4, based on the original infrared image dataset and the original SAR image dataset, the infrared image feature dataset and the SAR image feature dataset are obtained through inference by the visual encoder with the adaptive network, including: S41. Based on the original infrared image dataset and the original SAR image dataset, the first-stage infrared image feature dataset and the first-stage SAR image feature dataset are obtained through the basic visual feature extraction of the visual encoder with the adaptive network. S42. Based on the first-stage infrared image feature dataset and the first-stage SAR image feature dataset, the second-stage infrared image feature dataset and the second-stage SAR image feature dataset are obtained through filtering and key weighted optimization. S43. Based on the second-stage infrared image feature dataset and the second-stage SAR image feature dataset, the third-stage infrared image feature dataset and the third-stage SAR image feature dataset are obtained through a mixture of local window attention and a small amount of global attention. S44. Based on the third-stage infrared image feature dataset and the third-stage SAR image feature dataset, spatial features are enhanced by two-dimensional rotational position coding to obtain the infrared image feature dataset and the SAR image feature dataset.

3. The multimodal semantic generation method for infrared and SAR images according to claim 2, characterized in that, In step S41, based on the original infrared image dataset and the original SAR image dataset, the first-stage infrared image feature dataset and the first-stage SAR image feature dataset are obtained through basic visual feature extraction by the visual encoder with the adaptive network, including: S411. Based on the original infrared image dataset and the original SAR image dataset, obtain the infrared image patch dataset and the SAR image patch dataset through image segmentation; S412. Based on the infrared image patch dataset and the SAR image patch dataset, the first-stage infrared image feature dataset and the first-stage SAR image feature dataset are obtained through the embedding layer mapping of the visual encoder with the adaptive network.

4. The multimodal semantic generation method for infrared and SAR images according to claim 1, characterized in that, In step S5, based on the perception annotation library, the infrared image feature dataset, the SAR image feature dataset, and the semantic requirement instructions, a structured semantic dataset is obtained through manual annotation and automatic generation of structured semantics, including: S51. Based on the aforementioned perceptual annotation library, an artificially structured semantic description text library is obtained through artificial natural language organization; S52. The perception annotation library, the infrared image feature dataset, the SAR image feature dataset, and the semantic requirement instructions are used to train a multimodal large model under the supervision of the artificially structured semantic description text library to obtain a multimodal structured semantic automatic generation model. S53. Based on the perception annotation library, the infrared image feature dataset, the SAR image feature dataset, and the semantic requirement instruction, a structured semantic dataset is obtained through reasoning using the multimodal structured semantic automatic generation model.

5. The multimodal semantic generation method for infrared and SAR images according to claim 4, characterized in that, In step S53, based on the perception annotation library, the infrared image feature dataset, the SAR image feature dataset, and the semantic requirement instruction, a structured semantic dataset is obtained through reasoning using the multimodal structured semantic automatic generation model, including: S531. Based on the aforementioned perception annotation library, through statistical processing, the target layer semantics are obtained, including: a set of target types, a set of target boxes, a set of target locations, and a set of target quantities; S532. Based on the target layer semantics, a relation layer semantic vector is obtained by calculating the geometric relationships based on the target box set, including: regional distribution features, clustering features, and arrangement pattern features; S533. Based on the infrared image feature dataset, the SAR image feature dataset, the relational layer semantic vector, and the semantic requirement instruction, a structured semantic dataset is obtained through reasoning using the multimodal structured semantic automatic generation model.

6. The multimodal semantic generation method for infrared and SAR images according to claim 1, characterized in that, In step S6, based on the infrared image feature dataset and the SAR image feature dataset, a dual-modal aggregated visual feature dataset is obtained through fusion and local aggregation processing, including: S61. Based on the infrared image feature dataset and the SAR image feature dataset, a dual-modal fused visual feature dataset is obtained through transformation and fusion; S62. Based on the bimodal fusion visual feature dataset, a bimodal aggregated visual feature dataset is obtained through local aggregation processing.

7. The multimodal semantic generation method for infrared and SAR images according to claim 1, characterized in that, In step S7, based on the bimodal aggregated visual feature dataset, a multimodal mapping network is trained under the supervision of the structured semantic dataset to obtain a cross-domain conversion network model, including: S71. Based on the structured semantic dataset, obtain word-level semantic vectors through the word embedding mechanism of the language model; S72. Based on the bimodal aggregated visual feature dataset, train the multimodal mapping network through the word-level semantic vector supervision to obtain the cross-domain conversion network model.

8. A multimodal semantic generation system for infrared and SAR images, used to implement the multimodal semantic generation method for infrared and SAR images as described in any one of claims 1-7, characterized in that, The system includes: The dual-modal image acquisition module is used to acquire infrared and SAR images using remote sensing sensors to obtain raw infrared and SAR image datasets. The perception annotation module is used to obtain a perception annotation library by manually annotating the original infrared image dataset and the original SAR image dataset. The visual encoder adaptation module is used to train a visual encoder adaptation network based on the infrared image raw dataset and the SAR image raw dataset through the perceptual annotation library to obtain a visual encoder with an adaptation network. The dual-modal visual feature acquisition module is used to obtain the infrared image feature dataset and the SAR image feature dataset by reasoning through the visual encoder with the adaptive network based on the original infrared image dataset and the original SAR image dataset. The structured semantic automatic generation module is used to obtain a structured semantic dataset by means of manual annotation and automatic generation of structured semantics, based on the perception annotation library, the infrared image feature dataset, the SAR image feature dataset and semantic requirement instructions. The dual-modal aggregated visual feature acquisition module is used to obtain a dual-modal aggregated visual feature dataset by fusing and local aggregation processing based on the infrared image feature dataset and the SAR image feature dataset. The cross-domain transformation network training module is used to train a multimodal mapping network under the supervision of the structured semantic dataset based on the bimodal aggregated visual feature dataset to obtain a cross-domain transformation network model. The visual semantic transformation module is used to transform and align the bimodal aggregated visual feature dataset through the cross-domain transformation network model to obtain a visual semantic vector dataset. The efficient parameter tuning module for the large language model is used to train the parameters of the large language model under supervision through the structured semantic dataset based on the visual semantic vector dataset, the semantic requirement instructions and modal cue vectors, so as to obtain the efficient parameter tuning module for the large language model. The fine-grained semantic text generation module is used to obtain fine-grained semantic text through reasoning of the efficient parameter-tuned large language model based on the visual semantic vector dataset, the semantic requirement instructions and modal cue vectors.