Fine-grained remote sensing target detection method and device based on visual language instance fusion

By constructing a fine-grained remote sensing object detection method that integrates visual language instances, the deep fusion of visual and language information is used to solve the problem of weak inter-class differences in fine-grained remote sensing object detection, and the accurate detection of fine-grained targets is achieved.

CN120259795AActive Publication Date: 2025-07-04NAT UNIV OF DEFENSE TECH
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510746824.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-07-04
Estimated Expiration
2045-06-05

AI Technical Summary

Technical Problem

In the prior art, in fine-grained remote sensing target detection, weak inter-class differences lead to low classification accuracy, making it difficult to accurately identify fine-grained targets.

Method used

A fine-grained remote sensing object detection method based on visual language instance fusion is constructed, features are extracted through visual encoder and language encoder, and features are interacted and enhanced by visual language instance fusion module and depth enhancement encoder to achieve deep fusion of visual and language information and improve detection accuracy.

Benefits of technology

Through the deep fusion of visual and linguistic information, the accuracy of fine-grained remote sensing object detection is improved, the target characteristics can be better reflected, adapt to the changes and diversity of different input images, and the detection accuracy is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259795A_ABST
    Figure CN120259795A_ABST
Patent Text Reader

Abstract

The invention relates to a fine-grained remote sensing target detection method and device based on visual language instance fusion. The method comprises the steps that remote sensing image features are extracted according to a visual encoder, and category embedding features are extracted through a language encoder; inputting the target instance and the remote sensing image features into a visual language instance fusion module, obtaining visual instance features through an instance feature extractor, and storing the visual instance features in an instance feature memory area; updating the average instance feature stored in the instance feature memory area, inputting the category embedded feature and the visual average instance feature into a visual language fusion layer for interaction and fusion, and updating the instance feature stored in the instance feature memory area according to the feature after interaction and fusion; and enhancing the remote sensing image features and the interactively fused features according to a visual language depth enhancement encoder, and inputting the enhanced remote sensing image features and category embedded features into a detection head to obtain a detection result. By adopting the method, the fine-grained target can be detected more accurately.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of image processing, and in particular, to a fine-grained remote sensing target detection method and device based on visual-language instance fusion. Background Art

[0002] Fine-grained remote sensing target detection aims to accurately identify and distinguish the subtle subclasses of targets (such as aircraft models, ship types, crop varieties) from high-resolution remote sensing images. Its core challenges lie in the tiny target size, weak inter-class differences, and complex background interference. This technology is widely used in precision agriculture (monitoring crop disease subtypes) and ecological protection (tracking endangered species subgroups). The weak inter-class differences are the most core difficulty in fine-grained remote sensing target detection.

[0003] Existing methods generally introduce contrastive learning losses to increase the feature differences between different classes, but the accuracy of fine-grained classification in this way is low. Summary of the Invention

[0004] Based on this, in view of the above technical problems, it is necessary to provide a fine-grained remote sensing target detection method and device based on visual-language instance fusion that can effectively fuse visual instance features and language category concepts to achieve more accurate detection of fine-grained targets.

[0005] A fine-grained remote sensing target detection method based on visual-language instance fusion, the method comprising: Obtaining an input remote sensing image and category text; the category text is a string obtained by concatenating all the names of the categories to be detected; Constructing a fine-grained remote sensing image target detection network model; the fine-grained remote sensing image target detection network model includes a visual encoder, a language encoder, a visual-language instance fusion module, a visual-language depth enhancement encoder, and a detection head; Inputting the input remote sensing image and category text into the fine-grained remote sensing image target detection network model, extracting remote sensing image features according to the visual encoder, and extracting category embedding features through the language encoder; inputting the target instance and remote sensing image features into the visual-language instance fusion module through an instance feature extractor to obtain visual instance features and storing them in an instance feature memory area; updating the average instance features stored in the instance feature memory area to obtain visual average instance features; inputting the category embedding features and visual average instance features into a visual-language fusion layer for interaction and fusion to obtain interactively fused features; updating the instance features stored in the instance feature memory area according to the interactively fused features; enhancing the remote sensing image features and the interactively fused features according to the visual-language depth enhancement encoder to obtain enhanced remote sensing image features and category embedding features; inputting the enhanced remote sensing image features and category embedding features into the detection head to obtain a detection result.

[0006] A fine-grained remote sensing target detection device based on visual-language instance fusion, the device comprising: A data acquisition module, configured to acquire an input remote sensing image and class text; the class text is a string obtained by splicing all the names of classes to be detected; A model construction module, configured to construct a fine-grained remote sensing image target detection network model; the fine-grained remote sensing image target detection network model includes a visual encoder, a language encoder, a visual-language instance fusion module, a visual-language depth enhancement encoder, and a detection head; A fine-grained remote sensing target detection module, configured to input the input remote sensing image and class text into the fine-grained remote sensing image target detection network model, extract remote sensing image features according to the visual encoder, and extract class embedding features through the language encoder; input the target instance and remote sensing image features into the visual-language instance fusion module through an instance feature extractor to obtain visual instance features and store them in an instance feature memory area; update the average instance features stored in the instance feature memory area to obtain visual average instance features; input the class embedding features and the visual average instance features into a visual-language fusion layer for interaction and fusion to obtain interactively fused features; update the instance features stored in the instance feature memory area according to the interactively fused features; enhance the remote sensing image features and the interactively fused features according to the visual-language depth enhancement encoder to obtain enhanced remote sensing image features and class embedding features; input the enhanced remote sensing image features and class embedding features into the detection head to obtain a detection result.

[0007] The above-mentioned fine-grained remote sensing target detection method and device based on visual-language instance fusion. In this application, a fine-grained remote sensing image target detection network model is constructed, which includes multiple modules. The visual encoder and the language encoder are responsible for extracting remote sensing image features and category embedding features respectively, realizing the preliminary extraction of information from two modalities of vision and language, and providing rich basic features for subsequent fusion operations. Among them, the instance feature extractor in the visual-language instance fusion module extracts visual instance features from the remote sensing image features and stores them in the instance feature memory area. This design can specifically capture and save the instance features of the target, providing a data basis for subsequent utilization of these features. The average instance features stored in the instance feature memory area are updated to obtain visual average instance features. Through continuous updating, the model can adapt to different input images, better capture the changes and diversities of the target, and improve the model's representation ability for fine-grained features. The category embedding features and the visual average instance features are input into the visual-language fusion layer for interaction and fusion. Through methods such as the bidirectional multi-head attention mechanism, the deep fusion of visual and language information is realized. This fusion method can make full use of the language category concept to guide and supplement the visual instance features, and at the same time let the visual instance features refine and enhance the language category embedding, so that the fused features contain both rich visual details and accurate category concept information. The visual-language deep enhancement encoder enhances the remote sensing image features and the features after interaction and fusion, further improving the quality and representation ability of the features. The enhanced features can better reflect the characteristics of fine-grained targets, provide more accurate inputs for the detection head, and thus improve the detection accuracy. The entire model can automatically learn how to effectively fuse visual instance features and language category concepts from the input remote sensing images and category texts through end-to-end training, and achieve accurate detection of fine-grained targets. During the training process, the model can continuously adjust the parameters of each module according to the loss function, optimize the process of feature extraction and fusion, and adapt to different dataset and task requirements. Description of the Drawings

[0008] Figure 1 It is a schematic flow chart of a fine-grained remote sensing target detection method based on visual-language instance fusion in an embodiment; Figure 2 It is a schematic diagram of a fine-grained remote sensing image target detection network model in an embodiment; Figure 3 It is a schematic diagram of a fine-grained remote sensing target detection device based on visual-language instance fusion in an embodiment; Figure 4 It is an internal structure diagram of a computer device in an embodiment. Detailed Embodiments

[0009] To make the objectives, technical solutions, and advantages of this application more clear and understandable, the following further elaborates on this application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely used to explain this application and are not used to limit this application.

[0010] In one embodiment, as Figure 1 shown, a fine-grained remote sensing object detection method based on visual-language instance fusion is provided, including the following steps: Step 102: Obtain an input remote sensing image and class text; the class text is a string obtained by concatenating all the class names to be detected.

[0011] The class text is a string obtained by concatenating all the class names to be detected with ".".

[0012] Step 104: Construct a fine-grained remote sensing image object detection network model; the fine-grained remote sensing image object detection network model includes a visual encoder, a language encoder, a visual-language instance fusion module, a visual-language depth enhancement encoder, and a detection head.

[0013] The structure of the constructed fine-grained remote sensing image object detection network model is as Figure 2 shown, where the visual-language instance fusion module includes the instance extractor and the bi-directional multi-head attention layer in the figure.

[0014] Step 106: Input the input remote sensing image and class text into the fine-grained remote sensing image object detection network model. Extract the remote sensing image features through the visual encoder, and extract the class embedding features through the language encoder; input the target instance and the remote sensing image features into the visual-language instance fusion module through the instance feature extractor to obtain visual instance features and store them in the instance feature memory area; update the average instance features stored in the instance feature memory area to obtain visual average instance features; input the class embedding features and the visual average instance features into the visual-language fusion layer for interaction and fusion to obtain the interactively fused features; update the instance features stored in the instance feature memory area according to the interactively fused features; enhance the remote sensing image features and the interactively fused features according to the visual-language depth enhancement encoder to obtain the enhanced remote sensing image features and class embedding features; input the enhanced remote sensing image features and class embedding features into the detection head to obtain the detection results.

[0015] Through the visual encoder of the network Extract the remote sensing image features , through the language encoder Extract the class embedding features , Denotes the maximum number of tokens, Denotes the hidden dimension. The target instance GT annotation box and the feature map The instance feature extractor in the input visual language instance fusion module extracts visual instance features , represents the number of instances, represents the number of feature channels, represents the size of the instance feature. Update the average instance feature stored in the instance feature memory area: , where represents the average instance feature of the -th iteration process for the -th category, represents the category 's -th instance feature, represents the category instance number. In particular, ; . Embed the category ( ( is the number of tokens for the category ) and the corresponding visual average instance feature are interacted and fused through a bi-directional multi-head attention layer; first, and are layer-normalized respectively, and then and are projected through linear layers to obtain the corresponding key vectors and , value vectors and , query vectors and ; calculate the attention matrix ; obtain the interacted embedding and feature: , where represents the linear layer; finally, the interacted feature is fused into the original feature , where is a learnable parameter. Use to update the instance features stored in the instance feature memory area again.

[0016] Input the image feature and the category embedding feature into the visual language global fusion encoder to obtain the enhanced image feature and the category embedding feature ; where the visual language global fusion encoder includes 6 network blocks, and each network block consists of a cross-modal multi-head attention module, a DyHead module, and a BERT layer. The cross-modal multi-head attention module is used to achieve cross-modal feature interaction, and the DyHead module is used to encode visual features and extract image region features , the BERT layer is used to encode language features. The image region features and the class embedding features are input into the detection head. The detection head includes two branches: localization and classification. The localization branch consists of one layer of convolutional layer. The classification branch involves calculating the similarity between the class embedding features and the image region features; the region features obtain box predictions through the localization branch, and the similarity between the class embedding features and the image region features is calculated in the classification branch: , this similarity is the classification score for each class. During the model training process, the output results of the classification branch are calculated in combination with the ground truth Focal Loss , and the output results of the localization branch are calculated IoU Loss to optimize the network parameters.

[0017] In the above process, based on the single-stage detector GLIP, it includes a parallel visual encoder and a language encoder, which are respectively used to extract the input image feature map and the class text embedding; a visual-linguistic instance fusion module, which is used to extract visual instance features from the feature map, continuously store and update the average instance features of each class, and perform two-way multi-head attention interaction fusion with the corresponding class embeddings; a visual-linguistic global fusion encoder, which is used to fuse the global feature map and all class embeddings; and a detection head, which is used to generate detection results, including two parts: classification and localization. In particular, the visual-linguistic instance fusion module consists of an instance feature extractor, an instance feature memory area, and a visual-linguistic fusion layer. The instance feature extractor is used to extract instance features from the feature map, the instance feature memory area is used to store the class average instance features, and the visual-linguistic fusion layer is used to fuse the visual instance features and the language class embeddings. Since the classification result is obtained by calculating the similarity between the class embedding and the image region features, in this application, by extracting visual instance features and continuously storing and updating, the visual instance features are fused with the language class embeddings, thereby realizing learning rich fine-grained information from the conceptual level and the instance level, further calculating the similarity between the fused language class embeddings and the image region features during the classification process, realizing the utilization of the visual instance fine-grained information introduced in the class embeddings, and thus improving the fine-grained remote sensing object detection accuracy.

[0018] In the above-mentioned fine-grained remote sensing target detection method based on visual-language instance fusion, the present application constructs a fine-grained remote sensing image target detection network model, which includes multiple modules. The visual encoder and the language encoder are respectively responsible for extracting remote sensing image features and category embedding features, realizing the preliminary extraction of information from two modalities of vision and language, and providing rich basic features for subsequent fusion operations. Among them, the instance feature extractor in the visual-language instance fusion module extracts visual instance features from the remote sensing image features and stores them in the instance feature memory area. This design can specifically capture and save the instance features of the target, providing a data basis for subsequent utilization of these features. The average instance features stored in the instance feature memory area are updated to obtain visual average instance features. Through continuous updating, the model can adapt to different input images, better capture the changes and diversity of the target, and improve the model's representation ability for fine-grained features. The category embedding features and the visual average instance features are input into the visual-language fusion layer for interaction and fusion. Through methods such as the bidirectional multi-head attention mechanism, deep fusion of visual and language information is achieved. This fusion method can make full use of language category concepts to guide and supplement visual instance features, and at the same time let visual instance features refine and enhance language category embeddings, so that the fused features contain both rich visual details and accurate category concept information. The visual-language deep enhancement encoder enhances the remote sensing image features and the features after interaction and fusion, further improving the quality and representation ability of the features. The enhanced features can better reflect the characteristics of fine-grained targets, provide more accurate inputs for the detection head, and thus improve the detection accuracy. The entire model can automatically learn how to effectively fuse visual instance features and language category concepts from the input remote sensing images and category texts through end-to-end training, and achieve accurate detection of fine-grained targets. During the training process, the model can continuously adjust the parameters of each module according to the loss function, optimize the process of feature extraction and fusion, and adapt to different dataset and task requirements.

[0019] In one embodiment, updating the average instance features stored in the instance feature memory area to obtain visual average instance features includes: Updating the average instance features stored in the instance feature memory area to obtain visual average instance features as

[0020] Wherein, represents the average instance feature of the -th category in the -th iteration process, represents the -th instance feature of category , represents category Number of instances, ; 。

[0021] In one embodiment, the category embedding features and the visual average instance features are input into a bidirectional multi-head attention layer for interaction and fusion to obtain the features after interaction and fusion, including: The category embedding features and the visual average instance features are respectively layer-normalized, and then the category embedding features and the visual average instance features are respectively projected through a linear layer to obtain the corresponding key vectors and 、value vectors and 、query vectors and , where, represents the th category; Calculate the attention matrix , and then calculate the embedded and features after interaction according to the attention matrix; fuse the features after interaction into the original features to obtain the features after interaction and fusion.

[0022] In one embodiment, calculating the embedded and features after interaction according to the attention matrix includes: The embedded and features after interaction calculated according to the attention matrix are respectively ; where, represents the linear layer.

[0023] In one embodiment, fusing the features after interaction into the original features to obtain the features after interaction and fusion includes: Fusing the features after interaction into the original features, and the features after interaction and fusion obtained are ; where, is a learnable parameter.

[0024] In one embodiment, the vision-language depth enhancement encoder includes 6 network blocks, each network block consists of a cross-modal multi-head attention module, a DyHead module, and a BERT layer. The cross-modal multi-head attention module is used to achieve cross-modal feature interaction, the DyHead module is used to encode visual features and extract image region features, and the BERT layer is used to encode language features.

[0025] In one embodiment, the detection head includes two branches of localization and classification. The localization branch consists of one layer It consists of a convolutional layer. The classification branch involves calculating the similarity between the class embedding features and the image region features. The enhanced remote sensing image features and the class embedding features are input into the detection head to obtain the detection results, including: The enhanced remote sensing image features The box prediction is obtained through the localization branch, and the class embedding features Calculate the similarity with the enhanced remote sensing image features in the classification branch: , and the similarity is the classification score for each class. The class with the highest similarity is used as the classification result.

[0026] It should be understood that although Figure 1 each step in the flowchart of Figure 1 is shown in sequence according to the indication of the arrow, these steps are not necessarily executed in the order indicated by the arrow. Unless otherwise clearly stated in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover,

[0027] In one embodiment, as Figure 3 shown, a fine-grained remote sensing target detection device based on visual language instance fusion is provided, including: a data acquisition module 302, a model construction module 304, and a fine-grained remote sensing target detection module 306, where: The data acquisition module 302 is used to acquire the input remote sensing image and the class text; the class text is a string obtained by concatenating the names of all classes to be detected; The model construction module 304 is used to construct a fine-grained remote sensing image target detection network model; the fine-grained remote sensing image target detection network model includes a visual encoder, a language encoder, a visual language instance fusion module, a visual language depth enhancement encoder, and a detection head; The fine-grained remote sensing target detection module 306 is used to input the input remote sensing image and category text into the fine-grained remote sensing image target detection network model, extract the remote sensing image features according to the visual encoder, and extract the category embedding features through the language encoder; input the target instance and the remote sensing image features into the visual-language instance fusion module through the instance feature extractor, obtain the visual instance features and store them in the instance feature memory area; update the average instance features stored in the instance feature memory area to obtain the visual average instance features; input the category embedding features and the visual average instance features into the visual-language fusion layer for interaction and fusion to obtain the interactively fused features; update the instance features stored in the instance feature memory area according to the interactively fused features; enhance the remote sensing image features and the interactively fused features according to the visual-language deep enhancement encoder to obtain the enhanced remote sensing image features and category embedding features; input the enhanced remote sensing image features and category embedding features into the detection head to obtain the detection results.

[0028] For the specific limitations of the fine-grained remote sensing target detection device based on visual-language instance fusion, reference can be made to the limitations of the fine-grained remote sensing target detection method based on visual-language instance fusion in the above text, which will not be elaborated here. Each module in the above fine-grained remote sensing target detection device based on visual-language instance fusion can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor in the computer device in hardware form or be independent of it, or be stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0029] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 4 shown. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it realizes a fine-grained remote sensing target detection method based on visual-language instance fusion. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad set on the shell of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0030] Those skilled in the art can understand, Figure 4The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0031] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in this application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0032] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0033] The above-described embodiments only represent several implementation manners of this application, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of this application. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of this application should be subject to the appended claims.

Claims

1. A fine-grained remote sensing target detection method based on visual language instance fusion, characterized in that, The method includes: Obtaining an input remote sensing image and class text; the class text is a string obtained by splicing all the class names to be detected; Constructing a fine-grained remote sensing image object detection network model; the fine-grained remote sensing image object detection network model includes a visual encoder, a language encoder, a visual-language instance fusion module, a visual-language depth enhancement encoder, and a detection head; Inputting the input remote sensing image and class text into the fine-grained remote sensing image object detection network model, extracting remote sensing image features according to the visual encoder, and extracting class embedding features through the language encoder; inputting the target instance and remote sensing image features into the visual-language instance fusion module through the instance feature extractor to obtain visual instance features and storing them in the instance feature memory area; updating the average instance features stored in the instance feature memory area to obtain visual average instance features; inputting the class embedding features and visual average instance features into the visual-language fusion layer for interaction and fusion to obtain the interactively fused features; updating the instance features stored in the instance feature memory area according to the interactively fused features; enhancing the remote sensing image features and the interactively fused features according to the visual-language depth enhancement encoder to obtain the enhanced remote sensing image features and class embedding features; inputting the enhanced remote sensing image features and class embedding features into the detection head to obtain the detection result.

2. The method according to claim 1, characterized in that, Updating the average instance features stored in the instance feature memory area to obtain visual average instance features, including: Updating the average instance features stored in the instance feature memory area, and the obtained visual average instance features are Among them, represents the average instance feature of the -th category in the -th iteration process, represents the -th instance feature of the category , represents the number of instances of the category , ; .

3. The method according to claim 1, wherein Inputting the class embedding features and visual average instance features into a bidirectional multi-head attention layer for interaction and fusion to obtain the interactively fused features, including: Category embedding features and visual average instance features are respectively subjected to layer normalization, and then the category embedding features and visual average instance features are respectively projected through a linear layer to obtain corresponding key vectors and , value vectors and , query vectors and , where represents the th category; Calculate the attention matrix , and then calculate the interacted embeddings and features based on the attention matrix; fuse the interacted features into the original features to obtain the interactively fused features.

4. The method according to claim 3, wherein Calculating the interacted embeddings and features according to the attention matrix, including: The interacted embeddings and features calculated according to the attention matrix are respectively Among them, represents a linear layer.

5. The method according to claim 4, wherein Fusing the interacted features into the original features to obtain the interactively fused features, including: Fusing the interacted features into the original features, and the obtained interactively fused features are Among them, are learnable parameters.

6. The method according to claim 1, wherein The visual-language depth enhancement encoder includes 6 network blocks, each network block consists of a cross-modal multi-head attention module, a DyHead module, and a BERT layer. The cross-modal multi-head attention module is used to achieve cross-modal feature interaction, the DyHead module is used to encode visual features and extract image region features, and the BERT layer is used to encode language features.

7. The method according to claim 1, characterized in that The detection head includes two branches: positioning and classification. The positioning branch consists of one convolutional layer. The classification branch involves calculating the similarity between the class-embedded features and the image region features. The enhanced remote sensing image features and the class-embedded features are input into the detection head to obtain the detection results, including: The enhanced remote sensing image features The box prediction and class embedding features are obtained through the localization branch The similarity is calculated between the class embedding features and the enhanced remote sensing image features in the classification branch: , and the similarity is the classification score for each class. The class with the highest similarity is used as the classification result.

8. A fine-grained remote sensing target detection device based on visual language instance fusion, characterized in that The device includes: A data acquisition module, configured to acquire an input remote sensing image and class text; the class text is a string obtained by splicing all the class names to be detected; A model construction module, configured to construct a fine-grained remote sensing image object detection network model; the fine-grained remote sensing image object detection network model includes a visual encoder, a language encoder, a visual-language instance fusion module, a visual-language depth enhancement encoder, and a detection head; A fine-grained remote sensing target detection module is used to input the input remote sensing image and class text into a fine-grained remote sensing image target detection network model, extract remote sensing image features according to the visual encoder, and extract class embedding features through the language encoder; input the target instance and remote sensing image features into the visual-language instance fusion module through the instance feature extractor to obtain visual instance features and store them in the instance feature memory area; update the average instance features stored in the instance feature memory area to obtain visual average instance features; input the class embedding features and visual average instance features into the visual-language fusion layer for interaction and fusion to obtain the interactively fused features; update the instance features stored in the instance feature memory area according to the interactively fused features; enhance the remote sensing image features and the interactively fused features according to the visual-language depth enhancement encoder to obtain enhanced remote sensing image features and class embedding features; input the enhanced remote sensing image features and class embedding features into the detection head to obtain the detection results.

Citation Information

Patent Citations

  • Remote sensing image semantic segmentation method and system based on adaptive feature fusion

    CN118710914A

  • Feature mean fusion enhancement network based on attention mechanism

    CN118840767A

  • Multi-label image classification method and device, equipment, storage medium and program product

    CN119169339A

  • Remote sensing image target recognition method and system based on visual language model

    CN119625527A

  • Zero-sample food image detection method based on visual semantic bidirectional guidance

    CN119649365A