Method for improving fine-grained remote sensing target detection precision by utilizing attribute description
By constructing a remote sensing object detection model, using attribute description to enhance category embedding features, realizing the deep fusion of visual and language features, solving the problem of failing to effectively utilize attribute knowledge in the existing technology, and improving the accuracy of remote sensing object detection.
Patent Information
- Application Number
- CN202510746854.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-05
AI Technical Summary
The existing fine-grained remote sensing object detection methods fail to effectively utilize attribute knowledge, resulting in insufficient detection accuracy and difficulty in achieving fine-grained classification of targets.
A remote sensing object detection model is built, combining visual encoder, pre-trained language encoder, visual language depth enhancement encoder and bidirectional multi-head attention layer, enhance category embedding features through attribute description, realize the deep fusion of visual and language features, and use pre-trained language encoder to extract categories and attribute embeddings, and inject attribute knowledge into category embeddings through bidirectional multi-head attention layer to enrich the network's cognitive dimension of fine-grained targets.
The accuracy of remote sensing object detection is improved, allowing the model to identify targets more accurately based on attribute characteristics like human experts, improving the accuracy of fine-grained detection.
Smart Images

Figure CN120259796A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image detection technology, and particularly to a method for improving the accuracy of fine-grained remote sensing target detection by using attribute descriptions. Background Art
[0002] Remote sensing target detection is a core task at the intersection of computer vision and remote sensing technology, and has important applications in fields such as urban planning and ecological protection. Conventional remote sensing target detection only needs to locate the target and identify the large categories of the target, such as airplanes, ships, etc.; however, in practical applications, it is often necessary to understand the fine-grained classification of the target, such as the category or specific model of an airplane or a ship, in order to achieve more refined decision-making. Existing fine-grained remote sensing target detection methods usually improve the detection accuracy of fine-grained categories through methods such as multi-scale feature enhancement, contrast learning, and data augmentation. However, these methods only achieve fine-grained detection by optimizing visual features, and have never explored using language knowledge to enhance the detector's understanding of fine-grained targets. When human experts learn to identify fine-grained airplane models, they can often summarize the attribute characteristics of a certain model of airplane, such as the wing structure of the airplane, the shape of the tail, the number of engines, etc. More precisely, they can also understand data such as the wingspan and fuselage length of the airplane. These attribute knowledge helps human experts accurately identify airplane models. However, existing methods have not utilized attribute knowledge for fine-grained target detection. Summary of the Invention
[0003] Based on this, it is necessary to provide a method for improving the accuracy of fine-grained remote sensing target detection by using attribute descriptions for the above technical problems.
[0004] A method for improving the accuracy of fine-grained remote sensing target detection by using attribute descriptions, the method includes: Obtain an input remote sensing image, category text, and attribute text; construct a remote sensing target detection model; the remote sensing target detection model includes a visual encoder, a pre-trained language encoder, a visual-language depth enhancement encoder, a multi-head bidirectional attention layer, and a detection head; Extract remote sensing image features of the input remote sensing image according to the visual encoder; extract category embedding features and attribute embedding features of the category text and the attribute text by using the pre-trained language encoder; Input the category embedding features of each category and the corresponding attribute embedding features into the multi-head bidirectional attention layer for enhancement to obtain enhanced category embedding features; Input the remote sensing image features and the enhanced category embedding features into the visual-language depth enhancement encoder to obtain enhanced image features and secondarily enhanced category embedding features; Input the enhanced image features and the secondarily enhanced category embedding features into the detection head to obtain a detection result.
[0005] The method for improving the accuracy of fine-grained remote sensing target detection using attribute descriptions breaks through the limitations of traditional methods that only rely on visual features in the present application. It uses a pre-trained language encoder to extract class and attribute embeddings, and injects attribute knowledge into the class embeddings through a bidirectional multi-head attention layer, enriching the network's cognitive dimension of fine-grained targets. This enables the model to identify targets more accurately based on attribute characteristics, just like human experts. Secondly, a mechanism for deep fusion of vision and language is constructed. The language encoder is used to extract attribute annotation embedding features, which are then fused with class embedding features through bidirectional attention. In the detection head, the similarity between visual region features and class embedding features is calculated to improve the accuracy of fine-grained detection. At the same time, an innovative masked attribute reconstruction pre-training method enables the language encoder to understand attribute descriptions in combination with visual instances. The pre-trained parameters play a stable role in the detection network, ensuring the reliability of the utilization of attribute knowledge. Finally, by constructing multi-component collaborative operations, using visual encoders, language encoders, bidirectional multi-head attention layers, etc., it ensures that attribute descriptions are deeply integrated into the entire detection process, comprehensively improving the model's detection ability for fine-grained targets, and having important application value in fields such as urban planning. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] Figure 1 FIG. is a schematic flowchart of a method for improving the accuracy of fine-grained remote sensing target detection using attribute descriptions in one embodiment; Figure 2 FIG. is a schematic diagram of a remote sensing target detection model in one embodiment; Figure 3 FIG. is a schematic diagram of the pre-training process of a language encoder in one embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0007] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0008] In one embodiment, as Figure 1 shown, a method for improving the accuracy of fine-grained remote sensing target detection using attribute descriptions is provided, including the following steps: Step 102, obtain an input remote sensing image, class text, and attribute text; construct a remote sensing target detection model; the remote sensing target detection model includes a visual encoder, a pre-trained language encoder, a vision-language deep enhancement encoder, a bidirectional multi-head attention layer, and a detection head.
[0009] Obtain an input remote sensing image, class text, and attribute text, where the class text is a string obtained by concatenating all the names of the classes to be detected with ".", and the attribute text is the attribute descriptions of all classes, and each class attribute is a string. Construct a remote sensing target detection model asFigure 2 As shown in the figure, it includes a visual encoder, a pre-trained language encoder, a visual-language depth enhancement encoder, a bidirectional multi-head attention layer, and a detection head.
[0010] Step 104: Extract remote sensing image features of the input remote sensing image according to the visual encoder; use the pre-trained language encoder to extract the category embedding features and attribute embedding features of the category text and the attribute text.
[0011] Extract remote sensing image features through the visual encoder of the network , and extract category embedding features through the pre-trained language encoder and attribute embedding features . represents the maximum number of tokens, represents the hidden dimension, represents the number of categories.
[0012] Moreover, the present invention designs a masked attribute reconstruction pre-training method for the language encoder, including: Input the remote sensing image, the calibration box of each target therein, and the attribute text into the network. The attribute of each target is a string.
[0013] Randomly mask the attribute text. Specifically: randomly select 15% of the tokens; randomly select 10% of the tokens from the above 15% of the tokens and replace them with a random token; randomly select 80% of the tokens from the above 15% of the tokens and replace them with [MASK]; the remaining 10% of the tokens remain unchanged. Thus, the masked attribute and the original attribute are obtained.
[0014] Extract remote sensing image features through the visual encoder of the network , and extract masked attribute embedding features through the language encoder . represents the number of targets.
[0015] The instance feature extractor extracts the instance features of all targets from the image features through the annotation box to obtain . is the number of instance feature channels, is the instance feature size.
[0016] Map the attribute embedding features to the visual instance feature channels through a linear layer to obtain , and adjust the dimension of the instance features to obtain ; On the second dimension, for and Perform splicing, input the spliced features into the vision-language joint encoder for fusion, and obtain .
[0017] Input into the prediction head to generate attribute predictions, where the prediction head consists of two layers of linear layers.
[0018] Calculate the cross-entropy loss between the predicted attributes and the original attributes to optimize the network. Load the language encoder parameters obtained after the above pre-training process into the language encoder in the remote sensing object detection model, and freeze the language encoder parameters.
[0019] During the pre-training process, the network includes structures such as a vision encoder and a language encoder. By combining vision instances to predict attributes, a language encoder with a specific understanding of attribute descriptions is trained. When training the object detection network, load the language encoder parameters obtained from pre-training and cancel their gradient updates, only updating other network parameters. This ensures that the understanding of attribute descriptions learned by the language encoder can be stably applied to the detection model, providing a reliable basis for the model to utilize attribute knowledge during the detection process, thereby improving the fine-grained remote sensing object detection accuracy.
[0020] Step 106: Input the class embedding features and corresponding attribute embedding features of each class into the bidirectional multi-head attention layer for enhancement to obtain enhanced class embedding features.
[0021] Enhance the embedding features of each class and its corresponding attribute embedding features through the bidirectional multi-head attention layer. Taking class as an example, this process is expressed as: , where and represent the bidirectional multi-head attention layer and layer normalization respectively, is the class embedding feature of class , is the attribute embedding feature of class , is the number of tokens of class ; then enhance the class embedding features through residual connection: , where is the network learnable parameter, is the enhanced class embedding feature.
[0022] First, use the language encoder to extract the attribute annotation embedding features and perform bidirectional attention fusion with the class embedding features. This fusion method can better combine the attribute information with the class information, enabling the class concept to incorporate the details of the attribute description. Thus, when calculating the similarity between the visual region features and the class embedding features in the detection head, the judgment can be made based on a richer and more accurate class representation, improving the fine-grained detection accuracy.
[0023] Meanwhile, use the pre-trained language encoder to extract the class embedding and class attribute embedding, and inject the class attribute knowledge into the class embedding through the bidirectional multi-head attention layer. This enables the network to absorb and utilize the attribute knowledge, changing the previous limitation of relying only on visual features and enriching the understanding of fine-grained objects at the knowledge level. For example, human experts can more accurately identify the aircraft model by summarizing the attribute characteristics such as the wing structure and tail shape of the aircraft. Similarly, this method allows the network to learn this attribute knowledge, laying the foundation for more accurate detection.
[0024] Step 108: Input the remote sensing image features and the enhanced class embedding features into the visual-language deep enhancement encoder to obtain the enhanced image features and the secondarily enhanced class embedding features.
[0025] Input the image features and the class embedding features into the visual-language deep enhancement encoder to obtain the enhanced image features and the class embedding features ; where the visual-language deep enhancement encoder includes 6 network blocks, and each network block consists of a cross-modal multi-head attention module, a DyHead module, and a BERT layer. The cross-modal multi-head attention module is used to achieve cross-modal feature interaction, the DyHead module is used to encode visual features and extract image region features , and the BERT layer is used to encode language features.
[0026] Input the remote sensing image features and the enhanced class embedding features into the visual-language deep enhancement encoder to further enhance the interaction between visual and language features. Through this deep interaction, the image features can obtain the supplementary information brought by the attribute knowledge, and at the same time, the class embedding features can better match the target features in the image, enabling the model to more accurately understand the features of fine-grained objects, thereby improving the detection accuracy. Step 110: Input the enhanced image features and the secondarily enhanced class embedding features into the detection head to obtain the detection results.
[0027] The image region features and the class embedding features are input into the detection head. The detection head includes two branches: localization and classification. The localization branch consists of one layer The convolutional layer is composed. The classification branch involves the calculation of the similarity between the class embedding features and the image region features; the region features The box prediction is obtained through the localization branch, and the similarity between the class embedding features and the image region features is calculated in the classification branch: This similarity is the classification score for each class.
[0028] When the model starts training, the pre-trained parameters of the visual encoder and the language encoder are loaded, and the parameters of the language encoder are frozen. During the model training process, the Focal Loss is calculated for the output results of the classification branch combined with the ground truth, and the IoU Loss is calculated for the output results of the localization branch to optimize the network parameters.
[0029] For the method of improving the fine-grained remote sensing object detection accuracy by using attribute descriptions above, this application breaks through the limitation of traditional methods that only rely on visual features. It uses a pre-trained language encoder to extract class and attribute embeddings, and injects attribute knowledge into class embeddings through a bidirectional multi-head attention layer, enriching the network's cognitive dimension of fine-grained objects, enabling the model to identify objects more accurately based on attribute characteristics like human experts. Secondly, a deep fusion mechanism of vision and language is constructed. The language encoder is used to extract attribute annotation embedding features, which are fused with class embedding features through bidirectional attention, and then the similarity between visual region features and class embedding features is calculated in the detection head to improve the fine-grained detection accuracy. At the same time, an innovative masked attribute reconstruction pre-training method enables the language encoder to understand attribute descriptions in combination with visual instances, and the pre-trained parameters play a stable role in the detection network, ensuring the reliability of the utilization of attribute knowledge. Finally, by constructing multi-component collaborative operations, using visual encoders, language encoders, bidirectional multi-head attention layers, etc., to ensure that attribute descriptions are deeply integrated into the entire detection process, comprehensively improving the model's detection ability for fine-grained objects, which has important application value in fields such as urban planning.
[0030] In one embodiment, the pre-training process of the language encoder includes: Obtain the input remote sensing image, the calibration box and attribute text of each target in it; randomly mask the attribute text to obtain masked attributes and original attributes; use the visual encoder to extract the remote sensing image features, and use the language encoder to extract the masked attribute embedding features; extract the instance features of all targets from the remote sensing image features through the calibration box according to the instance feature extractor; Map the masked attribute embedding features to the number of channels of the visual instance features through a linear layer to obtain the mapped features, adjust the dimension of the instance features, and then splice the mapped features and the adjusted instance features in the second dimension. Input the spliced features into the visual-language joint encoder for fusion to obtain the fused features.
[0031] Input the fused features into the prediction head to generate attribute predictions, calculate the cross-entropy loss between the predicted attributes and the original attributes, optimize the network, and obtain the trained language encoder.
[0032] In one embodiment, randomly mask the attribute text to obtain the masked attributes and the original attributes, including: Randomly select 15% of the tokens from the attribute text, randomly select 10% of the 15% tokens and replace them with a random token; randomly select 80% of the 15% tokens and replace them with [MASK]; keep the remaining 10% of the tokens unchanged to obtain the masked attributes and the original attributes.
[0033] In one embodiment, map the masked attribute embedding features to the number of visual instance feature channels through a linear layer to obtain the mapped features, adjust the dimensions of the instance features, and concatenate the mapped features and the adjusted instance features in the second dimension, including: The masked attribute embedding features Are mapped to the number of visual instance feature channels through a linear layer to obtain the mapped features , adjust the dimensions of the instance features to obtain the adjusted instance features ; concatenate And In the second dimension to obtain the concatenated features; where Represents the number of targets, Represents the maximum number of tokens, Represents the hidden dimension, Is the number of instance feature channels.
[0034] In one embodiment, the category text is a string obtained by concatenating all the names of the categories to be detected with "."; the attribute text is the attribute descriptions of all the categories, and each category attribute is a string.
[0035] In one embodiment, input the category embedding features and the corresponding attribute embedding features of each category into a bidirectional multi-head attention layer for enhancement to obtain the enhanced category embedding features and attribute embedding features, including: Input the category embedding features and the corresponding attribute embedding features of each category into a bidirectional multi-head attention layer for enhancement, and the enhanced category embedding features obtained are: ; ; Where And Respectively represent the bidirectional multi-head attention layer and layer normalization, Is the category Of the category embedding features, is the category 's attribute embedding feature, is the number of lemmas of the category , is the network learnable parameter, is the enhanced category embedding feature.
[0036] In one embodiment, the vision-language depth enhancement encoder includes 6 network blocks, each network block consists of a cross-modal multi-head attention module, a DyHead module, and a BERT layer. The cross-modal multi-head attention module is used to achieve cross-modal feature interaction, the DyHead module is used to encode visual features and extract image region features, and the BERT layer is used to encode language features.
[0037] In one embodiment, the detection head includes two branches: localization and classification. The localization branch consists of one layer convolutional layer. The classification branch involves calculating the similarity between the category embedding feature and the image region feature. Input the enhanced image feature and the secondarily enhanced category embedding feature into the detection head to obtain the detection result, including: Enhanced remote sensing image feature Obtain the box prediction through the localization branch, and calculate the similarity between the secondarily enhanced category embedding feature and the enhanced remote sensing image feature in the classification branch: , the similarity is the classification score for each category, and the category with the highest similarity is used as the classification result.
[0038] In a specific embodiment, during the experiment, two parts of the aircraft data in the remote sensing aircraft detection dataset MAR20 and the remote sensing fine-grained object detection dataset FAIR1M are used to construct a remote sensing aircraft detection dataset containing various specific models of civilian aircraft. This dataset contains 30 types of aircraft, and the instance positions are annotated with rotated bounding boxes; each type of aircraft is annotated with fine-grained attributes, including wing settings, tail shapes, fuselage lengths, etc. The analysis metrics used for the object detection method are: mean Average Precision (mAP), with an IoU threshold of 0.5; the analysis metrics used for the mask attribute reconstruction method are: Accuracy. Mask attribute reconstruction pre-training is performed on the large remote sensing aircraft detection dataset RarePlanes, and the pre-training results are shown in Table 1. It can be seen that the pre-trained model can accurately predict the attributes by combining aircraft instances. The experimental results on the constructed 30-type aircraft dataset are shown in Table 2. It can be seen that after fusing attribute knowledge, the fine-grained detection accuracy of the network can be improved, and after loading the mask attribute reconstruction pre-trained language encoder, the detection accuracy further increases.
[0039] Table 1
[0040] Table 2
[0041] It should be understood that although Figure 1 the steps in the flowchart are shown sequentially according to the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless otherwise clearly stated in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 1 at least a part of the steps in can include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.
[0042] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0043] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A method for improving the accuracy of fine-grained remote sensing target detection by using attribute description, characterized in that The method includes: Obtaining an input remote sensing image, category text, and attribute text; constructing a remote sensing object detection model; the remote sensing object detection model includes a visual encoder, a pre-trained language encoder, a visual-language depth enhancement encoder, a bidirectional multi-head attention layer, and a detection head; Extracting remote sensing image features of the input remote sensing image according to the visual encoder; extracting category embedding features and attribute embedding features of the category text and the attribute text by using the language encoder; Inputting the category embedding features of each category and the corresponding attribute embedding features into the bidirectional multi-head attention layer for enhancement to obtain enhanced category embedding features; Inputting the remote sensing image features and the enhanced category embedding features into the visual-language depth enhancement encoder to obtain enhanced image features and secondarily enhanced category embedding features; Inputting the enhanced image features and the secondarily enhanced category embedding features into the detection head to obtain a detection result.
2. The method according to claim 1, wherein The pre-training process of the language encoder includes: Obtaining an input remote sensing image and the calibration box and attribute text of each target therein; randomly masking the attribute text to obtain masked attributes and original attributes; extracting remote sensing image features by using the visual encoder, and extracting masked attribute embedding features by using the language encoder; extracting instance features of all targets from the remote sensing image features through the calibration box according to the instance feature extractor; Mapping the masked attribute embedding features to the number of channels of the visual instance features through a linear layer to obtain the mapped features, adjusting the dimension of the instance features, then concatenating the mapped features and the adjusted instance features in the second dimension, and inputting the concatenated features into the visual-language joint encoder for fusion to obtain fused features; Inputting the fused features into the prediction head to generate attribute predictions, calculating the cross-entropy loss between the predicted attributes and the original attributes, and optimizing the network to obtain a trained language encoder.
3. The method according to claim 2, wherein Randomly masking the attribute text to obtain masked attributes and original attributes, including: Randomly selecting 15% of the tokens from the attribute text, randomly selecting 10% of the 15% tokens and replacing them with a random token; randomly selecting 80% of the 15% tokens and replacing them with [MASK]; keeping the remaining 10% of the tokens unchanged to obtain masked attributes and original attributes.
4. The method according to claim 2, characterized in that, Mapping the masked attribute embedding features to the number of channels of the visual instance features through a linear layer to obtain the mapped features, adjusting the dimension of the instance features, then concatenating the mapped features and the adjusted instance features in the second dimension, including: Embed the mask attribute into the feature Map it to the number of visual instance feature channels through a linear layer to obtain the mapped feature , adjust the dimension of the instance feature to obtain the adjusted instance feature ; On the second dimension and are concatenated to obtain the concatenated feature; among them, represents the number of targets, represents the maximum number of tokens, represents the hidden dimension, is the number of instance feature channels.
5. The method according to claim 1, wherein The category text is a string obtained by concatenating all the category names to be detected with "."; the attribute text is the attribute descriptions of all categories, and each category attribute is a string.
6. The method according to claim 1, characterized in that, Inputting the category embedding features of each category and the corresponding attribute embedding features into the bidirectional multi-head attention layer for enhancement to obtain enhanced category embedding features and attribute embedding features, including: Inputting the category embedding features of each category and the corresponding attribute embedding features into the bidirectional multi-head attention layer for enhancement, and the enhanced category embedding features obtained are: Among them, and represent a bidirectional multi-head attention layer and layer normalization respectively, is the category 's category embedding feature, is the category 's attribute embedding feature, is the number of tokens of the category , are network learnable parameters, is the enhanced category embedding feature.
7. The method according to claim 1, wherein The visual language depth enhancement encoder includes 6 network blocks, each of which consists of a cross-modal multi-head attention module, a DyHead module, and a BERT layer. The cross-modal multi-head attention module is used to achieve cross-modal feature interaction, the DyHead module is used to encode visual features and extract image region features, and the BERT layer is used to encode language features.
8. The method according to claim 1, wherein the detection head comprises two branches, namely a positioning branch and a classification branch. The positioning branch consists of one layer of convolutional layer. The classification branch involves calculating the similarity between the class embedding features and the image region features. The enhanced image features and the secondarily enhanced class embedding features are input into the detection head to obtain detection results, including: The enhanced image features The box prediction obtained through the localization branch and the class embedding features after secondary enhancement Calculate the similarity with the enhanced remote sensing image features in the classification branch: , the similarity is the classification score for each category, and the category with the highest similarity is used as the classification result.
Citation Information
Patent Citations
Dependency attribute enhanced text image search pedestrian re-identification method
CN117828121A
Image recognition method and device, equipment, storage medium and program product
CN118658035A
Remote sensing image change detection method based on language guidance
CN119169449A
Remote sensing image scene classification method based on attribute guidance
CN119723183A
Instance level scene recognition with a vision language model
US11978271B1