Remote sensing image fine-grained target recognition method based on text guidance
By using a generative large language model in remote sensing image recognition to generate category description text and adjust image features in combination with interactive modules, the problem of fine-grained target recognition of remote sensing images is solved, and the recognition accuracy is significantly improved.
Patent Information
- Application Number
- CN202510223380.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-20
AI Technical Summary
The prior art is difficult to effectively perform fine-grained target recognition in remote sensing images, especially when the target background is complex and the target diversity of fine-grained multi-category targets.
Design a text-guided method, generates category description text through a generative large language model (LLM), and encodes the image and text using an encoder. The design interactive module calculates the importance score of image features relative to the description text, and adjusts image features for detection as attention weight.
The design of the category description text and interactive module generated by LLM significantly improves the accuracy of fine-grained target recognition of remote sensing images, and can better handle complex background and diversity goals.
Smart Images

Figure CN120182573A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the cross - technical field of computer vision and natural language processing, and particularly relates to a method for fine - grained object recognition in remote sensing images based on text guidance. Background Art
[0002] Quickly and accurately locating and identifying objects of interest from satellite images covering millions of square kilometers is the core problem in the field of intelligent interpretation of remote sensing data. With the continuous development of satellite imaging technology and computer vision technology, the high - resolution remote sensing image interpretation algorithm has gradually evolved from object detection to fine - grained recognition.
[0003] Object detection and fine - grained recognition of remote sensing images are key image - processing tasks in the field of computer vision. Its purpose is to process and analyze remote sensing images through algorithmic techniques to determine the spatial position and category of objects, and further subdivide them into specific sub - categories, providing an effective means for the application of remote sensing technology in fields such as urban planning, agricultural production, and environmental monitoring. The difficulty of fine - grained object recognition lies in that different fine - grained categories belonging to the same coarse - grained category usually have the characteristics of high intra - class variability and low inter - class distinguishability. Therefore, it is necessary to extract the detailed differences of fine - grained categories as much as possible for object recognition. However, different from the diverse perspectives of natural images, the objects in remote sensing images only have a top - down view, with a single perspective, and a large number of differential features that can be effectively used for fine - grained classification are blocked. This makes it difficult to recognize fine - grained objects in remote sensing images.
[0004] Different from image features, text descriptions can well summarize the features of specific categories. Such descriptive texts are independent of the number of instances in the dataset and the imaging angle, and are an effective supplement to the image features of fine - grained categories. Text descriptions can provide information on various aspects such as the shape, color, texture, and function of the object, helping the model better understand the subtle differences of the object.
[0005] In previous methods, category description texts were usually manually written by several experts. However, this method has obvious limitations. Firstly, there is a domain limitation. Experts are usually only familiar with the objects of a specific coarse - grained category and cannot cover all fine - grained categories. For example, an expert in the aviation field may be very familiar with commercial airliners, but not with the detailed features of military aircraft or ships. Secondly, it is inefficient. The process of manually writing text descriptions is time - consuming and costly, especially when dealing with large - scale datasets, and it is difficult to meet the requirement of quickly and efficiently generating descriptive texts. Finally, there are large differences in expression styles. The expression styles of different experts may vary greatly and be subjective, resulting in a lack of consistency in the generated text descriptions and affecting the training effect of the model. These limitations restrict the large - scale application of traditional methods, especially in the case of dealing with a large number of remote sensing images.
[0006] In recent years, the research on large language models (LLMs) has developed rapidly. LLMs are trained on massive amounts of Internet data and can effectively summarize visual features across categories. Models such as ChatGPT and Qwen have demonstrated the powerful generation capabilities of LLMs. Using LLMs can also generate category description texts. Compared with manually written texts, the advantages of using LLMs lie first in their cross-domain applicability. LLMs are not restricted by specific domains and can generate detailed description texts for different types of fine-grained remote sensing targets, covering multiple fields such as aviation, transportation, agriculture, and architecture. Additionally, through carefully designed prompt templates, LLMs can quickly generate text descriptions with the same format, with high generation efficiency, and avoid the expression differences in manual writing, improving the stability of model training. According to specific task requirements, users can flexibly adjust the prompts to generate description texts suitable for different application scenarios, with high flexibility. Summary of the Invention
[0007] To overcome the deficiencies of the above-mentioned prior art, the purpose of the present invention is to provide a text-guided fine-grained target recognition method for remote sensing images. First, specific prompt templates are designed, and category description texts that meet the requirements are generated through LLMs. Then, the encoder is used to encode the image and text respectively, and an interaction module is designed to calculate the importance score of each image feature relative to the description text, which is used as the attention weight to adjust the image features for detection. Finally, the model is trained on the target dataset. Category description texts are generated through LLMs, and an interaction module is designed to enhance the image features using the category descriptions, improving the fine-grained recognition performance of the detector. The training set and test set come from the mainstream remote sensing image fine-grained target recognition dataset FAIR1M. FAIR1M subdivides the targets (including airplanes, ships, vehicles, stadiums, roads, etc.) into categories, including 37 fine-grained categories, with more than 1 million finely labeled and multi-angled distributed targets. The biggest challenge of this dataset is the complexity of the target background and the diversity of fine-grained multi-class targets. Improve the accuracy of fine-grained target recognition on remote sensing images.
[0008] To achieve the above purpose, the technical solution adopted by the present invention is:
[0009] A text-guided fine-grained target recognition method for remote sensing images, characterized by comprising the following steps:
[0010] Step 1, design a prompt template and use a generative large language model to obtain description texts for each category;
[0011] 1) For the prompt template, its purpose is to obtain feature descriptions for distinguishing different fine-grained categories. First, analyze all category labels in the dataset and conduct a preliminary study on each category. Divide all fine-grained categories into several coarse-grained categories, understand the general features, common locations, and any unique attributes that can be used to distinguish this category from others of each coarse-grained category. Then, define a set of key attributes for each coarse-grained category. These attributes include color, shape, size, texture, background, etc., especially those features that help distinguish fine-grained differences. Finally, based on the above analysis, create a general prompt framework and customize it for each coarse-grained category according to the key attributes to ensure that the prompt clearly indicates the type of information expected from the large language model.
[0012] 2) For the descriptive text, generate prompts for each category according to the prompt template and category labels, input the prompts into the large language model to obtain the descriptive text. There are no special restrictions on the large language model, and it can be flexibly selected according to the data type and language preference. Further improve and optimize the prompts according to the generation effect and the requirements for text length and text format. Save the descriptive text as a.txt file.
[0013] Step 2, use the encoder to encode the image and text respectively.
[0014] 1) For the image, select a CNN-based detector according to the task objective, such as Oriented RCNN, PCLDet. Keep the original architecture of the detector unchanged, and input the picture to be processed sequentially through the backbone network, FPN, RoI and other detector components to obtain image features. ; When used for the regression branch, it remains unchanged and is directly fed into the regression head to predict its spatial position. When used for the classification branch, it needs to interact with the text features first, and the obtained is fed into the classification head to predict the target category.
[0015] 2) For the text, use a pre-trained language model (such as BERT, RoBERTa) to encode the descriptive text. Freeze the text encoder during the encoding process, and record the encoded descriptive text as , save it as a.pt file for subsequent training.
[0016] Step 3, design an interaction module to calculate the importance score of each image feature relative to the descriptive text, and use it as the attention weight to adjust the image features for object recognition.
[0017] 1) Flatten the image features obtained after being processed by the backbone network, FPN, RoI and other detector components to obtain ;
[0018] 2) Process the image features by means of global average pooling, projection, and stacking of 3×3 convolutions to obtain different image feature representations mapped to the same feature space, denoted as , and . The tensor shapes of all three are . The calculation formula is as follows:
[0019]
[0020]
[0021]
[0022] Among them, MultiConv represents the stacking of N 3×3 convolutions. Denote the input of each stacking unit as , then the output , represents the convolution operation with a convolution kernel size of 3; the stacking number N is related to the shape of the image feature ; for a feature map with a shape of , after passing through each stacking unit, its feature map shape becomes , and after passing through N stacking units, it becomes a feature map of ; if H and W are inconsistent or not odd, padding calculation needs to be performed first;
[0023] 3) Through a trainable projection layer, map the text feature to the same feature space as the image features , and to obtain ;
[0024] 4) Use cosine similarity to calculate the similarity between the text feature and the image features , , . Weight the similarity scores through 3 trainable weight parameters , and , and use Softmax to normalize the weighted similarity scores to obtain the final attention weights ; the calculation formula is as follows:
[0025]
[0026] 5) Use the attention weights to process the text feature , and flatten the image feature Concatenate with the processed text features and map them back to the same feature space as through a fully connected layer to obtain ; The calculation method is as follows:
[0027]
[0028] where represents the fully connected layer, represents the concatenation operation; the calculated is used as the image feature after image-text interaction and sent to the classification head for class prediction;
[0029] Step 4, train using the improved detector on the target dataset;
[0030] Based on Oriented RCNN implementation, use the pre-trained ResNet 50 as the backbone network, BERT as the text encoder, and GPT 4 for the generation of descriptive text; use the SGD optimizer for training, with a momentum of 0.9 and a weight decay of 0.0001; GPT 4 and the text encoder can be performed before model training; after training is completed, only the encoded file of the descriptive text and the detector model with the added interaction module are used in the inference stage.
[0031] The beneficial effects of the present invention are:
[0032] First, design a specific prompt template to generate the required category description text through the LLM. Then, encode the image and text respectively through the encoder, design an interaction module, calculate the importance score of each image feature relative to the descriptive text, and use it as the attention weight to adjust the image features for detection. Finally, train the model on the target dataset. Generate the category description text through the LLM, design an interaction module, and use the category description to enhance the image features, improving the fine-grained recognition performance of the detector. The training set and test set come from the mainstream remote sensing image fine-grained target recognition dataset FAIR1M. FAIR1M subdivides the targets (including airplanes, ships, vehicles, stadiums, roads, etc.) into categories, including 37 fine-grained categories, with more than 1 million finely annotated and multi-angled distributed targets. The biggest challenge of this dataset is the complexity of the target background and the diversity of fine-grained multi-class targets. Improve the fine-grained target recognition accuracy on remote sensing images. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 is the flow schematic diagram of the present invention;
[0034] Figure 2 is the flow schematic diagram of the category description text generation;
[0035] Figure 3 Schematic diagram of the prompting template construction process;
[0036] Figure 4 Schematic diagram of the interaction module. Detailed implementation manners
[0037] The present invention will be further described below in conjunction with the accompanying drawings, but the present invention is not limited to the following embodiments.
[0038] As shown in Figure 1 , 2 , 3, and 4. A text-guided remote sensing image fine-grained target recognition method, characterized by including the following steps:
[0039] Step 1, design a prompting template, and use a generative large language model to obtain descriptive texts for each category;
[0040] Step 2, use an encoder to encode the image and the text respectively;
[0041] Step 3, design an interaction module, calculate the importance score of each image feature relative to the descriptive text, and use it as an attention weight to adjust the image features for target recognition;
[0042] Step 4, use the improved detector to train on the target dataset
[0043] Furthermore, for Step 1, the design of the prompting template and the use of the generative large language model to obtain descriptive texts for each category specifically include:
[0044] 1) For the prompting template, its purpose is to obtain feature descriptions for distinguishing different fine-grained categories. First, analyze all category labels of the dataset and conduct a preliminary study on each category. Divide all fine-grained categories into several coarse-grained categories, understand the general features, common locations, and any unique attributes that can be used to distinguish this category from other categories of each coarse-grained category. Then, define a set of key attributes for each coarse-grained category. These attributes include color, shape, size, texture, background, etc., especially those features that help distinguish fine-grained differences. Finally, based on the above analysis, create a general prompting framework and customize and adjust it for each coarse-grained category according to the key attributes to ensure that the prompting clearly indicates the type of information that is expected to be provided by the large language model.
[0045] 2) For the descriptive text, generate a prompt for each category according to the prompting template and the category label, input the prompt into the large language model to obtain the descriptive text. There is no special limitation on the large language model, and it can be flexibly selected according to the data type and language preference. Further improve and optimize the prompt according to the generation effect and the requirements for the length and format of the text. Save the descriptive text as a.txt file.
[0046] Further, in step 2, the use of encoders to encode images and texts respectively specifically includes:
[0047] 1) For images, select a CNN - type detector according to the task objective, such as Oriented RCNN, PCLDet. Keep the original architecture of the detector unchanged. The input image is sequentially processed by detector components such as the backbone network, FPN, and RoI to obtain image features . When used for the regression branch, it remains unchanged and is directly fed into the regression head to predict its spatial position. When used for the classification branch, it needs to interact with the text features first, and the obtained is fed into the classification head to predict the target category.
[0048] 2) For texts, use a pre - trained language model (such as BERT, RoBERTa) to encode the descriptive text. Freeze the text encoder during the encoding process. The encoded descriptive text is denoted as and saved as a.pt file for subsequent training.
[0049] Further, in step 3, the involved design of the interaction module calculates the importance score of each image feature relative to the descriptive text, and adjusts the image features as attention weights for object recognition. Specifically, it includes:
[0050] 1) Flatten the image features obtained after being processed by detector components such as the backbone network, FPN, and RoI to obtain .
[0051] 2) Process the image features through global average pooling, projection, and stacking of 3×3 convolutions to obtain different image feature representations mapped to the same feature space, denoted as , and , and the tensor shapes of all three are . The calculation formula is as follows:
[0052]
[0053]
[0054]
[0055] Among them, MultiConv represents the stacking of N 3×3 convolutions. Denote the input of each stacking unit as , then the output , represents the convolution operation with a convolution kernel size of 3. The stacking number N is related to the image features is related to the shape. For the feature map with the shape of , after passing through each stacking unit, the shape of its feature map becomes , and after passing through N stacking units, it becomes the feature map of . If H and W are inconsistent or not odd, padding calculation needs to be performed first.
[0056] 3) Through a trainable projection layer, project the text feature into the same feature space as the image features , and to obtain .
[0057] 4) Use cosine similarity to calculate the similarity between the text feature and the image features , , . Weight the similarity scores through three trainable weight parameters , and , and use Softmax to normalize the weighted similarity scores to obtain the final attention weights . The calculation formula is as follows:
[0058]
[0059] 5) Use the attention weights to process the text feature , splice the flattened image feature and the processed text feature, and map it back to the same feature space as through a fully connected layer to obtain . The calculation method is as follows:
[0060]
[0061] Among them, represents the fully connected layer, and represents the splicing operation. The calculated is used as the image feature after text-image interaction and sent to the classification head for class prediction.
[0062] Furthermore, in step 4, the model training adopts a multi-stage training strategy, which mainly includes:
[0063] This method can be implemented based on Oriented RCNN, using the pre-trained ResNet 50 as the backbone network, BERT as the text encoder, and GPT 4 for the generation of descriptive text. Training is carried out using the SGD optimizer with a momentum of 0.9 and a weight decay of 0.0001. GPT 4 and the text encoder can be performed before model training. After training is completed, only the encoded file of the descriptive text is used in the inference stage. and the detector model with an added interaction module.
Claims
1. A text-guided remote sensing image fine-grained target recognition method, characterized in that: The following steps are involved: Step 1: Design a prompt word template and use a generative large language model to obtain description text for each category; 1) First, analyze all the category labels of the dataset and conduct preliminary research on each category; Divide all fine-grained categories into several coarse-grained categories, understand the general characteristics of each coarse-grained category, common locations, and any unique attributes that can be used to distinguish this category from other categories; then, define a set of key attributes for each coarse-grained category; finally, based on the above analysis, create a general prompt word framework and customize it for each coarse-grained category based on the key attributes to ensure that the prompt word clearly indicates the type of information you want the large language model to provide; 2) For the description text, generate prompt words for each category based on the prompt word template and category label, input the prompt words into the large language model to obtain the description text; the large language model has no special restrictions and can be flexibly selected according to the data type and language preference; further improve and optimize the prompt words according to the generation effect and the requirements for the text length and text format; save the description text as a .txt file; Step 2, use the encoder to encode the image and text separately; 1) For images, select CNN detectors such as Oriented RCNN and PCLDet according to the task objectives; keep the original architecture of the detector unchanged, and process the input image through the backbone network, FPN, RoI and other detector components in turn to obtain image features ; When used for the regression branch, it remains unchanged and is directly sent to the regression head to predict its spatial position; When used for classification branches, it is necessary to interact with text features first. Send to the classification head to predict the target category; 2) For text, use a pre-trained language model (such as BERT, RoBERTa) to encode the description text. During the encoding process, freeze the text encoder. The encoded description text is recorded as , save as a .pt file for subsequent training; Step 3: Design an interactive module to calculate the importance score of each image feature relative to the description text, and use it as an attention weight to adjust the image features for target recognition. 1) The image features obtained after being processed by the backbone network, FPN, RoI and other detector components Flatten to get ; 2) The image features are processed by global average pooling, projection, and 3×3 convolution stacking to obtain different image feature representations mapped to the same feature space, denoted as , and , the tensor shapes of the three are , the calculation formula is as follows: Among them, MultiConv represents the stacking of N 3×3 convolutions, and the input of each stacking unit is recorded as , then the output , Represents a convolution operation with a convolution kernel size of 3; the number of stacks N and image features For the shape of The feature map of each stacking unit changes into , after N stacking units, becomes feature map; if H and W are inconsistent or not an odd number, padding calculation needs to be performed first; 3) Through a trainable projection layer, the text features Mapping to image features , and In the same feature space, we get ; 4) Use cosine similarity to calculate text features and image features , , The similarity between them is determined by three trainable weight parameters , and Weight the similarity scores and use Softmax to normalize the weighted similarity scores to get the final attention weights ; The calculation formula is as follows: 5) Using attention weights Text features Processing, flattening the image features It is concatenated with the processed text features and mapped back to The same feature space, ; The calculation method is as follows: in, represents the fully connected layer, Represents a concatenation operation; the calculated As the image features after the image-text interaction, it is sent to the classification head for category prediction; Step 4: Use the improved detector to train on the target dataset; Based on Oriented RCNN implementation, using pre-trained ResNet 50 as the backbone network, BERT as the text encoder, and GPT 4 for description text generation; using SGD optimizer for training, with momentum of 0.9 and weight decay of 0.0001; GPT 4 and text encoder can be performed before model training; after training, only the encoded file of description text is used in the inference phase And the detector model after adding the interaction module.
2. According to claim 1, a remote sensing image fine-grained target recognition method based on text guidance is characterized in that: The key attributes mentioned include color, shape, size, texture, background features, especially features that help distinguish fine-grained differences.
Citation Information
Cited By
Zero sample oncomelania identification method and system based on structured knowledge base and attention guidance
CN122090186A