Method for realizing automatic detection of remote sensing image based on RSODTR model

Through the automatic remote sensing image detection method based on the RSODTR model, the problem of poor results in the prior art when detecting small targets and multi-scale targets is solved, real-time and efficient detection results are achieved, and computing costs and resource utilization are reduced.

CN119964009APending Publication Date: 2025-05-09LANZHOU JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510373940.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

Existing remote sensing image automatic detection technology is difficult to effectively extract image information of low resolution and noise characteristics, especially when detecting small targets and multi-scale targets, and there are problems such as high computing costs and large resource utilization.

Method used

Using an RSODTR model-based approach, real-time object detection is achieved by introducing efficient hybrid encoder and dynamically adjusting the number of decoder layers, and reducing computational costs with smaller feature maps and fewer attention heads.

Benefits of technology

Real-time and efficient automatic detection of remote sensing images are realized, and can accurately detect multi-scale targets, reducing calculation costs and resource usage, while improving detection accuracy and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964009A_ABST
    Figure CN119964009A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of remote sensing image automatic detection, and discloses a method for realizing remote sensing image automatic detection based on an RSODTR model. Comprising the following steps: S1, starting PyCharm software, selecting a data set, and taking an image in the data set as an input image; s2, automatically detecting a target ground feature in the input image by using an RSODTR model, and automatically calculating the speed of the input image during automatic detection and the precision of the type of the ground feature so as to obtain the position, the range and the type of the target ground feature; s3, storing the position, range and category of the target ground feature; remote sensing image detection software can update the latest algorithm in time to carry out ground object automatic detection, and the detection efficiency and accuracy are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of automatic detection of remote sensing images, and more specifically, to a method for realizing automatic detection of remote sensing images based on a RSODTR model. Background Art

[0002] With the rapid innovation and development of earth observation technology, the resolution of remote sensing images is constantly improving, and the information about objects in the images is becoming more and more abundant. As an important research topic in the automatic interpretation of remote sensing images, Remote_sensing_object_detection (RSOD) can predict the location and category of objects of interest in remote sensing images through specific algorithms, thereby accurately detecting multiple objects in the images. Therefore, RSOD is widely used in many fields.

[0003] In recent years, some scholars have applied deep learning methods to RSOD, extracting feature information of objects from remote sensing images through deep neural networks. The following defects still exist: (1) Compared with natural images, remote sensing images inherently carry fewer pixels and less obvious feature information, and have low resolution and noise characteristics. However, the pooling layer in DNN further compresses the amount of information, which leads to the fact that feature information about small targets is easily ignored or lost when using DNN to extract image features. (2) Remote sensing images usually contain objects of different scales, and the aspect ratios of different objects are also different. Unlike image classification tasks, target detection usually requires predicting objects of different scales, and the default feature extraction method of DNN may not be able to effectively extract feature information of objects at different scales, resulting in poor detection results. (3) Objects that are easy to distinguish in natural images may show high inter-class similarity in remote sensing images. For example, the features of roads and bridges in remote sensing images are very similar, and their similar background information leads to small differences in inter-class features. (4) Compared with natural images, the background complexity of remote sensing images is higher. Objects of the same type may be in different backgrounds and change with the seasons and sensor viewing angles. In addition, the proportion of background images in remote sensing images is often much larger than that of foreground images, resulting in large differences in features within the same type of objects. Therefore, it is quite difficult to distinguish object categories and background information under different backgrounds and scales. The mixing of heterogeneous instance features with irrelevant noise and high response in background areas may lead to false detections and missed detections.

[0004] At present, most algorithms with good detection effects in RSOD tasks use the CNN framework. However, CNN-based object detection models easily lead to an increase in model size and inevitably produce spatial artifacts, making feature maps more susceptible to spatial bias and making it difficult for the network to detect small targets in images. In addition, CNNs usually have high requirements for memory and computing power, which makes them difficult to deploy in edge computing devices with limited computing power and power.

[0005] The existing DETR is a milestone innovation of the transformer architecture in the field of target detection. It combines the encoder-decoder architecture of CNN and Transformer. The encoder encodes the pixel clusters in the sliding window and extracts contextual information from them using the self-attention mechanism based on the transformer architecture. The decoder uses the cross-attention mechanism to generate the detection box of the target to be detected based on the contextual information obtained by the encoder. However, the training cost of DETR is high. Compared with the CNN-based detector, it takes longer to train to converge. In addition, DETR cannot detect the feature information of multi-scale targets and tends to prioritize long-range semantic information while ignoring important local features, which leads to poor performance of DETR when detecting small targets.

[0006] In order to solve the above problems, Real-Time_Detection_Transformer (RT-DETR) came into being. RT-DETR is an end-to-end real-time target detector that can realize real-time target detection and solve the problem of high training cost of DETR. RT-DETR introduces an efficient hybrid encoder that dynamically adjusts the inference speed by adjusting the number of layers of the decoder to adapt to different application scenarios without retraining. Compared with DETR, RT-DETR uses smaller feature maps to reduce computational costs and uses fewer attention heads to reduce the number of parameters in the model. RT-DETR achieves a balance between target detection performance and efficiency, has high resource utilization, does not require post-processing optimization of detection results such as Confidence_threshold and NMS, and gives full play to the advantages of end-to-end detection. However, there are currently few target detection methods that apply RT-DETR to RSOD.

[0007] In view of this, the present invention proposes a method for automatic detection of remote sensing images based on the RSODTR model to solve the above problems. Summary of the invention

[0008] In order to overcome the above defects of the prior art and to achieve the above objectives, the present invention provides the following technical solution: a method for realizing automatic detection of remote sensing images based on the RSODTR model, comprising the following steps:

[0009] S1: Start PyCharm software, select the dataset and use the image in the dataset as the input image;

[0010] S2: Use the RSODTR model to automatically detect the target object in the input image, and automatically calculate the rate of automatic detection of the input image and the accuracy of the object type, and then obtain the location, range and category of the target object;

[0011] S3: Save the location, range and category of the target object;

[0012] S4: The RSODTR model identifies and displays the target objects in the input image after automatic detection.

[0013] Further, the starting of the PyCharm software, selecting a data set and using the image in the data set as data of the input image includes:

[0014] After starting PyCharm, create a new project and configure the virtual environment of Python software;

[0015] Use PyCharm software to select a dataset containing several remote sensing images, use the images in the dataset as input images and divide them into training set, validation set and test set in a 1:1:1 ratio. The images in the training set and validation set are annotated with corresponding labels, while the images in the test set are not annotated with labels.

[0016] The training process of the RSODTR model includes: first, the RSODTR model is trained with a training set and a validation set. After the training is completed, the images in the test set are input into the RSODTR model for detection.

[0017] Select a data set through PyCharm software, and use the images in the data set as input images, where the data set includes n remote sensing images.

[0018] Furthermore, the method of automatically detecting the target object in the input image by using the RSODTR model, and automatically calculating the rate and accuracy of the object type when the input image is automatically detected, and then obtaining the position, range and category of the target object includes:

[0019] The RSODTR model is built and trained through the backbone network, encoder, decoder and detection head, and the input image can be input into the trained RSODTR model;

[0020] The trained RSODTR model extracts the feature information of the ground objects from the input image through the fast basic module and residual connection technology of the backbone network, and then inputs the feature maps of different scales into the encoder, where the encoder is composed of CGA, CCFF and SSFF;

[0021] The trained RSODTR model further encodes and contextualizes the output features of the last layer of the backbone through CGA, and fuses feature maps of different sizes through CCFF and SSFF to convert them into a series of feature encodings as the initial target query of the decoder;

[0022] The trained RSODTR model introduces a sparse query mechanism and learnable position encoding through the decoder to optimize the target query, thereby obtaining the optimized feature encoding;

[0023] The trained RSODTR model converts the decoder-optimized feature encoding into the predicted target object category and predicted target object location through the classification head and regression head of the detection head. The predicted target object category and predicted target object location are the detection results of the target object. The detection head includes a classification head and a regression head. The classification head is used to predict the category of the target object, and the regression head is used to predict the location of the target object.

[0024] During the detection process, the detection rate of the input image and the accuracy of the object type are automatically calculated;

[0025] Finally, the trained RSODTR model outputs the location, range, and category of the target object.

[0026] Furthermore, when the trained RSODTR model automatically detects the input image, it uses non-maximum suppression to filter repeated bounding boxes in the input image.

[0027] Furthermore, the RSODTR model is trained for 150 batches, each batch size is 32 images, the image size is 800×800 pixels, the AdamW optimizer is used, and the learning rate is set to 0.0003. During the training process, the RSODTR model has a 50% probability of performing up-down or left-right flipping data enhancement on the input data.

[0028] Furthermore, the method of saving the location, range and category of the target object includes:

[0029] Output data structure that defines the location, range, and category of target features;

[0030] The location, range and category of the target feature are saved according to the defined output data structure.

[0031] Furthermore, the RSODTR model identifies and displays the target object in the input image after automatic detection, including:

[0032] Use OpenCV to draw detection boxes and labels to identify target objects in the automatically detected input image;

[0033] The input image after automatic detection is displayed through the Sc iView window built into the PyCharm software.

[0034] The technical effects and advantages of the method for realizing automatic detection of remote sensing images based on RSODTR model of the present invention are as follows:

[0035] 1. Enable remote sensing image detection software to update the latest algorithm in a timely manner for automatic detection of ground objects, ensuring detection efficiency and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 A schematic diagram of a process for realizing automatic detection of remote sensing images based on the RSODTR model of the present invention;

[0037] Figure 2 It is a partial detection result diagram of the RSODTR model of the present invention on the D IOR dataset;

[0038] Figure 3 This is a partial detection result diagram of the RSODTR model of the present invention on the DOTA dataset. DETAILED DESCRIPTION

[0039] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0040] Example 1

[0041] See also Figures 1 to 3 As shown, the method for realizing automatic detection of remote sensing images based on the RSODTR model described in this embodiment includes the following steps:

[0042] S1: Start PyCharm software, select the dataset and use the image in the dataset as the input image;

[0043] S2: Use the RSODTR model to automatically detect the target object in the input image, and automatically calculate the rate of automatic detection of the input image and the accuracy of the object type, and then obtain the location, range and category of the target object;

[0044] S3: Save the location, range and category of the target object;

[0045] S4: The RSODTR model identifies and displays the target objects in the input image after automatic detection.

[0046] Further, the starting of the PyCharm software, selecting a data set and using the image in the data set as data of the input image includes:

[0047] After starting PyCharm, create a new project and configure the virtual environment of Python software;

[0048] Use PyCharm software to select a dataset containing several remote sensing images, use the images in the dataset as input images and divide them into training set, validation set and test set in a 1:1:1 ratio. The images in the training set and validation set are annotated with corresponding labels, while the images in the test set are not annotated with labels.

[0049] The training process of the RSODTR model includes: first, the RSODTR model is trained with a training set and a validation set. After the training is completed, the images in the test set are input into the RSODTR model for detection.

[0050] Specifically, the python script can edit, update and apply the RSODTR model, so that the remote sensing image detection software can timely update the latest algorithm for automatic detection of ground objects to ensure detection efficiency and accuracy.

[0051] Furthermore, the method of automatically detecting the target object in the input image by using the RSODTR model, and automatically calculating the rate and accuracy of the object type when the input image is automatically detected, and then obtaining the position, range and category of the target object includes:

[0052] The RSODTR model is built and trained through the backbone network, encoder, decoder and detection head, and the input image can be input into the trained RSODTR model;

[0053] The trained RSODTR model extracts the feature information of the ground objects from the input image through the fast basic module and residual connection technology of the backbone network, and then inputs the feature maps of different scales into the encoder. The encoder consists of CGA (Cascaded-Group-Attention) CCFF (Cross-scale-Feature-Fusion) and SSFF (Scale-Sequence-Feature-Fusion);

[0054] The trained RSODTR model further encodes and contextualizes the output features of the last layer of the backbone through CGA, and fuses feature maps of different sizes through CCFF and SSFF to convert them into a series of feature encodings as the initial target query of the decoder;

[0055] The trained RSODTR model introduces a sparse query mechanism and learnable position encoding through the decoder to optimize the target query, thereby obtaining the optimized feature encoding;

[0056] The trained RSODTR model converts the decoder-optimized feature encoding into the predicted target object category and predicted target object location through the classification head and regression head of the detection head. The predicted target object category and predicted target object location are the detection results of the target object. The detection head includes a classification head and a regression head. The classification head is used to predict the category of the target object, and the regression head is used to predict the location of the target object.

[0057] During the detection process, the detection rate of the input image (how many images are processed per second) and the accuracy of the object type (the proportion of correctly detected objects to the total number of detections) are automatically calculated;

[0058] Finally, the trained RSODTR model outputs the location, range, and category of the target object.

[0059] Specifically, the RSODTR model is applied to automatically detect and process remote sensing images. PyCharm has advanced code editing functions, such as syntax highlighting, automatic completion, and code formatting. Its smart prompts can provide suggestions for variables, functions, and modules based on the context, which helps to speed up the code writing process and reduce errors. PyCharm provides powerful code navigation and search functions, allowing developers to quickly locate and browse code. It supports operations such as jumping to function definitions, finding references, and finding specific symbols, providing a convenient code navigation experience. PyCharm integrates a comprehensive debugger that supports setting breakpoints, single-step debugging, and variable viewing, helping developers quickly locate and fix problems. In addition, PyCharm also supports unit testing, which is convenient for writing, running, and analyzing test cases. The above advantages of PyCharm serve the editing and updating of algorithms well.

[0060] Furthermore, when the trained RSODTR model automatically detects the input image, it uses non-maximum suppression to filter repeated bounding boxes in the input image.

[0061] Specifically, after obtaining multiple detection results, non-maximum suppression (NMS) is used to filter out duplicate or overlapping bounding boxes to ensure that each object is detected only once.

[0062] Furthermore, the RSODTR model is trained for 150 batches, each batch size is 32 images, the image size is 800×800 pixels, the AdamW optimizer is used, and the learning rate is set to 0.0003. During the training process, the RSODTR model has a 50% probability of performing up-down or left-right flipping data enhancement on the input data.

[0063] Specifically, the above techniques can further improve the performance of the RSODTR model.

[0064] Furthermore, the method of saving the location, range and category of the target object includes:

[0065] Output data structure that defines the location, range, and category of target features;

[0066] The location, range and category of the target feature are saved according to the defined output data structure.

[0067] Furthermore, the RSODTR model identifies and displays the target object in the input image after automatic detection, including:

[0068] Use OpenCV to draw detection boxes and labels to identify target objects in the automatically detected input image;

[0069] The input image after automatic detection is displayed through the SciView window built into the PyCharm software.

[0070] In this embodiment, the algorithm performance, including the rate and accuracy of the object type, can be intuitively demonstrated by using the test set images in the data set through PyCharm software.

[0071] It should be noted that the rate is measured by the number of frames per second (FPS), which represents how many images the model can detect per second, that is, the detection speed is evaluated based on the number of images that can be detected per second. The shorter the time, the faster the speed. The accuracy of different objects is measured by the average precision (AP), which is an indicator used to evaluate the performance of tasks such as information retrieval systems and target detection. It is the area under the PR curve. For a model, the higher the AP value, the better the detection accuracy.

[0072] For example, in Figure 2 The D IOR dataset includes 20 ground feature categories, including airplane (APL), airport (APO), baseball field (BD), basketball court (BC), bridge (BR), chimney (CH), highway service area (ESA), highway toll station (ETS), dam (Dam), golf course (GF), ground track and field (GTF), port (HA), overpass (OP), ship (SH), stadium (STA), storage tank (STO), tennis court (TC), train station (TS), vehicle (VE) and windmill (WM);

[0073] exist Figure 3In DOTA, there are 15 types of features, including airplane (PL), baseball field (BD), bridge (BR), track and field (GTF), small vehicle (SV), large vehicle (LV), ship (SH), tennis court (TC), basketball court (BC), oil storage tank (ST), football field (SBF), roundabout (RA), port (HA), swimming pool (SP), and helicopter (HC).

[0074] The speed and accuracy of feature types when using PyCharm for automatic image detection.

[0075] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed in the present invention can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0076] In the several embodiments provided by the present invention, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of the units is only one, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0077] The above description is only a specific implementation mode of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.

[0078] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for automatic detection of remote sensing images based on RSODTR model, characterized in that: The following steps are involved: S1: Start PyCharm software, select the dataset and use the image in the dataset as the input image; S2: Use the RSODTR model to automatically detect the target object in the input image, and automatically calculate the rate of automatic detection of the input image and the accuracy of the object type, and then obtain the location, range and category of the target object; S3: Save the location, range and category of the target object; S4: The RSODTR model identifies and displays the target objects in the input image after automatic detection.

2. According to claim 1, a method for realizing automatic detection of remote sensing images based on RSODTR model is characterized in that: The step of starting the PyCharm software, selecting a data set and using the image in the data set as the data of the input image includes: After starting PyCharm, create a new project and configure the virtual environment of Python software; Use PyCharm software to select a dataset containing several remote sensing images, use the images in the dataset as input images and divide them into training set, validation set and test set in a 1:1:1 ratio. The images in the training set and validation set are annotated with corresponding labels, while the images in the test set are not annotated with labels. The training process of the RSODTR model includes: first, the RSODTR model is trained with a training set and a validation set. After the training is completed, the images in the test set are input into the RSODTR model for detection.

3. The method for realizing automatic detection of remote sensing images based on RSODTR model according to claim 1, characterized in that: The method of automatically detecting the target object in the input image by using the RSODTR model, and automatically calculating the rate and accuracy of the object type when the input image is automatically detected, and then obtaining the position, range and category of the target object includes: The RSODTR model is built and trained through the backbone network, encoder, decoder and detection head, and the input image can be input into the trained RSODTR model; The trained RSODTR model extracts the feature information of the ground objects from the input image through the fast basic module and residual connection technology of the backbone network, and then inputs the feature maps of different scales into the encoder, where the encoder is composed of CGA, CCFF and SSFF; The RSODTR model further encodes and contextualizes the output features of the last layer of the backbone through CGA, and fuses feature maps of different sizes through CCFF and SSFF to convert them into a series of feature encodings as the initial target query of the decoder; The RSODTR model introduces a sparse query mechanism and learnable position encoding through the decoder to optimize the target query, thereby obtaining the optimized feature encoding; The RSODTR model converts the feature encoding optimized by the decoder into the predicted category and predicted location of the target object through the classification head and regression head of the detection head. The predicted category and predicted location of the target object are the detection results of the target object. The detection head includes a classification head and a regression head. The classification head is used to predict the category of the target object, and the regression head is used to predict the location of the target object. During the detection process, the detection rate of the input image and the accuracy of the object type are automatically calculated; Finally, the RSODTR model outputs the location, range, and category of the target object.

4. The method for realizing automatic detection of remote sensing images based on RSODTR model according to claim 3 is characterized in that: The trained RSODTR model uses non-maximum suppression to filter repeated bounding boxes in the input image when automatically detecting the input image.

5. The method for realizing automatic detection of remote sensing images based on RSODTR model according to claim 3 is characterized in that: The RSODTR model is trained for 150 batches, each batch size is 32 images, the image size is 800×800 pixels, the AdamW optimizer is used, and the learning rate is set to 0.0003. During the training process, the RSODTR model has a 50% probability of performing up-down or left-right flipping data enhancement on the input data.

6. The method for realizing automatic detection of remote sensing images based on RSODTR model according to claim 1, characterized in that: The method of saving the location, range and category of the target object includes: Output data structure that defines the location, range, and category of target features; The location, range and category of the target feature are saved according to the defined output data structure.

7. The method for realizing automatic detection of remote sensing images based on RSODTR model according to claim 1, characterized in that: The RSODTR model identifies and displays the target object in the input image after automatic detection, including: Use OpenCV to draw detection boxes and labels to identify target objects in the automatically detected input image; The input image after automatic detection is displayed through the SciView window built into the PyCharm software.