Remote sensing image target detection method, system and device based on deep learning
By adopting multi-scale depth separation convolution and strip convolution attention mechanisms in remote sensing image processing, the problems of large calculation overhead and inference delay in the prior art are solved, and efficient remote sensing image object detection is achieved.
Patent Information
- Application Number
- CN202411328327.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-24
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-09-24
AI Technical Summary
In the existing remote sensing image object detection methods, the computing overhead and inference delay of the backbone network are relatively large, making it difficult to effectively process the difference in target scales and complex background information in remote sensing images.
Using a deep learning-based method, the remote sensing image data can be processed through the depth of multiple convolution kernels of different scales and sizes. Combined with the strip convolution attention mechanism, features are extracted and enhanced to achieve feature fusion and object detection.
It effectively alleviates the problems of target scale differences and complex background information in target detection, and achieves high detection accuracy and high inference speed.
Smart Images

Figure CN119091322B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of remote sensing image processing, and specifically relates to a remote sensing image target detection method, system, and device based on deep learning. Background Art
[0002] Arbitrary Oriented Object Detection (AOOD) has been widely used to detect objects in different directions in remote sensing images. The backbone networks of existing methods mostly use convolutional neural networks (CNN) and Transformer. However, the huge computational overhead and inference delay brought to the model by the multi-layer convolution stacking in CNN and the attention mechanism of Transformer hinder their application in practice. In addition, in remote sensing images, most objects are in the shape of bars, such as vehicles, ships, ports, etc., and adaptive object detection methods are needed in terms of object scale differences and complex background information. Summary of the invention
[0003] In order to solve the problems raised by the background technology, the present invention provides a remote sensing image target detection method, system and device based on deep learning.
[0004] The technical solution of the present invention is as follows:
[0005] The present invention provides a remote sensing image target detection method based on deep learning, comprising the following steps:
[0006] S1: After obtaining the remote sensing image data to be processed, it is processed by depthwise separable convolution of five convolution kernels of different scales, and the obtained features of different scales are added to obtain the first feature;
[0007] The first feature is processed by batch normalization, the first activation function, deep convolution, and batch normalization in sequence to obtain the second feature;
[0008] S2: The second feature is extracted and enhanced to obtain the third feature;
[0009] The third feature is sequentially downsampled, extracted and enhanced to obtain the fourth feature;
[0010] The fourth feature is sequentially downsampled, feature extracted and enhanced to obtain the fifth feature;
[0011] The fifth feature is sequentially downsampled, feature extracted and enhanced to obtain the sixth feature;
[0012] The extraction and enhancement of the features include the following operations:
[0013] After the input features are copied, they are recorded as initial features and features to be processed;
[0014] The features to be processed are processed by layer normalization and linear processing in turn to obtain the processed features;
[0015] After data replication, the processed features are recorded as features to be extracted and features to be enhanced;
[0016] The features to be extracted are processed sequentially by deep convolution, second activation function, selective scanning mechanism, first activation function, and layer normalization to obtain the extracted features;
[0017] After the features to be enhanced are processed by the strip convolution attention, they are processed by the second activation function to obtain enhanced features;
[0018] After the extracted features and enhanced features are multiplied, they are added to the initial features to obtain the output features;
[0019] S3: The sixth feature is processed by convolution and feature fusion is performed. The obtained fused feature is processed by classification to generate a rotation positioning frame of the target; the rotation positioning frame is mapped to the fused feature to generate a candidate frame. After classification and regression, the position and category of the target are obtained.
[0020] In step S1, after obtaining the remote sensing image data to be processed, the data is subjected to depthwise separable convolution processing using convolution kernels of five different scales, and the obtained features of different scales are added together to obtain the first feature, which is specifically:
[0021] The remote sensing image data to be processed is processed by a depth-separable convolution of a first convolution kernel to obtain a first scale feature;
[0022] The remote sensing image data to be processed is processed by depth-separable convolution of a second convolution kernel to obtain a second scale feature;
[0023] The remote sensing image data to be processed is processed by depth-separable convolution of a third convolution kernel to obtain a third scale feature;
[0024] The remote sensing image data to be processed is processed by depth-separable convolution of a fourth convolution kernel to obtain a fourth scale feature;
[0025] The remote sensing image data to be processed is processed by depth-separable convolution of the fifth convolution kernel to obtain the fifth scale feature;
[0026] The first scale feature, the second scale feature, the third scale feature, the fourth scale feature and the fifth scale feature are added to obtain the first feature.
[0027] In step S2, the feature to be enhanced is processed by the strip convolution attention and then processed by the second activation function to obtain the enhanced feature, specifically:
[0028] The features to be enhanced are activated by the first convolution, the second convolution, the third convolution, the first convolution, and the second activation function in sequence to obtain enhanced features;
[0029] The convolution kernel of the first convolution is 1×1, the convolution kernel of the second convolution is 1×3, and the convolution kernel of the third convolution is 3×1.
[0030] In step S3, the sixth feature is processed by convolution to perform feature fusion, specifically:
[0031] After the sixth feature is processed by convolutions of four different scales, four features of different scales are obtained; each feature is processed by convolution and downsampling, and the four processed features are fused.
[0032] In step S3, the obtained fusion features are classified and processed to generate a rotation positioning frame of the target, specifically:
[0033] After the bounding box is determined by the obtained fusion features, classification processing is performed, and the position offset of the classified bounding box is calculated. After screening, the rotation positioning frame of the target is generated.
[0034] In step S3, the rotation positioning frame is mapped to the fusion feature to generate a candidate frame. After classification and regression, the position and category of the target are obtained, specifically:
[0035] Get the coordinate values of the rotation positioning frame, map the coordinate values to the fusion features, and get the mapping area; divide the mapping area into several areas, perform pooling on each area, and generate a candidate frame; after classification and regression through the fully connected layer, get the position and category of the target.
[0036] The first convolution kernel is 3×3, the second convolution kernel is 5×5, the third convolution kernel is 7×7, the fourth convolution kernel is 9×9, and the fifth convolution kernel is 11×11.
[0037] The first activation function is the GELU activation function; the second activation function is the SiLU activation function.
[0038] The present invention also provides a remote sensing image target detection system based on deep learning, comprising:
[0039] Preprocessing module: After obtaining the remote sensing image data to be processed, it is processed by depth-separable convolution of five convolution kernels of different scales. After adding the features of different scales, the first feature is obtained;
[0040] The first feature is processed by batch normalization, the first activation function, deep convolution, and batch normalization in sequence to obtain the second feature;
[0041] Feature extraction and enhancement module: used to extract and enhance the second feature to obtain the third feature;
[0042] The third feature is sequentially downsampled, extracted and enhanced to obtain the fourth feature;
[0043] The fourth feature is sequentially downsampled, feature extracted and enhanced to obtain the fifth feature;
[0044] The fifth feature is sequentially downsampled, feature extracted and enhanced to obtain the sixth feature;
[0045] The extraction and enhancement of the features include the following operations:
[0046] After the input features are copied, they are recorded as initial features and features to be processed;
[0047] The features to be processed are processed by layer normalization and linear processing in turn to obtain the processed features;
[0048] After data replication, the processed features are recorded as features to be extracted and features to be enhanced;
[0049] The features to be extracted are processed sequentially by deep convolution, second activation function, selective scanning mechanism, first activation function, and layer normalization to obtain the extracted features;
[0050] After the features to be enhanced are processed by the strip convolution attention, they are processed by the second activation function to obtain enhanced features;
[0051] After the extracted features and enhanced features are multiplied, they are added to the initial features to obtain the output features;
[0052] Target detection module: The sixth feature is processed by convolution and feature fusion is performed. The obtained fused feature is processed by classification to generate the rotation positioning frame of the target; the rotation positioning frame is mapped to the fused feature to generate a candidate frame. After classification and regression, the position and category of the target are obtained.
[0053] The present invention also provides a remote sensing image target detection device based on deep learning, comprising a processor and a memory, wherein the processor implements the remote sensing image target detection method based on deep learning when executing a computer program stored in the memory.
[0054] Beneficial effects: The present invention performs depth-separable convolutions of multiple convolution kernels of different scales on the remote sensing image data to be processed to capture information of different scales and effectively expand the receptive field; a strip convolution attention mechanism is adopted, and a convolution combination including a 1×3 convolution kernel and a 3×1 convolution kernel is used to perform feature enhancement on horizontal strip targets and vertical strip targets respectively. The strip convolution combination can not only help understand contextual information, but also enhance the sensitivity to strip objects; the obtained features effectively alleviate the problems of target scale differences and complex background information in target detection tasks, thereby achieving high detection accuracy and maintaining high reasoning speed.
[0055] The method provided by the present invention only needs to process the remote sensing image data to be processed once through depth-separable convolution of multiple convolution kernels of different scales. The processed data is then passed to the backbone network for layer-by-layer processing, instead of performing multi-core convolution in the backbone network. This method can not only directly process global information, but also effectively control the overall computational overhead. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 It is a flowchart of a remote sensing image target detection method based on deep learning in an embodiment. DETAILED DESCRIPTION
[0057] The following examples are intended to illustrate the present invention rather than to further limit the present invention.
[0058] The present invention provides a remote sensing image target detection method based on deep learning, such as Figure 1 As shown, the following steps are included:
[0059] Considering that the incoming data is global data, it is crucial for subsequent feature extraction and feature fusion. Therefore, extracting global information as fully as possible in the preprocessing stage is crucial to the overall performance of the subsequent feature extraction network.
[0060] S1: After obtaining the remote sensing image data to be processed, it is processed by depth-separable convolution with five convolution kernels of different scales, and the obtained features of different scales are added to obtain the first feature. Specifically:
[0061] The remote sensing image data to be processed is processed by a depth-separable convolution of a first convolution kernel to obtain a first scale feature;
[0062] The remote sensing image data to be processed is processed by depth-separable convolution of a second convolution kernel to obtain a second scale feature;
[0063] The remote sensing image data to be processed is processed by depth-separable convolution of a third convolution kernel to obtain a third scale feature;
[0064] The remote sensing image data to be processed is processed by depth-separable convolution of a fourth convolution kernel to obtain a fourth scale feature;
[0065] The remote sensing image data to be processed is processed by depth-separable convolution of the fifth convolution kernel to obtain the fifth scale feature;
[0066] The first scale feature, the second scale feature, the third scale feature, the fourth scale feature and the fifth scale feature are added to obtain the first feature.
[0067] In order to fully extract the background features around the target to assist detection and classification, preferably, the first convolution kernel is 3×3, the second convolution kernel is 5×5, the third convolution kernel is 7×7, the fourth convolution kernel is 9×9, and the fifth convolution kernel is 11×11.
[0068] The above operation can be used to extract multi-scale features through the following formula:
[0069] ,
[0070] In the formula, is the first feature; is the remote sensing image data to be processed; Indicates that the convolution kernel is a 3×3 depth-separable convolution; Indicates that the convolution kernel is a 5×5 depth-separable convolution; Indicates that the convolution kernel is 7×7 depth-separable convolution; Indicates that the convolution kernel is a 9×9 depth-separable convolution; It indicates a depth-wise separable convolution with a convolution kernel of 11×11.
[0071] After the multi-core depthwise separable convolution processing in the preprocessing stage, the first feature has rich multi-scale information.
[0072] Then, according to the formula: The first feature is processed by batch standardization, the first activation function, deep convolution, and batch standardization in sequence to obtain the second feature.
[0073] In the formula, is the second feature; is the first feature; Indicates batch normalization processing; Indicates that the convolution kernel is a 3×3 depth convolution; Indicates that the GELU activation function is used as the first activation function.
[0074] Afterwards, the second feature containing rich multi-scale information and fully integrated is passed into the feature extraction network for processing.
[0075] In the present invention, multiple depth-separable convolutions of convolution kernels of different scales are performed on the remote sensing image data to be processed, which can capture information of different scales and effectively expand the receptive field. The dynamic receptive field can adapt to target detection of different scales and fully extract the background feature information of the target. The method provided by the present invention only needs to process the remote sensing image data to be processed once, and the processed data is then passed to the backbone network for layer-by-layer processing, rather than performing multi-core convolution in the backbone network. This method can not only directly process global information, but also effectively control the overall computational overhead.
[0076] S2: The second feature is extracted and enhanced to obtain the third feature.
[0077] Considering that there are still deficiencies in processing target scale differences and complex background information in remote sensing image processing, the present invention chooses to apply the attention mechanism locally on this basis.
[0078] The extraction and enhancement of the features include the following operations:
[0079] After the input features are copied, they are recorded as initial features and features to be processed;
[0080] The features to be processed are processed by layer normalization and linear processing in turn to obtain the processed features;
[0081] After data replication, the processed features are recorded as features to be extracted and features to be enhanced;
[0082] The features to be extracted are processed sequentially by deep convolution, second activation function, selective scanning mechanism, first activation function, and layer normalization to obtain the extracted features;
[0083] After the features to be enhanced are processed by the strip convolution attention, they are processed by the second activation function to obtain enhanced features;
[0084] After the extracted features and enhanced features are multiplied, they are added to the initial features to obtain the output features.
[0085] In the step of "extracting and enhancing the second feature to obtain the third feature", specifically:
[0086] After the data is copied, the second feature is recorded as the initial feature and the feature to be processed;
[0087] The features to be processed are processed by layer normalization and linear processing in turn to obtain the processed features; the processed features are reduced in dimension, thus removing redundant features and reducing the computational overhead of subsequent processing.
[0088] After data replication, the processed features are recorded as features to be extracted and features to be enhanced;
[0089] According to the formula: The features to be extracted are processed sequentially by deep convolution, second activation function, selective scanning mechanism, first activation function, and layer normalization to obtain the extracted features.
[0090] In the formula, To extract features; is the feature to be extracted; Representation layer normalization (LayerNormalization) processing; represents a first activation function, preferably a GELU activation function; Indicates the selective scanning mechanism processing (2D-Selective-Scan), which can perform all-round feature scanning and feature extraction; represents the second activation function; Represents depthwise convolution.
[0091] After the features to be enhanced are processed by the strip convolution attention, they are processed by the second activation function to obtain enhanced features;
[0092] After the extracted features and the enhanced features are multiplied, they are added to the initial features to obtain the third features. That is, the third features include the initial features, the extracted features, and the enhanced features.
[0093] Specifically, it can be based on the formula: , , and obtain the third feature;
[0094] In the formula, The feature is the multiplication of the extracted feature and the enhanced feature; is the feature to be enhanced; To extract features; is the initial feature; is the third characteristic; represents a second activation function, preferably a SiLU activation function; represents strip convolution attention processing; Represents element-wise multiplication.
[0095] Based on the comprehensive scanning and extraction of features by the selective scanning mechanism (SS2D), in order to control the computational cost and achieve local feature enhancement, a lightweight and efficient attention mechanism needs to be selected when enhancing features.
[0096] Since most targets in remote sensing images are strip-shaped, such as vehicles, ships, ports, etc. Therefore, the features of strip-shaped objects are taken as the main targets of local feature enhancement.
[0097] In addition, in order to make full use of background features and further extract the extracted background information, the strip convolution attention mechanism is introduced to enhance feature expression using strip convolution.
[0098] Preferably, in step S2, the feature to be enhanced is processed by the strip convolution attention and then processed by the second activation function to obtain the enhanced feature, specifically:
[0099] The features to be enhanced are processed by the first convolution, the second convolution, the third convolution, the first convolution, and the second activation function in sequence to obtain enhanced features;
[0100] The convolution kernel of the first convolution is 1×1, the convolution kernel of the second convolution is 1×3, and the convolution kernel of the third convolution is 3×1. Before and after the strip convolution combination (i.e., the second convolution and the third convolution), the convolution kernel of the first convolution is 1×1, which can optimize feature extraction.
[0101] The present invention adopts a strip convolution attention mechanism, and uses a convolution combination including a 1×3 convolution kernel and a 3×1 convolution kernel to enhance the features of horizontal strip targets and vertical strip targets respectively. The strip convolution combination can not only help understand the context information, but also enhance the sensitivity to strip objects.
[0102] The third feature is sequentially downsampled, extracted and enhanced to obtain the fourth feature;
[0103] The fourth feature is sequentially downsampled, feature extracted and enhanced to obtain the fifth feature;
[0104] The fifth feature is sequentially downsampled, feature extracted and enhanced to obtain the sixth feature;
[0105] Similarly, in the steps of obtaining the fourth feature, the fifth feature, and the sixth feature, the "feature extraction and enhancement" is the same as the processing operation of obtaining the third feature.
[0106] The third feature obtained has a resolution of ; The resolution after downsampling is , the resolution of the fourth feature is ; The resolution after downsampling is , the resolution of the fifth feature is ; The downsampled resolution is , the resolution of the sixth feature is .in, , , and Indicates different number of channels; , They are height and width respectively.
[0107] S3: The sixth feature is processed by convolution and feature fusion is performed. The obtained fused feature is processed by classification to generate a rotation positioning frame of the target; the rotation positioning frame is mapped to the fused feature to generate a candidate frame. After classification and regression, the position and category of the target are obtained.
[0108] Preferably, in step S3, the sixth feature is subjected to convolution processing to perform feature fusion, specifically as follows:
[0109] After the sixth feature is processed by convolutions of four different scales, four features of different scales are obtained, and the specifications are as follows: , , , ; Each feature is processed by convolution and downsampling, and the four features after processing are fused.
[0110] Preferably, in step S3, the obtained fusion features are classified and processed to generate a rotation positioning frame of the target, specifically:
[0111] After the bounding box is determined by the obtained fusion features, classification processing is performed, such as using a softmax classifier to classify the bounding box into positive and negative, calculating the position offset of the classified bounding box, and after correction and screening, generating a rotation positioning frame of the target, thereby completing the positioning of the target.
[0112] Preferably, the rotation positioning frame is mapped to the fusion feature to generate a candidate frame. After classification and regression, the position and category of the target are obtained, specifically:
[0113] Get the coordinate value of the rotation positioning frame, map the coordinate value to the fusion feature, and get the mapping area; divide the mapping area into several areas, perform pooling on each area, and generate a candidate frame; after classification and regression through the fully connected layer, get the position and category of the target, and realize target detection.
[0114] The present invention performs multiple depth-separable convolutions with different kernels on the remote sensing image data to be processed to capture information of different scales and effectively expand the receptive field; adopts a strip convolution attention mechanism, in which the strip convolution combination can not only help understand contextual information, but also enhance the sensitivity to strip objects; the obtained features effectively alleviate the problems of target scale differences and complex background information in target detection tasks, achieving both high detection accuracy and high inference speed.
[0115] The present invention also provides a remote sensing image target detection system based on deep learning, comprising:
[0116] Preprocessing module: After obtaining the remote sensing image data to be processed, it is processed by depth-separable convolution of five convolution kernels of different scales. After adding the features of different scales, the first feature is obtained;
[0117] The first feature is processed by batch normalization, the first activation function, deep convolution, and batch normalization in sequence to obtain the second feature;
[0118] Feature extraction and enhancement module: used to extract and enhance the second feature to obtain the third feature;
[0119] The third feature is sequentially downsampled, extracted and enhanced to obtain the fourth feature;
[0120] The fourth feature is sequentially downsampled, feature extracted and enhanced to obtain the fifth feature;
[0121] The fifth feature is sequentially downsampled, feature extracted and enhanced to obtain the sixth feature;
[0122] The extraction and enhancement of the features include the following operations:
[0123] After the input features are copied, they are recorded as initial features and features to be processed;
[0124] The features to be processed are processed by layer normalization and linear processing in turn to obtain the processed features;
[0125] After data replication, the processed features are recorded as features to be extracted and features to be enhanced;
[0126] The features to be extracted are processed sequentially by deep convolution, second activation function, selective scanning mechanism, first activation function, and layer normalization to obtain the extracted features;
[0127] After the features to be enhanced are processed by the strip convolution attention, they are processed by the second activation function to obtain enhanced features;
[0128] After the extracted features and enhanced features are multiplied, they are added to the initial features to obtain the output features;
[0129] Target detection module: The sixth feature is processed by convolution and feature fusion is performed. The obtained fused feature is processed by classification to generate the rotation positioning frame of the target; the rotation positioning frame is mapped to the fused feature to generate a candidate frame. After classification and regression, the position and category of the target are obtained.
[0130] The present invention also provides a remote sensing image target detection device based on deep learning, comprising a processor and a memory, wherein the processor implements the remote sensing image target detection method based on deep learning when executing a computer program stored in the memory.
Claims
1. A remote sensing image target detection method based on deep learning, characterized in that: The following steps are involved: S1: After obtaining the remote sensing image data to be processed, it is processed by depthwise separable convolution of five convolution kernels of different scales, and the obtained features of different scales are added to obtain the first feature; The first feature is processed by batch normalization, the first activation function, deep convolution, and batch normalization in sequence to obtain the second feature; S2: The second feature is extracted and enhanced to obtain the third feature; The third feature is sequentially downsampled, extracted and enhanced to obtain the fourth feature; The fourth feature is sequentially downsampled, feature extracted and enhanced to obtain the fifth feature; The fifth feature is sequentially downsampled, feature extracted and enhanced to obtain the sixth feature; The extraction and enhancement of the features include the following operations: After the input features are copied, they are recorded as initial features and features to be processed; The features to be processed are processed by layer normalization and linear processing in turn to obtain the processed features; After data replication, the processed features are recorded as features to be extracted and features to be enhanced; The features to be extracted are processed sequentially by deep convolution, second activation function, selective scanning mechanism, first activation function, and layer normalization to obtain the extracted features; After the features to be enhanced are processed by the strip convolution attention, they are processed by the second activation function to obtain enhanced features; After the extracted features and enhanced features are multiplied, they are added to the initial features to obtain the output features; S3: The sixth feature is processed by convolution and feature fusion is performed. The obtained fused feature is processed by classification to generate a rotation positioning frame of the target; the rotation positioning frame is mapped to the fused feature to generate a candidate frame. After classification and regression, the position and category of the target are obtained.
2. The remote sensing image target detection method based on deep learning according to claim 1, characterized in that: In step S1, after obtaining the remote sensing image data to be processed, the data is subjected to depthwise separable convolution processing using convolution kernels of five different scales, and the obtained features of different scales are added together to obtain the first feature, which is specifically: The remote sensing image data to be processed is processed by a depth-separable convolution of a first convolution kernel to obtain a first scale feature; The remote sensing image data to be processed is processed by depth-separable convolution of a second convolution kernel to obtain a second scale feature; The remote sensing image data to be processed is processed by depth-separable convolution of a third convolution kernel to obtain a third scale feature; The remote sensing image data to be processed is processed by depth-separable convolution of a fourth convolution kernel to obtain a fourth scale feature; The remote sensing image data to be processed is processed by depth-separable convolution of the fifth convolution kernel to obtain the fifth scale feature; The first scale feature, the second scale feature, the third scale feature, the fourth scale feature and the fifth scale feature are added to obtain the first feature.
3. The remote sensing image target detection method based on deep learning according to claim 1, characterized in that: In step S2, the feature to be enhanced is processed by the strip convolution attention and then processed by the second activation function to obtain the enhanced feature, specifically: The features to be enhanced are activated by the first convolution, the second convolution, the third convolution, the first convolution, and the second activation function in sequence to obtain enhanced features; The convolution kernel of the first convolution is 1×1, the convolution kernel of the second convolution is 1×3, and the convolution kernel of the third convolution is 3×1.
4. The remote sensing image target detection method based on deep learning according to claim 1, characterized in that: In step S3, the sixth feature is processed by convolution to perform feature fusion, specifically: After the sixth feature is processed by convolutions of four different scales, four features of different scales are obtained; each feature is processed by convolution and downsampling, and the four processed features are fused.
5. The remote sensing image target detection method based on deep learning according to claim 1, characterized in that: In step S3, the obtained fusion features are classified and processed to generate a rotation positioning frame of the target, specifically: After the bounding box is determined by the obtained fusion features, classification processing is performed, and the position offset of the classified bounding box is calculated. After screening, the rotation positioning frame of the target is generated.
6. The remote sensing image target detection method based on deep learning according to claim 1, characterized in that: In step S3, the rotation positioning frame is mapped to the fusion feature to generate a candidate frame. After classification and regression, the position and category of the target are obtained, specifically: Get the coordinate values of the rotation positioning frame, map the coordinate values to the fusion features, and get the mapping area; divide the mapping area into several areas, perform pooling on each area, and generate a candidate frame; after classification and regression through the fully connected layer, get the position and category of the target.
7. The remote sensing image target detection method based on deep learning according to claim 2, characterized in that: The first convolution kernel is 3×3, the second convolution kernel is 5×5, the third convolution kernel is 7×7, the fourth convolution kernel is 9×9, and the fifth convolution kernel is 11×11.
8. The remote sensing image target detection method based on deep learning according to claim 1, characterized in that: The first activation function is the GELU activation function; the second activation function is the SiLU activation function.
9. A remote sensing image target detection system based on deep learning, characterized in that: include: Preprocessing module: After obtaining the remote sensing image data to be processed, it is processed by depth-separable convolution of five convolution kernels of different scales. After adding the features of different scales, the first feature is obtained; The first feature is processed by batch normalization, the first activation function, deep convolution, and batch normalization in sequence to obtain the second feature; Feature extraction and enhancement module: used to extract and enhance the second feature to obtain the third feature; The third feature is sequentially downsampled, extracted and enhanced to obtain the fourth feature; The fourth feature is sequentially downsampled, feature extracted and enhanced to obtain the fifth feature; The fifth feature is sequentially downsampled, feature extracted and enhanced to obtain the sixth feature; The extraction and enhancement of the features include the following operations: After the input features are copied, they are recorded as initial features and features to be processed; The features to be processed are processed by layer normalization and linear processing in turn to obtain the processed features; After data replication, the processed features are recorded as features to be extracted and features to be enhanced; The features to be extracted are processed sequentially by deep convolution, second activation function, selective scanning mechanism, first activation function, and layer normalization to obtain the extracted features; After the features to be enhanced are processed by the strip convolution attention, they are processed by the second activation function to obtain enhanced features; After the extracted features and enhanced features are multiplied, they are added to the initial features to obtain the output features; Target detection module: The sixth feature is processed by convolution and feature fusion is performed. The obtained fused feature is processed by classification to generate the rotation positioning frame of the target; the rotation positioning frame is mapped to the fused feature to generate a candidate frame. After classification and regression, the position and category of the target are obtained.
10. A remote sensing image target detection device based on deep learning, characterized in that: The method comprises a processor and a memory, wherein when the processor executes the computer program stored in the memory, the remote sensing image target detection method based on deep learning as described in any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Vehicle target detection method for complex road scene
CN118506319A
System and method for machine learning application for providing medical test results using visual indicia
US20190027251A1