Remote sensing target detection method, device and equipment based on large-kernel strip convolution and medium
By using large-core strip-shaped convolutional backbone network and feature pyramid network for multi-scale feature fusion and using decoupled detection heads for prediction, the problem of poor detection of large aspect ratio target detection in the prior art is solved, and efficient and accurate target detection is achieved.
Patent Information
- Application Number
- CN202510015815.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-06
AI Technical Summary
Existing remote sensing object detection technology is difficult to effectively detect various aspect ratio targets, especially large aspect ratio targets, and the calculation burden is high and low efficiency is used to use the method of large convolution kernel.
The remote sensing image is encoded and processed by a backbone network based on large core strip convolution, combined with the feature pyramid network for multi-scale feature fusion, and a decoupled detection head is used for target positioning and classification prediction.
Effective detection of various aspect ratio targets, especially large aspect ratio targets, improve detection accuracy and efficiency, and reduce calculation burden.
Smart Images

Figure CN119942332A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a remote sensing target detection method, device, equipment and medium based on large kernel strip convolution. Background Art
[0002] Remote sensing object detection is an important task in computer vision, especially when dealing with remote sensing and aerial images, which have a wide range of applications. These images usually contain various types of objects, such as vehicles, ships, bridges, and buildings. These objects are usually densely arranged and vary in size, scale, and orientation, which poses a huge challenge to object detection. The complex object features in remote sensing images include the following aspects: First, the wide coverage of remote sensing images leads to large scale variations of objects, such as ships and ports coexisting in the scene. Under the imaging perspective from top to bottom, the objects usually present a disordered directional arrangement. Therefore, the detection model needs to be sensitive to not only scale, but also direction perception. Second, the size of some objects may be very small, even occupying only a few pixels. Such objects occupy a small proportion of the entire image, and it becomes more difficult to extract features from limited pixels. Third, there may be highly similar objects in remote sensing images, such as tennis courts and baseball fields, or roads and bridges. The extracted similar features may confuse the detector and lead to wrong judgments. Finally, remote sensing images may contain special categories with extreme aspect ratios, such as mountain roads and sea-crossing bridges, whose elongated appearance makes it difficult for detectors to accurately identify their features.
[0003] However, in the field of remote sensing target detection, although the existing technology has made many progresses, there are still many problems. On the one hand, the targets in remote sensing images are often densely arranged, and their sizes, scales and directions are different, which requires the detection model to be sensitive to scale and have direction perception capabilities; at the same time, some targets are extremely small, occupying only a few pixels, and it is difficult to extract features from them. In addition, there are highly similar targets that are easy to confuse the detector, and special categories such as extreme aspect ratios make it difficult for the detector to accurately identify their features. On the other hand, in the related methods using large convolution kernels, the paradigm of using multiple large convolution kernels in parallel will increase the computational burden, and the complex block design makes the model inefficient. In addition, how to efficiently use large convolution kernels to cope with the high changes in the aspect ratio of objects is still an open problem to be solved. Summary of the invention
[0004] Based on this, it is necessary to provide a remote sensing target detection method, device, equipment and medium based on large kernel strip convolution that can effectively detect targets of various aspect ratios, especially targets with large length and aspect ratios, in order to address the above technical problems.
[0005] A remote sensing target detection method based on large kernel strip convolution, the method comprising:
[0006] Acquisition of remote sensing images;
[0007] The remote sensing image is encoded using a backbone network constructed based on large-core strip convolution to obtain a feature image, wherein the backbone network includes a normalization layer, a strip convolution unit, a normalization layer, and a feedforward network connected in sequence, the strip convolution unit includes a fully connected layer, an activation layer, a strip network, and a fully connected layer connected in sequence, and the strip network includes a square convolution layer, a horizontal strip convolution layer, a vertical strip convolution layer, and a point convolution layer connected in sequence;
[0008] A feature pyramid network is used to perform feature fusion on the feature image to obtain a multi-scale feature map set at different resolution levels;
[0009] Extracting a region of interest in each scale feature map in the multi-scale feature map set, and fusing the extracted regions of interest of different scales to obtain a fused feature;
[0010] A decoupled detection head is used to perform positioning prediction and classification prediction of the target based on the fusion features, wherein the position size and angle of the detection frame are predicted respectively when positioning the target to achieve target detection in the remote sensing image.
[0011] In one embodiment, in the backbone network:
[0012] Taking the remote sensing image as input data, sequentially passing through the normalization layer and the strip convolution unit to obtain first processed data;
[0013] Adding the first processed data to the original input data element by element to obtain second processed data;
[0014] After the second processed data is processed by the normalization layer and the feedforward network in sequence, the processing result is added element by element to the second processed data to obtain the feature image.
[0015] In one embodiment, in the strip network:
[0016] After the input data passes through the square convolution layer, the horizontal strip convolution layer, the vertical strip convolution layer and the point convolution layer, it is multiplied element by element with the original input data to obtain the output of the strip network.
[0017] In one embodiment, the convolution kernel size of the strip network is 5X5.
[0018] In one embodiment, the square convolutional layer is represented as:
[0019] The horizontal strip convolution layer is expressed as:
[0020] The vertical strip convolution layer is expressed as:
[0021] The point convolution layer is expressed as: Yc(i,j)=Kc(i,j)·Xc(i,j);
[0022] In the above formula, Xc represents the channel c of the input image X, Kc represents the channel c of the convolution kernel K, and K H represents the length of the convolution kernel, Kw represents the width of the convolution kernel, Kc(m,n) represents the value of the convolution kernel at the (m,n) position of channel c, Xc(m,n) represents the value of the input image X at the (m,n) position of channel c. Zc(i,j) represents the value of the feature map at the (i,j) position of channel c after square convolution. Indicates the value of the feature map after horizontal strip convolution at the (i, j) position of channel c, It represents the value of the feature map after vertical strip convolution at the (i, j) position of channel c, and Yc(i, j) represents the value of the feature map after point convolution at the (i, j) position of channel c.
[0023] In one embodiment, the decoupling detection head includes two branches, namely a first branch and a second branch;
[0024] The first branch is used for target classification prediction and target detection frame position and size prediction;
[0025] The second branch is used for angle prediction of the target detection frame, wherein the second branch includes a strip convolution unit.
[0026] In one embodiment, the fused features are used as input data of the first branch and the second branch respectively;
[0027] On the first branch, the fused features pass through two fully connected layers in sequence and then pass through two more fully connected layers respectively, and output the target classification prediction and the position and size prediction of the target detection box;
[0028] On the second branch, the fused features pass through the convolution layer, the strip convolution module and the fully connected layer in sequence to output the angle prediction of the target detection box.
[0029] The present application also provides a remote sensing target detection device based on large kernel strip convolution, the device comprising:
[0030] A remote sensing image acquisition module, used for acquiring remote sensing images;
[0031] A feature encoding module, used for encoding the remote sensing image using a backbone network constructed based on large-core strip convolution to obtain a feature image, wherein the backbone network includes a normalization layer, a strip convolution unit, a normalization layer, and a feedforward network connected in sequence, the strip convolution unit includes a fully connected layer, an activation layer, a strip network, and a fully connected layer connected in sequence, and the strip network includes a square convolution layer, a horizontal strip convolution layer, a vertical strip convolution layer, and a point convolution layer connected in sequence;
[0032] A multi-scale feature map acquisition module is used to perform feature fusion on the feature image using a feature pyramid network to obtain a set of multi-scale feature maps at different resolution levels;
[0033] A multi-scale feature fusion module, used to extract the region of interest in each scale feature map in the multi-scale feature map set, and fuse the extracted regions of interest of different scales to obtain a fused feature;
[0034] The target detection module is used to use a decoupled detection head to perform target positioning prediction and classification prediction based on the fusion features, wherein when performing target positioning prediction, the position size and angle of the detection frame are predicted separately to achieve target detection in the remote sensing image.
[0035] A computer device comprises a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0036] Acquisition of remote sensing images;
[0037] The remote sensing image is encoded using a backbone network constructed based on large-core strip convolution to obtain a feature image, wherein the backbone network includes a normalization layer, a strip convolution unit, a normalization layer, and a feedforward network connected in sequence, the strip convolution unit includes a fully connected layer, an activation layer, a strip network, and a fully connected layer connected in sequence, and the strip network includes a square convolution layer, a horizontal strip convolution layer, a vertical strip convolution layer, and a point convolution layer connected in sequence;
[0038] A feature pyramid network is used to perform feature fusion on the feature image to obtain a multi-scale feature map set at different resolution levels;
[0039] Extracting a region of interest in each scale feature map in the multi-scale feature map set, and fusing the extracted regions of interest of different scales to obtain a fused feature;
[0040] A decoupled detection head is used to perform positioning prediction and classification prediction of the target based on the fusion features, wherein the position size and angle of the detection frame are predicted respectively when positioning the target to achieve target detection in the remote sensing image.
[0041] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the following steps:
[0042] Acquisition of remote sensing images;
[0043] The remote sensing image is encoded using a backbone network constructed based on large-core strip convolution to obtain a feature image, wherein the backbone network includes a normalization layer, a strip convolution unit, a normalization layer, and a feedforward network connected in sequence, the strip convolution unit includes a fully connected layer, an activation layer, a strip network, and a fully connected layer connected in sequence, and the strip network includes a square convolution layer, a horizontal strip convolution layer, a vertical strip convolution layer, and a point convolution layer connected in sequence;
[0044] A feature pyramid network is used to perform feature fusion on the feature image to obtain a multi-scale feature map set at different resolution levels;
[0045] Extracting a region of interest in each scale feature map in the multi-scale feature map set, and fusing the extracted regions of interest of different scales to obtain a fused feature;
[0046] A decoupled detection head is used to perform positioning prediction and classification prediction of the target based on the fusion features, wherein the position size and angle of the detection frame are predicted respectively when positioning the target to achieve target detection in the remote sensing image.
[0047] The above-mentioned remote sensing target detection method, device, equipment and medium based on large-core strip convolution encode the remote sensing image by using a backbone network constructed based on large-core strip convolution to obtain a feature image, wherein the backbone network includes a normalization layer, a strip convolution unit, a normalization layer and a feedforward network connected in sequence, the strip convolution unit includes a fully connected layer, an activation layer, a strip network and a fully connected layer connected in sequence, and the strip network includes a square convolution layer, a horizontal strip convolution layer, a vertical strip convolution layer and a point convolution layer connected in sequence, and then a feature pyramid network is used to perform feature fusion on the feature image to obtain a multi-scale feature map set of different resolution levels, extract the region of interest in each scale feature map in the multi-scale feature map set, and fuse the extracted regions of interest of different scales to obtain a fusion feature, and finally use a decoupled detection head to perform target positioning prediction and classification prediction according to the fusion feature, wherein when the target is positioned, the position size and angle of the detection frame are predicted respectively to achieve target detection of the remote sensing image. This method can be used to effectively detect targets with various aspect ratios, especially targets with large length-to-width ratios. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 Schematic diagram of a process of a remote sensing target detection method based on large kernel strip convolution in one embodiment;
[0049] Figure 2 A schematic diagram of the structure of a backbone network constructed based on large-core strip convolution in one embodiment;
[0050] Figure 3 A schematic diagram of the structure of a decoupling detection head in one embodiment;
[0051] Figure 4 This is a schematic diagram for comparing the detection results of multiple methods in an experimental simulation;
[0052] Figure 5 A visualization diagram of the spatial sensitivity of different methods in an experimental simulation;
[0053] Figure 6 is a structural block diagram of a remote sensing target detection device based on large kernel strip convolution in one embodiment;
[0054] Figure 7 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0055] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0056] For existing remote sensing image target detection, there is still a problem that high aspect ratio targets cannot be successfully detected. In this application, Figure 1 As shown, a remote sensing target detection method based on large kernel strip convolution is provided, which specifically includes the following steps:
[0057] Step S100, acquiring a remote sensing image.
[0058] Step S110, using a backbone network constructed based on large-kernel strip convolution to encode the remote sensing image to obtain a feature image, wherein the backbone network includes a normalization layer, a strip convolution unit, a normalization layer and a feedforward network connected in sequence, the strip convolution unit includes a fully connected layer, an activation layer, a strip network and a fully connected layer connected in sequence, and the strip network includes a square convolution layer, a horizontal strip convolution layer, a vertical strip convolution layer and a point convolution layer connected in sequence.
[0059] Step S120, using a feature pyramid network to perform feature fusion on the feature image to obtain a set of multi-scale feature maps at different resolution levels.
[0060] Step S130, extracting the region of interest in each scale feature map in the multi-scale feature map set, and fusing the extracted regions of interest of different scales to obtain a fused feature.
[0061] Step S140, using a decoupled detection head, performs location prediction and classification prediction of the target based on the fusion features, wherein when performing location prediction of the target, the position size and angle of the detection frame are predicted separately to achieve target detection of the remote sensing image.
[0062] In this application, a sequential combination of square convolution and strip convolution is designed to effectively extract the features of objects with different aspect ratios. The accuracy of target positioning is improved by decoupling the detection head, and the prediction of angle information is assisted by coupling the classification prediction branch and the angle prediction branch, further improving the prediction ability of the detection head.
[0063] In step S100, remote sensing images refer to images obtained by remote sensing technology that record information about electromagnetic waves reflected or emitted by objects on the earth's surface. The targets in remote sensing images have the characteristics of large scale changes, diverse directions, dense distribution, and similar features. When there are high aspect ratio targets in remote sensing images, it is difficult to locate and classify such targets due to the difficulty in feature extraction, complex target representation and positioning, and data imbalance and sample scarcity during neural network training.
[0064] Specifically, the targets with high aspect ratio in the remote sensing image to be detected may be bridges, railways, roads, oil or gas pipelines, high-voltage towers, power lines, rivers, streams, mountains or ridgelines, etc.
[0065] In step S110, the specific structure of the backbone network constructed based on large-core strip convolution is as follows: Figure 2 As shown. In this backbone network: the remote sensing image is used as input data, and it passes through the normalization layer SyncBN and the strip convolution unit in sequence. The convolution kernel size of the horizontal strip convolution unit is 1×19, and the convolution kernel size of the vertical strip convolution unit is 19×1 to obtain the first processed data. After the first processed data is added element by element to the original input data, the second processed data is obtained. The second processed data is processed by the normalization layer SyncBN and the feedforward network in sequence. The feedforward network contains two linear layers and an activation layer. The processing result is added element by element to the second processed data to obtain the feature image.
[0066] Furthermore, the large square kernel convolution in the above backbone network provides the necessary long-range context information for remote sensing applications, but may contain irrelevant features from the background area. Therefore, strip convolution, i.e., strip network, is introduced in the model design, but it still relies on parallel large-size square kernel convolution, which brings problems of computational burden and feature redundancy. The goal of this method is to efficiently extract the key features of objects with different aspect ratios. To this end, a sequence paradigm is proposed in the strip network, which can effectively combine the advantages of standard convolution and strip convolution without the need for additional information fusion modules.
[0067] like Figure 2 As shown in the figure, in the strip network: after the input data passes through the square convolution layer, the horizontal strip convolution layer, the vertical strip convolution layer and the point convolution layer, it is multiplied element by element with the original input data to obtain the output of the strip network.
[0068] Specifically, given an input tensor X with 3 channels, a depthwise separable convolution with a square kernel K, i.e., a square convolution layer, is first used to extract local context features. In practical applications, the convolution kernel size can be set to 5×5.
[0069] Furthermore, after the initial depth-wise separable convolution, two sequential large-scale strip convolutions, horizontal and vertical, are used in the strip network to better extract spatial information. Unlike standard convolutions that extract features from a square area each time, large-scale strip convolutions enable the network to focus more on features along the horizontal or vertical axis. The combined use of horizontal and vertical large-scale strip convolutions, i.e., horizontal strip convolution layers and vertical strip convolution layers, enables the network to collect directional features in the spatial dimension and enhance the expression of slender or narrow structures. In order to further enhance the feature interaction in the channel dimension, a simple point convolution layer is applied so that the network can learn the dependencies between different channels.
[0070] The backbone network based on large-kernel strip convolution proposed in this method is much simpler than other existing remote sensing target detectors using large-kernel convolution. In this method, the basic block design does not use any spatial or channel attention mechanism, nor does it use composite fusion operations of different types of large-kernel convolutions, which makes the backbone network structure of this method very simple, but performs well in different remote sensing detection benchmarks. At the same time, in the strip network, there is no significant difference between applying horizontal strip convolution first and then vertical convolution, or vice versa, and both technical methods are effective.
[0071] Furthermore, the square convolutional layer is expressed as:
[0072] The horizontal strip convolution layer is represented as:
[0073] The vertical strip convolution layer is expressed as:
[0074] The point convolution layer is expressed as: Yc(i,j)=Kc(i,j)·Xc(i,j).
[0075] In the above formula, Xc represents the channel c of the input image X, Kc represents the channel c of the convolution kernel K, and K H represents the length of the convolution kernel, Kw represents the width of the convolution kernel, Kc(m,n) represents the value of the convolution kernel at the (m,n) position of channel c, Xc(m,n) represents the value of the input image X at the (n,n) position of channel c. Zc(i,j) represents the value of the feature map at the (i,j) position of channel c after square convolution. Indicates the value of the feature map after horizontal strip convolution at the (i, j) position of channel c, It represents the value of the feature map after vertical strip convolution at the (i, j) position of channel c, and Yc(i, j) represents the value of the feature map after point convolution at the (i, j) position of channel c.
[0076] In step S120, a feature pyramid network is used to fuse the feature image to obtain a multi-scale feature map set at different resolution levels. For high aspect ratio targets, such as cross-sea bridges, different parts of the target have features of different scales in the image. The feature pyramid network (FPN) can fuse high-level semantic features (low resolution, but contains strong semantic information) and low-level detail features (high resolution, rich details). For example, when detecting a cross-sea bridge, high-level features can help understand the overall category information of the bridge (such as whether it is a highway bridge or a railway bridge), and low-level features can capture details such as the edge and piers of the bridge, which is conducive to the complete detection of different parts of the high aspect ratio target. At the same time, the length and width directions of high aspect ratio targets have large scale differences. Taking a slender oil pipeline as an example, its length direction may span multiple regions of different resolutions, while the width direction may only be better perceived at a certain resolution level. By fusing multi-scale features, FPN can effectively extract the length and width direction features of the pipeline at different resolution levels, so that the model can better adapt to this scale difference and improve the accuracy of detection.
[0077] Furthermore, the feature pyramid network can also be used to locate the boundaries of high aspect ratio targets, and FPN provides great help. For example, when detecting high-voltage tower lines, it is critical to accurately find their boundary positions because the lines are long and thin. By fusing feature maps of different resolution levels, the precise edge position of the line can be determined at the low-resolution level with rich details, and the direction of the line can be determined by using high-level semantic information, so as to more accurately locate the boundaries of high aspect ratio targets and reduce positioning errors. At the same time, high aspect ratio targets may have various directional changes in the image. The multi-scale features in the FPN structure can assist in judging the direction of the target from different angles. Taking a curved river as an example, the overall flow direction of the river can be grasped in the high-resolution high-level feature map, and the details and directional changes of the river bend can be captured in the high-resolution low-level feature map. In this way, the integration of information at different scales helps to accurately detect high aspect ratio targets in different directions.
[0078] In the localization task, the model should be sensitive to transformations, because the accuracy of localization depends on the position of the input object. Previous strong remote sensing object detectors adopted the popular Oriented R-CNN framework, whose detection heads share the same fully connected layers for classification and localization tasks. However, the spatial correlation of fully connected layers is limited, which makes them insensitive to transformations and unsuitable for precise localization.
[0079] In step S130, in order to better solve the above problem, a better solution is adopted, that is, to decouple the classification and positioning tasks and use small kernel convolution in the positioning branch. However, analysis of the spatial correlation graph shows that small kernel convolution only captures short-range spatial correlation, which is insufficient for the long-range dependency required to accurately locate slender objects. In order to effectively locate objects of different aspect ratios, it is believed that the positioning head should be able to capture long-range dependencies, similar to those processed by the backbone network. Large strip convolution captures horizontal and vertical features across wide spatial areas, providing the extended spatial correlation required for enhanced positioning. Therefore, in the present application, large strip convolution is used in the positioning branch to improve the positioning capability of the detector.
[0080] Specifically, in the target detection method in this method, the parameters that need to be predicted include: x, y are the center coordinates of the target detection box, w and h are the width and height of the target detection box, and the rotation angle θ of the target detection box, which may lead to the problem of feature coupling. In order to solve this problem, the idea of decoupling the prediction of θ from other parameters (x, y, w, h) is proposed. Based on this decoupling method, the spatial sensitive areas of classification and angle prediction in the output feature map are proposed. In the spatial sensitivity map, the first column represents the classification sensitivity, and the second column shows the angle sensitivity. The spatial sensitivity of classification is concentrated in the central area of the object, while the sensitivity of angle is mainly concentrated near the boundary of the object. The sensitive areas of classification and angle prediction have some overlap and show complementary characteristics. Based on this, in this method, a shared fully connected layer is used for classification and angle prediction.
[0081] In this embodiment, the decoupled detection head includes two branches, namely a first branch and a second branch. The first branch is used for target classification prediction and position and size prediction of the target detection frame, and the second branch is used for angle prediction of the target detection frame, wherein the second branch includes a strip convolution unit.
[0082] like Figure 3 As shown in the figure, the fused features are used as the input data of the first branch and the second branch respectively. In the first branch, the fused features pass through two fully connected layers in sequence, and then pass through two fully connected layers (i.e., the classification head and the positioning head respectively), and output the target classification prediction and the position and size prediction of the target detection frame. In the second branch, the fused features pass through the convolution layer, the strip convolution module and the fully connected layer (i.e., the angle head) in sequence, and output the angle prediction of the target detection frame.
[0083] Furthermore, the classification head simply uses two fully connected layers with an output dimension of 1024. The localization head starts with a standard 3×3 convolution to extract local features. Then a strip network is added, followed by a fully connected layer to collect long-range spatial dependencies. And the angle head: three fully connected layers are used to effectively estimate the angle information.
[0084] It is important to note here that the fully connected layers in the classification head, localization head, and angle head share parameters. Note that the first two fully connected layers share parameters with the classification head.
[0085] In this embodiment, the backbone network constructed based on large-kernel strip convolution, the feature pyramid network, the region of interest extraction, and the decoupled detection head can be regarded as a complete remote sensing target detection network. The acquired remote sensing image can be used as the input of the network to obtain the target detection result in the remote sensing image.
[0086] In this embodiment, all three heads, including the classification head, the localization head, and the angle estimation head, are jointly trained end-to-end, wherein the classification head uses a cross entropy loss function, and the localization head and the angle estimation head use a Smooth L1 loss function.
[0087] In this paper, simulation results are also given to prove the effectiveness of this method. Figure 4 As shown in the simulation experiment, this method is applied to multiple remote sensing data sets for experimental verification. Figure 4 It can be seen that the first to third columns are the detection results and heat map visualization results of some of the most advanced methods in the past, and the fourth column is the detection results and heat map visualization results of this method. From the detection results, it can be seen that the most advanced methods in the past have missed detections and wrong detections for long objects, while this method has a good detection effect on such objects. From the heat map, it can be seen that this method has a good response result for such long objects, indicating that the network built by this method can pay better attention to such objects.
[0088] In the above-mentioned remote sensing target detection method based on large-core strip convolution, the remote sensing image is encoded by using a backbone network constructed based on large-core strip convolution to obtain a feature image, wherein the backbone network includes a normalization layer, a strip convolution unit, a normalization layer and a feedforward network connected in sequence, the strip convolution unit includes a fully connected layer, an activation layer, a strip network and a fully connected layer connected in sequence, and the strip network includes a square convolution layer, a horizontal strip convolution layer, a vertical strip convolution layer and a point convolution layer connected in sequence, and then a feature pyramid network is used to perform feature fusion on the feature image to obtain a multi-scale feature map set of different resolution levels, extract the region of interest in each scale feature map in the multi-scale feature map set, and fuse the extracted regions of interest of different scales to obtain a fusion feature, and finally use a decoupled detection head to perform target positioning prediction and classification prediction according to the fusion feature, wherein the position size and angle of the detection frame are predicted respectively when the target is positioned, so as to realize the target detection of the remote sensing image. The detection head in this method fuses the proposed strip unit. This design greatly improves the detection capability relative to the original detection head. As can be seen from the visualization of the spatial sensitivity of different methods, Figure 5 As shown, the left side is the original detection head and the right side is the detection head of this method. The spatial correlation patterns of the original detection heads on the output feature maps are similar, but the detection head that integrates strip units in this method shows more spatial correlation on the output feature maps, which enables this method to better locate objects at higher aspect ratios.
[0089] It should be understood that although Figure 1 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 1 At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.
[0090] In one embodiment, Figure 6 As shown, a remote sensing target detection device based on large kernel strip convolution is provided, comprising: a remote sensing image acquisition module 200, a feature encoding module 210, a multi-scale feature map acquisition module 220, a multi-scale feature fusion module 230 and a target detection module 240, wherein:
[0091] A remote sensing image acquisition module 200 is used to acquire remote sensing images;
[0092] A feature encoding module 210 is used to encode the remote sensing image using a backbone network constructed based on large-core strip convolution to obtain a feature image, wherein the backbone network includes a normalization layer, a strip convolution unit, a normalization layer, and a feedforward network connected in sequence, the strip convolution unit includes a fully connected layer, an activation layer, a strip network, and a fully connected layer connected in sequence, and the strip network includes a square convolution layer, a horizontal strip convolution layer, a vertical strip convolution layer, and a point convolution layer connected in sequence;
[0093] A multi-scale feature map obtaining module 220 is used to perform feature fusion on the feature image using a feature pyramid network to obtain a set of multi-scale feature maps at different resolution levels;
[0094] A multi-scale feature fusion module 230 is used to extract the region of interest in each scale feature map in the multi-scale feature map set, and fuse the extracted regions of interest of different scales to obtain a fused feature;
[0095] The target detection module 240 is used to use a decoupled detection head to perform target positioning prediction and classification prediction based on the fusion features, wherein when performing target positioning prediction, the position size and angle of the detection frame are predicted separately to achieve target detection in the remote sensing image.
[0096] For the specific limitations of the remote sensing target detection device based on large kernel strip convolution, please refer to the limitations of the remote sensing target detection method based on large kernel strip convolution above, which will not be repeated here. Each module in the above-mentioned remote sensing target detection device based on large kernel strip convolution can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0097] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 7As shown. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a remote sensing target detection method based on large-kernel strip convolution is implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a key, trackball or touchpad set on the computer device housing, or an external keyboard, touchpad or mouse, etc.
[0098] Those skilled in the art will understand that Figure 7 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0099] In one embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the following steps are implemented:
[0100] Acquisition of remote sensing images;
[0101] The remote sensing image is encoded using a backbone network constructed based on large-core strip convolution to obtain a feature image, wherein the backbone network includes a normalization layer, a strip convolution unit, a normalization layer, and a feedforward network connected in sequence, the strip convolution unit includes a fully connected layer, an activation layer, a strip network, and a fully connected layer connected in sequence, and the strip network includes a square convolution layer, a horizontal strip convolution layer, a vertical strip convolution layer, and a point convolution layer connected in sequence;
[0102] A feature pyramid network is used to perform feature fusion on the feature image to obtain a multi-scale feature map set at different resolution levels;
[0103] Extracting a region of interest in each scale feature map in the multi-scale feature map set, and fusing the extracted regions of interest of different scales to obtain a fused feature;
[0104] A decoupled detection head is used to perform positioning prediction and classification prediction of the target based on the fusion features, wherein the position size and angle of the detection frame are predicted respectively when positioning the target to achieve target detection in the remote sensing image.
[0105] In one embodiment, a computer readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:
[0106] Acquisition of remote sensing images;
[0107] The remote sensing image is encoded using a backbone network constructed based on large-core strip convolution to obtain a feature image, wherein the backbone network includes a normalization layer, a strip convolution unit, a normalization layer, and a feedforward network connected in sequence, the strip convolution unit includes a fully connected layer, an activation layer, a strip network, and a fully connected layer connected in sequence, and the strip network includes a square convolution layer, a horizontal strip convolution layer, a vertical strip convolution layer, and a point convolution layer connected in sequence;
[0108] A feature pyramid network is used to perform feature fusion on the feature image to obtain a multi-scale feature map set at different resolution levels;
[0109] Extracting a region of interest in each scale feature map in the multi-scale feature map set, and fusing the extracted regions of interest of different scales to obtain a fused feature;
[0110] A decoupled detection head is used to perform positioning prediction and classification prediction of the target based on the fusion features, wherein the position size and angle of the detection frame are predicted respectively when positioning the target to achieve target detection in the remote sensing image.
[0111] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0112] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0113] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.
Claims
1. A remote sensing target detection method based on large kernel strip convolution, characterized in that: The method comprises: Acquisition of remote sensing images; The remote sensing image is encoded using a backbone network constructed based on large-core strip convolution to obtain a feature image, wherein the backbone network includes a normalization layer, a strip convolution unit, a normalization layer, and a feedforward network connected in sequence, the strip convolution unit includes a fully connected layer, an activation layer, a strip network, and a fully connected layer connected in sequence, and the strip network includes a square convolution layer, a horizontal strip convolution layer, a vertical strip convolution layer, and a point convolution layer connected in sequence; A feature pyramid network is used to perform feature fusion on the feature image to obtain a multi-scale feature map set at different resolution levels; Extracting a region of interest in each scale feature map in the multi-scale feature map set, and fusing the extracted regions of interest of different scales to obtain a fused feature; A decoupled detection head is used to perform positioning prediction and classification prediction of the target based on the fusion features, wherein the position size and angle of the detection frame are predicted respectively when positioning the target to achieve target detection in the remote sensing image.
2. The remote sensing target detection method according to claim 1, characterized in that: In the backbone network: Taking the remote sensing image as input data, sequentially passing through the normalization layer and the strip convolution unit to obtain first processed data; Adding the first processed data to the original input data element by element to obtain second processed data; After the second processed data is processed by the normalization layer and the feedforward network in sequence, the processing result is added element by element to the second processed data to obtain the feature image.
3. The remote sensing target detection method according to claim 2, characterized in that: In the strip network: After the input data passes through the square convolution layer, the horizontal strip convolution layer, the vertical strip convolution layer and the point convolution layer, it is multiplied element by element with the original input data to obtain the output of the strip network.
4. The remote sensing target detection method according to claim 3, characterized in that: The convolution kernel size of the strip network is 5X5.
5. The remote sensing target detection method according to claim 4, characterized in that: The square convolutional layer is expressed as: The horizontal strip convolution layer is expressed as: The vertical strip convolution layer is expressed as: The point convolution layer is expressed as: Yc(i,j)=Kc(i,j)·Xc(i,j); In the above formula, Xc represents the channel c of the input image X, Kc represents the channel c of the convolution kernel K, and K H represents the length of the convolution kernel, Kw represents the width of the convolution kernel, Kc(m,n) represents the value of the convolution kernel at the (m,n) position of channel c, Xc(m,n) represents the value of the input image X at the (m,n) position of channel c. Zc(i,j) represents the value of the feature map at the (i,j) position of channel c after square convolution. Indicates the value of the feature map after horizontal strip convolution at the (i, j) position of channel c, It represents the value of the feature map after vertical strip convolution at the (i, j) position of channel c, and Yc(i, j) represents the value of the feature map after point convolution at the (i, j) position of channel c.
6. The remote sensing target detection method according to any one of claims 1 to 5, characterized in that: The decoupling detection head includes two branches, namely a first branch and a second branch; The first branch is used for target classification prediction and target detection frame position and size prediction; The second branch is used for angle prediction of the target detection frame, wherein the second branch includes a strip convolution unit.
7. The remote sensing target detection method according to any one of claims 1 to 5, characterized in that: The fused features are used as input data of the first branch and the second branch respectively; On the first branch, the fused features pass through two fully connected layers in sequence and then pass through two more fully connected layers respectively, and output the target classification prediction and the position and size prediction of the target detection box; On the second branch, the fused features pass through the convolution layer, the strip convolution module and the fully connected layer in sequence to output the angle prediction of the target detection box.
8. A remote sensing target detection device based on large kernel strip convolution, characterized in that: The device comprises: A remote sensing image acquisition module, used for acquiring remote sensing images; A feature encoding module, used for encoding the remote sensing image using a backbone network constructed based on large-core strip convolution to obtain a feature image, wherein the backbone network includes a normalization layer, a strip convolution unit, a normalization layer, and a feedforward network connected in sequence, the strip convolution unit includes a fully connected layer, an activation layer, a strip network, and a fully connected layer connected in sequence, and the strip network includes a square convolution layer, a horizontal strip convolution layer, a vertical strip convolution layer, and a point convolution layer connected in sequence; A multi-scale feature map acquisition module is used to perform feature fusion on the feature image using a feature pyramid network to obtain a set of multi-scale feature maps at different resolution levels; A multi-scale feature fusion module, used to extract the region of interest in each scale feature map in the multi-scale feature map set, and fuse the extracted regions of interest of different scales to obtain a fused feature; The target detection module is used to use a decoupled detection head to perform target positioning prediction and classification prediction based on the fusion features, wherein when performing target positioning prediction, the position size and angle of the detection frame are predicted separately to achieve target detection in the remote sensing image.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Rotating target detection method and device, equipment and medium
CN117853706A
Low-visibility environment traffic sign detection method based on deep learning
CN118799841A