Lightweight target detection method and device for remote sensing image, equipment and medium
By building a lightweight object detection model, using frequency-direction sensitive convolution layer and frequency-direction attention convolution layer, the problem of large model parameters and high computational overhead in remote sensing image object detection is solved, and efficient object detection and resource conservation is achieved.
Patent Information
- Application Number
- CN202510764868.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-10
AI Technical Summary
When the existing remote sensing image object detection model runs on edge devices with limited computing resources, there are problems such as large number of model parameters and high computing overhead, resulting in low detection efficiency.
A lightweight object detection model is built, using parallel set frequency-direction sensitive convolution layer and cascading frequency-direction attention convolution layer, image features are captured from different frequency domain angles through wavelet transformation, and features are refined and enhanced layer by layer, reducing model parameters and calculation amount.
It improves the accuracy of remote sensing image object detection, reduces missed detection and misdetection, realizes the lightweight design of the model, reduces the computing resource requirements, and expands the application scenarios of edge devices.
Smart Images

Figure CN120279260A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of object detection, and in particular, to a lightweight object detection method, device, equipment and medium for remote sensing images. Background Technique
[0002] With the rapid development of neural network technology and deep learning technology, many applications have penetrated into all aspects of the industrial and life fields, such as spectral data applications, millimeter wave imaging, medical diagnosis, traffic rescue, and so on. Among them, remote sensing technology is also increasingly relying on deep learning models. Especially with the wide application of unmanned aerial vehicle and satellite technologies, it has become relatively easy to obtain a large number of high-resolution remote sensing images.
[0003] At present, many researchers are committed to improving the accuracy and efficiency of object detection in remote sensing images. For example, super-resolution, feature fusion, data augmentation, semi-supervised learning, etc. can significantly improve the detection accuracy of object detection in remote sensing images. However, the multi-scale feature fusion structure may increase the parameters of the model, and data augmentation and semi-supervised methods may increase the time overhead of data preprocessing and training. Convolution operation is one of the most basic operations used for feature extraction in neural networks. However, the standard convolution operation requires a large amount of computational overhead, and the number of parameters will increase exponentially with the increase of the convolution kernel size.
[0004] However, for the application of object detection in remote sensing images, in many cases, it is run on edge devices with limited computing resources. Therefore, model lightweighting is one of the most important research directions in the object detection task of remote sensing images. Summary of the Invention
[0005] Based on this, in view of the above technical problems, it is necessary to provide a lightweight object detection method, device, equipment and medium for remote sensing images that are more suitable for lightweight object detection tasks in remote sensing images.
[0006] A lightweight object detection method for remote sensing images, the method includes: Construct a lightweight object detection model, the backbone network of the lightweight object detection model includes an initialization module and two or more sequentially connected feature extraction layers; wherein, the initialization module includes two or more frequency-direction sensitive convolution layers arranged in parallel, and the feature extraction layer includes two or more frequency-direction attention convolution layers connected in cascade; Input the obtained remote sensing image into the initialization module, and perform feature extraction on the remote sensing image through the frequency-direction sensitive convolution layers arranged in parallel to obtain a first output feature; Input the first output feature into a sequentially connected feature extraction layer. In each feature extraction layer, perform feature extraction through two or more cascaded frequency-direction attention convolutional layers, and use the output feature of the previous feature extraction layer as the input feature of the next feature extraction layer to finally obtain a second output feature; Process the second output feature through other modules of the lightweight object detection model to output an object detection image.
[0007] A lightweight object detection device for remote sensing images, the device includes: A model construction module for constructing a lightweight object detection model. The backbone network of the lightweight object detection model includes an initialization module and two or more sequentially connected feature extraction layers; wherein, the initialization module includes two or more parallel frequency-direction sensitive convolutional layers, and the feature extraction layer includes two or more cascaded frequency-direction attention convolutional layers; A first output feature extraction module for inputting the obtained remote sensing image into the initialization module, and performing feature extraction on the remote sensing image through parallel frequency-direction sensitive convolutional layers to obtain a first output feature; A second output feature extraction module for inputting the first output feature into a sequentially connected feature extraction layer. In each feature extraction layer, perform feature extraction through two or more cascaded frequency-direction attention convolutional layers, and use the output feature of the previous feature extraction layer as the input feature of the next feature extraction layer to finally obtain a second output feature; An object detection image output module for processing the second output feature through other modules of the lightweight object detection model to output an object detection image.
[0008] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the lightweight object detection method for remote sensing images are implemented.
[0009] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the lightweight object detection method for remote sensing images are implemented.
[0010] The above lightweight object detection method, device, equipment, and medium for remote sensing images construct a lightweight object detection model. The backbone network of the lightweight object detection model includes an initialization module and more than two sequentially connected feature extraction layers. Among them, the initialization module includes more than two frequency-direction sensitive convolutional layers arranged in parallel, and the feature extraction layer includes more than two cascaded frequency-direction attention convolutional layers. The obtained remote sensing image is input into the initialization module, and the frequency-direction sensitive convolutional layers arranged in parallel are used to extract features from the remote sensing image to obtain a first output feature. The first output feature is input into the sequentially connected feature extraction layers. In each feature extraction layer, feature extraction is performed through more than two cascaded frequency-direction attention convolutional layers, and the output feature of the previous feature extraction layer is used as the input feature of the next feature extraction layer, and finally a second output feature is obtained. The second output feature is processed by other modules of the lightweight object detection model to output an object detection image.
[0011] The beneficial effects of the present invention are as follows: By reconstructing the backbone network and setting more than two parallel frequency-direction sensitive convolutional layers in the initialization module, with the help of wavelet transform, it can capture the features of remote sensing images from different frequency domain angles and improve the frequency domain feature capture ability. More than two feature extraction layers are sequentially connected, and the frequency-direction attention convolutional layers within each layer are cascaded, which can refine and enhance the features layer by layer, making the final output feature more discriminative.
[0012] The combination of the frequency-direction sensitive convolutional layer and the frequency-direction attention convolutional layer. The former obtains rich frequency domain features, and the latter focuses on key features through the attention mechanism. The two cooperate to improve the detection accuracy of targets in remote sensing images, reduce missed detections and false detections, and thus can capture target features more accurately.
[0013] Using the frequency-direction sensitive convolutional layer and the frequency-direction attention convolutional layer to construct the backbone network, compared with the traditional complex convolutional structure, it can reduce the number of model parameters and the amount of calculation, reduce the model complexity, and achieve the lightweight design of the model structure. On the premise of ensuring the detection accuracy, the demand for computing resources is reduced, and it can be deployed on resource-constrained edge devices to expand the application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the structures shown in these drawings.
[0015] Figure 1 It is a schematic flowchart of the lightweight object detection method for remote sensing images provided in Embodiment 1; Figure 2Schematic diagram of the backbone network structure framework provided in Embodiment 1; Figure 3 Schematic diagram of the initialization module structure framework provided in Embodiment 1; Figure 4 Schematic diagram of the feature extraction layer structure framework provided in Embodiment 1; Figure 5 Schematic diagram of the frequency-sensitive convolution layer structure framework provided in Embodiment 1; Figure 6 Schematic diagram of the frequency attention convolution layer structure framework provided in Embodiment 1; Figure 7 Schematic diagram of the first group of detection effect comparison provided in Embodiment 1, Figure 7 (a) Schematic diagram of the detection result obtained by using the default ResNet-50, Figure 7 (b) Schematic diagram of the detection result obtained by using the method proposed in the present invention; Figure 8 Schematic diagram of the second group of detection effect comparison provided in Embodiment 1, Figure 8 (a) Schematic diagram of the detection result obtained by using the default ResNet-50, Figure 8 (b) Schematic diagram of the detection result obtained by using the method proposed in the present invention; Figure 9 Structure block diagram of the lightweight object detection device for remote sensing images provided in Embodiment 2; Figure 10 Internal structure diagram of the computer device provided in Embodiment 3.
[0016] The realization of the object, functional features and advantages of the present invention will be further described in conjunction with the embodiments with reference to the accompanying drawings. Detailed implementation manners
[0017] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0018] It can be understood that in the present invention, descriptions such as "first" and "second" are only for descriptive purposes, and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present invention, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0019] In addition, the technical solutions between the various embodiments of the present invention can be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.
[0020] Next, the embodiments of the present invention will be described in detail in conjunction with the accompanying drawings in the embodiments of the present invention.
[0021] Embodiment 1 This embodiment discloses a lightweight object detection method for remote sensing images. By reconstructing the backbone network and setting two or more parallel frequency-direction sensitive convolutional layers in the initialization module, with the help of wavelet transform, it can capture the features of remote sensing images from different frequency domain angles and improve the ability to capture frequency domain features. Two or more feature extraction layers are connected in sequence, and the frequency-direction attention convolutional layers within each layer are cascaded, which can refine and enhance the features layer by layer, making the finally output features more discriminative.
[0022] The frequency-direction sensitive convolutional layer is combined with the frequency-direction attention convolutional layer. The former obtains rich frequency domain features, and the latter focuses on key features through the attention mechanism. The two cooperate to improve the detection accuracy of objects in remote sensing images, reduce missed detections and false detections, and thus can capture target features more accurately.
[0023] Using the frequency-direction sensitive convolutional layer and the frequency-direction attention convolutional layer to construct the backbone network, compared with the traditional complex convolutional structure, it can reduce the number of model parameters and the amount of calculation, reduce the model complexity, achieve the lightweight design of the model structure, reduce the demand for computing resources on the premise of ensuring the detection accuracy, and can be deployed on resource-constrained edge devices to expand the application scenarios.
[0024] As Figure 1 shown, the lightweight object detection method for remote sensing images provided by this embodiment includes the following steps: Step 201, construct a lightweight object detection model. The backbone network of the lightweight object detection model includes an initialization module and two or more sequentially connected feature extraction layers; wherein, the initialization module includes two or more parallelly arranged frequency-direction sensitive convolutional layers, and the feature extraction layer includes two or more cascaded frequency-direction attention convolutional layers.
[0025] Step 202, input the obtained remote sensing image into the initialization module, and perform feature extraction on the remote sensing image through the parallelly arranged frequency-direction sensitive convolutional layers to obtain the first output feature.
[0026] Step 203, input the first output feature into the feature extraction layers connected in sequence, and in each feature extraction layer, perform feature extraction through two or more cascaded frequency-directed attention convolutional layers, and use the output feature of the previous feature extraction layer as the input feature of the next feature extraction layer, and finally obtain the second output feature.
[0027] Step 204: Process the second output feature through other modules of the lightweight target detection model to output a target detection image.
[0028] It can be understood that the lightweight target detection model structure constructed by this embodiment is mainly a lightweight design of the backbone network structure in the model, and other modules and loss functions reuse the design of the existing basic detection model. By seamlessly replacing the backbone network in the existing basic detector with the backbone network constructed in this embodiment, and then performing fine-tuning training, the lightweight target detection model constructed by this embodiment is obtained. Therefore, this embodiment mainly describes the backbone network, and other modules and loss functions are not described in detail.
[0029] like Figure 2 As shown, the backbone network (FOSNet for short) provided in this embodiment includes an initialization module (MSLayer for short) and more than two feature extraction layers. The initialization module is located in the first layer of the backbone network, followed by the feature extraction layer, which has the same structure and is arranged in sequence from top to bottom.
[0030] like Figure 3 As shown, the initialization module provided in this embodiment includes a downsampling layer, two or more frequency-sensitive convolution layers (FOSConv for short) arranged in parallel, and an average pooling layer, and the two or more frequency-sensitive convolution layers arranged in parallel are located between the downsampling layer and the average pooling layer. The frequency-sensitive convolution layers have the same structure, and the difference lies in the different scales of the convolution kernels.
[0031] like Figure 4 As shown, the feature extraction layer provided in this embodiment includes a downsampling layer and two or more cascaded frequency-directed attention convolution layers (RABlock for short). The frequency-directed attention convolution layers have the same structure and are connected behind the downsampling module.
[0032] It is worth noting that the number of feature extraction layers, frequency-sensitive convolution layers, and frequency-attention convolution layers is set according to the needs. For the convenience of subsequent explanation, in this implementation, the number of feature extraction layers is set to 3; the number of frequency-sensitive convolution layers is set to 5, and the convolution kernel scales are respectively and ; In the three feature extraction layers, the number of frequency-oriented attention convolutional layer cascades set They are 3, 4, and 6 respectively, and the number of channels of the output vector in each stage is 128, 256, and 512 respectively, and the output resolutions are one-fourth, one-eighth, and one-sixteenth of the output image respectively. It should be noted that the above specific quantity settings are only one illustration given for the convenience of understanding and explaining the solution, and do not constitute a specific limitation to the present invention. The same meaning applies to subsequent specific quantity settings, so no further elaboration will be made.
[0033] In one embodiment, as Figure 5 shown, it is a schematic structural framework diagram of the frequency-direction sensitive convolution layer provided by this embodiment. The frequency-direction sensitive convolution layer includes a wavelet transform layer and an inverse wavelet transform layer. There are two or more parallel first convolution layers arranged between the wavelet transform layer and the inverse wavelet transform layer; a second convolution layer is also arranged after the inverse wavelet transform layer.
[0034] Specifically, the wavelet transform layer is mainly used to perform wavelet transform on the input feature map, decompose it into different frequency bands, and obtain frequency submaps The quantity is set according to requirements. In this embodiment, 4 wavelet transform coefficients are set in the wavelet transform layer, that is, and , and 4 frequency submaps are output. The first convolution layer is mainly used to extract features from the frequency submaps , and its quantity corresponds to the quantity of the frequency submaps . Therefore, the first convolution layer is also set to 4. The convolution kernel size is set according to the characteristics of the frequency submaps. In this embodiment, the convolution kernel sizes of the first convolution layer are divided into , , and ; the inverse wavelet transform layer is mainly used to reconstruct features and is the inverse process of wavelet transform; the second convolution layer is set to 1, mainly used to restore the number of channels, and the convolution kernel size is , where is the convolution kernel scale set for the frequency-direction sensitive convolution layer. In this embodiment, the wavelet transform layer can be constructed using db1 wavelet, db2 wavelet, db3 wavelet, db4 wavelet, etc., and preferably db1 wavelet is used. Through the frequency-direction sensitive convolution layer, the input features can be decomposed into different frequencies and directions, and specific-shaped convolutions are used to extract features. While reducing parameters and time overhead, better feature extraction capabilities can be achieved.
[0035] It should be noted that since the frequency-direction sensitive convolution layer is used in both the initialization module and the frequency-direction attention convolution layer of the feature extraction layer, the feature maps input to the frequency-direction sensitive convolution layer are different in different stages. Figure 5 What is shown in is the input In the frequency - direction attention convolutional layer, the input of the frequency - direction sensitive convolutional layer is the second sub - feature map , and the output is the fifth feature map .
[0036] In one embodiment, before inputting the acquired remote - sensing image into the initialization module and performing feature extraction on the remote - sensing image through the parallel - set frequency - direction sensitive convolutional layers, it further includes: downsampling the remote - sensing image to obtain the first feature map .
[0037] In one embodiment, inputting the acquired remote - sensing image into the initialization module and performing feature extraction on the remote - sensing image through the parallel - set frequency - direction sensitive convolutional layers to obtain the first output feature includes: Inputting the first feature map into the parallel - set frequency - direction sensitive convolutional layers respectively for feature extraction.
[0038] In each frequency - direction sensitive convolutional layer, first perform wavelet transform on the first feature map through the wavelet transform layer to obtain more than two frequency sub - maps .
[0039] Input the frequency sub - maps into the corresponding first convolutional layer respectively for feature extraction to obtain more than two second feature maps .
[0040] Input more than two second feature maps into the inverse wavelet transform layer for inverse wavelet transform, and then input them into the second convolutional layer for processing to obtain the third feature map .
[0041] Add the third feature maps and then perform average pooling processing to obtain the first output feature .
[0042] Specifically, the remote - sensing image is input into the initialization module, first passes through a downsampling layer to obtain the first feature map , and then the first feature map is input into 5 parallel frequency - direction sensitive convolutional layers respectively for feature extraction to obtain the third feature map . The third feature maps output by the 5 frequency - direction sensitive convolutional layers are respectively denoted as 、 、 、 、 . Add the feature maps 、 、 , , After addition, through an average pooling layer, the first output feature is output, that is: .
[0043] In each frequency-direction sensitive convolutional layer, first, the input first feature map is subjected to wavelet transform through a wavelet transform layer, and the obtained frequency sub-maps are respectively the low-frequency sub-map , the horizontal high-frequency sub-map , the vertical high-frequency sub-map and the diagonal high-frequency sub-map . These frequency sub-maps are respectively input into the corresponding first convolutional layer for feature extraction to obtain the second feature maps , , and . Through the operation of the first convolutional layer, the number of channels of the frequency sub-map is reduced to one-fourth of the original, thus reducing the number of network parameters. Then, an inverse wavelet transform is performed through an inverse wavelet transform layer, and the number of channels is restored to the original using the second convolutional layer, and the third feature map is output.
[0044] In one embodiment, as Figure 6 shown, it is a schematic structural framework diagram of the frequency-direction attention convolutional layer provided by this embodiment. The frequency-direction attention convolutional layer includes an attention layer and a frequency-direction sensitive convolutional layer arranged in parallel. At the output ends of the attention layer and the frequency-direction sensitive convolutional layer, a third convolutional layer is connected; and a residual connection is made between the input end and the output end of the frequency-direction attention convolutional layer.
[0045] Specifically, the attention layer includes channel attention and spatial attention, mainly used to enhance the ability to focus on key information. The frequency-direction sensitive convolutional layer is arranged in parallel with the attention layer, and the convolutional kernel scale is . The third convolutional layer is set to 1, mainly used to restore the number of channels, and the convolutional kernel size is .
[0046] In one embodiment, before feature extraction through two or more cascaded frequency-direction attention convolutional layers in each feature extraction layer, it further includes: downsampling the first output feature to obtain the fourth feature map .
[0047] In one embodiment, feature extraction through two or more cascaded frequency-direction attention convolutional layers in each feature extraction layer includes: In the first frequency-oriented attention convolutional layer, the fourth feature map is sliced into a first sub-feature map and a second sub-feature map .
[0048] The first sub-feature map is input into the attention layer for feature extraction to obtain an attention vector ; the second sub-feature map is input into the frequency-oriented sensitive convolutional layer for feature extraction to obtain a fifth feature map .
[0049] The attention vector is multiplied by the fifth feature map and then input into the third convolutional layer to output a sixth feature map .
[0050] A residual operation is performed on the fourth feature map and the sixth feature map to obtain a seventh feature map output by each frequency-oriented attention convolutional layer .
[0051] The seventh feature map is input into the next frequency-oriented attention convolutional layer for processing to obtain the output features of the feature extraction layer; The output features of the previous feature extraction layer are used as the input features of the next feature extraction layer, and finally the second output features are obtained.
[0052] It can be understood that for sequentially connected feature extraction layers, the input features of each feature extraction layer are the features output by the previous feature extraction layer . The feature is input into the th feature extraction layer. First, it passes through a downsampling layer to obtain a fourth feature map , and then the fourth feature map is input into two or more cascaded frequency-oriented attention convolutional layers for feature extraction to output the second output features . Among them, in the th feature extraction layer, the number of cascaded frequency-oriented attention convolutional layers is .
[0053] Two or more cascaded frequency-oriented attention convolutional layers are regarded as a group of frequency-oriented attention convolutional modules. In a group of frequency-oriented attention convolutional modules, the fourth feature map is input into the first frequency-oriented attention convolutional layer. First, it is evenly divided along the channel dimension into a first sub-feature map and a second sub-feature map Among them, the first sub-feature map enters the input attention layer for feature extraction to obtain an attention vector . Since the attention layer includes channel attention and spatial attention, the attention vectors are respectively denoted as and ; The second sub-feature map enters a frequency-direction sensitive convolutional layer for feature extraction to obtain the fifth feature map .
[0054] After multiplying the attention vectors , and the fifth feature map , and then passing through a third convolutional layer to restore its number of channels, the sixth feature map is output.
[0055] Perform a residual operation on the fourth feature map and the sixth feature map to obtain the seventh feature map output by each frequency-direction attention convolutional layer. The expression is:[[]] .
[0056] By reconstructing the backbone network, the present invention can be easily applied to existing basic detectors. Among them, the frequency-direction sensitive convolutional layer, as the core component, simultaneously considers the characteristics of frequency and direction, decomposes the input features into different frequencies and directions, and uses a convolutional layer with a specific shape to extract features, which can not only achieve better performance, but also significantly reduce the parameters and time overhead. Based on this, a backbone network including an initialization module and a feature extraction layer is constructed, which can be used to replace the backbone network in the traditional model, so as to better cope with the large-scale scale changes of target objects in remote sensing images, further reduce the network parameters, not only ensure the detection accuracy, but also achieve a lightweight design.
[0057] In one of the embodiments, in order to intuitively compare the improvement in the target detection accuracy rate brought by applying the method proposed by the present invention, taking the classic two-stage general target detector Faster R-CNN as the benchmark, training and testing are carried out on the remote sensing image target detection dataset AI-TOD, and the detection effects of using the default ResNet-50 and using the FOSNet proposed by the present invention as the backbone network are respectively compared. Except for the backbone network, other settings of the model are the same. The specific detection effects are as Figure 7 and Figure 8 shown, where Figure 7 (a) and Figure 8 (a) are the detection effects of using the default ResNet-50Figure 7 (b) And Figure 8 (b) shows the detection effect using the FOSNet proposed by the present invention as the backbone network. Among them, the green, blue, and red rectangular boxes represent the correctly detected, misdetected, and undetected target objects respectively. By comparison, it can be intuitively found that the method proposed by the present invention has brought a significant improvement in the detection accuracy rate.
[0058] In addition to the visual comparison, detailed experimental verification was also carried out on the AI-TOD dataset. The specific results are shown in Table 1. Among them, the first to third rows are the detection results on the commonly used one-stage object detectors, the fourth to seventh rows are the detection results on the commonly used two-stage object detectors, the eighth to twelfth rows are the detection results on some of the latest proposed backbone networks for remote sensing image object detection and conventional image object detection tasks, and the thirteenth and fourteenth rows are the detection results based on the method proposed by the present invention. For each row of results, in addition to comparing the detection accuracy rate, the performance of each detector and backbone network in terms of the number of parameters was also compared, which is an important indicator reflecting the computational efficiency of the detector and backbone network.
[0059] The FOSNet proposed by the present invention can be easily replaced with the backbone network in the commonly used basic detectors. From the results in Table 1, it can be seen that compared with the basic detectors Faster R-CNN and Cascade R-CNN, the FOSNet proposed by the present invention has respectively brought a 5.5 and 4.6 percentage point increase in the accuracy rate on the premise of reducing the number of parameters by 20M. In addition, a comparison of the number of parameters and the accuracy rate was also made with the latest proposed backbone networks ARC-R50, LSKNet-S, PKINet-S for remote sensing image object detection and the lightweight networks FastViT-T12, RepViT-M1.1 for conventional object detection. From the experimental results, it can be seen that the FOSNet proposed by the present invention has achieved satisfactory results in both the number of parameters and the accuracy rate, which fully demonstrates the high efficiency and effectiveness of the FOSNet proposed by the present invention.
[0060] Table 1 Comparison of the accuracy rate of the test results on the AI-TOD dataset
[0061] In summary, the method proposed by the present invention can well handle the remote sensing image object detection task, not only can achieve accurate detection effects, but also can significantly reduce the number of parameters of the model.
[0062] Although this embodiment Figure 1The steps in [description] are shown in sequence according to the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 1 At least some of the steps in [description] may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least some of the sub-steps or stages of other steps.
[0063] Embodiment 2 Based on the lightweight object detection method for remote sensing images in Embodiment 1, this embodiment discloses a lightweight object detection device for remote sensing images, as Figure 9 shown. The lightweight object detection device for remote sensing images includes: a model construction module 401, a first output feature extraction module 402, a second output feature extraction module 403, and an object detection image output module 404, where: The model construction module 401 is used to construct a lightweight object detection model. The backbone network of the lightweight object detection model includes an initialization module and two or more sequentially connected feature extraction layers; among them, the initialization module includes two or more frequency-direction sensitive convolutional layers arranged in parallel, and the feature extraction layers include two or more cascaded frequency-direction attention convolutional layers.
[0064] The first output feature extraction module 402 is used to input the obtained remote sensing image into the initialization module, and perform feature extraction on the remote sensing image through the frequency-direction sensitive convolutional layers arranged in parallel to obtain a first output feature; The second output feature extraction module 403 is used to input the first output feature into the sequentially connected feature extraction layers. In each feature extraction layer, perform feature extraction through two or more cascaded frequency-direction attention convolutional layers, and use the output feature of the previous feature extraction layer as the input feature of the next feature extraction layer, and finally obtain a second output feature.
[0065] The object detection image output module 404 is used to process the second output feature through other modules of the lightweight object detection model and output an object detection image.
[0066] In this embodiment, the specific working processes and principles of the model construction module 401, the first output feature extraction module 402, the second output feature extraction module 403, and the target detection image output module 404 are the same as those of the method in Embodiment 1. Therefore, they will not be elaborated in this embodiment. Each of these unit modules can be implemented in whole or in part by software, hardware, or a combination thereof. Each unit module can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form to facilitate the processor to call and execute the operations corresponding to each of the above unit modules.
[0067] Embodiment 3 As Figure 10 shown, a terminal device disclosed in this embodiment includes a transmitter, a receiver, a memory, and a processor. Among them, the transmitter is used to send instructions and data, the receiver is used to receive instructions and data, the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions stored in the memory to implement the method in Embodiment 1 above.
[0068] It should be noted that the above memory can be either independent or integrated with the processor. When the memory is set independently, the terminal device further includes a bus for connecting the memory and the processor.
[0069] Embodiment 4 This embodiment discloses a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the method in Embodiment 1 above.
[0070] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0071] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0072] The above-described embodiments merely represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the appended claims.
Claims
1. A lightweight object detection method for remote sensing images, characterized in that, The method includes: Constructing a lightweight object detection model, the backbone network of the lightweight object detection model includes an initialization module and more than two sequentially connected feature extraction layers; wherein, the initialization module includes more than two parallel frequency-direction sensitive convolutional layers, and the feature extraction layers include more than two cascaded frequency-direction attention convolutional layers; Inputting the obtained remote sensing image into the initialization module, and performing feature extraction on the remote sensing image through the parallel frequency-direction sensitive convolutional layers to obtain a first output feature; Inputting the first output feature into the sequentially connected feature extraction layers, in each feature extraction layer, performing feature extraction through more than two cascaded frequency-direction attention convolutional layers, and using the output feature of the previous feature extraction layer as the input feature of the next feature extraction layer, and finally obtaining a second output feature; Processing the second output feature through other modules of the lightweight object detection model to output an object detection image.
2. The lightweight object detection method for remote sensing images according to claim 1, wherein, The frequency-direction sensitive convolutional layer includes a wavelet transform layer and an inverse wavelet transform layer, and more than two parallel first convolutional layers are arranged between the wavelet transform layer and the inverse wavelet transform layer; a second convolutional layer is further arranged after the inverse wavelet transform layer.
3. The lightweight object detection method for remote sensing images according to claim 2, characterized in that Before inputting the obtained remote sensing image into the initialization module and performing feature extraction on the remote sensing image through the parallel frequency-direction sensitive convolutional layers, it further includes: Downsample the remote sensing image to obtain a first feature map .
4. The lightweight object detection method for remote sensing images according to claim 3, characterized in that Inputting the obtained remote sensing image into the initialization module, and performing feature extraction on the remote sensing image through the parallel frequency-direction sensitive convolutional layers to obtain a first output feature, including: Input the first feature map into the frequency-direction sensitive convolution layers set in parallel respectively for feature extraction; In each frequency-direction sensitive convolutional layer, first, the first feature map is subjected to wavelet transform through a wavelet transform layer to obtain more than two frequency sub-maps ; Input the frequency sub-graph into the corresponding first convolutional layer for feature extraction respectively to obtain more than two second feature maps ; Input two or more second feature maps into the inverse wavelet transform layer for inverse wavelet transform, and then input them into the second convolutional layer for processing to obtain the third feature map ; Add the third feature maps and perform average pooling to obtain the first output feature .
5. The lightweight object detection method for remote sensing images according to claim 1, characterized in that, The frequency-direction attention convolutional layer includes a parallel attention layer and a frequency-direction sensitive convolutional layer, and a third convolutional layer is connected to the output ends of the attention layer and the frequency-direction sensitive convolutional layer; and a residual connection is made between the input end and the output end of the frequency-direction attention convolutional layer.
6. The lightweight object detection method for remote sensing images according to claim 5, wherein Before performing feature extraction through two or more cascaded frequency-direction attention convolutional layers in each feature extraction layer, it further includes: performing downsampling on the first output feature to obtain a fourth feature map .
7. The lightweight object detection method for remote sensing images according to claim 6, wherein In each feature extraction layer, performing feature extraction through more than two cascaded frequency-direction attention convolutional layers, including: In the first frequency-direction attention convolutional layer, the fourth feature map is split into a first sub-feature map and a second sub-feature map ; Input the first sub-feature map into the attention layer for feature extraction to obtain an attention vector ; Input the second sub-feature map into the frequency-direction sensitive convolution layer for feature extraction to obtain the fifth feature map ; Multiply the attention vector with the fifth feature map and input the result into the third convolutional layer to output a sixth feature map ; Perform a residual operation on the fourth feature map and the sixth feature map to obtain the seventh feature map output by each frequency-direction attention convolutional layer ; Input the seventh feature map into the next frequency-direction attention convolution layer for processing to obtain the output features of the feature extraction layer; Use the output features of the previous feature extraction layer as the input features of the next feature extraction layer, and finally obtain the second output features .
8. A lightweight object detection device for remote sensing images, characterized in that, The device includes: A model construction module, configured to construct a lightweight object detection model, the backbone network of the lightweight object detection model includes an initialization module and more than two sequentially connected feature extraction layers; wherein, the initialization module includes more than two parallel frequency-direction sensitive convolutional layers, and the feature extraction layers include more than two cascaded frequency-direction attention convolutional layers; A first output feature extraction module, configured to input the obtained remote sensing image into the initialization module, and perform feature extraction on the remote sensing image through the parallel frequency-direction sensitive convolutional layers to obtain a first output feature; A second output feature extraction module, configured to input the first output feature into the sequentially connected feature extraction layers, in each feature extraction layer, perform feature extraction through more than two cascaded frequency-direction attention convolutional layers, and use the output feature of the previous feature extraction layer as the input feature of the next feature extraction layer, and finally obtain a second output feature; An object detection image output module, configured to process the second output feature through other modules of the lightweight object detection model to output an object detection image.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the lightweight object detection method for remote sensing images according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the lightweight object detection method for remote sensing images according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Method for estimating 3D posture of a human body combining densely connecting attention pyramid residual network and equidistance restriction
CN108710830A
Convolutional neural network panchromatic sharpening method based on wavelet layer
CN114663301A
Wavelet-based direction perception attention image rain removal method, system and equipment
CN114782263A
Remote sensing image typical ground feature extraction method based on spectrum enhancement and two-way coding
CN117152616A
Light-weight remote sensing image target detection method based on deep learning
CN118334313A