A wide remote sensing sparse target fast small sample detection method
By employing methods such as image slicing and sparse convolution, the accuracy and efficiency issues of sparse and small-sample target detection in remote sensing images were addressed, achieving highly efficient detection results.
Patent Information
- Application Number
- CN202310316727.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-29
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2043-03-29
AI Technical Summary
Existing technologies struggle to meet the required accuracy for detecting sparse and small-sample targets in remote sensing images, resulting in significant waste of computer resources and low detection efficiency.
The remote sensing image is sliced, and features are extracted and aggregated using a trained feature processing module. Combined with sparse convolution, the resolution is gradually adjusted until the preset conditions are met, and then the target is detected.
It improves detection accuracy, reduces computational load, saves computer resources, and increases detection efficiency.
Smart Images

Figure CN117036229B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the field of remote sensing image processing, and particularly relates to a fast small sample detection method for sparse targets in wide remote sensing, an electronic device and a storage medium. BACKGROUND
[0002] With the continuous development of satellite remote sensing and other earth observation technologies, the number and quality of available remote sensing images have improved.
[0003] However, in the process of implementing the concept of the present disclosure, the inventors have found that in the case of using a traditional detection network to detect targets, for targets with a sparse distribution specific to a remote sensing scene and small sample size targets, the detection accuracy is difficult to meet the requirements, and detecting the entire image excessively wastes computer resources and has low detection efficiency. SUMMARY
[0004] In view of the above problems, the present disclosure provides a fast small sample detection method for sparse targets in wide remote sensing and an electronic device.
[0005] According to a first aspect of the present disclosure, a fast small sample detection method for sparse targets in wide remote sensing is provided, comprising: performing slice processing on a remote sensing image to obtain a plurality of slice remote sensing images, wherein the remote sensing image comprises a target object; processing the plurality of slice remote sensing images using a trained first feature processing module to obtain a first intermediate object region feature map corresponding to the target object; performing sparse convolution on the first intermediate object region feature map using a trained second feature processing module to obtain an object region feature map corresponding to the target object, wherein the resolution of the object region feature map satisfies a first preset condition; and performing target detection on the object region feature map to obtain a detection result corresponding to the target object.
[0006] According to an embodiment of the present disclosure, the sparse convolution of the first intermediate object region feature map by the trained second feature processing module to obtain the object region feature map corresponding to the target object comprises repeatedly performing the following operations until the resolution of the object region feature map satisfies the first preset condition: in the case where the resolution of the first intermediate object region feature map does not satisfy the first preset condition, processing the first intermediate object region feature map using the trained second feature processing module to obtain a new first intermediate object region feature map; and determining the resolution of the new first intermediate object region feature map.
[0007] According to an embodiment of the present disclosure, the trained second feature processing module is used to process the first intermediate object region feature map to obtain a new first intermediate object region feature map, including: predicting a target object region from the first intermediate object region feature map; improving the resolution of the first intermediate object region feature map to obtain a second intermediate object region feature map; and performing sparse convolution on the second intermediate object region feature map according to the target object region to obtain the new first intermediate object region feature map.
[0008] According to an embodiment of the present disclosure, the trained first feature processing module includes a trained feature extraction unit, a feature aggregation unit, and a feature reweighting unit; the trained first feature processing module is used to process a plurality of slice remote sensing images to obtain a first intermediate object region feature map corresponding to a target object, including: using the feature extraction unit to process the plurality of slice remote sensing images to obtain a plurality of first remote sensing image feature maps corresponding to the plurality of slice remote sensing images respectively, wherein the resolutions of the plurality of first remote sensing image feature maps corresponding to the slice remote sensing images are different from each other; using the feature aggregation unit to process the plurality of first remote sensing image feature maps corresponding to the plurality of slice remote sensing images respectively to obtain a plurality of second remote sensing image feature maps corresponding to the plurality of slice remote sensing images respectively, wherein the resolutions of the plurality of second remote sensing image feature maps corresponding to the slice remote sensing images are different from each other; using the feature reweighting unit to process the plurality of second remote sensing image feature maps corresponding to the plurality of slice remote sensing images respectively to obtain a plurality of third remote sensing image feature maps corresponding to the plurality of slice remote sensing images respectively, wherein the resolutions of the plurality of third remote sensing image feature maps corresponding to the slice remote sensing images are different from each other; and obtaining the first intermediate object region feature map corresponding to the target object according to the plurality of third remote sensing image feature maps corresponding to the plurality of slice remote sensing images respectively.
[0009] According to an embodiment of the present disclosure, the first remote sensing image feature map includes a plurality of object region features; the feature aggregation unit is used to process the plurality of first remote sensing image feature maps corresponding to the plurality of slice remote sensing images respectively, including: for a first remote sensing image feature map in the plurality of first remote sensing image feature maps corresponding to the plurality of slice remote sensing images respectively, for an object feature in the plurality of object features, determining a similarity between the object feature and each of the plurality of object region features to obtain a plurality of similarities corresponding to the object feature, wherein the object feature corresponds to a target object in the remote sensing image; obtaining an aggregated object feature corresponding to the object feature according to the plurality of similarities corresponding to the object feature and the object feature; obtaining an aggregated object feature map corresponding to the first remote sensing image feature map according to the aggregated object features corresponding to the plurality of object features respectively; and using the aggregated object feature map corresponding to the first remote sensing image feature map to process the first remote sensing image feature map to obtain a second remote sensing image feature map corresponding to the first remote sensing image feature map.
[0010] According to an embodiment of the present disclosure, the remote sensing image is sliced to obtain a plurality of sliced remote sensing images, including: processing the remote sensing image by using a sliding window slicing method to obtain a plurality of intermediate sliced remote sensing images; and processing the plurality of intermediate sliced remote sensing images by using a normalization method to obtain the plurality of sliced remote sensing images.
[0011] According to an embodiment of the present disclosure, the trained first feature processing module and the trained second feature processing module are obtained by the following method: constructing a first sample set and a second sample set according to a sample remote sensing image and a second preset condition, wherein the second preset condition is determined according to a category of a sample object in the sample remote sensing image; training the first feature processing module and the second feature processing module by using the first sample set to obtain an intermediate first feature processing module and an intermediate second feature processing module; and training the intermediate first feature processing module and the intermediate second feature processing module by using a third sample set to obtain the trained first feature processing module and the trained second feature processing module, wherein the third sample set is obtained based on the first sample set and the second sample set.
[0012] According to an embodiment of the present disclosure, training the first feature processing module and the second feature processing module by using the first sample set to obtain the intermediate first feature processing module and the intermediate second feature processing module includes: obtaining a support set and a query set corresponding to the first sample set according to the first sample set, wherein the support set is used to train the first feature processing module and the second feature processing module, and the query set is used to verify the training effect of the intermediate first feature processing module and the intermediate second feature processing module; and training the first feature processing module and the second feature processing module by using the support set and the query set to obtain the intermediate first feature processing module and the intermediate second feature processing module.
[0013] The second aspect of the present disclosure provides an electronic device, including: one or more processors; a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors execute the above method.
[0014] The third aspect of the present disclosure also provides a computer-readable storage medium having stored executable instructions, which are executed by a processor to make the processor execute the above method.
[0015] By slicing the remote sensing image, and then processing the plurality of sliced remote sensing images by using the trained first feature processing module to obtain the intermediate object region feature map corresponding to the target object, only the object region can be detected, the detection accuracy is improved, and then the resolution of the first intermediate object region feature map is adjusted and sparse convolution is performed by using the trained second feature processing module to obtain the object region feature map with a resolution meeting the requirement, the calculation amount is reduced, the computer resources are saved, and the detection efficiency of the remote sensing image is improved. Attached Figure Description
[0016] The foregoing contents, as well as other objects, features, and advantages of this disclosure, will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:
[0017] Figure 1 This illustration schematically depicts an application scenario of a small-sample target detection method for remotely sensed images according to embodiments of the present disclosure.
[0018] Figure 2 A flowchart illustrating a method for detecting small-sample targets in remotely sensed images according to embodiments of the present disclosure is shown schematically.
[0019] Figure 3 A schematic diagram of a target detection network according to an embodiment of the present disclosure is shown;
[0020] Figure 4 A flowchart illustrating the training of the first feature processing module and the second feature processing module according to an embodiment of the present disclosure is shown schematically.
[0021] Figure 5 A flowchart illustrating a training method utilizing a support set and a query set according to an embodiment of the present disclosure is shown schematically.
[0022] Figure 6 A schematic diagram illustrating the structure of a small-sample target detection apparatus for remotely sensed images according to embodiments of the present disclosure is shown; and
[0023] Figure 7 A block diagram schematically illustrates an electronic device suitable for implementing a small-sample target detection method for remotely sensed images according to embodiments of the present disclosure. Detailed Implementation
[0024] The embodiments of the present disclosure will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the disclosure. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concepts of the present disclosure.
[0025] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0026] All terms used herein, including technical and scientific terms, have the meanings as commonly understood by one of ordinary skill in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning that is consistent with the context of the specification, and should not be interpreted in an idealized or overly formal way.
[0027] In the case of using expressions such as "at least one of A, B, and C, etc.", it should generally be interpreted that the meaning of the expression is at least one of A, B, or C (for example, "a system having at least one of A, B, and C" should include but not be limited to a system having A alone, a system having B alone, a system having C alone, a system having both A and B, a system having both A and C, a system having both B and C, and / or a system having A, B, and C, etc.).
[0028] In the technical solutions of the present disclosure, the collection, storage, use, processing, transmission, provision, disclosure, and application of user personal information comply with relevant laws and regulations, necessary security measures are taken, and do not violate public order and good customs.
[0029] In the technical solutions of the present disclosure, the acquisition, collection, storage, use, processing, transmission, provision, disclosure, and application of data comply with relevant laws and regulations, necessary security measures are taken, and do not violate public order and good customs.
[0030] Since the deep learning method has good recognition rate and generalization, the deep learning method is commonly used for training detection network in the field of target detection of remote sensing images. However, the inventors found that the detection network trained by the deep learning method is difficult to meet the detection requirements.
[0031] Due to some characteristics of remote sensing images different from natural scene images, there are the following problems in detecting targets in remote sensing images.
[0032] The detection method of deep learning is limited by the memory occupied by the detection network, and has requirements for the size of the input image. Since there are many large scenes in remote sensing images, the steps for target detection of remote sensing images are relatively cumbersome. Second, some remote sensing targets are relatively rare, and data is difficult to obtain, so there are only a small number of labeled samples, which does not meet the requirement of deep learning for a large number of samples for training. Third, due to the wide width of remote sensing images, the objects in remote sensing images are distributed sparsely, such as stadiums, distant ships, etc. Detecting the entire remote sensing image will consume a lot of computing power.
[0033] In the related art, detecting large-area remote sensing images usually consumes a large amount of computing power and has low detection efficiency, and if only a single resolution remote sensing image is detected, the accuracy of the detection result is difficult to meet the demand. In order to solve the above problems, the inventors find that a low-resolution feature map can roughly indicate the target position, and a high-resolution feature map can accurately regress the target boundary, and by detecting only the key positions on the high-resolution feature map, the effect of saving computing power and improving detection efficiency can be achieved.
[0034] Therefore, embodiments of the present disclosure provide a fast small sample detection method for sparse targets in wide remote sensing, which can quickly complete the detection of remote sensing images under the condition of a small amount of samples, avoiding the problems of insufficient sample quantity and wasting computing power by detecting all feature maps, and improving detection accuracy and speed.
[0035] Specifically, embodiments of the present disclosure provide a fast small sample detection method for sparse targets in wide remote sensing, which does not need to directly detect the image with the highest resolution, but only needs to adjust the resolution and then process the image with the resolution meeting the demand, thereby avoiding wasting computing power by directly detecting the remote sensing image with the highest resolution and improving detection efficiency.
[0036] Specifically, embodiments of the present disclosure provide a fast small sample detection method for sparse targets in wide remote sensing, which includes: performing slice processing on a remote sensing image to obtain a plurality of slice remote sensing images, wherein the remote sensing image includes a target object; processing the plurality of slice remote sensing images by using a trained first feature processing module to obtain a first intermediate object region feature map corresponding to the target object; performing sparse convolution on the first intermediate object region feature map by using a trained second feature processing module to obtain an object region feature map corresponding to the target object, wherein the resolution of the object region feature map meets a first preset condition; and performing target detection on the object region feature map to obtain a detection result corresponding to the target object.
[0037] Figure 1 An application scenario diagram of a small sample target detection method for a remote sensing image according to an embodiment of the present disclosure is schematically shown.
[0038] As shown in Figure 1 The application scenario 100 according to this embodiment can include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 can include various connection types, such as wired, wireless communication links, or optical fiber cables, etc.
[0039] The user can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).
[0040] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with display screens and supporting web browsing, including but not limited to smartphones, tablet computers, laptop computers, desktop computers, etc.
[0041] The server 105 can be a server providing various services, such as a background management server providing support for websites browsed by the user using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (only as examples). The background management server can analyze and process received user requests, etc., and feed back the processing results (such as web pages, information, or data, etc. obtained or generated according to user requests) to the terminal device.
[0042] It should be noted that the small sample target detection method for remote sensing images provided by the embodiments of the present disclosure can generally be executed by the server 105. Correspondingly, the small sample target detection device for remote sensing images provided by the embodiments of the present disclosure can generally be arranged in the server 105. The small sample target detection method for remote sensing images provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105. Correspondingly, the small sample target detection device for remote sensing images provided by the embodiments of the present disclosure can also be arranged in a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105.
[0043] It should be understood that Figure 1 The number of terminal devices, networks, and servers in the above-mentioned scenario is only illustrative. According to the needs of implementation, there can be any number of terminal devices, networks, and servers.
[0044] The small sample target detection method for remote sensing images provided by the embodiments of the present disclosure will be described in detail below based on the scenario described above. Figure 1 Figures 2-5
[0045] Figure 2 A flowchart of a small sample target detection method of a remote sensing image is shown schematically according to an embodiment of the present disclosure.
[0046] As shown in Figure 2 The small sample target detection method of the remote sensing image of this embodiment includes operations S210-S240.
[0047] In operation S210, the remote sensing image is sliced to obtain a plurality of sliced remote sensing images, wherein the remote sensing image includes a target object.
[0048] According to an embodiment of the present disclosure, for example, the remote sensing image can include a target object and a background. The target object and the background can be divided according to categories. The target object can be an object of a target category. The background can be an object of a category other than the target category. For example, the target object can be an airplane, and the background can be vehicles and buildings located in the same remote sensing image, etc.
[0049] According to an embodiment of the present disclosure, the remote sensing image is sliced to obtain a plurality of sliced remote sensing images, for example, the remote sensing image can be sliced using a sliding window slicing method to obtain a plurality of sliced remote sensing images; also for example, the remote sensing image can be sliced using a sliding window slicing method first, and then the sliced remote sensing image is processed using a normalization method to obtain a plurality of sliced remote sensing images.
[0050] According to an embodiment of the present disclosure, the sliced remote sensing image can be input into a trained target detection network, which can be composed of a trained first feature processing module and a trained second feature processing module. By inputting the sliced remote sensing image into the trained target detection network, the sliced remote sensing image can be processed using the trained first feature processing module and the trained second feature processing module.
[0051] According to an embodiment of the present disclosure, since directly detecting the target object in the remote sensing image will consume a large amount of computer resources, the remote sensing image can be sliced, and then a plurality of sliced remote sensing images are processed and target detection is performed.
[0052] In operation S220, the plurality of sliced remote sensing images are processed using the trained first feature processing module to obtain a first intermediate object region feature map corresponding to the target object.
[0053] According to an embodiment of the present disclosure, the trained first feature processing module can be used to strengthen the object region features in the slice remote sensing image. The object region can be a region where an object in the remote sensing image is located, and can include a target object region and a background object region. The target object region can be a region where a target object is located. For example, the trained first feature processing module can be used to extract features in the remote sensing image, and aggregate the extracted features to strengthen the object region features in the remote sensing image. For another example, the trained first feature processing module can be used to extract object region features in a plurality of slice remote sensing images, aggregate the extracted features, and reweight the aggregated features to strengthen the object region features to obtain a first intermediate object region feature map.
[0054] By using the trained first feature processing module to process a plurality of slice remote sensing images, the object region features can be strengthened, and the detection efficiency and the accuracy of the detection result can be improved.
[0055] According to an embodiment of the present disclosure, for example, the trained first feature processing module can output a plurality of first intermediate object region feature maps with different resolutions from low to high according to each slice remote sensing image, which can be 8 times, 16 times, 32 times and 64 times of the resolution of the slice remote sensing image, respectively. For another example, the trained first feature processing module can output a first intermediate object region feature map with a single resolution, which is not limited by the present disclosure.
[0056] In operation S230, the trained second feature processing module is used to perform sparse convolution on the first intermediate object region feature map to obtain an object region feature map corresponding to the target object, wherein the resolution of the object region feature map satisfies a first preset condition.
[0057] According to an embodiment of the present disclosure, the first preset condition can correspond to the resolution of the first intermediate object region feature map. For example, the first preset condition can be a preset resolution, which can be a resolution corresponding to 64 times of the resolution of the slice remote sensing image, or a resolution corresponding to 32 times of the resolution of the slice remote sensing image. In the case where the resolution of the first intermediate object region feature map satisfies the preset resolution, the object region feature map can be obtained.
[0058] According to an embodiment of the present disclosure, the object region feature map can be a feature map whose resolution of the target object region meets the requirements, and can include features of the target object region whose resolution meets the requirements. For example, in the case where the resolution of the object region feature map satisfies the first preset condition, it can be determined that the resolution of the target object region in the object region feature map meets the requirements.
[0059] According to an embodiment of the present disclosure, for example, the first intermediate object region feature map can be sparsely convoluted by using the trained second feature processing module, and then the resolution of the object region can be improved, and the sparsely convolution can be performed on the object region with improved resolution, and the above process can be repeated until the resolution meets the first preset condition, and the object region feature map is obtained.
[0060] For example, the first intermediate object region feature map can be sparsely convoluted by using the trained second feature processing module, and then the resolution of the first intermediate object region feature map can be improved, and the sparsely convolution can be performed on the first intermediate object region feature map with improved resolution, and the above process can be repeated until the resolution meets the first preset condition, and the object region feature map is obtained. For example, the first intermediate object region feature map with different resolutions can be processed in order of low to high resolution by using the trained second feature processing module. For example, the first intermediate object region feature map with low resolution can be processed by using the second feature processing module to obtain the target object region corresponding to the first intermediate object region feature map with low resolution; the second feature processing module can be used to further process the target object region obtained from the first intermediate object region feature map with higher resolution in the target object region corresponding to the first intermediate object region feature map with low resolution, to obtain the target object region corresponding to the first intermediate object region feature map with higher resolution. The above process can be repeated until the resolution of the first intermediate object region feature map meets the first preset condition, and then the object region feature map can be obtained according to the target object region corresponding to the first intermediate object region feature map with resolution meeting the first preset condition.
[0061] Since the target object region with different resolutions may include different degrees of background region, the first intermediate object region feature map can be sparsely convoluted in order of low to high resolution to sequentially process the object region. In the case where the resolution meets the first preset condition, the target object can be detected, which can improve the detection efficiency and the accuracy of the detection result.
[0062] In operation S240, the object region feature map is detected to obtain a detection result corresponding to the target object.
[0063] According to an embodiment of the present disclosure, for example, the detection result can include the category and position of the target object.
[0064] According to an embodiment of the present disclosure, by slicing the remote sensing image, and then processing the plurality of sliced remote sensing images by using the trained first feature processing module, the intermediate object region feature map corresponding to the target object is obtained, so that only the object region can be detected, and the detection accuracy is improved. By using the trained second feature processing module, the resolution of the first intermediate object region feature map is adjusted and sparse convolution is performed to obtain the object region feature map with the required resolution, thereby reducing the calculation amount, saving computer resources, and improving the detection efficiency of the remote sensing image.
[0065] According to an embodiment of the present disclosure, the first intermediate object region feature map is processed by using the trained second feature processing module to obtain the object region feature map corresponding to the target object, wherein the resolution of the object region feature map satisfies a first preset condition. In the case that the resolution of the first intermediate object region feature map does not satisfy the first preset condition, the first intermediate object region feature map is processed by using the trained second feature processing module to obtain a new first intermediate object region feature map, and the resolution of the new first intermediate object region feature map is determined.
[0066] According to an embodiment of the present disclosure, for example, the trained second feature processing module can be used to predict each object region in the first intermediate object region feature map, and then classify each object region to filter the target object region from the features and determine the target object region. In the case that the target object region is filtered out, in the first intermediate object region feature map with a higher resolution, further filtering is performed according to the filtered target object region. The above operation is repeatedly performed until the resolution of the first intermediate object region feature map satisfies the first preset condition, and then the object region feature map can be obtained according to the target object region in the first intermediate object region feature map that satisfies the first preset condition. Since in the process of predicting the target object region in the first intermediate object region feature map from low resolution to high resolution, the target object region predicted in the last time is predicted each time, the efficiency can be improved.
[0067] According to an embodiment of the present disclosure, since the second feature processing module is repeatedly used to process the object region with different resolutions until the resolution satisfies the first preset condition, the detection efficiency is improved.
[0068] According to an embodiment of the present disclosure, the first intermediate object region feature map is processed by using the trained second feature processing module to obtain a new first intermediate object region feature map, including: predicting a target object region from the first intermediate object region feature map; improving the resolution of the first intermediate object region feature map to obtain a second intermediate object region feature map; and performing sparse convolution on the second intermediate object region feature map according to the target object region to obtain the new first intermediate object region feature map.
[0069] According to an embodiment of the present disclosure, for example, the second feature processing module can include a low-resolution probe head, a medium-resolution probe head, and a high-resolution probe head. The probe head can be a lightweight probe head. Each probe head can be used to predict a plurality of object regions in the first intermediate object region feature map, classify objects corresponding to the predicted object regions, filter out target object regions corresponding to target objects from the predicted object regions according to the categories, improve the resolution of the target object regions, further predict new target object regions from the target object regions with improved resolution, and perform sparse convolution on the new target object regions.
[0070] According to an embodiment of the present disclosure, the target object region is predicted from the first intermediate object region feature map. For example, it can include determining a plurality of object regions in the first intermediate object region feature map, and determining the target object region from the plurality of object regions.
[0071] For example, determining the target object region from the plurality of object regions can include determining the categories of the objects corresponding to the plurality of object regions, and filtering out the target object region according to the target category from the plurality of object regions. For example, the target category can be a category that needs to be detected, and can include vehicles, ships, trains, and the like.
[0072] According to an embodiment of the present disclosure, the resolution of the first intermediate object region feature map is improved to obtain a second intermediate object region feature map. For example, it can be to improve the resolution of the plurality of object regions included in the first intermediate object region feature map to obtain a second intermediate object region feature map with higher resolution. Also for example, it can be to improve the resolution of the predicted target object region in the first intermediate object region feature map to obtain a second intermediate object region feature map with higher resolution.
[0073] According to an embodiment of the present disclosure, the second intermediate object region feature map is sparsely convolved according to the target object region to obtain a new first intermediate object region feature map. For example, it can include determining the corresponding target object region in the high-resolution second intermediate object region feature map according to the target object region in the first intermediate object region feature map, and performing sparse convolution on the target object region in the second intermediate object region feature map to obtain a new first intermediate object region feature map.
[0074] According to an embodiment of the present disclosure, by predicting the target object region, the resolution of the first intermediate object region feature map is improved, and then according to the predicted target object region, only the target object region in the high-resolution first intermediate object region feature map is sparsely convolved, which reduces the consumption of computer resources and improves the detection efficiency.
[0075] According to an embodiment of the present disclosure, the trained first feature processing module comprises a trained feature extraction unit, a feature aggregation unit and a feature re-weighting unit. Wherein, the trained first feature processing module is used to process the plurality of sliced remote sensing images to obtain the first intermediate object region feature map corresponding to the target object, comprising: the feature extraction unit is used to process the plurality of sliced remote sensing images to obtain a plurality of first remote sensing image feature maps corresponding to the plurality of sliced remote sensing images respectively, wherein the resolutions of the plurality of first remote sensing image feature maps corresponding to the sliced remote sensing images are different from each other; the feature aggregation unit is used to process the plurality of first remote sensing image feature maps corresponding to the plurality of sliced remote sensing images respectively to obtain a plurality of second remote sensing image feature maps corresponding to the plurality of sliced remote sensing images respectively, wherein the resolutions of the plurality of second remote sensing image feature maps corresponding to the sliced remote sensing images are different from each other; the feature re-weighting unit is used to process the plurality of second remote sensing image feature maps corresponding to the plurality of sliced remote sensing images respectively to obtain a plurality of third remote sensing image feature maps corresponding to the plurality of sliced remote sensing images respectively, wherein the resolutions of the plurality of third remote sensing image feature maps corresponding to the sliced remote sensing images are different from each other; and the first intermediate object region feature map corresponding to the target object is obtained according to the plurality of third remote sensing image feature maps corresponding to the plurality of sliced remote sensing images respectively.
[0076] According to an embodiment of the present disclosure, the feature extraction unit can be used to extract the features of the sliced remote sensing image, for example, the feature extraction unit can be a ResNet-50 feature extraction backbone network, but is not limited to the ResNet-50 feature extraction backbone network. The ResNet-50 feature extraction backbone network can be used to extract the object region features in the sliced remote sensing image to obtain a plurality of first remote sensing image feature maps corresponding to the plurality of sliced remote sensing images. For example, each first remote sensing image feature map can include the corresponding object region in the plurality of sliced remote sensing images.
[0077] According to an embodiment of the present disclosure, the resolutions of the plurality of first remote sensing image feature maps corresponding to the sliced remote sensing images are different from each other, for example, in the case of obtaining 4 first remote sensing image feature maps, the resolutions of the 4 first remote sensing image feature maps can be 8 times, 16 times, 32 times and 64 times of the resolution of the sliced remote sensing image respectively.
[0078] According to an embodiment of the present disclosure, the feature aggregation unit can be used to aggregate the respective object region features in the first remote sensing image feature map. For example, for each first remote sensing image feature map in the plurality of first remote sensing image feature maps, the similarity between the respective object region features in the first remote sensing image feature map can be calculated; and the object region features are aggregated according to the similarity.
[0079] For example, a plurality of object features corresponding to the target object in the remote sensing image can be pre-stored, and the object region features of the first remote sensing image feature map can be aggregated by using the plurality of object features to strengthen the object region features in the first remote sensing image feature map.
[0080] For example, for each of the plurality of first remote sensing image feature maps, the similarity between each of the plurality of object features and each of the object region features included in the first remote sensing image feature map can be determined. For example, the plurality of object features can include two object features A1 and A2, and the first remote sensing image feature map can include two object region features B1 and B2. Thus, the similarity C1 between A1 and B1, the similarity C2 between A1 and B2, the similarity C3 between A2 and B1, and the similarity C4 between A2 and B2 can be determined. The size relationship between C1, C2, C3, and C4 can be C1>C2>C3>C4.
[0081] The plurality of object features can be processed by using the similarity to obtain an aggregated object feature map. For example, for each object feature, the similarity can be multiplied by the feature value of the corresponding object feature to obtain an aggregated object feature. The feature map corresponding to the first remote sensing image feature map can be constructed according to a plurality of aggregated object features corresponding to the object features to obtain the aggregated object feature map, so as to aggregate the first remote sensing image feature map by using the aggregated object feature map.
[0082] For example, the similarity C1 can be multiplied by the feature value of A1 to obtain an aggregated object feature D1 corresponding to A1 and B1; the similarity C2 can be multiplied by the feature value of A1 to obtain an aggregated object feature D2 corresponding to A1 and B2; the similarity C3 can be multiplied by the feature value of A2 to obtain an aggregated object feature D3 corresponding to A2 and B1; and the similarity C4 can be multiplied by the feature value of A2 to obtain an aggregated object feature D4 corresponding to A2 and B2. Thus, the aggregated object feature map can be constructed according to D1, D2, D3, and D4.
[0083] The object region features in the first remote sensing image feature map can be aggregated by using the aggregated object feature map. For example, D1 and D3 can be aggregated into the object region corresponding to B1 in the first remote sensing image feature map; and D2 and D4 can be aggregated into the object region corresponding to B2 in the first remote sensing image feature map, so as to strengthen the object region features B1 and B2 in the first remote sensing image feature map.
[0084] The second remote sensing image feature map can be obtained by aggregating the aggregated object feature map and the first remote sensing image feature map. Thus, a plurality of second remote sensing image feature maps corresponding to the resolution of the first image feature map can be obtained.
[0085] According to an embodiment of the present disclosure, the resolutions of the plurality of second remote sensing image feature maps corresponding to the slice remote sensing image are different from each other, for example, in the case of obtaining 4 second remote sensing image feature maps, the resolutions of the 4 second remote sensing image feature maps can be 8 times, 16 times, 32 times and 64 times of the resolution of the slice remote sensing image, respectively.
[0086] According to an embodiment of the present disclosure, the feature reweighting unit can be configured to reweight each object region feature in the second remote sensing image feature map, for example, each object region feature in the second remote sensing image feature map can be reweighted according to the pre-stored weight to strengthen the object region feature.
[0087] According to an embodiment of the present disclosure, the resolutions of the plurality of third remote sensing image feature maps corresponding to the slice remote sensing image are different from each other, for example, in the case of obtaining 4 third remote sensing image feature maps, the resolutions of the 4 third remote sensing image feature maps can be 8 times, 16 times, 32 times and 64 times of the resolution of the slice remote sensing image, respectively.
[0088] According to an embodiment of the present disclosure, the feature reweighting unit is configured to process the plurality of second remote sensing image feature maps corresponding to the plurality of slice remote sensing images respectively to obtain a plurality of third remote sensing image feature maps corresponding to the plurality of slice remote sensing images respectively, for example, the channel weight vector can be multiplied with all object region features included in the second remote sensing image feature map to reweight the object region features included in the second remote sensing image feature map to obtain the third remote sensing image feature map. Thus, the third remote sensing image feature map corresponding to the second remote sensing image feature map can be obtained. The channel weight vector can be obtained in the training process of the feature reweighting unit. Thus, the object region feature can be strengthened by using the channel weight vector, and the accuracy of the detection result is improved.
[0089] According to an embodiment of the present disclosure, the first intermediate object region feature map corresponding to the target object is obtained according to the plurality of third remote sensing image feature maps corresponding to the plurality of slice remote sensing images respectively, for example, the first intermediate object region feature map including the reweighted object region feature can be obtained according to the reweighted object region feature in the third remote sensing image feature map.
[0090] According to an embodiment of the present disclosure, for example, the feature aggregation unit, the feature reweighting unit and the sparse convolution unit described above can be obtained according to a Retinanet detection network. A lightweight target detection network can be constituted according to a ResNet-50 feature extraction backbone network, a feature aggregation unit, a feature reweighting unit and a lightweight sparse convolution detection head.
[0091] The second feature processing module can include sparse convolution detection heads of different resolutions from low to high, to detect the first intermediate object region feature maps of different resolutions. For example, the second feature processing module can include a low-resolution detection head, a medium-resolution detection head and a high-resolution detection head, to detect the first intermediate object region feature map of low resolution, the object region feature map of medium resolution and the object region feature map of high resolution.
[0092] For example, each detection head can include a sparse convolution unit and a target region prediction unit. The target region prediction unit can be used to predict the object region in the first intermediate object region feature map, classify the predicted object region, and screen the target object region from the predicted object region; the sparse convolution unit can be used to perform sparse convolution according to the screened target object region. For example, the target region prediction unit can first perform the above-mentioned prediction process, screening process and classification process on the target region to obtain the target object region, and then the sparse convolution unit can be used to perform sparse convolution on the screened target object region.
[0093] For another example, the target region prediction unit can first perform the above-mentioned prediction process, screening process and classification process on the target region to obtain the target object region, and then the sparse convolution unit can be used to perform sparse convolution on the corresponding object region in the feature map of the intermediate object region of higher resolution according to the screened target object region.
[0094] Figure 3 A schematic diagram of a target detection network according to an embodiment of the present disclosure is schematically shown.
[0095] As shown in Figure 3 The ResNet-50 feature extraction backbone network 310, the feature aggregation unit 320, the feature reweighting unit 330 and the second feature processing module 340 described above can be combined to constitute a target detection network. The slice remote sensing image can be input into the target detection network to obtain the detection result described above.
[0096] Since the ResNet-50 feature extraction backbone network outputs first remote sensing image feature maps of multiple levels of resolution, in the related art, in order to save computing power, detection is usually performed only based on the first remote sensing image feature map of the highest resolution, so that the accuracy of the detection result meets the demand. However, by using the target detection method of the embodiments of the present disclosure, the first remote sensing image feature maps of multiple levels of resolution output by the ResNet-50 feature extraction backbone network can be fully utilized, and the second feature processing module is used to detect the target object region level by level according to the resolution, compared with the traditional method, the accuracy of the detection result is further improved, and the computing power is saved.
[0097] Therefore, processing the first intermediate object feature map by using the trained second feature processing module can process the object region in the first intermediate object region feature map according to different resolutions to obtain an object region feature map. By using the trained feature extraction unit, feature aggregation unit, and feature reweighting unit, third remote sensing image feature maps of different resolutions are obtained, the object region features are strengthened, and the trained second feature processing module is used to process the third remote sensing image feature maps of different resolutions level by level according to the resolution, so that the accuracy of the detection result meets the demand to the greatest extent, and the computing power for detecting the remote sensing image is saved.
[0098] According to the embodiments of the present disclosure, the first remote sensing image feature map includes a plurality of object region features; processing the plurality of first remote sensing image feature maps corresponding to the plurality of sliced remote sensing images by using the feature aggregation unit includes: for a first remote sensing image feature map in the plurality of first remote sensing image feature maps corresponding to the plurality of sliced remote sensing images, for an object feature in the plurality of object features, determining a similarity between the object feature and each of the plurality of object region features to obtain a plurality of similarities corresponding to the object feature, wherein the object feature corresponds to a target object in the remote sensing image; obtaining an aggregated object feature corresponding to the object feature according to the plurality of similarities corresponding to the object feature and the object feature; obtaining an aggregated object feature map corresponding to the first remote sensing image feature map according to the aggregated object features corresponding to the plurality of object features; and processing the first remote sensing image feature map by using the aggregated object feature map corresponding to the first remote sensing image feature map to obtain a second remote sensing image feature map corresponding to the first remote sensing image feature map.
[0099] According to an embodiment of the present disclosure, similarity between each of the object feature and the plurality of object region features is determined, and a plurality of similarities corresponding to the object feature are obtained. For example, the object feature can be A, and the plurality of object region features can include object region feature B1, object region feature B2, object region feature B3, and object region feature B4, and similarity between each of the object feature A and the object region feature B1 is C1; similarity between each of the object feature A and the object region feature B2 is C2; similarity between each of the object feature A and the object region feature B3 is C3; and similarity between each of the object feature A and the object region feature B4 is C4.
[0100] According to an embodiment of the present disclosure, the aggregated object feature can be used to aggregate the first remote sensing image feature map to strengthen the object region feature in the first remote sensing image feature map. According to the plurality of similarities corresponding to the object feature and the object feature, an aggregated object feature corresponding to the object feature is obtained. For example, the similarity can be multiplied by the feature value of the object feature to obtain the aggregated object feature. For example, the feature value of the object feature A and the similarity C1 can be multiplied to obtain the aggregated object feature AC1; the feature value of the object feature A and the similarity C2 can be multiplied to obtain the aggregated object feature AC2; the feature value of the object feature A and the similarity C3 can be multiplied to obtain the aggregated object feature AC3; and the feature value of the object feature A and the similarity C4 can be multiplied to obtain the aggregated object feature AC4.
[0101] According to an embodiment of the present disclosure, according to the aggregated object features corresponding to each of the plurality of object features, an aggregated object feature map corresponding to the first remote sensing image feature map is obtained. For example, the aggregated object feature map can be composed of the aggregated object feature AC1, the aggregated object feature AC2, the aggregated object feature AC3, and the aggregated object feature AC4. The aggregated object feature AC1, the aggregated object feature AC2, the aggregated object feature AC3, and the aggregated object feature AC4 can be included in the aggregated object feature map.
[0102] According to an embodiment of the present disclosure, the first remote sensing image feature map is processed by using the aggregated object feature map corresponding to the first remote sensing image feature map, and a second remote sensing image feature map corresponding to the first remote sensing image feature map is obtained. For example, the aggregated object feature AC1 can be aggregated into the region corresponding to C1 in the corresponding first remote sensing image feature map; the aggregated object feature AC2 can be aggregated into the region corresponding to C2 in the corresponding first remote sensing image feature map; the aggregated object feature AC3 can be aggregated into the region corresponding to C1 in the corresponding first remote sensing image feature map; and the aggregated object feature AC4 can be aggregated into the region corresponding to C4 in the corresponding first remote sensing image feature map. Thus, the second remote sensing image feature map aggregated with AC1, AC2, AC3, and AC4 is obtained.
[0103] According to an embodiment of the present disclosure, each object region feature included in the plurality of first remote sensing image feature maps is corresponded to the plurality of object features, which not only realizes the aggregation of the first remote sensing image feature map, but also can associate the object region features included in the first remote sensing image feature map through the plurality of object features, avoid the problem of reduced detection quality caused by the features of the detected object being located in different slice remote sensing images due to the slicing processing, and further improve the detection quality.
[0104] According to an embodiment of the present disclosure, the slicing processing of the remote sensing image to obtain the plurality of slice remote sensing images comprises:
[0105] The remote sensing image is processed by using a sliding window slicing method to obtain a plurality of intermediate slice remote sensing images, and the plurality of intermediate slice remote sensing images are processed by using a normalization method to obtain the plurality of slice remote sensing images.
[0106] According to an embodiment of the present disclosure, the remote sensing image is divided into a plurality of slices by using the sliding window slicing method, and the intermediate slice remote sensing images obtained by slicing are normalized, which can reduce the occupation of computer resources in the processing process and improve the efficiency of detecting target objects in the remote sensing image.
[0107] According to an embodiment of the present disclosure, the small sample target detection method of the remote sensing image further comprises: processing the detection result by using a non-maximum suppression method to obtain a target detection result.
[0108] According to an embodiment of the present disclosure, since the target object region is labeled by using the detection frame in the process of processing the first intermediate object region feature map by using the trained second feature processing module, the detection result output by the trained second feature processing module has the detection frame, and the non-maximum suppression method can be used to filter the redundant detection frame.
[0109] According to an embodiment of the present disclosure, for the relatively rare detection target in the remote sensing image, the data is difficult to obtain, so only a small amount of samples are available, which is difficult to meet the requirement of training the detection network by using the deep learning method, and a large amount of training samples are required. Therefore, how to combine the characteristics of the remote sensing image, train a remote sensing target detection model by using a small amount of samples, and realize fast and accurate detection has important value and significance. Based on this, an embodiment of the present disclosure trains a model by using the following method to solve the above problems.
[0110] Figure 4 A flowchart for training the first feature processing module and the second feature processing module according to an embodiment of the present disclosure is schematically shown.
[0111] As Figure 4As shown, the trained first feature processing module and the trained second feature processing module are trained in the following manner. The training method comprises operations S410-S430.
[0112] In operation S410, a first sample set and a second sample set are constructed according to the sample remote sensing image and a second preset condition, wherein the second preset condition is determined according to the category of the sample object in the sample remote sensing image.
[0113] According to an embodiment of the present disclosure, for example, the second preset condition can be the number of samples corresponding to each category. A plurality of image samples can be obtained according to the sample remote sensing image. The category whose number of image samples satisfies the second preset condition can be determined as a base category, and the category whose number of samples does not satisfy the second preset condition can be determined as a new category. The first sample set can be constructed according to the base category, and the second sample set can be constructed according to the new category.
[0114] In operation S420, the first feature processing module and the second feature processing module are trained using the first sample set to obtain an intermediate first feature processing module and an intermediate second feature processing module.
[0115] According to an embodiment of the present disclosure, by training the first feature processing module and the second feature processing module using the first sample set whose number of samples of each category meets the requirement, the training effect can meet the requirement.
[0116] In operation S430, the intermediate first feature processing module and the intermediate second feature processing module are trained using a third sample set to obtain a trained first feature processing module and a trained second feature processing module, wherein the third sample set is obtained based on the first sample set and the second sample set.
[0117] According to an embodiment of the present disclosure, for example, the same number of samples can be selected from the first sample set and the second sample set to construct the third sample set. The intermediate first feature processing module and the intermediate second feature processing module are trained by the third sample set to reduce the tendency of the classification ability of the trained first feature processing module and the classification ability of the trained second feature processing module to the first sample set. The processing ability of the trained first feature processing module and the trained second feature processing module to the new category can be improved.
[0118] According to an embodiment of the present disclosure, for example, the sample remote sensing image used by the present disclosure can be selected from dior (a large-scale dataset for optical remote sensing image target detection), 15 categories of airplanes, baseball fields, basketball fields, bridges, chimneys, dams, ports, golf courses, ground tracks, skywalks, ships, storage tanks, tennis courts, vehicles, and windmills in dior can be taken as base categories, and a first sample set can be obtained according to all image samples included in the base categories. Airports, highway service areas, stadiums, railway stations, and highway toll stations can be taken as new categories, and a second sample set can be obtained according to all image samples included in the new categories. A third sample set can be constructed according to the first sample set and the second sample set, the third sample set including the same number of categories of the base categories and the new categories, and the number of image samples in each category can also be the same. For example, in the case that the second sample set includes 5 categories, 5 categories of image samples can be randomly selected from the first sample set, and then 20 image samples can be selected from each category of the first sample set and each category of the second sample set to construct the third sample set.
[0119] According to an embodiment of the present disclosure, for example, the first feature processing module and the second feature processing module can be trained by using the first sample set by means of meta-learning, and the intermediate first feature processing module and the intermediate second feature processing module can be adjusted by using the third sample set to obtain the trained first feature processing module and the trained second feature processing module, thereby realizing small sample training of the feature processing module and the second feature processing module.
[0120] According to an embodiment of the present disclosure, since the first sample set is used for training first, and then the intermediate first feature processing module and the intermediate second feature processing module are adjusted by using the third sample set by means of meta-learning, the processing capability of the trained module is prevented from excessively deviating to the base categories, the adaptability of the module to the new categories is improved, and the trained module can also have a processing capability meeting the requirements in the case of training for the new categories.
[0121] According to an embodiment of the present disclosure, the trained first feature processing module and the trained second feature processing module can be trained by using the second sample set to improve the processing capability of the trained first feature processing module and the trained second feature processing module to the objects of the new categories.
[0122] According to an embodiment of the present disclosure, for example, a detection frame can be labeled on a sample remote sensing image, and then the sample remote sensing image can be cropped according to the detection frame to obtain an image block including a target, which can be used as an image sample to construct a first sample set and a second sample set. By using the detection frame, the image block is labeled to be located in a region of the sample remote sensing image, which is beneficial to subsequent training of the first feature processing module and the second feature processing module. In order to make the image sample include certain information related to the target object, for example, object information around the target object, the image sample can include a part of the background outside the target frame. Each image sample can correspond to a binary mask generated according to the labeling information, which is used to indicate the range occupied by the real target object.
[0123] According to an embodiment of the present disclosure, the first feature processing module and the second feature processing module are trained by using the first sample set to obtain an intermediate first feature processing module and an intermediate second feature processing module, including:
[0124] According to the first sample set, a support set and a query set corresponding to the first sample set are obtained, wherein the support set is used to train the first feature processing module and the second feature processing module, and the query set is used to verify the training effect of the intermediate first feature processing module and the intermediate second feature processing module; the first feature processing module and the second feature processing module are trained by using the support set and the query set to obtain the intermediate first feature processing module and the intermediate second feature processing module.
[0125] According to an embodiment of the present disclosure, for example, the query set and the support set can be input into the first feature processing module and the second feature processing module at the same time to simulate the detection process while training. By detecting the target object in the query set, a detection result is obtained. Then, the training process is adaptively adjusted according to the detection result to make the training effect meet the requirements.
[0126] According to an embodiment of the present disclosure, for example, 20 image samples can be randomly selected from each category included in the first sample set to form the support set, and a required number of image samples can be randomly selected from each category corresponding to the first sample set to form the query set.
[0127] According to embodiments of this disclosure, for example, preprocessing and data augmentation can be performed on the support set and query set. Preprocessing may include image cropping and normalization. This could involve cropping the image to 800×800 pixels and then normalizing the cropped image using the RGB channel mean and variance statistically obtained from the first sample set. Data augmentation may include random flipping and multi-scale training. Each image is randomly flipped horizontally and vertically according to preset probabilities, and the annotation information is also flipped to expand the sample set. Multi-scale training may involve randomly adjusting the size of each image between 600 and 1000 pixels based on the input samples, improving the module's robustness to target detection with large scale variations.
[0128] According to embodiments of this disclosure, for example, k samples can be randomly selected from a first sample set according to each category to form a support set. Image samples from the query set can be... Image samples of the support set Inputting the ResNet-50 feature extraction backbone network yields feature maps of the first remote sensing image at different resolutions for the query set and support set, respectively. These four stages of feature maps output by the ResNet-50 feature extraction backbone network can be denoted as... , among which, respectively Corresponding to 1 / 8 of the original image size, Corresponding to 1 / 16 of the original image size, Corresponding to 1 / 32 of the original image size, This corresponds to 1 / 64 of the original image size.
[0129] The first remote sensing image feature map output by the ResNet-50 feature extraction backbone network in the final stage. For example, let's denote the first remote sensing image feature map corresponding to the query set as . The first remote sensing image feature map corresponding to the support set is The following uses the query set feature map to represent the first remote sensing image feature map corresponding to the query set, and the support set feature map to represent the first remote sensing image feature map corresponding to the support set. Since each category of the support set has k image samples, the support set feature map corresponding to the sample of the j-th category is... It can include ,in, This represents the feature map of the i-th image sample in the j-th category. It is the support set feature map obtained from the input corresponding image binary mask pair. Perform global mask pooling to obtain the support set feature map. The corresponding image features can be denoted as: The image feature matrix obtained from all support set patterns is P= The obtained image feature matrix P is compared with the query set feature map. Calculate cross attention.
[0130] The process of calculating cross-attention may include: calculating the image feature matrix P and the query set feature map. The similarity matrix S. ,in This represents matrix multiplication. Each element in S represents a feature map of the query set. The similarity between a location and a feature of the image feature matrix P is calculated, and Softmax is used as a normalization function to normalize each row of the result; then, the similarity in the similarity matrix S and the image feature matrix P are calculated, i.e. ) It can be calculated ; can With query set feature map The images are stitched together to obtain the aggregated second remote sensing image feature map. : =[ , ] Through the above process, the feature aggregation module can be trained to obtain an intermediate feature aggregation module. The image feature matrix P obtained through the above process can be stored. The features in the image feature matrix P can be mapped to the features of the multiple objects mentioned above, and used for aggregation of the first remote sensing image.
[0131] The aggregated second remote sensing image feature map The input is fed into the feature reweighting unit, which can consist of multiple convolutional and pooling layers. The channel weight vector is obtained based on the output of each channel from the final global average pooling. This channel weight vector can then be compared with the feature map of the second remote sensing image. By performing channel-by-channel multiplication, a reweighted third remote sensing image feature map is obtained. Through the above process, the feature reweighting unit can be trained, resulting in an intermediate feature reweighting module. The channel weight vectors obtained from the above process can be stored.
[0132] According to embodiments of this disclosure, for example, the reweighted third remote sensing image feature map can be... The input is fed into the second feature processing module to train it. Corresponding reweighted third-party remote sensing image feature maps at different resolutions , as the input of training the second feature processing module. On the basis of the detection head of the Retinanet detection network, a target region prediction branch parallel to the regression and classification is added, so that the second feature processing module can be obtained. By adding the target region prediction branch, the second feature processing module can predict the length, width and position of the detection box, and whether there is a target object in the detection box. In the input third remote sensing image feature map In the case of forward propagation in the second feature processing module, the low-resolution third remote sensing image feature map Starts classification regression, and predicts the third remote sensing image feature map The approximate region of the target object is represented by the foreground probability and the background probability, wherein the foreground probability can represent the probability corresponding to the target object. In the case of forward propagation from the third remote sensing image feature map The upper level resolution third remote sensing image feature map The corresponding position predicted as the foreground is sparse convoluted to classify, regress and predict the potential region of the target object, and the third remote sensing image feature map is processed from low resolution to high resolution. For example, the reweighted third remote sensing image feature map of different resolutions can be input into the corresponding detection head. The low-resolution third remote sensing image feature map can be input into the low-resolution detection head, the medium-resolution third remote sensing image feature map can be input into the medium-resolution detection head, and the high-resolution third remote sensing image feature map can be input into the high-resolution detection head, and then the low-resolution detection head, the medium-resolution detection head and the high-resolution detection head are used for processing in turn. Through the above process, the training of the second feature processing module can be completed, and the intermediate second feature processing module can be obtained.
[0133] According to embodiments of the present disclosure, for example, after the training process of the first feature processing module and the second feature processing module is completed, the non-maximum suppression method can be used to filter out the redundant detection boxes in the detection results output by the second feature processing module in the training process. Then, all the output detection results are integrated, and the first feature processing module and the second feature processing module are adjusted if the detection results do not meet the requirements.
[0134] Figure 5 A schematic diagram of training a target detection network according to an embodiment of the present disclosure is schematically shown.
[0135] As Figure 5As shown, the support set 501 and the query set 502 can be input into a shared feature extraction unit 503, and the support set corresponding image feature matrix and the query set corresponding first remote sensing image feature map of different resolutions are output. The image feature matrix and the first remote sensing image feature map of different resolutions are input into a feature aggregation unit 504, and a second remote sensing image feature map of different resolutions is output. The second remote sensing image feature map of different resolutions is input into a feature reweighting unit 505, and a third remote sensing image feature map of different resolutions is output. The third remote sensing image feature map of different resolutions is input into a low-resolution probe head 506, a medium-resolution probe head 507, and a high-resolution probe head 508 according to the resolutions. The low-resolution probe head 507, the medium-resolution probe head 507, and the high-resolution probe head 508 are used to process the predicted object region in the third remote sensing image feature map layer by layer according to different resolutions.
[0136] According to an embodiment of the present disclosure, for example, when training the first feature processing module and the second feature processing module on the support set and the query set, the loss function used can be , wherein is a Focal Loss classification function in a Retinanet detection network, is a bounding box regression function in the Retinanet detection network, is a cross-entropy loss function of the target region prediction branch. The label of this branch can be defined as follows: the label of this branch can be defined as a binary mask. Let the size of the predicted target of the i-th layer feature map be between , then the binary mask region label corresponding to all the detection boxes with a size less than is 1, and the rest is 0. That is, each layer of feature map is responsible for predicting the potential region of all targets with a size less than the prediction threshold of this layer. A person skilled in the art can set the values of and according to needs.
[0137] According to an embodiment of the present disclosure, training the first feature processing module and the second feature processing module using the support set and the query set to obtain the intermediate first feature processing module and the intermediate second feature processing module can include: training the first feature processing module and the second feature processing module using the support set and the query set to obtain target parameter information, wherein the target parameter information includes target parameters of the intermediate first feature processing module and target parameters of the intermediate second feature processing module; and obtaining the intermediate first feature processing module and the intermediate second feature processing module according to the target parameter information.
[0138] According to an embodiment of the present disclosure, by utilizing the support set and the query set to the first feature processing module and the second feature processing module, the target parameters corresponding to the target parameters of the first feature processing module and the second feature processing module are obtained, and the target parameters in the target parameter information can be used as the parameters of the first feature processing module and the second feature processing module, so that the intermediate first feature processing module and the intermediate second feature processing module are obtained.
[0139] Based on the above small sample target detection method of remote sensing image, the present disclosure further provides a small sample target detection device of remote sensing image. The following will be combined with Figure 6 The device will be described in detail.
[0140] Figure 6 The structure block diagram of the small sample target detection device of remote sensing image according to the embodiment of the present disclosure is schematically shown.
[0141] As Figure 6 shown, the small sample target detection device 600 of remote sensing image of the embodiment includes a slicing module 610, a first processing module 620, a second processing module 630 and a detection module 640.
[0142] The slicing module 610 is configured to perform slicing processing on the remote sensing image to obtain a plurality of sliced remote sensing images, wherein the remote sensing image includes a target object. In an embodiment, the slicing module 610 can be configured to perform the operation S210 described above, and details are not repeated here.
[0143] The first processing module 620 is configured to process the plurality of sliced remote sensing images by using the trained first feature processing module to obtain a first intermediate object region feature map corresponding to the target object. In an embodiment, the first feature processing module 620 can be configured to perform the operation S220 described above, and details are not repeated here.
[0144] The second processing module 630 is configured to perform sparse convolution on the first intermediate object region feature map by using the trained second feature processing module to obtain an object region feature map corresponding to the target object, wherein the resolution of the object region feature map satisfies a first preset condition. In an embodiment, the second feature processing module 630 can be configured to perform the operation S230 described above, and details are not repeated here.
[0145] The detection module 640 is configured to perform target detection on the object region feature map to obtain a detection result corresponding to the target object. In an embodiment, the detection module 640 can be configured to perform the operation S240 described above, and details are not repeated here.
[0146] According to an embodiment of the present disclosure, the second processing module 630 comprises a first processing submodule and a determination submodule. The first processing submodule is configured to, in a case where the resolution of the first intermediate object region feature map does not satisfy a first preset condition, process the first intermediate object region feature map by using the trained second feature processing module to obtain a new first intermediate object region feature map; and the determination submodule is configured to determine the resolution of the new first intermediate object region feature map.
[0147] According to an embodiment of the present disclosure, the first processing submodule comprises a prediction unit, an improvement unit and a convolution unit. The prediction unit is configured to predict a target object region from the first intermediate object region feature map; the improvement unit is configured to improve the resolution of the first intermediate object region feature map to obtain a second intermediate object region feature map; and the convolution unit is configured to perform sparse convolution on the second intermediate object region feature map according to the target object region to obtain the new first intermediate object region feature map.
[0148] According to an embodiment of the present disclosure, the first processing module 620 comprises a second processing submodule, a third processing submodule, a fourth processing submodule and an acquisition submodule. The second processing submodule is configured to process the plurality of slice remote sensing images by using the feature extraction unit to obtain a plurality of first remote sensing image feature maps corresponding to the plurality of slice remote sensing images respectively, wherein the resolutions of the plurality of first remote sensing image feature maps corresponding to the slice remote sensing images are different from each other; the third processing submodule is configured to process the plurality of first remote sensing image feature maps corresponding to the plurality of slice remote sensing images respectively by using the feature aggregation unit to obtain a plurality of second remote sensing image feature maps corresponding to the plurality of slice remote sensing images respectively, wherein the resolutions of the plurality of second remote sensing image feature maps corresponding to the slice remote sensing images are different from each other; the fourth processing submodule is configured to process the plurality of second remote sensing image feature maps corresponding to the plurality of slice remote sensing images respectively by using the feature reweighting unit to obtain a plurality of third remote sensing image feature maps corresponding to the plurality of slice remote sensing images respectively, wherein the resolutions of the plurality of third remote sensing image feature maps corresponding to the slice remote sensing images are different from each other; and the acquisition submodule is configured to obtain the first intermediate object region feature map corresponding to the target object according to the plurality of third remote sensing image feature maps corresponding to the plurality of slice remote sensing images respectively.
[0149] According to an embodiment of the present disclosure, the third processing submodule comprises a determination unit, a first acquisition unit, a second acquisition unit and a processing unit. The determination unit is configured to determine, for a first remote sensing image feature map in a plurality of first remote sensing image feature maps corresponding to a plurality of slice remote sensing images respectively, and for an object feature in a plurality of object features, a similarity between the object feature and each of a plurality of object region features, to obtain a plurality of similarities corresponding to the object feature, wherein the object feature corresponds to a target object in a remote sensing image. The first acquisition unit is configured to obtain, according to the plurality of similarities corresponding to the object feature and the object feature, an aggregated object feature corresponding to the object feature. The second acquisition unit is configured to obtain, according to an aggregated object feature corresponding to each of the plurality of object features, an aggregated object feature map corresponding to the first remote sensing image feature map. The processing unit is configured to process the first remote sensing image feature map using the aggregated object feature map corresponding to the first remote sensing image feature map to obtain a second remote sensing image feature map corresponding to the first remote sensing image feature map.
[0150] According to an embodiment of the present disclosure, the slicing module 610 comprises a slicing submodule and a fifth processing submodule. The slicing submodule is configured to process a remote sensing image using a sliding window slicing method to obtain a plurality of intermediate slice remote sensing images. The fifth processing submodule is configured to process the plurality of intermediate slice remote sensing images using a normalization method to obtain the plurality of slice remote sensing images.
[0151] According to an embodiment of the present disclosure, the small sample target detection device for remote sensing images further comprises a construction module, a first training module and a second training module. The construction module is configured to construct a first sample set and a second sample set according to a sample remote sensing image and a second preset condition, wherein the second preset condition is determined according to a category of a sample object in the sample remote sensing image. The first training module is configured to train the first feature processing module and the second feature processing module using the first sample set to obtain an intermediate first feature processing module and an intermediate second feature processing module. The second training module is configured to train the intermediate first feature processing module and the intermediate second feature processing module using a third sample set to obtain a trained first feature processing module and a trained second feature processing module, wherein the third sample set is obtained based on the first sample set and the second sample set.
[0152] According to an embodiment of the present disclosure, the first training module comprises an acquisition submodule and a training submodule. The acquisition submodule is configured to obtain, according to the first sample set, a support set and a query set corresponding to the first sample set, wherein the support set is used to train the first feature processing module and the second feature processing module, and the query set is used to verify the training effect of the intermediate first feature processing module and the intermediate second feature processing module. The training submodule is configured to train the first feature processing module and the second feature processing module using the support set and the query set to obtain the intermediate first feature processing module and the intermediate second feature processing module.
[0153] According to an embodiment of the present disclosure, the training sub-module comprises a training unit and an obtaining unit. The training unit is configured to train the first feature processing module and the second feature processing module by using the support set and the query set to obtain target parameter information, wherein the target parameter information comprises target parameters of the intermediate first feature processing module and target parameters of the intermediate second feature processing module; and the obtaining unit is configured to obtain the intermediate first feature processing module and the intermediate second feature processing module according to the target parameter information.
[0154] According to an embodiment of the present disclosure, any of the slice module 610, the first processing module 620, the second processing module 630 and the detection module 640 can be combined in one module, or any of them can be split into multiple modules. Alternatively, at least part of the function of one or more of these modules can be combined with at least part of the function of other modules, and implemented in one module. According to an embodiment of the present disclosure, at least one of the slice module 610, the first processing module 620, the second processing module 630 and the detection module 640 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on board, a system on package, an application specific integrated circuit (ASIC), or any other reasonable way of hardware or firmware that can be integrated or packaged, or any one of software, hardware and firmware or any appropriate combination of any of them. Alternatively, at least one of the slice module 610, the first processing module 620, the second processing module 630 and the detection module 640 can be at least partially implemented as a computer program module which can perform corresponding functions when it is run.
[0155] Figure 7 A block diagram of an electronic device suitable for implementing the small sample target detection method of remote sensing images according to an embodiment of the present disclosure is schematically shown.
[0156] As Figure 7 shown, the electronic device 700 according to an embodiment of the present disclosure comprises a processor 701 which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 702 or loaded from a storage portion 708 into a random access memory (RAM) 703. The processor 701 may, for example, comprise a general-purpose microprocessor (such as a CPU), an instruction set processor and / or a related chipset, and / or a special-purpose microprocessor (such as an application specific integrated circuit (ASIC)), etc. The processor 701 can also comprise an on-board memory for cache use. The processor 701 can comprise a single processing unit or multiple processing units for performing different actions of the method processes according to an embodiment of the present disclosure.
[0157] In the RAM 703, various programs and data required for the operation of the electronic device 700 are stored. The processor 701, the ROM 702, and the RAM 703 are connected to each other via the bus 704. The processor 701 performs various operations of the method flow according to the embodiments of the present disclosure by executing the programs in the ROM 702 and / or the RAM 703. It should be noted that the programs can also be stored in one or more memories other than the ROM 702 and the RAM 703. The processor 701 can also perform various operations of the method flow according to the embodiments of the present disclosure by executing the programs stored in the one or more memories.
[0158] According to an embodiment of the present disclosure, the electronic device 700 can further include an input / output (I / O) interface 705, which is also connected to the bus 704. The electronic device 700 can further include one or more of the following components connected to the I / O interface 705: an input portion 706 including a keyboard, a mouse, etc.; an output portion 707 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage portion 708 including a hard disk, etc.; and a communication portion 709 including a network interface card such as a LAN card, a modem, etc. The communication portion 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the I / O interface 705 as necessary. A removable medium 711 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is attached to the drive 710 as necessary, so that a computer program read out therefrom is installed in the storage portion 708 as necessary.
[0159] The present disclosure also provides a computer readable storage medium, which can be included in the device / apparatus / system described in the above embodiments; or can exist separately without being assembled into the device / apparatus / system. The above computer readable storage medium carries one or more programs, when the one or more programs are executed, the method according to the embodiments of the present disclosure is implemented.
[0160] According to an embodiment of the present disclosure, the computer readable storage medium can be a nonvolatile computer readable storage medium, for example, can include but is not limited to: a portable computer diskette, a hard disk, a random access memory (RAM), a read only memory (ROM), an erasable programmable read only memory (EPROM or flash memory), a portable compact disc read only memory (CD-ROM), an optical storage device, a magnetic storage device, or any appropriate combination thereof. In the present disclosure, the computer readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present disclosure, the computer readable storage medium can include the ROM 702 and / or the RAM 703 described above and / or one or more memories other than the ROM 702 and the RAM 703.
[0161] Embodiments of the present disclosure also include a computer program product, which includes a computer program containing program codes for executing the methods shown in the flowcharts. When the computer program product is run in a computer system, the program codes are used to make the computer system implement the small sample target detection method of the remote sensing image provided by the embodiments of the present disclosure.
[0162] The above functions defined in the system / device of the embodiments of the present disclosure are performed when the computer program is executed by the processor 701. According to an embodiment of the present disclosure, the system, device, module, unit, etc. described above can be implemented by computer program modules.
[0163] In one embodiment, the computer program can rely on tangible storage media such as optical storage media, magnetic storage media, etc. In another embodiment, the computer program can also be transmitted, distributed, downloaded and installed in the form of signals on network media. The program codes contained in the computer program can be transmitted by any appropriate network media, including but not limited to wireless, wired, etc., or any appropriate combination thereof.
[0164] In such an embodiment, the computer program can be downloaded and installed from the network by the communication part 709, and / or installed from the detachable medium 711. When the computer program is executed by the processor 701, the above functions defined in the system of the embodiments of the present disclosure are performed. According to an embodiment of the present disclosure, the system, device, apparatus, module, unit, etc. described above can be implemented by computer program modules.
[0165] According to embodiments of this disclosure, program code for executing the computer programs provided in embodiments of this disclosure can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C", or similar programming languages. The program code can execute entirely on a user's computing device, partially on a user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0166] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0167] Those skilled in the art will understand that the features described in the various embodiments and / or claims of this disclosure can be combined or combined in various ways, even if such combinations or combinations are not explicitly described in this disclosure. In particular, the features described in the various embodiments and / or claims of this disclosure can be combined or combined in various ways without departing from the spirit and teachings of this disclosure. All such combinations and / or combinations fall within the scope of this disclosure.
[0168] The above describes embodiments of the present disclosure. However, these embodiments are merely for illustrative purposes, and are not intended to limit the scope of the present disclosure. Although each embodiment is described above separately, this does not mean that the measures in each embodiment cannot be used advantageously in combination. The scope of the present disclosure is defined by the appended claims and their equivalents. Those skilled in the art can make various substitutions and modifications without departing from the scope of the present disclosure, and these substitutions and modifications should all fall within the scope of the present disclosure.
Claims
1. A fast small-sample detection method for sparse targets in wide-swath remote sensing, comprising: The remote sensing image is sliced to obtain multiple sliced remote sensing images, wherein the remote sensing images include the target object; The trained first feature processing module is used to process the multiple sliced remote sensing images to obtain a first intermediate object region feature map corresponding to the target object; The first intermediate object region feature map is sparsely convolved using a trained second feature processing module to obtain an object region feature map corresponding to the target object, wherein the resolution of the object region feature map satisfies a first preset condition. Target detection is performed on the feature map of the object region to obtain the detection result corresponding to the target object; The trained first feature processing module includes a trained feature extraction unit, a feature aggregation unit, and a feature reweighting unit. The step of processing the multiple sliced remote sensing images using a trained first feature processing module to obtain a first intermediate object region feature map corresponding to the target object includes: The feature extraction unit processes the plurality of sliced remote sensing images to obtain a plurality of first remote sensing image feature maps corresponding to each of the plurality of sliced remote sensing images, wherein the resolutions of the plurality of first remote sensing image feature maps corresponding to the sliced remote sensing images are different from each other; The feature aggregation unit processes multiple first remote sensing image feature maps corresponding to each of the multiple slice remote sensing images to obtain multiple second remote sensing image feature maps corresponding to each of the multiple slice remote sensing images, wherein the resolutions of the multiple second remote sensing image feature maps corresponding to the slice remote sensing images are different from each other. The feature reweighting unit is used to process multiple second remote sensing image feature maps corresponding to each of the multiple slice remote sensing images to obtain multiple third remote sensing image feature maps corresponding to each of the multiple slice remote sensing images, wherein the resolutions of the multiple third remote sensing image feature maps corresponding to the slice remote sensing images are different from each other. Based on the feature maps of multiple third remote sensing images corresponding to each of the multiple slice remote sensing images, a first intermediate object region feature map corresponding to the target object is obtained.
2. The method according to claim 1, wherein, The step of using a trained second feature processing module to perform sparse convolution on the first intermediate object region feature map to obtain an object region feature map corresponding to the target object includes repeatedly performing the following operations until the resolution of the object region feature map meets the first preset condition: If the resolution of the feature map of the first intermediate object region does not meet the first preset condition. The trained second feature processing module is used to process the first intermediate object region feature map to obtain a new first intermediate object region feature map. Determine the resolution of the new first intermediate object region feature map.
3. The method according to claim 2, wherein, The step of processing the first intermediate object region feature map using the trained second feature processing module to obtain a new first intermediate object region feature map includes: Predict the target object region from the feature map of the first intermediate object region; Increase the resolution of the first intermediate object region feature map to obtain the second intermediate object region feature map; Based on the target object region, a sparse convolution is performed on the feature map of the second intermediate object region to obtain the new feature map of the first intermediate object region.
4. The method according to any one of claims 1 to 3, wherein, The first remote sensing image feature map includes multiple object region features; The step of processing multiple first remote sensing image feature maps corresponding to each of the multiple sliced remote sensing images using the feature aggregation unit includes: For each of the multiple first remote sensing image feature maps corresponding to the multiple slice remote sensing images, For object features among multiple object features, Determine the similarity between the object feature and each of the plurality of object region features to obtain a plurality of similarities corresponding to the object feature, wherein the object feature corresponds to the target object in the remote sensing image; Based on multiple similarities corresponding to the object features and the object features, aggregated object features corresponding to the object features are obtained; Based on the aggregated object features corresponding to the respective multiple object features, an aggregated object feature map corresponding to the first remote sensing image feature map is obtained; The first remote sensing image feature map is processed using the aggregated object feature map corresponding to the first remote sensing image feature map to obtain a second remote sensing image feature map corresponding to the first remote sensing image feature map.
5. The method according to any one of claims 1 to 3, wherein, The process of slicing the remote sensing image yields multiple sliced remote sensing images, including: The remote sensing image is processed using a sliding window slicing method to obtain multiple intermediate slice remote sensing images; The multiple intermediate slice remote sensing images are processed using a normalization method to obtain multiple slice remote sensing images.
6. The method according to any one of claims 1 to 3, wherein, The trained first feature processing module and the trained second feature processing module are trained in the following manner: Based on the sample remote sensing images and the second preset conditions, a first sample set and a second sample set are constructed, wherein the second preset conditions are determined according to the category of the sample objects in the sample remote sensing images; The first feature processing module and the second feature processing module are trained using the first sample set to obtain the intermediate first feature processing module and the intermediate second feature processing module. The intermediate first feature processing module and the intermediate second feature processing module are trained using a third sample set to obtain the trained first feature processing module and the trained second feature processing module, wherein the third sample set is obtained based on the first sample set and the second sample set.
7. The method according to claim 6, wherein, The step of training a first feature processing module and a second feature processing module using the first sample set to obtain an intermediate first feature processing module and an intermediate second feature processing module includes: Based on the first sample set, a support set and a query set corresponding to the first sample set are obtained, wherein the support set is used to train the first feature processing module and the second feature processing module, and the query set is used to test the training effect of the intermediate first feature processing module and the intermediate second feature processing module. The first feature processing module and the second feature processing module are trained using the support set and the query set to obtain the intermediate first feature processing module and the intermediate second feature processing module.
8. An electronic device, comprising: One or more processors; Storage device for storing one or more programs. Wherein, when the one or more programs are executed by the one or more processors, the one or more processors perform the method according to any one of claims 1 to 7.
9. A computer-readable storage medium having stored thereon executable instructions that, when executed by a processor, cause the processor to perform the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Remote sensing image small target detection method based on resolution enhancement
CN111709307A
High-resolution remote sensing image target detection method and device and computer equipment
CN112132093A