Target detection system, target detection method and corresponding training method
By adopting a multi-stage refinement network and detection network object detection system in remote sensing image object detection, the problems of limited accuracy of regional suggestions and insufficient generalization capabilities are solved, and higher detection accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510038094.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-01-09
AI Technical Summary
Existing deep learning methods have problems in remote sensing image object detection, limited accuracy of regional suggestions and insufficient generalization ability of small sample data.
A target detection system is adopted, which includes a backbone network, a neck network and a head network. The backbone network extracts multi-layer feature maps, the neck network performs feature fusion, and the head network generates high-quality regional suggestions through multi-stage refinement networks, and uses detection networks for location regression and classification.
By combining feature information at different scales, the accuracy of object detection and system robustness are improved, and the ability to generalize small sample data is enhanced.
Smart Images

Figure CN119963814A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of target detection technology, and in particular to a target detection system, a target detection method, a corresponding training method, an electronic device, and a readable storage medium. Background Art
[0002] Object detection is a core task in computer vision, which aims to accurately locate and identify specific objects in static images or dynamic video frames. This task requires the model to not only accurately delineate the bounding boxes of objects, but also correctly classify the objects within these bounding boxes. In the field of remote sensing, the complexity of object detection is further increased. Remote sensing images are usually acquired by satellites or unmanned aerial vehicles (UAVs), cover a wide geographical area, and contain high resolution, multi-scale features, and a variety of ground objects. These properties make the task of object detection in remote sensing images more challenging and complex. For example, the scale differences of objects in remote sensing images can be very large, and even the resolution of the same object in different images can vary significantly. In addition, the complexity of ground object types often leads to blurred boundaries and severe occlusions. Coupled with complex geographical environments and changing lighting conditions, these factors further aggravate the difficulty of the object detection task. Therefore, achieving efficient and accurate remote sensing image object detection has become an important problem that needs to be solved urgently.
[0003] Taking advantage of the powerful feature representation capabilities of deep learning methods, many researchers have made significant progress in the task of object detection in remote sensing images. These methods have been widely used in disaster monitoring, environmental protection, urban planning and other fields, greatly improving the ability to monitor surface changes, ecological environment and urban development. At present, remote sensing image object detection methods based on deep learning, similar to those in computer vision, are mainly divided into single-stage detection models and two-stage detection models. Single-stage detection models usually integrate target localization and classification tasks within one network, with high detection speed suitable for real-time applications. The two-stage detection segment model separates the region proposal generation from the classification process, provides higher detection accuracy, and is more suitable for more detailed target detection tasks.
[0004] However, all of the above models generate region proposals in one step, and the accuracy of region proposals is limited. Therefore, how to obtain more accurate region proposals to improve detection accuracy has become a research hotspot. In addition, the high performance of these deep learning methods often relies on a large amount of labeled data. In remote sensing image processing, especially in scenes involving novel or rare objects, obtaining sufficient labeled data can be extremely difficult and time-consuming. This problem is not only because remote sensing images usually cover a wide geographical area and require manual annotation of many objects frame by frame, but also because many remote sensing applications involve a wide variety of types and extremely scarce objects, making the annotation process more complex and difficult. Due to these limitations, the generalization ability of existing deep learning methods in remote sensing target detection is limited to a certain extent. Summary of the invention
[0005] In order to overcome the multiple challenges faced by target detection in remote sensing images and improve the generalization ability of existing methods under small sample conditions, the present application provides a target detection system, a detection method, a training method, an electronic device and a readable storage medium.
[0006] According to a first aspect of the present application, a target detection system is provided, which includes: a backbone network, a neck network and a head network implemented based on a computer program, the backbone network is configured to extract different scale features of an input image to generate a multi-layer feature map; the neck network is configured to perform feature fusion on the multi-layer feature map to generate a multi-layer feature fusion map; the head network includes a multi-stage refinement network and a detection network, wherein the multi-stage refinement network is configured to generate an initial region proposal for a detection target based on each layer of the feature fusion map in the multi-layer feature fusion map, and each initial region proposal is iteratively refined through the multi-stage refinement network in turn to obtain a final region proposal for the detection target in each layer of the feature fusion map; the detection network is configured to perform position regression and classification on the detection target based on the final region proposal to obtain the position and category of the detection target in the input image.
[0007] According to a second aspect of the present application, a target detection method is provided, which includes: using a backbone network to extract different scale features of an input image to generate a multi-layer feature map; using a neck network to perform feature fusion on the multi-layer feature map to generate a multi-layer feature fusion map; using a multi-stage refinement network to generate an initial region proposal for a detection target based on each layer of the feature fusion map in the multi-layer feature fusion map, and iteratively refining each initial region proposal through the multi-stage refinement network in turn to obtain a final region proposal for the detection target in each layer of the feature fusion map; using the detection network to perform position regression and classification on the detection target based on the final region proposal to obtain the position and category of the detection target in the input image.
[0008] According to a third aspect of the present application, a training method for a target detection system described in any one of the embodiments of the present application is provided, and the training method includes: a basic training step and a fine-tuning step; the basic training step includes: using basic class data to train the network parameters of the backbone network, the neck network and the head network in the initial target detection system to obtain the target detection system after basic training; the fine-tuning step includes: using the basic class data and a small sample data containing at least one new category of object to train the network parameters of the neck network and the head network except the backbone network in the target detection system after basic training to obtain the target detection system after fine-tuning training.
[0009] According to a fourth aspect of the present application, an electronic device is provided, comprising a processor, a memory, and a program stored in the memory and capable of running on the processor, wherein when the program is executed by the processor, the program implements the steps of any one of the target detection methods provided in the embodiments of the present application, or implements the steps of any one of the training methods for a target detection system provided in the embodiments of the present application.
[0010] According to a fifth aspect of the present application, a computer-readable storage medium is provided, on which instructions are stored. When the instructions are executed by a processor, the steps of any one of the target detection methods provided in the embodiments of the present application are implemented, or the steps of any one of the training methods for a target detection system provided in the embodiments of the present application are implemented.
[0011] Advantageously, the target detection system, target detection method, corresponding training method, electronic device and computer storage medium provided herein have at least the following beneficial effects:
[0012] The input image can be processed by the backbone network to obtain a multi-layer feature map, and the multi-layer feature map can be refined by the neck network to obtain a multi-layer feature fusion map. In this way, it can be ensured that each layer of the feature fusion map contains both high-level semantic information and fine-grained spatial details, thereby improving the system's ability to detect targets at different scales. The multi-layer feature fusion map is input into the multi-stage refinement network in turn, thereby improving the quality and effectiveness of the region proposal through a multi-stage optimization strategy, making the region proposal more stable and reliable, and enhancing the robustness of the system. In addition, the detection network can accurately detect the position and category of the detection target in the input image under multiple final region proposals. In this way, by combining feature information at different scales, the target detection system can more accurately capture the features of the detection target, thereby improving the accuracy of detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In order to more clearly illustrate the specific implementation methods of the present application, the specific implementation methods will be briefly introduced below in conjunction with the drawings. The drawings in the following description are only some implementation methods of the scheme of the present application. For those skilled in the art, other implementation methods can be obtained based on these drawings without paying any creative work.
[0014] Figure 1 A schematic diagram showing the structure of a target detection system provided by an embodiment of the present application is shown;
[0015] Figure 2 A schematic diagram showing the structure of another target detection system provided by an embodiment of the present application is shown;
[0016] Figure 3 A schematic diagram showing a flow chart of a target detection method provided by an embodiment of the present application;
[0017] Figure 4 A schematic diagram showing a flow chart of a training method for a target detection system provided by an embodiment of the present application;
[0018] Figure 5 A schematic structural diagram of an electronic device provided according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0019] In order to make the above and other features and advantages of the present application more clear, the present application is further described below in conjunction with the accompanying drawings. It should be understood that the specific embodiments given herein are for the purpose of explaining to those skilled in the art and are only exemplary and not restrictive.
[0020] In the following description, many specific details are set forth to provide a thorough understanding of the present application. However, it is apparent to those skilled in the art that specific details need not be adopted to practice the present application. In other cases, well-known steps or operations are not described in detail to avoid blurring the present application.
[0021] An embodiment of the present application provides a target detection system. Figure 1 A schematic diagram of the structure of a target detection system provided in an embodiment of the present application is shown as follows: Figure 1 As shown, the target detection system 100 includes: a backbone network 10, a neck network 20 and a head network 30 implemented based on a computer program.
[0022] The backbone network 10 is configured to extract features of different scales of the input image to generate a multi-layer feature map.
[0023] The backbone network involved in the embodiments of this article can be composed of a deep convolutional neural network or other feature extraction networks with similar architectures. The different scale features of the input image can refer to the features of the input image at different resolutions and different scales.
[0024] In some embodiments, the input image may be a remote sensing image acquired through remote sensing technology, wherein the scales of the objects contained in the remote sensing image vary greatly and the background is complex.
[0025] In one embodiment, the backbone network may scale the input image at different scales to form a series of multi-layer images with gradually decreasing resolutions, discard the bottom layer image, and extract features from each of the other layers of images, thereby obtaining a multi-layer feature map. For example, the number of multi-layer feature maps may be, for example, at least 2, such as 3, 4, 5 or more. In an embodiment where the number of multi-layer feature maps is 4, the multi-layer feature maps may include a second layer feature map C2, a third layer feature map C3, a fourth layer feature map C4, and a fifth layer feature map C5.
[0026] In one embodiment, the backbone network may scale the input image to different scales. Specifically, the backbone network may use a preset multiple downsampling operation to scale the input image multiple times. Optionally, the preset multiple may be 2, i.e., 2x downsampling. 2x downsampling may mean reducing the height and width of the image by half.
[0027] The neck network 20 is configured to perform feature fusion on the multi-layer feature maps to generate a multi-layer feature fusion map.
[0028] In one embodiment, the number of multi-layer feature fusion maps is equal to the number of multi-layer feature maps. For example, if the number of multi-layer feature maps is 4, the number of multi-layer feature fusion maps is also 4. The hierarchy of the multi-layer feature fusion map corresponds to the hierarchy of the multi-layer feature map. For example, the multi-layer feature map may include a second-layer feature map C2, a third-layer feature map C3, a fourth-layer feature map C4, and a fifth-layer feature map C5, and the multi-layer feature fusion map may include a second-layer feature fusion map P2, a third-layer feature fusion map P3, a fourth-layer feature fusion map P4, and a fifth-layer feature fusion map P5.
[0029] In one embodiment, the neck network 20 can use the attention mechanism to obtain the highest-level feature fusion map corresponding to the highest-level feature map, and can use the intra-layer fusion mechanism and the cross-layer fusion mechanism to obtain the feature fusion map corresponding to each other layer feature map.
[0030] The head network 30 may be configured to perform target detection on the multi-layer feature fusion map to obtain the position and category of the detected target in the input image.
[0031] The detection target involved in the embodiments of this document may refer to a target object in an input image that can be identified by the target detection system. The position may refer to the position of the detection target in the input image. The category may refer to the category to which the detection target belongs.
[0032] like Figure 1 As shown, the head network 30 may include a multi-stage refinement network 31 and a detection network 32 .
[0033] Among them, the multi-stage refinement network 31 is configured to generate an initial region proposal for the detection target based on each layer of the feature fusion map in the multi-layer feature fusion map, and iteratively refine each initial region proposal through the multi-stage refinement network in turn to obtain the final region proposal for the detection target in each layer of the feature fusion map. The initial region proposal may include information about multiple initial bounding boxes that may surround the detection target. Among them, the information of the bounding box may include multiple groups of bounding box parameters and corresponding classification scores. A set of bounding box parameters may correspond to a bounding box. The classification score may be the probability that the candidate area corresponding to the bounding box contains the detection target. The final region proposal may include information about multiple final bounding boxes that may surround the detection target. The information of the final bounding box may include multiple groups of final bounding box parameters and corresponding category scores. The category score may refer to the probability that the detection target surrounded by the candidate area corresponding to the final bounding box belongs to a certain category.
[0034] In one embodiment, the multi-stage refinement network 31 can be called a multi-stage refinement region proposal network, which first generates an initial region proposal for each feature fusion map, and then refines each initial region proposal through multiple iterations to output a final region proposal corresponding to each initial region proposal.
[0035] For example, multiple feature fusion maps are represented by P2, P3, P4 and P5, and the multi-stage refinement network 31 can output the final region proposal corresponding to P2, the final region proposal corresponding to P3, the final region proposal corresponding to P4, and the final region proposal corresponding to P3.
[0036] The detection network 32 is configured to perform position regression and classification on the detection target based on the final region proposal to obtain the position and category of the detection target in the input image.
[0037] In one embodiment, the detection network may combine each layer of feature fusion map with the corresponding final region proposal and integrate the feature fusion map after multiple layers are combined, perform position regression and classification on the integrated feature fusion map, and output the position and category of the detection target in the input image. The position of the detection target is the coordinate of the bounding box. Classification may refer to identifying the category of the detection target contained in the candidate region corresponding to each bounding box in the integrated region proposal. Position regression may refer to adjusting the position of each bounding box in the integrated region proposal.
[0038] Through the above embodiments, the input image can be processed by the backbone network to obtain a multi-layer feature map, and the multi-layer feature map can be processed by the neck network to obtain a refined multi-layer feature fusion map. In this way, it can be ensured that each layer of the feature fusion map contains both high-level semantic information and fine-grained spatial details, thereby improving the system's ability to detect targets at different scales. The multi-layer feature fusion map is input into the multi-stage refinement network in turn, thereby improving the quality and effectiveness of the region proposal through a multi-stage optimization strategy, making the region proposal more stable and reliable, and enhancing the robustness of the system. In addition, the detection network can accurately detect the position and category of the detection target in the input image under multiple final region proposals. In this way, by combining feature information at different scales, the target detection system can more accurately capture the characteristics of the detection target, thereby improving the accuracy of detection.
[0039] In addition, in some embodiments, since remote sensing images often contain complex geometric features, the inventors have thought of forming a multi-stage network using an expanded convolutional network that can capture contextual features around a bounding box and a variable convolutional network that captures complex and detailed geometric features.
[0040] In one embodiment, the multi-stage refinement network may be composed of at least five types of convolutional layers, multiple classification heads, and multiple regression heads. Each convolutional layer is composed of a convolutional neural network. The at least five types of convolutional layers may include, but are not limited to, dilated convolutional layers, standard convolutional layers, variable convolutional layers, classification convolutional layers, and regression convolutional layers. Dilated convolutional layers and standard convolutional layers may be used to enhance feature map features. The dilated convolutional layer may divide the input feature map into multiple feature regions of fixed size, and perform convolution operations on each feature region to expand the receptive field, aggregate more context information, and output a feature map containing context enhancement.
[0041] The standard convolution layer can further focus on the refinement of local areas and emphasize the key features of edge feature areas based on the context enhancement of the feature map. The variable convolution layer can be used to adaptively adjust the receptive field to capture more complex and detailed geometric features in the feature map input in the final stage.
[0042] It should be noted that the feature map input to the first stage is the feature fusion map of each layer, and the feature map input to each stage except the first stage can be the feature map output by the previous stage.
[0043] The classification head can be used to predict the classification score of the candidate region corresponding to the bounding box, that is, the probability of detecting the existence of the target in the candidate region corresponding to the bounding box. The regression head can be used to predict the regression offset of the bounding box.
[0044] It should be noted that, in some embodiments, the classification head and the regression head can be implemented by existing classification heads and regression heads. For example, the classification head may include a fully connected layer and an activation function. The regression head may include multiple fully connected layers.
[0045] The regression convolution layer is a type of convolutional neural network that can focus on the regression task and optimize the bounding box to generate the final region proposal with precise positioning. The classification convolution layer is a type of convolutional neural network that can focus on the classification task and determine the category of the detection target in the candidate region corresponding to the proposed bounding box.
[0046] Figure 2 A schematic diagram of the structure of another target detection system provided by an embodiment of the present application is shown as follows: Figure 2 As shown, the multi-stage refinement network 31 in the target detection system may include M cascaded stage networks 311, where M is greater than or equal to 2.
[0047] The M cascaded stage networks include at least a first stage network and a final stage network.
[0048] It should be noted that, in the M cascaded stage networks, except for the final stage network, each of the other stage networks 311 includes at least one dilated convolution layer, one standard convolution layer, one classification head and one regression head. Among them, the extended receptive field provided by the dilated convolution layer is crucial for accurately capturing the boundaries of objects, especially in scenarios with limited training data, such as target detection in remote sensing small samples. By cascading the dilated convolution layer with the standard convolution layer, global and local information can be more effectively aggregated, compensating for the information loss caused by insufficient samples, thereby enhancing the robustness of the system and significantly improving the detection accuracy under limited data conditions.
[0049] In one embodiment, the first stage network may include a first dilated convolution layer, a first standard convolution layer, a first classification head, and a first regression head. The first dilated convolution layer is configured to perform context feature extraction on the Nth layer feature fusion map in the multi-layer feature fusion map to obtain a first stage enhanced feature map. Wherein N is greater than or equal to 2.
[0050] The Nth layer feature fusion map can be any layer feature fusion map in the multi-layer feature fusion map. In other words, the input of the first stage network is the Nth layer feature fusion map.
[0051] In one embodiment, the first dilated convolutional layer can be expressed as the following formula.
[0052] DilConv(P n )=∑ k P n [x+r·k x,y +r·k y ]·w[kx , k y ] (1)
[0053] Among them, (x, y) is the coordinate of a specific point on the input N-th layer feature fusion map Pn, r is the expansion rate that determines the spacing between kernel elements, (k x , k y ) is the index of the convolution kernel, identifying a specific element in the kernel.
[0054] The first standard convolutional layer is configured to refine the local features of the first-stage enhanced feature map to obtain the first-stage optimized feature map.
[0055] It should be noted that the input of the first dilated convolutional layer also includes preset anchor box information. Since the first-stage network does not have the region proposal input of the previous-stage network, the anchor box is used as the initialized bounding box input.
[0056] Compared with the Nth layer feature fusion map, the first stage enhanced feature map integrates the context features around the anchor box area into the feature representation. Compared with the first stage enhanced feature map, the first stage optimized feature map increases the refined features of the local area and emphasizes the key features of the anchor box area. Among them, the key features can be edge features and texture features.
[0057] The first regression head is configured to obtain an initial bounding box about the detection target based on the first stage optimized feature map, and the first classification head is configured to predict the classification score of the candidate region corresponding to the initial bounding box based on the first stage optimized feature map. In this way, the first stage initial region proposal of the detection target output by the first stage network is obtained. It should be noted that the initial region proposal includes multiple initial bounding box positions and sizes and corresponding classification scores.
[0058] In one embodiment, the final stage network may include a variable convolution layer, a regression convolution layer, and a classification convolution layer. The variable convolution layer is configured to extract features from the optimized feature map of the previous stage to obtain the feature map of the final stage. It should be noted that by adaptively adjusting the receptive field through the variable convolution layer, more complex and detailed geometric features in the optimized feature map of the previous stage can be captured. This is crucial for accurately identifying targets of different shapes and scales. The regression convolution layer is configured to perform position regression of the bounding box of the detection target for the previous stage region proposal based on the final stage feature map, and the classification convolution layer is configured to perform classification prediction on the candidate region corresponding to the bounding box of the detection target after position regression based on the final stage feature map. In this way, the final region proposal of the detection target of the Nth layer feature fusion map output by the final stage network is obtained.
[0059] The classification prediction in this article can be used to predict the category of the object in the candidate area, and the final region proposal can include the location, size and category of the bounding box.
[0060] In one embodiment, the M cascaded stage networks further include at least one intermediate stage network, and the intermediate stage network includes, for example, a second stage network, a third stage network to an M-1th stage network.
[0061] In addition, each stage network in the intermediate stage network may include an intermediate dilated convolution layer, an intermediate standard convolution layer, an intermediate classification head, and an intermediate regression head. It should be noted that the input of the intermediate stage network is the optimized feature map output by the previous stage network and the previous stage region proposal. The intermediate dilated convolution layer is configured to perform context feature extraction on the previous stage optimized feature map to obtain the intermediate stage enhanced feature map. The intermediate standard convolution layer is configured to refine the candidate region features of the intermediate stage enhanced feature map to obtain the intermediate stage optimized feature map. It should be noted that the candidate region features are the regions corresponding to each bounding box in the previous stage region proposal.
[0062] The intermediate classification head and the intermediate regression head are configured to update the previous stage region proposal based on the intermediate stage optimized feature map to obtain the intermediate stage region proposal. In one embodiment, the intermediate regression head is specifically configured to regress the position of the bounding box in the previous stage region proposal based on the intermediate stage optimized feature map, that is, to update the position of the bounding box in the previous stage region proposal. And the intermediate classification head is specifically used to predict the classification score of the bounding box in the previous stage region proposal based on the intermediate stage optimized feature map, that is, to update the classification score of the candidate region corresponding to the bounding box in the previous stage region proposal.
[0063] In the above embodiment, the multi-stage refinement network can implement multiple stages of iterative refinement to improve the quality of region proposals through M cascaded stage networks. Unlike the traditional single-stage method that usually generates region proposals in one step, the multi-stage refinement network gradually improves the accuracy and quality of the proposals through a series of progressive refinement processes. In addition, through the dilated convolutional layers and standard convolutional layers of each stage network and the variable convolutional layers of the final stage network, features at different resolution levels can be better integrated and adaptive refinement strategies can be implemented, so that the system can more effectively generalize on objects of different scales and appearances, significantly improving the overall generalization ability of the enhanced system, which is particularly beneficial when addressing the challenges of small-sample remote sensing target detection tasks.
[0064] In some embodiments, Figure 2As shown, the neck network 20 may include a dual attention module 21. The dual attention module 21 may be configured to obtain a highest level channel attention map based on a channel attention mechanism for the highest level feature map, and to generate a highest level spatial attention map based on a spatial attention mechanism; and to fuse the highest level channel attention map with the highest level spatial attention map to generate a highest level feature fusion map, wherein the highest level feature map is a feature map corresponding to a maximum value of N.
[0065] In one embodiment, the dual attention module 21 can be a convolutional block attention module constructed by combining a channel attention mechanism and a spatial attention mechanism. The dual attention module 21 performs a combination of global average pooling and global maximum pooling on the highest-level feature map, and then generates the highest-level channel attention map through a shared multi-layer perceptron.
[0066] The dual attention module 21 then multiplies the highest level channel attention map by the highest level feature map element by element to obtain the highest level optimized feature map. The dual attention module 21 performs a combination of maximum pooling and global average pooling on the highest level optimized feature map to obtain a connection feature map, and performs a convolution operation on the connection feature map to obtain the highest level spatial attention map.
[0067] The dual attention module 21 finally uses an element-by-element multiplication operation to fuse the highest-level channel attention map with the highest-level spatial attention map to generate a highest-level feature fusion map.
[0068] It should be noted that the highest level channel attention map is a single channel attention map, which emphasizes the channel with the richest information, suppresses the less relevant channels, and enhances the feature representation. The highest level spatial attention map can optimize the target area detection by focusing on the most significant spatial area in the feature map.
[0069] Taking the highest-level feature map C5 as an example, the calculation formula of the highest-level channel attention map Mc(C5) can be expressed as the following formula.
[0070] Mc(C5)=σ(MLP(AvgPool(C5))+MLP(MaxPool(C5))) (2)
[0071] Among them, σ is the sigmoid activation function, MLP is a shared multi-layer perceptron, AvgPool represents the global average pooling operation, and MaxPool represents the global maximum pooling operation.
[0072] The calculation formula of the highest level optimized feature map C5′ can be expressed as the following formula.
[0073]
[0074] in, Represents element-wise multiplication.
[0075] The highest level spatial attention map M s The calculation formula of (C5′) can be expressed as the following formula.
[0076] M s (C5′)=σ(Conv 7×7 ([AvgPool(C5′);MaxPool(C5′)])) (4)
[0077] Among them, Conv 7×7 represents a 7×7 convolution kernel, σ represents the sigmoid activation function, AvgPool represents the global average pooling operation, and MaxPool represents the global maximum pooling operation.
[0078] The calculation formula of the highest-level feature fusion map P5 can be expressed as the following formula.
[0079]
[0080] In the above embodiment, the dual attention module can ensure that more attention is paid to the areas containing key information in the highest-level feature map by applying the dual attention mechanism to the highest-level feature map, so as to improve the overall accuracy of target detection.
[0081] In some embodiments, Figure 2 As shown, the neck network 20 also includes a plurality of intra-layer feature fusion modules 22, a plurality of cross-layer feature fusion modules 23 and a plurality of upsampling modules 24 respectively distributed on the lowest layer, the highest layer and the middle layer.
[0082] In one embodiment, the number of intra-layer feature fusion modules 22 is equal to the number of multi-layer feature maps, and the number of cross-layer feature fusion modules and the number of sampling modules is one less than the number of intra-layer feature fusion modules.
[0083] It should be noted that although each layer of feature maps corresponds to an intra-layer feature fusion module 22, Figure 2 The intra-layer feature fusion module 22 corresponding to the highest layer in the is configured to multiply the highest layer channel attention map by the highest layer feature map element by element to obtain the highest layer optimized feature map so as to obtain the highest layer spatial attention map. Each layer of feature map except the lowest layer feature map corresponds to a cross-layer feature fusion module 23 and an upsampling module 24.
[0084] For each layer except the highest layer, the intra-layer feature fusion module 22 of each layer is configured to perform intra-layer feature fusion on the feature fusion map of each layer and the feature map after convolution operation based on the first spatial weight and the second spatial weight to obtain the intra-layer feature fusion map of each layer.
[0085] That is to say, the intra-layer feature fusion module 22 can perform a convolution operation on the feature map of each layer, and then perform weighted fusion on it with the corresponding feature fusion map.
[0086] For each layer except the lowest layer, the sampling module 24 of each layer is configured to upsample the intra-layer feature fusion map of each layer to obtain a sampled feature map so that the adopted feature map resolution of each layer is equal to the feature map resolution of the next layer.
[0087] It should be noted that the sampling multiple of the up-sampling module 24 is the same as the down-sampling multiple of the backbone network 10 , that is, if the backbone network adopts 2-fold down-sampling, then the up-sampling module 24 adopts 2-fold up-sampling.
[0088] For each layer except the lowest layer, the cross-layer feature fusion module 32 of each layer is configured to perform cross-layer feature fusion on the sampling feature map of each layer with the feature map of the next layer based on the third spatial weight to obtain the feature fusion map of the next layer.
[0089] The levels involved in the neck network involved in the embodiment of the present application correspond to the levels of the multi-layer feature map. That is, if the levels of the feature map are 2 to 5, then the levels included in the neck network are also 2 to 5. The next layer can be the level corresponding to any feature map except the highest feature map.
[0090] In one embodiment, the feature fusion map of the next layer except the highest layer feature fusion map is obtained by convolving the feature map of the previous layer of the next layer and fusing it with the feature fusion map, upsampling the obtained fusion map so that the sampled feature map matches the resolution of the feature map of the next layer, and fusing it with the feature map of the next layer that has undergone the convolution operation.
[0091] Specifically, the fusion of the feature map of the previous layer and the feature fusion map is achieved through weighted fusion, that is, the first spatial weight multiplied by the feature fusion map of the previous layer plus the second spatial weight multiplied by the feature map of the previous layer. In this way, this weighted fusion effectively captures and emphasizes the features from P n+1 The high-level semantic features and C n+1 The most relevant information in the detail spatial features.
[0092] The fusion of the sampled feature map and the feature map of the current layer after the operation can also be achieved through weighted fusion, that is, the sampled feature map is added to the third spatial weight and multiplied by the feature map of the next layer.
[0093] It should be noted that the first spatial weight, the second spatial weight and the third spatial weight corresponding to each layer need to be set according to the importance of the tasks at each level.
[0094] In this way, the adaptive spatial weighting mechanism enables the neck network to dynamically balance the contributions of these different feature sources according to the importance of the current level task.
[0095] The n-th layer feature fusion graph P n The calculation formula can be expressed as the following formula.
[0096] P n =U(α n ·P n+1 +β n Conv(C n+1 ))+γ n Conv(C n ) (6)
[0097] Among them, U represents upsampling, α n represents the first spatial weight, β n represents the second spatial weight, γ n represents the third space weight, C n It represents the feature map of the nth layer, and Conv represents convolution.
[0098] That is to say, after obtaining the highest-level feature fusion map, the neck network generates lower-level feature fusion maps in a top-down manner. For example, the neck network first generates P5, then P4, P3, and finally P2.
[0099] In the above embodiment, by utilizing the cross-layer adaptive fusion strategy, the neck network effectively enhances the richness and accuracy of the generated feature map. And this robust multi-scale representation significantly enhances the generalization ability of the system in the target detection task of remote sensing small samples, so that it can accurately detect and classify targets of various scales and complexities even when the training data is limited.
[0100] In addition, the neck network enhances the generalization ability of the system by introducing an attention mechanism on the top-level feature map and integrating multi-scale features through a cross-layer fusion process. This neck network is crucial in improving the system's ability to accurately detect objects at a wide range of scales, especially in the complex and diverse environments presented by remote sensing images. By effectively capturing and integrating multi-scale information, the neck network ensures that the system can robustly handle object size and appearance variations that are common in remote sensing data.
[0101] In some embodiments, the sum of the first spatial weight, the second spatial weight, and the third spatial weight is 1, which can be expressed as the following formula.
[0102] α n +β n +γ n =1 (7)
[0103] In the above embodiment, the constraint of the sum of the three weights ensures that the combination of features from different sources remains balanced and no single source dominates the fusion process. The adaptability of the three spatial weights enables the neck network to dynamically emphasize the most relevant features based on the context of each spatial position.
[0104] In some embodiments, Figure 2 As shown, the detection network 32 further includes a pooling layer 321 , a fully connected layer 322 , a final classification layer 323 , and a final regression layer 324 .
[0105] The pooling layer 321 is configured to perform a pooling operation based on the final region proposal and the feature fusion map of each layer to obtain a plurality of standardized feature maps.
[0106] That is, the pooling layer 321 is used to map the bounding box of each final region proposal to each layer of the feature fusion map, and perform a pooling operation on the mapped feature fusion map so that the output size of each layer of the feature fusion map after the pooling operation is consistent.
[0107] The fully connected layer 322 is configured to perform a full connection on each standardized feature map to obtain a feature vector of a fixed dimension corresponding to each standardized feature map.
[0108] That is, the fully connected layer 322 is used to map each standardized feature map to a feature vector of a fixed dimension, each feature vector of a fixed dimension contains high-level features of the detected target in the input image. It should be noted that the fixed dimension can be set according to demand. For example, the fixed dimension can be 512.
[0109] The final classification layer 323 is configured to perform classification based on a plurality of feature vectors of fixed dimensions to determine the category of the detected target in the input image.
[0110] That is, the last classification layer is used to 323 predict the probability that the object enclosed by the bounding box in each final region proposal belongs to each category by analyzing the feature vector of each fixed dimension, and merge the classification results to determine the category of each detected object in the input image.
[0111] The final regression layer 324 is configured to perform position regression based on a plurality of fixed-dimensional feature vectors to determine the position of the detected target in the input image.
[0112] That is, the final regression layer 324 adjusts the position and size of the bounding box in each final region proposal by analyzing each fixed-dimensional feature vector, and merges the adjustment results to determine the position of the detected target in the input image.
[0113] In some embodiments, the last classification layer 323 is further provided with placeholder nodes.
[0114] Specifically, a key challenge in few-shot object detection is that the system tends to overfit to the basic categories during the basic training phase, which affects its ability to adapt to new categories in subsequent stages. To address this problem, specialized nodes, called placeholder nodes, are introduced in the final classification layer of the system. Placeholder nodes are reserved specifically for new categories that the system has not encountered during the basic training phase.
[0115] In some embodiments, the object detection system 100 further includes a loss component. The loss component is used to train the object detection system. The loss component includes the total base loss and the total fine-tuning loss in the following embodiments.
[0116] On the other hand, the present application provides a target detection method, which can be applied to the target detection system 100 described in any of the aforementioned embodiments. Figure 3 A schematic diagram of a target detection method provided in an embodiment of the present application is shown as follows: Figure 3 As shown, the target detection method may include the following steps.
[0117] S41, uses the backbone network to extract different scale features of the input image to generate a multi-layer feature map.
[0118] S42, using the neck network to perform feature fusion on the multi-layer feature map to generate a multi-layer feature fusion map.
[0119] S43, using a multi-stage refinement network to generate an initial region proposal for the detection target based on each layer of the feature fusion map in the multi-layer feature fusion map, and iteratively refine each initial region proposal through the multi-stage refinement network in turn to obtain a final region proposal for the detection target in each layer of the feature fusion map.
[0120] S44, using the detection network to perform position regression and classification on the detection target based on the final region proposal, to obtain the position and category of the detection target in the input image.
[0121] In the above embodiment, the input image can be processed by the backbone network to obtain a multi-layer feature map, and the multi-layer feature map can be refined by the neck network to obtain a multi-layer feature fusion map. In this way, it can be ensured that each layer of the feature fusion map contains both high-level semantic information and fine-grained spatial details, thereby improving the system's ability to detect targets at different scales. The multi-layer feature fusion map is input into the multi-stage refinement network, thereby improving the quality and effectiveness of the region proposal through a multi-stage optimization strategy. And through the detection network, the position and category of the detection target in the input image can be accurately detected under multiple final region proposals.
[0122] In some embodiments, the multi-stage refinement network includes M cascaded stage networks, where M is greater than or equal to 2, and the M cascaded stage networks include at least a first stage network and a final stage network, wherein the first stage network includes a first dilated convolution layer, a first standard convolution layer, a first classification head, and a first regression head, and the final stage network includes a variable convolution layer, a regression convolution layer, and a classification convolution layer. S43 may include the following steps.
[0123] Use the first dilated convolution of the first-stage network to perform contextual feature extraction on the Nth layer feature fusion map in the multi-layer feature fusion map to obtain a first-stage enhanced feature map, use the first standard convolution layer to refine the local features of the first-stage enhanced feature map to obtain a first-stage optimized feature map, use the first regression head to obtain an initial bounding box about the detection target based on the first-stage optimized feature map, and use the first classification head to predict the classification score of the candidate area corresponding to the initial bounding box based on the first-stage optimized feature map, so as to obtain a first-stage initial region proposal for the detection target, where N is greater than or equal to 2.
[0124] The variable convolution layer of the final stage network is used to extract features from the optimized feature map of the previous stage to obtain the feature map of the final stage, and the regression convolution layer is used to regress the position of the bounding box of the detection target for the region proposal of the previous stage based on the feature map of the final stage, and the classification convolution layer is used to classify and predict the candidate region corresponding to the bounding box of the detection target after position regression based on the feature map of the final stage, so as to obtain the final region proposal of the detection target of the Nth layer feature fusion map.
[0125] In some embodiments, the M cascaded stage networks also include at least one intermediate stage network, wherein the intermediate stage network includes an intermediate dilated convolution layer, an intermediate standard convolution layer, an intermediate classification head, and an intermediate regression head. S43 may also include: using the intermediate dilated convolution layer to extract context features from the previous stage optimized feature map to obtain an intermediate stage enhanced feature map, using the intermediate standard convolution layer to refine the candidate region features of the intermediate stage enhanced feature map to obtain an intermediate stage optimized feature map, using the intermediate classification head and the intermediate regression head to update the previous stage region proposal based on the intermediate stage optimized feature map to obtain an intermediate stage region proposal.
[0126] In some embodiments, the neck network includes a dual attention module. S42 may include: using the dual attention module to obtain a highest level channel attention map based on a channel attention mechanism for the highest level feature map, and generating a highest level spatial attention map based on a spatial attention mechanism; fusing the highest level channel attention map with the highest level spatial attention map to generate a highest level feature fusion map, wherein the highest level feature map is a feature map corresponding to a maximum value of N.
[0127] In some embodiments, the neck network further includes multiple intra-layer feature fusion modules, multiple cross-layer feature fusion modules, and multiple upsampling modules respectively distributed on the lowest layer, the highest layer, and the middle layer. S42 may also include the following steps.
[0128] For each layer except the highest layer, an intra-layer feature fusion module of each layer is used to perform intra-layer feature fusion on the feature fusion map of each layer and the feature map after convolution operation based on the first spatial weight and the second spatial weight to obtain the intra-layer feature fusion map of each layer.
[0129] For each layer except the lowest layer, the upsampling module of each layer is used to upsample the intra-layer feature fusion map of each layer to obtain a sampled feature map, so that the resolution of the sampled feature map of each layer is equal to the resolution of the feature map of the corresponding next layer.
[0130] For each layer except the lowest layer, a cross-layer feature fusion module of each layer is used to perform cross-layer feature fusion on the sampling feature map of each layer and the feature map of the corresponding next layer based on the third spatial weight to obtain a feature fusion map of the next layer.
[0131] In some embodiments, the sum of the first spatial weight, the second spatial weight, and the third spatial weight is 1.
[0132] In some embodiments, S44 may include the following steps:
[0133] A pooling layer is used to perform a pooling operation based on the final region proposal and each layer of feature fusion map to obtain multiple standardized feature maps.
[0134] Each standardized feature map is fully connected using a fully connected layer to obtain a feature vector of fixed dimension corresponding to each standardized feature map.
[0135] The final classification layer is used to perform classification based on multiple fixed-dimensional feature vectors to determine the category of the detected target in the input image.
[0136] The final regression layer is configured to perform position regression based on a plurality of feature vectors of fixed dimensions to determine the position of the detected target in the input image.
[0137] In the actual test, airplanes, baseball fields, and tennis courts were selected as new categories, and the remaining categories were used as basic categories. The performance of the target detection system provided in this application was evaluated on 3-shot, 5-shot, 10-shot, and 20-shot detection tasks. For baseline comparison, we selected Meta-RCNN, Fs-DetView, TFA, P-CNN, FSOD, FSCE, ICPE, VFA, and SAE-FSDet as comparison methods. The test results show that the target detection system provided in this application always achieves the highest target detection performance for small samples under all settings, which is 2.47, 2.55, and 3.44 percentage points higher than the best baseline model in the 3-shot, 5-shot, and 10-shot settings, respectively.
[0138] On the other hand, the present application provides a training method for a target detection system, which can be applied to train any of the target detection systems described above. Figure 4 FIG. 1 is a flow chart showing a method for training a target detection system provided by an embodiment of the present application. Figure 4 As shown, the training method of the target detection system may include a basic training step and a fine-tuning step.
[0139] S51, the basic training step includes: using basic class data to train the network parameters of the backbone network, the neck network and the head network in the initial target detection system to obtain the target detection system after basic training.
[0140] For example, the base class data may include a large number of images with labeled base class objects. The base class data may be selected from the DIOR dataset, a large-scale benchmark dataset for target detection in optical remote sensing images.
[0141] In one embodiment, the network parameters of the backbone network, the network parameters of the neck network and the network parameters of the head network in the initial target detection system are initial values. The target detection system after basic training learns a universal feature representation, which can be used to identify the location of the basic category target in the input image and the basic category to which it belongs.
[0142] S52, the fine-tuning step includes: using basic category data and small sample data containing at least one new category object to train network parameters of the neck network and the head network in the target detection system after basic training except the backbone network, to obtain the target detection system after fine-tuning training.
[0143] The small sample data in this article includes multiple images that mark new category objects. The data volume of the small sample data is much smaller than the data volume of the basic class data (for example, less than one-third or less than the data volume of the basic class data). For example, the basic class data has 100 images or more, such as 120 or 150 images or more, and the small sample data has, for example, 20 images, 15 images, 10 images or less, as long as the data volume of the small sample data is much smaller than the data volume of the basic class data. New category objects are objects with different categories from the basic category. For example, new category objects include baseball fields, basketball courts, bridges, chimneys, and ships, and the remaining categories are basic categories.
[0144] In one embodiment, the object detection system after fine-tuning training has fine-tuned network parameters of the neck network and the head network compared to the object detection system after basic training. The object detection system after fine-tuning training is more adaptable to detecting new categories.
[0145] In some embodiments, S51, the basic training step may further include applying a sparse activation mechanism to the placeholder nodes in the last classification layer in the initial object detection system to limit the activity of the placeholder nodes.
[0146] In one embodiment, a sparse activation mechanism can be applied to all placeholder nodes through an L1 regularization mechanism. Applying a sparse activation mechanism can ensure that the placeholder nodes maintain a minimum activation, that is, the output of the placeholder nodes is 0 or close to 0, thereby preventing them from inadvertently learning features related to the base category. In this way, by keeping all placeholder nodes inactive, their flexibility is retained, so that they can be effectively activated when new categories are introduced in the subsequent fine-tuning stage.
[0147] In some embodiments, the network parameters of the backbone network, the neck network and the head network in the initial target detection system are trained using the basic class data to obtain the target detection system after basic training, further comprising: inputting the basic class data into the initial target detection system, iteratively updating the network parameters of the backbone network, the neck network and the head network in the initial target detection system to obtain the minimized total basic loss, thereby obtaining the target detection system after basic training.
[0148] The total base loss includes base classification loss and base regression loss. The base classification loss includes the generalized classification loss of the base categories generated by the detection network and the classification loss of the base categories generated by the final stage network of the multi-stage refinement network.
[0149] The generalized classification loss for the base class generated by the detection network consists of the standard cross entropy loss for the base class and the sparsity regularization loss applied to the placeholder nodes in the classification layer in the original object detection system.
[0150] Generalized classification loss for base classes generated by the detection network It can be expressed as the following formula.
[0151]
[0152] Among them, L base represents the standard cross entropy loss for the base class, L placeholder represents the sparse regularization loss applied to the placeholder nodes, λ placeholder is the regularization coefficient that controls the degree of this sparsity constraint.
[0153] In one embodiment, the classification loss of the base category generated by the final stage network of the multi-stage refinement network can evaluate the accuracy of the bounding box prediction, which can be represented by the cross entropy loss. The cross entropy loss is a measure of the category probability distribution predicted by the final stage network and the category probability distribution of the labeled data.
[0154] The base regression loss includes the regression loss generated by the detection network and the regression loss generated by each stage of the multi-stage refinement network.
[0155] In one embodiment, the regression loss generated by each stage network can be represented by an intersection-over-union loss. The intersection-over-union loss can be used to measure the degree of overlap between the predicted bounding box of the stage network and the true bounding box.
[0156] The regression loss generated by the multi-stage refinement network can be expressed as the following formula.
[0157]
[0158] Among them, L τ MRRPN_reg represents the regression loss of the τth stage, α τ represents the weight of the regression loss of the τth stage network. λ is the weight for measuring the classification loss and regression loss of the multi-stage refinement network.
[0159] In one embodiment, the regression loss generated by the detection network can also be expressed as an intersection-over-union loss.
[0160] In the above embodiment, the system trained by the total base loss including the generalized classification loss can further improve the performance in the small sample classification task.
[0161] In some embodiments, the network parameters of the neck network and the head network except the backbone network in the basic trained target detection system are trained using the basic class data and the small sample data containing at least one new category object to obtain the fine-tuned trained target detection system, including: freezing the parameters of the backbone network in the basic trained target detection system; activating the placeholder nodes in the classification layer of the basic trained target detection system; inputting the basic class data and the small sample data containing at least one new category object into the activated basic trained target detection system, and iteratively updating the parameters of the unfrozen networks in the activated basic trained target detection system to obtain the minimized total combined loss, thereby obtaining the fine-tuned trained target detection system.
[0162] Among them, the total fine-tuning loss includes the classification loss and regression loss for the base category and the new category.
[0163] The classification loss of the base and new categories includes the generalized classification loss of the classification loss of the base and new categories generated by the detection network and the classification loss of the base and new categories generated by the final stage network of the multi-stage refinement network.
[0164] Generalized classification loss for classification loss of base class and new class It includes the classification loss of the base categories generated by the detection network, the classification loss of the new categories, and the sparse regularization loss of the placeholder nodes.
[0165]
[0166] Among them, L base represents the classification loss of the base category, L novel represents the classification loss of the new category, L regularization represents the regularization loss, λ regularization represents the regularization coefficient.
[0167] The classification loss of the base category and the new category generated by the final stage network of the multi-stage refinement network can be expressed as a cross entropy loss.
[0168] In one embodiment, the classification loss of the base category and the classification loss of the new category can be represented by cross entropy loss. The classification loss of the base category can ensure that the system maintains the classification ability of the base category, thereby retaining the previously learned knowledge. The classification loss of the new category can be used to fine-tune the placeholder nodes, so that the system adapts to the classification requirements of the new category. The sparse regularization loss of the placeholder nodes can avoid overfitting the system to the new category.
[0169] The regression loss of the base category and the new category includes the regression loss of the base category and the new category generated by the detection network and the regression loss of the base category and the new category generated by each stage network of the multi-stage refinement network.
[0170] Among them, the regression loss of the basic category and the new category generated by each stage network of the multi-stage refinement network can be expressed as Formula 9. The regression loss of the basic category and the new category generated by the detection network can be expressed as the intersection-over-union loss.
[0171] In the above embodiment, by introducing placeholder nodes and regularization terms and generalized classification losses to train the system, not only the basic classification ability of the model is retained, but also its generalization ability is significantly enhanced. The trained target detection system can effectively adapt to new categories in scenarios with a small number of samples, thereby improving its detection accuracy and generalization performance in remote sensing small sample target detection tasks.
[0172] It should be understood that the specific features, operations and details described hereinabove with respect to the method of the present application may also be similarly applied to the device and system of the present application, or vice versa. In addition, each step of the method of the present application described above may be performed by a corresponding component or unit of the device or system of the present application.
[0173] It should be understood that each module / unit of the device of the present application can be implemented in whole or in part by software, hardware, firmware or a combination thereof. Each module / unit can be embedded in the processor of the electronic device or independent of the processor in the form of hardware or firmware, or can be stored in the memory of the electronic device in the form of software for the processor to call to perform the operation of each module / unit. Each module / unit can be implemented as an independent component or module, or two or more modules / units can be implemented as a single component or module.
[0174] In yet another aspect of the present application, an electronic device is provided. Figure 5 A schematic diagram of the structure of an electronic device provided according to an embodiment of the present application is shown. Figure 5 As shown, the electronic device 60 includes a processor 61, a memory 62, and a program stored in the memory and capable of running on the processor. When the program is executed by the processor, the steps of the target detection method provided in any of the above embodiments are implemented, or the steps of the training method of the target detection system provided in any of the above embodiments are implemented.
[0175] The electronic device 60 can be broadly a server, a terminal, or any other electronic device with necessary computing and / or processing capabilities.
[0176] In one embodiment, the electronic device 60 may include a processor, a memory, a network interface, a communication interface, etc. connected through a system bus. The processor of the electronic device 60 may be used to provide necessary computing, processing and / or control capabilities. The memory of the electronic device 60 may include a non-volatile storage medium and an internal memory. The non-volatile storage medium may store an operating system, a computer program, etc. The internal memory may provide an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface and the communication interface of the electronic device 60 may be used to connect and communicate with external devices through a network.
[0177] On the other hand, the present application provides a computer-readable storage medium, on which instructions are stored, wherein when the instructions are executed by a processor, the steps of the target detection method provided in any of the above embodiments, or the steps of the training method of the target detection system provided in any of the above embodiments are implemented.
[0178] Those skilled in the art will appreciate that the method steps of the present application can be completed by instructing related hardware such as electronic devices or processors through a computer program, and the computer program can be stored in a non-temporary computer-readable storage medium, and the computer program causes the steps of the present application to be executed when it is executed. Depending on the circumstances, any reference to memory, storage or other media herein may include non-volatile or volatile memory. Examples of non-volatile memory include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory, magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk, etc. Examples of volatile memory include random access memory (RAM), external cache memory, etc.
[0179] The various technical features described above can be combined arbitrarily. Although all possible combinations of these technical features are not described, any combination of these technical features should be considered to be covered by this specification as long as there is no contradiction in such combination.
[0180] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein by equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A target detection system, characterized in that: The target detection system includes: a backbone network, a neck network and a head network implemented based on a computer program, The backbone network is configured to extract different scale features of the input image to generate a multi-layer feature map; The neck network is configured to perform feature fusion on the multi-layer feature map to generate a multi-layer feature fusion map; The head network includes a multi-stage refinement network and a detection network, wherein: The multi-stage refinement network is configured to generate an initial region proposal for the detection target based on each layer of the feature fusion map in the multi-layer feature fusion map, and iteratively refine each initial region proposal through the multi-stage refinement network in turn to obtain a final region proposal for the detection target in each layer of the feature fusion map; The detection network is configured to perform position regression and classification on the detection target based on the final region proposal to obtain the position and category of the detection target in the input image.
2. The target detection system according to claim 1, characterized in that: The multi-stage refinement network includes M cascaded stage networks, where M is greater than or equal to 2, and the M cascaded stage networks include at least a first stage network and a final stage network, wherein: The first stage network includes a first dilated convolution layer, a first standard convolution layer, a first classification head and a first regression head, wherein the first dilated convolution layer is configured to perform context feature extraction on the Nth layer feature fusion map in the multi-layer feature fusion map to obtain a first stage enhanced feature map, the first standard convolution layer is configured to refine the local features of the first stage enhanced feature map to obtain a first stage optimized feature map, the first regression head is configured to obtain an initial bounding box about the detection target based on the first stage optimized feature map, and the first classification head is configured to perform classification score prediction on the candidate area corresponding to the initial bounding box based on the first stage optimized feature map, thereby obtaining a first stage initial area proposal about the detection target, wherein N is greater than or equal to 2; and / or The final stage network includes a variable convolution layer, a regression convolution layer and a classification convolution layer, wherein the variable convolution layer is configured to perform feature extraction on the optimized feature map of the previous stage to obtain the final stage feature map, the regression convolution layer is configured to perform position regression of the bounding box of the detection target for the previous stage region proposal based on the final stage feature map, and the classification convolution layer is configured to perform classification prediction on the candidate region corresponding to the bounding box of the detection target after position regression based on the final stage feature map, thereby obtaining the final region proposal of the detection target of the Nth layer feature fusion map.
3. The target detection system according to claim 2, characterized in that: The M cascaded stage networks also include at least one intermediate stage network, wherein: The intermediate stage network includes an intermediate dilated convolution layer, an intermediate standard convolution layer, an intermediate classification head, and an intermediate regression head. The intermediate dilated convolution layer is configured to perform context feature extraction on the previous stage optimized feature map to obtain the intermediate stage enhanced feature map. The intermediate standard convolution layer is configured to refine the candidate region features of the intermediate stage enhanced feature map to obtain the intermediate stage optimized feature map. The intermediate classification head and the intermediate regression head are configured to update the previous stage region proposal based on the intermediate stage optimized feature map to obtain the intermediate stage region proposal.
4. The target detection system according to claim 1, characterized in that: The neck network includes a dual attention module, where The dual attention module is configured to obtain a highest-level channel attention map based on a channel attention mechanism for the highest-level feature map, and to generate a highest-level spatial attention map based on a spatial attention mechanism; and to fuse the highest-level channel attention map with the highest-level spatial attention map to generate a highest-level feature fusion map, wherein the highest-level feature map is a feature map corresponding to N being the maximum value.
5. The target detection system according to claim 4, characterized in that: The neck network also includes a plurality of intra-layer feature fusion modules, a plurality of cross-layer feature fusion modules and a plurality of upsampling modules respectively distributed on the lowest layer, the highest layer and the middle layer; Wherein, for each layer except the highest layer, the intra-layer feature fusion module of each layer is configured to perform intra-layer feature fusion on the feature fusion map of each layer and the feature map after the convolution operation based on the first spatial weight and the second spatial weight to obtain the intra-layer feature fusion map of each layer; For each layer except the lowest layer, the upsampling module of each layer is configured to upsample the intra-layer feature fusion map of each layer to obtain a sampled feature map, so that the resolution of the sampled feature map of each layer is equal to the resolution of the feature map of the corresponding next layer; For each layer except the lowest layer, the cross-layer feature fusion module of each layer is configured to perform cross-layer feature fusion on the sampling feature map of each layer and the feature map of the corresponding next layer based on the third spatial weight to obtain the feature fusion map of the next layer.
6. The target detection system according to claim 5, characterized in that: The sum of the first spatial weight, the second spatial weight and the third spatial weight is 1.
7. The target detection system according to claim 1, characterized in that: The detection network further includes a pooling layer, a fully connected layer, a final classification layer, and a final regression layer; wherein, The pooling layer is configured to perform a pooling operation based on the final region proposal and each layer of feature fusion map to obtain a plurality of standardized feature maps; The fully connected layer is configured to perform a full connection on each standardized feature map to obtain a feature vector of a fixed dimension corresponding to each standardized feature map; The last classification layer is configured to perform classification based on a plurality of feature vectors of fixed dimensions to determine the category of the detection target in the input image; The last regression layer is configured to perform position regression based on the plurality of feature vectors of the fixed dimension to determine the position of the detection target in the input image.
8. A target detection method, characterized in that: The target detection method comprises: Use the backbone network to extract different scale features of the input image to generate multi-layer feature maps; Using the neck network to perform feature fusion on the multi-layer feature map to generate a multi-layer feature fusion map; Using a multi-stage refinement network to generate an initial region proposal for the detection target based on each layer of the feature fusion map in the multi-layer feature fusion map, and iteratively refine each initial region proposal through the multi-stage refinement network in turn to obtain a final region proposal for the detection target in each layer of the feature fusion map; A detection network is used to perform position regression and classification on the detection target based on the final region proposal to obtain the position and category of the detection target in the input image.
9. The target detection method according to claim 8, characterized in that: The multi-stage refinement network includes M cascaded stage networks, where M is greater than or equal to 2, and the M cascaded stage networks include at least a first stage network and a final stage network, wherein the first stage network includes a first dilated convolution layer, a first standard convolution layer, a first classification head, and a first regression head, and the final stage network includes a variable convolution layer, a regression convolution layer, and a classification convolution layer, The method of using the multi-stage refinement network to generate an initial region proposal for the detection target based on each layer of the feature fusion map in the multi-layer feature fusion map, and iteratively refining each initial region proposal through the multi-stage refinement network in turn to obtain a final region proposal for the detection target in each layer of the feature fusion map, includes: Using the first dilated convolution of the first-stage network to perform context feature extraction on the Nth layer feature fusion map in the multi-layer feature fusion map to obtain a first-stage enhanced feature map, using the first standard convolution layer to refine the local features of the first-stage enhanced feature map to obtain a first-stage optimized feature map, using the first regression head to obtain an initial bounding box about the detection target based on the first-stage optimized feature map, and using the first classification head to predict the classification score of the candidate area corresponding to the initial bounding box based on the first-stage optimized feature map, thereby obtaining a first-stage initial region proposal about the detection target, wherein N is greater than or equal to 2; Use the variable convolution layer of the final stage network to extract features from the optimized feature map of the previous stage to obtain the final stage feature map, use the regression convolution layer to regress the position of the bounding box of the detection target for the region proposal of the previous stage based on the final stage feature map, and use the classification convolution layer to classify and predict the candidate region corresponding to the bounding box of the detection target after position regression based on the final stage feature map, so as to obtain the final region proposal of the detection target of the Nth layer feature fusion map.
10. The target detection method according to claim 9, characterized in that: The M cascaded stage networks also include at least one intermediate stage network, wherein: The intermediate stage network includes an intermediate dilated convolution layer, an intermediate standard convolution layer, an intermediate classification head, and an intermediate regression head. The method further comprises: generating an initial region proposal for the detection target based on each layer of the feature fusion map in the multi-layer feature fusion map using the multi-stage refinement network, and iteratively refining each initial region proposal through the multi-stage refinement network in turn to obtain a final region proposal for the detection target in each layer of the feature fusion map. An intermediate dilated convolution layer is used to extract contextual features of the optimized feature map of the previous stage to obtain an intermediate stage enhanced feature map, an intermediate standard convolution layer is used to refine the candidate region features of the intermediate stage enhanced feature map to obtain an intermediate stage optimized feature map, an intermediate classification head and an intermediate regression head are used to update the previous stage region proposal based on the intermediate stage optimized feature map to obtain an intermediate stage region proposal.
11. The target detection method according to claim 9, characterized in that: The neck network includes a dual attention module, wherein the neck network is used to perform feature fusion on the multi-layer feature map to generate a multi-layer feature fusion map, including: A dual attention module is used to obtain a highest-level channel attention map based on a channel attention mechanism for the highest-level feature map, and a highest-level spatial attention map is generated based on a spatial attention mechanism; the highest-level channel attention map is fused with the highest-level spatial attention map to generate a highest-level feature fusion map, wherein the highest-level feature map is a feature map corresponding to a maximum value of N.
12. The target detection method according to claim 11, characterized in that: The neck network further includes a plurality of intra-layer feature fusion modules, a plurality of cross-layer feature fusion modules and a plurality of upsampling modules respectively distributed on the lowest layer, the highest layer and the middle layer, wherein the neck network is used to perform feature fusion on the multi-layer feature map to generate a multi-layer feature fusion map, and further includes: For each layer except the highest layer, an intra-layer feature fusion module of each layer is used to perform intra-layer feature fusion on the feature fusion map of each layer and the feature map after the convolution operation based on the first spatial weight and the second spatial weight to obtain an intra-layer feature fusion map of each layer; For each layer except the lowest layer, use the upsampling module of each layer to upsample the feature fusion map within each layer to obtain a sampled feature map, so that the resolution of the sampled feature map of each layer is equal to the resolution of the feature map of the corresponding next layer; For each layer except the lowest layer, a cross-layer feature fusion module of each layer is used to perform cross-layer feature fusion on the sampling feature map of each layer and the feature map of the corresponding next layer based on the third spatial weight to obtain a feature fusion map of the next layer.
13. The target detection method according to claim 12, characterized in that: The sum of the first spatial weight, the second spatial weight and the third spatial weight is 1.
14. A training method for a target detection system according to any one of claims 1 to 7, characterized in that: The training method comprises: a basic training step and a fine-tuning step; The basic training step includes: using basic class data to train network parameters of a backbone network, a neck network, and a head network in the initial target detection system to obtain a target detection system after basic training; The fine-tuning step includes: using the basic class data and small sample data containing at least one new category object to train the network parameters of the neck network and the head network in the target detection system after the basic training except the backbone network, to obtain the target detection system after fine-tuning training.
15. The training method according to claim 14, characterized in that: The basic training step further includes: applying a sparse activation mechanism to the placeholder nodes in the last classification layer in the initial object detection system to limit the activity of the placeholder nodes.
16. The training method according to claim 14, characterized in that: The method further includes: using the basic class data to train the network parameters of the backbone network, the neck network and the head network in the initial target detection system to obtain the target detection system after basic training; Inputting the basic class data into the initial target detection system, iteratively updating the network parameters of the backbone network, the neck network and the head network in the initial target detection system to obtain the minimized total basic loss, thereby obtaining the target detection system after basic training; The total basic loss includes a basic classification loss and a basic regression loss. The basic classification loss includes a generalized classification loss of a basic category generated by a detection network and a classification loss of a basic category generated by a final stage network of a multi-stage refinement network. The generalized classification loss of the basic category includes a standard cross entropy loss of the basic category and a sparse regularization loss applied to placeholder nodes in a classification layer in the initial target detection system. The basic regression loss includes a regression loss generated by a detection network and a regression loss generated by each stage network of the multi-stage refinement network.
17. The training method according to any one of claims 14 to 16, characterized in that: The method of using the basic class data and the small sample data containing at least one new class object to train the network parameters of the neck network and the head network except the backbone network in the target detection system after the basic training to obtain the target detection system after fine-tuning training includes: Freeze the parameters of the backbone network in the object detection system after basic training; Activate the placeholder nodes in the classification layer of the object detection system after base training; The base class data and the small sample data containing at least one new category object are input into the activated basic trained target detection system, and the parameters of the unfrozen network in the activated basic trained target detection system are iteratively updated to minimize the total fine-tuning loss, thereby obtaining the fine-tuned trained target detection system.
18. The training method according to claim 17, characterized in that: The total fine-tuning loss includes classification loss and regression loss for base categories and new categories, the classification loss for base categories and new categories includes the generalized classification loss of the classification loss for base categories and new categories generated by the detection network and the classification loss for base categories and new categories generated by the final stage network of the multi-stage refinement network, the generalized classification loss of the classification loss for base categories and new categories includes the classification loss for base categories generated by the detection network, the classification loss for new categories and the sparse regularization loss for placeholder nodes; the regression loss for base categories and new categories includes the regression loss for base categories and new categories generated by the detection network and the regression loss for base categories and new categories generated by each stage network of the multi-stage refinement network.
19. An electronic device, characterized in that: The invention comprises a processor, a memory and a program stored in the memory and capable of being run on the processor, wherein when the program is executed by the processor, the steps of the target detection method as described in any one of claims 8 to 13 are implemented, or the steps of the training method for a target detection system as described in any one of claims 14 to 18 are implemented.
20. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, and when the instructions are executed by the processor, the steps of the target detection method described in any one of claims 8 to 13 are implemented, or the steps of the training method of the target detection system described in any one of claims 14 to 17 are implemented.
Citation Information
Patent Citations
General target detection method based on remote sensing image and medium
CN118230183A
Skin disease detection method based on YOLO-IRLSK lightweight model
CN118396933A
Tongue body segmentation and tongue coating dryness moistening identification method based on multi-modal image fusion
CN119007236A
Multi-scale aware pedestrian detection method based on improved full convolutional network
US20210056351A1
Cited By
Online intelligent image encryption method and system
CN120263915A
Electric power target detection method based on space interaction and segmentation attention under small sample
CN120807886A
Power target detection method based on spatial interaction and segmentation attention under small sample
CN120807886B
Animal behavior frame-level time sequence segmentation method fusing skeleton and visual information
CN121121841A