Target detection method and device based on adaptive anchor alignment and storage medium
The self-adaptive anchor alignment method enhances target detection precision by aligning anchor boxes to fit target shapes and positions, improving classification and localization accuracy.
Patent Information
- Application Number
- CN202510796932.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-06-16
AI Technical Summary
The existing target detection methods use the preset uniform anchor frame mechanism to achieve low target detection accuracy, and cannot effectively cover the shape and size of some real targets in the image, resulting in positioning deviations and missed detection, affecting the accuracy of target classification and positioning tasks.
The target detection method based on adaptive anchor alignment is adopted, and the preset uniform distribution of anchor boxes is adjusted through the adaptive anchor alignment network. The target distribution and scale are extracted step by step by step by step by step by step by step by adaptive alignment module, and adaptive anchor boxes that are highly matched with the targets in the image to improve classification and positioning accuracy.
It improves the accuracy and efficiency of target detection, and can be better applied to various target classification and positioning tasks, especially when detecting complex shape targets, which significantly improves the detection effect.
Smart Images

Figure CN120318500A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image target detection, and in particular to a target detection method, device and computer-readable storage medium based on adaptive anchor alignment. Background Art
[0002] Target detection of natural images is an important means to obtain image information. Through target detection, the targets in the image can be located and classified. At present, target detection technology plays an important role in various fields. For example, in the field of intelligent transportation and autonomous driving, by performing target detection on road images, intelligent vehicles, pedestrians and obstacles in the images are identified and located, so as to calculate the distances between various targets, plan driving routes for intelligent vehicles, control intelligent vehicles to avoid obstacles and pedestrians, thereby reducing the accident rate; in the field of biodiversity research, by performing fine-grained recognition on images containing multi-species information, the distribution characteristics and quantity information of different species in the images can be accurately counted, thus providing key data for biodiversity protection and ecosystem research.
[0003] Existing target detection methods use a preset uniform anchor box mechanism to locate targets. First, feature extraction is performed on the image to obtain a classification feature map containing the semantic information of each target and a localization feature map containing the position information of each target. Then, multiple groups of fixed-size and uniformly distributed anchor boxes are generated on the classification feature map and the localization feature map to cover the possible targets in the image. Then, by predicting and classifying each anchor box in the classification feature map and the localization feature map containing the anchor boxes, it is determined whether each anchor box contains a target and the target category contained. Finally, the anchor boxes containing the targets and the target categories are output to obtain the target prediction category and the predicted localization box. However, since the anchor boxes are all preset with fixed sizes, this will result in the mismatch between the shapes and sizes of the anchor boxes and some real targets in the image. For example, when detecting a slender bridge in a drone image, the preset anchor boxes with a ratio of 2:1 cannot effectively cover the slender bridge, which will result in missed detection or positioning deviation; at the same time, the uniformly distributed anchor boxes will be strictly aligned with the feature map grid points, while the center of the real localization box of the target may be located in the grid gap, which results in a natural offset between the initial position of the anchor box and the real localization box, thus resulting in a large error in the final localization of the target and reducing the accuracy of target detection.
[0004] In summary, the existing target detection methods have the problem of low target detection accuracy, which affects the accuracy of various target classification and target localization tasks. Summary of the Invention
[0005] Therefore, the technical problem to be solved by the present invention is to overcome the problem that the existing target detection methods have low target detection accuracy, which affects the accuracy of various target classification and target localization tasks.
[0006] To solve the above technical problems, the present invention provides an object detection method based on adaptive anchor alignment, including: Input the image to be detected into the feature extraction network, and output the localization feature map and the classification feature map; Input the localization feature map and the classification feature map into the adaptive anchor alignment network. The adaptive anchor alignment network includes N cascaded adaptive alignment modules. Each adaptive alignment module performs adaptive alignment on the input feature map, including: Perform standard convolution on the input classification feature map to output an offset; obtain a bias feature map based on the offset and a preset dilated prior; Perform deformable convolution on the bias feature map and the input classification feature map simultaneously, and output the classification feature map to the next adaptive alignment module; Perform deformable convolution on the input localization feature map and the offset simultaneously, and output the localization feature map to the next adaptive alignment module; Take the classification feature map, the localization feature map, and the bias feature map output by the last adaptive alignment module as the object classification feature map, the object localization feature map, and the object bias feature map respectively; generate adaptive anchors based on the object bias feature map and the preset uniformly distributed anchors; Input the adaptive anchors, the object localization feature map, and the object classification feature map into the decoder to obtain the predicted localization boxes and predicted categories of each object in the image to be detected.
[0007] Preferably, each adaptive alignment module performs standard convolution on the input classification feature map to output an offset, including: performing standard convolution on the classification feature map using a first convolution unit and a second convolution unit with different convolutional kernel sizes in series, and obtaining the offset based on the output of the second convolution unit.
[0008] Preferably, the offset output by each adaptive alignment module is expressed as: , where, represents the offset output by the th adaptive alignment module, represents the number of adaptive alignment modules; represents the standard convolution unit for performing standard convolution; represents the classification feature map input to the th adaptive alignment module, that is, the classification feature map output by the th adaptive alignment module. When , is the classification feature map output by the feature extraction network; The bias feature map obtained by each adaptive alignment module is expressed as: , Among them, represents the bias feature map obtained by the th adaptive alignment module; represents a preset hole prior; The classification feature map output by each adaptive alignment module is expressed as: , where represents the classification feature map output by the th adaptive alignment module; represents a deformable convolution unit that performs deformable convolution; The localization feature map output by each adaptive alignment module is expressed as: , where represents the localization feature map output by the th adaptive alignment module; represents the localization feature map input to the th adaptive alignment module, that is, the localization feature map output by the th adaptive alignment module. When , is the localization feature map output by the feature extraction network.
[0009] Preferably, inputting the image to be detected into the feature extraction network, the output of the localization feature map and the classification feature map includes: Inputting the image to be detected into the encoder for multi-level feature extraction, and outputting a multi-scale feature map; Inputting the multi-scale feature maps into the parallel classification feature extraction network and localization feature extraction network respectively, and outputting the classification feature map and the localization feature map.
[0010] Preferably, generating adaptive anchors based on the target bias feature map and the preset uniform distribution anchors includes: jointly inputting the target bias feature map and the preset uniform distribution anchors into the alignment anchor generator, and outputting adaptive anchors; among them, the adaptive anchors are expressed as: , where represents the adaptive anchor; represents pixel-by-pixel addition; represents the anchor point list of the preset uniform distribution anchors; represents the downsampling rate of each scale feature map in the multi-scale feature map, , represents the number of scale feature maps; represents the average operator.
[0011] Preferably, when training the feature extraction network, the adaptive anchor alignment network, and the decoder using the images in the training set, the construction process of the object detection loss function includes: Using the sample assignment strategy paradigm, each object in the image is divided based on the predicted localization box and predicted category of each object in the image to obtain a set of positive object samples and a set of negative object samples; Based on the distance metric values between the predicted categories of each object in the set of positive object samples and the set of negative object samples and the object classification feature map of the image, a classification loss function is constructed; Based on the distance metric values between the predicted localization boxes of each object in the set of positive object samples and their true localization boxes, a localization loss function is constructed; Based on the weighted sum of the classification loss function and the localization loss function, an object detection loss function is constructed.
[0012] Preferably, the classification loss function is expressed as: , where, represents the classification loss function; represents a function for calculating the distance metric values between the predicted categories of each object in the set of positive object samples and the set of negative object samples and the object classification feature map of the image; represents the set of positive object samples; represents the set of negative object samples; represents the object classification feature map; The localization loss function is expressed as: , where, represents the localization loss function; represents a function for calculating the distance metric values between the predicted localization boxes of each object in the set of positive object samples and their true localization boxes; represents the predicted localization box; represents the true localization box; represents filtering; The object detection loss function is expressed as: , where, represents the object detection loss function; represents the weight of the localization loss function.
[0013] Preferably, the adaptive anchor alignment network includes 3 cascaded adaptive alignment modules.
[0014] The present invention also provides an object detection device based on adaptive anchor alignment, including: A feature extraction module, configured to input an image to be detected into a feature extraction network and output a localization feature map and a classification feature map; An adaptive alignment module, configured to input the localization feature map and the classification feature map into an adaptive anchor alignment network. The adaptive anchor alignment network includes N cascaded adaptive alignment modules. Each adaptive alignment module performs adaptive alignment on the input feature map, including: Performing standard convolution on the input classification feature map to output an offset; obtaining a bias feature map based on the offset and a preset atrous prior; Performing deformable convolution on the bias feature map and the input classification feature map simultaneously to output a classification feature map to the next adaptive alignment module; Performing deformable convolution on the input localization feature map and the offset simultaneously to output a localization feature map to the next adaptive alignment module; An adaptive anchor generation module, configured to use the classification feature map, the localization feature map, and the bias feature map output by the last adaptive alignment module as a target classification feature map, a target localization feature map, and a target bias feature map respectively; generating adaptive anchors based on the target bias feature map and a preset uniformly distributed anchor; A target detection module, configured to input the adaptive anchors, the target localization feature map, and the target classification feature map into a decoder to obtain predicted localization boxes and predicted categories of each target in the image to be detected.
[0015] The present invention also provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, the steps of the above target detection method based on adaptive anchor alignment are implemented.
[0016] The target detection method based on adaptive anchor alignment provided by this application has the following beneficial effects: After the present application performs preliminary feature extraction on an image to obtain a classification feature map and a localization feature map, instead of directly generating anchors with a fixed size and evenly distributed, it first inputs the classification feature map and the localization feature map into an adaptive anchor alignment network. The adaptive anchor alignment network includes multiple cascaded adaptive alignment modules. When each adaptive alignment module performs a sliding window operation on the input classification feature map using a standard convolution, it can output offsets such as the offset size of the sampling region with respect to the center of the target region and the scaling ratio of the sampling width and height with respect to the true width and height of the target. These offsets can guide the subsequent deformable convolution to offset the sampling positions on the feature map. At the same time, a basic sampling grid is provided through a preset dilation prior, so that the deformable convolution can fine-tune the sampling region in the classification feature map based on this sampling grid in combination with the offsets, obtaining feature information that is more focused on the target region. Then, deformable convolution is performed on the offsets and the input localization feature map to fine-tune the sampling region of the localization feature map, also making the output localization feature map more focused on the target region; through multiple cascaded adaptive alignment modules, the target distribution and scale in the classification feature map and the localization feature map are gradually extracted, and finally, a bias feature map containing the position of the target core region, a target classification feature map containing depth feature information, and a target localization feature map are output. Since the target bias feature map contains the position of the target core region, and the preset evenly distributed anchors serve as the basic anchor grid to provide an initial set of anchor boxes that evenly cover the image, therefore, by using the target bias feature map to adjust the evenly distributed adaptive anchors, anchor boxes (i.e., adaptive anchors) that highly match the scale, shape, and position of the target in the image can be generated. The adaptive anchors can guide the model to focus on the target region when performing classification and localization based on the target classification feature map and the target localization feature map, thereby improving the classification accuracy and localization accuracy of the model. The present application uses an adaptive anchor alignment network to generate adaptive anchors that converge to the target region, thereby improving the target detection accuracy and being better applicable to various target classification and target localization tasks. Description of the Drawings
[0017] To make the content of the present invention easier to understand clearly, the following further details the present invention according to specific embodiments of the present invention in combination with the drawings, where: Figure 1 Flowchart of the target detection method based on adaptive anchor alignment provided by the present application; Figure 2 Flowchart of the feature extraction of the feature map by the adaptive alignment module provided by the present application; Figure 3 Schematic diagram of the adaptive anchor alignment network structure provided by the present application; Figure 4 Schematic diagram of the adaptive alignment module structure provided by the present application; Figure 5 Schematic diagram of the convergence comparison between the adaptive anchor points obtained in this application and the traditional uniformly distributed anchor points; among them, Figure 5 in (a) is a schematic diagram of the traditional uniformly distributed anchor points generated on the image to be detected, Figure 5 in (b) is a schematic diagram of the object detection result obtained by detecting the object detection model in the prior art for the figure (a), Figure 5 in (c) is a schematic diagram of the adaptive anchor points generated on the image to be detected when using 1 adaptive alignment module, Figure 5 in (d) is the object detection result obtained by detecting the figure (c) by the object detection model provided in this application that includes 1 adaptive alignment module Figure 5 in; Figure 5 in (e) is a schematic diagram of the adaptive anchor points generated on the image to be detected when using 2 cascaded adaptive alignment modules, Figure 5 in (f) is the object detection result obtained by detecting the figure (e) by the object detection model provided in this application that includes 2 cascaded adaptive alignment modules Figure 5 in; Figure 5 in (g) is a schematic diagram of the adaptive anchor points generated on the image to be detected when using 3 adaptive alignment modules, Figure 5 in (h) is the object detection result obtained by detecting the figure (g) by the object detection model provided in this application that includes 3 adaptive alignment modules Figure 5 in; Figure 6 is a loss curve diagram when the object detection model provided in this application and the object detection model in the prior art detect the same image; among them, Figure 6 in (a) represents the localization loss curve diagram, Figure 6 in (b) represents the total detection loss curve diagram based on the localization loss and the classification loss; Figure 7 is a comparison diagram of the detection results when the object detection model provided in this application and the object detection model in the prior art detect different images; among them, Figure 7 in (a) is a schematic diagram of the result of detecting the first image using the object detection model in the prior art, Figure 7 in (b) is a schematic diagram of the result of detecting the first image using the object detection model of this application, Figure 7 in (c) is a schematic diagram of the result of detecting the second image using the object detection model in the prior art, Figure 7 in (d) is a schematic diagram of the result of detecting the second image using the object detection model of this application, Figure 7 in (e) is a schematic diagram of the result of detecting the third image using the object detection model in the prior art, Figure 7In (f), it is a schematic diagram of the detection result of the third image using the object detection model of the present application. Figure 7 In (g), it is a schematic diagram of the detection result of the fourth image using the object detection model in the prior art. Figure 7 In (h), it is a schematic diagram of the detection result of the fourth image using the object detection model of the present application. Specific embodiments
[0018] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it, but the examples given are not intended to limit the present invention.
[0019] Please refer to Figure 1 , Figure 1 As shown, it is a flowchart of an object detection method based on adaptive anchor alignment provided by the present application. The method specifically includes: S10: Input the image to be detected into the feature extraction network, and output the localization feature map and the classification feature map.
[0020] S20: Input the localization feature map and the classification feature map into the adaptive anchor alignment network. Among them, the adaptive anchor alignment network includes N cascaded adaptive alignment modules.
[0021] As Figure 2 shown, the adaptive alignment of the input feature map by each adaptive alignment module includes: S200: Perform standard convolution on the input classification feature map to output the offset.
[0022] S201: Perform deformable convolution on the bias feature map and the input classification feature map simultaneously, and output the classification feature map to the next adaptive alignment module.
[0023] S202: Perform deformable convolution on the input localization feature map and the offset simultaneously, and output the localization feature map to the next adaptive alignment module.
[0024] S30: Respectively use the classification feature map, the localization feature map, and the bias feature map output by the last adaptive alignment module as the target classification feature map, the target localization feature map, and the target bias feature map; generate adaptive anchors based on the target bias feature map and the preset uniformly distributed anchors.
[0025] S40: Input the adaptive anchors, the target localization feature map, and the target classification feature map into the decoder to obtain the predicted localization boxes and predicted categories of each target in the image to be detected.
[0026] Furthermore, in some embodiments of the present application, the specific implementation manner of step S10 is: S100: Input the image to be detected into the encoder for multi-level feature extraction, and output multi-scale feature maps.
[0027] Optionally, an encoder composed of a backbone network and a neck structure can be used as the multi-scale feature extractor, and its feature extraction formula is expressed as: , where, represents the multi-scale feature map, , represents the th scale feature map, represents the th height of the scale feature map, represents the th width of the scale feature map, represents the number of scale feature maps; represents the image to be detected, , represents the height of the image to be detected, represents the width of the image to be detected.
[0028] S101: Input the multi-scale feature maps into the parallel classification feature extraction network and localization feature extraction network respectively, and output the classification feature map and the localization feature map.
[0029] Specifically, the classification feature map output by the classification feature extraction network is expressed as: , where, represents the convolutional layer of the classification feature extraction network; The localization feature map output by the localization feature extraction network is expressed as: , where, represents the convolutional layer of the localization feature extraction network.
[0030] During image object detection, the model evenly distributes some anchor boxes of a fixed size on the image based on the localization feature map and the classification feature map. Subsequently, by extracting features from these anchor boxes, the probability and specific category of the presence of an object within each anchor box are predicted, and the offset of the anchor box with respect to the true object localization box is obtained, enabling the model to adjust the anchor box containing the object into a more accurate predicted localization box. In the prior art, since the anchor boxes are all preset with a fixed size, this can lead to a mismatch between the shape and size of the anchor boxes and some true objects in the image. For example, when detecting a slender bridge in a drone image, the preset anchor boxes with a 2:1 ratio cannot effectively cover the slender bridge, resulting in missed detections or localization deviations. At the same time, the evenly distributed anchor boxes are strictly aligned with the feature map grid points, but the center of the true object localization box may be located in the grid gap, which causes a natural offset between the initial position of the anchor box and the true object localization box, also leading to a large error in the final localization of the object. In addition, the number of evenly distributed anchor boxes in an image may be as high as tens of thousands, but only a few anchor boxes contain real objects. This not only increases the time for object detection, affecting the efficiency of object detection, but also makes the model more affected by the background (anchor boxes without objects), reducing the accuracy of object detection. Therefore, directly using the existing preset evenly distributed anchor boxes for object detection will affect the efficiency and accuracy of object detection. In practical applications, when it is necessary to perform object detection on an aerial vehicle and adjust the flight path and flight speed of the aerial vehicle based on the object detection results, such low-efficiency and low-accuracy detection results will bring incorrect data guidance to air traffic control.
[0031] For the above reasons, the present application designs an adaptive anchor alignment network to adjust the preset evenly distributed anchor boxes. Specifically, as Figure 3 shown, the adaptive anchor alignment network is composed of N adaptive alignment modules connected in series. The four convolutional blocks in the figure are the convolutional layers in the classification feature extraction network and the localization feature extraction network respectively. The last adaptive alignment module outputs the object classification feature map, the object offset feature map, and the object localization feature map.
[0032] The multi-stage cascaded adaptive alignment modules are used to gradually extract the object distribution and scale in the classification feature map and the localization feature map, thereby generating adaptive anchors that converge at the center position of the object. By using the adaptive anchors to adjust the preset evenly distributed anchor boxes, the natural spatial misalignment problem between the object localization box and the adjusted anchor boxes is gradually eliminated. At the same time, the multi-stage adaptive alignment modules continuously perform feature transformation on the classification feature map and the localization feature map, and can also gradually enhance the feature details of the object, extracting more detailed object feature information.
[0033] Specifically, the adaptive anchor alignment network is expressed as: , Among them, represents the first adaptive alignment module; represents the second adaptive alignment module; represents the th adaptive alignment module; represents a preset atrous prior, which is used to expand the receptive field without increasing the number of parameters based on the prior assumption of atrous convolution.
[0034] For example, Figure 4 As shown in the structural schematic diagram of the adaptive alignment module, it can be seen from the figure that each adaptive alignment module includes a standard convolution and a deformable convolution. The steps for performing deep feature extraction on the input classification feature map and localization feature map are as follows: Step 1: The standard convolution first performs a standard convolution operation on the input classification feature map and outputs an offset, which is expressed as: , Among them, represents the offset output by the th adaptive alignment module, represents the number of adaptive alignment modules; represents the standard convolution unit used to perform the standard convolution; represents the classification feature map input to the th adaptive alignment module, that is, the classification feature map output by the th adaptive alignment module. When , is the classification feature map output by the feature extraction network.
[0035] Optionally, each adaptive alignment module performs a standard convolution on the input classification feature map, and the output offset includes: performing a standard convolution on the classification feature map using a first convolution unit and a second convolution unit with different convolutional kernel sizes in series, and obtaining the offset based on the output of the second convolution unit.
[0036] Exemplarily, the convolutional kernel size of the first convolution unit is 3*3, and the convolutional kernel size of the second convolution unit is 1*1.
[0037] First, since the classification feature map is used to encode the semantic information of the target, such as the category of the target being a person, a vehicle, etc., the offset generated based on the classification feature map can guide the sampling area to adjust towards the area where the target of a specific category is located, enhancing the purposefulness of the anchor box adjustment, rather than guiding the anchor box to randomly move towards the areas where each target is located. Therefore, in this application, the offset is directly obtained based on the classification feature map.
[0038] Further, when performing a sliding window operation on the input classification feature map using standard convolution, offset amounts such as the offset size of the sampling region with respect to the center of the target region and the scaling ratio of the sampling width and height with respect to the true width and height of the target can be output. These offset amounts can then guide the sampling positions of the subsequent deformable convolution on the feature map to be offset.
[0039] Step 2: Add the offset amount to the preset hole prior to obtain a biased feature map, which is expressed as: , where, represents the biased feature map obtained by the th adaptive alignment module; represents the preset hole prior.
[0040] Step 3: The first deformable convolution simultaneously performs a deformable convolution operation on the biased feature map and the input classification feature map, outputs the classification feature map of the current adaptive alignment module, and uses it as the input for the next adaptive alignment module. The classification feature map output by the current adaptive alignment module is expressed as: , where, represents the classification feature map output by the th adaptive alignment module; represents the deformable convolution unit that performs the deformable convolution.
[0041] Since the role of the hole prior is to expand the receptive field without increasing the number of parameters based on the prior assumption of dilated convolution, a basic sampling grid is provided through the preset hole prior, so that the deformable convolution can fine-tune the sampling region in the classification feature map in combination with the offset amount on the basis of this sampling grid. For example, when detecting vehicles with different aspect ratios in an image, the hole prior selects the initial grid according to the target scale, and the offset amount offsets the grid points to key parts such as the front and rear of the vehicle, so that the output classification feature map is more focused on the core region of the target.
[0042] Step 4: The second deformable convolution simultaneously performs a deformable convolution operation on the offset amount and the input localization feature map, outputs the localization feature map of the current adaptive alignment module, and uses it as the input for the next adaptive alignment module. The localization feature map output by the current adaptive alignment module is expressed as: , where, represents the localization feature map output by the th adaptive alignment module; represents the localization feature map input to the th adaptive alignment module, that is, the The positioning feature map output by the adaptive alignment module is hour, It is the positioning feature map output by the feature extraction network.
[0043] Specifically, the adjustment of the positioning feature map only requires accurately regressing the area where the target is located in the positioning feature map to the real positioning frame, and only fine-tuning is required based on the offset. There is no need to introduce a hole prior to expand the receptive field. For example, when adjusting the sampling area containing the target in the positioning feature map, introducing a hole prior may cause the sampling area in the adjusted positioning feature map to exceed the range of the real positioning frame, thereby affecting the positioning accuracy.
[0044] Furthermore, the target bias feature map and the preset uniformly distributed anchor are input into the alignment anchor generator to output the adaptive anchor. The adaptive anchor is expressed as: , in, represents an adaptive anchor; Indicates pixel-by-pixel addition; A list of anchor points representing preset uniformly distributed anchors; represents the downsampling rate of each scale feature map in the multi-scale feature map, , Indicates the number of scale feature maps; Represents the average operator.
[0045] Since the target bias feature map is generated by combining the hole prior and the offset, it contains the core area position of the target, and the preset uniformly distributed anchors as the basic anchor frame grid provide an initial anchor frame set that uniformly covers the image. Therefore, the target bias feature map is used to adjust the uniformly distributed adaptive anchors to generate an anchor frame (i.e., adaptive anchor) that is highly matched with the scale, shape, and position of the target in the image. The adaptive anchor can guide the model to focus on the target area when performing classification and positioning based on the classification feature map and the positioning feature map, thereby improving the classification accuracy and positioning accuracy of the model.
[0046] It is worth noting that as the number of adaptive alignment modules increases, the anchor points of the generated adaptive anchors will converge more and more to the target center, thereby improving the target detection accuracy. However, this will also increase the complexity of the model and the computational cost, thereby affecting the target detection efficiency. Therefore, as a preferred option, the number of adaptive modules can be set to 3, which ensures the target detection accuracy while taking into account the detection efficiency as much as possible.
[0047] Before using an object detection model including a feature extraction network, an adaptive anchor alignment network, and a decoder to perform object detection on an image to be detected, it is also necessary to use the images in the training set to train the object detection model. In this application, a positive and negative sample assignment strategy is introduced when constructing an object detection loss function for training the model, so that the model can learn the common features of the same type of object samples from positive object samples, and learn the boundaries of different types of objects from negative object samples, thereby capturing the subtle differences between different types of objects.
[0048] Specifically, the construction process of the object detection loss function includes: 1. Using a sample assignment strategy paradigm to divide each object in the image based on the predicted localization box and predicted category of each object in the image, obtaining a positive object sample set and a negative object sample set.
[0049] Optionally, in some embodiments, the step of dividing objects based on the sample assignment strategy paradigm may be: obtaining the intersection over union (IoU) between the predicted localization box of each object and the ground truth localization box of its corresponding object, and the predicted category confidence of each object; based on the objects with an IoU greater than or equal to a preset IoU threshold and a predicted category confidence greater than or equal to a preset confidence threshold, obtaining a positive object sample set, and obtaining a negative object sample set based on the objects outside the positive object sample set. Specifically, the ground truth labels of the image include the ground truth localization boxes and ground truth categories of each object, and the center coordinates, width, height, and rotation angle information can be obtained based on the coordinate information of its ground truth localization box.
[0050] Exemplarily, objects can also be divided only based on the predicted localization box and the ground truth localization box of the object, and its formula is expressed as: .
[0051] 2. Based on the distance metric values between the predicted categories of each object in the positive object sample set and the negative object sample set and the object classification feature map of the image, constructing a classification loss function.
[0052] Specifically, the classification loss function is expressed as: , where, represents the classification loss function; represents a function for calculating the distance metric values between the predicted categories of each object in the positive object sample set and the negative object sample set and the object classification feature map of the image; represents the positive object sample set; represents the negative object sample set; represents the object classification feature map.
[0053] 3. Based on the distance metric values between the predicted localization boxes and the true localization boxes of each target in the positive target sample set, a localization loss function is constructed.
[0054] Specifically, the localization loss function is expressed as: , where, represents the localization loss function; represents a function for calculating the distance metric values between the predicted localization boxes and the true localization boxes of each target in the positive target sample set; represents the predicted localization box; represents the true localization box; represents filtering, which means filtering out the predicted localization boxes corresponding to each target in the positive target sample set from all the predicted localization boxes in the image.
[0055] 4. Based on the weighted sum of the classification loss function and the localization loss function, an object detection loss function is constructed.
[0056] Specifically, the object detection loss function is expressed as: , where, represents the object detection loss function; represents the weight of the localization loss function.
[0057] To verify the effectiveness of the adaptive anchor alignment network provided by this application, this embodiment also provides a comparison graph of the adaptive anchor points obtained by this application using the adaptive anchor alignment network and the traditional preset uniformly distributed anchors, as Figure 5 shown in the schematic diagram of the convergence comparison between the adaptive anchor points obtained by this application and the traditional uniformly distributed anchor points; where, Figure 5 in (a) is the schematic diagram of the traditional uniformly distributed anchor points generated on the image to be detected, Figure 5 in (b) is the schematic diagram of the object detection result obtained by detecting the image in (a) using the object detection model in the prior art, Figure 5 in (c) is the schematic diagram of the adaptive anchor points generated on the image to be detected when using 1 adaptive alignment module, Figure 5 in (d) is the schematic diagram of the object detection result obtained by detecting the image in Figure 5 (c) using the object detection model provided by this application including 1 adaptive alignment module, Figure 5 in (e) is the schematic diagram of the adaptive anchor points generated on the image to be detected when using 2 cascaded adaptive alignment modules, Figure 5 in (f) is the schematic diagram of the object detection result obtained by detecting the image in Figure 4 (e) using the object detection model provided by this application including 2 cascaded adaptive alignment modules,Figure 5 In (g), it is a schematic diagram of the adaptive anchor points generated on the image to be detected when using three adaptive alignment modules. Figure 5 In (h), it is a schematic diagram of the object detection result obtained by the object detection model provided by this application, which contains three adaptive alignment modules, for Figure 5 the detection of (g) in it.
[0058] From Figure 5 it can be seen from (a) and (b) in that the traditional uniformly distributed anchors do not focus on the area where the target is located in the image. From Figure 5 it can be seen from (c) and (d) in that the adaptive anchor points generated on the image to be detected when using one adaptive alignment module begin to gradually focus on the target in the image. From Figure 5 it can also be seen from (e), (f), (g) and (h) in that as the number of adaptive alignment modules increases, the generated adaptive anchor points become more and more focused on the target center. Therefore, the adaptive anchor alignment network introduced in this application can effectively make the anchor points converge to the area where the target is located, and it is also proved again that as the number of adaptive alignment modules increases, this convergence effect will be more obvious, so that the model can pay more attention to the area with dense anchor points (i.e., the target area) during the detection process, thereby improving the accuracy of object detection.
[0059] Figure 6 The following shows the loss curve graphs when the object detection model provided by this application and the object detection model in the prior art detect the same image; among them, Figure 6 in (a) represents the localization loss curve graph, Figure 6 in (b) represents the total detection loss curve graph based on the localization loss and the classification loss; the abscissa in the graph represents the number of training epochs, and the ordinate represents the loss value. It can be seen from the graph that whether it is the localization loss or the total detection loss, the object detection model provided by this application has a significantly faster convergence speed and a more stable and lower convergence result compared with the object detection model in the prior art, indicating the effectiveness of the model provided by this application for the object detection process.
[0060] As Figure 7 the following shows the comparison graph of the detection results when the object detection model provided by this application and the object detection model in the prior art detect different images; among them, Figure 7 in (a) is a schematic diagram of the result of detecting the first image using the object detection model in the prior art, Figure 7 in (b) is a schematic diagram of the result of detecting the first image using the object detection model of this application, Figure 7 in (c) is a schematic diagram of the result of detecting the second image using the object detection model in the prior art, Figure 7In (d), it is a schematic diagram of the detection result of using the object detection model of this application to detect the second image. Figure 7 In (e), it is a schematic diagram of the detection result of using the object detection model in the prior art to detect the third image. Figure 7 In (f), it is a schematic diagram of the detection result of using the object detection model of this application to detect the third image. Figure 7 In (g), it is a schematic diagram of the detection result of using the object detection model in the prior art to detect the fourth image. Figure 7 In (h), it is a schematic diagram of the detection result of using the object detection model of this application to detect the fourth image.
[0061] Through Figure 7 It can be seen that the method provided by this application has a high detection accuracy when detecting objects in different images. For example, by comparing Figure 7 (e) and (f) in it, the number of objects that can be detected by this application is much larger than that of the existing method in the case of a large number of objects gathering.
[0062] Furthermore, this embodiment also compared the detection results of multiple existing baseline models for different objects before and after introducing the adaptive anchor alignment network. Specifically, an adaptive alignment module (A 3 M) was introduced respectively on three baseline models of the RetinaNet network, the adaptive training sample selection method, and the Faster-RCNN. The detection results are shown in Table 1: Table 1
[0063] It can be seen from Table 1 that the performance of the RetinaNet network and the adaptive training sample selection method has increased by more than 3mAP after adding A 3 M. Specifically, for the three types of relatively small-scale and densely arranged objects, namely small vehicles, ships, and helicopters, there are obvious improvements. And, when deploying A 3 M to the multi-stage detector Faster-RCNN, an improvement of about 2mAP is also observed.
[0064] This embodiment also conducted ablation experiments on the number of standard convolutional units, the number of adaptive alignment modules, and the dilation rate of a single adaptive alignment module in the adaptive anchor alignment network. The experimental data are shown in Table 2: Table 2
[0065] As can be seen from the data in Table 2, under the parameter configuration where the number of adaptive alignment modules is 2, the number of standard convolutional units is 1, and the dilation rate of a single adaptive module is 2, the object detection model provided by this application has an mAP improvement of 2.6 compared to the baseline model in the prior art, and the number of parameters only increases by 1.1MB. This also indicates that the method provided by this application can improve the object detection accuracy without excessively increasing the computational cost, that is, it will not have a great impact on the object detection efficiency.
[0066] Based on the object detection method based on adaptive anchor alignment provided in the above embodiments, an embodiment of this application further provides an object detection device based on adaptive anchor alignment. The device specifically includes: A feature extraction module, configured to input the image to be detected into a feature extraction network and output a localization feature map and a classification feature map; An adaptive alignment module, configured to input the localization feature map and the classification feature map into an adaptive anchor alignment network. The adaptive anchor alignment network includes N serially connected adaptive alignment modules. Each adaptive alignment module performs adaptive alignment on the input feature map, including: Performing standard convolution on the input classification feature map to output an offset; obtaining a bias feature map based on the offset and a preset dilation prior; Performing deformable convolution on the bias feature map and the input classification feature map simultaneously to output the classification feature map to the next adaptive alignment module; Performing deformable convolution on the input localization feature map and the offset simultaneously to output the localization feature map to the next adaptive alignment module; An adaptive anchor generation module, configured to use the classification feature map, the localization feature map, and the bias feature map output by the last adaptive alignment module as the target classification feature map, the target localization feature map, and the target bias feature map respectively; generating adaptive anchors based on the target bias feature map and a preset uniformly distributed anchor; An object detection module, configured to obtain the predicted localization boxes and predicted categories of each object in the image to be detected based on the adaptive anchors, the target localization feature map, and the target classification feature map.
[0067] An embodiment of this application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above object detection method based on adaptive anchor alignment are implemented.
[0068] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0069] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0070] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means realizes the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0071] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0072] Obviously, the above embodiments are only examples given for clear illustration and are not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. And the obvious changes or modifications derived therefrom are still within the protection scope of the present invention.
Claims
1. An object detection method based on adaptive anchor alignment, characterized in that Including: Input the image to be detected into the feature extraction network, and output the localization feature map and the classification feature map; Input the localization feature map and the classification feature map into the adaptive anchor alignment network. The adaptive anchor alignment network includes N cascaded adaptive alignment modules. Each adaptive alignment module performs adaptive alignment on the input feature map, including: Perform standard convolution on the input classification feature map to output an offset; obtain a bias feature map based on the offset and a preset dilation prior; Perform deformable convolution on the bias feature map and the input classification feature map simultaneously, and output the classification feature map to the next adaptive alignment module; Perform deformable convolution on the input localization feature map and the offset simultaneously, and output the localization feature map to the next adaptive alignment module; Take the classification feature map, the localization feature map, and the bias feature map output by the last adaptive alignment module as the target classification feature map, the target localization feature map, and the target bias feature map respectively; generate adaptive anchors based on the target bias feature map and the preset uniformly distributed anchors; Input the adaptive anchors, the target localization feature map, and the target classification feature map into the decoder to obtain the predicted localization boxes and predicted classes of each target in the image to be detected.
2. The object detection method based on adaptive anchor alignment according to claim 1, characterized in that Each adaptive alignment module performs standard convolution on the input classification feature map to output an offset, including: performing standard convolution on the classification feature map using a first convolutional unit and a second convolutional unit with different convolutional kernel sizes in series, and obtaining the offset based on the output of the second convolutional unit.
3. The object detection method based on adaptive anchor alignment according to claim 1, wherein The offset output by each adaptive alignment module is expressed as: , Among them, represents the offset output by the -th adaptive alignment module, represents the number of adaptive alignment modules; represents the standard convolution unit used to perform standard convolution; represents the classification feature map input to the -th adaptive alignment module, that is, the classification feature map output by the -th adaptive alignment module. When , is the classification feature map output by the feature extraction network; The bias feature map obtained by each adaptive alignment module is expressed as: , Among them, represents the bias feature map obtained by the th adaptive alignment module; represents a preset hole prior; The classification feature map output by each adaptive alignment module is expressed as: , Among them, represents the classification feature map output by the th adaptive alignment module; represents the deformable convolution unit that performs deformable convolution; The localization feature map output by each adaptive alignment module is expressed as: , Among them, represents the localization feature map output by the th adaptive alignment module; represents the localization feature map input to the th adaptive alignment module, that is, the localization feature map output by the th adaptive alignment module. When , is the localization feature map output by the feature extraction network.
4. The object detection method based on adaptive anchor alignment according to claim 1, characterized in that Input the image to be detected into the feature extraction network, and output the localization feature map and the classification feature map, including: Input the image to be detected into the encoder for multi-level feature extraction, and output multi-scale feature maps; Input the multi-scale feature maps into the parallel classification feature extraction network and localization feature extraction network respectively, and output the classification feature map and the localization feature map.
5. The object detection method based on adaptive anchor alignment according to claim 4, wherein Generating adaptive anchors based on the target bias feature map and the preset uniformly distributed anchors includes: jointly inputting the target bias feature map and the preset uniformly distributed anchors into the alignment anchor generator to output adaptive anchors; where the adaptive anchors are expressed as: , Among them, represents an adaptive anchor; represents pixel-by-pixel addition; represents a list of anchor points for preset uniformly distributed anchors; represents the downsampling rate of each scale feature map in the multi-scale feature map, , represents the number of scale feature maps; represents the average operator.
6. The object detection method based on adaptive anchor alignment according to claim 1, wherein When training the feature extraction network, the adaptive anchor alignment network, and the decoder using the images in the training set, the construction process of the object detection loss function includes: Using the sample assignment strategy paradigm to divide each target in the image based on the predicted localization boxes and predicted classes of each target in the image, and obtain the positive target sample set and the negative target sample set; Construct a classification loss function based on the distance metric values between the predicted classes of each target in the positive target sample set and the negative target sample set and the target classification feature map of the image; Construct a localization loss function based on the distance metric values between the predicted localization boxes of each target in the positive target sample set and their true localization boxes; Construct the object detection loss function based on the weighted sum of the classification loss function and the localization loss function.
7. The object detection method based on adaptive anchor alignment according to claim 6, wherein The classification loss function is expressed as: , Among them, represents the classification loss function; represents a function for calculating the distance metric values between the predicted categories of each target in the positive target sample set and the negative target sample set and the target classification feature map of the image; represents the positive target sample set; represents the negative target sample set; represents the target classification feature map; The localization loss function is expressed as: , Among them, represents the positioning loss function; represents a function for calculating the distance metric value between the predicted positioning bounding box and the true positioning bounding box of each target in the positive target sample set; represents the predicted positioning bounding box; represents the true positioning bounding box; represents filtering; The object detection loss function is expressed as: , Among them, represents the object detection loss function; represents the weight of the localization loss function.
8. The object detection method based on adaptive anchor alignment according to claim 1, wherein, The adaptive anchor alignment network includes 3 cascaded adaptive alignment modules.
9. An object detection device based on adaptive anchor alignment, characterized in that, It includes: A feature extraction module, which is used to input the image to be detected into the feature extraction network and output a localization feature map and a classification feature map; An adaptive alignment module, which is used to input the localization feature map and the classification feature map into the adaptive anchor alignment network. Among them, the adaptive anchor alignment network includes N cascaded adaptive alignment modules. Each adaptive alignment module performs adaptive alignment on the input feature map, including: Performing standard convolution on the input classification feature map to output an offset; obtaining a biased feature map based on the offset and a preset atrous prior; Performing deformable convolution on the biased feature map and the input classification feature map simultaneously, and outputting the classification feature map to the next adaptive alignment module; Performing deformable convolution on the input localization feature map and the offset simultaneously, and outputting the localization feature map to the next adaptive alignment module; An adaptive anchor generation module, which is used to use the classification feature map, the localization feature map, and the biased feature map output by the last adaptive alignment module as the target classification feature map, the target localization feature map, and the target biased feature map respectively; generating adaptive anchors based on the target biased feature map and a preset uniformly distributed anchor; A target detection module, which is used to input the adaptive anchors, the target localization feature map, and the target classification feature map into a decoder to obtain the predicted localization boxes and predicted categories of each target in the image to be detected.
10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, the steps of the object detection method based on adaptive anchor alignment according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Task decoupling and adaptive point set strategy-based rotating target detection method
CN116935223A
Lake image segmentation method and computer readable storage medium
CN119107334A
Training method and training apparatus for rotating-ship target detection model, and storage medium
WO2023116631A1