Time-sensitive target detection method, device and equipment and computer readable storage medium
Through the combination of backbone network and feature pyramid network, multi-scale features are extracted using hollow convolution, which solves the problem of insufficient time-sensitive small-object detection accuracy, and achieves more efficient object detection and recognition.
Patent Information
- Application Number
- CN202510036017.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-16
AI Technical Summary
In complex battlefield environments, it is difficult for the bomb-loading platform to effectively detect and identify time-sensitive small targets with small target pixels and fast scale changes, resulting in insufficient detection accuracy.
The backbone network is used to extract the detection images at different scales, and the multi-scale feature maps are characterized by combining the feature pyramid network. The context information in different receptive fields is extracted using hollow convolutions with different hollow rates, and finally the target detection is performed through the detection head.
It improves the detection accuracy and performance of time-sensitive small targets, reduces the possibility of missed detection of extremely small targets, and enhances the detection ability of small targets.
Smart Images

Figure CN120014230A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of missile-borne target detection and identification, and specifically to a time-sensitive target detection method, device, equipment and computer-readable storage medium. Background Art
[0002] Facing the future combat methods under the new situation, the combat mode of weapon systems and the target requirements have also been newly expanded. Among them, ballistic weapons are required to gradually develop from achieving large-scale distributed ground fixed target matching and recognition to detecting and classifying time-sensitive targets such as aircraft on the apron, ships stationed in ports, and moving targets such as sea surface ship formations in complex and changeable battlefield environments.
[0003] Since small targets among time-sensitive targets (i.e. time-sensitive small targets) have the characteristics of few target pixels, fast target scale changes, and high real-time requirements for identification and tracking in the application environment of missile-borne platforms, how to effectively improve the detection accuracy of time-sensitive small targets is a problem that needs to be solved urgently. Summary of the invention
[0004] The present application provides a time-sensitive target detection method, device, equipment and computer-readable storage medium, which can effectively improve the detection accuracy of time-sensitive small targets.
[0005] In a first aspect, an embodiment of the present application provides a time-sensitive target detection method, the time-sensitive target detection method comprising:
[0006] Through the preset backbone network, features of different scales are extracted from the image to be detected to obtain a multi-scale feature map;
[0007] Based on a preset feature pyramid network, feature fusion is performed on the multi-scale feature map to obtain a target feature map, wherein the feature pyramid network includes a target aggregation module, and the target aggregation module includes a plurality of dilated convolutions with different dilated rates, and the plurality of dilated convolutions are used to extract context information under different receptive fields;
[0008] Target detection is performed on the target feature map according to a preset detection head to output a time-sensitive target detection result.
[0009] In combination with the first aspect, in one embodiment, the target aggregation module includes a first aggregation branch, a second aggregation branch, a third aggregation branch, a fourth aggregation branch and a fusion module, the fusion module is used to fuse the output results of the four aggregation branches, the first aggregation branch includes a convolutional layer, the second aggregation branch, the third aggregation branch and the fourth aggregation branch are all composed of sequentially connected convolutional layers, void convolutions and convolutional layers, wherein the void rates corresponding to the void convolutions in the second aggregation branch, the third aggregation branch and the fourth aggregation branch increase sequentially, and the convolutional layer is used to perform convolution processing on the input feature map to reduce the number of channels.
[0010] In combination with the first aspect, in one implementation, the backbone network is a Darknet53 network, and the multi-scale feature map includes a shallow feature map, a mid-level feature map, and a deep feature map.
[0011] In combination with the first aspect, in one embodiment, the target aggregation module is used to extract contextual information of shallow feature maps under different receptive fields and fuse them to output an enhanced shallow feature map, and use the enhanced shallow feature map as one of the target feature maps.
[0012] In combination with the first aspect, in one implementation, the number of the target aggregation modules is three;
[0013] The first target aggregation module is used to extract the context information of shallow feature maps under different receptive fields and fuse them to output enhanced shallow feature maps;
[0014] The second target aggregation module is used to extract the context information of the middle-level feature maps under different receptive fields and fuse them to output the enhanced middle-level feature maps;
[0015] The third target aggregation module is used to extract the context information of the deep feature maps under different receptive fields and fuse them to output the enhanced deep feature maps;
[0016] Among them, the enhanced shallow feature map, the enhanced middle feature map and the enhanced deep feature map are all target feature maps.
[0017] In combination with the first aspect, in one implementation, the detection head includes a basic convolution module and a convolution layer connected in sequence, and the basic convolution module includes a convolution layer, a batch normalization layer, and an activation layer connected in sequence.
[0018] In a second aspect, an embodiment of the present application provides a time-sensitive target detection device, the time-sensitive target detection device comprising:
[0019] The backbone network is used to extract features of different scales from the image to be detected to obtain a multi-scale feature map;
[0020] A feature pyramid network is used to perform feature fusion on multi-scale feature maps to obtain a target feature map. The feature pyramid network includes a target aggregation module, and the target aggregation module includes multiple dilated convolutions with different dilated rates. The multiple dilated convolutions are used to extract context information under different receptive fields.
[0021] A detection head is used to perform target detection on the target feature map to output a time-sensitive target detection result.
[0022] In combination with the second aspect, in one embodiment, the target aggregation module includes a first aggregation branch, a second aggregation branch, a third aggregation branch, a fourth aggregation branch and a fusion module, the fusion module is used to fuse the output results of the four aggregation branches, the first aggregation branch includes a convolutional layer, the second aggregation branch, the third aggregation branch and the fourth aggregation branch are all composed of sequentially connected convolutional layers, void convolutions and convolutional layers, wherein the void rates corresponding to the void convolutions in the second aggregation branch, the third aggregation branch and the fourth aggregation branch increase sequentially, and the convolutional layer is used to perform convolution processing on the input feature map to reduce the number of channels.
[0023] In combination with the second aspect, in one implementation, the backbone network is a Darknet53 network, and the multi-scale feature map includes a shallow feature map, a mid-level feature map, and a deep feature map.
[0024] In combination with the second aspect, in one embodiment, the target aggregation module is used to extract contextual information of shallow feature maps under different receptive fields and fuse them to output an enhanced shallow feature map, and use the enhanced shallow feature map as one of the target feature maps.
[0025] In conjunction with the second aspect, in one implementation, the number of the target aggregation modules is three;
[0026] The first target aggregation module is used to extract the context information of shallow feature maps under different receptive fields and fuse them to output enhanced shallow feature maps;
[0027] The second target aggregation module is used to extract the context information of the middle-level feature maps under different receptive fields and fuse them to output the enhanced middle-level feature maps;
[0028] The third target aggregation module is used to extract the context information of the deep feature maps under different receptive fields and fuse them to output the enhanced deep feature maps;
[0029] Among them, the enhanced shallow feature map, the enhanced middle feature map and the enhanced deep feature map are all target feature maps.
[0030] In combination with the second aspect, in one implementation, the detection head includes a basic convolution module and a convolution layer connected in sequence, and the basic convolution module includes a convolution layer, a batch normalization layer, and an activation layer connected in sequence.
[0031] In the third aspect, an embodiment of the present application provides a time-sensitive target detection device, which includes a processor, a memory, and a time-sensitive target detection program stored in the memory and executable by the processor, wherein when the time-sensitive target detection program is executed by the processor, the steps of the aforementioned time-sensitive target detection method are implemented.
[0032] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a time-sensitive target detection program is stored, wherein when the time-sensitive target detection program is executed by a processor, the steps of the aforementioned time-sensitive target detection method are implemented.
[0033] The beneficial effects brought by the technical solution provided in the embodiments of the present application include:
[0034] Through the backbone network, the image to be detected is subjected to feature extraction at different scales to obtain a multi-scale feature map, so as to adapt to the multi-scale changes of the detection target, so that the detection of objects with fast scale changes is more robust; then, the multi-scale feature map is subjected to feature fusion based on the feature pyramid network, so as to transmit the semantic information of the top layer from top to bottom to the lower layer, and then enhance the feature expression ability of the lower layer, thereby improving the detection ability of the lower layer features for the target; wherein, the context information of the feature map under different receptive fields is extracted through multiple hole convolutions with different hole rates, that is, the receptive fields of different scales are obtained, so as to enhance the feature information of the feature map while keeping the resolution of the feature map unchanged, thereby enhancing the detection ability of small targets; finally, the target feature map output by the feature pyramid network is subjected to target detection by the detection head, and the time-sensitive target detection result can be output. It can be seen that the detection accuracy and performance of time-sensitive small targets can be effectively improved through this application, so as to reduce the possibility of missing detection of extremely small targets. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 This is a flow chart of an embodiment of the time-sensitive target detection method of the present application;
[0036] Figure 2 This is a schematic diagram of the overall architecture of time-sensitive target detection involved in the embodiment of the present application;
[0037] Figure 3 This is a schematic diagram of the basic convolution module structure involved in the embodiment of the present application;
[0038] Figure 4 This is a schematic diagram of the residual block structure involved in the embodiment of the present application;
[0039] Figure 5 A schematic diagram of the residual unit structure involved in the embodiment of the present application;
[0040] Figure 6 This is a schematic diagram of feature pyramid network feature processing involved in the embodiment of the present application;
[0041] Figure 7 This is a schematic diagram of the MDCFB structure involved in the embodiment of the present application;
[0042] Figure 8 A schematic diagram of a feature pyramid network structure involved in an embodiment of the present application;
[0043] Fig. 9 This is a schematic diagram of another feature pyramid network structure involved in the embodiment of the present application;
[0044] Fig.10 A schematic diagram of information calculation involved in the embodiment of the present application;
[0045] Fig.11 This is a schematic diagram of the hardware structure of the time-sensitive target detection device involved in the embodiment of the present application. DETAILED DESCRIPTION
[0046] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0047] In order to make the objectives, technical solutions and advantages of the present application clearer, the implementation methods of the present application will be further described in detail below with reference to the accompanying drawings.
[0048] In a first aspect, an embodiment of the present application provides a time-sensitive target detection method.
[0049] In one embodiment, referring to Figure 1 , Figure 1 This is a flow chart of an embodiment of the time-sensitive target detection method of this application. Figure 1 As shown, the time-sensitive target detection method includes:
[0050] Step S10: extracting features of different scales from the image to be detected through a preset backbone network to obtain a multi-scale feature map; wherein the backbone network is a Darknet53 network, and the multi-scale feature map includes a shallow feature map, a middle feature map, and a deep feature map.
[0051] For example, it should be noted that the image to be detected refers to an image containing a time-sensitive target that needs to be detected; in this embodiment, the Darknet53 network without a fully connected layer (i.e., Darknet-53without FClayer) can be preferably used as the backbone network to realize feature extraction at different scales, and then obtain multi-scale feature maps including shallow feature maps, middle feature maps and deep feature maps to adapt to the multi-scale changes of the detection target, so that the detection of objects with fast scale changes is more robust.
[0052] See also Figure 2 As shown, the Darknet53 network includes a basic convolution module CBL and a series of residual blocks Res; see Figure 3 As shown in the figure, the basic convolution module CBL (i.e., Bcse Conv Block) is composed of a convolution layer Conv, a batch normalization layer BN, and an activation layer LeakyRelu (a nonlinear activation function) connected in sequence. The convolution layer Conv is used to perform convolution processing on the input feature map to extract local features in the image; the batch normalization layer BN is used to normalize the output of the convolution layer Conv to improve convergence and reduce the risk of overfitting; the activation layer LeakyRelu is used to introduce nonlinear factors to activate the output of the batch normalization layer BN, thereby better retaining feature information and helping to alleviate the problem of gradient disappearance.
[0053] See also Figure 4 As shown in , the residual block Res consists of a basic convolution module CBL and a residual unit Res_unit, where n in Res×n represents the number of residual units Res_unit. For example, a residual block Res×5 consists of a basic convolution module CBL and 5 residual units Res_unit; see Figure 5 As shown, the residual unit Res_unit includes two basic convolution modules CBL connected in sequence, so that the initial input features are processed in sequence by the two basic convolution modules CBL, and the output results thereof are residually connected with the initial input features to enhance the feature extraction capability.
[0054] For details, see Figure 2 As shown in the figure, after the image to be detected is input into the Darknet53 network, it will be processed by the basic convolution module CBL, residual block Res×1, residual block Res×2, residual block Res×8, residual block Res×8 and residual block Res×4 in sequence, and then the deep feature map P5 will be output; among them, after the feature extraction of the first residual block Res×8, the shallow feature map P3 will be output, and after the feature extraction of the second residual block Res×8, the middle feature map P4 will be output.
[0055] Step S20: performing feature fusion on the multi-scale feature map based on a preset feature pyramid network to obtain a target feature map, wherein the feature pyramid network includes a target aggregation module, and the target aggregation module includes multiple dilated convolutions with different dilated rates, and the multiple dilated convolutions are used to extract contextual information under different receptive fields.
[0056] For example, in this embodiment, the multi-scale feature graph output by the Darknet53 network is subjected to feature fusion using a feature pyramid structure, so as to transmit the semantic information of the top layer to the lower layer from top to bottom, thereby enhancing the feature expression capability of the lower layer, thereby improving the detection capability of the lower layer features for the target. Figure 6 The feature pyramid network FPN shown in the figure realizes feature fusion. Specifically, after feature enhancement of the deep feature map P5, an enhanced deep feature map P5′ is obtained, and the enhanced deep feature map P5′ is scaled up by the upsampling module Upsample and then fused and feature enhanced with the middle feature map P4 to obtain the enhanced middle feature map P4′; similarly, the enhanced middle feature map P4′ is scaled up by Upsample and then fused and feature enhanced with the shallow feature map P3 layer to obtain the enhanced shallow feature map P3′, so as to finally output the three target feature maps P3′, P4′ and P5′.
[0057] It should be understood that the feature pyramid network in this embodiment includes a multi-branch context information aggregation module (Multi-branch Dilate Convolution Context Fusion Block, MDCFB, i.e., target aggregation module) based on dilated convolution, and the MDCFB includes multiple dilated convolutions with different dilated rates, so as to extract context information in feature maps under different receptive fields through multiple dilated convolutions with different dilated rates, and then obtain receptive fields of different scales, thereby enhancing the feature information of the feature map while keeping the resolution of the feature map unchanged, thereby enhancing the detection capability of small targets. It should be noted that the specific values of the dilated rates corresponding to different dilated convolutions can be determined according to actual needs, and are not limited here, as long as different dilated convolutions have different dilated rates.
[0058] Furthermore, in one embodiment, the target aggregation module includes a first aggregation branch, a second aggregation branch, a third aggregation branch, a fourth aggregation branch and a fusion module, the fusion module is used to fuse the output results of the four aggregation branches, the first aggregation branch includes a convolutional layer, the second aggregation branch, the third aggregation branch and the fourth aggregation branch are all composed of sequentially connected convolutional layers, hole convolutions and convolutional layers, wherein the hole rates corresponding to the hole convolutions in the second aggregation branch, the third aggregation branch and the fourth aggregation branch increase sequentially, and the convolutional layer is used to perform convolution processing on the input feature map to reduce the number of channels.
[0059] For example, see Figure 7 As shown in the figure, MDCFB includes a first aggregation branch, a second aggregation branch, a third aggregation branch, a fourth aggregation branch and a fusion module; wherein the first aggregation branch includes a convolution layer with a convolution kernel of 1×1, and the second aggregation branch, the third aggregation branch and the fourth aggregation branch are composed of a convolution layer with a convolution kernel of 1×1, a dilated convolution with a convolution kernel of 3×3 (i.e., DilateConv) and a convolution layer with a convolution kernel of 1×1 connected in sequence; it should be understood that the dilated rate r corresponding to the dilated convolution in the second aggregation branch, the third aggregation branch and the fourth aggregation branch increases in sequence. For example, the dilation rate r of the dilated convolution in the second aggregation branch is set to 1 to retain the original receptive field of the feature map, the dilation rate r of the dilated convolution in the third aggregation branch is set to 3 to increase the receptive field of the feature map, and the dilation rate of the dilated convolution in the fourth aggregation branch can be set to 5 to further increase the receptive field of the feature map; based on this, the feature information of the feature map can be effectively enhanced to enhance the detection capability of small targets; the fusion module is used to fuse the output results of the four aggregation branches and unify the channels through a convolution layer with a convolution kernel of 1×1 to output the enhanced feature map. It should be noted that the convolution layers in MDCFB are all used to perform convolution processing on their inputs to reduce the number of channels.
[0060] Furthermore, in one embodiment, the target aggregation module is used to extract contextual information of shallow feature maps under different receptive fields and fuse them to output an enhanced shallow feature map, and use the enhanced shallow feature map as one of the target feature maps.
[0061] For example, see Figure 8 As shown, in this embodiment, MDCFB can be preferably applied to the shallow output of feature fusion, that is, using dilated convolutions with different dilation rates to obtain and fuse the contextual information under different receptive fields in the shallow feature map P3 while keeping the feature map resolution unchanged, so as to enhance the contextual information contained in the shallow feature map to obtain P3′; and for the deep feature map P5 and the middle feature map P4, the CBL×5 structure can be used to achieve feature enhancement and fusion.
[0062] Specifically, first use CBL×5 to extract the deep feature P5 to output the enhanced deep feature map P5′; then the enhanced deep feature map P5′ is processed by the CBL module, and then input into the Upsample module for scale enlargement and fusion splicing with the middle feature map P4, and then the fusion splicing result is processed by CBL×5 to output the enhanced middle feature map P4′; similarly, the enhanced middle feature map P4′ is processed by the CBL module, and then input into the Upsample module for scale enlargement and fusion splicing with the shallow feature map P3, and finally, the contextual information of the shallow fusion splicing result under different receptive fields is extracted and fused based on the MDCFB module to output the enhanced shallow feature map P3′.
[0063] It should be understood that since the dilated convolution in MDCFB has different dilation rates, it can obtain image features under different receptive fields, adapt to the multi-scale changes of the detection target, and improve the robustness of network detection; and multi-scale fusion of features extracted under different receptive fields to obtain more image context information, thereby achieving the effect of feature enhancement. In addition, compared with the number of channels used by the five-layer convolution (i.e., CBL×5) in the basic network, the four aggregation branches in MDCFB can preferably adopt a parallel structure, which can use fewer feature channels to reduce the parameters of the backbone network, thereby effectively reducing the amount of algorithm calculation.
[0064] Further, in one embodiment, the number of the target aggregation modules is three;
[0065] The first target aggregation module is used to extract the context information of shallow feature maps under different receptive fields and fuse them to output enhanced shallow feature maps;
[0066] The second target aggregation module is used to extract the context information of the middle-level feature maps under different receptive fields and fuse them to output the enhanced middle-level feature maps;
[0067] The third target aggregation module is used to extract the context information of the deep feature maps under different receptive fields and fuse them to output the enhanced deep feature maps;
[0068] Among them, the enhanced shallow feature map, the enhanced middle feature map and the enhanced deep feature map are all target feature maps.
[0069] For example, see Fig. 9As shown, this embodiment can preferably apply MDCFB to all feature fusion layer outputs, that is, using dilated convolutions with different dilation rates to obtain contextual information under different receptive fields in the shallow feature map P3, the deep feature map P5 and the middle feature map P4 while keeping the feature map resolution unchanged, so as to enhance the contextual information contained in the feature maps of each level.
[0070] Specifically, first, MDCFB is used to extract and fuse the context information of the deep feature P5 under different receptive fields to output the enhanced deep feature map P5′; then, the enhanced deep feature map P5′ is processed by the CBL module, and then input into the Upsample module for scale enlargement and fusion splicing with the middle feature map P4, and then the fusion splicing result is extracted and fused by MDCFB under different receptive fields, and the enhanced middle feature map P4′ is output; similarly, the enhanced middle feature map P4′ is processed by the CBL module, and then input into the Upsample module for scale enlargement and fusion splicing with the shallow feature map P3, and finally, the shallow fusion splicing result is extracted and fused under different receptive fields based on MDCFB, and the enhanced shallow feature map P3′ is output. It should be understood that by replacing the feature fusion shallow output branch DBL×5 structure with MDCFB, more context information can be obtained during feature fusion to enhance the feature information of the shallow feature map, thereby enhancing the detection capability of small targets.
[0071] Step S30: performing target detection on the target feature map according to a preset detection head to output a time-sensitive target detection result; wherein the detection head includes a basic convolution module and a convolution layer connected in sequence, and the basic convolution module includes a convolution layer, a batch normalization layer and an activation layer connected in sequence.
[0072] Exemplarily, in this embodiment, the detection head is used to detect the feature maps of three layers with different resolutions output by the feature pyramid network FPN to output the time-sensitive target detection results (i.e., y1 to y3). In order to accurately detect the extracted features, in this embodiment, the detection head may preferably be composed of a DBL module and a convolution with a convolution kernel of 1×1 to output the target information such as the horizontal coordinate, vertical coordinate, width, and height of the center point of the prediction result, and the sigmoid function is used to directly map the output target confidence and category confidence to between 0 and 1, and the target information, target confidence, and category confidence are used as the time-sensitive target detection results, thereby completing the detection of time-sensitive targets.
[0073] In summary, this embodiment uses the backbone network to extract features of different scales on the image to be detected to obtain a multi-scale feature map to adapt to the multi-scale changes of the detection target, so that the detection of objects with fast scale changes is more robust; then the multi-scale feature map is feature fused based on the feature pyramid network to pass the semantic information of the top layer from top to bottom to the lower layer, and then enhance the feature expression ability of the lower layer, thereby improving the detection ability of the lower layer features for the target; wherein, the context information of the feature map under different receptive fields is extracted through multiple hole convolutions with different hole rates, that is, the receptive fields of different scales are obtained, so as to enhance the feature information of the feature map while keeping the resolution of the feature map unchanged, thereby enhancing the detection ability of small targets; finally, the target feature map output by the feature pyramid network is detected by the detection head, and the time-sensitive target detection result can be output. It can be seen that the detection accuracy and performance of time-sensitive small targets can be effectively improved through this embodiment to reduce the possibility of missing detection of extremely small targets.
[0074] It should be understood that the Darknet53 network, the feature pyramid network FPN including MDCFB, and the detection head in this embodiment constitute an end-to-end time-sensitive target detection model. The following embodiments will explain the training and generation process and principles of the time-sensitive target detection model.
[0075] First, a data set is constructed; specifically, the drone aerial photography data set VisDrone can be used. This data set contains a total of 10 categories. However, this embodiment only requires drone aerial photography images with large depth of field corresponding to different scenes and different densities. Therefore, large depth of field images corresponding to 4 types of targets (i.e., cars, vans, trucks, and buses) can be selected to construct a training set, and the image resolution can be uniformly changed to 1024×1024 to obtain a 4-classification large depth of field VisDrone data set, in which the number of channels of each image is 3, i.e., including three channels of R, G, and B. For the large depth of field VisDrone data set, it can be divided into a training set and a test set in a ratio of 8:2.
[0076] Among them, in order to increase the number of expanded samples, this embodiment may preferably use a mosaic method to perform data augmentation on the training set data, so as to expand the number of multi-scale target samples in the same image; specifically, a plurality of photos are randomly selected from the training set, and the plurality of photos are subjected to image preprocessing (such as rotation, cropping, enlargement, reduction, etc.) and then randomly spliced into a single detection image, and the target label on the single detection image should be consistent with the target label of the original image, thereby increasing the style of targets of different scales on the same image, and enhancing the robustness of the network to scale change scene detection.
[0077] Secondly, an optimized network architecture is constructed; specifically, the network structure of the Yolov3 algorithm can be optimized, that is, the MDCFB structure is used to replace the DBL×5 structure at the shallow output branch of the feature fusion in the Yolov3 algorithm to obtain an optimized multi-branch context network structure based on dilated convolution, and then an improved detection network based on dilated convolution is obtained; among them, CBL is the basic component, which consists of a convolution with a convolution kernel of 3×3, a batch normalization layer BN and a LeakyRelu activation layer; the MDCFB module has four branches, three of which use dilated convolutions with dilated rates of 1, 3, and 5 respectively to obtain contextual information under three receptive fields of 3, 7, and 11, thereby enhancing the contextual information contained in the shallow feature map.
[0078] Next, the improved detection network based on dilated convolution can be trained based on the Pytorch platform, and the position loss L can be constructed. Ioc , category loss L cls , target loss L obj The loss function consists of three parts, and the optimization direction is constrained by the loss function to obtain the weight file with the highest average accuracy. Specifically, in the training stage, the image data of size 1024×1024×3 in the training set is input into the improved detection network based on dilated convolution for training; among them, the Darknet53 network will extract features from the input image; the feature pyramid network performs feature fusion on the output of the Darknet53 network to pass the top-level semantic information from top to bottom to the lower layer, which can enhance the feature expression ability of the lower layer and improve the detection ability of the lower layer features for the target, and obtain more contextual information through MDCFB during feature fusion to enhance the feature information of the shallow feature map, thereby enhancing the detection ability of small targets; finally, the network detection head will perform feature detection on the output of the feature pyramid network to output the horizontal coordinate, vertical coordinate, width, height and other information of the center point of the prediction result.
[0079] It can be understood that in the feature detection part, three layers of feature maps with different resolutions are used for detection. Specifically, the entire image is divided into multiple grid areas, and each area is responsible for detecting the target in its center. Specifically, three anchor boxes are generated at the center of the area, and detection output is performed based on these anchor boxes. Since the regional grid points are used to predict the target, the coordinate information output is based on the center offset of each point and the scale ratio of length and width. Fig.10 As shown, the dotted box is the position of the anchor box, the dotted box is the position of the prediction box, and c x 、c y is the coordinate information of the grid where the prediction point is located, p w 、p h is the width and height of the corresponding anchor box, b x , by , b w , b h is the position information after information solution, and the corresponding calculation formulas are:
[0080] b x =sigmoid(t x )+c x
[0081] b y =sigmoid(t y )+c y
[0082]
[0083] Where, t x ,t y ,t w ,t h is the coordinate information output by the current grid, and sigmoid is the activation function. For the target confidence and category confidence output by the network, this embodiment directly uses the sigmoid function to map them to between 0 and 1.
[0084] Then, the samples are allocated; specifically, after the output feature map is solved, the obtained results will be selected for positive and negative samples, where the positive sample selection process is: first find out on which area the center point of the target's true position coordinates (GroundTruth, GT) is located, and then calculate the IoU (Intersection over Union) of the 9 anchor boxes (anchor) generated in the area and the GT, and record the anchor with the largest IoU as the positive sample; the negative sample selection process is: calculate the IoU of each anchor and the GT of all targets, and select the part of the anchor whose maximum GT IoU is less than the threshold (generally set to 0.5) and is not a positive sample as a negative sample; for anchors that are not positive samples and have a large IoU with the GT, they are recorded as ignored samples and are not used for subsequent calculations; then use the loss function optimization to update the network parameters.
[0085] Among them, for the loss function, the loss function (Loss) of this embodiment is preferably composed of the position loss L Ioc , category loss L cls , target loss L obj It consists of three parts, and is calculated using the mean square error function, the cross entropy error function, and the cross entropy error function respectively. The specific calculation formulas are:
[0086]
[0087]
[0088] Loss = λ loc ×L loc +L cls +L obj
[0089] In the formula, S 2 Represents the size of the current feature map; B is the number of detection boxes generated on each feature map grid; To determine whether the jth detection box in the i-th feature map grid is a positive sample, if it is, it is 1, otherwise it is 0; To determine whether the jth detection box in the i-th feature map grid is a negative sample, if it is, it is 1, otherwise it is 0; D i,j is the scale compensation coefficient of the target; x i ,y i 、w i 、h i 、c i They are the center point coordinates, width, height and confidence of GT in the current feature map grid; The center point coordinates, width, height and confidence of the detection box in the current feature map grid; is the category probability of GT; p i (c) is the category probability output by the detection box; λ loc , cls , noobj They are the proportional weights of position loss, category loss, and negative sample confidence loss, except λ loc The default value is 0.5, and the default value for other parameters is 1.
[0090] After the model training is completed, the time-sensitive target detection model can be generated, and then the test phase can be entered, that is, the target detection is performed on the data in the test set through the time-sensitive target detection model; specifically, the image data in the test set is input into the trained time-sensitive target detection model to output the information of the prediction box (obj,t x ,t y ,f h ,t w ,c i ), where obj represents the target type; then the prediction boxes with low scores are filtered out through a preset threshold, and the retained prediction boxes are processed by non-maximum suppression (NMS) to obtain the final detection result.
[0091] In order to further verify the algorithm's detection performance for small targets and its adaptability to multi-scale target detection, this embodiment will also evaluate the performance of the time-sensitive target detection model; among them, the Average Precision (AP) indicator can be used to evaluate the performance of the detection algorithm, and the model parameter amount and calculation amount are used to evaluate the model's parameter size and the model's computational complexity, and the number of frames per second (FPS) is used to evaluate the model's detection speed.
[0092] It should be understood that for the target detection task, the model outputs the location, category and confidence of the target in the image. Therefore, the intersection of the output target box and the real target box GT is used. The size of the output confidence can reflect the credibility of the output result. Therefore, each detection result can be divided into positive and negative samples based on IoU and confidence. The precision (Precision, P) and recall (Recall, R) of the detection result are calculated. The definitions of P and R can be expressed as:
[0093]
[0094] Where D TP Represents the number of correctly detected targets; D FN Represents the number of undetected targets; D FP Represents the number of targets detected incorrectly; using different confidence thresholds to obtain different P and R values can draw a PR curve, and the performance AP of the detector is represented based on the area of the PR curve. The calculation formula of AP is:
[0095]
[0096] In addition, it should be understood that the model parameter quantity refers to the number of weight values used by the model, mainly including the number of parameters of the convolutional layer in the model and the number of parameters contained in the BN layer, and the actual size of the model is often the model parameter quantity multiplied by the number of bytes occupied by the model parameter storage. Therefore, the model parameter quantity can be used as a benchmark to measure the size of the model; the model computational amount is the number of calculations performed by the model during forward reasoning, usually expressed in floating point operations (FLOPs), which can be used to measure the complexity of the model; FPS (Frames Per Second) is the number of images that the model can detect per second, which can measure the detection speed of the model.
[0097] Based on this, this embodiment conducts a performance comparison experiment on a Yolov3-based model and a MDCFB-based time-sensitive target detection model. The performance comparison results are shown in Tables 1 and 2.
[0098] Table 1 Comparison of model indicators
[0099]
[0100] Table 2. Comparison of complexity and speed
[0101] Model Parameter quantity / M Computing capacity / GFLOPS FPS Basic Model 61.54 207.48 34.7 MDCFB 61.21 193.21 35.1
[0102] It can be seen from Table 1 and Table 2 that, compared with the basic model, the model provided by this embodiment has advantages in average precision AP, parameter amount, calculation amount and transmission frame number per second FPS; it can be seen that this embodiment not only provides an end-to-end time-sensitive target detection algorithm that can adapt to large changes in target scale and a large number of concentrated small targets, improves the performance of small target detection, but also reduces the model parameter amount and calculation amount, effectively improves the detection speed, and can meet the needs of intelligent time-sensitive target detection and recognition under missile-borne platform conditions.
[0103] In a second aspect, an embodiment of the present application also provides a time-sensitive target detection device.
[0104] In one embodiment, the time-sensitive target detection device includes:
[0105] The backbone network is used to extract features of different scales from the image to be detected to obtain a multi-scale feature map;
[0106] A feature pyramid network is used to perform feature fusion on multi-scale feature maps to obtain a target feature map. The feature pyramid network includes a target aggregation module, and the target aggregation module includes multiple dilated convolutions with different dilated rates. The multiple dilated convolutions are used to extract context information under different receptive fields.
[0107] A detection head is used to perform target detection on the target feature map to output a time-sensitive target detection result.
[0108] Furthermore, in one embodiment, the target aggregation module includes a first aggregation branch, a second aggregation branch, a third aggregation branch, a fourth aggregation branch and a fusion module, the fusion module is used to fuse the output results of the four aggregation branches, the first aggregation branch includes a convolutional layer, the second aggregation branch, the third aggregation branch and the fourth aggregation branch are all composed of sequentially connected convolutional layers, hole convolutions and convolutional layers, wherein the hole rates corresponding to the hole convolutions in the second aggregation branch, the third aggregation branch and the fourth aggregation branch increase sequentially, and the convolutional layer is used to perform convolution processing on the input feature map to reduce the number of channels.
[0109] Furthermore, in one embodiment, the backbone network is a Darknet53 network, and the multi-scale feature map includes a shallow feature map, a middle feature map, and a deep feature map.
[0110] Furthermore, in one embodiment, the target aggregation module is used to extract contextual information of shallow feature maps under different receptive fields and fuse them to output an enhanced shallow feature map, and use the enhanced shallow feature map as one of the target feature maps.
[0111] Further, in one embodiment, the number of the target aggregation modules is three;
[0112] The first target aggregation module is used to extract the context information of shallow feature maps under different receptive fields and fuse them to output enhanced shallow feature maps;
[0113] The second target aggregation module is used to extract the context information of the middle-level feature maps under different receptive fields and fuse them to output the enhanced middle-level feature maps;
[0114] The third target aggregation module is used to extract the context information of the deep feature maps under different receptive fields and fuse them to output the enhanced deep feature maps;
[0115] Among them, the enhanced shallow feature map, the enhanced middle feature map and the enhanced deep feature map are all target feature maps.
[0116] Furthermore, in one embodiment, the detection head includes a basic convolution module and a convolution layer connected in sequence, and the basic convolution module includes a convolution layer, a batch normalization layer and an activation layer connected in sequence.
[0117] Among them, the functional implementation of each module in the above-mentioned time-sensitive target detection device corresponds to the various steps in the above-mentioned time-sensitive target detection method embodiment, and its functions and implementation processes will not be repeated here one by one.
[0118] In a third aspect, an embodiment of the present application provides a time-sensitive target detection device, which may be a personal computer (PC), a laptop computer, a server, or other device with data processing capabilities.
[0119] Reference Fig.11 , Fig.11 Schematic diagram of the hardware structure of the time-sensitive target detection device involved in the embodiment of the present application. In the embodiment of the present application, the time-sensitive target detection device may include a processor, a memory, a communication interface and a communication bus.
[0120] The communication bus may be of any type and is used to interconnect the processor, the memory, and the communication interface.
[0121] The communication interface includes input / output (I / O) interface, physical interface and logical interface, etc., which are used to realize the interconnection of devices inside the time-sensitive target detection device, and the interface used to realize the interconnection of the time-sensitive target detection device with other devices (such as other computing devices or user devices). The physical interface can be an Ethernet interface, a fiber optic interface, an ATM interface, etc.; the user device can be a display, a keyboard, etc.
[0122] The memory can be various types of storage media, such as random access memory (RAM), read-only memory (ROM), non-volatile RAM (NVRAM), flash memory, optical storage, hard disk, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), etc.
[0123] The processor may be a general-purpose processor, and the general-purpose processor may call the time-sensitive target detection program stored in the memory and execute the time-sensitive target detection method provided in the embodiment of the present application. For example, the general-purpose processor may be a central processing unit (CPU). The method executed when the time-sensitive target detection program is called may refer to the various embodiments of the time-sensitive target detection method of the present application, and will not be repeated here.
[0124] Those skilled in the art will understand that Fig.11 The hardware structure shown in the figure does not constitute a limitation on the present application, and may include more or less components than shown in the figure, or combine certain components, or arrange the components differently.
[0125] In a fourth aspect, an embodiment of the present application also provides a computer-readable storage medium.
[0126] A time-sensitive target detection program is stored on a readable storage medium of the present application, wherein when the time-sensitive target detection program is executed by a processor, the steps of the time-sensitive target detection method as described above are implemented.
[0127] Among them, the method implemented when the time-sensitive target detection program is executed can refer to the various embodiments of the time-sensitive target detection method of the present application, and will not be repeated here.
[0128] It should be noted that the serial numbers of the above-mentioned embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments.
[0129] The terms "including" and "having" and any variations thereof in the specification and claims of this application and the above-mentioned drawings are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally includes steps or units that are not listed, or optionally includes other steps or units inherent to these processes, methods, products or devices. The terms "first", "second" and "third" are used to distinguish different objects, etc., and do not represent a sequence, nor do they limit "first", "second" and "third" to different types.
[0130] In the description of the embodiments of the present application, "exemplary", "for example" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary", "for example" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary", "for example" or "for example" is intended to present related concepts in a specific way.
[0131] In the description of the embodiments of the present application, unless otherwise specified, “ / ” means or, for example, A / B can mean A or B; the “and / or” in the text is merely a description of the association relationship of associated objects, indicating that three relationships may exist, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of the present application, “multiple” refers to two or more than two.
[0132] In some processes described in the embodiments of the present application, multiple operations or steps that appear in a specific order are included, but it should be understood that these operations or steps may not be executed in the order in which they appear in the embodiments of the present application or in parallel, and the sequence number of the operation is only used to distinguish the different operations, and the sequence number itself does not represent any execution order. In addition, these processes may include more or fewer operations, and these operations or steps may be executed in sequence or in parallel, and these operations or steps may be combined.
[0133] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, disk, CD) as described above, and includes a number of instructions for a terminal device to execute the methods described in each embodiment of the present application.
[0134] The above are only preferred embodiments of the present application, and are not intended to limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A time-sensitive target detection method, characterized in that: The time-sensitive target detection method comprises: Through the preset backbone network, features of different scales are extracted from the image to be detected to obtain a multi-scale feature map; Based on a preset feature pyramid network, feature fusion is performed on the multi-scale feature map to obtain a target feature map, wherein the feature pyramid network includes a target aggregation module, and the target aggregation module includes a plurality of dilated convolutions with different dilated rates, and the plurality of dilated convolutions are used to extract context information under different receptive fields; Target detection is performed on the target feature map according to a preset detection head to output a time-sensitive target detection result.
2. The time-sensitive target detection method according to claim 1, characterized in that: The target aggregation module includes a first aggregation branch, a second aggregation branch, a third aggregation branch, a fourth aggregation branch and a fusion module, wherein the fusion module is used to fuse the output results of the four aggregation branches, wherein the first aggregation branch includes a convolutional layer, and the second aggregation branch, the third aggregation branch and the fourth aggregation branch are all composed of sequentially connected convolutional layers, dilated convolutions and convolutional layers, wherein the dilated rates corresponding to the dilated convolutions in the second aggregation branch, the third aggregation branch and the fourth aggregation branch increase sequentially, and the convolutional layer is used to perform convolution processing on the input feature map to reduce the number of channels.
3. The time-sensitive target detection method according to claim 1, characterized in that: The backbone network is a Darknet53 network, and the multi-scale feature map includes a shallow feature map, a middle feature map and a deep feature map.
4. The time-sensitive target detection method according to claim 3, characterized in that: The target aggregation module is used to extract context information of shallow feature maps under different receptive fields and fuse them to output an enhanced shallow feature map, and use the enhanced shallow feature map as one of the target feature maps.
5. The time-sensitive target detection method according to claim 3, characterized in that: The number of the target aggregation modules is three; The first target aggregation module is used to extract the context information of shallow feature maps under different receptive fields and fuse them to output enhanced shallow feature maps; The second target aggregation module is used to extract the context information of the middle-level feature maps under different receptive fields and fuse them to output the enhanced middle-level feature maps; The third target aggregation module is used to extract the context information of the deep feature maps under different receptive fields and fuse them to output the enhanced deep feature maps; Among them, the enhanced shallow feature map, the enhanced middle feature map and the enhanced deep feature map are all target feature maps.
6. The time-sensitive target detection method according to claim 1, characterized in that: The detection head includes a basic convolution module and a convolution layer connected in sequence, and the basic convolution module includes a convolution layer, a batch normalization layer and an activation layer connected in sequence.
7. A time-sensitive target detection device, characterized in that: The time-sensitive target detection device comprises: The backbone network is used to extract features of different scales from the image to be detected to obtain a multi-scale feature map; A feature pyramid network is used to perform feature fusion on multi-scale feature maps to obtain a target feature map. The feature pyramid network includes a target aggregation module, and the target aggregation module includes multiple dilated convolutions with different dilated rates. The multiple dilated convolutions are used to extract context information under different receptive fields. A detection head is used to perform target detection on the target feature map to output a time-sensitive target detection result.
8. The time-sensitive target detection device according to claim 7, characterized in that: The target aggregation module includes a first aggregation branch, a second aggregation branch, a third aggregation branch, a fourth aggregation branch and a fusion module, wherein the fusion module is used to fuse the output results of the four aggregation branches, wherein the first aggregation branch includes a convolutional layer, and the second aggregation branch, the third aggregation branch and the fourth aggregation branch are all composed of sequentially connected convolutional layers, dilated convolutions and convolutional layers, wherein the dilated rates corresponding to the dilated convolutions in the second aggregation branch, the third aggregation branch and the fourth aggregation branch increase sequentially, and the convolutional layer is used to perform convolution processing on the input feature map to reduce the number of channels.
9. A time-sensitive target detection device, characterized in that: The time-sensitive target detection device includes a processor, a memory, and a time-sensitive target detection program stored in the memory and executable by the processor, wherein when the time-sensitive target detection program is executed by the processor, the steps of the time-sensitive target detection method as described in any one of claims 1 to 6 are implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a time-sensitive target detection program, wherein when the time-sensitive target detection program is executed by a processor, the steps of the time-sensitive target detection method according to any one of claims 1 to 6 are implemented.