Object detection method based on multi-branch parallel hybrid hole coding neural network
Through the object detection method based on multi-branch parallel hybrid hollow coding neural network, the accuracy problem of drone target detection in dense and occlusion scenarios is solved, and the high-precision and real-time object detection effect is achieved.
Patent Information
- Application Number
- CN202210319406.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-29
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-03-29
AI Technical Summary
The existing drone target detection technology has problems such as low detection accuracy, easy to miss detection and miss detection when processing drone ground surveillance images, especially in dense and occlusion scenarios.
The object detection method based on a multi-branch parallel hybrid hollow coding neural network is adopted. This method extracts the image feature through a parallel hybrid hollow coding neural network, and uses the decoding prediction network and the attention-free anchor prediction network to obtain two detection results, and obtains the final detection results through fusion.
The drone's detection accuracy of small targets on the ground is improved, and the missed detection and missed detection of small targets in dense and occluded scenarios are reduced, while the detection model is lightweight and real-time.
Smart Images

Figure CN114821462B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of unmanned aerial vehicle target detection, and in particular relates to a target detection method based on a multi-branch parallel hybrid hole coding neural network. Background Art
[0002] In order to strengthen security management, installing a network monitoring system, strengthening monitoring of key areas and linking with the alarm system has become one of the commonly used security measures. Fixed surveillance cameras are now commonly used for monitoring, which inevitably leads to blind spots and blind areas, and is easily interfered by the outside world. Fixed surveillance cameras have problems such as narrow viewing angles, unclear images, and difficulty in obtaining first-hand on-site image data. They cannot meet the increasingly complex security monitoring needs, and consume a lot of manpower to view monitoring in real time. It is also difficult to easily obtain useful information when calling surveillance videos.
[0003] Drone monitoring can quickly, accurately and efficiently identify detection targets in scenes with dense buildings, wide distribution of personnel and vehicles, and uneven distribution, conduct real-time monitoring of large-scale activities, and quickly locate pedestrians and various types of vehicles to check for dangerous factors. Therefore, by making full use of the advantages of drone target detection systems such as low labor cost, strong mobility, clear imaging, and wide coverage, effective monitoring can be carried out to achieve intelligent control and management of security monitoring.
[0004] Pedestrian and vehicle detection is an essential part of drone monitoring tasks. However, due to the characteristics of drone images and complex road scenes, there are three major difficulties in target detection: (1) The targets in drone ground monitoring images are small in scale, lack appearance information, and have few available feature points, resulting in low detection accuracy. (2) In scenes with dense urban buildings, targets are easily blocked, which can easily lead to missed detections. (3) When traffic congestion occurs on the road, a large number of targets gather, resulting in low detection accuracy and easy false detections. Summary of the invention
[0005] In order to solve the above problems existing in the prior art, the present invention provides a target detection method based on a multi-branch parallel hybrid hole coding neural network. The technical problem to be solved by the present invention is achieved by the following technical solutions:
[0006] The present invention provides a target detection method based on a multi-branch parallel hybrid hole coding neural network, the method comprising:
[0007] Acquire the image to be tested;
[0008] Inputting the image to be tested into a trained multi-branch parallel hybrid hole coding neural network to obtain a final detection result of the image to be tested;
[0009] The multi-branch parallel hybrid hole coding neural network is trained based on multiple training samples, and the training samples include images taken by a drone and their corresponding category labels;
[0010] The multi-branch parallel hybrid hole coding neural network includes: a plurality of sequentially connected downsampling modules, the output end of each of the downsampling modules is connected to a parallel hybrid hole coding neural network to form a parallel branch structure, the output end of the first parallel hybrid hole coding neural network of the branch structure is connected to the decoding prediction network; the output ends of the other parallel hybrid hole coding neural networks of the branch structure are all connected to the attention anchor-free prediction network;
[0011] The downsampling module is used to downsample the input image to obtain a low-level feature map; the parallel hybrid hole coding neural network is used to extract features from the low-level feature map to obtain an enhanced feature map; the decoding prediction network is used to classify and detect the input enhanced feature map to obtain a first detection result; the attention anchor-free prediction network is used to classify and detect the input enhanced feature map to obtain a second detection result; and the final detection result of the image to be tested is obtained based on the first detection result and the second detection result.
[0012] In one embodiment of the present invention, the parallel hybrid hole coding neural network comprises: a parallel first branch link and a second branch link, wherein:
[0013] The first branch link captures image context information of the input low-level feature map at multiple scales to obtain an initial feature map; the second branch link is used to perform global-aware attention weight allocation on the input low-level feature map to obtain an attention-weighted feature map; the initial feature map and the attention-weighted feature map are merged by concat to obtain the enhanced feature map.
[0014] In one embodiment of the present invention, the first branch link includes a plurality of first convolution units connected in sequence, the first convolution unit includes a first dilated convolution layer, a first BN layer and a first Mish activation function layer connected in sequence, and the first dilated convolution layer in each first convolution unit has a different dilation rate and a different convolution kernel size;
[0015] The second branch link includes a first LayerNorm layer, a first multi-head attention module, a first Dropout layer, a second LayerNorm layer, a first feedforward neural network and a second Dropout layer connected in sequence, the output of the first LayerNorm layer is multiplied by the output of the first Dropout layer as the input of the second LayerNorm layer, and the input of the second LayerNorm layer is multiplied by the output of the feedforward neural network as the input of the second Dropout layer.
[0016] In one embodiment of the present invention, the decoding prediction network includes an attention module, an encoding-decoding module and a classification prediction module connected in sequence, wherein:
[0017] The attention module is used to perform global perception attention weight allocation on the input enhanced feature map, and infer in sequence along two independent dimensions of channel and space to obtain an attention feature map, and multiply the attention feature map with the enhanced feature map input by the decoding prediction network to achieve adaptive feature refinement;
[0018] The encoding-decoding module is used to encode the feature-refined attention feature map into a coding information matrix, and fuse and decode it with the enhanced feature map output by the parallel hybrid hole coding network to obtain a fused decoding feature map;
[0019] The classification prediction module is used to perform a convolution operation on the fused decoding feature map to obtain the first detection result, which includes the category probability and target frame coordinate information of the image to be tested.
[0020] In one embodiment of the present invention, the attention module includes a second multi-head attention module, a first channel and spatial attention module, a third LayerNorm layer and a third Dropout layer connected in sequence, wherein the output of the third Dropout layer is multiplied by the input of the attention module as the output of the attention module;
[0021] The encoding-decoding module includes an Encoder-Decoder attention module, a fourth LayerNorm layer, a second feedforward neural network and a fourth Dropout layer connected in sequence, wherein the output of the fourth LayerNorm layer is multiplied by the output of the second feedforward neural network and used as the input of the fourth Dropout layer;
[0022] The classification prediction module includes a first FFN unit and a second FFN unit, wherein the first FFN unit and the second FFN unit are both connected to the output end of the fourth Dropout layer, and the first FFN unit and the second FFN unit correspondingly output the category probability and target frame coordinate information of the image to be tested.
[0023] In one embodiment of the present invention, a target detection method based on a multi-branch parallel hybrid hole coding neural network is characterized in that the attention-free anchor prediction network includes a plurality of parallel-connected attention hybrid hole convolution modules, and the attention hybrid hole convolution modules are correspondingly connected to the parallel hybrid hole coding neural network of the branch structure;
[0024] The attention mixed hole convolution module performs classification detection on the input enhanced feature map to obtain the category probability, target frame information and target frame coordinate information of the image to be tested respectively; the two-dimensional feature vector obtained by concat merging and reshaping the category probability, the target frame information and the target frame coordinate information is used as the second detection result.
[0025] In one embodiment of the present invention, the attention mixed hole convolution module includes a convolution layer, a second channel and spatial attention module, and a plurality of second convolution units connected in sequence, wherein the second convolution unit includes a second hole convolution layer, a second BN layer, and a second Mish activation function layer connected in sequence, and the second hole convolution layer in each second convolution unit has different hole rates and different convolution kernel sizes.
[0026] In one embodiment of the present invention, obtaining a final detection result of the image to be detected according to the first detection result and the second detection result includes:
[0027] The first detection result and the second detection result are fused through the NMS operation, the category probabilities in the first detection result and the second detection result are sorted, and the category corresponding to the maximum category probability and the target frame coordinate information are used as the final detection result of the image to be tested.
[0028] In one embodiment of the present invention, during the training of the multi-branch parallel hybrid hole coding neural network, the loss function used is:
[0029] L=λ giou ·L giou +λ fl ·L fl +λ bce ·L bce ;
[0030] Where, L giouThe loss function L represents the position deviation between the predicted target box and the true labeled box. bce The loss function that represents the size deviation between the predicted target box and the true labeled box, L fl Represents the loss function of the predicted classification result and the true category label, λ giou , fl , bce Represents the corresponding coefficients of each loss function component.
[0031] Compared with the prior art, the present invention has the following beneficial effects:
[0032] 1. The target detection method based on a multi-branch parallel hybrid void coding neural network of the present invention performs target detection on the image to be tested taken by a drone through a trained multi-branch parallel hybrid void coding neural network. The multi-branch parallel hybrid void coding neural network uses a parallel hybrid void coding neural network to extract features of the input image, and simultaneously uses a decoding prediction network and an attention-free anchor prediction network to obtain two detection results. The final detection result is obtained by fusing the two detection results. The target detection method can improve the detection accuracy of drones for small ground targets, especially for missed detection and false detection of small targets in dense and occluded scenes.
[0033] 2. The target detection method based on the multi-branch parallel hybrid void coding neural network of the present invention improves the detection accuracy and the detection model is lightweight and real-time. The method of the present invention can be used to realize an unmanned aerial vehicle intelligent detection and monitoring system platform, optimize equipment resources and human resource allocation, and reduce the operating cost of monitoring.
[0034] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present invention more obvious and easy to understand, the following specifically cites a preferred embodiment and describes it in detail with the accompanying drawings as follows. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 is a schematic diagram of a target detection method based on a multi-branch parallel hybrid hole coding neural network provided by an embodiment of the present invention;
[0036] Figure 2 is a schematic diagram of the structure of a multi-branch parallel hybrid hole coding neural network provided by an embodiment of the present invention;
[0037] Figure 3 is a schematic diagram of the structure of a parallel hybrid hole coding neural network provided by an embodiment of the present invention;
[0038] Figure 4is a schematic diagram of the structure of a decoding prediction network provided by an embodiment of the present invention;
[0039] Figure 5 It is a structural diagram of an attention-free anchor prediction network provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0040] In order to further explain the technical means and effects adopted by the present invention to achieve the predetermined purpose of the invention, the following is a detailed description of a target detection method based on a multi-branch parallel hybrid void coding neural network proposed in accordance with the present invention in combination with the accompanying drawings and specific implementation methods.
[0041] The above and other technical contents, features and effects of the present invention are clearly presented in the following detailed description of the specific implementation modes in conjunction with the accompanying drawings. Through the description of the specific implementation modes, the technical means and effects adopted by the present invention to achieve the predetermined purpose can be more deeply and specifically understood. However, the attached drawings are only for reference and explanation purposes and are not used to limit the technical solutions of the present invention.
[0042] Embodiment 1
[0043] See also Figure 1 , Figure 1 1 is a schematic diagram of a target detection method based on a multi-branch parallel hybrid hole coding neural network provided by an embodiment of the present invention. As shown in the figure, the target detection method includes the following steps:
[0044] Step 1: Get the image to be tested;
[0045] In this embodiment, the image to be tested is a small target image captured by a drone equipped with a camera, and the image to be tested contains small targets to be identified and classified, including but not limited to one or more of eight categories: pedestrians, bicycles, cars, vans, trucks, tricycles, awning tricycles and irrelevant areas.
[0046] Step 2: Input the image to be tested into the trained multi-branch parallel hybrid hole coding neural network to obtain the final detection result of the image to be tested.
[0047] Among them, the multi-branch parallel hybrid void coding neural network is trained based on multiple training samples, and the training samples include images taken by drones and their corresponding category labels.
[0048] See also Figure 2 , Figure 2It is a structural schematic diagram of a multi-branch parallel hybrid hole coding neural network provided by an embodiment of the present invention. As shown in the figure, the multi-branch parallel hybrid hole coding neural network of this embodiment includes: a number of downsampling modules connected in sequence, and the output end of each downsampling module is connected to a parallel hybrid hole coding neural network to form a parallel branch structure. The output end of the parallel hybrid hole coding neural network of the first branch structure is connected to the decoding prediction network; the output ends of the parallel hybrid hole coding neural networks of other branch structures are all connected to the attention anchor-free prediction network.
[0049] Specifically, the downsampling module is used to downsample the input image to obtain a low-level feature map; the parallel hybrid hole coding neural network is used to extract features from the low-level feature map to obtain an enhanced feature map; the decoding prediction network is used to classify and detect the input enhanced feature map to obtain a first detection result; the attention anchor-free prediction network is used to classify and detect the input enhanced feature map to obtain a second detection result; and the final detection result of the image to be tested is obtained based on the first detection result and the second detection result.
[0050] In this embodiment, four downsampling modules are provided, and the output end of each downsampling module is connected to a parallel hybrid hole coding neural network to form four parallel branch structures. Specifically, in the first branch structure, the input image to be tested is downsampled once to obtain the first low-level feature map F 1 ; In the second branch structure, the input image to be tested is downsampled twice to obtain the second low-level feature map F 2 In the third branch structure, the input image to be tested is downsampled three times to obtain the third low-level feature map F 3 In the fourth branch structure, the input image to be tested is downsampled four times to obtain the fourth low-level feature map F 4 .
[0051] The specific downsampling process is: use 3*3 maximum pooling operation for downsampling, take the maximum value of the feature points in the neighborhood of the feature map, assume that N is the output feature map size, W is the input feature map size, F is the convolution kernel size, P is the size of the filling value, S is the step size, and the downsampling process is described as: N=(W-F+2P) / S+1.
[0052] Furthermore, the first low-level feature map F corresponding to the four-branch structure is obtained by a parallel hybrid hole coding neural network. 1 , the second lowest level feature map F 2 , the third low-level feature map F 3 And the fourth low-level feature map F 4 Perform feature extraction to obtain the corresponding first enhanced feature map F 1 , the second enhanced feature map F 2, the third enhanced feature map F 3 , the fourth enhanced feature map F 4 .
[0053] Furthermore, the first enhanced feature map F is decoded and predicted by the network 1 Perform detection to obtain the first detection result; at the same time, the second enhanced feature map F is predicted by the attention-free anchor prediction network. 2 , the third enhanced feature map F 3 and the fourth enhanced feature map F 4 Perform detection to obtain a second detection result. Merge the first detection result and the second detection result to obtain a final detection result of the image to be detected.
[0054] Please refer to Figure 3 , Figure 3 This is a schematic diagram of the structure of a parallel hybrid hole coding neural network provided by an embodiment of the present invention. As shown in the figure, the parallel hybrid hole coding neural network of this embodiment includes: a parallel first branch link and a second branch link, wherein the first branch link captures the image context information of the input low-level feature map at multiple ratios to obtain an initial feature map; the second branch link is used to perform global perception attention weight allocation on the input low-level feature map to obtain an attention weighted feature map. The specific weight allocation process is:
[0055] Assume the input is a 1 , a 2 , a 1 and a 2 With the weight parameter matrix WqW k W v Multiply to generate the weighted q 1 ,k 1 ,v 1 and q 2 ,k 2 ,v 2 , as described in the following formula, and then use the generated weighted q 1 ,k 1 ,v 1 and q 2 ,k 2 ,v 2 Perform weighting to obtain the attention weighted feature map:
[0056] q i =a i W q (1);
[0057] k i =a i W k (2);
[0058] vi =a i W v (3).
[0059] The initial feature map and the attention weighted feature map are combined by concat to obtain an enhanced feature map. Assume that the channels of the two inputs are X i and Y i , then the single output channel through concat is:
[0060]
[0061] Among them, * represents convolution and K represents the convolution kernel.
[0062] Specifically, the first branch link includes a plurality of first convolution units connected in sequence, the first convolution unit includes a first hole convolution layer, a first BN layer and a first Mish activation function layer connected in sequence, and the first hole convolution layer in each first convolution unit has a different hole rate and a different convolution kernel size. The second branch link includes a first LayerNorm layer, a first multi-head attention module, a first Dropout layer, a second LayerNorm layer, a first feedforward neural network and a second Dropout layer connected in sequence, the output of the first LayerNorm layer is multiplied by the output of the first Dropout layer as the input of the second LayerNorm layer, and the input of the second LayerNorm layer is multiplied by the output of the feedforward neural network as the input of the second Dropout layer.
[0063] In this embodiment, the first feedforward neural network is MLP, and the first branch link includes three first convolutional units connected in sequence. The void rates of the first hole convolutional layers in the three first convolutional units are 1, 2, and 3, respectively, and the convolution kernel sizes are 3*3, 5*5, and 9*9, respectively.
[0064] In this embodiment, the first branch link can be used to capture the contextual information of the image at multiple scales for the input low-level feature map, and the continuously increasing hole rate increases the receptive field while avoiding the problems of discontinuous receptive field and local information loss. The second branch link can be used to allocate global perceived attention weights to the input low-level feature map to enhance the expressiveness of the input features.
[0065] Please refer to Figure 4 , Figure 4It is a structural schematic diagram of a decoding prediction network provided by an embodiment of the present invention. As shown in the figure, the decoding prediction network of this embodiment includes an attention module, an encoding-decoding module and a classification prediction module connected in sequence, wherein the attention module is used to perform global perception attention weight allocation on the input enhanced feature map, and perform inference in sequence along two independent dimensions of channel and space to obtain an attention feature map, and multiply the attention feature map with the enhanced feature map input by the decoding prediction network to achieve adaptive feature refinement; the encoding-decoding module is used to encode the attention feature map after feature refinement into a coding information matrix, and fuse and decode it with the enhanced feature map output by the parallel hybrid void coding network to obtain a fused decoding feature map; the classification prediction module is used to perform a convolution operation on the fused decoding feature map to obtain a first detection result, and the first detection result includes the category probability and target frame coordinate information of the image to be tested.
[0066] Specifically, the attention module includes a second multi-head attention module, a first channel and spatial attention module, a third LayerNorm layer and a third Dropout layer connected in sequence, wherein the output of the third Dropout layer is multiplied by the input of the attention module as the output of the attention module.
[0067] In this embodiment, first, the first enhanced feature map F 1 After a multi-head attention module, the calculation process is as follows:
[0068] MultiHead(Q,K,V)=Concat(head 1 ,...,head h )W O (5);
[0069] Specifically, the weight vectors obtained by self-attention are concatenated, the first subscripts are concatenated together, and then a learnable parameter W is used. O The spliced data are further fused in the form of matrix multiplication.
[0070] Secondly, through a channel and spatial attention module (CBAM), inference is performed along two independent dimensions, channel and space, to obtain the attention feature map, and then a LayerNorm layer and a Dropout layer are connected in sequence to facilitate subsequent training.
[0071] Furthermore, the encoding-decoding module includes an Encoder-Decoder attention module, a fourth LayerNorm layer, a second feedforward neural network and a fourth Dropout layer connected in sequence, wherein the output of the fourth LayerNorm layer is multiplied by the output of the second feedforward neural network and used as the input of the fourth Dropout layer.
[0072] In this embodiment, the Encoder-Decoder attention module is used to encode the feature-refined attention feature map into a coding information matrix, and fuse and decode it with the enhanced feature map output by the parallel hybrid hole coding network for subsequent convolution operations; the fourth LayerNorm layer is used to normalize the fused decoded feature map; the second feedforward neural network is an MLP, which includes two fully connected layers to solve nonlinear problems.
[0073] Furthermore, the classification prediction module includes a first FFN unit and a second FFN unit, wherein the first FFN unit and the second FFN unit are both connected to the output end of the fourth Dropout layer, and the first FFN unit and the second FFN unit correspondingly output the category probability and target frame coordinate information of the image to be tested.
[0074] Please refer to Figure 5 , Figure 5 It is a structural schematic diagram of an attention-free anchor prediction network provided by an embodiment of the present invention. As shown in the figure, the attention-free anchor prediction network of this embodiment includes a plurality of parallel-connected attention mixed void convolution modules, and the attention mixed void convolution modules are correspondingly connected to the parallel mixed void coding neural network with a branch structure; the attention mixed void convolution module performs classification detection on the input enhanced feature map, and obtains the category probability, target frame information and target frame coordinate information of the image to be tested respectively; the category probability, target frame information and target frame coordinate information are merged through concat and reshape operation to obtain a two-dimensional feature vector as the second detection result.
[0075] Specifically, the attention mixed dilated convolution module includes a sequentially connected convolution layer, a second channel and spatial attention module (CBAM), and several second convolution units. The second convolution unit includes a sequentially connected second dilated convolution layer, a second BN layer, and a second Mish activation function layer. The second dilated convolution layer in each second convolution unit has different dilation rates and different convolution kernel sizes.
[0076] In this embodiment, the attention-free anchor prediction network includes three parallel-connected attention-mixed hole convolution modules, which are respectively connected to the parallel mixed hole coding neural networks in the three branch structures, and are used to input the second enhanced feature map F 2 , the third enhanced feature map F 3 , the fourth enhanced feature map F 4 The target category, foreground and background scores, and target box coordinates are predicted accordingly.
[0077] In this embodiment, three second convolution units are arranged in the attention mixed dilated convolution module, and the dilation rates of the second dilated convolution layers in the three second convolution units are 1, 2, and 3, respectively, and the convolution kernel sizes are 3*3, 5*5, and 9*9, respectively.
[0078] The attention-free anchor prediction network predicts the following: The second enhanced feature map F 2 , the third enhanced feature map F 3 and the fourth enhanced feature map F 4 The outputs of the three branches are obtained through the attention-free anchor prediction network. In this embodiment, the output channels of the three branches are 12, 1, and 4, respectively. The output with a channel number of 12 performs category judgment on 12 categories of targets to obtain the corresponding category probabilities; the output with a channel number of 1 mainly judges whether the target box is foreground or background; the output with a channel number of 4 mainly predicts the target box coordinate information (x, y, w, h). The last three outputs are fused together through Concat to obtain a feature vector with a channel number of 17, which contains the category probability, target box information, and target box coordinate information of the image to be tested, and then the feature map is transformed into a two-dimensional feature vector through the Reshape operation.
[0079] Furthermore, the final detection result of the image to be tested is obtained according to the first detection result and the second detection result, including: fusing the first detection result and the second detection result through the NMS operation, sorting the category probabilities in the first detection result and the second detection result, and taking the category corresponding to the maximum category probability and the target frame coordinate information as the final detection result of the image to be tested.
[0080] In order to make the scheme clearer, the training process of the multi-branch parallel hybrid void coding neural network is illustrated as follows: first, an image data set is obtained, and the original image is collected by shooting with a drone equipped with a camera. The targets in the original image include eight categories: pedestrians, bicycles, cars, vans, trucks, tricycles, awning tricycles and irrelevant areas; each collected image is annotated, and an XML format annotation file containing semantic information is generated for the target to be detected in the image, and the data set is divided into a training set, a verification set and a test set according to a ratio of 1:1:2, which are used for pre-training, verification and testing of the multi-branch parallel hybrid void coding neural network respectively.
[0081] Then, the training images are input into the multi-branch parallel hybrid hole coding neural network in batches to obtain the prediction results. The prediction results are specifically the target prediction box and the category probability of the target. The prediction box and the true label box are associated to obtain the positive sample prediction box, and the loss function is calculated by the final prediction box and the true label box. During the training process, the loss function used is:
[0082] L=λ giou ·L giou +λ fl ·L fl +λ bce ·L bce (6);
[0083] Where, L giou The loss function L represents the position deviation between the predicted target box and the true labeled box. bce The loss function that represents the size deviation between the predicted target box and the true labeled box, L fl Represents the loss function of the predicted classification result and the true category label, λ giou , fl , bce Represents the corresponding coefficients of each loss function component.
[0084] Finally, according to the calculated loss value, the optimizer is used to optimize the network parameters; when the loss value calculated after a batch of training images is input into the multi-branch parallel hybrid void coding neural network is less than the preset threshold, the network is considered to have converged and the training is completed.
[0085] The target detection method based on a multi-branch parallel hybrid hole coding neural network of this embodiment performs target detection on the image to be tested taken by a drone through a trained multi-branch parallel hybrid hole coding neural network. The multi-branch parallel hybrid hole coding neural network uses a parallel hybrid hole coding neural network to extract features of the input image, and simultaneously uses a decoding prediction network and an attention-free anchor prediction network to obtain two detection results. The final detection result is obtained by fusing the two detection results. The target detection method can improve the detection accuracy of drones for small ground targets, especially for missed detection and false detection of small targets in dense and occluded scenes.
[0086] The target detection method based on a multi-branch parallel hybrid void coding neural network in this embodiment improves the detection accuracy and at the same time the detection model is lightweight and real-time. The method of the present invention can be used to implement an intelligent detection and monitoring system platform for unmanned aerial vehicles, optimize equipment resources and human resource allocation, and reduce the operating cost of monitoring.
[0087] Based on the same inventive concept, an embodiment of the present invention also provides an electronic device, including a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; the memory is used to store computer programs; when the processor is used to execute the program stored in the memory, it implements any of the method steps of the above-mentioned target detection method based on a multi-branch parallel hybrid hole coding neural network, or implements the functions implemented by any of the above-mentioned multi-branch parallel hybrid hole coding neural networks.
[0088] The embodiment of the present invention also provides a computer-readable storage medium. A computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, the method steps of any of the above-mentioned target detection methods based on a multi-branch parallel hybrid hole coding neural network are implemented, or the functions implemented by any of the above-mentioned multi-branch parallel hybrid hole coding neural networks are implemented.
[0089] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants are intended to cover non-exclusive inclusion, so that an article or device including a series of elements includes not only those elements, but also other elements that are not explicitly listed. In the absence of further restrictions, the elements defined by the statement "including one..." do not exclude the existence of other identical elements in the article or device including the elements. Similar words such as "connect" or "connected" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect.
[0090] The above contents are further detailed descriptions of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as falling within the protection scope of the present invention.
Claims
1. A target detection method based on multi-branch parallel hybrid hole coding neural network, It is characterized in that include: Acquire the image to be tested; The image to be tested is input into a trained multi-branch parallel hybrid hole coding neural network to obtain a final detection result of the image to be tested; wherein the parallel hybrid hole coding neural network includes: a parallel first branch link and a second branch link, wherein the first branch link captures image context information of the input low-level feature map at multiple ratios to obtain an initial feature map; the second branch link is used to perform global perception attention weight allocation on the input low-level feature map to obtain an attention weighted feature map; the initial feature map and the attention weighted feature map are merged by concat to obtain an enhanced feature map; The multi-branch parallel hybrid hole coding neural network is trained based on multiple training samples, and the training samples include images taken by a drone and their corresponding category labels; The multi-branch parallel hybrid hole coding neural network includes: a plurality of sequentially connected downsampling modules, the output end of each of the downsampling modules is connected to a parallel hybrid hole coding neural network to form a parallel branch structure, the output end of the first parallel hybrid hole coding neural network of the branch structure is connected to the decoding prediction network; the output ends of the other parallel hybrid hole coding neural networks of the branch structure are all connected to the attention anchor-free prediction network; wherein the decoding prediction network includes an attention module, an encoding-decoding module and a classification prediction module connected in sequence, wherein the attention module is used to perform global perception of attention weight allocation on the input enhanced feature map, and sequentially infer along two independent dimensions of channel and space to obtain an attention feature map, and multiply the attention feature map with the enhanced feature map input by the decoding prediction network to achieve adaptive feature refinement; The encoding-decoding module is used to encode the feature-refined attention feature map into a coding information matrix, and fuse and decode it with the enhanced feature map output by the parallel hybrid hole coding network to obtain a fused decoding feature map; The classification prediction module is used to perform a convolution operation on the fused decoding feature map to obtain a first detection result, where the first detection result includes the category probability and target frame coordinate information of the image to be tested; The attention-free anchor prediction network includes a plurality of parallel-connected attention-mixed hole convolution modules, and the attention-mixed hole convolution modules are correspondingly connected to the parallel-mixed hole coding neural network of the branch structure; The attention mixed hole convolution module performs classification detection on the input enhanced feature map to obtain the category probability, target frame information and target frame coordinate information of the image to be tested respectively; the two-dimensional feature vector obtained by concat merging and reshaping the category probability, the target frame information and the target frame coordinate information is used as the second detection result The downsampling module is used to downsample the input image to obtain a low-level feature map; the parallel hybrid hole coding neural network is used to extract features from the low-level feature map to obtain an enhanced feature map; the decoding prediction network is used to classify and detect the input enhanced feature map to obtain a first detection result; the attention anchor-free prediction network is used to classify and detect the input enhanced feature map to obtain a second detection result; and the final detection result of the image to be tested is obtained based on the first detection result and the second detection result.
2. The target detection method based on a multi-branch parallel hybrid hole coding neural network according to claim 1, It is characterized in that The first branch link includes a plurality of first convolution units connected in sequence, the first convolution unit includes a first dilated convolution layer, a first BN layer and a first Mish activation function layer connected in sequence, and the first dilated convolution layer in each first convolution unit has a different dilation rate and a different convolution kernel size; The second branch link includes a first LayerNorm layer, a first multi-head attention module, a first Dropout layer, a second LayerNorm layer, a first feedforward neural network and a second Dropout layer connected in sequence, the output of the first LayerNorm layer is multiplied by the output of the first Dropout layer as the input of the second LayerNorm layer, and the input of the second LayerNorm layer is multiplied by the output of the feedforward neural network as the input of the second Dropout layer.
3. The target detection method based on multi-branch parallel hybrid hole coding neural network according to claim 1, It is characterized in that The attention module includes a second multi-head attention module, a first channel and spatial attention module, a third LayerNorm layer and a third Dropout layer connected in sequence, wherein the output of the third Dropout layer is multiplied by the input of the attention module as the output of the attention module; The encoding-decoding module includes an Encoder-Decoder attention module, a fourth LayerNorm layer, a second feedforward neural network and a fourth Dropout layer connected in sequence, wherein the output of the fourth LayerNorm layer is multiplied by the output of the second feedforward neural network and used as the input of the fourth Dropout layer; The classification prediction module includes a first FFN unit and a second FFN unit, wherein the first FFN unit and the second FFN unit are both connected to the output end of the fourth Dropout layer, and the first FFN unit and the second FFN unit correspondingly output the category probability and target frame coordinate information of the image to be tested.
4. The target detection method based on a multi-branch parallel hybrid hole coding neural network according to claim 1, It is characterized in that The attention mixed hole convolution module includes a convolution layer, a second channel and a spatial attention module, and a plurality of second convolution units connected in sequence. The second convolution unit includes a second hole convolution layer, a second BN layer, and a second Mish activation function layer connected in sequence. The second hole convolution layer in each second convolution unit has different hole rates and different convolution kernel sizes.
5. The target detection method based on multi-branch parallel hybrid hole coding neural network according to claim 1, It is characterized in that Obtaining a final detection result of the image to be detected according to the first detection result and the second detection result, including: The first detection result and the second detection result are fused through the NMS operation, the category probabilities in the first detection result and the second detection result are sorted, and the category corresponding to the maximum category probability and the target frame coordinate information are used as the final detection result of the image to be tested.
6. The target detection method based on multi-branch parallel hybrid hole coding neural network according to claim 1, It is characterized in that In the process of training the multi-branch parallel hybrid hole coding neural network, the loss function used is: L=λ giou ·L giou +λ fl ·L fl +λ bce ·L bce ; Where, L giou The loss function L represents the position deviation between the predicted target box and the true labeled box. bce The loss function that represents the size deviation between the predicted target box and the true labeled box, L fl Represents the loss function of the predicted classification result and the true category label, λ giou , fl , bce Represents the corresponding coefficient of each loss function component.