A multi-branch target detection method, system and device with multi-feature fusion
By adopting a multi-branch object detection method with multi-feature fusion in the object detection network, feature processing is performed using hollow convolution and CA attention mechanism, and multiple feature fusions are performed through the neck network of the fused feature pyramid, the problem of information retention and detection accuracy when detecting different size targets in the prior art is solved, and high-accuracy multi-size object detection is achieved.
Patent Information
- Application Number
- CN202310948382.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-31
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2043-07-31
AI Technical Summary
Existing target detection networks find it difficult to meet the needs of information retention and detection accuracy when detecting targets of different sizes, resulting in reduced detection accuracy for small targets or excessive detection information for large targets.
A multi-branch object detection method with multi-feature fusion is proposed. Multi-scale feature extraction is performed through the backbone feature extraction network, hollow convolution module is used to increase the receptive field, and feature weighting is performed through the CA attention mechanism. Then, the neck network of the fusion feature pyramid is used for multiple feature fusions, and the fusion feature map is finally input to the head network for multi-branch object detection.
While avoiding information loss, this method enhances the detailed information content, improves the accuracy of detection of targets of each size, and can detect small, medium and large image targets simultaneously.
Smart Images

Figure CN116863289B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of target detection technology, and in particular to a multi-branch target detection method, system and device with multi-feature fusion. Background Art
[0002] In the object detection network, it is roughly divided into three or four parts, namely the backbone feature extraction backbone network responsible for feature extraction, the head network responsible for feature detection, and the neck network between the backbone feature extraction network and the head network. The main purpose of the neck network is to fuse the features extracted by the feature extraction network at multiple scales and make better use of the features extracted by the backbone feature extraction network. However, if only the feature map output by the backbone feature extraction network is used, the feature Figure 1 Generally, multiple levels of convolution have been performed, and the final output feature map has only a very low resolution. Although it has a large receptive field, it has inevitably lost a lot of information because it has undergone multiple layers of maximum pooling downsampling and convolution. This is very unfriendly to target detection, and there is no way to simultaneously meet the target detection of targets of different sizes or with large size differences. If a model for small targets is used to detect large target data, it will cause problems such as too much information and too small receptive field; if a detection for large targets is used to detect small targets, too many picture details will be lost, making it difficult to detect complete small target objects, reducing the accuracy of small target detection. Summary of the invention
[0003] The present invention aims to solve at least one of the technical problems existing in the prior art. To this end, the present invention proposes a multi-branch target detection method, system and device with multi-feature fusion, which can enhance the content of each detail information accordingly while avoiding information loss, thereby improving the accuracy of target detection of various sizes.
[0004] In a first aspect, an embodiment of the present invention provides a multi-branch target detection method with multi-feature fusion, and the multi-branch target detection method with multi-feature fusion includes:
[0005] Obtain target detection image dataset;
[0006] Inputting the target detection image data set into a backbone feature extraction network to perform multi-scale feature extraction to obtain first feature maps of multiple scales;
[0007] Using a dilated convolution module to increase the receptive field of the first feature maps of the multiple scales to obtain second feature maps of the multiple scales;
[0008] The second feature maps of the multiple scales are weighted by using a CA attention mechanism to obtain weighted feature maps of the multiple scales;
[0009] The neck network of the fused feature pyramid is used to perform multiple feature fusions on the weighted feature maps of the multiple scales to obtain a multi-scale fused feature map;
[0010] The multi-scale fusion feature map is input into the head network to perform multi-branch target detection, and target detection results of images of various sizes are obtained.
[0011] Compared with the prior art, the first aspect of the present invention has the following beneficial effects:
[0012] The method inputs the target detection image data set into the backbone feature extraction network for multi-scale feature extraction to obtain first feature maps of multiple scales, adopts a hole convolution module to increase the receptive field of the first feature maps of multiple scales, and obtains second feature maps of multiple scales. The hole convolution module increases the receptive field of the first feature maps of multiple scales and enhances the detail information, thereby improving the semantic representation ability; the second feature maps of multiple scales are feature-weighted by using the CA attention mechanism to obtain weighted feature maps of multiple scales, and the feature maps output by the backbone feature extraction network are feature-weighted by using the CA attention mechanism, which can improve the feature extraction ability and obtain better features; the neck network of the fused feature pyramid is used to perform multiple feature fusions on the weighted feature maps of multiple scales to obtain multi-scale fused feature maps, and the multi-scale fused feature maps are input into the head network for multi-branch target detection to obtain target detection results of images of multiple sizes, and the neck network of the fused feature pyramid is used to fuse high-level features with high semantic information but low resolution with low-level features with high resolution but low semantic information, thereby avoiding the loss of a lot of information and meeting the needs of target detection. Therefore, the multi-scale fusion feature map is input into the head network for multi-branch target detection to improve the accuracy of target detection for images of each size. In addition, multi-branch target detection is performed through the head network to simultaneously obtain target detection results for images of different sizes, such as tiny images, medium images, and large images.
[0013] According to some embodiments of the present invention, the backbone feature extraction network includes multiple feature layers and multiple convolutional layers, and each of the feature layers includes downsampling and stacked convolutions integrated with dilated convolutions.
[0014] According to some embodiments of the present invention, the step of inputting the target detection image dataset into a backbone feature extraction network to perform multi-scale feature extraction to obtain first feature maps of multiple scales includes:
[0015] Inputting the target detection image data set into multiple convolutional layers for feature extraction to obtain a convolutional feature map;
[0016] The convolution feature map is input into the first feature layer, the second feature layer, the third feature layer, and the fourth feature layer for feature extraction to obtain first feature maps of various scales; wherein the first feature layer adds a prediction head for tiny object detection.
[0017] According to some embodiments of the present invention, the first feature map output by the fourth feature layer is size-fixed using a spatial pyramid pooling module, and the spatial pyramid pooling module uses multiple SoftPool pooling layers.
[0018] According to some embodiments of the present invention, before the neck network of the fused feature pyramid is used to perform multiple feature fusions on the weighted feature maps of multiple scales, the multi-branch target detection method of multi-feature fusion further includes:
[0019] Using a graph neural network to structurally prune the neck network of the fused feature pyramid to obtain a pruned neck network;
[0020] A learning weight parameter is set in the pruned neck network to perform feature learning on the weighted feature maps of the multiple scales.
[0021] According to some embodiments of the present invention, the neck network using the fused feature pyramid performs multiple feature fusions on the weighted feature maps of multiple scales to obtain a multi-scale fused feature map, including:
[0022] A three-branch structured atrous convolution is used to train the feature map after the neck network of the fused feature pyramid performs feature fusion on the weighted feature maps of the multiple scales each time, so as to obtain a plurality of trained first fused feature maps;
[0023] The plurality of trained first fused feature maps are inferred by using structural reparameterization to obtain a multi-scale fused feature map.
[0024] According to some embodiments of the present invention, the head network adopts a WIOU loss function.
[0025] In a second aspect, an embodiment of the present invention further provides a multi-branch target detection system with multi-feature fusion, wherein the multi-branch target detection system with multi-feature fusion includes:
[0026] A data acquisition unit, used to acquire a target detection image data set;
[0027] A feature extraction unit, used for inputting the target detection image data set into a backbone feature extraction network to perform multi-scale feature extraction to obtain first feature maps of multiple scales;
[0028] A receptive field increasing unit, configured to increase the receptive fields of the first feature maps of the plurality of scales by using a dilated convolution module to obtain second feature maps of the plurality of scales;
[0029] A feature weighting unit, used for performing feature weighting on the second feature maps of multiple scales by using a CA attention mechanism to obtain weighted feature maps of multiple scales;
[0030] A feature fusion unit, used to perform multiple feature fusions on the weighted feature maps of multiple scales using a neck network of a fused feature pyramid to obtain a multi-scale fused feature map;
[0031] The target detection unit is used to input the multi-scale fusion feature map into the head network to perform multi-branch target detection and obtain target detection results for images of various sizes.
[0032] In the third aspect, an embodiment of the present invention also provides a multi-branch target detection device with multi-feature fusion, comprising at least one control processor and a memory for communicating with the at least one control processor; the memory stores instructions executable by the at least one control processor, and the instructions are executed by the at least one control processor so that the at least one control processor can execute a multi-feature fusion multi-branch target detection method as described above.
[0033] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute a multi-branch target detection method with multi-feature fusion as described above.
[0034] It can be understood that the beneficial effects of the second to fourth aspects compared with the related art are the same as the beneficial effects of the first aspect compared with the related art. Please refer to the relevant description in the first aspect, and no further details will be given here. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] The above and / or additional aspects and advantages of the present invention will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which:
[0036] Figure 1 is a flow chart of a multi-branch target detection method of multi-feature fusion according to an embodiment of the present invention;
[0037] Figure 2 is a flow chart of a multi-branch target detection method of multi-feature fusion according to another embodiment of the present invention;
[0038] Figure 3 is a schematic diagram of a backbone feature extraction network according to an embodiment of the present invention;
[0039] Figure 4 is a schematic diagram of MP structure downsampling according to an embodiment of the present invention;
[0040] Figure 5 is a schematic diagram of DP structure downsampling according to an embodiment of the present invention;
[0041] Figure 6 is a schematic diagram of a spatial pyramid pooling module according to an embodiment of the present invention;
[0042] Figure 7 is a schematic diagram of a dilated convolution module according to an embodiment of the present invention;
[0043] Figure 8 is a schematic diagram of a CA attention mechanism according to an embodiment of the present invention;
[0044] Fig. 9 is a schematic diagram of integrating the CA attention mechanism into the feature layer of the backbone feature extraction output according to an embodiment of the present invention;
[0045] Fig.10 Schematic diagram of the difference between ordinary convolution and dilated convolution according to an embodiment of the present invention;
[0046] Fig.11 is a schematic diagram of a parallel atrous convolution structure with three branches according to an embodiment of the present invention;
[0047] Fig.12 is a schematic diagram of a training structure according to an embodiment of the present invention;
[0048] Fig.13 is a schematic diagram of an inference structure of an embodiment of the present invention;
[0049] Fig.14 It is a structural diagram of a multi-branch target detection system of multi-feature fusion according to an embodiment of the present invention. DETAILED DESCRIPTION
[0050] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be understood as limiting the present invention.
[0051] In the description of the present invention, if there is a description of first, second, etc., it is only for the purpose of distinguishing technical features, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features or implicitly indicating the order of the indicated technical features.
[0052] In the description of the present invention, it should be understood that descriptions involving orientation, such as orientation or positional relationship indicated as up, down, etc., are based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as a limitation on the present invention.
[0053] In the description of the present invention, it should be noted that, unless otherwise clearly defined, terms such as setting, installing, connecting, etc. should be understood in a broad sense, and technicians in the relevant technical field can reasonably determine the specific meanings of the above terms in the present invention in combination with the specific content of the technical solution.
[0054] In the object detection network, it is divided into three or four parts, namely the backbone feature extraction backbone network responsible for feature extraction, the head network responsible for feature detection, and the neck network between the backbone feature extraction network and the head network. The main purpose of the neck network is to fuse the features extracted by the feature extraction network at multiple scales and make better use of the features extracted by the backbone feature extraction network. However, if only the feature map output by the backbone feature extraction network is used, the feature Figure 1 Generally, multiple levels of convolution have been performed, and the final output feature map has only a very low resolution. Although it has a large receptive field, it has inevitably lost a lot of information because it has undergone multiple layers of maximum pooling downsampling and convolution. This is very unfriendly to target detection in multi-size datasets, and there is no way to simultaneously meet the target detection of targets of different sizes or with large size differences. If a model for small targets is used to detect large targets, it will cause problems such as too much information and too small receptive field, reducing the accuracy of large target detection; if a model for large targets is used to detect small targets, too many picture details will be lost, making it difficult to detect complete small target objects, reducing the accuracy of small target detection.
[0055] To solve the above problems, the present invention inputs the target detection image data set into the backbone feature extraction network for multi-scale feature extraction, obtains first feature maps of multiple scales, adopts a hole convolution module to increase the receptive field of the first feature maps of multiple scales, obtains second feature maps of multiple scales, and increases the receptive field and enhances the detail information of the first feature maps of multiple scales through the hole convolution module, thereby improving the semantic representation ability; the second feature maps of multiple scales are feature-weighted by using the CA attention mechanism to obtain weighted feature maps of multiple scales, and the feature maps output by the backbone feature extraction network are feature-weighted by using the CA attention mechanism, thereby improving the feature extraction ability and obtaining better features; the neck network that integrates the feature pyramid is used to perform multi-scale weighted feature weighting on the weighted feature maps of multiple scales. The multi-scale fusion feature map is then input into the head network for multi-branch target detection to obtain target detection results for images of various sizes. The high-level features with high semantic information but low resolution are fused with the low-level features with high resolution but low semantic information through the neck network of the fusion feature pyramid, which can avoid losing a lot of information. It not only meets the needs of large target detection and classification, but also meets the needs of target detection. Then, the multi-scale fusion feature map is input into the head network for multi-branch target detection to improve the accuracy of target detection for images of each size. In addition, multi-branch target detection is performed through the head network to simultaneously obtain target detection results for images of different sizes, such as tiny images, medium images, and large images.
[0056] Reference Figure 1 to Figure 2 The embodiment of the present invention provides a multi-branch target detection method by multi-feature fusion. The multi-branch target detection method by multi-feature fusion includes but is not limited to steps S100 to S600, wherein:
[0057] Step S100, obtaining a target detection image dataset;
[0058] Step S200: input the target detection image data set into the backbone feature extraction network to perform multi-scale feature extraction to obtain first feature maps of multiple scales;
[0059] Step S300: using a dilated convolution module to increase the receptive field of first feature maps of various scales to obtain second feature maps of various scales;
[0060] Step S400: performing feature weighting on the second feature maps of multiple scales using a CA attention mechanism to obtain weighted feature maps of multiple scales;
[0061] Step S500, using the neck network of the fused feature pyramid to perform multiple feature fusions on weighted feature maps of multiple scales to obtain a multi-scale fused feature map;
[0062] Step S600: Input the multi-scale fusion feature map into the head network to perform multi-branch target detection to obtain target detection results for images of various sizes.
[0063] In steps S100 to S600 of some embodiments, in order to increase the receptive field and enhance the detail information and improve the semantic representation ability, this embodiment obtains a target detection image data set, inputs the target detection image data set into the backbone feature extraction network for multi-scale feature extraction, obtains first feature maps of multiple scales, and uses a hole convolution module to increase the receptive field of the first feature maps of multiple scales to obtain second feature maps of multiple scales; in order to improve the feature extraction ability and obtain better features, this embodiment uses the CA attention mechanism to perform feature weighting on the second feature maps of multiple scales to obtain weighted feature maps of multiple scales; in order to avoid losing a lot of information and improve the accuracy of target detection, this embodiment uses the neck network of the fused feature pyramid to perform multiple feature fusions on the weighted feature maps of multiple scales to obtain multi-scale fused feature maps, and inputs the multi-scale fused feature maps into the head network for multi-branch target detection to obtain target detection results of images of multiple sizes. This embodiment performs multi-branch target detection through the head network, and can simultaneously obtain target detection results of images of different sizes such as tiny images, medium images, and large images.
[0064] In some embodiments, the backbone feature extraction network includes multiple feature layers and multiple convolutional layers, each feature layer includes downsampling and stacked convolutions incorporating dilated convolutions.
[0065] In this embodiment, the dilated convolution is used to replace the downsampling structure in the backbone feature extraction network, which can preserve more image details.
[0066] In some embodiments, the target detection image dataset is input into a backbone feature extraction network to perform multi-scale feature extraction to obtain first feature maps of multiple scales, including:
[0067] Input the target detection image dataset into multiple convolutional layers for feature extraction to obtain a convolutional feature map;
[0068] The convolution feature map is input into the first feature layer, the second feature layer, the third feature layer, and the fourth feature layer for feature extraction to obtain first feature maps of various scales; wherein the first feature layer adds a prediction head for tiny object detection.
[0069] In this embodiment, a prediction head for tiny object detection is added, which can have high-resolution characteristics, be more sensitive to tiny objects, improve the accuracy and completeness of feature extraction, and maintain the idea of divide and conquer at all levels to perform target detection of images with different resolutions, such as tiny targets, small targets, large targets, etc.
[0070] In some embodiments, the first feature map output by the fourth feature layer is size-fixed using a spatial pyramid pooling module, and the spatial pyramid pooling module uses multiple SoftPool pooling layers.
[0071] In this embodiment, the SoftPool pooling layer is used to perform pooling of the spatial pyramid pooling module in the network, which can minimize the loss of information.
[0072] In some embodiments, before the neck network of the fused feature pyramid is used to perform multiple feature fusions on weighted feature maps of multiple scales, the multi-branch target detection method of multi-feature fusion further includes:
[0073] A graph neural network is used to structurally prune the neck network of the fused feature pyramid to obtain the pruned neck network.
[0074] The learning weight parameters are set in the pruned neck network to perform feature learning on weighted feature maps of multiple scales.
[0075] In this embodiment, a graph neural network is used to perform structural pruning on the neck network of the fusion feature pyramid, which can reduce more computational workload without affecting network learning. Learning weight parameters are set to perform feature learning on weighted feature maps of multiple scales, so that the network can learn the importance of features of different sizes in each layer, and use different weights for detection targets of different sizes, thereby achieving divide and conquer.
[0076] In some embodiments, the neck network of the fused feature pyramid is used to perform multiple feature fusions on weighted feature maps of multiple scales to obtain a multi-scale fused feature map, including:
[0077] A three-branch structured atrous convolution is used to train the feature maps of the neck network of each fused feature pyramid on the weighted feature maps of multiple scales after feature fusion, and obtain multiple trained first fused feature maps;
[0078] Structural reparameterization is used to infer multiple trained first fused feature maps to obtain multi-scale fused feature maps.
[0079] In this embodiment, a three-branch dilated convolution is used to modify the multi-branch stacking module in the neck network, thereby increasing the receptive field without increasing the number of parameters and keeping the multi-branch stacking module in the backbone feature extraction unchanged.
[0080] In some embodiments, the head network adopts the WIOU loss function.
[0081] In this embodiment, the use of the WIOU loss function can effectively solve the category imbalance problem in target detection, thereby improving the accuracy of target detection.
[0082] To facilitate understanding by those skilled in the art, a set of best embodiments is provided below:
[0083] (1) Obtain a set of images, preprocess the images, and then use the backbone feature extraction network to extract multi-scale features from the preprocessed images, and output the first feature maps of four different scales. Figure 3 The backbone feature extraction network includes multiple feature layers and multiple convolutional layers. After feature extraction through multiple convolutional layers, the first feature maps of three layers with different scales are output through multiple feature layers, and the first feature map extracted by the shallowest feature layer of the prediction head for tiny target detection is added. In the last feature layer, the spatial pyramid pooling module (spp module) is used to adjust the feature size.
[0084] In the backbone feature extraction network, DP structure downsampling and stacking modules are used for feature extraction. The DP structure downsampling is formed by replacing the MP (Max-Pooling) structure with dilated / atrous convolutions of different sizes in parallel for downsampling. Specifically:
[0085] There are two main existing downsampling methods: one is to use a convolution layer with a stride of 2: the image becomes smaller due to the convolution process in order to extract features, the downsampling process is a process of information loss, and the pooling layer is not learnable. Using a learnable convolution layer with a stride of 2 to replace the pooling can achieve better results, but it also increases the amount of calculation. The second is to use a pooling layer with a stride of 2: pooling downsampling is to reduce the dimension of features, such as maximum pooling (Max-Pooling) and average pooling (Average-Pooling). Max-Pooling is currently commonly used because it is simple to calculate and can better retain texture features. The MP structure formed by combining the above two downsampling methods is as follows: Figure 4 However, Max-Pooling will still cause inevitable loss to the image, so different sizes of dilated / atrous convolution are used in parallel to replace Max-Pooling for downsampling, forming a DP structure downsampling, as shown in Figure 5 As shown, the convolution module includes a convolution layer, a BN normalization layer and a Silu activation function.
[0086] In the last feature layer, the spatial pyramid pooling module (spp module) is used to adjust the feature size. Specifically:
[0087] The SoftPool pooling layer can reduce the size of the feature map and the amount of network calculation. Currently, the network uses maximum pooling. Although the maximum pooling calculation is very fast and occupies little memory, it will greatly lose information in the network, which is not conducive to the network detection of target objects. Therefore, the pooling operation based on softmax is used to enhance the feature map.
[0088] Define the local area of the feature map a of size C×H×W as R, where R represents the 2D spatial area, and its size is equal to the pooling kernel size. The core idea of the SoftPool pooling layer is to use softmax to nonlinearly calculate the eigenvalue weights of the local area R based on the eigenvalues:
[0089]
[0090] Among them, a i and a j Represents the value of each square in the R region, with weight w i It can ensure the transmission of important features. The feature values in region R will have at least the preset minimum gradient when they are transmitted in the reverse direction. i After that, the output is obtained by weighting the feature values in the area R
[0091]
[0092] The SoftPool pooling layer can refer to the distribution of activation values in the region well and obey a certain probability distribution, while the output of the maximum pooling and average pooling methods is undistributed.
[0093] Using SoftPool to pool the SPP modules in the network can minimize the loss of information. The structure is as follows: Figure 6 shown.
[0094] It should be noted that this embodiment uses the existing technology to pre-process the image, mainly to enhance the image, which is not described in detail in this embodiment.
[0095] (2)Reference Figure 7 , a dilated convolution module is used to increase the receptive field of feature maps of different scales output by the backbone feature extraction network to obtain second feature maps of different scales. The dilated convolution module uses multiple convolution kernels with different expansion rates for dilated convolution.
[0096] (3) The CA attention mechanism is integrated into the output layer of the backbone feature extraction network, and the second feature maps of different scales are input into the CA attention mechanism to better focus on the features that the backbone feature extraction network needs to observe. Specifically:
[0097] The CA attention mechanism decomposes channel attention into two one-dimensional feature encoding processes, aggregating features along two spatial directions respectively. Among them, long-range dependencies can be captured along one spatial direction, while precise position information can be retained along another spatial direction. The generated feature maps are then encoded into a pair of direction-aware and position-sensitive attention maps, which can be complementarily applied to the input feature map to enhance the representation of the object of attention. The specific operation is divided into two steps: coordinate information embedding and coordinate attention generation. The structure is as follows: Figure 8 As shown in the figure, the probability of gradient explosion is reduced through the residual module, a deeper network is built, and batch sample normalization is used to effectively prevent gradient disappearance and gradient explosion, and prevent overfitting.
[0098] The CA attention mechanism is then integrated into the feature layer of the backbone feature extraction network output, so that the backbone feature extraction network can better focus on the desired part. The structure is as follows: Fig. 9 shown.
[0099] (4) Input the second feature maps of different scales into the neck network of the fusion feature pyramid module (i.e. Figure 3 In the neck network, after each concatenation (i.e., concat, used to aggregate multiple convolution blocks), the multi-branch stacking module improved by the dilated convolution is used for convolution and training, while the inference is performed using ordinary convolution to the end (i.e., structural reparameterization) to obtain a multi-scale fusion feature map. Specifically:
[0100] In neural networks, high-level networks often have the characteristics of high semantics and low details, and there is no way to better meet the needs of target detection. Therefore, the feature pyramid module will be integrated in the neck network to perform multi-level feature fusion, and the high-level features with high semantic information but low resolution will be integrated with the low-level features with high resolution but low semantic information to meet the needs of target detection.
[0101] The processing and representation of multi-scale features has always been a difficult point in target detection. The pyramid network (FPN) is a pioneering structure that proposes to combine multi-scale features through top-down feature fusion. Then the path aggregation network (PANet) adds bottom-up bidirectional feature fusion based on single feature fusion, which also proves its effectiveness.
[0102] However, since the computational complexity of bidirectional feature fusion is almost doubled, the graph neural network is used for structural pruning to modify the above bidirectional feature fusion network. For example, the node with only one input edge in the neural network is deleted. It is generally believed that the fewer input nodes have a smaller impact on the entire graph neural network. Therefore, deleting the node will not have a significant impact on the network, but it can reduce a considerable amount of computation. Similarly, the output features of different layers in the neural network have different resolutions, and different resolutions contribute differently to the detection of targets of different sizes. Therefore, a learnable weight is set to be shared in the network, so that the network can learn the importance of features of different sizes at each layer, and use different weights for detection targets of different sizes to achieve divide and conquer.
[0103] In the neck network, after each concat fusion, the multi-branch stacking module improved by dilated convolution is used for convolution and training. Specifically:
[0104] Dilated convolution is also called expanded convolution or dilated convolution. The use of dilated convolution can increase the receptive field while keeping the size of the feature map unchanged, while reducing the loss of image information. Unlike normal convolution, dilated convolution introduces a hyperparameter called "dilation rate", which defines the spacing between values when the convolution kernel processes data. The dilation rate is also called the number of holes in Chinese. Taking a 3*3 convolution as an example, the difference between ordinary convolution and dilated convolution is shown, as follows: Fig.12 shown.
[0105] Fig.10 The middle left picture shows the normal convolution process, with dilation rate = 1 and a receptive field of 3 after convolution. Fig.10 The middle right picture shows a dilated convolution with a dilation rate of 2. The receptive field after convolution is 5. It can be considered that ordinary convolution is a special case of dilated convolution. Fig.10 It can be seen from the middle-right figure that the same 3*3 dilated convolution can have the effect of a 5*5 ordinary convolution. The dilated convolution increases the receptive field without increasing the number of parameters or doing pooling, preventing overfitting due to excessive parameters and information loss in pooling. However, if the dilated convolution is only performed by multiple superpositions of 3*3 convolutions with a dilation rate of 2, the feature map after the convolution will be discontinuous because not all pixels are calculated. Therefore, HDC (hybrid dilated convolution) is used to solve the problem of discontinuity. Multiple convolution kernels with different dilation rates are used for dilated convolution instead of the pooling layer to realize the function of the pooling layer and extract features. The multi-branch stacked module is replaced by a three-branch parallel dilated convolution structure. The structure is as follows: Fig.11 shown.
[0106] For inference, ordinary convolution is used to perform a complete convolution (i.e., structural reparameterization) to obtain a multi-scale fusion feature map. Specifically:
[0107] Structural reparameterization refers to constructing a series of structures (generally used for training) first. For example, the training structure of the parallel atrous convolution structure with three branches is constructed as follows: Fig.12 As shown, its parameters are equivalently converted into another set of parameters (generally used for reasoning). Its reasoning structure is as follows Fig.13 As shown, this series of structures is equivalently converted into another series of structures. In real-world scenarios, training resources are generally relatively abundant, and the cost and performance during inference are more important. Therefore, the structure during training is larger and has a good property (higher accuracy or other useful properties, such as sparsity), and the converted structure during inference is smaller and retains this property (same accuracy or other useful properties). In other words, the term "structural reparameterization" means converting a set of parameters of a structure into another set of parameters, and using the converted parameters to parameterize another structure. As long as the conversion of parameters is equivalent, the replacement of the two structures is equivalent.
[0108] (5) Then, the multi-scale fusion feature map is input into the head network for classification and regression, and the loss function of the head network is updated to the WIoU loss function to output the classification results and position regression results. Specifically:
[0109] The head network adopts the WIoU loss function (i.e., bounding box loss based on dynamic non-monotonic focusing mechanism). The WIoU loss function proposes the attention-based loss WIoU v1, designs the monotonic FM (i.e., static focusing mechanism) WIoU v2 and the dynamic non-monotonic WIoU v3. Among them:
[0110] The calculation formula of WIoU v1 loss is as follows:
[0111] L IoU =1-IoU
[0112] L WIoU v1 =R WIoU L IoU
[0113]
[0114] Among them, W g and H g Respectively represent the width and height of the minimum bounding box, R WIoU represents the penalty term of the WIoU loss function. In order to prevent R WIoU Produces gradients that hinder convergence, x, y, and xgt ,y gt Respectively represent the position of the center point of the anchor box and the center point of the target box, W g and H g It is separated from the computation graph (superscript * denotes this operation). Because it effectively removes factors that hinder convergence, no new metrics such as aspect ratio are introduced.
[0115] WIoU v2 constructs L based on WIoU WIoU v1 The monotonic focusing coefficient The formula is as follows:
[0116]
[0117] During the training process, the gradient gain increases with L IoU The decrease of decreases leads to slow convergence speed, so the average value is introduced as a factor:
[0118]
[0119] WIoU v3 defines the outlier degree to describe the quality of the anchor box β:
[0120]
[0121] A small outlier degree means a high quality anchor box. A non-monotonic focusing coefficient r is constructed using the outlier degree and applied to WIoU v1. The calculation formula is:
[0122] L WIoU v3 =rL WIoU v1
[0123] in, Represents the gradient gain, and IoU represents the degree of overlap between the predicted box and the real box in the target detection task. represents the mean value of IoU, γ represents, r represents the gradient gain, α and δ represent hyper parameters.
[0124] Reference Fig.14 The embodiment of the present invention further provides a multi-branch target detection system with multi-feature fusion. The multi-branch target detection system with multi-feature fusion includes a data acquisition unit 100, a feature extraction unit 200, a receptive field increase unit 300, a feature weighting unit 400, a feature fusion unit 500 and a target detection unit 600, wherein:
[0125] The data acquisition unit 100 is used to acquire a target detection image data set;
[0126] A feature extraction unit 200 is used to input the target detection image data set into the backbone feature extraction network to perform multi-scale feature extraction and obtain first feature maps of multiple scales;
[0127] A receptive field increasing unit 300 is used to increase the receptive field of the first feature maps of various scales by using a dilated convolution module to obtain second feature maps of various scales;
[0128] A feature weighting unit 400 is used to perform feature weighting on the second feature maps of multiple scales using a CA attention mechanism to obtain weighted feature maps of multiple scales;
[0129] A feature fusion unit 500 is used to perform multiple feature fusions on weighted feature maps of multiple scales using a neck network of a fused feature pyramid to obtain a multi-scale fused feature map;
[0130] The target detection unit 600 is used to input the multi-scale fusion feature map into the head network to perform multi-branch target detection and obtain target detection results for images of various sizes.
[0131] It should be noted that since the multi-feature fusion multi-branch target detection system in this embodiment and the multi-feature fusion multi-branch target detection method mentioned above are based on the same inventive concept, the corresponding contents in the method embodiment are also applicable to the system embodiment and will not be described in detail here.
[0132] An embodiment of the present invention further provides a multi-branch target detection device with multi-feature fusion, comprising: at least one control processor and a memory for communicating with the at least one control processor.
[0133] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0134] The non-transient software program and instructions required to implement the multi-branch target detection method of multi-feature fusion in the above embodiment are stored in the memory. When executed by the processor, the multi-branch target detection method of multi-feature fusion in the above embodiment is executed, for example, the above described Figure 1 The method comprises steps S100 to S600.
[0135] The system embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0136] The embodiment of the present invention further provides a computer-readable storage medium, which stores computer-executable instructions. The computer-executable instructions are executed by one or more control processors, so that the one or more control processors can execute a multi-branch target detection method with multi-feature fusion in the above method embodiment, for example, executing the above described Figure 1 The functions of method steps S100 to S600 in the embodiment.
[0137] It will be appreciated by those skilled in the art that all or some of the steps and systems in the disclosed method above may be implemented as software, firmware, hardware and appropriate combinations thereof. Some physical components or all physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor or a microprocessor, or may be implemented as hardware, or may be implemented as an integrated circuit, such as an application specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or a non-transitory medium) and a communication medium (or a temporary medium). As known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, disk storage or other magnetic storage devices, or any other medium that may be used to store desired information and may be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically embodies computer readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.
[0138] The above is a specific description of the preferred implementation of the embodiments of the present application, but the embodiments of the present application are not limited to the above-mentioned implementation methods. Technical personnel familiar with the field can also make various equivalent modifications or substitutions without violating the spirit of the embodiments of the present application. These equivalent modifications or substitutions are all included in the scope defined by the claims of the embodiments of the present application.
Claims
1. A multi-branch target detection method with multi-feature fusion, characterized in that: The multi-branch target detection method of multi-feature fusion includes: Obtain target detection image dataset; Inputting the target detection image data set into a backbone feature extraction network to perform multi-scale feature extraction to obtain first feature maps of multiple scales; Using a dilated convolution module to increase the receptive field of the first feature maps of the multiple scales to obtain second feature maps of the multiple scales; The second feature maps of the multiple scales are weighted by using a CA attention mechanism to obtain weighted feature maps of the multiple scales; The neck network of the fused feature pyramid is used to perform multiple feature fusions on the weighted feature maps of multiple scales to obtain a multi-scale fused feature map; specifically: A three-branch structured dilated convolution is adopted to train the feature map after the neck network of the fused feature pyramid performs feature fusion on the weighted feature maps of multiple scales each time, so as to obtain a plurality of trained first fused feature maps; the three-branch structured dilated convolution is a plurality of stacked parallel dilated convolution structures; Using structural reparameterization to infer the plurality of trained first fused feature maps to obtain a multi-scale fused feature map; The multi-scale fusion feature map is input into the head network to perform multi-branch target detection, and target detection results of images of various sizes are obtained.
2. The multi-branch target detection method of multi-feature fusion according to claim 1 is characterized in that: The backbone feature extraction network includes multiple feature layers and multiple convolutional layers, each of which includes downsampling and stacked convolutions integrated with dilated convolutions.
3. The multi-branch target detection method of multi-feature fusion according to claim 2 is characterized in that: The target detection image data set is input into a backbone feature extraction network to perform multi-scale feature extraction to obtain first feature maps of multiple scales, including: Inputting the target detection image data set into multiple convolutional layers for feature extraction to obtain a convolutional feature map; The convolution feature map is input into the first feature layer, the second feature layer, the third feature layer, and the fourth feature layer for feature extraction to obtain first feature maps of various scales; wherein the first feature layer adds a prediction head for tiny object detection.
4. The multi-branch target detection method of multi-feature fusion according to claim 3 is characterized in that: The first feature map output by the fourth feature layer is size-fixed using a spatial pyramid pooling module, and the spatial pyramid pooling module uses multiple SoftPool pooling layers.
5. The multi-branch target detection method of multi-feature fusion according to claim 1 is characterized in that: Before the neck network of the fused feature pyramid is used to perform multiple feature fusions on the weighted feature maps of multiple scales, the multi-branch target detection method of multi-feature fusion further includes: Using a graph neural network to structurally prune the neck network of the fused feature pyramid to obtain a pruned neck network; A learning weight parameter is set in the pruned neck network to perform feature learning on the weighted feature maps of the multiple scales.
6. The multi-branch target detection method of multi-feature fusion according to claim 1 is characterized in that: The head network adopts the WIOU loss function.
7. A multi-branch target detection system with multi-feature fusion, characterized in that: The multi-branch target detection system with multi-feature fusion includes: A data acquisition unit, used to acquire a target detection image data set; A feature extraction unit, used for inputting the target detection image data set into a backbone feature extraction network to perform multi-scale feature extraction to obtain first feature maps of multiple scales; A receptive field increasing unit, configured to increase the receptive fields of the first feature maps of the plurality of scales by using a dilated convolution module to obtain second feature maps of the plurality of scales; A feature weighting unit, used for performing feature weighting on the second feature maps of multiple scales by using a CA attention mechanism to obtain weighted feature maps of multiple scales; The feature fusion unit is used to perform multiple feature fusions on the weighted feature maps of the multiple scales by using the neck network of the fused feature pyramid to obtain a multi-scale fused feature map; specifically: A three-branch structured dilated convolution is adopted to train the feature map after the neck network of the fused feature pyramid performs feature fusion on the weighted feature maps of multiple scales each time, so as to obtain a plurality of trained first fused feature maps; the three-branch structured dilated convolution is a plurality of stacked parallel dilated convolution structures; Using structural reparameterization to infer the plurality of trained first fused feature maps to obtain a multi-scale fused feature map; The target detection unit is used to input the multi-scale fusion feature map into the head network to perform multi-branch target detection and obtain target detection results for images of various sizes.
8. A multi-branch target detection device with multi-feature fusion, characterized in that: It includes at least one control processor and a memory for communicating with the at least one control processor; the memory stores instructions that can be executed by the at least one control processor, and the instructions are executed by the at least one control processor so that the at least one control processor can execute the multi-branch target detection method with multi-feature fusion as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the multi-branch target detection method with multi-feature fusion as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Target detection method and detection system
CN115527094A