A Real-time Object Detection Method Combining Convolutional Neural Network and Transformer Network
By combining convolutional neural network and Transformer network, the detection backbone, neck and head networks are designed to solve the problems of insufficient remote dependencies and high computational complexity in the object detection model, and efficient real-time object detection is achieved.
Patent Information
- Application Number
- CN202210508625.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-10
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-05-10
AI Technical Summary
The existing object detection model lacks remote dependencies during local feature extraction, resulting in insufficient detection accuracy, and the visual Transformer network is difficult to train on finite data sets and has high computational complexity.
Combining convolutional neural network and Transformer network, a detection backbone network is designed to perform local feature extraction, providing high-resolution and semantic features by detecting the neck network, and introducing a streamlined Transformer network to construct remote dependencies, and a nonlinear combination method is used to reduce false negative samples.
Real-time object detection that rapidly converges on finite data sets is realized, which improves detection accuracy and reduces model parameters and calculation complexity, and improves the ability to capture targets.
Smart Images

Figure CN114842316B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image processing and relates to a real-time object detection method combining a convolutional neural network and a Transformer network. Background Art
[0002] Object detection is an attractive and challenging topic in computer vision. Its attractiveness comes from a wide range of applications such as autonomous driving and robot navigation, while the challenges come from the changing scales, complex shapes, and multiple categories. With the rapid development of Convolutional Neural Networks (CNNs), the number of object detection models has increased rapidly. Although the models are diverse, they can all be divided into anchor-based methods or anchor-free methods through deep stacks of convolutional operations. These methods are sensitive to local regions of interest and require fewer parameters than Multi-Layer Perception (MLP). However, a significant feature of these methods is that the image features extracted from the detection network are only limited to local regions. This lacks long-range semantic correlations, and long-range dependencies are important for the network to focus on regions of interest and ignore noise in the entire feature map. In addition, other works have mathematically proven that the effective receptive field (RF) of the extracted features is much smaller than the theoretical one, which means that the deep stack mechanism of convolutional operations is unrealistic in establishing long-range dependencies between local image features.
[0003] Therefore, to overcome the limitations of the inherent locality of convolutional operations, some self-attention mechanisms based on local features have been proposed. On the other hand, the Transformer is a creative network mainly used for natural language processing, which parallelly mines multiple long-range correlations between time series information and has recently been introduced into the field of computer vision, achieving state-of-the-art results in many visual tasks. The success of various visual Transformer networks proves the necessity of building long-range dependencies.
[0004] However, compared with CNNs, due to the lack of inductive biases such as translational invariance and locality, visual Transformer networks cannot generalize well, which means that sufficient data quantity or a reasonable combination of training techniques is required during training. On the other hand, high-resolution images are processed under the accumulation of the Self-Attention Network (SAN) and MLP in visual Transformer networks, which will lead to a rapid increase in computational complexity. In addition, many object detection networks will generate many correct boxes with low scores and incorrect boxes with high scores during detection.
[0005] After retrieval, the application publication number is CN110765886A, a target detection method and device based on a convolutional neural network. The method includes: importing a real-time image into a target detection network and outputting the target objects included in the real-time image; the target detection network includes a convolutional layer, a transposed convolutional layer, a feature enhancement block, a feature fusion block, a first regressor, and a second regressor. Since the features extracted by the convolutional neural network are local, the features often only reflect the local region characteristics of the image, introducing biases for subsequent feature extraction and final prediction output. The Transformer network is a network that constructs long-range dependencies between features. Its characteristic is that the key-value pair self-attention mechanism is the core of feature extraction, and it has excellent performance in processing global information. However, since the Transformer network is a neural network model without inductive bias, more data volume and data augmentation methods are required to converge during the training process and have good generalization performance.
[0006] To solve the training and prediction biases introduced by the local features of the convolutional neural network and the problem of difficult convergence of the Transformer network in the training of a limited data set, the present invention proposes a method combining a convolutional neural network and a Transformer network to solve the above problems, thereby realizing real-time and accurate detection of objects. The target detection network model of the present invention includes a convolutional backbone network, a feature fusion network, and a lightweight Transformer detection head network. The convolutional neural network is located in the backbone network part, enabling the extracted features to have local and invariant characteristics, and constructing long-range dependencies between features through the subsequent Transformer network. The feature fusion network provides rich high-resolution semantic features for the detection head network. The lightweight Transformer detection head network uses the input with local and invariant characteristics, making the network easy to converge on a limited data set. In addition, the present invention reduces the number of parameters and computational complexity of the network by introducing a lightweight spatial reduction module (Lite Spatial Reduction, LSR) and a lightweight multi-layer perceptron (Lite MLP) into the lightweight Transformer detection head network, further improving the convergence performance of the network. Experiments confirm that the method proposed by the present invention can reduce the number of parameters of the network model while improving the detection accuracy. Finally, by proposing a non-linear combination method (i.e., using a logarithmic operation on the object confidence and classification score and introducing an additional hyperparameter), the capture ability of the detection model for the target is further improved. Summary of the Invention
[0007] The present invention aims to solve the above problems of the prior art. A real-time target detection method combining a convolutional neural network and a Transformer network is proposed. The technical solution of the present invention is as follows:
[0008] A real-time object detection method combining a convolutional neural network and a Transformer network, comprising the following steps:
[0009] S1: Input training image data into the network;
[0010] S2: Design a convolutional neural backbone network to extract features from the images for training, so that the extracted features have inductive bias characteristics, that is, the features extracted by the convolutional neural network have locality and translational invariance;
[0011] S3: Design a detection neck network to transition between the detection backbone network and the head network, provide high-resolution and high-semantic features for the detection head network; and compress the channel dimensions of some layers;
[0012] S4: Design a detection head network, introduce a Transformer network, that is, a full self-attention network, in the head network to build multiple long-range dependencies between the generated local features, and represent the target categories and coordinates existing in the image;
[0013] S5: Design a non-linear combination method. Since there are some false negative FN samples in the results output by the object detection model in S4, a non-linear combination method is designed to reduce the false negative samples and improve the target capture ability of the detection model;
[0014] S6: Perform detection on the natural dataset, and use the non-maximum suppression algorithm to screen the prediction results. Calculate the intersection over union IoU between the screened prediction results and the true target boxes and count the prediction results, and then obtain the average precision value AP, the average precision value AP50 under the condition that the IoU threshold is 0.5, and the average precision value AP75 under the condition that the IoU threshold is 0.75 as the evaluation results of the model;
[0015] Further, the step S1 inputs the image dataset to be trained, specifically including the following steps:
[0016] The training image dataset uses the PASCAL VOC and MS COCO datasets. The training batch size for each iteration is set to 24. Perform 50 multi-scale trainings (320, 352, 384, 416, 448, 480, 512, 544, 576, and 608) on PASCAL VOC, and the size of the image during testing is 448; perform 300 three-scale trainings (320, 352, and 384) on MS COCO, and the length and width of the input image during testing are both 320; use a post-processing algorithm to screen the output results to obtain the final prediction results.
[0017] Further, after the step S1, a post-processing algorithm is used to screen the output results to obtain the final prediction results, which specifically includes:
[0018] The prediction results are screened using the non-maximum suppression algorithm. For the screened results, the intersection over union (IoU) between the predicted results of the samples and the true target boxes is calculated. First, according to two preset thresholds, namely the PASCAL criterion and the standard MS COCO criterion, the sample attributes are determined. The PASCAL criterion is IoU > 0.5 and IoU > 0.7, and the MS COCO criterion is that the IoU threshold ranges from 0.5 to 0.95 with a step size of 0.05. All samples are sorted from high to low according to their classification results. The sorted samples are traversed, and the accuracy and recall are calculated for the traversed samples according to formulas (1) and (2):
[0019]
[0020]
[0021] where TP, FP, and FN represent true positives, false positives, and false negatives;
[0022] According to the accuracy and recall obtained from each traversal, a curve with recall as the X-axis and accuracy as the Y-axis is constructed. Finally, by calculating the area of the curve enclosed by the accuracy and recall, the average precision (AP) is obtained, and the mean average precision (mAP) is obtained by calculating the AP for each category. Three metrics, namely the average precision value AP, AP50 (average precision value under the condition that the IoU threshold is 0.5), and AP75 (average precision value under the condition that the IoU threshold is 0.75), are used as the model evaluation criteria. AP50 is the average precision value when the IoU threshold is 0.5, and AP75 is the average precision value when the IoU threshold is 0.75. In addition, for small, medium, and large-sized objects, AP small 、AP middle and AP large are also used for evaluation.
[0023] Further, the step S3 specifically includes:
[0024] The described neck network for detection includes a feature data compression part and a feature fusion part; the feature data compression part is located in the third and fourth network layers of the convolutional neural backbone network. The third and fourth networks provide rich semantic information, and the feature data compression part compresses the features extracted by the backbone network through depthwise separable convolution calculation. Then, bilinear interpolation is used to upsample the compressed features so that the spatial size of the features in the third and fourth network layers is the same as that of the features in the second network layer. Finally, the interpolated feature data is concatenated with the feature data in the second network layer in the channel dimension; the second network layer provides high-resolution low-level information, and the feature fusion part further fuses the concatenated data through depthwise separable convolution, that is, fuses the data features of the second, third, and fourth network layers.
[0025] Further, in step S4, a multi-branch detection head network is designed, which specifically includes the following steps:
[0026] At the input of each branch, dilated depth convolution is set to expand the receptive field RF of different head branches, and then a split operation is performed. In the split operation, the features in the feature fusion network FFN are divided into two parts in the channel dimension, so that the channel dimension of the split features is half of the original feature channel number. One part passes through LT, and the other part is concatenated with the output of LT in the channel dimension. Then, the fusion operation is responsible for fusing the concatenated features with a 1×1 convolutional kernel and outputting them after passing through the Leaky ReLU activation function. The corresponding patch sizes are present in the lightweight Transformer network LT in different detection head network branches. The dilation factors and patch size parameters on the detection head network branches for large, medium, and small scale objects will be set to 4, 2, and 1 respectively.
[0027] Further, in step S4, as Figure 4 The lightweight Transformer network LT is divided into three different parts. The first part divides the input features into non-overlapping patches and adds the learnable position vectors element-wise to the vector mappings of different patches to ensure the unique position of each patch vector mapping; in addition, lightweight convolution operations are introduced into the LT network to reduce the computational complexity in the process of obtaining patch vector mappings; each patch vector mapping also passes through the GELU activation function before output.
[0028] It can be expressed in formula (3):
[0029]
[0030] x represents the output data after passing through the lightweight patch vector mapping, and GELU(x) represents the output data through this activation function;
[0031] In the second part, after the position vector is combined with the tile vector mapping, each tile vector mapping query value, key value, and variable value will be calculated, and multiple global correlation information will be constructed in parallel in the multi-head spatial attention module (MSAN). The lightweight spatial reduction module (LSR) is used in both the key value and variable value calculation branches to traverse the key value and variable value through non-overlapping convolution operations to compress their sequence sizes. The global correlation information calculation complexity ratio between LT and ViT's multi-head self-attention layer is as follows:
[0032]
[0033] where p, N, and C represent the number of spatial reduction times, the total number of tiles, and the number of channel dimensions, respectively;
[0034] The parallel construction of multiple global correlation information in the multi-self-attention network (MSAN) can be expressed in formula (5):
[0035]
[0036] where q, k, and v represent the query value, key value, and variable value respectively, and d head represents the key value channel dimension number. The lightweight spatial reduction module is used in both the key value and variable value calculation branches, and finally, a multi-layer perceptron (MLP) and a shortcut connection are used for further global feature extraction.
[0037] Furthermore, step S5 specifically includes:
[0038] An additional hyperparameter is introduced, which can be expressed in formula (6) as:
[0039] R = log3(1 + α · C) · S (6)
[0040] where R represents the combined result, α is a hyperparameter that controls the prediction box result, C is the object confidence, and S is the category score with the highest probability.
[0041] The advantages and beneficial effects of the present invention are as follows:
[0042] The innovation of the present invention is mainly the combination of claims 2-9. The present invention proposes a real-time object detector Transformers Only Look Once (TOLO) that combines a convolutional neural network and a Transformer network, achieving real-time and accurate detection of objects. The detector mainly consists of a detection backbone network, a detection neck network, and a detection head network. The detection backbone network is used to extract image features, the detection neck network fuses features of different network layers, and the detection head network is used to represent the category and coordinates of the object. Among them, for the detection neck network, a feature fusion network is proposed to achieve feature fusion of different network layers in the same path in a lightweight manner, solving the problems of inconsistent feature interaction between different network layers and high computational complexity of the neck network; for the detection head network, it consists of three different streamlined Transformer head branches, which are used to detect large, medium, and small-scale objects respectively. The streamlined Transformer inside each branch is used to extract multiple long-range dependencies between features and achieve object detection with less memory overhead, solving the problems of low efficiency in constructing long-range dependencies in previous real-time detection models and high computational complexity of the vision Transformer network. In addition, in order to discover a large number of potential correct prediction samples during the prediction process, the present invention proposes a non-linear combination method between the object confidence and the classification score, improving the capture ability of the detection model for objects. Therefore, the present invention improves the model from three aspects: the detection neck network, the detection head network, and the non-linear combination method, solving the problems of low efficiency in constructing long-range dependencies in real-time detection models, excessive computational complexity, and too many false negative samples, reducing the number of parameters of the detection model, and improving the performance of real-time object detection. Description of the Drawings
[0043] Figure 1 It is a flowchart of three core modules and the overall structure of the preferred embodiment provided by the present invention.
[0044] Figure 2 It is the structure of the detection neck network.
[0045] Figure 3 It is the structure of the detection head network.
[0046] Figure 4 It is the structure of the LT network.
[0047] Figure 5 It is a comparison of the effects of linear combination and non-linear combination.
[0048] Figure 6 It is the visualization result of linear and non-linear combination. Detailed Embodiment
[0049] Next, the technical solutions in the embodiments of the present invention will be clearly and detailedly described in conjunction with the accompanying drawings in the embodiments of the present invention. The described embodiments are only a part of the embodiments of the present invention.
[0050] The technical solution of the present invention to solve the above technical problems is as follows:
[0051] A real-time object detection method combining a convolutional neural network and a Transformer network, the method comprising the following steps:
[0052] S1: Input training image data into the network;
[0053] S2: Design a convolutional neural backbone network to extract features from the images for training, so that the extracted features have inductive bias characteristics, that is, the features extracted by the convolutional neural network have locality and translational invariance;
[0054] S3: Design a detection neck network to transition between the detection backbone network and the head network, provide high-resolution and high-semantic features for the detection head network; and compress the channel dimensions of some layers;
[0055] S4: Design a detection head network, introduce a Transformer network, that is, a full self-attention network, in the head network to construct multiple long-range dependencies between the generated local features, and represent the target categories and coordinates existing in the image;
[0056] S5: Design a non-linear combination method. Since there are some false negative FN samples in the results output by the object detection model in S4, a non-linear combination method is designed to reduce the false negative samples and improve the target capture ability of the detection model;
[0057] S6: Perform detection on the natural dataset, and screen the prediction results using the non-maximum suppression algorithm. Calculate the intersection over union IoU between the screened prediction results and the true target boxes and count the prediction results, and then obtain the average precision value AP, the average precision value AP50 under the condition that the IoU threshold is 0.5, and the average precision value AP75 under the condition that the IoU threshold is 0.75 as the evaluation results of the model;
[0058] Optionally, the S1 specifically includes the following steps:
[0059] The training image data uses the PASCAL VOC and MS COCO datasets. The training batch size for each iteration is set to 24. In this invention, 50 multi-scale trainings (320, 352, 384, 416, 448, 480, 512, 544, 576, and 608) are carried out on PASCAL VOC, and the size of the image during testing is 448; 300 three-scale trainings (320, 352, and 384) are carried out on MS COCO, and the length and width of the input image during testing are both 320.
[0060] In S2, for the ViT (Vision Transformer) network with high computational costs, there is a lack of inductive bias. This invention remedies this shortcoming by combining the Convolutional Neural Network (CNN) with Transformer. The image passes through the convolutional neural backbone network, enabling the extracted features to have the characteristics of inductive bias, that is, the features extracted by the convolutional neural network have locality and translational invariance.
[0061] Preferably, in S3, as Figure 2 shown, the network layers in the last three stages of the backbone network are mainly concerned. For the features in the last two layers, this invention compresses the channel dimension to half of the original by using depthwise separable convolution operations. Because compared with traditional convolution operations, this convolution operation method requires fewer convolution kernel parameters and less computational effort. It should be noted that the compressed feature channel dimension number is related to the original channel number, rather than outputting a fused feature with a fixed channel dimension size. Therefore, this operation can not only better avoid a large amount of loss of semantic information (especially in the deep layers), but also ensure the fusion efficiency of features due to the lower computational complexity. After the features are compressed in the channel dimension, they are upsampled by bilinear interpolation to make the spatial size the same as that of the first-layer features. Then, the features in different layers are concatenated in the channel dimension and feature fusion is achieved through depthwise separable convolution to avoid a large increase in computational complexity. Therefore, the network layer features at each stage can directly interact with the features of other network layers through the same path, without going through a discontinuous layer-by-layer combination interaction method.
[0062] Preferably, in S4, as Figure 3As shown, the head network of the present invention consists of three different detection head networks (Lite Transformer Heads, LTHs), which respectively detect large, medium, and small-scale objects. For a large number of real-time object detectors, it is inefficient to establish long-range distance dependencies between local features. The present invention overcomes this defect by combining CNN and Transformer. For all LTHs, the internal Lite Transformer (LT) will be used to efficiently mine the global correlation information between features. The structures of the LTHs for different-scale objects are somewhat different. At the input of each branch, dilated depth convolution is set to further expand the receptive field (RF) of different head branches. The LT in different detection head network branches has its corresponding patch size. The dilation factors and patch size parameters on the detection head branch networks for large, medium, and small-scale objects will be set to 4, 2, and 1 respectively. For the large-scale detection head branch, rich channel information from the Feature Fusion Network (FFN) is mainly considered, and the image patch size in the LT is set larger; for the detection head branches of medium-scale and small-scale objects, since the local changes in the image become more important, more attention will be paid to the spatial local features from the Feature Fusion Network (FFN) during the design process. Therefore, the patch size is set relatively small. In addition, in the separation operation, the present invention divides the features in the FFN into two parts in the channel dimension, so that the channel dimension of the separated features is half of the original features. One part passes through the LT, and the other part is concatenated with the output of the LT in the channel dimension. Then, the fusion operation is responsible for fusing the concatenated features with a 1×1 convolution kernel and outputting after passing through the Leaky ReLU activation function. In the detection head network for small-scale objects, a compression operation is first used to reduce the channel dimension of the FFN features. In the detection head branch network for medium-scale objects, average pooling operation will be used to compress the spatial size of the local features. Finally, the mapping operation is used to adjust the channel dimension to adapt to the final prediction result. As Figure 4As shown, the present invention proposes a streamlined visual Transformer network for real-time target detection. Unlike previous visual Transformer networks, the LT network is lightweight, which means that the network model will be very fast in the training and inference stages. LT can be divided into three different parts. The first part focuses on dividing the input features into non-overlapping tiles and adding the learnable position vectors to the vector mappings of different tiles element by element to ensure the positional uniqueness of each tile vector mapping. In addition, lightweight convolution operations will be introduced into the LT network to reduce the computational complexity of obtaining tile vector mappings. Each tile vector mapping will also pass through a GELU activation function before output.
[0063]
[0064] x represents the output data after mapping by the simplified tile vector, and GELU(x) represents the output data through the activation function.
[0065] After combining the position vector with the tile vector mapping, in the second part, the query value, key value, and variable value of each tile vector mapping are calculated, and multiple global association information are constructed in parallel in the Multi-head Spatial Attention Module (MSAN). It can be clearly seen that we use the Lite Spatial Reduction (LSR) module in both the key value and variable value calculation branches. Its purpose is still to use lightweight as the main calculation method. Non-overlapping convolution operations are used to traverse the key and variable values to compress their sequence size, thereby reducing computational complexity. The computational complexity of the global association information of the multi-head self-attention layer between LT and ViT is:
[0066]
[0067] where p, N, and C represent the number of spatial reductions, the total number of tiles, and the number of channel dimensions, respectively. We omit the complexity statistics for computing the query value, key value, variable value, and final output here because they are the same for both visual Transformer network models. From Equation (2), we can see that the LSR module can reduce the complexity by a factor of P.
[0068] We use MLP and short-circuiting to further extract global features to avoid the degradation of the self-attention network's expressive power as the network depth increases. However, the large number of parameters in the MLP also reduces the network's feature extraction efficiency. Therefore, in LT, we set the number of hidden layer neurons of the proposed Lite MLP to be the same as the input channel dimension. The ratio of its computational complexity to the MLP in ViT can be calculated as:
[0069]
[0070] Among them, N and C represent the total number of patches and the channel dimension size respectively. It can be obtained from formula (3) that the Lite MLP in LT saves four times the parameters of the MLP in ViT. Continuing to compress the number of neurons in the hidden layer can further reduce the number of parameters and computational complexity. However, we found that the detection performance will also decline.
[0071] Preferably, in S5, instead of directly multiplying the object confidence and classification score commonly used in the Transformers Only Look Once (TOLO) network, an additional hyperparameter is introduced and the calculation method is changed, which can be expressed in formula (4) as:
[0072] R = log3(1 + α·C)·S (4)
[0073] Where R represents the combined result, α is the hyperparameter controlling the prediction box result, C is the object confidence, and S is the score of the most likely class. Figure 5 The effects of different combination methods are shown in. By assuming that the score of the most likely classification label is a random constant S independent of the confidence C, the dashed line in the figure is the standard linear combination method. It can be clearly seen that when α is 1 or 1.5, it has a strong inhibitory effect and the higher the confidence, the stronger the inhibitory effect. When α is 2, too low confidence will not have a great impact on the final result, indicating that the influence of true negative prediction boxes can be avoided. However, when the confidence is within 0.4 to 0.6, the combined result will be greatly improved. Therefore, the number of false negative prediction boxes can be reduced by increasing the value of α. The visualization of the results of the non-linear combination method and the linear combination method is as Figure 6 shown.
[0074] Preferably, in S6, the performance of the target detector and the non-linear combination method is evaluated on the PASCAL VOC 2007, PASCAL VOC 2012, and MS COCO 2017 datasets. The PASCAL VOC 2007 dataset has 20 categories and 9,962 images; the PASCAL VOC 2012 dataset has 20 categories and 22,531 images. These two datasets are divided into training set, validation set, and test set respectively. And the standard PASCAL VOC standard is followed, that is, the mean Average Precision (mAP) is used as the test evaluation metric. The present invention uses three metrics: Average Precision (AP), AP50 (the average precision value under the condition that the IoU threshold is 0.5), and AP75 (the average precision value under the condition that the IoU threshold is 0.75), which are the standard PASCAL criteria (i.e., Intersection over Union (IoU) > 0.5, IoU > 0.7) and the standard MS COCO criteria (i.e., calculate the mean mAP of IoU ∈ [0.5:0.05:0.95]). In addition, for small, medium, and large-sized objects, AP small , AP middle , and AP large are also used for evaluation. In terms of the training strategy, we use Stochastic Gradient Descent (SGD) to optimize our model, set the initial learning rate to 0.001, and train on two GPUs (GTX 3090). The cosine learning rate schedule is also set between 0.001 and 0.00001. The weight decay is 0.0005, and the momentum is 0.9. In addition, some training techniques such as MixUp and label smoothing are adopted to avoid overfitting and improve the generalization of the model.
[0075] Experimental Results
[0076] In this example, we evaluated the effectiveness of the proposed target detector on the PASCAL VOC and MS COCO datasets. The performance of TOLO was compared with other state-of-the-art detectors, real-time detectors, and some Transformer-based detectors that include first-level or second-level categories.
[0077] Table 1 Detection Results of Different Detectors on PASCAL VOC
[0078]
[0079]
[0080] Table 2 Detection Results of Real-Time Detectors on MS COCO
[0081]
[0082] Table 3 Detection Results of Transformer-Based Detectors on MS COCO
[0083]
[0084] The systems, devices, modules or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.
[0085] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, commodity or device including the element.
[0086] The above embodiments should be understood as being only for illustrative purposes of the present invention and not for limiting the protection scope of the present invention. After reading the content described in the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent changes and modifications also fall within the scope defined by the claims of the present invention.
Claims
1. A real-time object detection method combining a convolutional neural network and a Transformer network, characterized in that, It includes the following steps: S1: Input training image data into the network; S2: Design a convolutional neural backbone network to extract features from the images for training, enabling the extracted features to have inductive bias characteristics, that is, the features extracted by the convolutional neural network have locality and translational invariance; S3: Design a detection neck network to make a transition between the detection backbone network and the head network, providing high-resolution and high-semantic features for the detection head network; and compress the channel dimensions of some layers; S4: Design a detection head network, introduce a Transformer network, i.e., a full self-attention network, in the head network to build multiple long-range dependencies between the generated local features, and represent the target categories and coordinates existing in the image; S5: Design a non-linear combination method. Since there are some false negative (FN) samples in the results output in S4 by the target detection model, a non-linear combination method is designed to reduce false negative samples and improve the target capture ability of the detection model; S6: Conduct detection on the natural dataset and filter the prediction results using the non-maximum suppression algorithm; Calculate the intersection over union (IoU) between the filtered prediction results and the true target boxes and count the prediction results, and then obtain the average precision (AP) value, the average precision value AP50 under the condition that the IoU threshold is 0.5, and the average precision value AP75 under the condition that the IoU threshold is 0.75 as the evaluation results of the model; The step S1 inputs the image dataset to be trained, which specifically includes the following steps: The training image dataset uses the PASCAL VOC and MS COCO datasets. The training batch size for each iteration is set to 24. Conduct 50 multi-scale trainings on PASCAL VOC, and the multi-scale sizes are 320, 352, 384, 416, 448, 480, 512, 544, 576, and 608. The size of the image during testing is 448; Conduct 300 three-scale trainings on MS COCO, and the three-scale sizes are 320, 352, and 384. The length and width of the input image during testing are both 320; Use a post-processing algorithm to filter the output results to obtain the final prediction results; The step S1 uses a post-processing algorithm to filter the output results to obtain the final prediction results, which specifically includes: Filter the prediction results using the non-maximum suppression algorithm. Calculate the IoU between the filtered results and the true target boxes of the samples. First, determine the sample attributes according to two preset thresholds, i.e., two criteria: the PASCAL criterion and the standard MS COCO criterion; The PASCAL criterion is IoU > 0.5 and IoU > 0.7, and the MS COCO criterion is that the IoU threshold ranges from 0.5 to 0.95 with a step of 0.
05. Sort all samples from high to low according to their classification results; Traverse the sorted samples, and calculate the accuracy and recall rates of the traversed samples according to formula (1) and formula (2): Where TP, FP, and FN represent true positives, false positives, and false negatives; Based on the accuracy and recall obtained from each traversal, a curve is constructed with recall as the X-axis and accuracy as the Y-axis; finally, by calculating the area of the curve enclosed by the accuracy and recall, the average precision AP is obtained, and the mean average precision mAP is obtained by calculating the AP for each category; three metrics are used, namely the average precision value AP, AP50 (average precision value under the condition that the IoU threshold is 0.5), and AP75 (average precision value under the condition that the IoU threshold is 0.75) as the model evaluation criteria, AP50 is the average precision value under the condition that the IoU threshold is 0.5, and AP75 is the average precision value under the condition that the IoU threshold is 0.
75. In addition, for small, medium, and large-sized objects, AP small , AP middle , and AP large are also used for evaluation; The specific steps of step S3 include: The detected neck network includes a feature data compression part and a feature fusion part; the feature data compression part is located in the third and fourth network layers of the convolutional neural backbone network. The third and fourth networks provide rich semantic information. The feature data compression part compresses the features extracted by the backbone network through the depthwise separable convolution calculation method; then, bilinear interpolation is used to upsample the compressed features to make the spatial size of the features in the third and fourth network layers the same as that of the features in the second network layer; finally, the interpolated feature data and the feature data in the second network layer are concatenated in the channel dimension; the second network layer provides high-resolution underlying information, and the feature fusion part further fuses the concatenated data through depthwise separable convolution, that is, fuses the data features of the second, third, and fourth network layers; In step S4, a multi-branch detection head network is designed, which specifically includes the following steps: At the input of each branch, dilated depth convolution is set to expand the receptive field RF of different head branches, and then a split operation is performed. In the split operation, the features in the feature fusion network FFN are divided into two parts in the channel dimension, so that the channel dimension of the split features is half of the original feature channel number. One part passes through LT, and the other part is concatenated with the output of LT in the channel dimension. Then, the fusion operation is responsible for fusing the concatenated features with a 1×1 convolutional kernel and outputting them after passing through the Leaky ReLU activation function; the lightweight Transformer network LT in different detection head network branches has corresponding patch sizes, and the dilation factors and patch size parameters on the detection head network branches for large, medium, and small scale objects are set to 4, 2, and 1 respectively; In step S4, the lightweight Transformer network LT is divided into three different parts. The first part divides the input features into non-overlapping patches and adds the learnable position vectors element-wise to the vector mappings of different patches to ensure the unique position of each patch vector mapping; in addition, lightweight convolution operations are introduced into the LT network to reduce the computational complexity in the process of obtaining patch vector mappings; each patch vector mapping also passes through the GELU activation function before output; It is represented in formula (3): x represents the output data after passing through the lightweight patch vector mapping, and GELU(x) represents the output data through this activation function; In the second part, after the position vectors are combined with the patch vector mappings, the query values, key values, and variable values of each patch vector mapping are calculated, and multiple global correlation information is constructed in parallel in the multi-head spatial attention module MSAN. The lightweight spatial reduction module LSR is used in both the key value and variable value calculation branches to compress their sequence sizes by traversing the key values and variable values through non-overlapping convolution operations. The computational complexity ratio of the global correlation information calculation of the multi-head self-attention layer between LT and ViT is: where p, N, and C represent the number of spatial reduction times, the total number of patches, and the channel dimension number respectively; Constructing multiple global correlation information in parallel in the multi-head spatial attention module MSAN, which is expressed in formula (5): where q, k, and v represent the query value, key value, and variable value respectively, and d head represents the key-value channel dimension number. A lightweight spatial compression module is used in both the key-value and variable-value calculation branches, and finally, a multi-layer perceptron (MLP) and shortcut connection are used for further global feature extraction.
2. The real-time object detection method combining a convolutional neural network and a Transformer network according to claim 1, characterized in that The specific steps of step S5 include: An additional hyperparameter is introduced, which is expressed in formula (6) as: R = log3(1 + α·C)·S (6) where R represents the combined result, α is the hyperparameter controlling the prediction box result, C is the object confidence, and S is the category score with the highest probability.
Citation Information
Patent Citations
Road target detection method and device based on convolutional neural network
CN110765886A