Weak Pedestrian Detection Method Based on Skip Connection Context and Channel Attention

By introducing a jump connection context module and a channel attention module in the detection of weak pedestrian targets, the problems of high computational complexity and low feature transmission efficiency during the detection process are solved, and the detection accuracy of weak pedestrian targets is improved.

CN116721439BActive Publication Date: 2025-07-01XIDIAN UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310670161.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-07
Publication Date
2025-07-01
Estimated Expiration
2043-06-07

AI Technical Summary

Technical Problem

When detecting weak pedestrian targets, the prior art has high computational complexity and low feature delivery efficiency, and is easily covered by surrounding irrelevant background information, resulting in low detection accuracy.

Method used

The jump connection context module is used to improve the feature transmission efficiency between the expanded convolutional layers, and the channel attention module uses the average pooling of the channel dimension to calculate the channel attention to avoid the influence of irrelevant information in the spatial dimension.

Benefits of technology

The calculation complexity is reduced and the feature representation ability and detection accuracy of weak pedestrian targets are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116721439B_ABST
    Figure CN116721439B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for detecting small and weak pedestrians based on skip connection context and channel attention, which solves the problems of high computational complexity and limited feature transfer efficiency in the existing technology detection process, and the fact that small and weak pedestrian targets are easily covered by surrounding irrelevant background information, resulting in low detection accuracy of the detection network. The implementation steps of the present invention are as follows: constructing a skip connection context module; constructing a channel attention module; constructing a small and weak pedestrian target detection network; generating a pre-training set and a training set; training the small and weak pedestrian target detection network; and detecting small and weak pedestrian targets. Based on the constructed skip connection context module and temporal aggregation module, the present invention constructs a small and weak pedestrian detection network, which improves the detection efficiency and accuracy of small and weak pedestrian targets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image recognition or understanding, and further relates to a method for detecting small and weak pedestrian targets based on skip connection context and channel attention in the field of computer vision technology. The present invention can be used to detect small and weak pedestrian targets in RGB images. Background Art

[0002] The detection of small and weak pedestrian targets belongs to the intersection of the research directions of pedestrian target detection and small target detection, and is essential in many real-world applications, such as autonomous driving and intelligent monitoring. General pedestrian target detection networks can well complete the accurate detection tasks of medium and large-sized pedestrians. However, when these detection networks are applied to detect small and weak pedestrian targets, the detection accuracy is greatly reduced because the resolution of small and weak pedestrian targets is low and the information contained is relatively small, making it difficult for general pedestrian target detection networks to capture discriminative features. In the present invention, a small and weak pedestrian target refers to a pedestrian target with a height lower than 75 pixels in an image. Such a small target is easily overlooked when performing target detection in an image.

[0003] Cui L et al. disclosed a method for detecting small targets based on a context-aware network in their published paper "Context-Aware Block Net for Small Object Detection" (IEEE Transactions on Cybernetics, 2020, PP(99): 1-14). The context-aware network (Context-Aware Block Net, CAB Net) in this method includes a context-aware block (Context-Aware Block, CAB) and a backbone network (VGG) with a discarded part of the structure. Among them, the structure of CAB is a multi-branch in parallel, and each branch is sequentially composed of a specific number of dilated convolutional layers in series. The disadvantages of this method are that each branch is connected in parallel, and each branch outputs a context feature, which increases the computational complexity when dealing with multiple context features during the detection process; the dilated convolutional layers are connected in series within each branch, resulting in limited feature transfer efficiency between the dilated convolutional layers.

[0004] Nanjing University of Aeronautics and Astronautics discloses a pedestrian detection method in its patent document "A Pedestrian Detection Method of Lightweight YOLOv4 from the Perspective of Unmanned Aerial Vehicle" (Patent Application No.: CN 202211000295.3, Publication No. CN115359376 A). This method uses an improved MobileNetv3-YOLOv4 lightweight pedestrian target detection network, and a SESAM attention module is introduced into this network. Among them, the SESAM attention module includes a compression module and an excitation module. The compression module contains a global average pooling layer, which is used to compress the feature vector in the spatial dimension to obtain the global feature vector of each channel and input it into the excitation module; the excitation module contains two fully connected layers, which are used to perform non-linear transformation on the compressed feature vector. The disadvantage of this method is that the global average pooling layer is used to perform average pooling processing on the input features in the spatial dimension, so that small-sized pedestrian targets are easily covered by surrounding irrelevant background information, resulting in insufficient attention of the detection network to small-sized pedestrian targets and low detection accuracy. Summary of the Invention

[0005] The purpose of the present invention is to address the deficiencies of the above-mentioned existing technologies and propose a method for detecting small and weak pedestrians based on skip connection context and channel attention, which is used to solve the problems that the existing technology uses multiple parallel branches, each branch is composed of a specific number of dilated convolutional layers connected in series in sequence, resulting in high computational complexity and limited feature transfer efficiency during the detection process, and using average pooling in the spatial dimension to calculate channel attention, so that small and weak pedestrian targets are easily covered by surrounding irrelevant background information, resulting in insufficient attention of the detection network to small-sized pedestrian targets and low detection accuracy.

[0006] To achieve the above object, the idea of the present invention is that in the constructed skip connection context module, the input layer, the first convolutional layer, the first dilated convolutional layer, the second dilated convolutional layer, the third dilated convolutional layer, the adder, the third convolutional layer, and the output layer are connected in series in sequence; the first convolutional layer is connected to the adder; the output features of the first dilated convolutional layer and the second dilated convolutional layer are concatenated and then connected to the adder through the second convolutional layer. Among them, the kernel sizes of the first and second convolutional layers are 1×1×c2 and 1×1×c respectively, and the strides are both 1. c is the channel dimension value of the features input to the input layer. The kernel sizes of the first to third dilated convolutional layers are all 3×3×3, the strides are all 1, and the dilation rates are set to 1, 3, and 5 respectively. Using skip connections can improve the feature transmission efficiency between dilated convolutional layers; removing specific skip connections in dense connections can effectively prevent holes in the receptive field; the adder fuses the features of different receptive fields, so that only one fused feature needs to be processed during the detection process. Therefore, the skip connection context module solves the problems in the prior art that a parallel multi-branch structure is adopted, and each branch is composed of a specific number of dilated convolutional layers connected in series in sequence, resulting in high computational complexity and limited feature transfer efficiency during the detection process. The present invention uses the self-attention mechanism to process the input features to obtain the mutual relationship between channels, uses average pooling in the channel dimension to process the mutual relationship between channels to obtain the channel attention score vector, and uses the channel attention score to weight the input features in the channel dimension. Based on this, the constructed channel attention module avoids the influence of irrelevant information in the spatial dimension on small pedestrian targets, improves the feature representation ability of small pedestrian targets, and solves the problem in the prior art that average pooling in the spatial dimension is used to calculate channel attention, so that small pedestrian targets are easily covered by surrounding irrelevant background information, resulting in insufficient attention of the detection network to small-size pedestrian targets and low detection accuracy.

[0007] To achieve the above object, the specific implementation steps of the present invention are as follows:

[0008] Step 1, construct a skip connection context module:

[0009] Build a skip connection context module including an input layer, a first convolutional layer, a first dilated convolutional layer, a second dilated convolutional layer, a third dilated convolutional layer, a splicing unit, a second convolutional layer, an adder, a third convolutional layer, and an output layer; where: the input layer, the first convolutional layer, the first dilated convolutional layer, the second dilated convolutional layer, the third dilated convolutional layer, the adder, the third convolutional layer, and the output layer are connected in series in sequence; the first convolutional layer is connected to the adder; the output features of the first dilated convolutional layer and the second dilated convolutional layer are concatenated and then connected to the adder through the second convolutional layer; set the parameters in the skip connection context module;

[0010] Step 2, construct a channel attention module:

[0011] Build a channel attention module including an input layer, a convolutional layer, a first dimension recombination layer, a second dimension recombination layer, a transpose unit, a multiplication unit, a global pooling layer, a weighting unit, an adder, and an output layer; wherein: the input layer, the convolutional layer, the first dimension recombination layer, the transpose unit, the multiplication unit, the global pooling layer, the weighting unit, the adder, and the output layer are connected in series in sequence; the input layer is respectively connected to the weighting unit and the adder; the second dimension recombination layer is bridged between the convolutional layer and the multiplication unit; set the parameters in the channel attention module;

[0012] Step 3, construct a weak pedestrian target detection network:

[0013] Build a weak pedestrian target detection network including an input layer, a convolutional group, a jump connection context module, a selective search network, a first pooling layer, an adder, a detection head, a second pooling layer, and a channel attention module; wherein: the input layer, the convolutional group, the jump connection context module, the selective search network, the first pooling layer, the adder, and the detection head are connected in series in sequence; the jump connection context module is connected to the first pooling layer; the convolutional group and the selective search network are respectively connected to the second pooling layer and then connected to the adder through the channel attention module; set the parameters in the weak pedestrian target detection network;

[0014] Step 4, generate a pre-training set and a training set:

[0015] Select at least 3000 pedestrian target images to form a pre-training set. Among them, the sizes of the pedestrian targets in each pedestrian target image are different, and the height is greater than or equal to 20 pixels. Each image corresponds to a true bounding box label of a pedestrian target; select at least 200 weak pedestrian target images to form a training set. Among them, the height of the weak pedestrian target in each weak pedestrian target image is less than 75 pixels. Each image corresponds to a true bounding box label of a weak pedestrian target;

[0016] Step 5, train the weak pedestrian target detection network:

[0017] Use the pre-training set to pre-train the weak pedestrian target detection network to obtain a pre-trained weak pedestrian target detection network; adopt the same method as pre-training, and use the training set to fine-tune the network parameters of the pre-trained weak pedestrian target detection network to obtain a trained weak pedestrian target detection network;

[0018] Step 6, detect weak pedestrian targets:

[0019] Input the weak pedestrian target image to be detected into the trained weak pedestrian target detection network, and use the predicted bounding box label output by the weak pedestrian target detection network as the detection result.

[0020] Compared with the prior art, the present invention has the following advantages:

[0021] First, the present invention constructs a skip connection context module, which overcomes the defects of the prior art that uses multiple parallel branches, and each branch is composed of a specific number of dilated convolutional layers connected in series in sequence, resulting in high computational complexity and limited feature transfer efficiency during the detection process. The present invention improves the transfer efficiency between dilated convolutional layers, only needs to process a single output feature during the detection process, reduces the computational complexity, and improves the detection efficiency of small and weak pedestrian targets.

[0022] Second, the present invention constructs a channel attention module, which overcomes the problem of the prior art that calculates channel attention using average pooling in the spatial dimension, making small and weak pedestrian targets easily covered by surrounding irrelevant background information, resulting in insufficient attention of the detection network to small-sized pedestrian targets and low detection accuracy. The present invention can calculate channel attention using average pooling in the channel dimension, avoid the influence of irrelevant information in the spatial dimension on small and weak pedestrian targets, improve the feature representation ability of small and weak pedestrian targets, and improve the detection accuracy of small and weak pedestrian targets. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 is a flowchart of the present invention;

[0024] Figure 2 is a schematic structural diagram of the skip connection context module of the present invention;

[0025] Figure 3 is a schematic structural diagram of the channel attention module of the present invention;

[0026] Figure 4 is a schematic structural diagram of the small and weak pedestrian target detection network of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0027] The following further describes the specific steps of the present invention in conjunction with the attached Figure 1 drawings.

[0028] Step 1, construct a skip connection context module.

[0029] Refer to Figure 2 for a further description of the structure of the skip connection context module of the present invention.

[0030] Step 1.1, build a skip connection context module, which includes an input layer, a first convolutional layer, a first dilated convolutional layer, a second dilated convolutional layer, a third dilated convolutional layer, a splicing unit, a second convolutional layer, an adder, a third convolutional layer, and an output layer; where: the input layer, the first convolutional layer, the first dilated convolutional layer, the second dilated convolutional layer, the third dilated convolutional layer, the adder, the third convolutional layer, and the output layer are connected in series in sequence; the first convolutional layer is connected to the adder; the output features of the first dilated convolutional layer and the second dilated convolutional layer are spliced and then connected to the adder through the second convolutional layer.

[0031] Step 1.2, set the convolutional kernel sizes of the first to third convolutional layers to 1×1×c2, 1×1×c2, 1×1×c respectively, and the stride is set to 1 for all, where c is the channel dimension of the input layer; set the convolutional kernel sizes of the first to third dilated convolutional layers to 3×3×3, the stride is set to 1 for all, and the dilation rates are set to 1, 3, 5 respectively.

[0032] Use the first convolutional layer with a convolutional kernel size of 1×1×c2 to reduce the channel dimension of the input layer by half to reduce parameters. For the detection of small and weak pedestrian targets, increasing the size of the receptive field and reducing the loss of effective information are both very necessary for the localization and detection of small and weak pedestrian targets. Dilated convolution can increase the receptive field and capture context information at the same resolution without introducing additional parameters. In the skip connection context module constructed in the embodiment of the present invention, stacked first dilated convolutional layer, second dilated convolutional layer, and third dilated convolutional layer are introduced to increase the context information of small and weak pedestrian targets and enhance the network's feature representation ability for small and weak pedestrian targets. The first dilated convolutional layer, the second dilated convolutional layer, and the third dilated convolutional layer output features with different receptive field sizes. The first dilated convolutional layer outputs features with a receptive field size of 3×3, the second dilated convolutional layer outputs features with a receptive field size of 9×9, and the third dilated convolutional layer outputs features with a receptive field size of 19×19. Use the adder to fuse the features of the above three sizes of receptive fields respectively to generate multi-scale fusion features. Use the residual connection method to connect the first convolutional layer and the adder to prevent gradient disappearance. Use the second convolutional layer with a convolutional kernel size of 1×1×c to restore the channel dimension of the features processed by the adder to c.

[0033] Step 2, build a channel attention module.

[0034] Refer to Figure 3 , and further describe the structure of the channel attention module constructed by the present invention.

[0035] Step 2.1, construct a channel attention module including an input layer, a convolutional layer, a first dimension reorganization layer, a second dimension reorganization layer, a transpose unit, a multiplication unit, a global pooling layer, a weighting unit, an adder, and an output layer; where: the input layer, the convolutional layer, the first dimension reorganization layer, the transpose unit, the multiplication unit, the global pooling layer, the weighting unit, the adder, and the output layer are connected in series in sequence; the input layer is connected to the weighting unit and the adder respectively; the second dimension reorganization layer is connected across the convolutional layer and the multiplication unit.

[0036] Step 2.2, set the convolutional kernel size of the convolutional layer to 1×1 and the stride to 1; both the first and second dimension reorganization layers are implemented by the reshape function; the transpose unit is implemented by matrix transpose operation; the global pooling layer is implemented by the global pooling function and the softmax function.

[0037] The background of the small and weak pedestrian target image is complex, and the small and weak pedestrian target is easily covered by the background or other small and weak pedestrian targets. The channels are related to the components of the target, that is, the channels can locate the components. The embodiment of the present invention constructs a channel attention module to prompt the network to focus on the visible part of the small and weak pedestrian target information.

[0038] Given the input features of the input layer First, use a convolutional layer with a convolutional kernel size of 1×1 for linear mapping to obtain Use the reshape function to perform dimension reorganization on with dimensions of W×H×C and X, so that and the dimensions of X change to WH×C. Use the self-attention mechanism to calculate the dimension-reorganized and X to obtain the relationship β between channels. Use the global pooling function to perform global pooling on β in the channel dimension to obtain the channel attention vector, and use the softmax function to normalize the channel attention vector to obtain the channel attention score α. Use α to perform weighting on X in the channel dimension, and perform an element-wise addition operation on the weighted X and X.

[0039] Step 3, construct a small and weak pedestrian target detection network.

[0040] Refer to Figure 4 for a further description of the structure of the small and weak pedestrian target detection network of the present invention.

[0041] Build a small and weak pedestrian target detection network including an input layer, a convolutional group, a skip connection context module, a selective search network, a first pooling layer, an adder, a detection head, a second pooling layer, and a channel attention module; where: the input layer, the convolutional group, the skip connection context module, the selective search network, the first pooling layer, the adder, and the detection head are connected in series in sequence; the skip connection context module is connected to the first pooling layer; the convolutional group and the selective search network are respectively connected to the second pooling layer and then connected to the adder through the channel attention module.

[0042] The detection head includes a first fully connected layer, a second fully connected layer, a third fully connected layer, a fourth fully connected layer, a softmax layer, a first output layer, and a second output layer; where: the first fully connected layer, the second fully connected layer, the third fully connected layer, the softmax layer, and the first output layer are connected in series in sequence; the fourth fully connected layer is respectively connected to the second fully connected layer and the second output layer. The softmax layer is implemented by the softmax function, and the number of neurons in the first to fourth fully connected layers is set to 4096, 4096, 1, and 4 respectively.

[0043] Set the parameters in the small and weak pedestrian target detection network as follows: both the first and second pooling layers are implemented by RoI Pooling; the selective search network is implemented by RPN. The convolutional group is composed of the series connection of conv1, conv2_x, conv3_x, and conv4_x in the existing ResNet-50 network.

[0044] The structural parameters of conv1, conv2_x, conv3_x, conv4_x, and conv5_x in the ResNet-50 are shown in Table 1. conv1, conv2_x, conv3_x, conv4_x, and conv5_x are respectively composed of convolutional layers with different convolution kernels and parameters.

[0045] Table 1 List of structural parameters of conv1, conv2_x, conv3_x, conv4_x, and conv5_x

[0046]

[0047] Taking conv2_x as an example to illustrate the symbols in Table 1 and their corresponding meanings, 1×1×1,64 represents a convolutional layer composed of 64 convolutional kernels with the same structure, and the size of each convolutional kernel is 1×1×1. Similarly, 3×3×3,64 represents a convolutional layer composed of 64 convolutional kernels with the same structure, and the size of each convolutional kernel is 3×3×3. 1×1×1,256 represents a convolutional layer composed of 256 convolutional kernels with the same structure, and the size of each convolutional kernel is 1×1×1. [] means that all the convolutional layers inside are connected in series to form a convolutional group, and x3 means that 3 convolutional groups with the same structure are connected in series to form conv2_x.

[0048] Step 4, generate the pre-training set and the training set.

[0049] Select at least 3000 pedestrian target images to form the pre-training set. Among them, the sizes of the pedestrian targets in each image are different, the height is greater than or equal to 20 pixels, and each image corresponds to a true bounding box label.

[0050] Select the training set of pedestrian target images that only contain pedestrian labels in the KITTI dataset as the pre-training set for the following reasons: 1) Small and weak pedestrian targets often appear in traffic scenarios, and the KITTI dataset contains real images including urban, rural, and highway scenes. 2) The pedestrian target images in the KITTI dataset contain multi-scale pedestrian targets such as large, medium, and small. The true bounding box label includes 4 values, namely the abscissa and ordinate of the upper left corner of the true bounding box and the width and height of the true bounding box.

[0051] Select at least 200 small and weak pedestrian target images to form the training set. Among them, the height of the small and weak pedestrian target in each image is less than 75 pixels, and each image corresponds to a true bounding box label.

[0052] There is no clear definition for small and weak pedestrian targets, and the present invention adopts a self-defined form. Existing pedestrian target datasets all contain subsets of small, medium, and large pedestrian targets. For example, the CityPersons dataset includes the Small subset, the Medium subset, and the Large subset, and the Caltechs dataset includes the Far subset, the Middle subset, and the Near subset.

[0053] In order not to increase the workload, the present invention uses the definition of the pedestrian target size in the Small subset of the CityPersons dataset, that is, the height of the pedestrian target is less than 75 pixels, as the definition of the small and weak pedestrian target. The images of the training set are different from those of the training set. Specifically, 100 images are selected from the Small subset of the CityPersons dataset and 100 images are selected from the Far subset of the Caltechs dataset, and the 200 selected images are used as the training set.

[0054] Step 5, training the weak pedestrian target detection network:

[0055] Pre-train the weak pedestrian target detection network using the pre-training set to obtain a pre-trained weak pedestrian target detection network; adopt the same method as pre-training, and use the training set to fine-tune the network parameters of the pre-trained weak pedestrian target detection network to obtain a trained weak pedestrian target detection network.

[0056] The so-called pre-training means calculating the loss value between the predicted class label and the true class label of the output of the weak pedestrian target detection network using the cross-entropy loss function and the regression loss function respectively, and the regression value between the predicted bounding box label and the true bounding box label of the output of the weak pedestrian target detection network. Using the mini-batch stochastic gradient descent algorithm, iterate and update the network parameters of the weak pedestrian target detection network until the cross-entropy loss function and the regression loss function converge, and the pre-training is completed.

[0057] The cross-entropy loss function is as follows:

[0058]

[0059] where L C represents the loss value of the cross-entropy loss function, ∑ represents the summation operation, i represents the i-th predicted bounding box label, represents the true class label of the i-th predicted bounding box label, log(·) represents the natural logarithm operation with base e, and p i represents the predicted class label of the i-th predicted bounding box label.

[0060] Similar to the definition of the true bounding box label, the predicted bounding box label includes 4 values, namely the abscissa and ordinate of the upper left corner of the predicted bounding box and the width and height of the predicted bounding box. The value-taking rule is as follows: Calculate the intersection over union (IoU) between the predicted bounding box label and the true bounding box label. When the IoU is greater than 0.5, the predicted bounding box label is considered a correct detection result. When the IoU is less than or equal to 0.5, it is considered an incorrect detection result. When the predicted bounding box label is considered a correct detection result, the value is 1; otherwise, the value is 0.

[0061] The regression loss function is as follows:

[0062]

[0063] where L regrepresents the loss value of the regression loss function, ∑ represents the summation operation, i represents the i-th predicted bounding box label, ||·|| represents the SmoothL1 loss function, and b i represents the regression value of the i-th predicted bounding box label, represents the regression value of the ground truth bounding box label corresponding to the i-th predicted bounding box label. The regression value of the ground truth bounding box label is 0, and the regression value of the predicted bounding box label is 1 - IoU.

[0064] The SmoothL1 loss function is as follows:

[0065]

[0066] where L(x) represents the loss value of the SmoothL1 loss function, x represents the input value, and · represents the absolute value operation.

[0067] The parameter settings of the mini-batch stochastic gradient descent method are as follows: the batch size is set to 32, the learning rate is set to 0.01, the learning rate is reduced by 0.001 every 100 iterations, and the total number of iterations is 500.

[0068] Step 6, detect small and weak pedestrian targets.

[0069] Input the image of the small and weak pedestrian target to be detected into the trained small and weak pedestrian target detection network, and use the predicted bounding box label output by the small and weak pedestrian target detection network as the detection result.

[0070] Select the images of small and weak pedestrian targets in the test set of the Small subset in the CityPersons dataset and the test set of the Far subset in the Caltechs dataset as the images to be detected.

[0071] The following further illustrates the effect of the present invention in combination with simulation experiments.

[0072] 1. Simulation experiment conditions:

[0073] The hardware platform for the simulation experiment of the present invention is: the processor is Intel(R) Xeon(R) CPU E5-2640 v3, the main frequency is 2.60 GHz, and the memory is 128 GB.

[0074] The software platform for the simulation experiment of the present invention is: the pytorch deep learning framework and the Ubuntu18.04 operating system.

[0075] In the simulation experiment of the present invention, the KITTI dataset is used to pre-train the small and weak pedestrian target detection network. The KITTI dataset is a challenging dataset captured by an autonomous driving platform. This dataset contains real images including urban, rural, and highway scenes. According to different degrees of occlusion and truncation, the images are divided into three difficulties: easy, medium, and hard. The KITTI training set can be split into a training set and a validation set. There are 7,481 training images and 7,518 test images. Each image has at most 15 cars and 30 pedestrians.

[0076] In the simulation experiment of the present invention, the CityPersons dataset is used to train and test the small and weak pedestrian target detection network. The CityPersons dataset is constructed based on the Cityscape dataset. The Cityscape dataset records the landscapes of many regions in Europe, making the CityPersons dataset have rich scenes. In addition, the pedestrian density in this dataset is relatively high and it contains various pedestrian target states, such as various occlusions and different target sizes. There are 2,975 images for training, and 500 and 1,575 images for validation and testing respectively.

[0077] In the simulation experiment of the present invention, the Caltech dataset is used to train and test the small and weak pedestrian target detection network. The Caltech dataset is currently the most popular and largest-scale pedestrian dataset. The data consists of approximately 10 hours of videos captured from traffic records in an urban environment. This dataset has a total of 11 video groups. The first 6 groups are used for training, including 42,782 images, and the last 5 groups are used for testing, including 4,024 images. The occlusion degree of the images can be divided into reasonable occlusion, heavy occlusion, partial occlusion, and no occlusion; the images can be divided into three scales: near, medium, and far.

[0078] 2. Simulation content and its result analysis:

[0079] In the simulation experiment of the present invention, the small and weak pedestrian target images in the test set of the Small subset of the CityPersons dataset and the test set of the Far subset of the Caltechs dataset are detected respectively using the pedestrian target detection method of the present invention and six existing technologies (MS-CNN, FRCNN+seg, TLL+MRF, MagnifierNet, AP2M, MSAF-Net). For each small and weak pedestrian target image, after being processed by each of the above methods, a predicted bounding box label that is considered to be the correct detection result can be obtained. These predicted bounding box labels are the detection results of the small and weak pedestrian target image.

[0080] The prior art MS-CNN pedestrian object detection method refers to the pedestrian object detection method of the multi-scale deep convolutional network proposed by Cai Z et al. in "A Unified Multi-scale Deep Convolutional Neural Network for Fast Object Detection. European Conference on Computer Vision, 2016", abbreviated as MS-CNN.

[0081] The prior art FRCNN+seg pedestrian object detection method refers to the CityPersons pedestrian object detection method proposed by Zhang S et al. in "CityPersons: A Diverse Dataset for Pedestrian Detection. IEEE, 2017", abbreviated as FRCNN+seg.

[0082] The prior art TLL+MRF pedestrian object detection method refers to the small-scale pedestrian object detection method of somatic topology localization and temporal feature aggregation proposed by Song T et al. in "Small-scale pedestrian detection based on somatic topology localization and temporal feature aggregation, 2018", abbreviated as TLL+MRF.

[0083] The prior art MagnifierNet pedestrian object detection method refers to the small-scale pedestrian object detection method based on multiple dense regions proposed by Cheng Q et al. in "MagnififierNet: Learning efficient small-scale pedestrian detector towards multiple dense Regions, 2021", abbreviated as MagnifierNet.

[0084] The prior art AP2M pedestrian object detection method refers to the robust pedestrian detection method for adaptive pattern-parameter matching proposed by Liu M et al. in "Adaptive pattern-parameter matching for robust pedestrian detection. Proceedings of the AAAI Conference on Artificial Intelligence, 2021", abbreviated as AP2M.

[0085] The prior art MSAF-Net pedestrian target detection method refers to the pedestrian target detection method based on multi-scale feature extraction and attention feature fusion proposed by Xia H et al. in "Pedestrian detection algorithm based on multi-scale feature extraction and attention feature fusion. Digital Signal Processing, 2022", abbreviated as MSAF-Net.

[0086] To verify the simulation effect of the present invention, the following false positive average miss rate (log-average miss rate, MR -2 ) is used to verify the method adopted in the simulation experiment of the present invention and six different pedestrian target detection methods. After detecting all small and weak pedestrian target images in the test set of the Small subset in the CityPersons dataset and the test set of the Far subset in the Caltechs dataset, the MR -2 in each test set is calculated. All the calculation results are plotted in Table 2, and Our in Table 2 represents the simulation experiment results of the present invention.

[0087] MR -2 is the average value of the miss rate (Miss Rate, MR) at 9 false positives per image (FPPIs) values (evenly spaced in logarithmic space in the value range [0.01, 1.0]).

[0088]

[0089] MR = 1 - recall

[0090]

[0091] where, N FP is the number of false positive FP (False Positive) detection results, N is the total number of detected images, recall represents the recall rate, N TP is the number of true positive TP (True Positive) detection results, N FN is the number of false negative FN (False Negative) detection results.

[0092] The prediction result of TP needs to meet the following three conditions simultaneously:

[0093] Condition 1, the confidence score is greater than the confidence threshold;

[0094] Condition 2: The predicted category is consistent with the true category label;

[0095] Condition 3: The IoU between the predicted bounding box and the true bounding box is greater than or equal to a predetermined threshold.

[0096] The situation that satisfies Condition 1 but does not satisfy Condition 2 and Condition 3 is FP. When there are multiple predicted bounding box labels corresponding to the same true bounding box label, only the prediction result with the highest confidence score is considered TP, and the rest are considered FP.

[0097] FN corresponds to the situation where there is no corresponding predicted bounding box label for the true bounding box label, that is, the samples that should have been detected are not detected. This indicator reflects the missed detection rate, and the smaller this indicator is, the better.

[0098] Table 2 Detection performance list of the present invention and six prior art methods

[0099]

[0100] Combined with Table 2, it can be seen that the MR of the present invention -2 is lower than that of the prior art in both the Small subset of the CityPersons dataset and the Far subset of the Caltechs dataset -2 , which proves that the present invention can improve the detection accuracy of small and weak pedestrian targets.

Claims

1. A method for detecting small and weak pedestrians based on skip connection context and channel attention, characterized in that, Based on the construction of the skip connection context module and the channel attention module respectively, a small and weak pedestrian target detection network is constructed; the specific steps of this detection method are as follows: Step 1, construct the skip connection context module: Build a skip connection context module including an input layer, a first convolutional layer, a first dilated convolutional layer, a second dilated convolutional layer, a third dilated convolutional layer, a splicing unit, a second convolutional layer, an adder, a third convolutional layer, and an output layer; among them: the input layer, the first convolutional layer, the first dilated convolutional layer, the second dilated convolutional layer, the third dilated convolutional layer, the adder, the third convolutional layer, and the output layer are connected in series in sequence; the first convolutional layer is connected to the adder; the output features of the first dilated convolutional layer and the second dilated convolutional layer are spliced and then connected to the adder through the second convolutional layer; set the parameters in the skip connection context module; Step 2, construct the channel attention module: Build a channel attention module including an input layer, a convolutional layer, a first dimension rearrangement layer, a second dimension rearrangement layer, a transpose unit, a multiplication unit, a global pooling layer, a weighting unit, an adder, and an output layer; among them: the input layer, the convolutional layer, the first dimension rearrangement layer, the transpose unit, the multiplication unit, the global pooling layer, the weighting unit, the adder, and the output layer are connected in series in sequence; the input layer is connected to the weighting unit and the adder respectively; the second dimension rearrangement layer is bridged between the convolutional layer and the multiplication unit; set the parameters in the channel attention module; Step 3, construct the small and weak pedestrian target detection network: Build a small and weak pedestrian target detection network including an input layer, a convolutional group, a skip connection context module, a selective search network, a first pooling layer, an adder, a detection head, a second pooling layer, and a channel attention module; among them: the input layer, the convolutional group, the skip connection context module, the selective search network, the first pooling layer, the adder, and the detection head are connected in series in sequence; the skip connection context module is connected to the first pooling layer; the convolutional group and the selective search network are respectively connected to the second pooling layer and then connected to the adder through the channel attention module; set the parameters in the small and weak pedestrian target detection network; Step 4, generate the pre-training set and the training set: Select at least 3000 pedestrian target images to form the pre-training set. Among them, the sizes of the pedestrian targets in each pedestrian target image are different, and the height is greater than or equal to 20 pixels. Each image corresponds to a true bounding box label of a pedestrian target; select at least 200 small and weak pedestrian target images to form the training set. Among them, the height of the small and weak pedestrian targets in each small and weak pedestrian target image is less than 75 pixels. Each image corresponds to a true bounding box label of a small and weak pedestrian target; Step 5, train the small and weak pedestrian target detection network: Use the pre-training set to pre-train the small and weak pedestrian target detection network to obtain a pre-trained small and weak pedestrian target detection network; adopt the same method as pre-training, and use the training set to fine-tune the network parameters of the pre-trained small and weak pedestrian target detection network to obtain a trained small and weak pedestrian target detection network; Step 6, detect the small and weak pedestrian targets: Input the image of a small and weak pedestrian target to be detected into the trained small and weak pedestrian target detection network, and use the predicted bounding box labels output by the small and weak pedestrian target detection network as the detection results.

2. The weak and small pedestrian detection method based on skip connection context and channel attention according to claim 1, wherein The parameters in the jump connection context module set in step 1 are as follows: Set the convolutional kernel sizes of the first and second convolutional layers to 1×1×c / 2 and 1×1×c respectively, and the strides are both set to 1, where c is the channel dimension value of the features input to the input layer; set the convolutional kernel sizes of the first to third dilated convolutional layers to 3×3×3, the strides are both set to 1, and the dilation rates are set to 1, 3, and 5 respectively.

3. The weak and small pedestrian detection method based on skip connection context and channel attention according to claim 1, wherein The parameters in the channel attention module set in step 2 are as follows: Set the convolutional kernel size of the convolutional layer to 1×1 and the stride to 1; both the first and second dimension reorganization layers are implemented by the reshape function; the transpose unit is implemented by matrix transpose operation; the global pooling layer is implemented by the global pooling function and the softmax function.

4. The weak and small pedestrian detection method based on skip connection context and channel attention according to claim 1, characterized in that The detection head in step 3 includes a first fully connected layer, a second fully connected layer, a third fully connected layer, a fourth fully connected layer, a softmax layer, a first output layer, and a second output layer; among them: the first fully connected layer, the second fully connected layer, the third fully connected layer, the softmax layer, and the first output layer are connected in series in sequence; the fourth fully connected layer is connected to the second fully connected layer and the second output layer respectively.

5. The weak pedestrian detection method based on skip connection context and channel attention according to claim 1, wherein The parameters in the small and weak pedestrian target detection network set in step 3 are as follows: Both the first and second pooling layers are implemented by RoI Pooling; the selective search network is implemented by RPN.

6. The method for detecting small and weak pedestrians based on skip connection context and channel attention according to claim 1, wherein The pre-training in step 5 refers to calculating the loss value between the predicted class label and the true class label of the output of the small and weak pedestrian target detection network, and the regression value of the predicted bounding box label and the regression value of the true bounding box label of the output of the small and weak pedestrian target detection network respectively using the cross-entropy loss function and the regression loss function, and using the mini-batch stochastic gradient descent algorithm to iteratively update the network parameters of the small and weak pedestrian target detection network until the cross-entropy loss function and the regression loss function converge, thus completing the pre-training.

7. The weak pedestrian detection method based on skip connection context and channel attention according to claim 6, wherein, The cross-entropy loss function is as follows: Among them, L C represents the loss value of the cross-entropy loss function, ∑ represents the summation operation, i represents the i-th predicted bounding box label, represents the true class label of the i-th predicted bounding box label, log(·) represents the natural logarithm operation, p i represents the predicted class label of the i-th predicted bounding box label.

8. The weak and small pedestrian detection method based on skip connection context and channel attention according to claim 6, characterized in that The regression loss function is as follows: Among them, L reg represents the loss value of the regression loss function, ∑ represents the summation operation, i represents the i-th predicted bounding box label, ||·|| represents the Smooth L1 loss function, and b i represents the regression value of the i-th predicted bounding box label, represents the regression value of the ground truth bounding box label corresponding to the i-th predicted bounding box label.

9. The method for detecting weak and small pedestrians based on skip connection context and channel attention according to claim 8, wherein, The Smooth L1 loss function is as follows: Among them, L(x) represents the loss value of the Smooth L1 loss function, x represents the input value, and |·| represents the absolute value operation.

Citation Information

Patent Citations

  • Lightweight YOLOv4 pedestrian detection method under view angle of unmanned aerial vehicle

    CN115359376A

  • Small target detection method for context feature fusion screening based on attention mechanism

    CN111275688A

  • Pedestrian detection method based on convolutional neural network and double attention mechanism

    CN111680619A