Pedestrian detection method based on lightweight YOLO v5 network model and spatiotemporal memory mechanism
By applying a lightweight YOLO v5 network model and space-time memory mechanism in the monitoring of university wall boundaries, the problems of high error recognition rate and hardware overhead in monitoring of university wall boundaries are solved, and more efficient pedestrian detection and campus safety guarantees are achieved.
Patent Information
- Application Number
- CN202210756317.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-29
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2042-06-29
AI Technical Summary
The monitoring of university wall boundaries has problems such as high misidentification rate, large hardware overhead and prone to personnel errors, which makes it difficult to effectively ensure campus safety.
Pedestrian detection method based on lightweight YOLO v5 network model and space-time memory mechanism is adopted. By building a training database, building a lightweight YOLO v5 network model and combining a space-time memory mechanism, pedestrian position detection is reduced and misidentified and robustness is improved.
The error recognition rate of pedestrian detection is reduced from 7% to 1%, and the processing speed is improved from 56FPS to 74FPS, suitable for multi-channel video processing and model deployment.
Smart Images

Figure CN115116137B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of deep learning target detection, and specifically provides a pedestrian detection method based on a lightweight YOLO v5 network model and a spatiotemporal memory mechanism. Background Art
[0002] The tight staffing during the epidemic and the possible deviations in human visual recognition make it easy for people to slack off and make negligence and mistakes.
[0003] The following problems exist in the closed management model of colleges and universities, where a monitoring room is used to maintain campus security around the clock: (1) There are too many monitoring screens: The boundaries of most campus walls are several kilometers long, with dozens of cameras and monitoring screens. (2) Staff members are prone to making mistakes: Violations of discipline on campus walls are low-probability events, and people tend to be optimistic that low-probability events will not occur. Staff members are more likely to be lazy, negligent, and have visual fatigue, resulting in missed judgments. (3) High costs for the school: The security staff in the monitoring room have low salaries and cannot attract employees. However, the cost is too high for the school because the eight-hour work system requires hiring three times as many employees and three times as much expenditure. (4) Existing detection methods have a high false recognition rate and high hardware costs.
[0004] Therefore, it is urgent to carry out intelligent monitoring of the boundaries of university walls. In this case, developing a lightweight method and intelligent monitoring method to reduce pedestrian misidentification and improve robustness is conducive to the intelligent processing and rapid response of campus epidemic prevention and control, and better maintain campus safety. Summary of the invention
[0005] In view of the shortcomings of the prior art, the present invention provides a pedestrian detection method based on a lightweight YOLO v5 network model and a spatiotemporal memory mechanism. The detection method is based on the YOLO v5 deep network model. On the basis of the original model, a spatiotemporal memory mechanism is designed and a lightweight YOLOv5 model is constructed, thereby reducing the misidentification of pedestrian detection and improving the robustness of the system. The model is lightweighted by using channel shuffling and pruning model methods to speed up the processing speed of the model, which is more suitable for the processing of multi-channel videos and the deployment of the model.
[0006] The technical solution of the present invention to solve the technical problem is: designing a pedestrian detection method based on a lightweight YOLO v5 network model and a spatiotemporal memory mechanism, the method comprising the following steps:
[0007] Step 1: Build a training database
[0008] 1) Collect images of different scenes at the monitoring location, including sunny days, cloudy days, rainy days and nights; divide the human targets in the collected images into different sizes according to the distance between the pedestrians and the camera, among which the targets with an area less than 32*32 pixels are small-sized targets, those between 32*32 pixels and 96*96 pixels are medium-sized targets, and those larger than 96*96 pixels are large-sized targets; select images according to the number ratio of the above three types of human target sizes of 1:1:1, perform data enhancement on the selected images, and finally perform image size unification operation to obtain the original data set;
[0009] 2) Dataset division: The original dataset obtained in step 1) is manually labeled, and the people and walls in the image are marked with rectangular boxes, and the images in the original dataset are randomly divided into training set and validation set according to a certain number ratio;
[0010] Step 2: Build a lightweight YOLO v5 network model
[0011] 1) Preprocessing of the training set: Perform data augmentation on the training set obtained in step 2) of the first step;
[0012] 2) Build a lightweight YOLO v5 network model
[0013] The lightweight YOLO v5 network model is an improved structure of the YOLO v5 network model, specifically, the Foucs module, CBL module, CSP1_1 module, CBL module, CSP1_3 module, CBL module, CBL module, CSP1_3 module, CBL module, CBL module, CBL module, SPP module, and CBL module in the backbone network of the YOLO v5 network model are replaced with 2 CBL modules, SFB1 modules, 2 SFB2 modules, SFB1 modules, 7 SFB2 modules, SFB1 modules, SFB2 modules, and CBS modules in sequence; the input of the backbone network part of the lightweight YOLO v5 network model is first input into the first CBL module, and the output of the second SFB2 module is respectively input into the SFB1 module connected thereto and the first CBL module of the Neck network part, and the output of the CBS module is input into the CSP2_1 module in the backbone network of the YOLO v5 network model. The other parts of the structure of the lightweight YOLO v5 network model are the same as those of the YOLO The v5 network model is the same;
[0014] 3) Training the network
[0015] The pre-trained weights obtained on ImageNet are used to initialize the backbone network, the convolutional layer parameters are initialized using the Kaiming normal distribution, and the rest of the network is initialized using Xavier. The learning rate is set to decrease in steps with the number of training generations, and the backbone network parameters are frozen in the first 50 generations.
[0016] The training set preprocessed in step 1) is input into the initialized lightweight YOLO v5 network model, and the backbone network is used to extract and fuse features. The classification and regression network is used to obtain the position, category and confidence of the human target, as well as the position and confidence of the wall, which are compared with the true label to obtain the Loss value; the SGD optimizer is used according to the Loss value to perform back propagation to update the network parameters until the Loss drops to the preset value, and the network model training is completed;
[0017] 4) Network model verification
[0018] The verification set obtained in step 2) of the first step is input into the network model trained in step 3), and the detection label output by the network model is compared with the true label to obtain the misrecognition rate. When the misrecognition rate is not greater than 10%, the current parameters of the network model are saved, and the network model is a valid model; when the misrecognition rate is greater than 10%, the initial parameters of the network model are adjusted, and the network is retrained until the loss drops to the preset value, and the misrecognition rate of the verification set is not greater than 10%, the current parameters of the network model are saved, and the current network model is a valid model, and the network model verification is completed;
[0019] Step 3: Use the lightweight YOLO v5 network model and spatiotemporal memory mechanism module to detect pedestrian positions
[0020] 1) Obtain preliminary test results
[0021] Input the video stream captured by the camera frame by frame into the lightweight YOLO v5 network model verified in the second step to obtain the detection results of the frame sequence images of the video. The detection result of each image includes the position, category and confidence of the human target, as well as the position and confidence of the wall. This detection result is the preliminary detection result.
[0022] 2) Obtain the corrected test results
[0023] The preliminary detection results are input into the spatiotemporal memory mechanism module. The principle of the spatiotemporal memory mechanism is as follows:
[0024]
[0025] Where P n+1represents the confidence of the human target in the n+1th frame of the video sequence image; Δx and Δy represent the change values of the x-axis and y-axis of the closest human target between the n+1th frame and the i-th frame, respectively, and the value range is 0 to +∞; P i Represents the confidence of the person target in the i-th frame image output by the network model; Indicates rounding up;
[0026] In the above formula, f(x) is expressed as follows:
[0027]
[0028] The confidence of the predicted person target in the next frame output by the spatiotemporal memory mechanism module is replaced with the confidence in the preliminary detection result of the corresponding frame to obtain a revised detection result;
[0029] 3) Pedestrian position detection
[0030] According to the detected human targets and walls in the corrected detection results in step 2), when the trajectory of the human target intersects with the set wall warning line or is lower than a certain threshold, it can be determined as a wall climbing behavior or illegal wall-taking behavior; the upper left corner of the video screen is used as the coordinate origin, and the positive directions of the x-axis and y-axis are set to the right and downward respectively to establish a two-dimensional coordinate system. The position of the wall is given manually, and the wall warning line is approximately a straight line; x i ,y i Indicates the coordinates of the detected human target in the same coordinate system; the principle is as follows:
[0031] f(x,y)=Ax+By+C=0
[0032] Indicates the location of the fence cordon. If
[0033]
[0034] If
[0035] |f(x i ,y i )|<t
[0036] It indicates that there is suspicion of taking something through the wall; i ,y i It represents the position coordinates of the human target detected in the i-th frame image, t represents the set threshold, and A, B, and C are constant parameters calculated when the wall position is manually given.
[0037] Compared with the prior art, the beneficial effect of the present invention is that the pedestrian detection method of the present invention adopts a lightweight YOLOv5 network, replaces the Foucs layer of the original YOLO v5 model with a convolution layer, and replaces the convolution in the backbone network with a grouped convolution with random channel mixing to lightweight the model, and uses a spatiotemporal memory mechanism to correct the network detection results to reduce misidentification, and uses the corrected detection results to detect pedestrian positions, thereby reducing the misidentification rate of the model, reducing the hardware overhead of the model, and improving the processing speed. The test results of the pedestrian detection method of the present invention in the data set are: the misidentification rate is reduced from 7% to 1%, and the processing speed is increased from 56FPS to 74FPS. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings required in the embodiments or the prior art descriptions are briefly introduced below. The following drawings are for illustrating selected implementation cases rather than all possible implementation plans.
[0039] Figure 1 A lightweight YOLO v5 network model structure diagram of an embodiment of the pedestrian detection method of the present invention.
[0040] Figure 2 This is the YOLO v5 network model structure diagram.
[0041] Figure 3 for Figure 2 The Focus module structure diagram in .
[0042] Figure 4 for Figure 2 SPP module structure diagram in.
[0043] Figure 5 for Figure 2 CSP1_x module structure diagram in.
[0044] Figure 6 for Figure 2 CSP2_x module structure diagram in.
[0045] Figure 7 This is the structure diagram of the CBL module in the YOLO v5 network model.
[0046] Figure 8 This is the Res unit module structure diagram in the CSP1_x module.
[0047] Fig. 9 for Figure 1 SFB1 module structure diagram in.
[0048] Fig.10 for Figure 1SFB2 module structure diagram in.
[0049] Fig.11 for Figure 2 Schematic diagram of the Focus module in .
[0050] Fig.12 for Figure 1 Schematic diagram of the CBL*2 module in Figure 1.
[0051] Fig.13 This is a schematic diagram of the spatiotemporal memory mechanism of an embodiment of the pedestrian detection method of the present invention.
[0052] Fig.14 for
[0053]
[0054] Function graph.
[0055] Fig.15 A set of real-time video images of the monitoring room in the embodiment.
[0056] Fig.16 In order to use the pedestrian detection method of the present invention Fig.15 The test results. DETAILED DESCRIPTION
[0057] The following is a concise and clear description of the technical solutions in the embodiments of the present invention, and an explanation of the drawings in the embodiments of the invention. The described embodiments are only some embodiments of the invention, not all. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0058] The present invention provides a pedestrian detection method based on a lightweight YOLO v5 network model and a spatiotemporal memory mechanism (hereinafter referred to as a pedestrian detection method), which comprises the following steps:
[0059] Step 1: Build a training database
[0060] 1) Collect images of different scenes at the monitoring location, including sunny days, cloudy days, rainy days and nights. According to the distance between the pedestrians and the camera, the human targets in the collected images are divided into different sizes, among which the area less than 32*32 pixels is a small-sized target, the area between 32*32 pixels and 96*96 pixels is a medium-sized target, and the area larger than 96*96 pixels is a large-sized target. Select images with a quantity ratio of 1:1:1 for the above three types of human target sizes, perform data enhancement on the selected images, and finally perform image size unification to obtain the original data set. Data enhancement can enhance the generalization ability of the model, and the enhancement methods include mirroring, cropping, translation and scaling.
[0061] The different camera installation positions and the distance from pedestrians to the camera will change as pedestrians move, so the proportion of target pedestrians in the surveillance video will be different. In actual situations, large, medium, and small-sized target objects may appear. In order to ensure the detection effect of the model in actual situations, the proportion of large, medium, and small-sized target objects in the data is set to 1:1:1 to ensure the balance of the number of samples of different scales.
[0062] 2) Dataset division: The original dataset obtained in step 1) is manually labeled, and the people and walls in the image are marked with rectangular boxes respectively. The images in the original dataset are randomly divided into training set and validation set in a quantity ratio of 7:3.
[0063] Step 2: Build a lightweight YOLO v5 network model
[0064] 1) Preprocessing of the training set: Perform data enhancement on the training set obtained in step 2) of the first step. The enhancement methods include sharpening, histogram equalization, color space change, adding different types of noise and normalization methods, and mosaic data enhancement. Make the TSNE distribution of the test set as close as possible to the TSNE distribution of the training set to improve the recognition accuracy of the model.
[0065] 2) Build a lightweight YOLO v5 network model
[0066] Lightweight YOLO v5 network model (see Figure 1 ) is the YOLO v5 network model (see Figure 2 ), specifically, the Foucs module, CBL module, CSP1_1 module, CBL module, CSP1_3 module, CBL module, CBL module, CSP1_3 module, CBL module, CBL module, CBL module, SPP module, and CBL module in the backbone network (Backbone) of the YOLO v5 network model are replaced with 2 CBL modules, SFB1 module, 2 SFB2 modules, SFB1 module, 7 SFB2 modules, SFB1 module, SFB2 module, and CBS module in sequence. The input of the backbone network part of the lightweight YOLO v5 network model is first input into the first CBL module, the output of the second SFB2 module is respectively input into the SFB1 module connected thereto and the first CBL module of the Neck network part, the output of the CBS module is input into the CSP2_1 module in the backbone network (Backbone) of the YOLO v5 network model, and the other parts of the structure of the lightweight YOLO v5 network model are the same as those of the YOLO v5 network model.
[0067] The backbone network components of YOLO v5 include: Foucs, CBL, SPP, CSP1_x and CSP2_x. CBL stands for convolutional layer, normalization layer and activation layer. The Foucs layer downsamples the input feature map, halving the width and height of the feature layer while keeping the number of channels unchanged. The input image undergoes four-way Slice operation to form four feature layers that are downsampled at different positions, and then the results are concat operated. After that, they are processed by a CBL module and the output of the Foucs layer is output. SPP divides the input feature map into three groups according to the number of channels and performs Maxpool respectively, then the results are concat operated, and then they are processed by a CBL module and the output of the SPP module is output. The working principle of the CSP1_x module: First, the x residual components (CBL module + Res unit module + CONV) connected in series are paralleled with a convolutional layer, and then the outputs of the two are concat operated. After that, they are processed by the BN layer and the Leaky relu activation function in sequence, and the output of the CSP1_x module is output. The working principle of the CSP2_x module is: first, connect x CBLC modules (CBL module + CONV) in series and a convolutional layer in parallel, then perform a Concat operation on the outputs of the two, and then process them through the BN layer and the Leaky relu activation function in turn, and then output the output of the CSP2_x module. The backbone network of the YOLO v5 model starts with Foucs.
[0068] The main components of the backbone network of the Shufflenet network are SFB1, SFB2, and CBS modules. The SFB1 module is composed of a CBL module, a DWB module, a Concat module, and a Channel Shuffle module, where a DWB module is connected in series with a CBL module, and x components consisting of a CBL module, a DWB module, and a CBL module are connected in series. The above two series parts are parallel structures and their outputs are input to the Concat module; the Concat module processes the input of the two parts, and its output is input to the Channel Shuffle module, and the output of the Channel Shuffle module is the output of the SFB1 module.
[0069] The SFB2 module consists of a Channel split module, a CBL module, a DWB module, a CBL module, a Concat module, and a Channel Shuffle module. The input of the SFB2 module is first input into the Channel split module, and the Channelsplit module converts the multi-channel array into an independent single-channel array. The output of the Channel split module is input into the Concat module after being processed by x components consisting of a CBL module, a DWB module, and a CBL module in series. The output of the Channel split module is directly input into the Concat module. The Concat module processes the input of the two parts, and its output is input into the Channel Shuffle module. The output of the Channel Shuffle module is the output of the SFB2 module.
[0070] The DWB module represents Depthwise Separable Convolution and BN layers. DWB is different from traditional convolution operations. The convolution part is divided into two steps: the first step is to select M convolution kernels to perform one-to-one convolution on the original channels without summing to generate M channels; the second step is to select N 1*1 convolution kernels to perform convolution operations on the feature layers of the M channels in the first step to obtain N results. The SFB1 module does not change the number of input and output channels and the size of the feature map. When the number of input and output channels is the same, the memory access amount MAC is the smallest. The SFB2 module is downsampling, which halves the width and height of the feature map and doubles the number of channels.
[0071] The Channel Shuffle module means randomly shuffling the channels formed by grouped convolution. Grouped convolution is to divide the feature layer into several groups, perform convolution separately, and then perform Concat operation. Using grouped convolution alone will cause the features of each channel to propagate forward in their own channels without intersecting each other, which will produce boundary effects. The resulting feature map is relatively limited and not suitable for the extraction and fusion of overall features. Channel Shuffle allows the channels to be combined and fused in a certain arrangement after channel grouping, and enter the next grouped convolution. In the backbone network of the Shufflenet network, SFB1 and SFB2 are arranged and combined alternately.
[0072] The lightweight YOLO v5 network model replaces the Focus layer of the YOLO v5 model with a convolutional layer and replaces the rest of its backbone network with the Shufflenet backbone network, while retaining the data preprocessing method of the YOLO v5 model.
[0073] According to the third criterion proposed in Shufflenetv2, too many fragmentation operations will affect the parallel computing speed of the hardware. The Foucs layer forms a four-way feature layer by performing four-way Slice operations on the input image. This approach will result in too many paths and reduce the hardware parallelism. Therefore, it is replaced with a common convolution layer CBL. The operation of common convolution depends largely on the memory size and processing speed of the GPU. Under the premise of ensuring the detection accuracy, a certain degree of accuracy can be appropriately lost to improve the detection speed.
[0074] The backbone network of the original YOLO v5 model is replaced with the Shufflenet backbone network. The Shufflenet backbone network is mostly grouped convolution of mixed channels. Using grouped convolution alone will cause the features of each channel to propagate forward in their own channels without intersecting with each other, which is not suitable for the extraction and fusion of overall features. After the channels are grouped, the channels are combined and fused in a certain arrangement and enter the next grouped convolution.
[0075] 3) Training the network
[0076] The pre-trained weights obtained on ImageNet are used to initialize the backbone network, the convolutional layer parameters are initialized using the kaiming normal distribution, and the rest of the network is initialized using Xavier. The learning rate is set to decrease in steps with the number of training generations, so that the network can seek the optimal solution in the early stage of training and have better convergence in the later stage of training. The backbone network parameters are frozen for the first 50 generations.
[0077] The training set preprocessed in step 1) is input into the initialized lightweight YOLO v5 network model, and the backbone network is used for feature extraction and fusion. The classification and regression network is used to obtain the position, category and confidence of the human target, as well as the position and confidence of the wall, and the Loss value is obtained by comparing it with the true label. According to the Loss value, the SGD (gradient descent) optimizer is used to back-propagate and update the network parameters until the Loss drops to the preset value, and the network model training is completed.
[0078] 4) Network model verification
[0079] The validation set obtained in step 2) of the first step is input into the network model trained in step 3), and the detection label output by the network model is compared with the real label to obtain the misrecognition rate. When the misrecognition rate is not greater than 10%, the current parameters of the network model are saved, and the network model is a valid model. When the misrecognition rate is greater than 10%, the initial parameters of the network model are adjusted, and the network is retrained until the Loss drops to the preset value, and the misrecognition rate of the validation set is not greater than 10%, the current parameters of the network model are saved, and the current network model is a valid model, and the network model verification is completed.
[0080] Step 3: Use the lightweight YOLO v5 network model and spatiotemporal memory mechanism module to detect pedestrian positions
[0081] 1) Obtain preliminary test results
[0082] The video stream captured by the camera is input frame by frame into the lightweight YOLO v5 network model verified in the second step to obtain the detection results of the frame sequence images of the video. The detection results of each image include the position, category and confidence of the human target, as well as the position and confidence of the wall. This detection result is a preliminary detection result.
[0083] 2) Obtain the corrected test results
[0084] The preliminary detection results are input into the spatiotemporal memory mechanism module. The principle of the spatiotemporal memory mechanism is as follows:
[0085]
[0086] Where P n+1 represents the confidence of the human target in the n+1th frame of the video sequence image; Δx and Δy represent the change values of the x-axis and y-axis of the closest human target between the n+1th frame and the i-th frame, respectively, and the value range is 0 to +∞; P i Represents the confidence of the person target in the i-th frame image output by the network model. Indicates rounding up.
[0087] In the above formula, the expression of f(x) is as follows:
[0088]
[0089] The implementation idea of this spatiotemporal memory mechanism is that the confidence of the person target in the next frame is determined by the spatial position changes of the person target in the previous 150 frames of images in the time series. This mechanism integrates the temporal dimension and spatial dimension of the person target to predict the confidence of the person target in the next frame of the image.
[0090] The confidence of the predicted person target in the next frame output by the spatiotemporal memory mechanism module replaces the confidence in the preliminary detection result of the corresponding frame to obtain a revised detection result.
[0091] 3) Pedestrian position detection
[0092] According to the detected human targets and walls (generally, it is considered that the confidence level is greater than 0.5 for the target to exist) in the corrected detection results in step 2), when the trajectory of the human target intersects with the set wall warning line or is lower than a certain threshold, it can be determined as a wall climbing behavior or illegal wall-taking behavior. With the upper left corner of the video screen as the coordinate origin, the positive directions of the x-axis and y-axis are set to the right and downward respectively to establish a two-dimensional coordinate system. The position of the wall is given manually, and the wall warning line is approximately a straight line. i ,y i Indicates the coordinates of the detected human target in the same coordinate system. The principle is as follows:
[0093] f(x,y)=Ax+By+C=0
[0094] Indicates the location of the fence cordon. If
[0095]
[0096] If
[0097] |f(x i ,y i )|<t
[0098] It indicates that there is suspicion of taking something through the wall. i ,y i It represents the position coordinates of the human target detected in the i-th frame image, t represents the set threshold, and A, B, and C are constant parameters calculated when the wall position is manually given.
[0099] Example 1
[0100] This embodiment provides a pedestrian detection method based on a lightweight YOLO v5 network model and a spatiotemporal memory mechanism. The method is used for abnormality recognition of campus security intelligent monitoring. The method includes the following steps:
[0101] Step 1: Build a training database
[0102] 1) Collect images of different scenes at the monitoring location, including sunny days, cloudy days, rainy days and nights. According to the distance between the pedestrians and the camera, the human targets in the collected images are divided into different sizes, among which the targets with an area less than 32*32 pixels are small-sized targets, those between 32*32 pixels and 96*96 pixels are medium-sized targets, and those larger than 96*96 pixels are large-sized targets. Select images with a quantity ratio of 1:1:1 for the above three types of human target sizes, and perform data enhancement on the selected images to obtain the original data set. Data enhancement can enhance the generalization ability of the model, and the enhancement methods include mirroring, cropping, translation and scaling.
[0103] The different camera installation positions and the distance from pedestrians to the camera will change as pedestrians move, so the proportion of target pedestrians in the surveillance video will be different. In actual situations, large, medium, and small-sized target objects may appear. In order to ensure the detection effect of the model in actual situations, the proportion of large, medium, and small-sized target objects in the data is set to 1:1:1 to ensure the balance of the number of samples of different scales.
[0104] There are few surveillance cameras in the old campus of a certain school, and the security department of the monitoring room has a total of 24 channels of video. Because the analysis route of the original video cannot be obtained, an external industrial camera is used to collect the monitoring screen. Two high-definition 5-megapixel XW500 USB industrial cameras and two 5-megapixel machine vision industrial lenses are used. The image acquisition device has a line-by-line scanning sensor, no compression, no difference compensation, a high-speed 2.0 interface, and a transmission speed of 480Mb / s, which can realize the function of docking with a PC and real-time display of images. Images of different scenes are collected to build a database. The scenes of these data collection include sunny days, rainy days, cloudy days and nights. And according to the weather conditions in reality, the weights of the number of samples in different scenes are processed accordingly. The target instance area is less than 32*32 pixels for small-size targets, between 32*32 pixels and 96*96 pixels for medium-size targets, and larger than 96*96 pixels for large-size targets. The proportion of target objects of different sizes in the image is weighed according to the actual scene. The different camera installation positions and the distance from pedestrians to the camera will change as pedestrians move, so the proportion of target pedestrians in the surveillance video will be different. In actual situations, large, medium, and small target objects may appear. In order to ensure the detection effect of the model in actual situations, the proportion of large, medium, and small target objects in the data is set to 1:1:1. Ensure that samples of different sizes are balanced.
[0105] Since the images collected in a short time may not cover all the actual conditions throughout the year, the collected images are data enhanced to expand the database and enhance the generalization ability of the model. Enhancement methods include mirroring, cropping, translation and scaling. The image size is unified to obtain images of uniform scale to form the original data set. This implementation case collected a total of 2130 sample images, and after data enhancement, there are a total of 3500 sample images, that is, the original data set contains 3500 images.
[0106] 2) Dataset division: The original dataset obtained in step 1) is manually labeled, and the people and walls in the image are marked with rectangular boxes respectively. The images in the original dataset are randomly divided into training set and validation set in a quantity ratio of 7:3.
[0107] Step 2: Build a lightweight YOLO v5 network model
[0108] 1) Preprocessing of the training set: Mosaic data enhancement is performed on the training set obtained in step 2) of the first step, so that the TSNE distribution of the test set is as close as possible to the TSNE distribution of the training set, thereby improving the recognition accuracy of the model.
[0109] 2) Build a lightweight YOLO v5 network model
[0110] The lightweight YOLO v5 network model is an improved structure of the YOLO v5 network model. Specifically, the Foucs module, CBL module, CSP1_1 module, CBL module, CSP1_3 module, CBL module, CBL module, CSP1_3 module, CBL module, CBL module, CBL module, SPP module, and CBL module in the backbone network (Backbone) of the YOLO v5 network model are replaced with 2 CBL modules, SFB1 module, 2 SFB2 modules, SFB1 module, 7 SFB2 modules, SFB1 module, SFB2 module, and CBS module. The input of the backbone network of the lightweight YOLO v5 network model is first input into the first CBL module. The output of the second SFB2 module is respectively input into the SFB1 module connected to it and the first CBL module of the Neck network part. The output of the CBS module is input into the CSP2_1 module in the backbone network of the YOLO v5 network model. The other parts of the structure of the lightweight YOLO v5 network model are the same as those of the YOLO v5 network model.
[0111] The backbone network components of YOLO v5 include: Foucs, CBL, SPP, CSP1_x and CSP2_x. CBL stands for convolutional layer, normalization layer and activation layer. The Foucs layer downsamples the input feature map, halving the width and height of the feature layer while keeping the number of channels unchanged. The input image undergoes four-way Slice operation to form four feature layers that are downsampled at different positions, and then the results are concat operated. After that, they are processed by a CBL module and the output of the Foucs layer is output. SPP divides the input feature map into three groups according to the number of channels and performs Maxpool respectively, then the results are concat operated, and then they are processed by a CBL module and the output of the SPP module is output. The working principle of the CSP1_x module: First, the x residual components (CBL module + Res unit module + CONV) connected in series are paralleled with a convolutional layer, and then the outputs of the two are concat operated. After that, they are processed by the BN layer and the Leaky relu activation function in sequence, and the output of the CSP1_x module is output. The working principle of the CSP2_x module is: first, connect x CBLC modules (CBL module + CONV) in series and a convolutional layer in parallel, then perform a Concat operation on the outputs of the two, and then process them through the BN layer and the Leaky relu activation function in turn, and then output the output of the CSP2_x module. The backbone network of the YOLO v5 model starts with Foucs.
[0112] The main components of the backbone network of the Shufflenet network are SFB1, SFB2, and CBS modules. The SFB1 module is composed of a CBL module, a DWB module, a Concat module, and a Channel Shuffle module, where a DWB module is connected in series with a CBL module, and x components consisting of a CBL module, a DWB module, and a CBL module are connected in series. The above two series parts are parallel structures and their outputs are input to the Concat module; the Concat module processes the input of the two parts, and its output is input to the Channel Shuffle module, and the output of the Channel Shuffle module is the output of the SFB1 module.
[0113] The SFB2 module consists of a Channel split module, a CBL module, a DWB module, a CBL module, a Concat module and a Channel Shuffle module. The input of the SFB2 module is first input into the Channel split module. The output of the Channel split module is input into the Concat module after being processed by x components consisting of a CBL module, a DWB module and a CBL module in series. The output of the Channel split module is directly input into the Concat module. The Concat module processes the two parts of the input, and its output is input into the Channel Shuffle module. The output of the ChannelShuffle module is the output of the SFB2 module.
[0114] The DWB module represents Depthwise Separable Convolution and BN layers. DWB is different from traditional convolution operations. The convolution part is divided into two steps: the first step is to select M convolution kernels to perform one-to-one convolution on the original channels without summing to generate M channels; the second step is to select N 1*1 convolution kernels to perform convolution operations on the feature layers of the M channels in the first step to obtain N results. The SFB1 module does not change the number of input and output channels and the size of the feature map. When the number of input and output channels is the same, the memory access amount MAC is the smallest. The SFB2 module is downsampling, which halves the width and height of the feature map and doubles the number of channels.
[0115] The Channel Shuffle module means randomly shuffling the channels formed by grouped convolution. Grouped convolution is to divide the feature layer into several groups, perform convolution separately, and then perform Concat operation. Using grouped convolution alone will cause the features of each channel to propagate forward in their own channels without intersecting each other, which will produce boundary effects. The resulting feature map is relatively limited and not suitable for the extraction and fusion of overall features. Channel Shuffle allows the channels to be combined and fused in a certain arrangement after channel grouping, and enter the next grouped convolution. In the backbone network of the Shufflenet network, SFB1 and SFB2 are arranged and combined alternately.
[0116] The lightweight YOLO v5 network model replaces the Focus layer of the YOLO v5 model with a convolutional layer and replaces the rest of its backbone network with the Shufflenet backbone network, while retaining the data preprocessing method of the YOLO v5 model.
[0117] According to the third criterion proposed in Shufflenetv2, too many fragmentation operations will affect the parallel computing speed of the hardware. The Foucs layer forms a four-way feature layer by performing four-way Slice operations on the input image. This approach will result in too many paths and reduce the hardware parallelism. Therefore, it is replaced with a common convolution layer CBL. The operation of common convolution depends largely on the memory size and processing speed of the GPU. Under the premise of ensuring the detection accuracy, a certain degree of accuracy can be appropriately lost to improve the detection speed.
[0118] The backbone network of the original YOLO v5 model is replaced with the Shufflenet backbone network. The Shufflenet backbone network is mostly grouped convolution of mixed channels. Using grouped convolution alone will cause the features of each channel to propagate forward in their own channels without intersecting with each other, which is not suitable for the extraction and fusion of overall features. After the channels are grouped, the channels are combined and fused in a certain arrangement and enter the next grouped convolution.
[0119] In the feature fusion part, two modules, CSP1_x and CSP2_x, are used. The convolution layers in front of these two modules are for downsampling, with a convolution kernel of 3*3 and a step size of 2. In the middle of the CSP1_x module, x residual blocks are accumulated, and in the middle of the CSP2_x module, x CBLs are accumulated. In the feature fusion part, the feature maps at different positions in the backbone network are fused and sent to the classification regression network. The classification regression network uses GIOU_Loss as the loss function of the Bounding box and uses cross entropy as the classification loss, which effectively solves the problem of non-overlapping bounding boxes.
[0120] 3) Training the network
[0121] The pre-trained weights obtained on ImageNet are used to initialize the backbone network. The convolutional layer parameters are initialized using the kaiming normal distribution. The rest of the network is initialized using Xavier. The backbone network parameters are frozen for the first 50 generations. The learning rate is set to decrease in steps with the number of training generations. The initial learning rate is set to 0.002 so that the network can find the optimal solution in the early stage of training and have better convergence in the later stage of training. The batch_size is set to 16, the maximum number of iterations is 300 epochs, and the regularization term decay weight is 5×e -4 , freeze the backbone network parameters in the first 50 generations.
[0122] The training set preprocessed in step 1) is input into the initialized lightweight YOLO v5 network model, and the backbone network is used for feature extraction and fusion. The classification and regression network is used to obtain the position, category and confidence of the human target, as well as the position and confidence of the wall, and the Loss value is obtained by comparing it with the true label. According to the Loss value, the SGD (gradient descent) optimizer is used to back-propagate and update the network parameters until the Loss drops to the preset value, and the network model training is completed.
[0123] 4) Network model verification
[0124] The verification set obtained in step 2) of the first step is input into the network model trained in step 3), and the detection label output by the network model is compared with the true label. The misrecognition rate is 1%, and the AP value (average accuracy) is calculated to be 94.1%. The network model is a valid model, and the network model verification is completed.
[0125] Table 1 Model validation results
[0126] Images Labels Ap False recognition rate 1050 6481 94.1% 1%
[0127] Step 3: Use the lightweight YOLO v5 network model and spatiotemporal memory mechanism module to detect pedestrian positions
[0128] 1) Obtain preliminary test results
[0129] The video stream captured by the camera is input frame by frame into the lightweight YOLO v5 network model verified in the second step to obtain the detection results of the video frame sequence images. The detection results of each image include the position, category and confidence of the human target, as well as the position and confidence of the wall. This detection result is a preliminary detection result.
[0130] 2) Obtain the corrected test results
[0131] The preliminary detection results are input into the spatiotemporal memory mechanism module. The principle of the spatiotemporal memory mechanism is as follows:
[0132]
[0133] Where P n+1 represents the confidence of the human target in the n+1th frame of the video sequence image; Δx and Δy represent the change values of the x-axis and y-axis of the closest human target between the n+1th frame and the i-th frame, respectively, and the value range is 0 to +∞; P i Represents the confidence of the person target in the i-th frame image output by the network model. Indicates rounding up.
[0134] In the above formula, the expression of f(x) is as follows:
[0135]
[0136] The implementation idea of this spatiotemporal memory mechanism is that the confidence of the person target in the next frame is determined by the spatial position changes of the person target in the previous 150 frames of images in the time series. This mechanism integrates the temporal dimension and spatial dimension of the person target to predict the confidence of the person target in the next frame of the image.
[0137] The confidence of the predicted person target in the next frame output by the spatiotemporal memory mechanism module replaces the confidence in the preliminary detection result of the corresponding frame to obtain a revised detection result.
[0138] 3) Pedestrian position detection
[0139] According to the detected human targets and walls (generally, it is considered that the confidence level is greater than 0.5 for the target to exist) in the corrected detection results in step 2), when the trajectory of the human target intersects with the set wall warning line or is lower than a certain threshold, it can be determined as a wall climbing behavior or illegal wall-taking behavior. With the upper left corner of the video screen as the coordinate origin, the positive directions of the x-axis and y-axis are set to the right and downward respectively to establish a two-dimensional coordinate system. The position of the wall is given manually, and the wall warning line is approximately a straight line. i ,y i Indicates the coordinates of the detected human target in the same coordinate system. The principle is as follows:
[0140] f(x,y)=Ax+By+C=0
[0141] Indicates the location of the fence cordon. If
[0142]
[0143] If
[0144] |f(x i ,y i )|<t
[0145] It indicates that there is suspicion of taking something through the wall. i ,y i It represents the position coordinates of the human target detected in the i-th frame image, t represents the set threshold, and A, B, and C are constant parameters calculated when the wall position is manually given.
[0146] The pedestrian detection method proposed in the present invention has a low false recognition rate, a lightweight model and a fast inference speed. Test results show that the false recognition rate is reduced from 7% to 1%, and the processing speed is increased from 56FPS to 74FPS.
[0147] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the enlightenment of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the purpose of the present invention and the claims, which all fall within the protection of the present invention.
[0148] Any matters not described in the present invention are applicable to the prior art.
Claims
1. A pedestrian detection method based on a lightweight YOLO v5 network model and spatiotemporal memory mechanism, characterized in that: The method comprises the following steps: Step 1: Build a training database 1) Collect images of different scenes at the monitoring location, including sunny days, cloudy days, rainy days and nights; divide the human targets in the collected images into different sizes according to the distance between the pedestrians and the camera, among which the targets with an area less than 32*32 pixels are small-sized targets, those between 32*32 pixels and 96*96 pixels are medium-sized targets, and those larger than 96*96 pixels are large-sized targets; select images according to the number ratio of the above three types of human target sizes of 1:1:1, perform data enhancement on the selected images, and finally perform image size unification operation to obtain the original data set; 2) Dataset division: The original dataset obtained in step 1) is manually labeled, and the people and walls in the image are marked with rectangular boxes, and the images in the original dataset are randomly divided into training set and validation set according to a certain number ratio; Step 2: Build a lightweight YOLO v5 network model 1) Preprocessing of the training set: Perform data augmentation on the training set obtained in step 2) of the first step; 2) Build a lightweight YOLO v5 network model The lightweight YOLO v5 network model is an improved structure of the YOLO v5 network model. Specifically, the Foucs module, CBL module, CSP1_1 module, CBL module, CSP1_3 module, CBL module, CBL module, CSP1_3 module, CBL module, CBL module, CBL module, SPP module, and CBL module in the backbone network of the YOLO v5 network model are replaced with 2 CBL modules, SFB1 module, 2 SFB2 modules, SFB1 module, 7 SFB2 modules, SFB1 module, SFB2 module, and CBS module in sequence; the input of the backbone network part of the lightweight YOLO v5 network model is first input into the first CBL module, and the output of the second SFB2 module is respectively input into the SFB1 module connected thereto and the first CBL module of the Neck network part, and the output of the CBS module is input into the CSP2_1 module in the backbone network of the YOLO v5 network model. The other parts of the structure of the lightweight YOLO v5 network model are the same as those of the YOLO The v5 network model is the same; 3) Training the network The pre-trained weights obtained on ImageNet are used to initialize the backbone network, the convolutional layer parameters are initialized using the Kaiming normal distribution, and the rest of the network is initialized using Xavier. The learning rate is set to decrease in steps with the number of training generations, and the backbone network parameters are frozen in the first 50 generations. The training set preprocessed in step 1) is input into the initialized lightweight YOLO v5 network model, and the backbone network is used to extract and fuse features. The classification and regression network is used to obtain the position, category and confidence of the human target, as well as the position and confidence of the wall, which are compared with the true label to obtain the Loss value; the SGD optimizer is used according to the Loss value to perform back propagation to update the network parameters until the Loss drops to the preset value, and the network model training is completed; 4) Network model verification The verification set obtained in step 2) of the first step is input into the network model trained in step 3), and the detection label output by the network model is compared with the true label to obtain the misrecognition rate. When the misrecognition rate is not greater than 10%, the current parameters of the network model are saved, and the network model is a valid model; when the misrecognition rate is greater than 10%, the initial parameters of the network model are adjusted, and the network is retrained until the loss drops to the preset value, and the misrecognition rate of the verification set is not greater than 10%, the current parameters of the network model are saved, and the current network model is a valid model, and the network model verification is completed; Step 3: Use the lightweight YOLO v5 network model and spatiotemporal memory mechanism module to detect pedestrian positions 1) Obtain preliminary test results Input the video stream captured by the camera frame by frame into the lightweight YOLO v5 network model verified in the second step to obtain the detection results of the frame sequence images of the video. The detection results of each image include the position, category and confidence of the human target, as well as the position and confidence of the wall. This detection result is the preliminary detection result. 2) Obtain the corrected test results The preliminary detection results are input into the spatiotemporal memory mechanism module. The spatiotemporal memory mechanism works as follows: Where P n+1 represents the confidence of the human target in the n+1th frame of the video sequence image; Δx and Δy represent the change values of the x-axis and y-axis of the closest human target between the n+1th frame and the i-th frame, respectively, and the value range is 0 to +∞; P i Represents the confidence of the person target in the i-th frame image output by the network model; Indicates rounding up; In the above formula, f(x) is expressed as follows: The confidence of the predicted person target in the next frame output by the spatiotemporal memory mechanism module is replaced with the confidence in the preliminary detection result of the corresponding frame to obtain a revised detection result; 3) Pedestrian position detection According to the detected human targets and walls in the corrected detection results in step 2), when the trajectory of the human target intersects with the set wall warning line or is lower than a certain threshold, it is determined as a wall climbing behavior or illegal wall-taking behavior; the upper left corner of the video screen is used as the coordinate origin, and the positive directions of the x-axis and y-axis are set to the right and downward respectively to establish a two-dimensional coordinate system. The position of the wall is given manually, and the wall warning line is a straight line; x i ,y i Indicates the coordinates of the detected human target in the same coordinate system; The principle is as follows: f(x,y)=Ax+By+C=0 Indicates the location of the fence cordon. If If |f(x i ,y i )|<t It indicates that there is suspicion of taking something through the wall; i ,y i It represents the position coordinates of the human target detected in the i-th frame image, t represents the set threshold, and A, B, and C are constant parameters calculated when the wall position is manually given.
2. The pedestrian detection method based on the lightweight YOLO v5 network model and the spatiotemporal memory mechanism according to claim 1 is characterized in that: In step 1) of the first step, the image data is enhanced by mirroring, cropping, translation and scaling.
3. The pedestrian detection method based on the lightweight YOLO v5 network model and the spatiotemporal memory mechanism according to claim 1, characterized in that: In step 2) of the first step, the ratio of the number of images in the training set to that in the validation set is 7:
3.
4. The pedestrian detection method based on the lightweight YOLO v5 network model and the spatiotemporal memory mechanism according to claim 1, characterized in that: In step 1) of the second step, the data enhancement method for the training set is Mosaic data enhancement.
Citation Information
Patent Citations
Lightweight convolutional neural network pedestrian recognition method
CN110321874A
Safety suit detection method and system based on improved YOLO V5
CN113553979A