Detection Method for Intrusion of Weighbridge Personnel Based on PSPNet and Improved YOLOv4
By combining PSPNet and improved YOLOv4 model detection method, the problem of personnel intrusion identification during the floor weighing process is solved, accurate identification of floor weighing areas and real-time detection of personnel intrusion are achieved, and identification accuracy and detection speed are improved.
Patent Information
- Application Number
- CN202111517711.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-13
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2041-12-13
AI Technical Summary
The prior art is difficult to effectively identify and prevent personnel intrusion during the weighing process of floor scales, resulting in errors in weighing data, and traditional image segmentation and infrared radiation imaging technologies have problems of poor robustness and high cost.
The ground scale personnel intrusion detection method based on the PSPNet semantic segmentation model and the improved YOLOv4 target detection model is adopted. The ground scale area is accurately identified through PSPNet, and combined with the improved YOLOv4 model to detect personnel and vehicles in real time, we can determine whether personnel illegally invade the ground scale area.
It realizes accurate identification of floor scale areas and real-time detection of personnel intrusions, improves identification accuracy and detection speed, reduces design and maintenance costs, and meets the requirements of real-time and accuracy.
Smart Images

Figure CN114627286B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image recognition during the weighing process of weighbridges, and in particular to a detection method for the intrusion of personnel on weighbridges based on PSPNet and improved YOLOv4. Background Art
[0002] Electronic truck scales are accurate and convenient weighing and metering devices. Over the years, they have been increasingly applied to various industries such as logistics, steel, building materials, coal, and asphalt. The alias of an electronic truck scale is a weighbridge, which is an effective mechanical manual weighing instrument and plays a very important role in an unattended weighing system.
[0003] Since the weighing system is unattended, metering cheating behaviors driven by economic interests will occur. For example, the most common situation is that when weighing a vehicle, it is not fully on the weighbridge, resulting in a smaller tare weight of the truck than the actual one and a larger net weight. These illegal behaviors will cause significant economic losses to enterprises and customers. Currently, methods such as setting gratings before and after the weighbridge have been used to avoid this situation. During the weighing process, illegal intrusion of personnel into the weighbridge area or the driver staying on the weighbridge for a long time, which are cheating methods that cause incorrect recording of weighing data, are urgent problems to be solved.
[0004] To address the above problems, it is necessary to separately identify the accurate weighbridge area and the two target objects of vehicles and personnel. Currently, most solutions still use traditional image segmentation techniques and image recognition and detection techniques such as infrared radiation imaging. The most common in traditional image segmentation techniques are thresholding and edge detection. Only low-level semantic information of the image is utilized when segmenting the image. The object segmentation effect is acceptable in simple scenarios, but when it comes to complex background segmentation scenarios, it is necessary to extract medium- and high-level semantics of the image to improve the segmentation effect, and it is sensitive to noise and has poor robustness. For the detection of targets such as personnel, infrared thermal imaging technology is relatively popular, but the thermal imaging pictures have low contrast, poor ability to distinguish details, and the price and maintenance cost of infrared thermal imagers are relatively high; for the common traditional image target recognition, features are extracted by sliding manually designed feature extractors, and classifiers such as SVM are used for classification output. However, manually designing features is time-consuming and labor-intensive, the recognition accuracy for personnel and vehicles of different sizes fluctuates greatly, and the target detection performance is weak under a large amount of data, and it takes a long time in image processing and cannot meet the requirements of real-time detection. Summary of the Invention
[0005] To solve the above problems, the present invention provides a detection method for the intrusion of personnel on weighbridges based on the PSPNet semantic segmentation model and the improved YOLOv4 target detection model, which is used for real-time image processing during the weighbridge weighing process to improve the real-time performance and accuracy of identifying and detecting cheating behaviors of personnel intrusion.
[0006] The technical solution for achieving the purpose of the present invention is as follows:
[0007] A detection method for intrusion of weighbridge weighing personnel based on PSPNet and improved YOLOv4, comprising the following steps:
[0008] Step 1, collect pictures of weighbridges at different locations, as well as pictures of incoming and outgoing vehicles and personnel, and collect pictures of idle weighbridges from multiple angles as the original data set;
[0009] Step 2, perform enhancement processing on the collected images to obtain the final data set;
[0010] Step 3, manually annotate the anchor points of the weighbridge area for the collected weighbridge pictures to generate corresponding mask bitmaps, and organize them into a VOC format data set; manually annotate the collected pictures of personnel and vehicles to generate corresponding xml files, and organize them into a VOC format data set;
[0011] Step 4, input the weighbridge picture data set into the PSPNet semantic segmentation network with the initial training hyperparameters set for training. The trained PSPNet model identifies and segments the idle weighbridge area;
[0012] Step 5, input the personnel and vehicle data sets into the improved YOLOv4 object detection model with the initial training hyperparameters set for training. The trained YOLOv4 model identifies and marks the prediction boxes of vehicles and personnel;
[0013] The improved YOLOv4 is: replace the backbone feature extraction network in the original YOLOv4 object detection model with the lightweight MobileNetV2. The backbone feature extraction network extracts three effective feature layers. The last effective feature layer is connected to the SPP module after 3 convolutional block operations. After the channel splicing operation of the SPP module, it goes through 3 more convolutional block operations; the middle convolutional block of the 3 convolutional blocks before and after the SPP module is a 3×3 ordinary convolutional block for feature extraction; in the PANet module, during the process of repeatedly extracting features from the three effective feature layers, there are 5 convolutional block operations after the channel splicing operation; the second and fourth convolutional blocks among the 5 convolutional blocks are both 3×3 ordinary convolutional blocks for feature extraction;
[0014] Change the ordinary convolution operation in these ordinary convolutional blocks to depthwise separable convolution for each channel and each point, and change the ReLU activation function to the ReLU6 activation function.
[0015] Step 6. Pass the framed surveillance video images into the trained PSPNet model and the improved YOLOv4 model to identify illegal intrusion of people in the video: Under the weighing state, determine the relative position of the person prediction frame identified by the improved YOLOv4 and the weighbridge area segmented by PSPNet, cut out a rectangular frame of the human leg area from bottom to top for the person prediction frame, and determine the overlap between the rectangular frame and the idle weighbridge area. If the overlap exceeds the set person overlap threshold, it is determined that the person is staying on the weighbridge at this moment, and start a timer for the target person to record the stay time. If the target person's recorded stay time exceeds the set time threshold, the person is determined to have illegally intruded.
[0016] Compared with the prior art, the present invention has the following significant advantages:
[0017] (1) The PSPNet semantic segmentation network is used to identify the weighing scale area at the pixel level. The pixel recognition accuracy is high, the PA value reaches 94%, and the MIoU value reaches 83.9%. It can accurately segment the weighing scale area. Compared with the traditional method, it is not affected by complex backgrounds. It can accurately identify the weighing scale area in environments with poor lighting conditions such as night, rainy days, and foggy days, and has good robustness.
[0018] (2) The improved YOLOv4 model is used to identify vehicles and personnel. The detection accuracy of the target is high, the mAP value reaches 90.36%, and the average number of images detected per second is 36.01. While taking into account the detection accuracy, the detection speed is fast, which is more in line with the real-time requirements of intrusion detection.
[0019] (3) Using the improved YOLOv4 target detection network, when creating personnel and vehicle data sets, the k-means clustering algorithm is used to obtain anchor box parameter values that are more suitable for the data set, making recognition more efficient and accurate.
[0020] (4) Combining the PSPNet semantic segmentation model and the YOLOv4 target detection model, the system can analyze the weighing status information, the number of people in the weighbridge area, and the number of people in the surrounding environment in real time, so as to facilitate the measurement personnel to better trace the weighing data and reduce the design and maintenance costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 This is a flow chart of the identification processing method of the present invention.
[0022] Figure 2 Partial dataset consisting of multiple weighbridges and corresponding labels.
[0023] Figure 3 Part of the dataset consisting of factory personnel and weighing vehicles and their corresponding labels.
[0024] Figure 4It is the ground scale area mask bitmap.
[0025] Figure 5 It is the ground scale area map recognized by using the PSPNet model.
[0026] Figure 6 It is the schematic diagram of the modified position of the convolutional block of the YOLOv4 enhanced feature extraction network structure.
[0027] Figure 7 It is the schematic diagram of the modification of the YOLOv4 enhanced feature extraction network structure.
[0028] Figure 8 It is the schematic diagram for judging the illegal intrusion behavior of personnel.
[0029] Figure 9 It is the map of the positions of people and vehicles recognized by using the improved YOLOv4 model during normal weighing of the ground scale.
[0030] Figure 10 It is the map of the positions of people and vehicles recognized by using the improved YOLOv4 model when there are personnel intruding into the ground scale. Specific implementation manners
[0031] In order to more clearly illustrate the specific implementation manners and technical purposes of the present invention, the following will further introduce the implementation of the present invention in detail with reference to the drawings. Obviously, the described examples are partial examples of the present invention. For those of ordinary skill in the art, other examples can be obtained without creative efforts.
[0032] Combined with Figure 1 , the detection method for ground scale personnel intrusion based on PSPNet and improved YOLOv4 of the present invention includes the following steps:
[0033] Step 1: Collect historical monitoring videos through the monitoring cameras in each ground scale area of each factory area, and use the Opencv vision library to perform frame-by-frame processing on the videos to obtain ground scale pictures, pictures of people and vehicles at different angles in multiple ground scale monitoring scenarios. The sizes and appearances of multiple ground scales are different, but their contours and proportions are relatively single and fixed. Therefore, 426 idle ground scale pictures and 2860 pictures of people and vehicles are collected as the original data set.
[0034] Step 2: Perform multiple data augmentation methods such as rotation, translation, brightness transformation, blur processing, and random cropping on some pictures of the idle ground scale data set and the people and vehicle data set to expand the data set. The initial data set and the data set after data augmentation together form the final data set. The final idle ground scale data set and the people and vehicle data set have a total of 994 and 5246 respectively.
[0035] Step 3: Manually annotate the anchor points of the weighbridge contour area in the idle weighbridge dataset using the Labelme semantic segmentation annotation tool. Set the weighbridge label as "weighbridge", then generate the corresponding json file, convert the json file into a mask bitmap of the corresponding image, and finally organize the original image, mask image, and train and val description files into a dataset in the VOC format standard. The ratio of dividing the training set and the test set is 8:2. Manually annotate the weighing vehicles and plant personnel in the personnel and vehicle datasets using the LabelImg data annotation tool. Set the weighing vehicle label as "truck" and the plant personnel label as "person". The tool generates xml files for the corresponding images of the dataset. Finally, make the original images and xml files into a dataset in the VOC2007 format and divide it into a training set and a test set at a ratio of 8:2.
[0036] Step 4: Build a PSPNet semantic segmentation model, input the final idle weighbridge dataset in Step 3 into the PSPNet model for training, use the trained PSPNet model to semantically segment the idle weighbridge area, and use the mean intersection over union (MIoU), pixel classification accuracy (PA), and FPS as evaluation indicators for the effect of identifying the weighbridge area to evaluate the model performance.
[0037] The specific steps are as follows:
[0038] 4.1. Set the maximum number of iterations to 10,000 times. Freeze the backbone feature extraction network for the first 50 epochs for training. The learning rate (lr) during the training process is 0.0001, and the batch size is 16. After the model is unfrozen, train for another 50 epochs. The learning rate during the training process is 0.00001, and the batch size is 8. During the training process, use cross-entropy loss and dice loss to calculate the loss value. To facilitate the rapid convergence of the model, use the Adam optimizer to optimize the network parameters.
[0039] 4.1.1. Calculate the cross-entropy loss function:
[0040] L = y log y'+(1 - y) log(1 - y')
[0041] where L represents the cross-entropy loss function, y is the sample label, the positive class is 1, the negative class is 0, and y' is the probability of predicting a positive class sample.
[0042] 4.1.2. Calculate the Dice Loss:
[0043] The Dice loss takes the evaluation index of semantic segmentation as the Loss. The Dice coefficient is a set similarity measurement function used to calculate the similarity between two samples, with a value range of [0,1]. The calculation formula is as follows:
[0044]
[0045] Where X represents the prediction result, Y represents the true result, and S represents the Loss value. The larger S is, the greater the overlap between the prediction result and the true result. The Dice coefficient is better when it is larger, and as a Loss, it is better when it is smaller. Therefore, Dice loss = 1 - Dice is used as the loss function for semantic segmentation.
[0046] 4.2. Use the trained PSPNet model for semantic segmentation of the weighbridge area, which mainly includes the following steps.
[0047] 4.2.1. Preprocess the input image without distortion by adding gray bars, resize it to (3, 473, 473), and use the processed image as the input to the backbone feature extraction network Resnet50 to obtain Feature Maps of different scales.
[0048] 4.2.2. Divide the Feature Map extracted by the backbone feature extraction network into two parts. One part is used as the global feature, and the other part is passed into the enhanced feature extraction network for further feature extraction. PSPNet uses the pyramid pooling module as the enhanced feature extraction structure. This module divides the input feature layer into four regions of different scales: 6×6, 3×3, 2×2, and 1×1, and then performs average pooling on each region internally.
[0049] 4.2.3. Upsample the pooled pyramid feature map using bilinear interpolation to scale it to the size of the original feature map. Then use 3×3 convolution to integrate the features and fuse them into global prior information. Finally, use 1×1 convolution to adjust the channels and upsample to the final prediction segmentation map with the same width and height as the input image.
[0050] 4.3. Use the mean intersection over union (MIoU), pixel classification accuracy (PA), and FPS as evaluation metrics for identifying the effect of the weighbridge area, and obtain the performance of PSPNet semantic segmentation.
[0051] 4.3.1. MIoU is a standard metric for semantic segmentation models. It first calculates the IoU (intersection over union of the true label and the predicted label) for each class, and then takes the mean of the IoUs of all classes. The intersection over union is the intersection of the predicted expectation and the actual area divided by the union of the two. The MIoU calculation formula is as follows:
[0052]
[0053] In the formula: k is the number of classifications. Usually, there is a background class in semantic segmentation, so k + 1 is used, p ijrepresents the total number of pixels of the i-th true category that are mispredicted as the j-th category, p ji is the total number of pixels of the j-th true category that are correctly predicted as the i-th category, is the total number of pixels of the i-th category in the picture, that is, the marked area, is the total number of pixels predicted as the i-th category by the model in the picture, that is, the predicted area.
[0054] 4.3.2. PA is also a standard metric for semantic segmentation models, that is, the ratio of the number of pixels with correct prediction classification to the total number of pixels in the picture. The calculation formula is as follows:
[0055]
[0056] 4.3.3. The average accuracy of the PSPNet model on the training set is 95.5%, and the average accuracy on the test set is 93.9%. The highest average intersection over union ratio reaches 83.9%. Save the training weights when the training generations reach 30, 65, and 100. The indicators such as intersection over union, accuracy, and FPS of the corresponding weights are shown in Table 1.
[0057] Table 1 Training results of PSPNet model
[0058]
[0059] It can be seen from Table 1 that when training to the 65th generation, the comprehensive effect is relatively good. Actually, save the model weights of the 79th training epoch to identify the weighbridge area at the pixel level.
[0060] Step 5: Build a YOLOv4 object detection model and perform lightweight improvement. Input the final personnel and vehicle datasets in Step 3 into the improved YOLOv4 model for training. Use the trained improved YOLOv4 model to identify the weighing vehicles and personnel, and use the mean average precision (mAP) and FPS as the evaluation indicators for the recognition effects of personnel and vehicles to evaluate the model performance. The specific steps are as follows:
[0061] 5.1. Build the YOLOv4 object detection network model and perform lightweight improvement on it. Replace the backbone feature extraction network of the model with MobileNetV2. After extracting three effective feature layers, the last feature layer of 1024×13×13 is connected to the SPP module after passing through 3 convolutional blocks. After global pooling by SPP and concatenating the feature layers, it passes through another 3 convolutional blocks. The 3 convolutional blocks before and after the SPP module are used for feature extraction, and the middle one is a 3×3 ordinary convolutional block for feature extraction; in the PANet module, the three effective feature layers are used to repeatedly extract features through multiple upsampling, downsampling, feature layer concatenation, and convolutional block convolution operations. There are 5 convolutional blocks after each feature layer concatenation, and 2 of them are 3×3 convolutional blocks for feature extraction; the 3×3 convolutional block in the enhanced feature extraction network consists of ordinary convolution, BN normalization, and ReLU activation function. Based on depthwise separable convolution, this invention uses pointwise and depthwise convolutions to replace the ordinary convolution in the 3×3 convolutional block, and uses ReLU6 as the activation function. The replaced 3×3 convolutional block consists of depthwise separable convolution, BN normalization, and ReLU6 activation function. The number of parameters of the improved model is reduced to about one-sixth of the original YOLOv4. The number of parameters of the model before and after improvement is shown in Table 2.
[0062] Table 2 Comparison of the number of parameters of the YOLOV4 model before and after improvement
[0063]
[0064] 5.2. Set the maximum number of iterations to 25000, set the frozen training epoch to 50, the learning rate during frozen training to 0.001, the batch size to 32. After the model is unfrozen, train for another 50 epochs, the learning rate during unfrozen training is 0.0001, and the batch size is 16. The weight decay regularization coefficient for the entire training process is 0.0005, and the momentum coefficient is 0.9. The loss function used for training consists of the bounding box regression loss L ciou , confidence loss L conf and classification loss L class . These three parts. If there is no target in a certain prior box, only calculate the confidence loss; otherwise, calculate the three losses.
[0065] 5.2.1. The bounding box regression loss CIoU takes into account the center distance between the target and the anchor, the scale information of the aspect ratio, and the overlap degree of the bounding boxes on the basis of IoU, and will not have the problem of divergence during training like IoU. The CIoU formula is as follows:
[0066]
[0067]
[0068]
[0069] Among them, L ciou represents the value of the bounding box loss function, and L iou represents the value of the intersection over union loss function. ρ 2 (b, b gt ) represents the Euclidean distance between the centers of the predicted box and the ground truth box; c represents the diagonal distance of the smallest closed region that contains both the predicted box and the ground truth box; w p represents the width of the predicted box; h p represents the height of the predicted box; w gt represents the width of the ground truth box; h gt represents the height of the ground truth box. ν is used to measure the similarity of the aspect ratio, and α is the weight coefficient for balancing the ratio.
[0070] 5.2.2. Confidence Loss L conf It is calculated by the cross-entropy method, and the formula is as follows:
[0071]
[0072] In the formula, s 2 represents the number of grids into which the image is divided; b represents the number of prior boxes for each grid; and represent 1 and 0 respectively if the t-th prior box in the k-th grid has an object, and 0 and 1 respectively if there is no object; λ noobj represents the loss weight of the confidence of the bounding box that does not contain an object; c k and c k represent the predicted class and the actual class to which the k-th grid belongs respectively.
[0073] 5.2.3. Classification Loss L class It is calculated by the cross-entropy method, and the formula is as follows:
[0074]
[0075] In the formula, s 2 represents the number of grids into which the image is divided, represents whether the k-th grid contains an object, p k and p k (c) represent the predicted object probability and the actual object probability of the k-th grid respectively.
[0076] 5.3. Use the trained improved YOLOv4 model to perform object detection on the weighed vehicles and personnel, which mainly includes the following steps.
[0077] 5.3.1. Use the k-means clustering algorithm to perform clustering analysis on vehicles and personnel in the sample set, screen out the prior box sizes that better match the detection objects in the data set, and obtain 9 anchor boxes for predicting the target. Each yolo head feature map will correspond to 3 anchor boxes respectively.
[0078] 5.3.2. Preprocess the images after frame-by-frame segmentation of the surveillance video without distortion by adding gray bars, and uniformly adjust them to a size of (3, 608, 608). Then, use the processed images as inputs and pass them into the backbone feature extraction network MobileNetv2 to obtain three Feature Maps as effective feature layers.
[0079] 5.3.3. Use the SPP spatial pyramid pooling structure and the PANet path aggregation structure as the neck to obtain three yolo head feature maps of different sizes. The three yolo head feature maps in this example are (21, 76, 76), (21, 38, 38), and (21, 19, 19) respectively, corresponding to the positions of three prediction boxes on the grids where the pictures are divided into 76×76, 38×38, and 19×19. The first dimension 21 in the yolo head feature map represents 3×(4 + 1 + 2), where 3 represents the 3 preset prior boxes, 4 represents the adjustment parameters for the width, height, and center of the prior box, 1 represents whether there is a target, and 2 represents the two categories of vehicles and personnel to be detected.
[0080] 5.3.4. The three effective feature layers divide the picture into grids of 19×19, 38×38, and 76×76. Each grid is responsible for predicting a region. The prediction results of the feature layer correspond to the positions of three prediction boxes. YOLO adds the adjustment parameters of the corresponding prior box center to each grid point and then adjusts the width and height parameters to determine the length, width, and position of the prediction box. Finally, the prediction boxes of the target objects are sorted by confidence scores and screened by non-maximum suppression to obtain the prediction box closest to the target.
[0081] 5.4. Conduct object detection experiments on the original YOLOv4 algorithm, the lightweight version YOLOv4-tiny, and the improved YOLOv4 algorithm in this paper under the same software and hardware environment and data set, and compare the differences in accuracy, detection speed, and model size among them. The experimental results are shown in Table 3. The lightweight improved YOLOv4 of the present invention has decreased by 2.03% compared to the original YOLOv4, the weight size has decreased by 203MB, the detection speed has been greatly improved, and the number of pictures recognized per second has increased by nearly 14. Compared with YOLOv4-tiny, the detection speeds are similar, but the detection accuracy is much higher. Through comprehensive comparison, the improved YOLOv4 takes into account both detection speed and accuracy, meeting the accuracy and real-time requirements of the weighbridge personnel intrusion detection of the present invention.
[0082] Table 3 Comparison of Test Results of Improved YOLOV4 Algorithm
[0083]
[0084] Step 6: Identify the illegal intrusion behavior of personnel in the weighing state of the weighbridge and send an alarm message. The specific process is as follows.
[0085] 6.1 After the improved YOLOv4 model identifies the weighing vehicle, calculate the overlap degree between the vehicle and the weighbridge area segmented by PSPNet. If the overlap degree reaches the set vehicle overlap threshold, it is determined that the current weighbridge is in the weighing state.
[0086] 6.2 The improved YOLOv4 model identifies the prediction box of the personnel, intercepts the rectangular box of the personnel's leg area, and calculates the overlap degree between the rectangular box and the idle weighbridge area. If the overlap degree reaches the set personnel overlap threshold, it is determined that the personnel are staying on the weighbridge at this moment, and a timer is started to record the stay time of the personnel. If the stay time of the target personnel exceeds the set stay time threshold, the illegal intrusion behavior of the personnel is determined.
[0087] 6.3 Save the information such as the number of target personnel, the intrusion time recorded by the corresponding timer, and the weighbridge point number into the database, and push the alarm message to the relevant personnel through a Web pop-up window and WeChat service.
[0088] In view of the problems that the traditional image segmentation technology cannot extract medium and high-level semantics in the image to process complex backgrounds, the cost of infrared image detection is high, and the detection time is long, resulting in poor real-time performance, etc., the present invention proposes a detection method for weighbridge personnel intrusion that combines PSPNet semantic segmentation and improved YOLOv4 object detection. The PSPNet model performs pixel-level detection on the monitoring pictures to segment the accurate weighbridge area, and the improved Yolov4 model processes the monitoring video frame by frame to detect personnel and vehicles in real time. When in the weighing state, it can detect the illegally intruding personnel in time and immediately send an alarm message, and discover and prevent in time the cheating behavior of illegal personnel staying on the weighbridge during weighing, resulting in abnormal weighing data.
Claims
1. A detection method for intrusion of weighbridge weighing personnel based on PSPNet and improved YOLOv4, characterized in that, it includes the following steps: Step 1: Collect pictures of weighbridges at different locations, as well as pictures of incoming and outgoing vehicles and personnel, and collect pictures of idle weighbridges from multiple angles as the original data set; Step 2: Perform enhancement processing on the collected images to obtain the final data set; Step 3: Manually annotate the weighbridge area of the collected weighbridge pictures to generate corresponding mask bitmaps, and organize them into a VOC format data set; Manually annotate the collected pictures of personnel and vehicles to generate corresponding xml files, and organize them into a VOC format data set; Step 4: Input the weighbridge picture data set into the PSPNet semantic segmentation network with initial training hyperparameters set for training. The trained PSPNet model identifies and segments the idle weighbridge area; Train the PSPNet semantic segmentation model and identify the idle weighbridge area. The specific training hyperparameter settings and loss function are as follows: Set the maximum number of iterations for model training, the number of frozen training epochs, the learning rate during the frozen training process, the batch size during the frozen training process, the number of unfrozen training epochs, the learning rate during the unfrozen training process, and the batch size during the unfrozen training process; Set the model training loss function as the cross-entropy function and Dice Loss, and set the Adam optimizer to optimize the model parameters; The construction process of the PSPNet semantic segmentation model for pixel-level identification of the weighbridge area is as follows: (1) Preprocess the pictures after frame-by-frame splitting of the surveillance video without distortion by adding gray bars, and extract Feature Maps of different scales through the Resnet50 backbone feature extraction network; (2) Divide the extracted Feature Maps into two parts; one part is used as the global feature, and the other part is input into the pyramid pooling module to strengthen feature extraction. The pyramid pooling module divides this part of the feature layer into four scales of regions, and performs average pooling processing within each region; (3) Perform upsampling operation on the pooled feature layer using bilinear interpolation method, splice the channels with the global feature layer, use 3×3 convolution for the entire feature, 1×1 convolution to adjust the channels, and upsample to the final predicted segmentation map with the original width and height; Step 5: Input the personnel and vehicle data set into the improved YOLOv4 object detection model with initial training hyperparameters set for training. The trained YOLOv4 model identifies and marks the prediction boxes of vehicles and personnel; Improve YOLOv4 as follows: Replace the backbone feature extraction network in the original YOLOv4 object detection model with the lightweight MobileNetV2. The backbone feature extraction network extracts three effective feature layers. After the last effective feature layer undergoes 3 convolutional block operations, it is connected to the SPP module. After the channel concatenation operation of the SPP module, it undergoes another 3 convolutional block operations. The middle convolutional block of the 3 convolutional blocks before and after the SPP module is a 3×3 ordinary convolutional block for feature extraction. In the PANet module, during the process of repeatedly extracting features from the three effective feature layers, there are 5 convolutional block operations after the channel concatenation operation. The second and fourth convolutional blocks among the 5 convolutional blocks are both 3×3 ordinary convolutional blocks for feature extraction. Change the ordinary convolution operation in these ordinary convolutional blocks to depthwise separable convolution that is channel-wise and pointwise, and change the ReLU activation function to the ReLU6 activation function. Step 6: Input the frame-split surveillance video images into the trained PSPNet model and the improved YOLOv4 model to identify the illegal intrusion phenomenon of personnel in the video. In the weighing state, judge the relative position between the person prediction box identified by the improved YOLOv4 and the weighbridge area segmented by the PSPNet. Intercept the rectangular box of the human leg area from the bottom up of the person prediction box, and judge the coincidence degree between this rectangular box and the idle weighbridge area. If the coincidence degree exceeds the set person coincidence degree threshold, it is determined that the person is staying on the weighbridge at this moment, and a timer is started for this target person to record the stay time. If the recorded stay time of the target person exceeds the set time threshold, then the illegal intrusion behavior of the person is determined.
2. The detection method for weighbridge weighing personnel intrusion based on PSPNet and improved YOLOv4 according to claim 1, characterized in that, In step 5, the training of the YOLOv4 network model for target detection of personnel and vehicles is as follows. The specific training hyperparameters and loss function are as follows: Set the number of target categories for model training, the maximum number of iterations, the number of epochs for frozen training, the learning rate for frozen training, the batch size for frozen training, the number of epochs for unfrozen training, the learning rate for unfrozen training, the batch size for unfrozen training, the regularization coefficient of weight decay, and the momentum coefficient; set that the model training loss function consists of the bounding box regression loss L ciou , the confidence loss L conf and the classification loss L class and set the Adam optimizer to optimize the model parameters.
3. The detection method for weighbridge weighing personnel intrusion based on PSPNet and improved YOLOv4 according to claim 2, characterized in that, Bounding box regression loss L ciou The calculation formula is as follows: Among which L ciou represents the value of the bounding box loss function, and L iou represents the value of the intersection over union loss function, and ρ 2 (b, b gt ) represents the Euclidean distance between the centers of the predicted box and the ground truth box; c represents the diagonal distance of the smallest closed region that contains both the predicted box and the ground truth box; w p represents the width of the predicted box; h p represents the height of the predicted box; w gt represents the width of the ground truth box; h gt represents the height of the ground truth box, ν is used to measure the similarity of the aspect ratio, and α is the weight coefficient for balancing the ratio.
4. The detection method for weighbridge weighing personnel intrusion based on PSPNet and improved YOLOv4 according to claim 1, characterized in that, In step 5, the target detection of personnel and vehicles is carried out as follows: (1) Use the k-means clustering algorithm to perform clustering analysis on vehicles and personnel in the sample set to obtain multiple anchor boxes for predicting targets. (2) Preprocess the frame-split images of the surveillance video in a distortion-free manner by adding gray bars, and use the backbone feature extraction network MobileNet v2 to extract three effective feature layers. (3) Use the SPP spatial pyramid pooling structure and the PANet path aggregation structure as the enhanced feature extraction network to perform enhanced feature extraction to obtain the prediction parameters of the prior boxes, including the adjustment parameters of the prior box center, the width and height adjustment parameters, the classification, and the confidence of the target category. (4) Combine the adjustment parameters of the center of the obtained prior box with the width and height adjustment parameters to determine the length, width, and position of the prediction box, and perform confidence score sorting and non-maximum suppression screening on the prediction boxes of the target object to obtain the final prediction box closest to the target.
Citation Information
Patent Citations
Wagon balance human body intrusion detection method based on Mask Rcnn and SSD
CN112861631A
Outdoor parking lot unoccupied parking space detection method based on deep learning
CN113076904A