A method for detecting illegal spam delivery behavior based on deep learning
By installing a surveillance camera near the garbage delivery kiosk, using deep learning algorithms to detect the location and movement of pedestrians and garbage in real time, the problem of difficulty in accurately detecting illegal garbage delivery behavior in the existing technology is solved, and efficient and accurate garbage delivery monitoring is achieved.
Patent Information
- Application Number
- CN202210608558.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-31
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-05-31
AI Technical Summary
The prior art is difficult to accurately detect whether there are illegal waste delivery near the garbage delivery kiosk in real time, and manual monitoring costs are high and the effect is not good.
The object detection algorithm based on deep learning is used to obtain video images in real time through the surveillance camera, detect the location and actions of pedestrians, garbage bags and garbage cans, and combine the handover and key point detection to determine whether pedestrians carry garbage and perform delivery action recognition.
Real-time detection near garbage delivery kiosks is realized, and illegal garbage delivery behaviors are accurately identified, which reduces manual monitoring costs and improves detection efficiency and accuracy.
Smart Images

Figure CN115100588B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of deep learning, and relates to technologies such as deep neural networks and video image processing. In particular, it relates to a method for detecting illegal garbage delivery behavior based on deep learning. Background Art
[0002] In order to build a more environmentally friendly community and respond to the garbage classification policy, many communities have adopted a method of scheduled time-sharing garbage delivery to regularly dump the domestic garbage generated by residents in the community. At present, there are many households in the community that do not abide by the rules and deliver garbage on time. Although community staff are dispatched to squat near the garbage delivery kiosk to verbally persuade the households that deliver garbage illegally, this method has a huge labor cost, little effect, and there are no staff on duty at all times. Most illegal residents deliver garbage illegally with a fluke mentality.
[0003] The key to solving the above problems is to detect in real time whether there are residents coming to deliver garbage near the garbage delivery kiosk. During the detection process, it is necessary to accurately identify the illegal personnel who have the behavior of illegal garbage delivery and also distinguish the residents passing by with plastic bags. In order to better implement the garbage classification policy, it is crucial to detect illegal garbage delivery behavior in real time. Summary of the Invention
[0004] Aiming at the problems existing in the above background art introduction, the purpose of the present invention is to provide a method for detecting illegal garbage delivery behavior based on deep learning that can detect and judge in real time whether there is illegal garbage delivery behavior.
[0005] The technical solution adopted by the present invention is as follows:
[0006] A method for detecting illegal garbage delivery behavior based on deep learning, and its specific steps are as follows:
[0007] Step (1), obtain the images captured by the monitoring camera installed outside the garbage delivery kiosk during non-garbage delivery time periods;
[0008] Step (2), use a deep learning-based object detection algorithm to perform frame-by-frame detection on pedestrians p, garbage bags r, and trash cans e in the received video images, and obtain the detection box coordinate information and categories of pedestrians, garbage bags, and trash cans;
[0009] Step (3), if a pedestrian p is detected in the t-th frame image, intercept all the garbage bag and pedestrian detection box images in the t-th frame image where represents the i-th garbage bag detection box image in the t-th frame, represents the j-th pedestrian detection box image in the t-th frame; and takes the detection box images, detection box coordinate information, and categories of all pedestrians and garbage bags in this frame as the input of the tracking module. The tracking module assigns an id to each pedestrian and garbage bag;
[0010] Step (4): Input the pedestrian detection box image detected in the t-th frame into the human key point detection module, and the human key point detection module outputs the key point coordinate information of the human body in the image;
[0011] Step (5): Calculate the intersection over union (IoU) between all garbage bags and all pedestrians in the t-th frame. Based on the size of the IoU, determine whether a pedestrian is carrying a garbage bag into the monitoring area. If a pedestrian carrying a garbage bag is detected, perform an addition or update operation on the pedestrian and garbage bag information in the data table. If no pedestrian carrying a garbage bag is detected, go to step (2);
[0012] Step (6): After updating the data table, calculate the distance between each garbage bag carried by a pedestrian in the t-th frame and the trash can, and whether the pedestrian makes a delivery action to determine whether the pedestrian has the behavior of illegally throwing garbage.
[0013] Furthermore, when the target detection algorithm in step (2) fails to detect pedestrian p j for consecutive Η frames, it is considered that pedestrian p j has left the video monitoring range, and delete the information of pedestrian p j in the data table.
[0014] Furthermore, the deep learning-based target detection algorithm in step (2) is trained as follows:
[0015] Step (2.1) Data preparation: Use the monitoring camera installed outside the garbage delivery kiosk to take an image every ten seconds, and manually annotate the image. The annotation information is the detection box information of pedestrians, garbage bags, and trash cans, i.e., (f, x, y, w, h), where f represents the category. f = 0 represents a pedestrian, f = 1 represents a garbage bag, f = 2 represents a trash can, x and y represent the horizontal and vertical axis coordinates of the center point of the target, and w and h represent the length and width of the target; Divide the annotated data samples into a training set, a validation set, and a test set according to 8:1:1;
[0016] Step (2.2) Network structure design: The object detection network takes the images captured by the surveillance camera as input. First, it performs FOCUS slicing operation on the input image, and then passes through multiple convolutional groups A followed by batch normalization and activation functions, and convolutional group B containing residual components. After passing through the pooling group composed of convolutional and pooling layers, it is connected to convolutional group C and convolutional group A. In the backbone network, three detection branches of different scales are formed through skip connections. The three branches respectively splice and convolve the feature maps output from the backbone network, and finally use convolutional group C followed by a convolutional layer to obtain the coordinates, confidence levels, and class information of pedestrians, garbage bags, and trash cans in the figure;
[0017] Step (2.3) Network training: Input the training set in step (2.1) into the object detection network for training. During the training process, the learning rates of the weight layer, bias layer, and batch normalization layer are adjusted separately. At the beginning of training, one-dimensional linear interpolation is used to update the learning rate. After the learning rate warm-up is completed, the cosine annealing algorithm is used to update the learning rate;
[0018] Step (2.4) Model testing: Input the images captured by the surveillance camera and output the coordinate information, class, and confidence level of pedestrians, garbage bags, and trash cans.
[0019] Furthermore, the cosine annealing algorithm in step (2.3) is expressed as:
[0020]
[0021] where, new lr represents the newly obtained learning rate, initial lr represents the initial learning rate, eta min represents the minimum learning rate, cur epoch represents the current number of training epochs, T max represents the total number of training epochs;
[0022] The network loss function is expressed as:
[0023]
[0024]
[0025] where ρ() is the Euclidean distance, b and b gt respectively represent the center point coordinates of the predicted target box and the actual target box, c represents the diagonal distance of the minimum circumscribed rectangle of the predicted box B and the actual box B gt ; a represents the trade-off coefficient, u represents the aspect ratio consistency parameter, w gt and h gt represent the width and height of the actual target box, w and h represent the width and height of the predicted target box; IOU represents the predicted box B and the actual box Bgt Intersection over Union
[0026] Further, the tracking module in step (3) is implemented as follows:
[0027] Step (3.1) Feature extraction: Input all the garbage bag and pedestrian detection box images in the t-th frame into a lightweight image feature extraction network model formed by stacking convolutional layers. Use a convolutional neural network to extract the image features of all the garbage bag and pedestrian detection box images, and each image obtains a 512-dimensional feature vector
[0028] Step (3.2) Object matching: Input the garbage bag and pedestrian detection box images detected in frames t + 1, t + 2, and t + 3 into the feature extraction network in step (3.1) to obtain the corresponding feature vectors:
[0029]
[0030]
[0031]
[0032] Calculate the distance l between the feature vectors of adjacent two frames. If l is less than the threshold δ, then it is considered that and the pedestrians in are the same pedestrian; Predict the position of the remaining detection boxes in the previous frame in the next frame through the Kalman filter algorithm. If the intersection over union of the predicted detection box and the actual detection box in the next frame is greater than the threshold ε, then it is considered that the targets in these two detection boxes are the same target
[0033] Further, the specific steps for establishing the lightweight image feature extraction network model in step (3.1) are as follows:
[0034] Step (3.1.1) Data preparation: Train the image feature extraction network through a publicly available pedestrian re-identification dataset and a self-made garbage bag re-identification dataset. Each pedestrian and garbage bag in the dataset has multiple images taken from different angles. Divide the labeled data samples into a training set, a validation set, and a test set according to 8:1:1 respectively
[0035] Step (3.1.2) Network structure design: The backbone network takes an image as input. After passing through a convolutional layer with a convolutional kernel size of 3X3, the obtained feature map is input into an inverted residual structure composed of a 1X1 convolution and a DepthWise convolution with a convolutional kernel size of 3X3. Channel attention is formed through a pooling layer and a convolutional layer. After the feature map is input into the stacked inverted residual and attention structures, it passes through a 1X1 convolutional layer, a pooling layer, and two fully connected layers to finally output a vector of size 1X512
[0036] Step (3.1.3) Network training: Input multiple pedestrian images or garbage bag images, and use the two types of images to train a pedestrian image feature extraction model and a garbage bag image feature extraction model respectively.
[0037] Furthermore, in step (3.2), only the targets that are matched in consecutive k frames will be assigned a unified ID. And if a target with an existing ID fails to be matched continuously for multiple times, it is considered that the target has left the camera monitoring range, and the ID is deleted from the data table.
[0038] Furthermore, the specific steps for establishing the human key point detection module in step (4) are as follows:
[0039] Step (4.1) Data preparation: Train the model through a publicly available human key point data set. The annotation information of the images in the data set is the coordinate information and key point information of the pedestrians in the image; divide the labeled data samples into a training set, a validation set, and a test set according to 8:1:1 respectively.
[0040] Step (4.2) Network structure design: The backbone network takes the image as input. After passing through a convolutional layer with a convolutional kernel size of 3X3, the obtained feature map is input into an inverted residual structure composed of a 1X1 convolution and a DepthWise convolution with a convolutional kernel size of 3X3. Channel attention is formed through a pooling layer and a convolutional layer. After the feature map is input into the stacked inverted residual and attention structure, a 1X1 convolutional layer, a pooling layer, and a convolutional layer for regressing the key point coordinates are added.
[0041] Step (4.3) Network training: Input the labeled pedestrian images to train the human key point detection module.
[0042] Furthermore, the judgment of whether a pedestrian enters while carrying garbage in step (5) is carried out as follows:
[0043] Calculate the intersection over union iou between all garbage bags and all pedestrians in the t-th frame. If the intersection over union iou i between the garbage bag r j and the pedestrian p p,r is greater than the threshold α, and there is no information in the data table that the pedestrian p j carries the garbage bag r i , then it is considered that the pedestrian p j carries the garbage bag r i and enters the video monitoring range. Add the information that the pedestrian p j carries the garbage bag r i to the data table. If the information that the pedestrian p j carries the garbage bag r i already exists in the data table, then use the information in the t-th frame to update the data table to replace the original information.
[0044] Further, the determination of whether a pedestrian has committed an illegal garbage delivery behavior in step (6) is carried out as follows:
[0045] After updating the data table, calculate the distances between the left and right wrists of each pedestrian in the t-th frame and the garbage bags they carry to determine whether the pedestrian holds the garbage bag with the left hand or the right hand; if there is a pedestrian p j holds the garbage bag r with the right hand i , then calculate p j the angle d between the right arm and the body, p j and r i 's intersection over union iou p,r and r i 's intersection over union iou with the trash can e in the t-th frame r,e . If d is greater than the threshold β and iou p,r is equal to 0 and iou r,e is greater than the threshold γ, then it is considered that p j has committed an illegal garbage delivery behavior.
[0046] Compared with the prior art, the significant advantages of the present invention include: using partial community monitoring and deep learning algorithms to detect in real time residents carrying plastic bags near garbage delivery kiosks and determine whether they have committed illegal garbage delivery behaviors. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 is the block diagram of the detection system of the present invention;
[0048] Figure 2 is the schematic diagram of the target detection network structure of the present invention;
[0049] Figure 3 is the schematic diagram of the target detection network module structure of the present invention;
[0050] Figure 4 is the schematic diagram of the structures of the image feature extraction network and the human key point detection network of the present invention;
[0051] Figure 5 is the schematic diagram of the inverted residual and attention module structures in the image feature extraction network and the human key point detection network of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0052] The present invention will be further described below in conjunction with specific embodiments, but the present invention is not limited to these specific embodiments. Those skilled in the art should recognize that the present invention covers all alternative solutions, improvement solutions, and equivalent solutions that may be included within the scope of the claims.
[0053] See Figure 1 , this embodiment provides a method for detecting illegal garbage delivery behaviors based on deep learning, and its specific steps are as follows:
[0054] Step (1): Obtain the images captured by the monitoring camera installed outside the garbage delivery kiosk during non-garbage delivery time periods.
[0055] Step (2): Use a deep learning-based object detection algorithm to perform frame-by-frame detection on pedestrians p, garbage bags r, and trash cans e in the received video images, and obtain the detection box coordinate information (center point coordinates, rectangle length, and rectangle width) and categories of pedestrians, garbage bags, and trash cans.
[0056] The deep learning-based object detection algorithm is trained as follows:
[0057] Step (2.1) Data preparation: Take an image every ten seconds through the monitoring camera installed outside the garbage delivery kiosk, and manually annotate the images. The annotation information is the detection box information of pedestrians, garbage bags, and trash cans, i.e., (f, x, y, w, h), where f represents the category. f = 0 represents a pedestrian, f = 1 represents a garbage bag, and f = 2 represents a trash can. x and y represent the horizontal and vertical axis coordinates of the center point of the target, and w and h represent the length and width of the target. Divide the annotated data samples into a training set, a validation set, and a test set according to 8:1:1.
[0058] Step (2.2) Network structure design: The object detection network takes the images captured by the monitoring camera as input. First, perform a FOCUS slicing operation on the input image, and then pass through multiple convolutional groups A followed by batch normalization and activation functions, and convolutional group B containing residual components. After passing through the pooling group composed of convolutional and pooling layers, connect convolutional group C and convolutional group A. In the backbone network, three detection branches of different scales are formed through skip connections. The three branches respectively perform splicing and convolution on the feature maps output from the backbone network, and finally use convolutional group C followed by a convolutional layer to obtain the coordinates, confidence levels, and category information of pedestrians, garbage bags, and trash cans in the figure.
[0059] Step (2.3) Network training: Input the training set in step (2.1) into the object detection network for training. During the training process, the learning rates of the weight layer, bias layer, and batch normalization layer are adjusted separately. At the beginning of the training, one-dimensional linear interpolation is used to update the learning rate. After the learning rate warm-up is completed, the cosine annealing algorithm is used to update the learning rate.
[0060] The cosine annealing algorithm is expressed as:
[0061]
[0062] where, new lr represents the newly obtained learning rate, initial lr represents the initial learning rate, eta minDenote the minimum learning rate as cur epoch Denote the current training epoch as T max Denote the total number of training epochs;
[0063] The network loss function is expressed as:
[0064]
[0065] where ρ() is the Euclidean distance, b and b gt represent the center point coordinates of the predicted target box and the actual target box respectively, c represents the diagonal distance of the minimum circumscribed rectangle of the predicted box B and the actual box B gt ; a represents the trade-off coefficient, u represents the aspect ratio consistency parameter, w gt and h gt represent the width and height of the actual target box, w and h represent the width and height of the predicted target box; IOU represents the intersection over union of the predicted box B and the actual box B gt of the intersection over union.
[0066] Step (2.4) Model testing: Input the images captured by the monitor, and output the coordinate information, category, and confidence of pedestrians, garbage bags, and trash cans.
[0067] Step (3), if a pedestrian p is detected in the t-th frame, then intercept all the garbage bag and pedestrian detection box images in the t-th frame where represents the i-th garbage bag detection box image in the t-th frame, represents the j-th pedestrian detection box image in the t-th frame; and take all the detection box images, detection box coordinate information, and categories of pedestrians and garbage bags in this frame as the input of the tracking module, and the tracking module assigns an id to each pedestrian and garbage bag;
[0068] The tracking module is implemented as follows:
[0069] Step (3.1) Feature extraction, input all the garbage bag and pedestrian detection box images in the t-th frame into a lightweight image feature extraction network model formed by stacking convolutional layers, and use the convolutional neural network to extract the image features of all the garbage bag and pedestrian detection box images, and each image obtains a 512-dimensional feature vector
[0070] The specific steps for establishing the lightweight image feature extraction network model are as follows:
[0071] Step (3.1.1) Data Preparation: Train the image feature extraction network using publicly available person re-identification datasets and self-made garbage bag re-identification datasets. Each person and garbage bag in the dataset has multiple images taken from different angles. The labeled data samples are divided into a training set, a validation set, and a test set at a ratio of 8:1:1 respectively.
[0072] Step (3.1.2) Network Structure Design: The backbone network takes an image as input. After passing through a convolutional layer with a 3X3 convolutional kernel, the resulting feature map is input into an inverted residual structure composed of a 1X1 convolution and a DepthWise convolution with a 3X3 convolutional kernel. Channel attention is formed through a pooling layer and a convolutional layer. After the feature map is input into the stacked inverted residual and attention structures, it finally outputs a vector of size 1X512 through a 1X1 convolutional layer, a pooling layer, and two fully connected layers.
[0073] Step (3.1.3) Network Training: Input multiple person images or garbage bag images to train a person image feature extraction model and a garbage bag image feature extraction model using the two types of images respectively. The network sets the learning rate to 10 -4 , and the learning rate is reduced to half of the original at each epoch. The Adam optimizer is used for optimization.
[0074] Step (3.2) Target Matching: Input the detected garbage bag and person detection box images in frames t+1, t+2, and t+3 into the feature extraction network in step (3.1) to obtain the corresponding feature vectors:
[0075]
[0076]
[0077]
[0078] Calculate the distance l between the feature vectors of two adjacent frames. If l is less than the threshold δ, then it is considered that and the persons in are the same person; To avoid the inability to match features caused by a person turning around or changing posture, the remaining detection boxes in the previous frame are predicted to their positions in the next frame through the Kalman filtering algorithm. If the intersection over union of the predicted detection box and the actual detection box in the next frame is greater than the threshold ε, then it is considered that the targets within these two detection boxes are the same target. Only targets that are matched for k consecutive frames will be assigned a unified id, and if a target with an existing id fails to match continuously for multiple times, it is considered that the target has left the camera monitoring range and the id is deleted from the data table.
[0079] Step (4): The detected person detection box image in frame t It is input into the human key point detection module, and the human key point detection module outputs the coordinate information of the key points of the human body in the image (key points such as the top of the head, neck, shoulders, chest, elbow joints, wrists, etc.);
[0080] The specific steps for establishing the human key point detection module are as follows:
[0081] Step (4.1) Data preparation: Train the model through a publicly available human key point dataset. The annotation information of the images in the dataset is the coordinate information and key point information of the pedestrians in the image (the center point coordinates of the detection box, length and width, and the key point coordinates such as the top of the head, neck, shoulders, chest, elbow joints, wrists, etc.); The labeled data samples are divided into a training set, a validation set, and a test set according to 8:1:1 respectively;
[0082] Step (4.2) Network structure design: The backbone network takes the image as input. After passing through a convolutional layer with a convolutional kernel size of 3X3, the resulting feature map is input into an inverted residual structure composed of a 1X1 convolution and a DepthWise convolution (DW convolution) with a convolutional kernel size of 3X3. Channel attention is formed through a pooling layer and a convolutional layer. The feature map is input into a stacked inverted residual and attention structure, and then a 1X1 convolutional layer, a pooling layer, and a convolutional layer for regressing the key point coordinates are added;
[0083] Step (4.3) Network training: Input the pedestrian images with annotations (annotations: the center point coordinates of the pedestrian detection box, the length and width of the detection box, and the key point coordinates such as the top of the head, neck, shoulders, chest, elbow joints, wrists, etc.) to train the human key point detection module. The network sets the learning rate to 10 -4 , and the learning rate is reduced to half of the original value in each epoch. The Adam optimizer is used for optimization.
[0084] Step (5), Calculate the intersection over union iou between all garbage bags and all pedestrians in the t-th frame. According to the size of the iou, determine whether a pedestrian is carrying a garbage bag into the monitoring area. If a pedestrian carrying a garbage bag is detected, add or update the pedestrian and garbage bag information (pedestrian id, garbage bag id, pedestrian position information, garbage bag position information, and pedestrian key point information) in the data table. If a pedestrian carrying a garbage bag is not detected, go to step (2);
[0085] The determination of whether a pedestrian is carrying garbage is carried out as follows:
[0086] Calculate the intersection over union iou between all garbage bags and all pedestrians in the t-th frame. If the intersection over union iou i of the garbage bag r j and the pedestrian p p,r is greater than the threshold α, and there is no record in the data table that the pedestrian p j is carrying the garbage bag r i, then it is considered that pedestrian p j Bring trash bags i Enter the video surveillance range and add pedestrians to the grid data table j Bring trash bags i Information (information includes: pedestrian p in the tth frame j With garbage bag i id, coordinate information and pedestrian p j key point coordinates), if pedestrian p already exists in the data table j Bring trash bags i If the information of the tth frame is available, the data table is updated with the information of the tth frame to replace the original information.
[0087] Step (6), after updating the data table, calculate the distance between each garbage bag carried by the pedestrian and the garbage bin in the tth frame and whether the pedestrian has made a garbage delivery action to determine whether the pedestrian has violated the regulations in garbage delivery;
[0088] The judgment of whether a pedestrian has committed illegal garbage delivery is carried out as follows:
[0089] After updating the data table, calculate the distance between the left and right wrists of each pedestrian in the tth frame and the garbage bag they carry to determine whether the pedestrian is holding the garbage bag in the left or right hand; if there is a pedestrian p j The right hand holds the garbage bag i , then calculate p j The angles d and p between the right arm and the body j With r i The intersection and union ratio of iou p,r and r i The intersection and union of the bin e in the tth frame is iou r,e , if d is greater than the threshold β and iou p,r Equal to 0 and iou r,e If it is greater than the threshold γ, then p j Committing illegal garbage delivery.
[0090] In this embodiment, when the target detection algorithm in step (2) fails to detect the pedestrian p in consecutive H frames, j When the pedestrian p j Has left the video surveillance range, delete the pedestrian p in the data table j (The data table only contains information about pedestrians and garbage bags in the picture).
[0091] The present invention utilizes partial community monitoring and deep learning algorithms to detect residents carrying plastic bags near garbage delivery booths in real time and determine whether they have violated the regulations in garbage delivery.
Claims
1. A method for detecting illegal garbage delivery behavior based on deep learning, Characterized in that: The specific steps are as follows: Step (1), obtain the images captured by the monitoring camera installed outside the garbage delivery kiosk during the non-garbage delivery time period; Step (2): Use a deep learning-based object detection algorithm to detect pedestrians in the received video frames , garbage bags and trash cans frame by frame, and obtain the detection box coordinate information and categories of pedestrians, garbage bags, and trash cans; Step (3), if a pedestrian is detected in the frame image , then intercept all the garbage bag and pedestrian detection box images in the frame image, where represents the th garbage bag detection box image in the frame, and represents the th pedestrian detection box image in the frame; And use the detection box images, detection box coordinate information, and categories of all pedestrians and garbage bags in this frame as the input of the tracking module. The tracking module assigns ; Step (4), input the pedestrian detection box image detected in the frame into the human key point detection module, and the human key point detection module outputs the coordinate information of the human key points in the image; Step (5), calculate the intersection over union of all garbage bags and all pedestrians in the nth frame , and based on the size of, determine whether a pedestrian enters the monitoring area while carrying a garbage bag. If a pedestrian carrying a garbage bag is detected, add or update the pedestrian and garbage bag information in the data table. If a pedestrian carrying a garbage bag is not detected, go to step (2); Step (6), after updating the data table, calculate the distance between each garbage bag carried by a pedestrian and a trash can in the frame, and whether the pedestrian makes a delivery action to determine whether the pedestrian has the behavior of illegally throwing garbage; Among them, the judgment of whether a pedestrian makes an illegal garbage delivery behavior is carried out as follows: After updating the data table, calculate the The distance between the left and right wrists of each pedestrian in the frame and the garbage bag they carry is used to determine whether the pedestrian is holding the garbage bag in his left or right hand. The right hand holds the garbage bag , then calculate The angle between the right arm and the body , and The intersection ratio and With Trash can in frame The intersection ratio ,like Greater than threshold and equal and Greater than threshold , then it is believed that Committing illegal garbage delivery.
2. A method for detecting illegal garbage delivery behavior based on deep learning according to claim 1, Characterized in that: When the target detection algorithm in step (2) fails to detect pedestrians for consecutive frames it is considered that the pedestrian has left the video surveillance range, and the information of the pedestrian in the data table is deleted .
3. A method for detecting illegal garbage delivery behavior based on deep learning according to claim 1 or 2, Characterized in that: The object detection algorithm based on deep learning in step (2) is trained as follows: Step (2.1) Data preparation: A monitoring camera installed outside the garbage delivery kiosk takes an image every ten seconds, and the image is manually annotated. The annotation information is the detection box information of pedestrians, garbage bags, and trash cans, that is , where represents the category, being 0 represents a pedestrian, being 1 represents a garbage bag, being 2 represents a trash can, represents the horizontal and vertical axis coordinates of the center point of the target, represents the width and height of the target; Divide the labeled data samples into a training set, a validation set, and a test set according to 8:1:1; Step (2.2) Network structure design, the object detection network takes the images captured by the monitoring camera as input, first performs FOCUS slicing operation on the input image, and then passes through multiple convolutional groups A followed by batch normalization and activation functions and convolutional group B containing residual components. After passing through the pooling group composed of convolutional and pooling layers, it is connected to convolutional group C and convolutional group A. In the backbone network, three detection branches of different scales are formed through skip connections. The three branches respectively splice and convolve the feature maps output from the backbone network, and finally use convolutional group C followed by a convolutional layer to obtain the coordinates, confidence levels, and category information of pedestrians, garbage bags, and trash cans in the figure; Step (2.3) Network training, input the training set in step (2.1) into the object detection network for training. During the training process, the learning rates of the weight layer, bias layer, and batch normalization layer are adjusted separately. At the beginning of the training, one-dimensional linear interpolation is used to update the learning rate. After the learning rate warm-up is completed, the cosine annealing algorithm is used to update the learning rate; Step (2.4) Model testing, input the images captured by the monitoring, and output the coordinate information, category, and confidence level of pedestrians, garbage bags, and trash cans.
4. A method for detecting illegal garbage delivery behavior based on deep learning according to claim 3, Characterized in that: The cosine annealing algorithm in step (2.3) is expressed as: ; Among them, represents the newly obtained learning rate, represents the initial learning rate, represents the minimum learning rate, represents the current number of training epochs, represents the total number of training epochs; The network loss function is expressed as: ; ; ; ; where is the Euclidean distance, and represent the center point coordinates of the predicted target box and the center point coordinates of the actual target box respectively, represents the predicted target box and the diagonal distance of the minimum bounding rectangle of the actual target box; represents the trade-off coefficient, represents the aspect ratio consistency parameter, and represent the width and height of the actual target box, and represent the width and height of the predicted target box; represents the intersection over union of the predicted target box and the actual target box.
5. A method for detecting illegal garbage delivery behavior based on deep learning according to claim 1, Characterized in that: The tracking module in step (3) is implemented as follows: Step (3.1) Feature extraction. Input the images of all garbage bags and pedestrian detection frames in the th frame into a lightweight image feature extraction network model formed by stacking convolutional layers, and use a convolutional neural network to extract the image features of all garbage bag and pedestrian detection frame images. Each image obtains a 512-dimensional feature vector. ; Step (3.2) Target matching. Input the garbage bag and pedestrian detection box images detected in the frame into the feature extraction network in step (3.1) to obtain the corresponding feature vectors: ; ; ; Calculate the distances between the feature vectors of adjacent two frames in sequence When and the distance is less than the threshold it is considered that and the pedestrians in are the same pedestrian; where represents the image feature vector in the th frame and the th predicted target box, that is the feature vector corresponding to the image; Predict the positions of the remaining detection boxes in the previous frame in the next frame through the Kalman filter algorithm. If the intersection over union of the predicted detection box and the actual detection box in the next frame is greater than the threshold it is considered that the targets in these two detection boxes are the same target.
6. A method for detecting illegal garbage delivery behavior based on deep learning according to claim 5, Characterized in that: The specific steps for establishing the lightweight image feature extraction network model in step (3.1) are as follows: Step (3.1.1) Data preparation, train the image feature extraction network through the publicly available pedestrian re-identification dataset and the self-made garbage bag re-identification dataset. Each pedestrian and garbage bag in the dataset has multiple images taken from different angles. Divide the labeled data samples into a training set, a validation set, and a test set according to 8:1:1 respectively; Step (3.1.2) Network structure design: The backbone network takes an image as input. After passing through a convolutional layer with a convolutional kernel size of 3X3, the resulting feature map is input into an inverted residual structure composed of a 1X1 convolution and a DepthWise convolution with a convolutional kernel size of 3X3. Channel attention is formed through a pooling layer and a convolutional layer. After the feature map is input into the stacked inverted residual and attention structure, it passes through a 1X1 convolutional layer, a pooling layer, and two fully connected layers, and finally outputs a vector with a size of 1X512. Step (3.1.3) Network training: Input multiple pedestrian images or garbage bag images, and use the two types of images to train a pedestrian image feature extraction model and a garbage bag image feature extraction model respectively.
7. A method for detecting illegal garbage delivery behavior based on deep learning according to claim 5, characterized in that: Only the targets that are continuously matched in step (3.2) will be assigned a unified , and if an existing target fails to match continuously for multiple times, it is considered that the target has left the camera monitoring range, and the target will be deleted from the data table.
8. A method for detecting illegal garbage delivery behavior based on deep learning according to claim 1, characterized in that: The specific steps for establishing the human key point detection module in step (4) are as follows: Step (4.1) Data preparation: Train the model through a publicly available human key point data set. The annotation information of the images in the data set is the coordinate information and key point information of the pedestrians in the image; The labeled data samples are divided into a training set, a validation set, and a test set according to 8:1:1 respectively; Step (4.2) Network structure design: The backbone network takes an image as input. After passing through a convolutional layer with a convolutional kernel size of 3X3, the resulting feature map is input into an inverted residual structure composed of a 1X1 convolution and a DepthWise convolution with a convolutional kernel size of 3X3. Channel attention is formed through a pooling layer and a convolutional layer. After the feature map is input into the stacked inverted residual and attention structure, a 1X1 convolutional layer, a pooling layer, and a convolutional layer for regressing the key point coordinates are added; Step (4.3) Network training: Input labeled pedestrian images to train the human key point detection module.
9. A method for detecting illegal garbage delivery behavior based on deep learning according to claim 1, characterized in that: Judging whether a pedestrian enters with garbage in step (5) is carried out as follows: Calculate the The intersection of all garbage bags and all pedestrians in the frame If the garbage bag With pedestrians The intersection ratio Greater than threshold , and there are no pedestrians in the data table Bring a garbage bag information, then the pedestrian is considered Bring a garbage bag Enter the video surveillance range and add pedestrians to the data table Bring a garbage bag If the pedestrian already exists in the data table Bring a garbage bag , then use the information The frame information updates the data table to replace the original information.
Citation Information
Patent Citations
Garbage throwing behavior real-time detection methods
CN111178182A
Garbage throwing behavior detection method based on urban management monitoring video
CN111611970A