Barrier object passing behavior identification method combining human body posture and interaction area detection
By combining the methods of human posture and interactive area detection, and using technologies such as example segmentation network and YOLOv7 backbone network, the problem of difficult to identify the behavior of the partitioned objects in rail transit security scenarios in the prior art is solved, and the online recognition effect with high accuracy is achieved.
Patent Information
- Application Number
- CN202510285958.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-27
AI Technical Summary
Existing behavior recognition algorithms based on deep learning are difficult to effectively identify the behavior of the partition objects in rail transit security scenarios, especially in the case of high traffic, occlusion and complex environment.
The partition object behavior recognition method combining human posture and interactive area detection is adopted, and the railing is automatically positioned through the instance segmentation network and the least squares fitting algorithm. The features are extracted using the YOLOv7 backbone network and the SPPCSPC module, and combined with the character detection, human posture detection and character interaction area detection modules, a delivery object behavior detector is constructed, and the joint loss function is used to optimize the model.
It realizes high-accuracy online identification of partition objects, improves the accuracy and efficiency of identification, and reduces the dependence on the recognition area of manual labeling behavior.
Smart Images

Figure CN120220227A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of rail transit security, and particularly relates to a method for identifying the behavior of passing objects across a railing by combining human posture and interaction area detection. Background Art
[0002] The safe operation of urban rail transit largely depends on the security inspection at stations. In the public area before security inspection, due to the high mobility and high openness of people, and this area has not been inspected yet, this area is a weak link in security prevention. The public area before security inspection belongs to the non-security inspection area and is often separated from the security inspection area by a fence. People may pass uninspected items at the fence for convenience. If someone passes prohibited items, it will pose a great potential safety hazard to the operation of the station. Therefore, the identification of the behavior of passing objects across the fence is an indispensable part of public safety in the rail transit industry.
[0003] In recent years, with the wide application of behavior recognition in the field of intelligent security and the rapid development of deep learning algorithms, behavior recognition algorithms based on deep learning have emerged. Compared with traditional methods, behavior recognition algorithms based on deep learning have the advantages of strong robustness and high accuracy. The mainstream methods of behavior recognition based on deep learning include two-stream convolutional networks, 3D convolutional networks, long short-term memory recurrent networks, and behavior recognition based on images. The advantage of these methods is that they can achieve end-to-end learning. However, due to the complex model structure and large computational amount, it is difficult to apply them in actual scenarios. In addition, due to factors such as large traffic flow, occlusion, and complex environment in the monitoring scenario, the difficulty of behavior recognition is further increased. Summary of the Invention
[0004] In view of the above-mentioned disadvantages and deficiencies of the prior art, the present invention provides a method for identifying the behavior of passing objects across a railing by combining human posture and interaction area detection. This method has a high recognition accuracy and can identify the behavior of passing objects across the railing online.
[0005] In order to achieve the above object, the main technical solutions adopted by the present invention include: A method for identifying the behavior of passing objects across a railing by combining human posture and interaction area detection, comprising the following steps: Step S01: Obtain a real-time video stream from a security high-definition camera F t ( t = 1, 2,... T ), T representing a time variable; Step S02: Use the deep neural network MaskRCNN model as a segmentation model to perform instance segmentation on the input image, extract the set of upper edge contour points of the railing from the segmentation map, and linearly fit the set of edge contour points by the least square method to obtain the linear representation equation of the railingL and the coordinates of both ends of the railing 、 , so as to automatically locate the actual position of the railing on the image; Step S03: After preprocessing the image, send it into a feature extraction detector to obtain feature maps of different resolutions; The feature extraction detector uses the YOLOv7 backbone network as the basic network for feature extraction, and sends the feature maps of different resolutions into the SPPCSPC module for feature fusion to obtain a reconstructed multi-scale feature map; Step S04: Set up a handing behavior detector, including a person detection module, a human body pose detection module, and a person interaction area detection module, and the three module branches execute in parallel; Step S05: Send the feature map obtained in Step S03 into the person detection module, the human body pose detection module, and the person interaction area detection module in Step S04 in parallel, and the three modules simultaneously perform detection processing on the multi-scale feature map; Obtain the t person detection frame at the moment, including the set of human detection frames and the set of object detection frames , where the human detection frame 、the object detection frame , where x 、 y 、 w 、 h respectively represent the horizontal and vertical coordinates of the center point of the detection frame on the feature map and the width and height of the detection frame, represents at t the confidence score of the p th human detection frame at the moment, represents at t the confidence score of the o th object detection frame at the moment, n ps represents the number of human detection frames, n ob represents the number of object detection frames; Through the human body pose detection module, the human handing pose detection frame is detected for the feature map, where represents the human handing pose detection frame at t the moment, n d represents the number of human handing pose detection frames, represents at t the confidence score of the d th human handing pose detection frame at the moment; Through the person interaction area detection module, the person interaction area detection frame is obtained, where represents at tMoment human interaction area detection box Indicates at t moment the in confidence score of the detection box of the th human interaction area; Step S06: Conduct a preliminary screening of the human detection boxes. According to the set of human object - handing pose detection boxes obtained by the human pose detection module in step S05 and the set of human detection boxes obtained by the human detection module, take out a object - handing pose detection box and a human detection box from the two sets respectively, eliminate the human detection boxes that do not meet the requirements, and combine the detection boxes of different types in the two sets pairwise to obtain all the human detection box sets satisfying the object - handing pose, l where n l represents the number of human detection boxes satisfying the object - handing pose; Step S07: Pair the human detection boxes output in step S06 pairwise according to the adjacent principle, and form teams with the object detection boxes and the human interaction area detection boxes obtained in step S05 to calculate the overlap degree of the interaction areas at different scales and . When the following formula is satisfied, it is the detection box of the object - handing behavior . Retain the detection results with a confidence score higher than 0.5; ; where represents the indicator function. The above formula contains four indicator functions. When the candidate box simultaneously satisfies , , , the four conditions, then , which means that the candidate box is the detection box of the object - handing behavior; represents l 1 true human box, represents l 2 true human box, represents the true box of the object, represents the true box of the object - handing interaction area, represents covering and the smallest box of the two boxes as the candidate box, represents the candidate box and the ratio of the intersection of to represents the candidate box The threshold of the ratio of the intersection of the real human body box and the real object box to the real object box Indicates the candidate box and The intersection and The threshold of the ratio; Step S08, calculate the delivery behavior detection box And the railing edge line automatically located in step S02 L Two intersection points And , x 1、 y 1 represents the abscissa and ordinate of the intersection point p 1, x 2、 y 2 represents the abscissa and ordinate of the intersection point p 2. Determine whether the intersection point coordinates are within the two end points of the railing, that is , , then it is recognized as an over-railing delivery detection behavior, where x a 、 y a Represents L The abscissa and ordinate of one end point of x b 、 y b Represents L The abscissa and ordinate of the other end point; Step S09, combine the delivery behavior detector in step S04 with YOLOv7 object detection to construct an over-railing delivery model based on human pose and human interaction area detection, and optimize and train the model by backpropagation of the loss function of the delivery behavior detection.
[0006] Furthermore, in step S02, when detecting the real-time video stream, the trained segmentation model and the video frame are sent into the inference network, the feature map is extracted through the ResNet-50 network, and then the features are sent into the feature pyramid network to obtain the fused feature map. The fused features of the region of interest ROI are extracted according to the generated candidate region coordinates, the ROI features of different scales are aligned through the ROIAlign alignment module, and finally the railing segmentation map is predicted and output by sending into the mask branch; the edge point set of the railing is extracted from the segmentation map , n Represents the number of railing edge points x 1, x 2,…, x n Represents the abscissa of the railing edge point y 1, y 2,…, y nDenote the ordinate of the railing edge point. Use the least squares method to linearly fit the set of upper edge points of the railing. For the linear model, construct the objective function using the sum of squared errors as follows: ; where F represents the objective function, y r denote the abscissa and ordinate of the r th edge point of the railing, w denotes the weight, b denotes the bias. Obtain the linear representation equation L of the upper edge line of the railing and the position coordinates of the two end points of the railing and 、 .
[0007] Furthermore, in step S04, the object delivery behavior detector is based on the YOLOv7 network structure. On the basis of the original human detection module, two branch modules of human pose detection and human interaction area detection are added. Both modules use two-dimensional convolution with a kernel size of 3×3 to extract the target spatial features and sigmoid activation operation for detection regression.
[0008] Furthermore, in step S04, when setting up the object delivery behavior detector, collect object delivery behavior pictures from the subway video stream and public datasets. Use the labelImg software to select rectangular boxes to annotate the object delivery behavior pictures to obtain the object delivery behavior detection dataset. The labeled categories and the positions of the rectangular boxes are used as the ground truth for network training. The labeled categories include human detection categories: person, backpack, handbag, suitcase, human pose categories: object delivery, standing, other, and object delivery interaction area categories.
[0009] Furthermore, in step S06, the specific method for removing the human detection boxes that do not meet the requirements is to calculate the intersection over union (IOU) according to the following formula , if then retain the human detection box, if then remove the human detection box.
[0010] Furthermore, in step S07, when there are m sets of ground truth boxes for human interaction areas matching the same candidate box b j , define the level index of the overlap degree according to the following formula to make the candidate box correspond to at most one object delivery behavior: ; where represents the i th ground truth box for human interaction area the real bounding box of the human body related to the interaction area 、 and the real bounding box of the object with the candidate bounding box The degree of overlap, when the k th degree of overlap is greater than the degree of overlap , that is then the candidate bounding box is associated with .
[0011] Furthermore, in step S09, the loss function for detecting the object-passing behavior is the sum of the joint loss functions of the loss functions of the human detection module, the human pose detection module, and the human interaction area detection module: ; where the loss function of the human detection module includes human body regression loss, object regression loss, and object category loss; the loss function of the human pose detection module is equal to the sum of the pose regression loss and the category loss; the above regression loss uses the L1 norm loss, and the category loss uses the standard binary cross-entropy loss; the loss function of the human interaction area detection module is the ignorance loss function based on focal loss for the object-passing interaction area, and eliminates the influence of missed unlabeled interactions by eliminating the background loss.
[0012] The beneficial effects of the present invention are: 1. The present invention uses an instance segmentation network model and a least squares fitting algorithm to achieve automatic and accurate positioning of the railing, without the need for manual annotation of the behavior recognition area; 2. The present invention uses spatial fine visual features and position relationship modeling, and based on the convolutional network combined with human spatial features, human pose spatial features, and human interaction spatial features, can effectively enhance the feature expression ability; 3. The present invention uses a multi-scale interaction area decision function to comprehensively evaluate different types of visual features, and at the same time sets the overlap level of the interaction area, which helps to correctly detect the interaction area; 4. The present invention uses a joint loss function, which can optimize the model from multiple perspectives of human targets, human poses, and interaction areas, promote the effective utilization of features and the transmission of information, and helps to improve the generalization ability and training efficiency of the object-passing detection model. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 is a diagram of the recognition result of the behavior of passing an object across a railing in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0014] For better explanation and understanding of the present invention, the present invention will be described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0015] The present invention provides a method for identifying the behavior of passing objects across a railing by combining human posture and interaction area detection, including the following steps: Step S01: Obtain a real-time video stream from a security high-definition camera F t ( t = 1, 2,... T ) T represents a time variable.
[0016] Step S02: Use the deep neural network MaskRCNN model as a segmentation model to perform instance segmentation on the input image, extract the set of upper edge contour points of the railing from the segmentation map, and linearly fit the set of edge contour points by the least squares method to obtain the linear representation equation of the railing L and the coordinates of the two endpoints of the railing , , so as to automatically locate the actual position of the railing in the image.
[0017] In step S02, for the task of identifying the behavior of passing objects across a railing, the accurate position information of the railing needs to be determined first. Different from the object detection task, detecting the railing requires accuracy to the pixel level and is achieved through instance segmentation. To achieve railing positioning, a railing instance segmentation model is trained using the MaskRCNN network with a railing dataset. The railing dataset uses the polygon annotation method of the labelme annotation software for instance annotation, and the annotation categories are set to two categories: railing and background.
[0018] In the same scene, if there is no sudden situation (such as the camera rotating), the position of the railing will not change. Therefore, when performing real-time detection, it is not necessary to locate the railing for each frame of the image, and the railing positioning can be freely timed. When detecting the real-time video stream, the video frame is sent into the instance segmentation model obtained in step S02, and the feature map is extracted through the ResNet-50 network. Then, through a series of operations such as convolution and upsampling, these low-level to high-level features are sent into the feature pyramid network to obtain a fused feature map. According to the generated candidate region coordinates, the fused features of the region of interest ROI are extracted, and the ROI features of different scales are aligned through the ROIAlign alignment module, and finally sent into the mask branch to predict and output to obtain the railing segmentation map. The set of edge points of the railing is extracted from the segmentation map , n represents the number of railing edge points, x 1, x 2,..., x n represents the abscissa of the railing edge point, y 1,y 2, …, y n Denote the ordinate of the railing edge point. Use the least squares method to linearly fit the set of upper edge points of the railing. For the linear model, construct the objective function using the sum of squared errors as follows: ; where F represents the objective function, y r denotes the abscissa and ordinate of the r th edge point of the railing, w denotes the weight, b denotes the bias. Obtain the linear representation equation of the upper edge line L of the railing and the position coordinates of the two end points of the railing , .
[0019] By taking the derivative of the objective function F and setting it equal to zero, the weight w and the bias b can be obtained. In this way, the linear representation equation L of the upper edge line of the railing can be obtained. Substitute the abscissas x a and x b of the two end points of the railing into the linear representation equation to obtain the ordinates y a , y b。
[0020] Step S03: After preprocessing the image, send it into the feature extraction detector. The preprocessing includes normalization, resizing the image, image enhancement, etc., to obtain feature maps of different resolutions. The feature extraction detector uses the YOLOv7 backbone network as the basic network for feature extraction, and sends the feature maps of different resolutions into the SPPCSPC module for feature fusion to obtain the reconstructed multi-scale feature maps.
[0021] Step S04: Set up the object delivery behavior detector, which includes a person detection module, a human pose detection module, and a person interaction area detection module. The three modules share the feature map and are executed in parallel. The object delivery behavior detector uses the YOLOv7 network structure as the basic architecture, and adds two branch modules of human pose detection and person interaction area detection on the basis of the original person detection module. Both modules use a two-dimensional convolution with a kernel size of 3×3 to extract the target spatial features and perform sigmoid activation operations for detection and regression.
[0022] In the person detection module, interactions are not considered, and it is responsible for detecting a single human body or object. The human body pose detection module plays an auxiliary role in interaction classification. It predicts the object-passing actions of the human body, which helps to associate human interaction behaviors. The person interaction area detection module is an important branch of the over-the-barrier object-passing behavior recognition. It uses the subtle visual features of the interaction area to directly predict the detection box of the object-passing interaction area.
[0023] In step S04, when setting up the object-passing behavior detector, collect object-passing behavior pictures from the subway video stream and the public dataset. Use the labelImg software to select rectangular boxes to annotate the object-passing behavior pictures as the ground truth for network training. The annotation includes person detection categories: person, backpack, handbag, suitcase, etc., human body pose categories: object-passing, standing, others, and object-passing interaction area categories.
[0024] Step S05: Parallelly send the feature maps obtained in step S03 into the person detection module, human body pose detection module, and person interaction area detection module in step S04. The three modules simultaneously perform detection processing on the multi-scale feature maps; obtain the t person detection box at a certain moment , including the set of human body detection boxes and the set of object detection boxes , where the human body detection box , the object detection box , where x , y , w , h respectively represent the horizontal and vertical coordinates of the center point of the detection box on the feature map and the width and height of the detection box. represents at t the confidence score of the p th human body detection box at a certain moment, represents at t the confidence score of the o th object detection box at a certain moment, n ps represents the number of human body detection boxes, n ob represents the number of object detection boxes. Through the human body pose detection module, the human body object-passing pose detection box is detected for the feature map, where represents the human body object-passing pose detection box at t a certain moment, n d represents the number of human body object-passing pose detection boxes, represents at t the dThe confidence score of the human handover gesture detection frame. The human interaction area detection frame is obtained through the human interaction area detection module. ,in Indicated in t The detection frame of the character interaction area at each moment, Indicated in t Moment in The confidence score of the detection box of the interaction area of each person, , , , The value is 0.5.
[0025] Step S06: Preliminary screening of human body detection frames, based on the human body gesture detection frame set obtained by the human body gesture detection module in step S05 And the human detection frame set obtained by the human detection module , take out a handover gesture detection box from each of the two sets and a human detection frame , calculate the Intersection over Union (IOU) according to the following formula: ,like Then keep the human detection frame, if The human body detection frame is removed. The different types of detection frames in the two sets are combined two by two to obtain all the human body detection frame sets that meet the handover posture. , Indicates l Individual object pose human body detection frame, n l Indicates the number of human body detection frames in the object delivery posture; through the above screening of human body detection frames, the detection speed and accuracy can be improved.
[0026] Step S07: Pair the human body detection frames output from step S06 according to the adjacent principle, and team them with the object detection frame and the human interaction area detection frame obtained in step S05. , calculate the interaction areas at different scales and The overlap degree satisfies the following formula, which is the object delivery behavior detection frame , retain the detection results with confidence scores higher than 0.5; ; in Represents the indicator function. The above formula contains four indicator functions. When the candidate box satisfies , , , When the four conditions are met, , which indicates that the candidate box is a delivery behavior detection box; Indicates l 1 human body true box, Indicates l 2 human body true boxes, Indicates the true box of the object, Indicates the true box of the delivery interaction area, Indicates covering And The minimum box of the two boxes is used as the candidate box, Indicates the candidate box And The intersection with The threshold of the ratio, Indicates the candidate box The threshold of the ratio of the intersection of the candidate box and the human body true box to the object true box, Indicates the candidate box and The intersection with The threshold of the ratio.
[0027] When there is m A set of true boxes of personal interaction areas Matching the same candidate box b j At this time, the level index of the overlap degree is defined by the following formula, so that the candidate box corresponds to at most a unique delivery behavior: ; Where Indicates the i th true box of the personal interaction area And the human body true box related to this interaction area , And the object true box The overlap degree with the candidate box . When the k th overlap degree is greater than overlap degree , that is, then the candidate box is associated with .
[0028] Step S08, calculate the two intersection points of the delivery behavior detection box L and the railing edge line automatically located in step S02 and , x 1, y 1 indicates the abscissa and ordinate of the intersection point p 1, x 2, y2 represents the intersection point p The abscissa and ordinate of 2 are used to determine whether the intersection point coordinates are within the two end points of the railing, that is , , then it is recognized as an object-passing behavior detection across the railing, where x a 、 y a represents L the abscissa and ordinate of one end point of x b 、 y b represents L the abscissa and ordinate of the other end point of
[0029] Step S09: Combine the object-passing behavior detector in step S04 with YOLOv7 object detection to construct a model for detecting object-passing across the railing based on human pose and human interaction area detection, and optimize and train the model by backpropagation using the loss function of object-passing behavior detection
[0030] In the above step S09, the loss function of object-passing behavior detection is the sum of the combined loss functions of the loss functions of the human detection module, the human pose detection module, and the human interaction area detection module: ; Among them, the loss function of the human detection module includes human regression loss, object regression loss, and object class loss; the loss function of the human pose detection module is equal to the sum of the pose regression loss and the class loss; the above regression losses use the L1 norm loss, and the class loss uses the standard binary cross-entropy loss; the loss function of the human interaction area detection module is an ignorance loss function based on focal loss for the object-passing interaction area, which eliminates the influence of missed unlabeled interactions by eliminating background losses, that is, the detection boxes related to non-interaction areas will not take effect during learning, to solve the serious imbalance problem between the number of missed positive samples in the interaction area and the number of missed samples of human and pose during training. To solve the above problem, the losses of non-dominant non-interaction areas are ignored, and the class label of the candidate box needs to meet the following conditions: ; When the candidate box is an interaction behavior detection box and meets the condition, the class label of the candidate box is 1. When the candidate box is an interaction behavior detection box but not the largest , then the class label of the candidate box is 0. When the candidate box is a non-interaction behavior detection box, the loss of the candidate box is ignored. The above ignorance loss function formula is as follows: ; where α and γ represent balance factors. γ can reduce the loss contribution of easy samples while increasing the loss weight of hard samples. p represents the probability of the model predicting the behavior of passing objects, and its value range is 0-1. In the experiment, α = 0.25 and γ = 2.
[0031] To verify the effectiveness of the model of the present invention, the rail transit scene images are sent into the model for recognition, and the visualization effect is as Figure 1 shown. Figure 1 Shows the recognition result of the behavior of passing objects across the railing, indicating that the model can accurately recognize the behavior of passing objects across the railing.
[0032] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Modifications, alterations, substitutions, and variations made by those of ordinary skill in the art to the above embodiments all fall within the scope of the present invention.
Claims
1. A method for identifying barrier object delivery behavior by combining human posture and interaction area detection, characterized in that: The steps include: Step S01: Obtain real-time video stream from security HD camera F t ( t =1,2,... T ), T represents the time variable; Step S02: Use the deep neural network MaskRCNN model as the segmentation model to segment the input image instance, extract the upper edge contour point set of the railing from the segmentation map, linearly fit the edge contour point set through the least squares method, and obtain the linear representation equation of the railing L And the coordinates of the two end points of the railing , , thereby automatically locating the actual position of the railing on the image; Step S03, after preprocessing the image, send it to the feature extraction detector to obtain feature maps of different resolutions; the feature extraction detector uses the YOLOv7 backbone network as the basic network for extracting features, and sends the feature maps of different resolutions to the SPPCSPC module for feature fusion to obtain a reconstructed multi-scale feature map; Step S04, setting an object handover behavior detector, including a person detection module, a human posture detection module, and a person interaction area detection module, and the three module branches are executed in parallel; Step S05: The feature map obtained in step S03 is sent to the person detection module, the human posture detection module, and the person interaction area detection module in step S04 in parallel. The three modules simultaneously detect and process the multi-scale feature map; the person detection module obtains t People detection frame at all times , including the human detection box set And object detection box collection , where the human body detection box , object detection frame ,in x , y , w , h Respectively represent the horizontal and vertical coordinates of the center point of the detection frame on the feature map and the width and height of the detection frame, Indicated in t Moment p The confidence score of the individual human detection box, Indicated in t Moment o The confidence score of the object detection box, n ps Indicates the number of human detection frames, n ob Indicates the number of object detection frames; Through the human posture detection module, the feature map is detected to obtain the human handover posture detection frame ,in Indicated in t The human handover posture detection frame at all times, n d Indicates the number of human handover gesture detection frames, Indicated in t Moment d The confidence score of the human handover gesture detection frame; the human interaction area detection frame is obtained through the human interaction area detection module ,in Indicated in t The detection frame of the character interaction area at each moment, Indicated in t Moment in Confidence score of the detection box of each person’s interaction area; Step S06: Preliminary screening of human body detection frames, based on the human body gesture detection frame set obtained by the human body gesture detection module in step S05 And the human detection frame set obtained by the human detection module , take out a handover gesture detection box from each of the two sets and a human detection frame , remove the human body detection frames that do not meet the requirements, and combine the different types of detection frames in the two sets to obtain all the human body detection frames that meet the handover posture. , Indicates l Individual object pose human body detection frame, n l Indicates the number of human body detection frames in the object-handling posture; Step S07: Pair the human body detection frames output from step S06 according to the adjacent principle, and team them with the object detection frame and the human interaction area detection frame obtained in step S05. , calculate the interaction areas at different scales and The overlap degree satisfies the following formula, which is the object delivery behavior detection frame , retain the detection results with confidence scores higher than 0.5; ; in Represents the indicator function. The above formula contains four indicator functions. When the candidate box satisfies , , , When the four conditions are met, , which indicates that the candidate box is a delivery behavior detection box; express l 1 Human body real frame, express l 2 Human body real frame, represents the ground-truth box of the object, represents the real box of the object-handling interaction area, Indicates coverage and The minimum box of the two boxes is used as the candidate box. Represents a candidate box and Intersection and The threshold value of the ratio, Represents a candidate box The threshold of the ratio of the intersection of the human body real frame and the object real frame, Represents the candidate box and Intersection and The threshold value of the ratio of Step S08: Calculate the object delivery behavior detection frame The edge line of the railing automatically located in step S02 L The two intersection points and , x 1. y 1 indicates the intersection p 1's horizontal and vertical coordinates, x 2. y 2 represents the intersection p 2, and determine that the coordinates of the intersection point are within the two end points of the railing, that is, , , it is identified as a barrier object delivery detection behavior, where x a , y a express L The horizontal and vertical coordinates of one end point, x b , y b express L The horizontal and vertical coordinates of the other endpoint; Step S09, combining the object delivery behavior detector of step S04 with YOLOv7 target detection, constructing a barrier object delivery model based on human posture and character interaction area detection, and back-propagating the optimized training model with the loss function of the object delivery behavior detection.
2. The method for identifying barrier object delivery behavior by combining human posture and interaction area detection according to claim 1, characterized in that: In step S02, when detecting the real-time video stream, the trained segmentation model and the video frame are sent to the inference network, the feature map is extracted through the ResNet-50 network, and then the feature is sent to the feature pyramid network to obtain a fused feature map, and the fused features of the region of interest ROI are extracted according to the generated candidate region coordinates, and the ROI features of different scales are aligned through the ROIAlign alignment module, and finally sent to the mask branch prediction output to obtain the railing segmentation map; the edge point set of the railing is extracted from the segmentation map , n Indicates the number of railing edge points, x 1, x 2,…, x n Indicates the horizontal coordinate of the edge point of the railing, y 1, y 2,…, y n Represents the ordinate of the edge point of the railing. The least squares method is used to linearly fit the edge point set on the railing. For the linear model, the sum of squares of the errors is used to construct the objective function, as shown in the following formula: ; Where F represents the objective function, y r Indicates the railing r The horizontal and vertical coordinates of the edge points are w represents the weight, b Represents the bias, and the upper edge line of the railing is obtained by least squares fitting L The linear representation equation And the coordinates of the two end points of the railing , .
3. The method for identifying barrier object delivery behavior by combining human posture and interaction area detection according to claim 1, characterized in that: In step S04, the object delivery behavior detector uses the YOLOv7 network structure as the basic architecture, and adds two branch modules of human posture detection and human interaction area detection on the basis of the original human detection module. Both modules use a two-dimensional convolution with a convolution kernel size of 3×3 to extract target space features and a sigmoid activation operation for detection regression.
4. The method for identifying barrier object delivery behavior by combining human posture and interaction area detection according to claim 1, characterized in that: The step S04 is to set up an object delivery behavior detector, collect object delivery behavior pictures from subway video streams and public data sets, use labelImg software to select rectangular boxes to mark the object delivery behavior pictures, and obtain an object delivery behavior detection data set. The marked categories and rectangular box positions are used as the true values of network training. The marked categories include person detection categories: people, backpacks, handbags, suitcases, human posture categories: object delivery, standing, others, and object delivery interaction area categories.
5. The method for identifying barrier object delivery behavior by combining human posture and interaction area detection according to claim 1, characterized in that: The step S06 of eliminating the human body detection frames that do not meet the requirements is specifically to calculate the intersection-over-union (IOU) ratio according to the following formula: ,like Then keep the human detection frame, if Then remove the human detection frame.
6. The method for identifying barrier object delivery behavior by combining human posture and interaction area detection according to claim 1, characterized in that: In step S07, when there is m A collection of real frames of the interaction area of the characters Match the same candidate box b j When , the overlap level index is defined as follows, so that the candidate boxes correspond to at most unique delivery behaviors: ; in Indicates i The real frame of the character interaction area And the real human frame related to the interaction area , and the object's true frame With candidate box The overlap degree is k indivual Overlap Greater than Overlap ,Right now The candidate box and Make an association.
7. The method for identifying barrier object delivery behavior by combining human posture and interaction area detection according to claim 1, characterized in that: In step S09, the loss function of the object delivery behavior detection is the sum of the loss function of the person detection module, the loss function of the human posture detection module and the loss function of the person interaction area detection module. ; The loss function of the human detection module Including human regression loss, object regression loss and object category loss; Human posture detection module loss function =Equal to the sum of the posture regression loss and the category loss; the above regression loss uses the L1 norm loss, and the category loss uses the standard binary cross entropy loss; the loss function of the character interaction area detection module An ignorant loss function based on focal loss is used for the inter-object interaction regions, and the influence of missed unlabeled interactions is eliminated by removing the background loss.