A method for identifying jumping and squatting ticket evasion behaviors

By combining human key point extraction, multi-objective tracking, machine learning and time series analysis, jumping and squatting ticket evasion behaviors in subway stations are identified and monitored, and the problems of low identification efficiency and frequent misjudgment in the existing technology are solved, achieving more efficient monitoring and security guarantees.

CN114038056BActive Publication Date: 2025-05-06TONGJI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111276839.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-29
Publication Date
2025-05-06
Estimated Expiration
2041-10-29

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively identify and monitor jumping and squatting ticket evasion behaviors in subway stations, resulting in low monitoring efficiency, frequent misjudgment and misjudgment, and cannot meet the demand for video surveillance.

Method used

The human body key point extraction method, multi-objective tracking method, machine learning and time series algorithm is used to analyze the monitoring video of the subway station, and detect and identify jumping and squatting ticket evasion behaviors. Specific steps include obtaining surveillance video, detecting human skeleton data, performing multi-object tracking, feature extraction and action category detection.

Benefits of technology

Accurate identification and monitoring of jumping and squatting ticket evasion behaviors is achieved, errors and misjudgments of manual monitoring are reduced, and the safety and operation order of subway stations are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114038056B_ABST
    Figure CN114038056B_ABST
Patent Text Reader

Abstract

The present invention provides a method for identifying jumping and squatting fare evasion behaviors, including obtaining a surveillance video of the location of the subway gate entrance and converting it into an image frame; using a human key point detection algorithm to detect each frame of the image and obtain the human skeleton data of each frame of the image; using a multi-target tracking algorithm to perform multi-target tracking on pedestrians in consecutive multi-frame images to obtain a human skeleton sequence of each pedestrian; extracting features from the skeleton data at a certain moment in the human skeleton sequence; building a detection model for single-frame behaviors, inputting the obtained human skeleton features into the detection model, and obtaining the action category of the pedestrian; repeating the steps to obtain a sequence curve of the action category and time at each moment, and detecting whether the fare evasion behavior occurs based on the curve. Through this invention, jumping and squatting fare evasion behaviors in subway stations can be identified, and passengers' fare evasion behaviors can be captured and stopped in time, ensuring subway order and safe operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of rail transit, and in particular to a method for identifying jumping and squatting fare evasion behaviors. Background Art

[0002] The role of subway in people's daily life and work is becoming more and more important. It carries a large number of people every day, and its safety issues are of top priority. Any hidden dangers and dangers in subway stations will cause mass panic and easily spread, causing personal and property losses to people, and have an adverse impact on the operation of the entire subway line. At present, with the rapid development of computers, network transmission technology and image processing, video surveillance equipment has become a public area security infrastructure, and a large number of front-end cameras are distributed in traffic arteries and public areas with a large flow of people. In order to ensure the safety of subway stations, a large number of front-end cameras are naturally installed inside subway stations. With video surveillance, we can more conveniently monitor the internal situation of the station and leave video records.

[0003] Due to the lack of a sufficiently intelligent real-time monitoring system for abnormal behavior, most of these tasks are completed manually, and most monitoring systems can only provide video playback after the incident. Since people cannot fully focus on monitoring the video for a long time, they are often negligent due to human factors during real-time monitoring. For the post-event evidence call, this work requires a lot of human effort to check all the monitoring video data. This work not only consumes a lot of manpower costs, but also there will be a lot of misjudgments and missed judgments in the results of manual monitoring. Using manual methods to analyze video data cannot meet people's needs for video monitoring.

[0004] On the one hand, fare evasion is a common abnormal behavior in subway stations. Using computer vision technology to intelligently identify it will strongly crack down on fare evasion, enhance the rule of law and moral consciousness of the whole society, and have high practical application value and social significance. Fare evasion, a common abnormal behavior in subway stations, not only harms the public interest, but also brings chaos to the entire subway order and social civilization. The emergence of fare evasion is not only due to the low level of civilization of passengers themselves, but also due to lax supervision and insufficient punishment of subways. However, during rush hour and subway stations with large passenger flow, it is difficult for subway staff to supervise the entire subway station. If computer technologies such as image processing and machine learning can be used to identify pedestrians in subway station surveillance videos in real time and accurately capture fare evasion in the video, it can guide staff to stop it in time; even if it is not stopped in time, the identity of the fare evader can be verified by combining facial recognition technology in the later stage and included in the personal integrity record of citizens.

[0005] On the other hand, among the many abnormal behaviors in subway stations, the diversity and complexity of fare evasion make it difficult to identify. According to the different exit gates of subway stations, it can be divided into three types: jumping, squatting and trailing. Among them, jumping and squatting fare evasion behaviors involve the body movements of pedestrians, which cannot be identified only by target detection of pedestrians, and require in-depth analysis of the body movements of pedestrians; trailing fare evasion behaviors involve the interaction between two pedestrians, and only analyzing the state of a single pedestrian is not enough. Therefore, the identification of fare evasion behavior includes both individual abnormal behavior identification and interactive abnormal behavior identification, which is extremely challenging. At the same time, a large number of people often gather near the subway exit, making the occlusion between pedestrians more serious. These problems greatly increase the difficulty of video-based fare evasion behavior identification. Therefore, this patent takes fare evasion as an example to explore the abnormal behavior identification method in subway stations, and provides a reference for the identification of abnormal behaviors such as illegal intrusion into restricted areas and panic running. Summary of the invention

[0006] The present invention provides a method for identifying jumping and squatting fare evasion behaviors. Based on the video data of fare evasion behaviors in subway stations, a human body key point extraction method, a multi-target tracking method, and an algorithm combining machine learning and time series are used to detect and identify autonomous fare evasion behaviors in subway stations. Passengers' fare evasion behaviors can be captured in time through monitoring videos and stopped, thereby ensuring subway order and safe operation.

[0007] A method for identifying jumping and squatting ticket evasion behaviors, comprising the following steps:

[0008] (1) Obtain surveillance video of the subway gate entrance and convert it into image frames;

[0009] (2) Using the human key point detection algorithm to detect each frame of the image, obtain the human skeleton data of each frame of the image;

[0010] (3) Using a multi-target tracking algorithm to track pedestrians in multiple consecutive frames of images, obtain a human skeleton sequence for each pedestrian;

[0011] (4) performing feature extraction on the skeleton data of the human skeleton sequence at a certain moment in step (3);

[0012] (5) Building a single-frame behavior detection model, inputting the human skeleton features obtained in step (4) into the detection model to obtain the pedestrian's action category;

[0013] (6) Repeat steps (4) to (5) with the skeleton data of each moment of the skeleton sequence in step (3) to obtain a sequence curve of action category and time at each moment, and detect whether the fare evasion behavior occurs based on the curve.

[0014] Furthermore, the human body key point detection algorithm described in step (2) is the OpenPose algorithm, and the specific steps are as follows:

[0015] (2a): The first stage network generates a set of detection confidence maps S 1 =ρ 1 (F) and a set of joint affine field maps where ρ 1 and They are the CNN structures of the first stage. The input of each subsequent stage comes from the prediction results of the previous stage and the original image features F, producing more accurate prediction results.

[0016] (2b): According to the predicted confidence map, a discrete set of candidate joint points is obtained, and the formula is as follows:

[0017]

[0018] in,

[0019] represents the position of the mth joint point of the jth joint point, where j represents the joint point category;

[0020] (2c): Define the line set of all candidate joint points, the formula is as follows:

[0021]

[0022] in,

[0023] Candidate joint points Represents the position of the mth joint point of the j1th joint point;

[0024] Candidate joint points Represents the position of the nth joint point of the j2th joint point;

[0025] Represents candidate joint points and candidate joint points Is there a wired connection between them? Indicates that the two are connected by wire. Indicates that there is no connection between the two;

[0026] (2d): Consider all two candidate joint points j corresponding to limb c separately 1 and j 2 ,The optimal problem of graph matching with the highest total affinity value is solved by the Hungarian algorithm and the optimal matching ,is obtained. The optimization problem with the maximum total affinity is as follows:

[0027]

[0028] in,

[0029] E c represents affinity value, c represents each limb of human;

[0030] E mn Represents candidate joint points and candidate joint points Affinity between

[0031] (2e): When estimating the full body pose of multiple people, the K-point graph allocation problem is solved to assemble the links that share the same body parts into the full body pose of the human body. The formula is as follows:

[0032]

[0033] Among them, E represents the affinity value, c represents each human limb, and C represents the number of humans;

[0034] (2f): Set rules to clean the obtained human skeleton data. The formula is as follows:

[0035]

[0036] in,

[0037] Indicates the validity of the corresponding human skeleton data;

[0038] Indicates the number of joint points with a confidence greater than 0.2 in the human pose estimation result of the i-th pedestrian in the t-th frame image;

[0039] num J is the joint point number threshold.

[0040] Furthermore, the OpenPose algorithm described in step (2) uses the BODY_25 human skeleton model to detect key points of the human body and connect and group limbs to obtain human skeleton data.

[0041] Furthermore, the multi-target tracking algorithm described in step (3) is a Pose Flow algorithm, and the specific steps are as follows:

[0042] (3a) Establish a pedestrian candidate set, where each pedestrian in the candidate set includes the following attributes: pedestrian detection box position information, pedestrian detection box confidence, joint point position information, joint point confidence, pedestrian number, and last appearance time;

[0043] (3b) Feature extraction based on human posture data;

[0044] (3b1) Define the intersection-over-union (IOU) ratio between the detection boxes of pedestrians box , the formula is as follows:

[0045]

[0046] in,

[0047] A and B are the areas where the two pedestrian detection frames are located;

[0048] (3b2) Define the intersection-over-union (IOU) ratio of the joint detection boxes between pedestrians pose_box The specific calculation process is as follows: With the joint point as the center, a square with a side length of l is constructed as the detection frame of the joint point; the intersection and union ratio of the joint point detection frames of the two pedestrians corresponding to the joint points is calculated. Where i = 0, 1, ..., 24; the 25 joint point detection frame intersection and union ratio calculation results are sorted from large to small according to the confidence of the joint point detection results, and the corresponding elements with a confidence of 0 are directly eliminated; the average of the first num joint point detection frame intersection and union ratios after sorting is taken as the joint point detection frame intersection and union ratio IOU between pedestrian C1 and pedestrian P1 pose_box ;

[0049] (3b3) Define the intersection-over-union ratio (ORB) of the detection box feature points between pedestrians box , the formula is as follows:

[0050]

[0051] in,

[0052] Kp A represents the set of ORB feature points in the detection box A, Kp B represents the set of ORB feature points in the detection box B, KP A ∩Kp B Indicates the set of ORB feature points that are successfully matched between detection box A and detection box B, num(Kp A ∩Kp B ) represents the set Kp A ∩Kp B The number of feature points in ;

[0053] (3b4) Define the intersection-and-joint ratio (ORB) of the feature points of the joint detection boxes between pedestrians pose_box The calculation process is as follows: take the joint point as the center, construct a square with a side length of l as the detection frame of the joint point; calculate the intersection and union ratio of the feature points of the joint point detection frames of the two pedestrians corresponding to the joint points Where i = 0, 1, ..., 24; the calculated results of the intersection-and-union ratios of the feature points of the 25 joint point detection frames are sorted from large to small according to the confidence of the joint point detection results, and the corresponding elements with a confidence of 0 are directly eliminated; the average of the intersection-and-union ratios of the feature points of the first num joint point detection frames after sorting is taken as the intersection-and-union ratio of the feature points of the joint point detection frames between pedestrians (ORB) pse_box ;

[0054] (3c) Set the pedestrian matching evaluation index S, whose formula is as follows:

[0055]

[0056] in,

[0057] w 1 、w 2 、w 3 、w 4 Feature IOU box , IOU pose_box 、ORB box 、ORB pose_box The weight of

[0058] (3d) Pedestrian matching: The Hungarian algorithm is used to match the pedestrians detected in the current frame with the pedestrians in the pedestrian candidate set according to the cost matrix. The pedestrians in the current frame are matched with the pedestrians in the pedestrian candidate set based on the matching principle of the maximum matching index S of all pedestrians.

[0059] Assume that there are M pedestrians in the current frame and N pedestrians in the pedestrian candidate set. The matching index between the pedestrian Ci in the current frame and the pedestrian Pj in the pedestrian candidate set is s ij , we can get the pedestrian matching index matrix S:

[0060]

[0061] Use k ij represents the matching status of the pedestrian Ci in the current frame and the pedestrian Pj in the pedestrian candidate set, and then the pedestrian matching status matrix K can be obtained. The matrix K also contains M rows and N columns. ij The assignment rules are as follows:

[0062]

[0063] The pedestrian matching problem is transformed into the following optimization problem:

[0064]

[0065] By solving the above model through the Hungarian algorithm, the pedestrian matching state matrix K can be obtained, that is, the matching result between the pedestrian in the current frame and the pedestrian in the pedestrian candidate set can be obtained;

[0066] (3e) Pedestrian candidate set update, the content is mainly divided into the following three parts:

[0067] (3e1) If the pedestrian in the pedestrian candidate set matches the pedestrian in the current frame successfully, the pedestrian information in the current frame is used to replace the corresponding pedestrian information in the pedestrian candidate set. For example, in the above example, pedestrian C1 in the current frame matches pedestrian P1 in the pedestrian candidate set successfully, then the pedestrian detection box position information, pedestrian detection box confidence, joint point position information and other attributes of pedestrian C1 are used to replace the corresponding attributes of pedestrian P1;

[0068] (3e2) If there is a pedestrian Px in the pedestrian candidate set that has not been successfully matched with the pedestrian in the current frame, then it can be divided into two cases: ① Pedestrian Px has not been successfully matched for a long time, for example, pedestrian Px has not updated its status for 100 frames or 5s in a row. At this time, the pedestrian may have left the camera's monitoring area, so the pedestrian should be removed from the pedestrian candidate set; ② Pedestrian Px has not been successfully matched for a short period of time, which may be due to the target detection algorithm. The pedestrian is just not detected in the current frame. At this time, the attributes of pedestrian Px should be kept unchanged;

[0069] (3e3) If there is a pedestrian Cx in the current frame that is not successfully matched with any pedestrian in the pedestrian candidate set, this may be because this pedestrian enters the camera's monitoring area for the first time and there is no information about this pedestrian in the pedestrian candidate set. Therefore, this pedestrian needs to be added to the pedestrian candidate set, and its pedestrian number is the maximum pedestrian number that has not been added to the previous pedestrian candidate set + 1.

[0070] Furthermore, the characteristics of step (4) are mainly three characteristics: relative distance relationship, angle relationship and relative speed relationship, specifically:

[0071] (4a) Relative distance relationship: The distance relationship is used to describe the distance in space between any pair of joints in the human skeleton. It is specifically quantified by the Euclidean distance. The relative position relationship RD(i,j) is defined as:

[0072]

[0073] Where i represents the joint point J i , j represents the joint point J j ;

[0074] (x i ,y i ) and (x j ,y j) represent the joint point J i and joint point J j The position coordinates at the same time satisfy the conditions of i,j=0,1,…,14 and i≠j; dis_shoulder represents the left shoulder joint point J of the pedestrian at the same time 5 Right shoulder joint point J 2 Half the distance between the shoulders, or half the width of the shoulders;

[0075] (4b) Angular relationships: limbs With limbs The angle between Defined as:

[0076]

[0077] Among them, (x i ,y i )、(x j ,y j )、(x e ,y e )、(x f ,y f ) represent the joint point J i , J j , J e , J f Position coordinates at the same moment;

[0078] (4c) Relative speed relationship: Speed ​​is an important indicator of motion. It can provide key dynamic information for pedestrian motion recognition. It is usually measured by the joints. The average speed over time is taken as the instantaneous speed of the joint point at time t. i The trajectory of motion L i for:

[0079] L i ={(x i,0 ,y i,0 ),…,(x i,t-1 ,y i,t-1 ),(x i,t ,y i,t )}

[0080] Among them, (x i,t ,y i,t ) is the joint point J i Position coordinates at time t;

[0081] Joint point J i The instantaneous speed at time t for:

[0082]

[0083] Using the method used to calculate the relative distance relationship, the instantaneous speed will be obtained Divide by the pedestrian's half shoulder width dis_shoulder at time t t , convert it into a relative value, recorded as v i,t :

[0084]

[0085] Joint point J i At time t, along the spine The relative speed V i,t for:

[0086]

[0087] Furthermore, the pedestrians in step (5) are classified into four categories, namely normal state, jumping state, squatting state and transition state. In the normal state, the limbs of pedestrians are in a relatively relaxed and natural state, marked as "0"; pedestrians in the jumping state have the characteristics of arms stretched out and the joints of the lower limbs moving up quickly, marked as "1"; pedestrians in the squatting state are in a curled up state as a whole, marked as "2"; pedestrians in the transition state have some characteristics of the normal state and some characteristics of the abnormal state (jumping or squatting), and even the human eye cannot accurately judge their state, marked as "3".

[0088] Furthermore, the single-frame behavior detection model described in step (5) is a random forest model.

[0089] Furthermore, the random forest model is a parallel ensemble learning method based on decision trees. It makes the final decision by voting on the classification results of multiple independent decision trees. The randomness of random forests mainly consists of two aspects: random sample combinations and random feature sets. The random forest algorithm randomly selects a certain number of feature sets and a certain number of sample combinations from the training samples as training subsets to ensure that the training data used by each decision tree is different, but also overlaps with each other. This randomness also makes random forests have stronger generalization capabilities.

[0090] Furthermore, the time series curve described in step (6) is obtained by using human posture estimation and pedestrian tracking to obtain the skeleton sequence of the pedestrian. Each human skeleton in the human skeleton sequence is used as the input of the ticket evasion behavior detection model to obtain the change of the pedestrian state over time, that is, the time series curve of the pedestrian state.

[0091] Furthermore, in step (6), whether fare evasion occurs is detected based on the curve, and the steps are as follows:

[0092] (6a) One-hot encoding is performed on the time series. One-hot encoding uses n-bit binary code to represent n states, and only one bit is valid at any time, that is, only one bit has a value of 1 and the rest are 0;

[0093] Assume that the recognition result of the pedestrian state, that is, the time series of the pedestrian state is Where T represents the length of the pedestrian's time series, represents the prediction result of the model at time t; there are three types of pedestrian states, then The value of is:

[0094]

[0095] Since there are three types of pedestrian states, at least a 2-bit unique hot code can be used to represent the pedestrian state. Then the prediction result of the model at time t is It is represented by one-hot encoding in and is defined as follows:

[0096]

[0097] Through the above one-hot encoding method, the time series Y of the pedestrian status (pred) It becomes two time series O 1 and O 2 ;

[0098]

[0099] (6b) Perform Kalman filtering on the encoded time series. The Kalman filter includes two parts: prediction and update. The prediction part is defined as follows:

[0100]

[0101] Among them, A represents the state transfer matrix, which is set to 1. represents the optimal estimate at time t-1, B represents the gain of the optional control input u, and u t-1 represents the control gain at time t-1, represents the predicted value at time t; P t-1 represents the posterior estimated covariance at time t-1, Q is the covariance matrix of the predicted noise, and P′ t represents the prior estimated covariate at time t;

[0102] The update part of the Kalman filter is defined as follows:

[0103]

[0104] Among them, H is the measurement matrix, K t is the Kalman gain, R is the covariance matrix of the observation noise; y t is the observed value at time t, is the optimal estimate at time t; I is the unit matrix, P t represents the posterior estimated covariance at time t;

[0105] Using Kalman filter algorithm to analyze time series O 1 and O 2 Processing is performed to obtain the time series O_kf 1 and O_kf 2 ;

[0106]

[0107] (6c) Perform threshold segmentation on the filtered time series; O_kf is segmented by setting the threshold. 1 and O_kf 2 Discretization is performed, and the specific method is as follows:

[0108]

[0109] Among them, o_kf t is the optimal estimate obtained by the Kalman filter at time t, o_th t is the result obtained after the optimal estimate at time t is segmented by the threshold, and o_threshold is the threshold;

[0110] Through the above threshold segmentation method, the time series o_th of discrete variable type is obtained 1 and O_th 2 , recorded as:

[0111]

[0112] (6d) Perform one-hot inverse transformation on the segmented time series; according to the defined one-hot encoding rules, 1 and O_th 2 Perform the inverse transformation to obtain the pedestrian status time series Y (pred) The result after Kalman filtering The specific conversion rules are as follows:

[0113]

[0114] in, Represents the state of the pedestrian at time t after being processed by the Kalman filter.

[0115] The beneficial effects of the present invention are:

[0116] (1) A jumping and autonomous fare evasion behavior recognition method is proposed to address the shortcomings of manual fare evasion behavior recognition;

[0117] (2) Overcoming the interference of clothing and complex background;

[0118] (3) Make full use of spatiotemporal features to avoid the misjudgment caused by the fare evasion detection model based on single-frame skeleton information. BRIEF DESCRIPTION OF THE DRAWINGS

[0119] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0120] Figure 1 It is a flow chart of the method for identifying jumping and squatting ticket evasion behaviors according to the technical solution of the present invention;

[0121] Figure 2 is a schematic diagram of an OpenPose BODY_25 human skeleton model according to an embodiment of the present invention;

[0122] Figure 3 is a flowchart of pedestrian tracking based on the PoseFlow algorithm according to an embodiment of the present invention;

[0123] Figure 4 is a schematic diagram of data calibration results of a normal state, a squatting state, a jumping state and a transition state according to an embodiment of the present invention;

[0124] Figure 5 is the pedestrian posture feature importance ranking result (top 30) according to an embodiment of the present invention;

[0125] Figure 6 It is a flow chart of pedestrian state recognition based on time series;

[0126] Figure 7 It is the result of time series processing such as Kalman filtering. DETAILED DESCRIPTION

[0127] The technical solution of the present invention will be clarified below in conjunction with the accompanying drawings. It is obvious that the described embodiments are part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0128] In the description of the present invention, it should be noted that the terms “first”, “second” and “third” are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.

[0129] In addition, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0130] The present invention provides a method for identifying jumping and squatting ticket evasion behaviors, such as Figure 1 As shown, the process includes the following steps:

[0131] (1) Obtain surveillance video of the subway gate entrance and convert it into image frames;

[0132] (2) Using the human key point detection algorithm to detect each frame of the image, obtain the human skeleton data of each frame of the image;

[0133] (3) Using a multi-target tracking algorithm to track pedestrians in multiple consecutive frames of images, obtain a human skeleton sequence for each pedestrian;

[0134] (4) performing feature extraction on the skeleton data of the human skeleton sequence at a certain moment in step (3);

[0135] (5) Building a single-frame behavior detection model, inputting the human skeleton features obtained in step (4) into the detection model to obtain the pedestrian's action category;

[0136] (6) Repeat steps (4) to (5) with the skeleton data of each moment of the skeleton sequence in step (3) to obtain a sequence curve of action category and time at each moment, and detect whether the fare evasion behavior occurs based on the curve.

[0137] Through the above steps, based on the collected video data of fare evasion at subway gates, the above steps can accurately identify the status of pedestrians in complex environments, multi-person interactions, and scarce samples, thereby ensuring the stable order of subway stations and the personal safety of passengers.

[0138] Combine the following Figures 2 to 7 The optional embodiments of steps (1) to (6) of the present invention are described in detail.

[0139] The BODY_25 human skeleton model using the OpenPose algorithm described in step (2) detects key points of the human body and connects and groups limbs to obtain human skeleton data, and the specific steps are as follows:

[0140] (2a): The first stage network generates a set of detection confidence maps S 1 =ρ 1(F) and a set of joint affine field maps where ρ 1 and They are the CNN structures of the first stage. The input of each subsequent stage comes from the prediction results of the previous stage and the original image features F, producing more accurate prediction results.

[0141] (2b): According to the predicted confidence map, a discrete set of candidate joint points is obtained, and the formula is as follows:

[0142]

[0143] in,

[0144] represents the position of the mth joint point of the jth joint point, where j represents the joint point category;

[0145] (2c): Define the line set of all candidate joint points, the formula is as follows:

[0146]

[0147] in,

[0148] Candidate joint points Represents the position of the mth joint point of the j1th joint point;

[0149] Candidate joint points Represents the position of the nth joint point of the j2th joint point;

[0150] Represents candidate joint points and candidate joint points Is there a wired connection between them? Indicates that the two are connected by wire. Indicates that there is no connection between the two;

[0151] (2d): Consider all two candidate joint points j corresponding to limb c separately 1 and j 2 ,The optimal problem of graph matching with the highest total affinity value is solved by the Hungarian algorithm and the optimal matching ,is obtained. The optimization problem with the maximum total affinity is as follows:

[0152]

[0153] in,

[0154] E c represents affinity value, c represents each limb of human;

[0155] E mn Represents candidate joint points and candidate joint points Affinity between

[0156] (2e): When estimating the full body pose of multiple people, the K-point graph allocation problem is solved to assemble the links that share the same body parts into the full body pose of the human body. The formula is as follows:

[0157]

[0158] Among them, E represents the affinity value, c represents each human limb, and C represents the number of humans;

[0159] (2f): Set rules to clean the obtained human skeleton data. The formula is as follows:

[0160]

[0161] in,

[0162] Indicates the validity of the corresponding human skeleton data;

[0163] Indicates the number of joint points with a confidence greater than 0.2 in the human pose estimation result of the i-th pedestrian in the t-th frame image;

[0164] num J is the joint point number threshold.

[0165] Figure 2 This is a schematic diagram of the OpenPose BODY_25 human skeleton model.

[0166] The Pose Flow algorithm is used in step (3) to track pedestrians and obtain a human skeleton sequence. The specific steps are as follows:

[0167] (3a) Establish a pedestrian candidate set, where each pedestrian in the candidate set includes the following attributes: pedestrian detection box position information, pedestrian detection box confidence, joint point position information, joint point confidence, pedestrian number, and last appearance time;

[0168] (3b) Feature extraction based on human skeleton posture data.

[0169] (3b1) Define the intersection-over-union (IOU) ratio between the detection boxes of pedestrians box , the formula is as follows:

[0170]

[0171] in,

[0172] A and B are the areas where the two pedestrian detection frames are located;

[0173] (3b2) Define the intersection-over-union (IOU) ratio of the joint detection boxes between pedestrians pose_box The specific calculation process is as follows: With the joint point as the center, a square with a side length of l is constructed as the detection frame of the joint point; the intersection and union ratio of the joint point detection frames of the two pedestrians corresponding to the joint points is calculated. Where i = 0, 1, ..., 24; the 25 joint point detection frame intersection and union ratio calculation results are sorted from large to small according to the confidence of the joint point detection results, and the corresponding elements with a confidence of 0 are directly eliminated; the average of the first num joint point detection frame intersection and union ratios after sorting is taken as the joint point detection frame intersection and union ratio IOU between pedestrian C1 and pedestrian P1 pose_box .

[0174] (3b3) Define the intersection-over-union ratio (ORB) of the detection box feature points between pedestrians box , the formula is as follows:

[0175]

[0176] in,

[0177] Kp A represents the set of ORB feature points in the detection box A, Kp B represents the set of ORB feature points in the detection box B, Kp A ∩Kp B Indicates the set of ORB feature points that are successfully matched between detection box A and detection box B, num(Kp A ∩Kp B ) represents the set Kp A ∩Kp B The number of feature points in ;

[0178] (3b4) Define the intersection-and-joint ratio (ORB) of the feature points of the joint detection boxes between pedestrians pose_box The calculation process is as follows: take the joint point as the center, construct a square with a side length of l as the detection frame of the joint point; calculate the intersection and union ratio of the feature points of the joint point detection frames of the two pedestrians corresponding to the joint points Where i = 0, 1, ..., 24; the calculated results of the intersection-and-union ratios of the feature points of the 25 joint point detection frames are sorted from large to small according to the confidence of the joint point detection results, and the corresponding elements with a confidence of 0 are directly eliminated; the average of the intersection-and-union ratios of the feature points of the first num joint point detection frames after sorting is taken as the intersection-and-union ratio of the feature points of the joint point detection frames between pedestrians (ORB) pose_box .

[0179] (3c) Set the pedestrian matching evaluation index S, whose formula is as follows:

[0180]

[0181] in,

[0182] w 1 、w 2 、w 3 、w 4 Feature IOU box , IOU pose_box 、ORB box 、ORB pose_box The weight of

[0183] (3d) Pedestrian matching: The Hungarian algorithm is used to match the pedestrians detected in the current frame with the pedestrians in the pedestrian candidate set according to the cost matrix. The pedestrians in the current frame are matched with the pedestrians in the pedestrian candidate set based on the matching principle of the maximum matching index S of all pedestrians.

[0184] Assume that there are M pedestrians in the current frame and N pedestrians in the pedestrian candidate set. The matching index between the pedestrian Ci in the current frame and the pedestrian Pj in the pedestrian candidate set is s ij , we can get the pedestrian matching index matrix S:

[0185]

[0186] Use k ij represents the matching status of the pedestrian Ci in the current frame and the pedestrian Pj in the pedestrian candidate set, and then the pedestrian matching status matrix K can be obtained. The matrix K also contains M rows and N columns. ij The assignment rules are as follows:

[0187]

[0188] The pedestrian matching problem is transformed into the following optimization problem:

[0189]

[0190] By solving the above model through the Hungarian algorithm, the pedestrian matching state matrix K can be obtained, and the pedestrian matching results between the pedestrians in the current frame and the pedestrian candidate set can be obtained.

[0191] (3e) Pedestrian candidate set update, the content is mainly divided into the following three parts:

[0192] (3e1) If the pedestrian in the pedestrian candidate set matches the pedestrian in the current frame successfully, the pedestrian information in the current frame is used to replace the corresponding pedestrian information in the pedestrian candidate set. For example, in the above example, pedestrian C1 in the current frame matches pedestrian P1 in the pedestrian candidate set successfully, then the pedestrian detection box position information, pedestrian detection box confidence, joint point position information and other attributes of pedestrian C1 are used to replace the corresponding attributes of pedestrian P1;

[0193] (3e2) If there is a pedestrian Px in the pedestrian candidate set that has not been successfully matched with the pedestrian in the current frame, then it can be divided into two cases: ① Pedestrian Px has not been successfully matched for a long time, for example, pedestrian Px has not updated its status for 100 frames or 5s in a row. At this time, the pedestrian may have left the camera's monitoring area, so the pedestrian should be removed from the pedestrian candidate set; ② Pedestrian Px has not been successfully matched for a short period of time, which may be due to the target detection algorithm. The pedestrian is just not detected in the current frame. At this time, the attributes of pedestrian Px should be kept unchanged;

[0194] (3e3) If there is a pedestrian Cx in the current frame that is not successfully matched with any pedestrian in the pedestrian candidate set, this may be because this pedestrian enters the camera's monitoring area for the first time and there is no information about this pedestrian in the pedestrian candidate set. Therefore, this pedestrian needs to be added to the pedestrian candidate set, and its pedestrian number is the maximum pedestrian number that has not been added to the previous pedestrian candidate set + 1.

[0195] Figure 3 This is a flowchart of pedestrian tracking based on the PoseFlow algorithm.

[0196] In the present invention, the three posture characteristics of relative distance relationship, angle relationship and relative speed relationship defined in step (4) are specifically:

[0197] (4a) Relative distance relationship. The distance relationship is used to describe the distance in space between any pair of joints in the human skeleton. It is specifically quantified by the Euclidean distance. The relative position relationship RD(i,j) is defined as:

[0198]

[0199] Where i represents the joint point J i , j represents the joint point J j ;

[0200] (x i ,y i ) and (x j ,y j ) represent the joint point J i and joint point J j The position coordinates at the same time satisfy the conditions of i,j=0,1,…,14 and i≠j; dis_shoulder represents the left shoulder joint point J of the pedestrian at the same time 5 Right shoulder joint point J 2 Half the distance between the shoulders, that is, half the shoulder width.

[0201] (4b) Angle relationship. With limbs The angle between Defined as:

[0202]

[0203] Among them, (x i ,y i )、(x j ,y j )、(x e ,y e )、(x f ,y f ) represent the joint point J i , J j , J e , J f The position coordinates at the same moment.

[0204] (4c) Relative speed. Speed ​​is an important indicator of motion. It can provide key dynamic information for pedestrian motion recognition. It is usually measured by the joints. The average speed over time is taken as the instantaneous speed of the joint point at time t. i The trajectory of motion L i for:

[0205] L i ={(x i,0 ,y i,0 ),…,(x i,t-1 ,y i,t-1 ),(x i,t ,y i,t )}

[0206] Among them, (x i,t ,y i,t ) is the joint point J i The position coordinates at time t.

[0207] Joint point J i The instantaneous speed at time t for:

[0208]

[0209] Using the method used to calculate the relative distance relationship, the instantaneous speed will be obtained Divide by the pedestrian's half shoulder width dis_shoulder at time t t , convert it into a relative value, recorded as v i,t :

[0210]

[0211] In the detection of fare evasion at subway stations, we mainly distinguish between jumping fare evasion and squatting fare evasion. In jumping fare evasion, pedestrians will have an upward acceleration along the spine; while in squatting fare evasion, pedestrians will have a downward acceleration along the spine. In this scenario, it is more practical to calculate the speed of pedestrians along the spine. Then the joint point J i At time t, along the spine The relative speed V i,t for:

[0212]

[0213] In the present invention, the states of pedestrians are divided into four categories, namely normal state, jumping state, squatting state and transition state. In the normal state, the limbs of pedestrians are in a relatively relaxed and natural state, marked as "0"; pedestrians in the jumping state have the characteristics of arms stretched out and the joints of the lower limbs moving up quickly, marked as "1"; pedestrians in the squatting state are in a curled up state as a whole, marked as "2"; pedestrians in the transition state have some characteristics of the normal state and some characteristics of the abnormal state (jumping or squatting), even the human eye sometimes cannot accurately judge its state, marked as "3", and the data marked as the transition state will not participate in the subsequent model training.

[0214] Figure 4 It is a schematic diagram of the data calibration results of the normal state, squatting state, jumping state and transition state.

[0215] In the present invention, the jumping and squatting fare evasion behavior detection model based on machine learning described in step (5) is a random forest model.

[0216] In the present invention, the importance of pedestrian posture features is ranked, and the "Gini index" is used as the standard for decision tree selection and partitioning attributes. The purity of the data set D is expressed by the Gini index as follows:

[0217]

[0218] Among them, p k It represents the proportion of samples of the kth category in the data set D (k=1,2,…,|y|).

[0219] Generally speaking, Gini(D) reflects the probability that two samples are randomly drawn from the data set D and their categories are inconsistent. Therefore, the smaller Gini(D), the higher the purity of the data set D.

[0220] Assume there are m features, denoted as {x 1 ,x 2 ,…,x m}. Using feature x jThe change in the Gini index before and after the decision tree node t branches is used to represent its importance on the decision tree node t, denoted as

[0221]

[0222] Among them, Gini t Represents the Gini index of node t, Gini l and Gini r They represent the Gini index of the two new nodes after node t branches out.

[0223] Features j The nodes appearing in decision tree i constitute a set T, then x j The importance of the i-th tree is:

[0224]

[0225] If there are n decision trees in the random forest model, then feature x j The importance in this random forest is:

[0226]

[0227] After normalizing the result, feature x j The importance in this random forest is:

[0228]

[0229] Figure 5 It is the top 30 ranking results of pedestrian posture feature importance.

[0230] In the present invention, the steps of pedestrian state recognition based on time series are as follows:

[0231] (6a) One-hot encoding is performed on the time series. One-hot encoding uses n-bit binary code to represent n states, and only one bit is valid at any time, that is, only one bit has a value of 1 and the rest are 0.

[0232] Assume that the recognition result of the pedestrian state, that is, the time series of the pedestrian state is Where T represents the length of the pedestrian's time series, represents the prediction result of the model at time t. In this paper, there are three types of pedestrian states: The value of is:

[0233]

[0234] Since there are three types of pedestrian states, at least a 2-bit unique hot code can be used to represent the pedestrian state. Then the prediction result of the model at time t is It is represented by one-hot encoding in and is defined as follows:

[0235]

[0236] Through the above one-hot encoding method, the time series Y of the pedestrian status (pred) It becomes two time series O 1 and O 2 .

[0237]

[0238] (6b) Perform Kalman filtering on the encoded time series. The Kalman filter includes two parts: prediction and update. The prediction part is defined as follows:

[0239]

[0240] Among them, A represents the state transfer matrix, which is set to 1. represents the optimal estimate at time t-1, B represents the gain of the optional control input u, and u t-1 represents the control gain at time t-1, represents the predicted value at time t; P t-1 represents the posterior estimated covariance at time t-1, Q is the covariance matrix of the predicted noise, and P′ t represents the prior estimated covariance at time t.

[0241] The update part of the Kalman filter is defined as follows:

[0242]

[0243] Among them, H is the measurement matrix, K t is the Kalman gain, R is the covariance matrix of the observation noise; y t is the observed value at time t, is the optimal estimate at time t; I is the unit matrix, P t represents the posterior estimated covariance at time t.

[0244] Use the Kalman filter algorithm to get the time series O in 4.4.1 1 and O 2 Processing is performed to obtain the time series O_kf 1 and O_kf 2 .

[0245]

[0246] (6c) Perform threshold segmentation on the filtered time series. 1 and O_kf 2 Discretization is performed, and the specific method is as follows:

[0247]

[0248] Among them, o_kf t is the optimal estimate obtained by the Kalman filter at time t, o_th t is the result obtained after threshold segmentation of the optimal estimate at time t, o_threshold is the threshold, which is set to 0.5 in this paper.

[0249] Through the above threshold segmentation method, the time series O_th of discrete variable type is obtained 1 and O_th 2 , recorded as:

[0250]

[0251] (6d) Perform one-hot encoding inverse transformation on the segmented time series. 1 and O_th 2 Perform the inverse transformation to obtain the pedestrian status time series Y (pred) The result after Kalman filtering The specific conversion rules are as follows:

[0252]

[0253] in, Represents the state of the pedestrian at time t after being processed by the Kalman filter.

[0254] Figure 6 It is a flowchart of pedestrian state recognition based on time series.

[0255] Figure 7 It is the result of time series processing such as Kalman filtering.

[0256] After the above steps are performed, the jumping and squatting fare evasion behaviors at the subway station can be accurately identified and detected successfully.

[0257] It is known in the art that abnormal behaviors such as fare evasion at subway stations occur from time to time, threatening the safety of passengers and the operating order of subway stations. Subway stations are equipped with a large number of surveillance cameras, which can realize all-round, no-dead-angle, and full-time monitoring of subway stations. Using manpower to identify abnormal behaviors in surveillance videos often results in misjudgments and omissions. According to the method for identifying jumping and squatting fare evasion described in the method of the present invention, artificial intelligence can be used to automatically identify and detect squatting and jumping fare evasion behaviors, thereby ensuring the operating order of subway stations and the safety of passengers themselves.

[0258] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0259] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0260] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0261] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1The steps for the functions specified in one or more boxes.

[0262] Obviously, the above embodiments are merely examples for the purpose of clear explanation, and are not intended to limit the implementation methods. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the implementation methods here. The obvious changes or modifications derived therefrom are still within the scope of protection of the invention.

Claims

1. A method for identifying jumping and squatting ticket evasion behaviors, characterized in that: The following steps are involved: (1) Obtain surveillance video of the subway gate entrance and convert it into image frames; (2) Using the human key point detection algorithm to detect each frame of the image, obtain the human skeleton data of each frame of the image; (3) Using a multi-target tracking algorithm to track pedestrians in multiple consecutive frames of images, obtain a human skeleton sequence for each pedestrian; (4) performing feature extraction on the skeleton data of the human skeleton sequence at a certain moment in step (3); (5) Building a single-frame behavior detection model, inputting the human skeleton features obtained in step (4) into the detection model to obtain the pedestrian's action category; (6) Repeating steps (4) to (5) with the skeleton data of each moment of the skeleton sequence in step (3) to obtain a sequence curve of action category and time at each moment, and detecting whether the fare evasion behavior occurs based on the curve; The steps of pedestrian state recognition based on time series are as follows: (6a) One-hot encoding of the time series: One-hot encoding uses n-bit binary code to represent n states, and only one bit is valid at any time, that is, only one bit has a value of 1 and the rest are 0; Assume that the recognition result of the pedestrian state, that is, the time series of the pedestrian state is Where T represents the length of the pedestrian's time series, It represents the prediction result of the model at time t. There are three types of pedestrian states. The value of is: Since there are three types of pedestrian states, at least a 2-bit unique hot code is used to represent the pedestrian state. The prediction result of the model at time t is It is represented by one-hot encoding in and is defined as follows: Through the above one-hot encoding method, the time series Y of the pedestrian status (pred) It becomes two time series O 1 and O 2 ; (6b) Kalman filtering is performed on the encoded time series. The Kalman filter includes two parts: prediction and update. The prediction part is defined as follows: Among them, A represents the state transfer matrix, which is set to 1. represents the optimal estimate at time t-1, B represents the gain of the optional control input u, and u t-1 represents the control gain at time t-1, represents the predicted value at time t; P t-1 represents the posterior estimated covariance at time t-1, Q is the covariance matrix of the predicted noise, and P t ′ represents the prior estimated covariance at time t; The update part of the Kalman filter is defined as follows: Among them, H is the measurement matrix, K t is the Kalman gain, R is the covariance matrix of the observation noise; y t is the observed value at time t, is the optimal estimate at time t; I is the unit matrix, P t represents the posterior estimated covariance at time t; Using Kalman filter algorithm to analyze time series O 1 and O 2 Processing is performed to obtain the time series O_kf 1 and O_kf 2 ; (6c) Perform threshold segmentation on the filtered time series; O_kf is segmented by setting the threshold. 1 and O_kf 2 Discretization is performed, and the specific method is as follows: Among them, o_kf t is the optimal estimate obtained by the Kalman filter at time t, o_th t is the result obtained after the optimal estimate at time t is segmented by the threshold, and o_threshold is the threshold; Through the above threshold segmentation method, the time series O_th of discrete variable type is obtained 1 and O_th 2 , recorded as: (6d) Perform one-hot inverse transformation on the segmented time series; according to the defined one-hot encoding rules, 1 and O_th 2 Perform the inverse transformation, and then get the pedestrian state time series as Y (pred) The result after Kalman filtering The specific conversion rules are as follows: in, Represents the state of the pedestrian at time t after being processed by the Kalman filter.

2. The method for identifying jumping and squatting ticket evasion behaviors according to claim 1, characterized in that: The human key point detection algorithm described in step (2) is the OpenPose algorithm, and the specific steps are as follows: (2a): The first stage network generates a set of detection confidence maps S 1 =ρ 1 (F) and a set of joint affine field maps where ρ 1 and They are the CNN structures of the first stage. The input of each subsequent stage comes from the prediction results of the previous stage and the original image features F, producing more accurate prediction results. (2b): According to the predicted confidence map, a discrete set of candidate joint points is obtained, and the formula is as follows: in, represents the position of the mth joint point of the jth joint point, where j represents the joint point category; (2c): Define the line set of all candidate joint points, the formula is as follows: in, Candidate joint points Represents the position of the mth joint point of the j1th joint point; Candidate joint points Represents the position of the nth joint point of the j2th joint point; Represents candidate joint points and candidate joint points Is there a wired connection between them? Indicates that the two are connected by wire. Indicates that there is no connection between the two; (2d): Consider the two candidate joint points j1 and j2 corresponding to the limb c separately, and solve the optimal problem of the graph matching method with the highest total affinity value through the Hungarian algorithm to obtain the optimal matching. The optimization problem with the maximum total affinity is as follows: in, E c represents affinity value, c represents each limb of human; E mn Represents candidate joint points and candidate joint points Affinity between (2e): When estimating the full body pose of multiple people, the K-point graph allocation problem is solved to assemble the links that share the same body parts into the full body pose of the human body. The formula is as follows: Among them, E represents the affinity value, c represents each human limb, and C represents the number of humans; (2f): Set rules to clean the obtained human skeleton data. The formula is as follows: in, Indicates the validity of the corresponding human skeleton data; Indicates the number of joint points with a confidence greater than 0.2 in the human pose estimation result of the i-th pedestrian in the t-th frame image; num j is the joint point number threshold.

3. The method for identifying jumping and squatting ticket evasion behaviors according to claim 2, characterized in that: The OpenPose algorithm described in step (2) uses the BODY_25 human skeleton model to detect key points of the human body and connect and group limbs to obtain human skeleton data.

4. The method for identifying jumping and squatting ticket evasion behaviors according to claim 1, characterized in that: The multi-target tracking algorithm described in step (3) is the Pose Flow algorithm, and the specific steps are as follows: (3a) Establish a pedestrian candidate set, where each pedestrian in the candidate set includes the following attributes: pedestrian detection box position information, pedestrian detection box confidence, joint point position information, joint point confidence, pedestrian number, and last appearance time; (3b) Feature extraction based on human posture data; (3b1) Define the intersection-over-union (IOU) ratio between the detection boxes of pedestrians box , the formula is as follows: in, A and B are the areas where the two pedestrian detection frames are located; (3b2) Define the intersection-over-union (IOU) ratio of the joint detection boxes between pedestrians pose_box The specific calculation process is as follows: With the joint point as the center, a square with a side length of l is constructed as the detection frame of the joint point; the intersection and union ratio of the joint point detection frames of the two pedestrians corresponding to the joint points is calculated. Where i = 0, 1, ..., 24; the 25 joint point detection frame intersection and union ratio calculation results are sorted from large to small according to the confidence of the joint point detection results, and the corresponding elements with a confidence of 0 are directly eliminated; the average of the first num joint point detection frame intersection and union ratios after sorting is taken as the joint point detection frame intersection and union ratio IOU between pedestrian C1 and pedestrian P1 pose_box ; (3b3) Define the intersection-over-union ratio (ORB) of the detection box feature points between pedestrians box , the formula is as follows: in, Kp A represents the set of ORB feature points in the detection box A, Kp B represents the set of ORB feature points in the detection box B, Kp A ∩Kp B Indicates the set of ORB feature points that are successfully matched between detection box A and detection box B, num(Kp A ∩Kp B ) represents the set Kp A ∩Kp B The number of feature points in ; (3b4) Define the intersection-and-joint ratio (ORB) of the feature points of the joint detection boxes between pedestrians pose_box The calculation process is as follows: take the joint point as the center, construct a square with a side length of l as the detection frame of the joint point; calculate the intersection and union ratio of the feature points of the joint point detection frames of the two pedestrians corresponding to the joint points Where i = 0, 1, ..., 24; the calculated results of the intersection-and-union ratios of the feature points of the 25 joint point detection frames are sorted from large to small according to the confidence of the joint point detection results, and the corresponding elements with a confidence of 0 are directly eliminated; the average of the intersection-and-union ratios of the feature points of the first num joint point detection frames after sorting is taken as the intersection-and-union ratio of the feature points of the joint point detection frames between pedestrians (ORB) pose_box ; (3c) Set the pedestrian matching evaluation index S, whose formula is as follows: in, w1, w2, w3, w4 are feature IOUs respectively box , IOU pose_box 、ORB box 、ORB pose_box The weight of (3d) Pedestrian matching: The Hungarian algorithm is used to match the pedestrians detected in the current frame with the pedestrians in the pedestrian candidate set according to the cost matrix. The pedestrians in the current frame are matched with the pedestrians in the pedestrian candidate set based on the matching principle of the maximum matching index S of all pedestrians. Assume that there are M pedestrians in the current frame and N pedestrians in the pedestrian candidate set. The matching index between the pedestrian Ci in the current frame and the pedestrian Pj in the pedestrian candidate set is s ij , then we get the pedestrian matching index matrix S: Use k ij represents the matching status of the pedestrian Ci in the current frame and the pedestrian Pj in the pedestrian candidate set, and then obtains the pedestrian matching status matrix K, which also contains M rows and N columns; where k ij The assignment rules are as follows: The pedestrian matching problem is transformed into the following optimization problem: The above model is solved by the Hungarian algorithm to obtain the pedestrian matching state matrix K, that is, the matching result between the pedestrian in the current frame and the pedestrian in the pedestrian candidate set is obtained; (3e) Pedestrian candidate set update, which is divided into the following three parts: (3e1) If the pedestrian in the candidate set matches the pedestrian in the current frame, the pedestrian information in the current frame is used to replace the corresponding pedestrian information in the candidate set; (3e2) If there is a pedestrian Px in the pedestrian candidate set that has not been successfully matched with the pedestrian in the current frame, there are two cases: ① Pedestrian Px has not been successfully matched for a long time, and the pedestrian is removed from the pedestrian candidate set; ② Pedestrian Px has not been successfully matched for a short time, and the attributes of pedestrian Px remain unchanged; (3e3) If there is a pedestrian Cx in the current frame that has not been successfully matched with any pedestrian in the pedestrian candidate set, the pedestrian is added to the pedestrian candidate set, and its pedestrian number is the maximum pedestrian number that has not been added to the previous pedestrian candidate set + 1.

5. The method for identifying jumping and squatting ticket evasion behaviors according to claim 1, characterized in that: The characteristics of step (4) are three characteristics: relative distance relationship, angle relationship and relative speed relationship, specifically: (4a) Relative distance relationship: The distance relationship is used to describe the distance in space between any pair of joints in the human skeleton. It is specifically quantified by the Euclidean distance. The relative position relationship RD(i,j) is defined as: Where i represents the joint point J i , j represents the joint point J j ; (x i ,y i ) and (x j ,y j ) represent the joint point J i and joint point J j The position coordinates at the same time satisfy the conditions of i, j = 0, 1, ..., 14 and i ≠ j; dis_shoulder represents half of the distance between the left shoulder joint point J5 and the right shoulder joint point J2 of the pedestrian at the same time, that is, half shoulder width; (4b) Angular relationships: limbs With limbs The angle between Defined as: Among them, (x i ,y i )、(x j ,y j )、(x e ,y e )、(x f ,y f ) represent the joint point J i , J j , J e , J f Position coordinates at the same moment; (4c) Relative speed relationship: Speed ​​is an important indicator of motion, which can provide key dynamic information for pedestrian motion recognition. The average speed over time is taken as the instantaneous speed of the joint point at time t; the pedestrian joint point J i The trajectory of motion L i for: L i ={(x i,0 ,and i,0 ),…,(x i,t-1 ,and i,t-1 ),(x i,t ,and i,t )} Among them, (x i,t ,y i,t ) is the joint point J i Position coordinates at time t; Joint point J i The instantaneous speed at time t for: Using the method used to calculate the relative distance relationship, the instantaneous speed will be obtained Divide by the pedestrian's half shoulder width dis_shoulder at time t t , convert it into a relative value, recorded as v i,t : Joint point J i At time t, along the spine The relative speed V i,t for:

6. The method for identifying jumping and squatting ticket evasion behaviors according to claim 1, characterized in that: The pedestrians in step (5) are classified into four categories, namely normal state, jumping state, squatting state and transition state; the limbs of pedestrians in the normal state are in a relatively relaxed and natural state, marked as "0"; the pedestrians in the jumping state have the characteristics of arms outstretched and the joints of the lower limbs moving up quickly, marked as "1"; the pedestrians in the squatting state are in a curled up state as a whole, marked as "2"; the pedestrians in the transition state have some characteristics of the normal state and some characteristics of the abnormal state. Even the human eye sometimes cannot accurately judge their state, marked as "3".

7. The method for identifying jumping and squatting ticket evasion behaviors according to claim 1 or 6, characterized in that: The single-frame behavior detection model described in step (5) is a random forest model.

8. The method for identifying jumping and squatting ticket evasion behaviors according to claim 7, characterized in that: The random forest model is a parallel ensemble learning method based on decision trees. It makes the final decision by voting on the classification results of multiple independent decision trees. The random forest has two random aspects: random sample combination and random feature set. The random forest algorithm randomly selects a certain number of feature sets and a certain number of sample combinations from the training samples as training subsets to ensure that the training data used by each decision tree is different and overlaps with each other.

9. The method for identifying jumping and squatting ticket evasion behaviors according to claim 1, characterized in that: The time series curve described in step (6) is obtained by using human posture estimation and pedestrian tracking to obtain the skeleton sequence of the pedestrian, and each human skeleton in the human skeleton sequence is used as the input of the ticket evasion behavior detection model to obtain the change of the pedestrian state over time, that is, the time series curve of the pedestrian state.