Vision-based three-stage escalator pedestrian converse driving detection method

By adopting a three-stage visual-based detection method in the escalator scene, combining facial detection, tracking of key points of multiple human bones and pedestrian retrograde behavior recognition, the problems of high misjudgment rate and poor generalization ability in the existing technology are solved, and accurate and real-time detection of pedestrian retrograde behavior of escalator pedestrians is achieved.

CN120198949APending Publication Date: 2025-06-24NANJING TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510600212.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

When detecting pedestrians retrograde behavior in the prior art, the misjudgment rate is high, the complex action recognition ability is limited, and the model generalization ability is poor, making it difficult to adapt to the specific needs of escalator scenes.

Method used

A three-stage detection method based on vision is adopted. First, the pedestrian face is detected through the face detection model. If detected, the multi-human skeleton key point tracking model is activated, and the target matching is combined with Kalman filtering and Hungarian algorithm. Finally, the pedestrian retrograde behavior recognizer is used to identify retrograde behavior by analyzing the skeleton key point data in successive frames.

Benefits of technology

Accurate detection of pedestrians' retrograde behaviors is achieved, calculation overhead and false alarm rate are reduced, suitable for real-time operation, and the safety guarantee and operation efficiency of public places are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198949A_ABST
    Figure CN120198949A_ABST
Patent Text Reader

Abstract

The invention provides a three-stage escalator pedestrian converse driving detection method based on vision, which specifically comprises the following steps: arranging camera equipment at an entrance of an escalator, and collecting continuous video streams; detecting the image frame by frame by using a face detection model; if the face of the pedestrian is detected, starting a multi-human skeleton key point tracking model; detecting key points of the whole body of the pedestrian by using a multi-body skeleton key point tracking model, and calculating the ratio of left and right thighs of the pedestrian based on the skeleton key points; when continuous N frames of left and right thigh ratios are obtained, a pedestrian converse motion behavior recognizer is started; discrete Fourier transform is carried out on the continuous N frames of ratio sequences; and if the significant amplitude gain occurs at a certain non-zero frequency point, determining that the pedestrian converse motion behavior exists. According to the pedestrian escalator retrograde walking detection method, automatic and accurate detection of pedestrian retrograde walking of the escalator is achieved, good robustness and low calculation complexity are achieved, and the pedestrian escalator retrograde walking detection method is suitable for real-time scene monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision, behavior recognition, and intelligent monitoring, and particularly relates to a vision-based three-stage escalator pedestrian reverse detection method. Background Art

[0002] In crowded public places such as shopping malls and subway stations, escalators, as core transportation facilities, undertake a large number of personnel transportation tasks. However, pedestrian reverse behavior is likely to cause safety accidents such as falls and collisions, especially during peak hours, which is particularly harmful to vulnerable groups such as the elderly and children. At the same time, such behavior will also exacerbate the wear of escalator components, increase equipment failure rates, interfere with public order, and affect the overall operation efficiency. Therefore, it is necessary to effectively detect and warn against escalator reverse behavior. Currently, related research mostly relies on intelligent video monitoring and adopts detection methods based on object tracking and behavior analysis. However, such methods are prone to misjudgment of behaviors such as short-term stays and turning around, and have limited complex action recognition capabilities. At the same time, reverse events themselves are low-frequency behaviors, and data samples are scarce, resulting in poor model generalization capabilities. Moreover, existing publicly available pedestrian detection datasets are generally not applicable to escalator scenarios. In summary, it is necessary to design a pedestrian reverse detection method applicable to escalator scenarios, with high detection accuracy and strong real-time performance, to improve the safety guarantee level and operation efficiency of public places. Summary of the Invention

[0003] The present invention provides a vision-based three-stage escalator pedestrian reverse detection method, which has the advantages of small computational overhead and low false alarm rate, to solve the problems proposed in the above background art. A vision-based three-stage escalator pedestrian reverse detection method includes the following specific steps:

[0004] S1. Deploy a camera directly behind the up or down entrance of the escalator (i.e., install it in the same direction as the escalator running direction) to ensure that the camera's view covers the entire panorama of the escalator running area;

[0005] S2. Collect video images through the camera and set the image collection frequency (frame rate) to ensure smooth pictures;

[0006] S3. For each frame of video image collected in real time, use a face detection model to detect the faces of pedestrians;

[0007] S4. Determine whether a pedestrian face is detected. If a pedestrian face is detected, enter S5 and start a multi-person skeleton key point tracking model. Otherwise, return to S2 and continue to collect the video stream;

[0008] S5. Use the multi-person skeleton key point tracking model to detect the full-body detection frames and skeleton key points of pedestrians in the image, and track the skeleton key points of different individuals in consecutive frames;

[0009] S6, determining whether the number of consecutive image frames that successfully detect and track the key points of the human skeleton has reached a preset threshold value of N frames, if it has reached N frames, entering S7 to trigger the pedestrian retrograde behavior identifier, otherwise returning to S5 to continue detecting and tracking the key points of the pedestrian skeleton;

[0010] S7, based on the human skeleton key points of the continuous N frames of images, use the pedestrian retrograde behavior identifier to analyze, if retrograde behavior is detected, an alarm prompt will pop up; if no retrograde behavior is detected, return to S2 and continue to collect video streams;

[0011] S8. Repeat S2 to S7 to ensure continuous monitoring and analysis of pedestrian behavior in the escalator area until the video stream collection is completed.

[0012] Furthermore, the facial detection model is used to perform facial detection on pedestrians in step S3, which specifically includes the following steps:

[0013] S3-1: Process the collected training images to uniformly adjust them to a preset standard size, manually mark the positions of pedestrian faces in the training images, and establish a pedestrian face dataset;

[0014] S3-2: applying mosaic data enhancement technology to the pedestrian face dataset to generate new enhanced samples to expand the data volume of the original training set;

[0015] S3-3: The enhanced pedestrian face dataset is fed into the YOLOv5s target detection model for training. By iteratively optimizing the YOLOv5s network parameters, a pedestrian face detector is obtained that can accurately detect the position of pedestrian faces in the image.

[0016] S3-4: Send the image to be detected to the trained pedestrian face detector, and obtain the detection frame of the pedestrian face in the image through model reasoning.

[0017] Further, in step S5, the whole body detection frame and skeleton key points of pedestrians in the image are detected by using a multi-human skeleton key point tracking model, and the skeleton key points of different individuals are tracked in consecutive frames, which specifically includes the following steps:

[0018] S5-1: Send the image to be detected to the multi-human skeleton key point tracking model for detection, obtain the detection frame of the pedestrian's whole body in the image and the coordinate information of the skeleton key points of the whole body, and obtain the skeleton key point set P = {P i}, a=1,2,...,C,P i =(x i ,y i ) is the coordinate of the i-th bone key point, x i ,y iare the abscissa and ordinate of the skeletal key points respectively, and C is the number of skeletal key points;

[0019] S5-2: For each detected pedestrian target, create a corresponding identity ID and initialize the status of the created identity ID to the to-be-confirmed status;

[0020] S5-3: Based on the motion trend and posture change of the skeletal key points in the historical frames, use the Kalman filtering method to predict the approximate position and posture of the pedestrian in the current frame, and obtain the predicted key point coordinates and then obtain the skeletal key point set to assist in matching the pedestrians in the current frame with those in the historical frames;

[0021] S5-4: For the pedestrians with each ID in the to-be-confirmed status, compare the skeletal key point set detected in the t-th frame with the skeletal key point set predicted based on the historical frames

[0022] to get the matching score S: Euclid + w2·E time

[0023] where: is the average Euclidean distance between the predicted key point in the t-th frame and the detected key point and is the error between the displacement change of the detected key point and the displacement change of the predicted key point. w1 and w2 are weight coefficients, satisfying w1 + w2 = 1;

[0024] S5-5: Based on the matching score S in step S5-4, use the Hungarian algorithm to perform the target matching operation. The matching process will produce three types of results: (1) Identity mismatch: If the matching degree between the key points of all detected bounding boxes and the key points of the predicted bounding boxes in the current frame does not reach the preset threshold and the corresponding identity ID is in the to-be-confirmed status, directly remove the identity ID; (2) Detection mismatch: If the key points of the predicted bounding boxes of all identity IDs cannot form a matching relationship with the key points of the detected bounding boxes generated in the current frame, a new identity ID and the corresponding ID need to be configured for the target, and the initial status of the newly generated identity ID is set to the to-be-confirmed state; (3) Identity match: When the similarity between the detected bounding box and the key points of the predicted bounding box meets the matching condition, assign the identity ID corresponding to the predicted bounding box to the detected bounding box and update the motion model. If the identity match is achieved three times in a row, change the status of the identity ID from the to-be-confirmed state to the confirmed state;

[0025] S5-6: Repeat steps S5-1 to S5-5 until a confirmed-state identity ID appears;

[0026] S5-7: For the pedestrian with the confirmed ID, the skeleton key point set obtained by the detection of the tth frame And the set of skeleton key points predicted based on historical frames Match again and get the matching score S;

[0027] S5-8: If the matching score S calculated in step S5-7 exceeds the preset threshold Max_source, it is determined that the matching fails. If the matching fails for multiple consecutive times, it is determined that the target is lost and its identity ID is removed.

[0028] Furthermore, the step S7 uses a pedestrian wrong-way behavior identifier to identify wrong-way behavior, which specifically includes the following steps:

[0029] S7-1: After the multi-body skeleton key point tracking model draws the key points of the pedestrian, the coordinates of the left hip joint (x1, y1) and the coordinates of the left knee joint (x2, y2) are extracted, and the length of the left thigh is calculated using the following formula:

[0030]

[0031] S7-2: According to the formula for calculating the length of the left thigh in step S7-1, the length of the right thigh d can be obtained similarly. right , and get the length of the left thigh d left and the length of the right thigh d right Then, the ratio of the two is obtained using the following formula:

[0032]

[0033] S7-3: The ratios of the left and right thighs in the acquired N consecutive frames are used to form a sequence x[n] = {rate1, rate2, rate 3... .rate N};

[0034] S7-4: In order to eliminate the influence of excessively high DC components, the ratios in the sequence are averaged. Then subtract the average value from all the ratios in the sequence to get a new sequence

[0035] S7-5: For new sequences Perform discrete Fourier transform, the formula is as follows:

[0036] in Where M is the length of the discrete transformation interval, represents rounding down, X[k] is a sequence The discrete Fourier transform of

[0037] S7-6: Calculate the amplitude A[k] of X[k], and the formula is as follows:

[0038]

[0039] where Re(X[k]) is the real part of X[k], and Im(X[k]) is the imaginary part of X[k];

[0040] S7-7: Determine the maximum amplitude A peak and its corresponding frequency f peak respectively as follows:

[0041]

[0042] where is the index corresponding to the maximum value in the sequence A[k];

[0043] S7-8: If the maximum amplitude A peak obtained in step S7-7 exceeds the set threshold ε A , and the frequency f peak corresponding to the maximum amplitude > ∈ f , then it can be determined that a retrograde behavior has occurred.

[0044] Compared with the disadvantages and deficiencies of the prior art, the present invention has the following beneficial effects:

[0045] (1) The vision-based three-stage escalator pedestrian retrograde detection method proposed by the present invention realizes automatic and accurate detection of escalator pedestrian retrograde, and has good robustness;

[0046] (2) The vision-based three-stage escalator pedestrian retrograde detection method proposed by the present invention only activates the subsequent multi-human body skeleton key point tracking model and retrograde behavior recognizer when a pedestrian's face is detected, reducing the computational overhead and being suitable for real-time operation. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 is a flowchart of the method of the present invention;

[0048] Figure 2 is a detection result diagram of the skeleton key point detector and the face detector when a pedestrian retrogrades on an escalator;

[0049] Figure 3 is a time domain diagram of the continuous N-frame left and right thigh ratio sequence x[n] collected;

[0050] Figure 4 is the continuous N-frame left and right thigh ratio sequence of the time domain diagram;

[0051] Figure 5 For sequence The amplitude-frequency characteristic diagram of

[0052] Specific implementation cases

[0053] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0054] The present invention discloses a three-stage escalator pedestrian reverse detection method based on vision, the overall flow chart is as follows: Figure 1 As shown, the method specifically comprises the following steps:

[0055] S1. Deploy the camera directly behind the upward or downward exit of the escalator (i.e. install it in the same direction as the escalator's running direction) to ensure that the camera's field of view covers the entire escalator running area;

[0056] S2, collect video images through the camera, and set the image collection frequency (frame rate) to ensure the continuity of the picture;

[0057] S3, for each frame of video image collected in real time, using a facial detection model to perform facial detection on pedestrians;

[0058] S4, judging whether a pedestrian's face is detected, if a pedestrian's face is detected, entering S5, starting a multi-human skeleton key point tracking model, otherwise returning to S2, continuing to collect video streams;

[0059] S5. Detect the whole body detection frame and skeleton key points of pedestrians in the image using a multi-human skeleton key point tracking model, and track the skeleton key points of different individuals in consecutive frames;

[0060] S6, determining whether the number of consecutive image frames that successfully detect and track the key points of the human skeleton reaches a preset threshold of 150 frames, if it reaches 150 frames, entering S7 to trigger the pedestrian retrograde behavior identifier, otherwise returning to S5 to continue detecting and tracking the key points of the pedestrian skeleton;

[0061] S7, based on the key points of the human skeleton of 150 consecutive frames of images, use the pedestrian retrograde behavior identifier to analyze, if retrograde behavior is detected, an alarm prompt will pop up; if retrograde behavior is not detected, return to S2 and continue to collect video streams;

[0062] S8. Repeat S2 to S7 to ensure continuous monitoring and analysis of pedestrian behavior in the escalator area until the video stream collection is completed.

[0063] Further, in step S3, a face detection model is used for face detection, which specifically includes the following steps:

[0064] S3-1: In this example, 2,690 real-scene images with different pedestrian flow densities in the escalator scene inside the subway station provided by a certain subway company are collected. The size of the collected images is processed to 640×640, and the pedestrian faces in the collected pictures are labeled using the annotation tool Labelimg to establish a pedestrian face dataset.

[0065] S3-2: Use the mosaic data augmentation technique to expand the pedestrian face dataset.

[0066] S3-3: Feed the constructed pedestrian face dataset into the YOLOv5s object detection model for training. The training parameters are set as follows: batch size 16, learning rate 0.01, number of iterations 100, confidence loss coefficient 0.5, loss function type GIOU, and the optimization algorithm is stochastic gradient descent. After multiple rounds of training, a pedestrian face detector is obtained.

[0067] S3-4: Feed the image to be detected into the trained pedestrian face detector, and obtain the detection box of the pedestrian face in the image through model inference.

[0068] Further, in step S5, a multi-person skeleton key point tracking model is used to detect the full-body detection box and skeleton key points of pedestrians in the image, and track the skeleton key points of different individuals in consecutive frames, which specifically includes the following steps:

[0069] S5-1: Feed the image to be detected into the multi-person skeleton key point tracking model for detection, obtain the full-body detection box of pedestrians in the image and the coordinate information of the full-body skeleton key points, and obtain the skeleton key point set P = {P i}, i = 1, 2,..., 17, P i = (x i , y i ) is the coordinate of the i-th skeleton key point, and x i , y i are the abscissa and ordinate of the skeleton key point respectively.

[0070] S5-2: For each detected pedestrian target, create a corresponding identity ID, and initialize the status of the created identity ID to the to-be-confirmed status.

[0071] S5-3: Based on the motion trend and posture change of the skeleton key points in the historical frame, use the Kalman filtering method to predict the approximate position and posture of the pedestrian in the current frame, obtain the predicted key point coordinates and then obtain the skeleton key point set to assist in the matching of pedestrians in the current frame and pedestrians in the historical frame.

[0072] S5-4: For each pedestrian with an ID in the to-be-confirmed status, the set of skeleton key points detected in the t-th frame is matched with the set of skeleton key points predicted based on historical frames to obtain a matching score S:

[0073] S = w1·L Euclid + w z ·E time

[0074] where is the average Euclidean distance between the predicted key points in the t-th frame and the detected key points and is the error between the displacement change of the detected key points and the displacement change of the predicted key points. w1 and w2 are weight coefficients, w1 = 0.7, w2 = 0.3;

[0075] S5-5: Based on the matching score S in step S5-4, the Hungarian algorithm is used to perform the object matching operation. The matching process will produce three types of results: (1) Identity mismatch: If the matching degree between the key points of all detected boxes and the key points of the predicted boxes in the current frame does not reach the preset threshold and the corresponding identity label is in the to-be-confirmed status, the identity label is directly removed; (2) Detection mismatch: If the key points of the predicted boxes of all identity labels cannot form a matching relationship with the key points of the detected boxes generated in the current frame, a new identity label and the corresponding ID need to be configured for the target, and the initial state of the newly generated identity label is set to the to-be-confirmed state; (3) Identity match: When the similarity between the detected box and the key points of the predicted box meets the matching condition, the identity label ID corresponding to the predicted box is assigned to the detected box, and the motion model is updated. If the identity match is achieved three times in a row, the status of the identity label is changed from the to-be-confirmed state to the confirmed state;

[0076] S5-6: Repeat steps S5-1 to S5-5 until a confirmed-state identity label appears;

[0077] S5-7: For the pedestrian with a confirmed-state ID, the set of skeleton key points detected in the t-th frame is matched again with the set of skeleton key points predicted based on historical frames to obtain a matching score S;

[0078] S5-8: If the matching score S calculated in step S5-7 exceeds the preset threshold Max_source = 0.6, it is determined as a matching failure. If the matching failure occurs 4 times in a row, it is determined that the target is lost and its identity label ID is removed.

[0079] The results of using the multi-human skeleton key point tracking model to detect video images (frames 70, 128, 158, and 201) are as follows: Figure 2 When the pedestrian’s face is facing the camera, the multi-body skeleton key point tracking model can successfully identify the pedestrian’s body skeleton key points, and preliminarily determine that the pedestrian may be walking against the traffic.

[0080] Furthermore, the step S7 uses a pedestrian wrong-way behavior identifier to identify wrong-way behavior, which specifically includes the following steps:

[0081] S7-1: After the multi-body skeleton key point tracking model draws the key points of the pedestrian, the coordinates of the left hip joint (x1, y1) and the coordinates of the left knee joint (x2, y2) are extracted, and the length of the left thigh is calculated using the following formula:

[0082]

[0083] S7-2: According to the formula for calculating the length of the left thigh in step 1, the length of the right thigh d can be obtained by the same logic. right , and get the length of the left thigh d left and the length of the right thigh d right Then, the ratio of the two is obtained using the following formula:

[0084]

[0085] S7-3: The obtained 150 consecutive frames of left and right thigh ratios form a sequence x[n] = {rate1, rate2, rate3....rate 150};

[0086] S7-4: In order to eliminate the influence of excessively high DC components, the ratios in the sequence are averaged. Then subtract the average value from all the ratios in the sequence to get a new sequence

[0087] S7-5: For new sequences Perform discrete Fourier transform, the length M of the discrete transform interval is set to 150, and the formula is as follows:

[0088] where k = 0, 1, ..., 74

[0089] S7-6: For X[k], calculate its amplitude A[k], the formula is as follows:

[0090]

[0091] where Re(X[k]) is the real part of X[k], and Im(X[k]) is the imaginary part of X[k];

[0092] S7-7: Using the amplitude value calculated in step S7-6, determine the maximum amplitude A through the argmax operation peak and its corresponding frequency f peak are respectively:

[0093]

[0094] A peak = A[k peak

[0095] where is the index corresponding to the maximum value in the sequence A[k];

[0096] S7-8: If the maximum amplitude A obtained in step S7-7 peak exceeds the set threshold of 2, and the frequency f corresponding to the maximum amplitude peak > 0.01, then it can be determined that a retrograde behavior has occurred.

[0097] Further analysis Figure 2 finds that when a pedestrian walks on an escalator, the length of the femur of the lifted leg is significantly shorter than that of the supporting leg, and the left and right thighs show an obvious alternating swing pattern. Based on this phenomenon, the present invention extracts the bone length between the key points of the left and right thighs of the pedestrian and calculates their ratio, which is used as an auxiliary basis for further judging whether the pedestrian is retrograde. The obtained left and right thigh ratio sequence x[n] and its corresponding new sequence are plotted as time-domain curves, as shown in Figure 3 and Figure 4 respectively. It is found from the figure that the left and right thigh ratios show obvious periodic characteristics with time. Based on this characteristic, the time-domain signal is further subjected to frequency-domain transformation. The amplitude-frequency characteristic of the sequence is as shown in Figure 5 . Based on this characteristic, the time-domain signal is further subjected to frequency-domain transformation. The amplitude-frequency characteristic of the sequence is as shown in Figure 5 . It can be seen that the maximum amplitude A peak exceeds the set threshold of 2, and the corresponding frequency f peak > 0.01, then it can be determined that a retrograde behavior has occurred.

[0098] In summary, the method proposed by the present invention alleviates to a certain extent the problems of insufficient training data and weak model generalization ability caused by the scarcity of pedestrian retrograde behavior samples. By converting the time-domain signal to the frequency-domain and analyzing it in combination with frequency characteristics, the accuracy and robustness of retrograde behavior recognition are effectively improved.

[0099] ​The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A three-stage vision-based escalator pedestrian reverse detection method, the key is: The steps include: S1. Deploy the camera directly behind the upward or downward exit of the escalator (i.e. install it in the same direction as the escalator's running direction) to ensure that the camera's field of view covers the entire escalator running area; S2, collect video images through the camera, and set the image collection frequency (frame rate) to ensure the continuity of the picture; S3, for each frame of video image collected in real time, using a facial detection model to perform facial detection on pedestrians; S4, judging whether a pedestrian's face is detected, if a pedestrian's face is detected, entering S5, starting a multi-human skeleton key point tracking model, otherwise returning to S2, continuing to collect video streams; S5. Detect the whole body detection frame and skeleton key points of pedestrians in the image using a multi-human skeleton key point tracking model, and track the skeleton key points of different individuals in consecutive frames; S6, determining whether the number of consecutive image frames that successfully detect and track the key points of the human skeleton has reached a preset threshold value of N frames, if it has reached N frames, entering S7 to trigger the pedestrian retrograde behavior identifier, otherwise returning to S5 to continue detecting and tracking the key points of the pedestrian skeleton; S7, based on the human skeleton key points of the continuous N frames of images, use the pedestrian retrograde behavior identifier to analyze, if retrograde behavior is detected, an alarm prompt will pop up; if no retrograde behavior is detected, return to S2 and continue to collect video streams; S8. Repeat S2 to S7 to ensure continuous monitoring and analysis of pedestrian behavior in the escalator area until the video stream collection is completed.

2. The three-stage escalator pedestrian reverse detection method based on vision according to claim 1, wherein the key is that the face detection model is used to perform face detection on the pedestrian in step S3, and specifically comprises the following steps: S3-1: Process the collected training images to uniformly adjust them to a preset standard size, manually mark the positions of pedestrian faces in the training images, and establish a pedestrian face dataset; S3-2: applying mosaic data enhancement technology to the pedestrian face dataset to generate new enhanced samples to expand the data volume of the original training set; S3-3: The enhanced pedestrian face dataset is fed into the YOLOv5s target detection model for training. By iteratively optimizing the YOLOv5s network parameters, a pedestrian face detector is obtained that can accurately detect the position of pedestrian faces in the image. S3-4: Send the image to be detected to the trained pedestrian face detector, and obtain the detection frame of the pedestrian face in the image through model reasoning.

3. A three-stage vision-based escalator pedestrian reverse detection method according to claim 1, characterized in that: The step S5 uses a multi-human skeleton key point tracking model to detect the whole body detection frame and skeleton key points of pedestrians in the image, and tracks the skeleton key points of different individuals in consecutive frames, specifically including the following steps: S5-1: Send the image to be detected to the multi-human skeleton key point tracking model for detection, obtain the detection frame of the pedestrian's whole body in the image and the coordinate information of the skeleton key points of the whole body, and obtain the skeleton key point set P = {P i },i=1,2,...,C,P i =(x i ,y i ) is the coordinate of the i-th bone key point, x i ,y i are the horizontal and vertical coordinates of the skeleton key points, respectively, and C is the number of skeleton key points; S5-2: For each pedestrian target detected, a corresponding identity ID is created, and the state of the created identity ID is initialized to a pending confirmation state; S5-3: Based on the motion trend and posture changes of the skeleton key points in the historical frames, the Kalman filter method is used to predict the approximate position and posture of the pedestrian in the current frame to obtain the predicted key point coordinates Then we get the skeleton key point set To assist in matching pedestrians in the current frame with pedestrians in the historical frame; S5-4: For each pedestrian ID in the state to be confirmed, the skeleton key point set obtained by the detection of the tth frame And the set of skeleton key points predicted based on historical frames Perform matching and obtain the matching score S: S=w1·D Euclid +w2·E time in: Predict key points for frame t And detection key points The average Euclidean distance between The error between the displacement change of the detected key point and the displacement change of the predicted key point is w1 and w2, which are weight coefficients, satisfying w1+w2=1; S5-5: Based on the matching score S in step S5-4, the Hungarian algorithm is used to perform the target matching operation. The matching process will produce three types of results: (1) Identity mismatch: If the matching degree between the key points of all detection frames and the key points of the prediction frames in the current frame does not reach the preset threshold, and the corresponding identity identifier is in the pending confirmation state, the identity identifier is directly removed; (2) Detection mismatch: If the prediction frame key points of all identity identifiers cannot form a matching relationship with the detection frame key points generated by the current frame detection, a new identity identifier and corresponding ID must be configured for the target, and the initial state of the newly generated identity identifier is set to the pending confirmation state; (3) Identity match: When the similarity between the key points of the detection frame and the prediction frame meets the matching condition, the identity identifier ID corresponding to the prediction frame is assigned to the detection frame, and the motion model is updated. If identity matching is achieved three times in a row, the state of the identity identifier is changed from the pending confirmation state to the confirmed state; S5-6: Repeat steps S5-1 to S5-5 until a confirmation status identity appears; S5-7: For the pedestrian with the confirmed ID, the skeleton key point set obtained by the detection of the tth frame And the set of skeleton key points predicted based on historical frames Match again and get the matching score S; S5-8: If the matching score S calculated in step S5-7 exceeds the preset threshold Max_source, it is determined that the matching fails. If the matching fails for multiple consecutive times, it is determined that the target is lost and its identity ID is removed.

4. According to the three-stage vision-based escalator pedestrian retrograde detection method of claim 1, the key is to use the pedestrian retrograde behavior identifier in step S7 to identify the retrograde behavior, which specifically includes the following steps: S7-1: After the multi-body skeleton key point tracking model draws the key points of the pedestrian, the coordinates of the left hip joint (x1, y1) and the coordinates of the left knee joint (x2, y2) are extracted, and the length of the left thigh is calculated using the following formula: S7-2: According to the formula for calculating the length of the left thigh in step S7-1, the length of the right thigh d can be obtained similarly. right , and get the length of the left thigh d left and the length of the right thigh d right Then, the ratio of the two is obtained using the following formula: S7-3: The ratios of the left and right thighs in the acquired N consecutive frames are used to form a sequence x[n] = {rate1, rate2, rate3....rate N }; S7-4: In order to eliminate the influence of excessively high DC components, the ratios in the sequence are averaged. Then subtract the average value from all the ratios in the sequence to get a new sequence S7-5: For new sequences Perform discrete Fourier transform, the formula is as follows: in, M is the length of the discrete transformation interval, represents rounding down, X[k] is a sequence The discrete Fourier transform of S7-6: For X[k], calculate its amplitude A[k], the formula is as follows: Where Re(X[k]) is the real part of X[k], and Im(X[k]) is the imaginary part of X[k]; S7-7: Determine the maximum amplitude A by using the argmax operation based on the amplitude calculated in step S7-6 peak And its corresponding frequency f peak They are: A peak =A[k peak ] in is the index corresponding to the maximum value in the sequence A[k]; S7-8: If the maximum amplitude A obtained in step S7-7 peak Exceeding the set threshold ε A , and the frequency corresponding to the maximum amplitude is f peak >ε f , it can be determined that retrograde behavior has occurred.