Pedestrian passing intention multi-modal prediction method and prediction device

By incorporating multimodal features to predict pedestrian intent, the real-time performance and accuracy of turnstile systems in detecting abnormal passage behavior are addressed. This enables efficient identification and timely response to abnormal intent, adapting to various population states, especially special groups, thereby improving passage safety and efficiency.

CN120877412AActive Publication Date: 2025-10-31IVES (SUZHOU) SPECIAL EQUIP CO LTD

Patent Information

Application Number
CN202511385910.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-26
Publication Date
2025-10-31
Estimated Expiration
2045-09-26

AI Technical Summary

Technical Problem

Existing turnstile systems, especially normally open turnstiles, suffer from insufficient real-time performance and accuracy when detecting abnormal passage behavior. They struggle to accurately distinguish between special situations such as multiple people passing closely together, carrying luggage, or children, leading to missed detections or misjudgments.

Method used

By integrating multimodal features of pedestrians, such as gaze direction, action behavior, hand swiping action, and human body features, lightweight object detection, ReID technology, Kalman filtering, and ByteTrack algorithm are used for multi-target tracking. Key points are extracted by combining HRNet-W32 network to identify card swiping action, running behavior, and pedestrian category, and a decision model is used for comprehensive judgment.

Benefits of technology

It improves the accuracy and timeliness of abnormal behavior detection, can identify abnormal intentions in advance, reduce missed and false alarms, adapt to complex scenarios, especially special groups of people, ensure traffic safety, and complete processing in milliseconds to improve traffic efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120877412A_ABST
    Figure CN120877412A_ABST
Patent Text Reader

Abstract

The invention provides a pedestrian passing intention multi-modal prediction method and a prediction device, which are applied to the technical field of access control monitoring, extract pedestrian multi-modal features based on acquired image information, and judge whether current pedestrians are allowed to pass through or not through decision logic, and comprise the following steps: S1, acquiring an image sequence of pedestrians entering a station, executing multi-target detection and tracking, and determining whether the current pedestrians are allowed to pass through; identifying each pedestrian ID to form a bounding box sequence; s2, multi-modal feature extraction is carried out on each pedestrian, and the multi-modal features comprise card swiping actions, eye directions, running behaviors and pedestrian types; and S3, processing the multi-modal features, and inputting the processed multi-modal features into a decision model for judgment so as to control the state of the gate. According to the pedestrian passing intention multi-mode prediction method and the pedestrian passing intention multi-mode prediction device, the passing intention of the current passing person is comprehensively judged by fusing information such as the sight direction, the action behavior, the hand card swiping action and the human body characteristics of the pedestrian, and the accuracy and the timeliness of abnormal behavior detection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of access control monitoring technology, specifically relating to a multimodal prediction method and device for pedestrian passage intentions. Background Technology

[0002] In modern urban rail transit systems, the management of turnstiles is crucial for preventing fare evasion and ensuring passenger safety. Currently, commonly used turnstiles are divided into two types: normally closed turnstiles and normally open turnstiles. Normally closed turnstiles typically keep the gate closed when unauthorized access is granted, opening only after card authorization. Normally open turnstiles, on the other hand, keep the gate open under normal circumstances to improve throughput, only closing quickly to block unauthorized passage or abnormal situations. While normally open turnstiles increase throughput speed due to their default openness, they also place higher demands on abnormal passage detection.

[0003] Current methods for detecting abnormal passage primarily rely on infrared beam sensor arrays. While placing infrared beam sensors within turnstile channels can detect people passing by and the number of people, they have limitations in complex situations: when multiple people are tailgating, multiple infrared beams may be blocked simultaneously, making it difficult to accurately distinguish the number of people and thus failing to promptly identify tailgating. Furthermore, the installation location of infrared sensors within turnstile channels is currently limited, often resulting in misjudgments or missed detections in special circumstances such as those with luggage or children. For example, a small child might pass under the sensor without triggering it, causing fare evasion; conversely, passengers passing through legitimately might be mistakenly identified as violating regulations if their movements are slightly abnormal. Moreover, because sensors can only be deployed on the turnstile body, their detection area is limited, and the gate's reaction time is insufficient in the event of emergencies. Therefore, current infrared beam sensor methods cannot fully meet the real-time and accuracy requirements of normally open turnstiles.

[0004] With the development of machine vision technology, some turnstiles have begun to incorporate cameras for monitoring. For example, binocular cameras are used to acquire 3D depth information to determine the number of people in the passageway, or facial recognition is combined to verify the identity of people entering and exiting. These vision-based solutions have improved detection capabilities to some extent, such as recognizing faces or detecting multiple people passing through simultaneously within a certain range. However, current vision solutions mostly focus on identity authentication (such as facial comparison) or people counting, lacking in-depth analysis of the details of pedestrians' intentions. Furthermore, analysis of single visual features is still unreliable in some scenarios (such as judging tailgating solely by whether two people are detected), potentially leading to missed detections and false alarms.

[0005] In summary, existing technologies still have significant shortcomings in accurately predicting pedestrians' intentions and cannot fully meet the needs of normally open turnstiles for advance prediction and precise control. Summary of the Invention

[0006] In view of the above-mentioned problems in the prior art, the purpose of this invention is to provide a multimodal prediction method for pedestrian crossing intentions, which integrates information such as pedestrian's gaze direction, action behavior, hand swiping action and human body characteristics to comprehensively judge the current pedestrian's crossing intentions, thereby improving the accuracy and timeliness of abnormal behavior detection.

[0007] A multimodal prediction method for pedestrian crossing intentions extracts multimodal features of pedestrians based on image information acquired from camera equipment, and determines whether to allow the current pedestrian to cross through decision logic. The specific process includes the following steps: S1. Obtain the image sequence of pedestrians entering the station, perform multi-target detection and tracking, identify each pedestrian ID, and form a bounding box sequence with consistent pedestrian IDs; S2. For each pedestrian, perform multimodal feature extraction, including card swiping action, gaze direction, running behavior, and pedestrian category; S3. After processing the multimodal features, input them into the decision model for judgment to control the gate status; The decision-making process of the decision-making model is as follows: S31. Determine whether the pedestrian has swiped their card. If the pedestrian has swiped their card, the decision model outputs a probability of passage of 1. If the pedestrian has not swiped their card, proceed to the pedestrian intention analysis stage. S32. Determine whether the pedestrian's gaze is fixed on the card-swiping area. If the pedestrian's gaze is fixed on the card-swiping area and the probability of running is not greater than running probability one, then the decision model outputs passage probability two; if the pedestrian's gaze is fixed on the card-swiping area and the probability of running is greater than running probability one, then the decision model outputs passage probability three. If a pedestrian's gaze is not fixed on the card-swiping area, and the probability of running is greater than the probability of running 2, the decision model outputs the probability of passage 4; if a pedestrian's gaze is not fixed on the card-swiping area, and the probability of running is not greater than the probability of running 2, the model proceeds to the pedestrian category recognition stage. S33. Based on S32, determine whether the pedestrian is a special passenger. If the pedestrian is a special passenger, the decision model outputs the overall passage probability as probability five. If the pedestrian is not a special passenger, the decision model outputs the passage probability as probability six. S34. Based on S31-S33, drive the gate to remain open or immediately close and trigger an alarm according to the passage probability output by the decision model; wherein, if the overall passage probability output by the decision model is less than a preset value, drive the gate to close and trigger an alarm. If not, then keep the gate open.

[0008] Preferably, in step S1, the sequence of pedestrian images entering the station is acquired, multi-object detection and tracking are performed, each pedestrian ID is identified, and a sequence of bounding boxes with consistent pedestrian IDs is formed. The specific process is as follows: Pedestrian target detection: A lightweight object detection model is used to output pedestrian bounding boxes in each frame of the image. Where B represents the bounding box, t represents the frame number, and i represents the pedestrian ID; Pedestrian identity association: Pedestrian re-identification is performed using ReID technology, which extracts pedestrian features to form ReID feature vectors, used to distinguish different pedestrians. Pedestrian features include clothing, accessories, and body posture. The position / velocity of pedestrians is predicted using the Kalman filter algorithm, which is used to predict the position of pedestrians in the next frame; Multi-target tracking is performed using the ByteTrack algorithm, outputting the bounding box sequence of pedestrian i in T consecutive frames. The expression for the bounding box sequence is: .

[0009] Preferably, the recognition of a card swipe action includes the following steps: Set up corresponding ROI areas based on the card swiping area of ​​the turnstile equipment; Extract the two-dimensional coordinates and confidence scores of key points of the pedestrian skeleton in each frame of the image; In N consecutive frames, the two-dimensional coordinates of the pedestrian's wrist key points continuously fall within the ROI region, and the confidence level is not lower than the confidence threshold. If the pedestrian's elbow joint shows a tendency to move towards the ROI area, then the pedestrian is determined to have swiped a card.

[0010] Preferably, in each frame of the image, the HRNet-W32 network outputs the two-dimensional coordinates and confidence scores of 17 key points of the pedestrian, expressed as follows: ,in Let K be the two-dimensional coordinates of the Kth keypoint on the t-frame image. The corresponding confidence levels are: K=9 for the left wrist, k=10 for the right wrist, K=7 for the left elbow, and K=8 for the right elbow.

[0011] Preferably, the recognition of gaze direction includes the following steps: Based on the coordinates of the pedestrian's facial key points and the constructed 3D average face model, the rotation matrix of the pedestrian's head relative to the camera device is calculated using OpenCV's PnP or POSIT algorithm. and translation vector To obtain the pedestrian's head pose, where the rotation matrix... Including yaw angle Pitch angle Roll angle ; The system spatially matches the pedestrian's head posture with the card-swiping area determined by the gate device to determine whether the pedestrian's gaze is pointing towards the card-swiping area of ​​the gate device. When the pedestrian's gaze is continuously directed towards the card-swiping area for a preset minimum time, it is determined that the pedestrian has the intention to swipe the card.

[0012] Preferably, spatial matching of pedestrian head posture and card-swiping area is achieved by comparing angle thresholds. The specific process is as follows: Set the reference value for the head yaw angle required for pedestrians to normally look at the card swiping area. Head pitch angle reference value ; Set the maximum allowable deviation threshold, including yaw angle deviation. Pitch angle deviation ; When pedestrian i's head is facing the desired direction and At that time, it is determined that pedestrian i's line of sight falls into the direction of the card swiping area.

[0013] Preferably, the determination of running behavior is based on the coordinates and confidence level of the target pedestrian's skeletal key points, specifically including the following steps: Skeletal keypoint extraction: The bounding box sequence is input into the HRNet-W32 keypoint network, which outputs 17 skeletal keypoints, expressed as follows:

[0014] in, Let K be the two-dimensional coordinates of the Kth keypoint on the t-frame image. The corresponding confidence level; Temporal feature construction: A sliding window of length L is set for each pedestrian i to capture dynamic behavioral features within the sliding window. The expression for the sliding window is: ; A dimension matrix is ​​constructed based on 17 key points of each frame image. This forms a spatiotemporal trajectory map; Set normalized timestamp The time series feature matrix is ​​obtained. ; Calculate the displacement difference of key points in adjacent frames: , in, Let be the coordinates of the k-th keypoint in frame t.

[0015] Based on the above displacement difference characteristics, an enhanced temporal feature matrix is ​​constructed: ; in: Faux = , indicating the introduced keypoint displacement difference feature; This represents a temporal feature set with a length of L frames, where each frame contains original keypoints and auxiliary features; L is the length of the sliding window; i is the pedestrian number.

[0016] Behavior classification: based on temporal feature matrix Temporal convolutional networks are used to output running probabilities. ; Velocity fusion determination: Take the center position difference of pedestrian i bounding boxes spaced 4 frames apart as the displacement, and calculate the trajectory velocity. The expression is:

[0017] in, Let be the coordinates of the center point of the i-th pedestrian bounding box in frame t; when At that time, it was determined that pedestrian i posed a risk of running, among which, The running probability threshold, This is the speed threshold.

[0018] Preferably, the pedestrian's classification is based on visual recognition. Specifically, the key points of the human skeleton and the pedestrian's height are compared with the thresholds set by the system. If the result of the judgment is within the set thresholds, and the key point distribution and height are both within the normal range for adults, then the pedestrian is judged to be a normal adult. If the key points of the human skeleton and the pedestrian's height are lower than the thresholds set by the system, then the pedestrian is judged to be a child. If the pedestrian's walking speed is less than the walking threshold and the posture stability is less than the posture threshold, then the pedestrian is judged to be an elderly person. By combining pedestrian target detection model, head orientation detection model and object detection model, we can identify people with disabilities and their commonly used assistive tools, including wheelchairs, canes or walking aids. If an assistive tool is identified, the pedestrian is classified as a person with a disability. By combining pedestrian target detection model, head orientation detection model and object detection model, the system identifies the characteristics and shape of the person pushing the stroller and the stroller, including the stroller frame and wheels. When the stroller is detected, the pusher is associated with the stroller, and the pedestrian category is determined to be a person pushing a stroller.

[0019] The beneficial effects of this invention are: the multimodal prediction method and device for pedestrian crossing intentions, Multi-feature fusion improves accuracy: By fusing multiple visual features, including card-swiping actions, eye direction, running movements, and pedestrian personality, the system comprehensively judges pedestrians' intentions. Compared to solutions using a single sensor or a single algorithm, this multi-modal information fusion provides richer criteria for judgment, significantly improving the accuracy of abnormal behavior detection and reducing false negatives and missed positives.

[0020] Early prediction and timely control: By using deep learning models to predict pedestrian intentions, abnormal intentions can be identified before pedestrians complete their passage. Once the model determines that there may be abnormal behavior such as tailgating or attempting to break through the gate, the system can issue an instruction in advance to quickly close the normally open gate, intercepting unauthorized personnel before they enter the passage. Compared to the traditional method of responding only when the gate passage is occupied, this method can respond to potential risks more promptly and ensure passage safety.

[0021] Adaptable to complex scenarios and highly robust: By introducing human key point recognition and behavior analysis, the solution of this invention can adapt to various crowd states. For example, whether a pedestrian is running, walking slowly, carrying luggage, or pushing a stroller, this system can identify and judge them through posture and movement analysis. At the same time, features such as eye direction and hand movements enable the system to perceive whether a pedestrian is performing operations such as swiping a card. Therefore, it still has good robustness in complex scenarios such as dense crowds, changing lighting, and partial occlusion, significantly outperforming traditional solutions that rely solely on a single signal.

[0022] Intelligent identification for special groups: The system specifically considers special passage scenarios such as the elderly, children, and strollers. Through height judgment or facial feature analysis, it can identify children or elderly people with slow movement in accompanying persons and adjust the gate logic to adapt to these situations. For example: Elderly people with children: If an adult is detected traveling with a young child and the travel regulations are met, the system can determine that passage is normal and keep the gate open, avoiding misidentifying the child as a trailer and incorrectly closing the gate. Stroller passage: When a parent is detected pushing a stroller, the system automatically extends the open mode, not immediately closing the gate, allowing the parent to smoothly push the stroller into the passage before swiping the card, improving passage efficiency and eliminating the risk of the gate closing incorrectly due to an excessively long stroller. This intelligent identification capability for various special scenarios not only improves the system's user-friendliness but also expands its applicability, ensuring safe and barrier-free passage in various complex group combinations.

[0023] Real-time performance and efficient passage: Optimized feature extraction and decision-making modules enable processing within milliseconds, meeting the real-time requirements of peak passenger flow. Since normally open turnstiles do not require physical opening or closing under normal circumstances, this invention ensures that passage efficiency is not reduced while maintaining safety. Compared to normally closed turnstiles that require each person to stop and swipe their card to open the door, this solution eliminates the need for stopping in most normal passage situations, improving passenger experience and passage speed.

[0024] Easy to integrate and promote: The method and apparatus of this invention mainly rely on software algorithms and general-purpose camera hardware, thus exhibiting good versatility in implementation. Existing normally open turnstile systems can be upgraded by adding cameras and computing units without requiring major modifications to the original turnstile mechanism. Furthermore, this technical solution is not only applicable to turnstiles in subways, train stations, and other rail transit systems, but can also be extended to other access control scenarios, such as office building access control and scenic spot ticket gates, in systems that require distinguishing between normal passage and tailgating, demonstrating broad application value. Attached Figure Description

[0025] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart of the decision-making model judgment process of the present invention; Figure 2 This is a schematic block diagram of the prediction device of the present invention; Figure 3 This is a connection diagram of the camera device and the gate device of the present invention; Figure 4 This is a flowchart of the card swiping action recognition process of the present invention; Figure 5 This is a flowchart of the gaze recognition process of the present invention; Figure 6 This is a flowchart of the running motion recognition process of the present invention; Figure 7 This is a flowchart of the pedestrian category recognition process of the present invention. Detailed Implementation

[0026] Example 1 like Figure 1 As shown, a multimodal prediction method for pedestrian crossing intentions extracts multimodal features of pedestrians based on image information acquired from a camera device, and determines whether to allow the current pedestrian to cross through decision logic. The specific process includes the following steps: S1. Obtain the image sequence of pedestrians entering the station, perform multi-object detection and tracking, identify each pedestrian ID, and form a sequence of bounding boxes with consistent pedestrian IDs. The specific process is as follows: Pedestrian object detection: A lightweight object detection model is used to output pedestrian bounding boxes in each frame of the image acquired by the camera device. Where B represents the bounding box, t represents the frame number, and i represents the pedestrian ID. Specifically, a lightweight object detection model can be YOLOv8-Nano, which, as a real-time detection model, can obtain the location of all pedestrians in each frame of the image.

[0027] Pedestrian identity association: Pedestrian re-identification is performed through ReID technology, and pedestrian features are extracted to form ReID feature vectors, which are used to distinguish different pedestrians. Pedestrian features include clothing, accessories and body posture, etc.

[0028] The position / velocity of pedestrians is predicted using the Kalman filter algorithm, which is then used to predict the pedestrian's position in the next frame, thus achieving tracking continuity.

[0029] Multi-target tracking is performed using the ByteTrack algorithm, outputting the bounding box sequence of pedestrian i in T consecutive frames. The expression for the bounding box sequence is: The ByteTrack algorithm, as a multi-target tracking algorithm, can maintain the consistency of pedestrian IDs in complex occlusion scenarios, avoid feature confusion between different pedestrians, and improve the stability and accuracy of multimodal feature recognition.

[0030] The consistency of each pedestrian's identity across frames is ensured through the cooperation of ReID feature vectors, Kalman filtering algorithm, and ByteTrack algorithm.

[0031] S2. For each pedestrian, perform multimodal feature extraction, including card swiping action, gaze direction, running behavior, and pedestrian category.

[0032] The recognition of card swiping actions includes the following steps: Based on the card-swiping area, a corresponding ROI region is set: Since the position of the card reader in the gate is relatively fixed, the location range of the card reader can be determined manually or by using camera equipment calibration parameters, thereby setting the corresponding ROI region. This ROI region can be represented by a rectangular area R. .

[0033] Upper limb joint skeletal keypoint extraction: In each frame of the image, a human pose estimation model is used to extract keypoints of the hand skeleton. Specifically, the HRNet-W32 network is used to output the two-dimensional coordinates and confidence scores of 17 keypoints, expressed as follows: ,in Let K be the two-dimensional coordinates of the Kth keypoint on the t-frame image. For the corresponding confidence levels, specifically, the key point K=9 for the left wrist, K=10 for the right wrist, K=7 for the left elbow, and K=8 for the right elbow.

[0034] Card swipe action recognition determination: When the wrist key point of pedestrian i falls into the ROI region for N consecutive frames and the confidence level is not lower than the confidence threshold. If the elbow joint shows a clear tendency to move towards the ROI area, the card swiping action is considered complete; otherwise, the card swiping action is considered incomplete.

[0035] The specific implementation process is as follows: First, based on the extracted hand key points, it is determined whether the pedestrian's hand falls within the designated Region of Interest (ROI). Specifically, for each pedestrian i in frame t, it is checked whether the coordinates of its left and right wrist key points fall within the rectangular region R, and whether the corresponding confidence values ​​are not lower than the confidence threshold. .

[0036] For example, setting a confidence threshold =0.5, using an indicator function to represent the judgment of a single frame:

[0037] when If the value is positive, it means that the frame detected pedestrian i's wrist entering the card swiping area; otherwise, no card swiping action was detected.

[0038] Secondly, a time condition for continuous frame verification is set. When the number of consecutive frames N in which pedestrian i's wrist remains within the ROI region reaches the frame threshold, the card swipe action is considered to have been completed, where the frame threshold is set to N=3.

[0039] Since single-frame entry detection may be affected by jitter or momentary false detections, a continuous frame determination strategy is adopted to improve robustness. That is, when there are consecutive frames N=3 within a time period... Only then was it determined that pedestrian i had completed the card swipe action.

[0040] In addition, the movement direction of the elbow joint is introduced as a redundancy check to prevent misjudgments caused by accidental overlap of key wrist points. By observing the positional changes of the elbow key points corresponding to the wrist, it is ensured that the arm movement trend matches the card swiping action.

[0041] Specifically, the displacement vector of the elbow before and after entering the card swiping area is calculated. If the elbow position moves significantly towards the ROI area compared to several frames ago, it is determined whether the horizontal distance between the elbow and the ROI area has decreased. If it has decreased, it means that the arm has indeed extended towards the card swiping area. Conversely, if the wrist briefly extends into the ROI area but the elbow does not move accordingly, it may be an unintentional action that is not card swiping and can be regarded as a false trigger, which is not counted as a valid card swipe.

[0042] By combining the three factors of wrist ROI overlap, duration, and elbow coordination, the card-swiping action can be identified without introducing a specially trained gesture classification model.

[0043] In summary, this card-swiping action recognition process can identify whether a pedestrian's card-swiping action has occurred the instant they approach the card-swiping area, providing a crucial basis for distinguishing between normal passage and tailgating. This skeletal keypoint-based discrimination method has extremely low computational overhead (only a few keypoint coordinates are determined, with negligible time consumption) and does not rely on additional training of complex models, exhibiting good interpretability and engineering feasibility. On a practical turnstile embedded platform, this method can run synchronously with the video frame rate, achieving high-speed and reliable detection of card-swiping actions, providing important visual feature support for predicting abnormal behavior in turnstile passage.

[0044] Passengers using cards normally will focus their attention on the card reader on the gate to align and complete the ticket verification process; conversely, those attempting to sneak in often look straight ahead or glance around without paying attention to the card reader. Based on this behavioral characteristic, by detecting the direction of a passenger's head and gaze, we can identify whether their attention is focused on the card reader area, thereby inferring their card-swiping intention.

[0045] The process of recognizing the direction of gaze is as follows: 1) Head orientation detection: Based on the extracted facial landmark distribution and combined with the predefined 3D average face model coordinates, the rotation matrix of the head relative to the camera device is calculated using OpenCV's PnP algorithm or POSIT algorithm. and translation vector To obtain relevant information data for subsequent steps in determining pedestrian walking speed, a rotation matrix is ​​used. Head orientation parameters that can be converted to Euler angles, including yaw angles. Pitch angle Roll angle This is used for head roll correction. The yaw angle is one of the components. and pitch angle These are key parameters for determining the direction of the gaze, reflecting the angle at which the passenger is looking to the left / right and up / down.

[0046] 2) Gaze Area Matching: Spatially match the direction of a pedestrian's head with the set ROI area to determine whether the pedestrian's gaze is pointing towards the ROI area of ​​the gate, which helps to determine their card swiping intention.

[0047] There are two methods for matching gaze regions, and you can choose one of them to achieve gaze recognition.

[0048] The first method involves comparing angle thresholds to determine the result. Based on the relative orientation between the pedestrian's current position and the card reader's position, set reference values ​​for the head yaw and pitch angles required for normal gaze at the ROI area. .

[0049] For example, if the card reader is located slightly lower to the right front of the turnstile, passengers need to look down slightly to the right front when swiping their cards. In this case, a maximum allowable deviation threshold is defined, such as yaw angle deviation. Pitch angle deviation Then, when pedestrian i's head is facing the condition that... and At that time, it can be determined that their gaze is roughly falling in the direction of the ROI region. Furthermore, to improve reliability, the face detection confidence level can be required to be no lower than a confidence threshold, i.e., a confidence level of [missing information]. This is to avoid misjudgments caused by inaccurate angle estimation.

[0050] The second method involves detection and judgment through line-of-sight projection. 3) Combining camera calibration parameters, the spatial relationship between the line of sight and the ROI region is directly determined using geometric methods. Specifically, using the head pose estimation results, the pedestrian's line of sight is extended into a ray in three-dimensional space and projected onto the world coordinate system or image plane. The pedestrian's line of sight is determined by the head pose rotation matrix obtained in the previous steps. Determined by the relative position of the head; If the line of sight intersects with a predefined ROI (e.g., the line of sight falls within the ROI rectangle on the gate plane), it indicates that the passenger's gaze is indeed directed at the card reader, and the gaze can be considered valid.

[0051] Compared to simple angle thresholding, this method utilizes the absolute spatial positional relationship obtained from camera device calibration, resulting in more accurate judgment, but the calculation is also relatively complex. Considering real-time performance on embedded platforms, the angle thresholding method is usually preferred for rapid screening.

[0052] 4) Determining the intent to swipe the card: The gaze region is matched frame by frame. When a pedestrian's gaze is continuously directed toward the ROI region for a preset minimum time, it is determined that they have the intention to swipe a card.

[0053] For example, in this embodiment, the consecutive frame threshold N = 3 (approximately 0.1 seconds) is defined. If pedestrian i's head is pointing towards the ROI region in all N consecutive frames and the detection confidence level is met, it is determined that the pedestrian has focused its attention on the card reader and meets the card swiping intention determination condition; otherwise, if the pedestrian only briefly faces the card reader in a few frames (e.g., a quick glance) and does not meet the duration requirement, it is considered that there is no clear card swiping intention.

[0054] By setting timing constraints, jitter caused by misjudgment of a single frame can be greatly reduced, improving the accuracy of card swiping intent determination.

[0055] The determination of running behavior is based on the coordinates and confidence levels of the target pedestrian's skeletal key points, and includes the following steps: Skeletal keypoint extraction: Input the bounding box sequence with consistent pedestrian IDs into the HRNet-W32 keypoint network, and output 17 skeletal keypoints, expressed as:

[0056] in, Let K be the two-dimensional coordinates of the Kth keypoint on the t-frame image. The corresponding confidence level; Temporal feature construction: For each pedestrian i, a sliding window of length L is set, typically 8-12 frames, to capture dynamic behavioral features within 0.25 seconds. The expression for the sliding window is: .

[0057] A dimension matrix is ​​constructed based on 17 key points of each frame image. This forms a spatiotemporal trajectory map.

[0058] Set normalized timestamp The time series feature matrix is ​​obtained. This increases sensitivity to changes in the pace and rhythm of behavior.

[0059] Calculate the displacement difference of key points in adjacent frames: , in, Let be the coordinates of the k-th keypoint in frame t.

[0060] Based on the above displacement difference characteristics, an enhanced temporal feature matrix is ​​constructed: ; Where: Faux = , indicating the introduced keypoint displacement difference feature; This represents a temporal feature set with a length of L frames, where each frame contains original keypoints and auxiliary features; L is the length of the sliding window; i represents the pedestrian's ID number.

[0061] Behavior classification: Based on the temporal feature matrix obtained from the above steps Temporal convolutional networks are used to output running probabilities. .

[0062] Specifically, a temporal convolutional network was trained using a publicly available skeleton dataset combined with on-site sampling data, and a weighted cross-entropy loss function was used to improve the network's recall capability for the "running" category. Furthermore, each update of a sliding window completes one inference step, continuously updating the pedestrian's current running confidence.

[0063] Velocity fusion determination: Take the center position difference of pedestrian i bounding boxes spaced 4 frames apart as the displacement, and calculate the trajectory velocity. The expression is:

[0064] in, Let be the coordinates of the center point of the i-th pedestrian bounding box in frame t.

[0065] Specifically, a running probability threshold can be set. Trajectory velocity threshold In this embodiment, a temporal convolutional network is used to combine the running probability and speed in the output judgment. At that time, it was determined that pedestrian i posed a risk of running, among which, For the pre-set running probability threshold, These are pre-set speed thresholds. Both of these thresholds can be adaptively updated based on the station's passenger flow environment.

[0066] The pedestrian category determination is based on multimodal visual recognition methods, combined with deep learning models such as human attribute recognition, estimation, and assistive tool detection, to achieve high-precision identification of pedestrian categories. Pedestrian categories include children, the elderly, people with disabilities, and people pushing strollers, based on gaze region matching.

[0067] The system compares key points of the human skeleton and pedestrian height with a threshold set by the system. If the result is within the threshold, and the key point distribution and height are within the normal range for adults, the person is identified as a normal adult. If the key points of the human skeleton and pedestrian height are below the threshold set by the system (e.g., about 1.2 to 1.3 meters), the person is identified as a child. The method of fusing age and height features can effectively distinguish between preschool children and short adult passengers, improving the accuracy of identification.

[0068] The elderly identification process includes: if the pedestrian's walking speed is less than the walking threshold and the posture stability is less than the posture threshold, then the pedestrian is determined to be an elderly person; The process of identifying persons with disabilities includes: combining pedestrian target detection models, head orientation detection models, and object detection models to identify persons with disabilities and their commonly used assistive devices, such as wheelchairs, canes, or walking aids. If typical structural features of a wheelchair (such as the seat frame and wheel outline) are detected next to or below a passenger, the passenger is identified as a person with a disability who uses a wheelchair. Similarly, detecting the use of canes, walking aids, or other similar devices can also help identify a person with mobility impairments. By identifying these assistive devices, the system can promptly classify the corresponding passengers as persons with disabilities.

[0069] The process of identifying a person pushing a stroller includes: combining pedestrian target detection models, head orientation detection models, and object detection models to identify the distinctive shape of the person pushing the stroller and the stroller itself, including the stroller's frame and wheels. Specifically, when the distinctive shape of the stroller is identified in an image, its spatial relationship is used to associate it with the person pushing it immediately behind it, determining that the passenger is the parent pushing the stroller. The association between the stroller and the adult needs to consider the distance, relative position, and consistency of movement direction between them to avoid misassigning the stroller to the wrong pedestrian.

[0070] S3. After processing the multimodal features, input them into the decision model for judgment to control the gate status.

[0071] The decision-making process of the decision-making model is as follows: S31. Determine whether the pedestrian has swiped their card. If the pedestrian has swiped their card, the decision model outputs a probability of passage of 1. If the pedestrian has not swiped their card, proceed to the pedestrian intention analysis stage. S32. Determine whether the pedestrian's gaze is fixed on the card-swiping area. If the pedestrian's gaze is fixed on the card-swiping area and the probability of running is not greater than running probability one, then the decision model outputs passage probability two; if the pedestrian's gaze is fixed on the card-swiping area and the probability of running is greater than running probability one, then the decision model outputs passage probability three. If a pedestrian's gaze is not fixed on the card-swiping area, and the probability of running is greater than the probability of running 2, the decision model outputs the probability of passage 4; if a pedestrian's gaze is not fixed on the card-swiping area, and the probability of running is not greater than the probability of running 2, the model proceeds to the pedestrian category recognition stage. S33. Based on S32, determine whether the pedestrian is a special passenger. If the pedestrian is a special passenger, the decision model outputs the overall passage probability as probability five. If the pedestrian is not a special passenger, the decision model outputs the passage probability as probability six. S34. Based on S31-S33, drive the gate to remain open or immediately close and trigger an alarm according to the passage probability output by the decision model; wherein, if the overall passage probability output by the decision model is less than a preset value, drive the gate to close and trigger an alarm. In addition, such as Figure 2As shown, the camera device is positioned above the existing turnstile channel to monitor the entrance direction of the turnstile channel from a top-down angle. The camera range covers the pedestrian passage area in the entrance direction, collecting real-time traffic image data. Specifically, in this embodiment, the camera device is a binocular RGB camera, which is externally mounted on the crossbeam or bracket above the turnstile channel. This allows the binocular RGB camera to overlook the entire passage area in the entrance direction, ensuring comprehensive monitoring of pedestrian behavior entering the turnstile.

[0072] Example 2 In this multimodal prediction method for pedestrian crossing intentions, the values ​​of the crossing probability and the running probability are set by the user and can be modified. In this embodiment, a set of crossing probability values ​​is provided for reference, as follows: S31. Determine whether the pedestrian has swiped their card. If the pedestrian has swiped their card, the decision model outputs an overall passage probability of 0.98. If the pedestrian has not swiped their card, proceed to the pedestrian intention analysis stage. S32. If the pedestrian does not swipe their card, determine whether the pedestrian's gaze is fixed on the card swiping area. If the pedestrian's gaze is fixed on the card swiping area and the probability of running is less than a pre-set running threshold of 0.5, the decision model outputs an overall passage probability of 0.9. If the pedestrian's gaze is fixed on the card swiping area and the probability of running is greater than the pre-set running threshold of 0.5, the decision model outputs an overall passage probability of 0.2. If a pedestrian's gaze is not fixed on the card-swiping area and the probability of running is greater than the pre-set running threshold, the decision model outputs an overall passage probability of 0.02; if a pedestrian's gaze is not fixed on the card-swiping area and the probability of running is less than the pre-set running threshold of 0.5, the pedestrian category recognition stage begins. S33. Based on S32, determine whether the pedestrian is a special passenger. If the pedestrian is a special passenger, the decision model outputs an overall passage probability of 0.9. If the pedestrian is not a special passenger, the decision model outputs a passage probability of 0.2. S34. Based on S31-S33, drive the gate to remain open or immediately close and alarm according to the passage probability output by the decision model; wherein, if the overall passage probability output by the decision model is less than 0.8, drive the gate to close and alarm. If not, then keep the gate open.

[0073] Example 3 Because the predictive device is an independent module and connects to the turnstile via a standardized interface, it can be integrated without modifying the main structure of the turnstile during installation and deployment, making it compatible with turnstile equipment from different manufacturers or models. Furthermore, the device possesses excellent versatility and scalability, making it suitable not only for rail transit turnstiles but also for other access control scenarios, such as office building access control and scenic area ticket gates, serving as an independent, additional intelligent monitoring unit to enhance the pedestrian access monitoring capabilities of existing systems.

[0074] In addition, such as Figure 2 , Figure 3 As shown, the prediction device provides three standard connection methods: power interface, network interface, and serial communication interface. These are connected to the power supply, data communication, and control signal interfaces reserved inside the gate device through corresponding power lines, network cables, and serial cables, respectively, thereby realizing the electrical power supply, data transmission, and control signal interaction between the prediction device and the gate.

[0075] Through these connections, the predictive device can obtain stable power from the turnstile and conduct bidirectional data communication with the main control system of the turnstile equipment (e.g., sending pedestrian detection results and receiving turnstile status information). It can also output control commands to cooperate with the turnstile in performing opening and closing actions.

[0076] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A multimodal prediction method for pedestrian crossing intentions, characterized in that, Based on image information acquired from camera equipment, multimodal features of pedestrians are extracted, and decision logic is used to determine whether the current pedestrian is allowed to pass. The specific process includes the following steps: S1. Obtain the image sequence of pedestrians entering the station, perform multi-target detection and tracking, identify each pedestrian ID, and form a bounding box sequence with consistent pedestrian IDs; S2. For each pedestrian, perform multimodal feature extraction, including card swiping action, gaze direction, running behavior, and pedestrian category; S3. After processing the multimodal features, input them into the decision model for judgment to control the gate status; The decision-making process of the decision-making model is as follows: S31. Determine whether the pedestrian has swiped their card. If the pedestrian has swiped their card, the decision model outputs a probability of passage of 1. If the pedestrian has not swiped their card, proceed to the pedestrian intention analysis stage. S32. Determine whether the pedestrian's gaze is fixed on the card-swiping area. If the pedestrian's gaze is fixed on the card-swiping area and the probability of running is not greater than running probability one, then the decision model outputs passage probability two; if the pedestrian's gaze is fixed on the card-swiping area and the probability of running is greater than running probability one, then the decision model outputs passage probability three. If a pedestrian's gaze is not fixed on the card-swiping area, and the probability of running is greater than the probability of running 2, the decision model outputs the probability of passage 4; if a pedestrian's gaze is not fixed on the card-swiping area, and the probability of running is not greater than the probability of running 2, the model proceeds to the pedestrian category recognition stage. S33. Based on S32, determine whether the pedestrian is a special passenger. If the pedestrian is a special passenger, the decision model outputs the overall passage probability as probability five. If the pedestrian is not a special passenger, the decision model outputs the passage probability as probability six. S34. Based on S31-S33, drive the gate to remain open or immediately close and trigger an alarm according to the passage probability output by the decision model; wherein, if the overall passage probability output by the decision model is less than a preset value, drive the gate to close and trigger an alarm. If not, then keep the gate open.

2. The multimodal prediction method for pedestrian walking intentions according to claim 1, characterized in that, S1. Obtain the image sequence of pedestrians entering the station, perform multi-object detection and tracking, identify each pedestrian ID, and form a bounding box sequence with consistent pedestrian IDs. The specific process is as follows: Pedestrian target detection: A lightweight object detection model is used to output pedestrian bounding boxes in each frame of the image. Where B represents the bounding box, t represents the frame number, and i represents the pedestrian ID; Pedestrian identity association: Pedestrian re-identification is performed using ReID technology, which extracts pedestrian features to form ReID feature vectors, used to distinguish different pedestrians. Pedestrian features include clothing, accessories, and body posture. The position / velocity of pedestrians is predicted using the Kalman filter algorithm, which is used to predict the position of pedestrians in the next frame; Multi-target tracking is performed using the ByteTrack algorithm, outputting the bounding box sequence of pedestrian i in T consecutive frames. The expression for the bounding box sequence is: .

3. The multimodal prediction method for pedestrian walking intentions according to claim 1, characterized in that, The recognition of card swipe actions includes the following steps: Set up corresponding ROI areas based on the card swiping area of ​​the turnstile equipment; Extract the two-dimensional coordinates and confidence scores of key points of the pedestrian skeleton in each frame of the image; In N consecutive frames, the two-dimensional coordinates of the pedestrian's wrist key points continuously fall within the ROI region, and the confidence level is not lower than the confidence threshold. If the pedestrian's elbow joint shows a tendency to move towards the ROI area, then the pedestrian is determined to have swiped a card.

4. The multimodal prediction method for pedestrian walking intentions according to claim 3, characterized in that, In each frame of the image, the HRNet-W32 network outputs the two-dimensional coordinates and confidence scores of 17 key points of the pedestrian, expressed as follows: ,in Let K be the two-dimensional coordinates of the Kth keypoint on the t-frame image. The corresponding confidence levels are: K=9 for the left wrist, k=10 for the right wrist, K=7 for the left elbow, and K=8 for the right elbow.

5. The multimodal prediction method for pedestrian walking intentions according to claim 1, characterized in that, The identification of eye direction includes the following steps: Based on the extracted facial landmark distribution and combined with predefined 3D average face model coordinates, the rotation matrix of the pedestrian's head relative to the camera device is calculated using OpenCV's PnP or POSIT algorithm. and translation vector To obtain the pedestrian's head pose, where the rotation matrix... Including yaw angle Pitch angle Roll angle ; The system spatially matches the pedestrian's head posture with the card-swiping area determined by the gate device to determine whether the pedestrian's gaze is pointing towards the card-swiping area of ​​the gate device. When the pedestrian's gaze is continuously directed towards the card-swiping area for a preset minimum time, it is determined that the pedestrian has the intention to swipe the card.

6. The multimodal prediction method for pedestrian walking intentions according to claim 5, characterized in that, Spatial matching of pedestrian head posture and card-swiping area is achieved by comparing angle thresholds. The specific process is as follows: Set the reference value for the head yaw angle required for pedestrians to normally look at the card swiping area. Head pitch angle reference value ; Set the maximum allowable deviation threshold, including yaw angle deviation. Pitch angle deviation ; When pedestrian i's head is facing the desired direction and At that time, it is determined that pedestrian i's line of sight falls into the direction of the card swiping area.

7. The multimodal prediction method for pedestrian walking intentions according to claim 1, characterized in that, The determination of running behavior is based on the coordinates and confidence level of the target pedestrian's skeletal key points, and specifically includes the following steps: Skeletal keypoint extraction: The bounding box sequence is input into the HRNet-W32 keypoint network, which outputs 17 skeletal keypoints, expressed as follows: in, For the first K The two-dimensional coordinates of key points on frame t of the image. The corresponding confidence level; Temporal feature construction: A sliding window of length L is set for each pedestrian i to capture dynamic behavioral features within the sliding window. The expression for the sliding window is: ; A dimension matrix is ​​constructed based on 17 key points of each frame image. This forms a spatiotemporal trajectory map; Set normalized timestamp The time series feature matrix is ​​obtained. ; Calculate the displacement difference of key points in adjacent frames: , in, Let k be the coordinates of the k-th keypoint in frame t. Based on the above displacement difference characteristics, an enhanced temporal feature matrix is ​​constructed: ; Where: Faux = , indicating the introduced keypoint displacement difference feature; This represents a temporal feature set with a length of L frames, where each frame contains original keypoints and auxiliary features; L is the length of the sliding window; i is the pedestrian number; Behavior classification: based on temporal feature matrix Temporal convolutional networks are used to output running probabilities. ; Velocity fusion determination: Take the center position difference of pedestrian i bounding boxes spaced 4 frames apart as the displacement, and calculate the trajectory velocity. The expression is: in, Let be the coordinates of the center point of the i-th pedestrian bounding box in frame t; when At that time, it was determined that pedestrian i posed a risk of running, among which, The running probability threshold, This is the speed threshold.

8. The multimodal prediction method for pedestrian walking intentions according to claim 1, characterized in that, The judgment of a pedestrian's identity is based on visual recognition. Specifically, the key points of the human skeleton and the pedestrian's height are compared with the thresholds set by the system. If the judgment result is within the set thresholds, and the distribution of key points and height are within the normal range for adults, then the pedestrian is judged to be a normal adult. If the key points of the human skeleton and the height of the pedestrian are below the threshold set by the system, the person is identified as a child. If a pedestrian's walking speed is less than the walking threshold and their posture stability is less than the posture threshold, then the pedestrian is classified as an elderly person. By combining pedestrian target detection model, head orientation detection model and object detection model, we can identify people with disabilities and their commonly used assistive tools, including wheelchairs, canes or walking aids. If an assistive tool is identified, the pedestrian is classified as a person with a disability. By combining pedestrian target detection model, head orientation detection model and object detection model, the system identifies the characteristics and shape of the person pushing the stroller and the stroller, including the stroller frame and wheels. When the stroller is detected, the pusher is associated with the stroller, and the pedestrian category is determined to be a person pushing a stroller.

9. A multimodal prediction device for pedestrian crossing intentions, characterized in that, For implementing the multimodal prediction method for pedestrian walking intentions as described in any one of claims 1-8, the prediction device comprises: The camera equipment is installed above the turnstile channel, and the camera range covers the pedestrian passage area in the direction of entering the station, for real-time collection of traffic video data; A multimodal feature extraction module is used to extract multimodal features of pedestrians based on image information acquired by a camera device; the multimodal feature extraction module includes a card swiping action recognition unit, a gaze direction recognition unit, a running behavior recognition unit, and a pedestrian category recognition unit; The decision model, based on decision logic, judges whether to allow the current pedestrian to pass based on the feature information extracted by the multimodal feature extraction module; The main control module is used to control the state of the gate device according to the output of the decision model. The gate is in the normally open state. If the passage probability output by the decision model is less than the preset value, the gate is driven to close and an alarm is triggered. If the passage probability output by the decision model is not less than the preset value, the gate is kept in the normally open state.

Citation Information

Patent Citations

  • System and method for recording a person in a region of interest

    CA2635790A1

  • Human body activity identification method, device and equipment and readable storage medium

    CN116229579A

  • Subway passenger abnormal behavior video description method based on skeleton key point knowledge enhancement

    CN117557945A

  • Pedestrian crossing intention prediction method based on multi-self-attention mechanism fused with multi-source information

    CN117765568A

  • Pedestrian abnormal behavior video identification method based on attitude estimation

    CN117789255A

Cited By

  • Linkage calling pass checking method of clearance camera and gate equipment

    CN121708682A