Station passenger falling behavior identification method fusing human-material master-slave binary information

By integrating the identification method of binary information of people and objects, and using quantum entanglement theory and von Neumann entropy, the coupling strength between passengers and their belongings is identified, which solves the problems of false alarms and missed alarms in the detection of falling behavior of station passengers and achieves high-precision falling behavior identification.

CN120673481APending Publication Date: 2025-09-19ZHEJIANG RAIL TRANSIT OPERATION MANAGEMENT GROUP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510853942.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

The existing station passenger fall behavior detection system is prone to false alarms and missed alarms in complex environments, and it is difficult to meet the requirements of real-time, accuracy and robustness, especially under the influence of dynamic coupling between passengers and their belongings.

Method used

An identification method that integrates binary information of human and object master and slave is adopted. The Yolov8-Pose model and Yolov8-Detection+object posture regression network are used to identify the key points of the human body and objects. The vertical jitter intensity and posture tilt angle are calculated. The entangled state density operator is constructed and the von Neumann entropy is used to characterize the coupling strength between the human body and the object. Real-time recognition is performed in combination with the offline learned recognition threshold.

Benefits of technology

It achieves accurate recognition of passenger falling behavior in complex environments, improves the accuracy and reliability of the recognition system, and has good robustness and engineering promotion value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005465571420000023
    Figure BDA0005465571420000023
  • Figure BDA0005465571420000024
    Figure BDA0005465571420000024
  • Figure BDA0005465571420000033
    Figure BDA0005465571420000033
Patent Text Reader

Abstract

Along with the rapid development of urban rail transit, a station safety monitoring system puts forward higher requirements for the real-time performance and accuracy of passenger falling behavior recognition. The invention provides a station passenger falling behavior identification method fusing human-material master-slave binary information. Accurately extracting key skeleton points of a human trunk based on a Yov8-Pose model of station scene fine adjustment, and calculating a vertical direction variable quantity and a body posture inclination angle; key skeleton points of a carry-on object of a passenger are identified through target detection and an object attitude estimation network, and the vertical direction and topological structure variation of the key skeleton points are calculated; the method comprises the following steps of: mapping transient characteristics of a human body and an object into a binary quantum state by utilizing a quantum entanglement theory, calculating von Neumann entropy to quantify dynamic coupling strength between the human body and the object, and determining an entropy threshold value through a Bayesian decision boundary so as to accurately judge a tumble behavior. According to the invention, the accuracy and real-time performance of station passenger falling behavior identification can be obviously improved, and the emergency response efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of abnormal behavior recognition, and in particular to a method for recognizing falling behavior of passengers at a station. Background Art

[0002] With the rapid development of urban rail transit and public transportation networks, stations are experiencing dense crowds and complex environments. Accidental falls, such as slips and trips, often occur while passengers are waiting for trains, boarding or exiting trains, or walking quickly. This not only poses a serious threat to passengers' own safety but can also lead to secondary risks such as platform congestion and delayed rescue efforts. Currently, most fall detection systems monitor only vertical changes in human posture or motion trajectory. These systems identify falls by determining the magnitude of a drop in height from the top of the head or key skeletal points within a short period of time, while ignoring the dynamic coupling between passengers and the objects they carry (such as backpacks and luggage). These methods are prone to false positives and false negatives when encountering occlusion, interference from multiple people, or swaying luggage, making them difficult to meet the stringent requirements of station safety management for real-time, accuracy, and robustness. Summary of the Invention

[0003] In order to solve the above problems, the present invention provides a station passenger fall behavior recognition method that integrates human-object master-slave binary information. The method first uses the station monitoring equipment to collect high-frame rate continuous video of the passengers, and combines the scene understanding algorithm to automatically segment the long video into two types of scenes: "fall" and "normal walking", and construct a labeled dataset for training and evaluation; then, based on the Yolov8-Pose model and the Yolov8-Detection+ object posture regression network, the key points of the human body and objects are extracted in parallel, and the key points of the human torso (including the top of the head, neck, shoulders, and hips) are calculated. The vertical jitter intensity and overall plane tilt angle are used to characterize the transient changes of the main state of the human body; at the same time, the vertical jitter intensity and the topological structure angle change of the skeletal point connection line of the key points of the object (the four corners of the package and the handle connection point) are calculated to characterize the transient changes of the slave state of the object; then, the above two sets of master-slave features are normalized and mapped into binary quantum states, and the entangled state density operator is constructed to characterize the coupling strength between the human body and the object through von Neumann entropy; finally, the entropy value calculated online in real time is compared with the optimal recognition threshold obtained by offline learning. When the entropy value exceeds the threshold, the passenger fall alarm is triggered and the position of the fallen passenger is highlighted on the monitoring interface.

[0004] According to one aspect of the present application, a method for identifying falling behavior of passengers at a station by integrating binary information of person and object:

[0005] S1. First, cut the original video into image data, perform data enhancement on the data, and adjust Yolov8 parameters.

[0006] Specifically, step S1 is as follows:

[0007] S101. First, a station monitoring system is used to obtain a continuous video stream of passengers, and video clips containing falling behaviors and normal walking situations are screened to construct an original video database.

[0008] S102: Extract each video frame by frame at a fixed frame rate f to generate an image sequence. The images are uniformly normalized (e.g., scaled to 640×640 pixels) and color-normalized to meet the input requirements of the deep learning model.

[0009] S103. Apply image enhancement methods based on this, including but not limited to brightness perturbation, horizontal flipping, random cropping, and affine transformation operations, to expand sample diversity and improve model generalization capabilities.

[0010] S104. Split the preprocessed images into training and test sets according to the ratio R∈[0.7,0.9] to provide data support for the training and verification of the skeleton and object recognition model.

[0011] S2. Based on the image sequence processed by S1, identify the key skeleton points of the human body and extract the torso posture features.

[0012] Specifically, step S2 is as follows:

[0013] S201: Input the preprocessed image into the Yolov8-Pose model, which has been fine-tuned for the station scene. This model, based on the CSPDarknet backbone network, uses an FPN+PAN architecture to extract multi-scale features. It also adds a pose estimation branch at the output to predict a 2D coordinate heatmap of key points on the human body and regress the key point locations.

[0014] S202, the system focuses on extracting the five key skeleton points of the human body: the top of the head, the neck, the left and right shoulders, and the left and right hips, forming a set

[0015] S203, project each key point from the image coordinates (u, v) to three-dimensional coordinates (x, y, z) through the camera depth estimation module. The height change of the same skeleton point between adjacent frames t and t+1 is Perform statistics and define the vertical shaking intensity of the human body as:

[0016]

[0017] S204, further, fit the skeleton point set of the current frame into a spatial plane, set its normal vector as n, and the horizontal plane n0 = [0, 1, 0] T The angle θ(t) represents the overall tilt of the torso posture:

[0018]

[0019] This constitutes the transient feature vector of the human body:

[0020] f h (t)=[B vert (t),θ(t)]

[0021] S3. Identify the key skeleton points of the objects carried by the passengers and extract their posture and topological structure changes.

[0022] Specifically, step S3 is as follows:

[0023] S301, apply Yolov8 target detection model to locate the position of the passenger's belongings, extract the target ROI area, and input it into the posture estimation network to identify the key skeleton point set This includes the four corners of the object and the handle attachment points.

[0024] S302: Similarly, calculate the vertical height difference of key points in adjacent frames and define the vertical jitter intensity B of the object in the form of root mean square. obj_vert (t)

[0025] S303: Construct structural units (such as four-corner connecting lines) between key skeleton points of the object and calculate the angle φ between adjacent structural units. i , compare the angle difference between the current frame and the previous frame:

[0026]

[0027] Then quantify the change in the overall structural topology:

[0028]

[0029] The transient eigenvector of the object is:

[0030] f o (t)=[V obj_vert (t),T topo (t)]

[0031] S4. Based on the human-object features extracted in S2 and S3, the entangled quantum state is constructed and the von Neumann entropy is calculated to complete the training of the fall behavior recognition standard.

[0032] Specifically, step S4 is as follows:

[0033] S401: For all labeled video samples in the training set (including clear "fall" and "non-fall" labels), execute the key feature extraction processes of human body and objects described in S2 and S3 respectively to obtain the master-slave feature vector pair f for each frame.h (t) and f o (t).

[0034] S402, normalize the above eigenvectors to construct a two-dimensional complex amplitude quantum state, and form an entangled state representation of the master-slave quantum system:

[0035]

[0036] S403. Calculate the overall density matrix of the entangled state ρ(t)=|Ψ(t)><Ψ(t)|, and perform a detour operation on the object subsystem to obtain the reduced density matrix of the human body subsystem:

[0037] ρ h (t) = Tr o (ρ(t))

[0038] S404. Use the von Neumann entropy formula to calculate the coupling strength at each moment:

[0039] S(t)=-Tr(ρ h (t)logρ h (t))

[0040] S405. Based on the labels of each video in the training set, all entropy values ​​are clustered and density modeled to form the entropy probability distribution corresponding to the "falling state" and "non-falling state", namely: Entropy distribution of non-falling state: S normal ~P0(S); Entropy distribution of falling state: S fall ~P1(S)

[0041] S406, using the Bayesian decision boundary to determine the optimal recognition threshold τ * , that is, when the entropy value exceeds the threshold, it is identified as a fall behavior:

[0042] like

[0043] S407, the above standard entropy threshold τ * It is solidified as part of the model and used as the basis for recognition in the subsequent online real-time recognition system.

[0044] Beneficial effects of the present invention: The present invention proposes a method for identifying falling behaviors of passengers in stations by integrating the master-slave binary information of people and objects. It applies quantum entanglement theory to the field of video surveillance for the first time. By modeling the transient motion characteristics of the human body and the objects carried with them as master-slave quantum states, the von Neumann entropy is introduced to characterize their dynamic coupling strength, thereby achieving accurate identification of the falling behaviors of passengers. First, the Yolov8 algorithm is used to identify the trunk skeletal points of the passenger and the key skeletal points of the objects carried with them, and the vertical jitter intensity and posture tilt angle of the human body, the jitter amplitude and topological structure change of the object are extracted respectively to construct the master-slave feature vector. Secondly, the features are mapped into binary entangled states, and the von Neumann entropy distribution in the falling and non-falling states is statistically analyzed during the training phase. The optimal recognition threshold is determined by combining ROC analysis. Finally, in the online recognition process, the entanglement entropy of the human body and the object is calculated in real time and compared with the standard threshold, so as to accurately identify the falling behavior and trigger an alarm. This method not only has strong sensitivity to slight posture changes, but also entropy as a global quantity has good robustness in dealing with video noise and detection errors. It significantly improves the accuracy and reliability of the passenger fall behavior recognition system in the complex environment of the station, and has good practical application prospects and engineering promotion value.

[0045] From the following detailed description of the present invention in conjunction with the accompanying drawings, those skilled in the art will become more aware of the above and other objects, advantages and features of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 This is a flowchart for a passenger fall behavior recognition method that integrates human-object master-slave information. This diagram illustrates the core processing flow of this method, including video acquisition and preprocessing, human skeletal key point recognition and posture analysis, object key skeletal point recognition and topological structure change extraction, construction of human-object master-slave quantum states, and calculation of von Neumann entropy for fall recognition. Each step forms a closed-loop recognition chain from raw data perception, structural feature extraction, quantum modeling, to final decision-making, demonstrating strong engineering practicality and theoretical scalability.

[0047] Figure 2 This is a flow chart showing the complete process of identifying key skeletal points and calculating their vertical variation and their angle variation with a plane, according to one embodiment of the present application.

[0048] Figure 3 This is a flow chart showing the complete process of identifying key skeleton points of an object and calculating vertical changes and topological changes of the object in accordance with one embodiment of the present invention.

[0049] Figure 4The quantum entanglement theory of one embodiment of this application introduces the von Neumann entropy calculation process. This figure shows the complete process of calculating the von Neumann entropy of key skeletal points and calculated changes in the collected human body and objects.

[0050] Figure 5 This is an overall flow chart of an embodiment of the present application. This figure shows the complete system structure from video data extraction to fall behavior recognition. The left side of the figure is the training stage, in which the system extracts key skeletal points of the human body and objects from the original video, calculates vertical jitter, posture tilt and topological structure change respectively, and constructs a master-slave state based on quantum entanglement theory. The human-object coupling strength is characterized by von Neumann entropy, and finally the standard threshold for fall behavior recognition is obtained by statistics. The right side is the online recognition stage, in which the system receives image frames in real time and repeats the above calculation process, compares the current entropy value with the standard threshold to determine whether a fall behavior has occurred. This figure clearly reflects the functional logic and technical path of the present invention in the two stages of training and recognition. DETAILED DESCRIPTION

[0051] The implementation process mainly includes four steps: cutting the original video data to generate an image dataset, identifying the key skeletal points of the human body and calculating the vertical change and the change in angle with the plane, identifying the key skeletal points of the object and calculating the vertical change and the topological structure change, and calculating the von Neumann entropy based on the above characteristics and determining the standard entropy value accordingly.

[0052] S1. First, cut the original video into image data, perform data enhancement on the data, and adjust Yolov8 parameters.

[0053] Specifically, step S1 is as follows:

[0054] S101. Preprocess the continuous video stream data collected by the station monitoring system. First, perform scene segmentation on the long-term video and extract typical segments containing normal walking and falling behaviors of passengers to form a structured video sample set.

[0055] S102. Extract frames from the video clip at a fixed frame rate of f = 25 fps and convert them into image sequences. Unify the image size to 640 × 640 and perform channel normalization, brightness normalization, and mean-variance normalization (Zero-Center Normalization) to ensure that the image has a consistent numerical feature distribution at the model input stage.

[0056] S103. Enhancement processing is performed at the image level, including: Random Crop: Randomly select window areas within the image to retain key targets and improve the position invariance of the model; Affine Transform: Perform linear transformations such as rotation, translation, and scaling on the image to improve the model's tolerance to structural deformation; Color Jitter: Perturb the brightness and saturation distribution in the HSV color space to improve the model's robustness to lighting changes; Horizontal Flip: Randomly mirror the image to improve the model's left-right symmetric generalization ability.

[0057] S104. Divide the samples into a training set and a test set according to a ratio of R = 0.7 to 0.9, for modeling and generalization performance verification of the skeleton recognition and entropy calculation model.

[0058] S2. Based on the image sequence processed by S1, identify the key skeleton points of the human body and extract the torso posture features.

[0059] Specifically, step S2 is as follows:

[0060] S201: Input the preprocessed image into the Yolov8-Pose model, which has been fine-tuned for the station scene. This model uses CSPDarknet as its backbone network, integrating a Feature Pyramid Network (FPN) and a Path Aggregation Network (PAN). This model achieves high-precision object detection and keypoint localization through multi-scale feature enhancement. Its pose estimation branch uses heatmap regression to output a probabilistic heatmap of 17 keypoints in COCO format.

[0061] S202, focus on selecting five key points of the torso: top of the head, neck, left shoulder, right shoulder, and hip, and record them as Each key point extracts the maximum activation value position from the heat map and combines it with the depth estimation result to be projected into a three-dimensional coordinate (x i ,y i ,z i ), the height difference of each point in adjacent frames is defined as: The sum of squares of consecutive frame differences is taken as the average jitter intensity: In addition, the maximum vertical jump amplitude and instantaneous speed indicators are also defined:

[0062] S203, further, the skeleton point set B of the current frame h Fit it into a space plane, let its normal vector be n, and the horizontal plane n0 = [0,1,0] T The angle θ(t) represents the overall tilt of the torso posture:

[0063] S204. Construct human body transient feature vector: f h (t)=[V vert (t),θ(t),V max (t)]

[0064] S3. Identify the key skeleton points of the objects carried by the passengers and extract their posture and topological structure changes.

[0065] Specifically, step S3 is as follows:

[0066] S301, use the target detection model Yolov8 to locate the passenger's personal objects (such as backpacks, suitcases) in the image and extract the ROI area. Input the object pose estimation network (which can be a Lightweight Pose Estimation model) to predict the key skeleton point set Such as the four corners of the package, the vertices of the pull rod and the connection points.

[0067] S302: The vertical height difference between each key skeleton point and adjacent frames: The vertical jitter intensity is obtained by calculating the root mean square:

[0068]

[0069] Calculate the maximum jitter value and the maximum velocity estimate:

[0070]

[0071] S303. Define the angle formed by the structural connection lines, such as the diagonal change of the package body and the topological disturbance amount:

[0072] S304. Construct the transient feature vector of the object:

[0073] f o (t)=[V obj_vert (t),T topo (t),V obj_max (t)]

[0074] S4: Based on the human-object features extracted in S2 and S3, the entangled quantum state is constructed and the von Neumann entropy is calculated to complete the training of the fall recognition standard.

[0075] Specifically, step S4 is as follows:

[0076] S401, the human body and the object feature vector f h (t), f o (t) Normalize the mapping to the quantum amplitude components of the master state and the slave state to construct the entangled state:

[0077]

[0078] where α ij (t) = ψ hi (t)·ψ oj (t).ψ hi , ψ oj are the amplitude components of the main and slave states respectively, and the normalization satisfies: ∑ i , j|α ij (t)| 2 = 1.

[0079] S402. Construct the total system density matrix: ρ(t) = |Ψ(t)><Ψ(t)| and perform a partial trace operation on the object subsystem to obtain the reduced density matrix of the human subsystem: ρ h (t) = Tr o (ρ(t))

[0080] S403. The von Neumann entropy is used to characterize the entanglement strength between the main-slave systems: S(ρ h (t)) = -Tr(ρ h (t)· logρ h (t))

[0081] If the spectral decomposition of ρ h is ρ h = ∑ k λ k |k><k|, then the entropy can be simplified to:

[0082] S = -∑ k λ k logλ k

[0083] The larger the entropy value, the more uncertain the states of the human body and the object, the stronger the correlation, and the maximum coupling perturbation during a fall.

[0084] S404. Estimate the entropy value distribution in the training set, calculate the density functions in the fall and non-fall states respectively: p fall (S), p non (S) and establish a Bayesian discriminant function:

[0085]

[0086] S405. Determine the recognition threshold τ * :

[0087] τ * = argmin S [π fall (1 - TPR(S)) + π non·FPR(S)]

[0088] This threshold will serve as the standard entropy boundary value for identifying falling behavior in actual operation of the system.

[0089] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0090] The above is only the most effective implementation scheme of the present invention. It should be pointed out that for ordinary technicians in this technical field, appropriate improvements and modifications can be made without departing from the working principle of the present invention. These improvements and modifications should also be considered as the scope of protection of the present invention.

Claims

1. A method for identifying passenger falls at a station by integrating binary information of person and object, characterized by The method comprises: Continuous video data of passengers is collected through station monitoring equipment. Long videos are cut into segments that include passengers falling and those walking normally. The video collection is divided into training and test sets to prepare for model training and evaluation. The Yolov8 algorithm is used to identify key human skeleton points. These key skeleton points are a set of trunk skeleton points (including the top of the head, neck, shoulders, and hips), whose spatial distribution remains relatively stable over a short period of time. For this trunk skeleton point set, the vertical change and the change in the angle between the trunk and the plane are calculated. The Yolov8 algorithm combines target detection with object pose estimation to identify key skeletal points of an object. These are a set of landmark points with a relatively fixed spatial shape (the four corners of the bag and the handle or shoulder strap connection points). The vertical change and topological structure change of these key skeletal points are calculated. The mapping of key skeletal points on the human body to those on the object is constructed into a master-slave binary quantum state representation of the human-object pair. Based on quantum entanglement theory, the entropy of the master-slave quantum state is introduced and compared with a fall determination threshold. When the entropy exceeds the threshold, it is determined that the passenger has fallen.

2. The method for identifying key human skeleton points and calculating vertical direction variation and angle variation with a plane according to claim 1 is characterized in that: When realizing the automatic recognition of human torso skeleton points, the station surveillance video is first preprocessed: the original video is extracted at a fixed frame rate, and each frame is scaled (for example, scaled to 640×640 pixels) and color-normalized. Then, it is input into the Yolov8-Pose model that has been pre-trained and fine-tuned in the station scene. This model uses CSPDarknet as the backbone network, extracts multi-scale features through a feature pyramid (FPN+PAN), and adds a posture regression branch to the detection head. This branch outputs the heat map and regression coordinates of 17 key points of the human body. Among them, we focus on the torso skeleton points such as the top of the head, neck, shoulders, and hips to form a set. In the inference phase, the model first searches for the maximum activation position of each heat map to obtain the two-dimensional pixel coordinates of each key bone point, and combines it with the depth estimation z obtained by camera calibration. i (t), and project it into three-dimensional coordinates (x i (t),y i (t),z i (t)). For each set of torso bone points at time point t, its transient change in the vertical direction is defined as the height difference between the same bone point in adjacent frames t-1 and t: Δz i (t)=z i (t)-z i (t-1) The overall vertical jitter intensity is quantified by calculating the root mean square of the height differences of all N points. The formula for calculating the root mean square of the height difference of all N points is: This indicator can effectively reflect the rapid displacement of the passenger's torso in the direction of gravity during the process from standing to falling. At the same time, in order to capture the tilt change between the torso and the ground, all the 3D key skeleton points {(x i ,y i ,z i )} Fit to the plane using the least squares method: z=αx+βy+d in: z=(z1,…,z N ) T By calculating the angle between the normal vector n = (-α, -β, 1) of the fitting plane and the normal vector k = (0, 0, 1) of the horizontal plane: That is, it quantifies the overall tilt of the passenger’s body posture. vert The continuous time series changes of (t) and the tilt angle θ(t) can accurately capture the transient characteristics of the fall process and provide stable and distinguishable main state input for the subsequent von Neumann entropy judgment based on quantum entanglement.

3. The method of combining target detection and object pose estimation according to claim 1, identifying key skeletal points of an object and vertical changes, and calculating changes in its topological structure, is characterized by: The Yolov8 object detection model, which has been fine-tuned for the station scene, is applied to each video frame to accurately locate the two-dimensional bounding box (ROI) of the passenger's hand-held object in the image. The ROI area is then input into the object pose estimation network, which outputs the two-dimensional coordinates of the key skeleton points of the object in the pixel plane and combines them with the depth estimation to obtain the three-dimensional coordinates (x i (t),y i (t),z i (t)), where the key skeleton points of the object include the four corners P1 to P4 of the package body and the handle connection point P5. The height difference of the key skeleton points of the same object in adjacent frames t-1 and t is: Δz i (t)=z i (t)-z i (t-1) The transient jitter intensity of the object in the vertical direction is quantified by calculating the root mean square of the height difference of all K key points of the object. The root mean square calculation formula for the height difference of all K key points of objects is: At the same time, according to the topological structure of the object, the adjacent key points (P a ,P b ) is considered as a structural unit for the connection formed in frame t, which is the same as the connection (P b ,P c ) together form an angle: The structural change is measured by the difference in the corresponding angle of the previous frame (t-1). The calculation formula for the corresponding angle difference is: Df abc (t)=|φ abc (t)-φ abc (t-1)| Then perform the root mean square operation on the angle differences of all M structural units: The transient change of the topological structure of the object can be obtained. and topological structure change T topo (t) will serve together as the input features for calculating the quantum entanglement entropy of the physical state.

4. The method of claim 1 for determining passenger fall behavior by introducing von Neumann entropy through quantum entanglement theory, characterized in that: This patent not only captures the violent movements of the human body during a fall, but also quantifies the dynamic coupling relationship between the human body and the objects carried with it to improve the accuracy and robustness of fall detection. The system introduces quantum entanglement theory to map the transient characteristics of the human body and the object into a pair of master-slave quantum states. First, for each time point t, the human body transient feature vector is extracted through the aforementioned Yolov8 human posture estimation and object posture detection modules: f h (t)=[V vert (t),θ(t)] And the object transient eigenvector: Among them, V vert (t) with The vertical jitter intensity of the human body and the object is quantified, θ and T topo The change of overall tilt and structural topology is measured. Then, these two eigenvectors are respectively passed through L 2 Normalize so that ∑ i |f i h (t)| 2 =1 and ∑ j |f j o (t)| 2 =1, the normalized component is taken as the amplitude of the quantum state. On this basis, a pure binary quantum entangled state is constructed: And define the overall density operator: ρ(t)=|Ψ(t)><Ψ(t)| By performing trace operations on the object subsystem: ρ h (t)=Tr o [ρ(t)] That is, a reduced density matrix that only describes the human body subsystem is obtained, and then the von Neumann entropy is used to characterize the coupling strength of the master-slave quantum state. S(t)=-Tr[ρ h (t)logρ h (t)] In the offline stage, this method uses a large number of video samples with fall / non-fall labels to calculate the corresponding entropy distribution, and determines the optimal judgment threshold S through methods such as ROC curve, clustering and Gaussian mixture model. th When running online, when the entropy value calculated in real time satisfies S(t)>S th , indicating that the level of entanglement between a person and their belongings has exceeded the normal walking range, the system identifies the passenger as having fallen, automatically triggers an alarm, and highlights the passenger's location on the monitoring interface. By converting human-object interaction characteristics into quantum state coupling, this module is not only sensitive to subtle posture changes but also inherently robust to local noise and detection errors due to the "global" nature of entropy. This significantly improves the accuracy and reliability of passenger fall detection at stations.