A personnel trajectory drawing method and system based on an AI algorithm
Patent Information
- Application Number
- CN202511030340.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2026-08-18
- Estimated Expiration
- 2045-07-25
AI Technical Summary
[0003]鉴于上述问题,本发明提供了一种基于AI算法的人员轨迹绘制方法及系统,解决多源异构数据融合困难导致的轨迹断裂及精度不足的问题
[0062] This invention provides a method and system for personnel trajectory mapping based on AI algorithms. The method includes: acquiring signals from positioning devices and monitoring video data; performing dynamic Kalman filtering denoising on the positioning device signals to obtain first trajectory data, and performing adaptive illumination compensation enhancement processing on the monitoring video data to obtain first video data; performing trajectory prediction compensation on signal interruption segments based on an LSTM neural network; extracting personnel motion posture features and scene semantic features; constructing a spatiotemporal correlation graph using a graph convolutional neural network to obtain enhanced trajectory data; performing multi-scale trajectory reconstruction based on an improved Transformer architecture to generate final personnel trajectory information; and outputting a visualized trajectory report including trajectory accuracy assessment, abnormal behavior markers, and electronic fence trigger events. This invention achieves deep fusion of multi-source heterogeneous data and accurate trajectory reconstruction, significantly improving the continuity and reliability of personnel trajectory mapping in complex scenarios.
Smart Images

Figure CN120932297B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and specifically to a method and system for drawing personnel trajectories based on AI algorithms. Background Technology
[0002] In the field of personnel positioning and trajectory analysis, real-time location tracking typically relies on the fusion processing of signals from positioning devices and video surveillance data. The collaborative analysis capability of multi-source heterogeneous data directly impacts the accuracy and robustness of trajectory mapping. Traditional methods primarily employ a single data source for trajectory reconstruction or achieve multimodal fusion through simple data stitching. However, due to the complexity and dynamism of real-world application scenarios, these methods struggle to effectively address the spatiotemporal alignment issues between multi-source data, resulting in insufficient trajectory reconstruction accuracy and limited adaptability. While some recent research has attempted to introduce machine learning algorithms for trajectory optimization, systematic solutions are still lacking for typical scenarios such as missing positioning signals and visual occlusion. Summary of the Invention
[0003] In view of the above problems, the present invention provides a method and system for personnel trajectory drawing based on AI algorithm, which solves the problems of trajectory breakage and insufficient accuracy caused by the difficulty of multi-source heterogeneous data fusion.
[0004] To achieve the above objectives, in a first aspect, this application provides a method for personnel trajectory mapping based on AI algorithms, comprising:
[0005] Collect positioning device signals and monitoring video data. The positioning device signals include the real-time location information of the current user, and the monitoring video data includes multi-view spatiotemporal continuous images.
[0006] The positioning device signal is subjected to a first preprocessing to obtain first trajectory data. The first preprocessing is configured as signal denoising processing based on dynamic Kalman filtering.
[0007] In addition, the surveillance video data is subjected to a second preprocessing to obtain the first video data, and the second preprocessing is configured as multi-target tracking enhancement processing based on adaptive illumination compensation;
[0008] The spatiotemporal continuity trajectory features are extracted from the first trajectory data, and trajectory prediction compensation is performed on the signal interruption segment based on the LSTM neural network to obtain the compensated trajectory data.
[0009] In addition, human motion posture features and scene semantic features are extracted from the first video data;
[0010] A spatiotemporal correlation graph is constructed by using a graph convolutional neural network to combine the compensated trajectory data, personnel motion posture features and scene semantic features to obtain enhanced trajectory data. The spatiotemporal correlation graph is configured as a trajectory-scene joint optimization model.
[0011] Based on the improved Transformer architecture, multi-scale trajectory reconstruction is performed on the enhanced trajectory data to obtain the final personnel trajectory information. The multi-scale trajectory reconstruction is configured as an end-to-end trajectory generation that integrates positioning signals, visual observations, and environmental topology.
[0012] In addition, based on the final personnel trajectory information, a visual trajectory report is generated that includes trajectory accuracy assessment, abnormal behavior marking, and electronic fence triggering events. The visual trajectory report is configured to support spatiotemporal data analysis that supports real-time monitoring and historical backtracking.
[0013] In some embodiments, the positioning device signal undergoes a first preprocessing step to obtain first trajectory data. The first preprocessing step is configured as signal denoising processing based on dynamic Kalman filtering, including:
[0014] An adaptive Kalman filter model based on the motion state of the positioning device is constructed. The state vector of the adaptive Kalman filter model includes the device position coordinates, the device motion speed, and the device acceleration. The observation noise covariance matrix of the adaptive Kalman filter model is dynamically adjusted according to the current signal strength of the positioning device. The process noise covariance matrix of the adaptive Kalman filter model is used to represent the error of motion state prediction.
[0015] The original positioning signal is coarsely filtered to remove outliers that exceed the physical motion constraints, which include the maximum velocity threshold and the maximum acceleration threshold.
[0016] Motion trend parameters are calculated using a sliding window mechanism, and the window size of the sliding window mechanism is adaptively adjusted according to the positioning update frequency.
[0017] The process noise covariance matrix is updated based on motion trend parameters, and the weight of the angle component in the process noise covariance matrix is increased when a turning feature is detected.
[0018] Perform Kalman prediction-update iteration to output smoothed first trajectory data. The first trajectory data includes a denoised three-dimensional coordinate sequence, motion state confidence at each time step, signal quality evaluation index, and outlier marker information.
[0019] In some embodiments, the surveillance video data undergoes a second preprocessing step to obtain first video data. The second preprocessing step is configured as multi-target tracking enhancement processing based on adaptive illumination compensation, including:
[0020] An adaptive compensation model based on illumination perception is constructed. The adaptive compensation model includes an illumination component estimation module, a reflection component calculation module, and a dynamic compensation module. The dynamic compensation module is used to adjust the compensation intensity parameters according to changes in scene illumination.
[0021] Multi-target tracking processing is performed on the compensated video data to output the first video data. The first video data includes a video frame sequence with illumination compensation, multi-target tracking results with identification tags, target motion state data, and tracking quality evaluation results. The multi-target tracking processing includes:
[0022] The video frames are processed by a people detection network, and the detection results are output, including the coordinates of the people detection box and the confidence level.
[0023] A multi-feature fusion algorithm is used to perform target association on the detection results. Target association includes motion state matching and appearance feature association.
[0024] In some embodiments, spatiotemporal continuity trajectory features are extracted from the first trajectory data, and trajectory prediction compensation is performed on the signal interruption segment based on an LSTM neural network to obtain compensated trajectory data, including:
[0025] A trajectory prediction model based on LSTM is constructed. The trajectory prediction model includes a trajectory feature extraction module and a trajectory prediction compensation module. The trajectory prediction compensation module is used to predict the trajectory of the interrupted segment based on motion features.
[0026] Trajectory prediction compensation is performed on the signal interruption segment, and the compensated trajectory data is output. The compensated trajectory data includes a complete three-dimensional coordinate sequence, position compensation markers, prediction confidence scores, and correction records. The trajectory prediction compensation includes:
[0027] Trajectory features are extracted from the first trajectory data, including position coordinate sequence, velocity change trend, and motion direction angle;
[0028] The trajectory features are input into the LSTM neural network, which outputs a sequence of predicted trajectory coordinates.
[0029] When the signal is restored, the predicted trajectory is verified and corrected against the actual positioning data.
[0030] In some embodiments, extracting human motion posture features and scene semantic features from the first video data includes:
[0031] A deep learning-based 3D pose estimation model is constructed. The 3D pose estimation model includes a 2D keypoint detection network and a 3D pose reconstruction network. The 2D keypoint detection network is used to detect the 2D coordinates of each joint of the human body from video frames. The 3D pose reconstruction network is used to map the 2D joint coordinates to 3D space and calculate the relative motion relationship between the joints.
[0032] The first video data is processed by a three-dimensional pose estimation model. Human body detection and joint point localization are performed on the input video frames of the first video data, and the two-dimensional joint point coordinates including the main body parts such as the head and limbs are output.
[0033] Based on multi-view geometric constraints and prior knowledge of human skeleton, the coordinates of two-dimensional joint points are reconstructed into three-dimensional spatial coordinates.
[0034] Calculate the motion changes of joints between adjacent video frames and output the human motion posture features, which include a three-dimensional spatial coordinate sequence, joint angle changes, and overall motion velocity vector.
[0035] A scene semantic understanding model is constructed, which includes a scene classification network and a semantic segmentation network. The scene classification network is used to identify the scene category of the monitored area, and the semantic segmentation network is used to label the pixel-level boundaries of each functional area in the scene.
[0036] The first video data is processed by a scene semantic understanding model. The input video frames of the first video data are classified into scene partitions, which include corridors, rooms and stairwells.
[0037] Perform pixel-level segmentation on the rooms in the scene partition, and output a segmentation mask containing entrances / exits, rest areas, and work areas to obtain the segmentation result;
[0038] A scene topology graph is constructed based on the segmentation results, and scene semantic features are generated.
[0039] In some embodiments, a graph convolutional neural network is used to construct a spatiotemporal correlation graph from the compensated trajectory data, personnel motion posture features, and scene semantic features to obtain enhanced trajectory data, including:
[0040] A spatiotemporal correlation model based on graph neural networks is constructed. The spatiotemporal correlation model includes a node feature encoder and a graph convolution operation module. The node feature encoder is used to uniformly encode the compensated trajectory data, personnel motion posture features and scene semantic features into node feature vectors. The graph convolution operation module is used to model the spatiotemporal correlation relationship between nodes.
[0041] Using trajectory points, human body joints, and scene partitions as graph nodes, an initial graph structure containing spatially adjacent edges and temporally continuous edges is established.
[0042] Based on the scene topology, establish personnel-scene interaction edges, which include spatial inclusion relationships and movement direction relationships, and update personnel-scene interaction edges to spatial adjacency edges;
[0043] Additionally, establish trajectory continuity edges along the time dimension, which include velocity consistency and motion smoothness constraints, and update the trajectory continuity edges to the time continuity edges;
[0044] Generate the final graph structure;
[0045] The features of neighboring nodes in the final graph structure are aggregated by multi-layer graph convolution operation. The first weight coefficient of spatially adjacent edges and the second weight coefficient of temporally continuous edges are dynamically adjusted based on the attention mechanism, and the output is enhanced trajectory data containing spatial correction and temporal smoothing.
[0046] Enhanced trajectory data includes refined trajectory coordinates that integrate multi-source features, annotations of personnel-scene interaction states, trajectory reliability scores, and abnormal behavior detection markers.
[0047] In some embodiments, multi-scale trajectory reconstruction is performed on the enhanced trajectory data based on an improved Transformer architecture to obtain the final personnel trajectory information, including:
[0048] A multi-scale trajectory reconstruction model based on an improved Transformer is constructed. The multi-scale trajectory reconstruction model includes a spatiotemporal feature encoding module and a trajectory generation module. The spatiotemporal feature encoding module is used to extract trajectory features at different time and spatial scales, and the trajectory generation module is used to output the optimized trajectory sequence.
[0049] In the time dimension, long-time windows, medium-time windows, and short-time windows are used to extract motion features respectively;
[0050] In the spatial dimension, location features are extracted by establishing local motion neighborhoods and global scene ranges respectively;
[0051] The correlation weights between motion features and position features are calculated using a spatiotemporal attention mechanism to generate correlation weight information.
[0052] Based on the association weight information, a feature fusion network is used to perform feature fusion, and the final personnel trajectory information is output.
[0053] In some embodiments, a visual trajectory report is generated based on the final personnel trajectory information, including trajectory accuracy assessment, abnormal behavior markers, and electronic fence trigger events, comprising:
[0054] A visualization report generation model is constructed, which includes a trajectory analysis unit and an event marking unit. The trajectory analysis unit is used to calculate trajectory accuracy indicators, and the event marking unit is used to identify abnormal behavior and electronic fence events.
[0055] The trajectory accuracy assessment is calculated by the trajectory analysis unit, and the trajectory accuracy assessment result is obtained. The trajectory accuracy assessment result includes the positioning error index and the trajectory smoothness index.
[0056] Generate accuracy visualization markers based on trajectory accuracy assessment results;
[0057] Additionally, abnormal behavior events and electronic fence triggering events are detected through event tagging units to obtain detection results;
[0058] Generate behavioral anomaly markers and regional violation markers based on the detection results;
[0059] A visual trajectory report is generated based on accuracy visualization markers, anomaly markers, and regional violation markers.
[0060] In a second aspect, the present invention also provides a personnel trajectory mapping system based on an AI algorithm, which is applicable to the method described in the first aspect.
[0061] Unlike existing technologies, the above technical solution has the following beneficial effects:
[0062] This invention provides a method and system for personnel trajectory mapping based on AI algorithms. The method includes: acquiring signals from positioning devices and monitoring video data; performing dynamic Kalman filtering denoising on the positioning device signals to obtain first trajectory data, and performing adaptive illumination compensation enhancement processing on the monitoring video data to obtain first video data; performing trajectory prediction compensation on signal interruption segments based on an LSTM neural network; extracting personnel motion posture features and scene semantic features; constructing a spatiotemporal correlation graph using a graph convolutional neural network to obtain enhanced trajectory data; performing multi-scale trajectory reconstruction based on an improved Transformer architecture to generate final personnel trajectory information; and outputting a visualized trajectory report including trajectory accuracy assessment, abnormal behavior markers, and electronic fence trigger events. This invention achieves deep fusion of multi-source heterogeneous data and accurate trajectory reconstruction, significantly improving the continuity and reliability of personnel trajectory mapping in complex scenarios.
[0063] The above description of the invention is merely an overview of the technical solution of this application. In order to enable those skilled in the art to better understand the technical solution of this application and to implement it based on the description and drawings, and to make the above-mentioned objectives and other objectives, features and advantages of this application easier to understand, the following description is provided in conjunction with the specific embodiments and drawings of this application. Attached Figure Description
[0064] The accompanying drawings are only used to illustrate the principles, implementation methods, applications, features, and effects of specific embodiments of the present invention and other related contents, and should not be considered as limitations on this application.
[0065] In the accompanying drawings of the instruction manual:
[0066] Figure 1 This is a flowchart illustrating steps S101 to S105 of the drawing method described in the specific implementation.
[0067] Figure 2 The diagram illustrates steps S201 to S205 of the drawing method described in the specific implementation. Detailed Implementation
[0068] To illustrate the possible application scenarios, technical principles, implementable specific solutions, and achievable objectives and effects of this application in detail, the following description, in conjunction with the listed specific embodiments and accompanying drawings, provides a detailed explanation. The embodiments described herein are merely illustrative of the technical solutions of this application and are therefore intended to limit the scope of protection of this application.
[0069] In this document, the term "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The term "embodiment" appearing in various places throughout the specification does not necessarily refer to the same embodiment, nor does it specifically limit its independence or connection with other embodiments. In principle, in this application, as long as there are no technical contradictions or conflicts, the technical features mentioned in each embodiment can be combined in any way to form corresponding implementable technical solutions.
[0070] Unless otherwise defined, the technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the use of related terms herein is merely for the purpose of describing particular embodiments and is not intended to limit this application.
[0071] In the description of this application, the term "and / or" is used to describe the logical relationship between objects, indicating that three relationships can exist. For example, A and / or B means: A exists, B exists, and A and B exist simultaneously. Additionally, the character " / " in this document generally indicates that the preceding and following objects have an "or" logical relationship.
[0072] In this application, terms such as “first” and “second” are used only to distinguish one entity or operation from another, and do not necessarily require or imply any actual quantity, hierarchy or order relationship between these entities or operations.
[0073] Without further limitations, the use of terms such as “comprising,” “including,” “having,” or other similar open-ended expressions in this application is intended to cover non-exclusive inclusion, which does not exclude the presence of additional elements in a process, method, or product that includes the stated elements, such that a process, method, or product that includes a list of elements may include not only those defined elements but also other elements not expressly listed, or elements inherent to such a process, method, or product.
[0074] Similar to the understanding in the Examination Guidelines, in this application, expressions such as "greater than," "less than," and "exceeding" are understood to exclude the stated number; expressions such as "above," "below," and "within" are understood to include the stated number. Furthermore, in the description of the embodiments in this application, "multiple" means two or more (including two), and similar expressions related to "multiple" are also understood in this way, such as "multiple groups" and "multiple times," unless otherwise explicitly specified.
[0075] In the description of the embodiments of this application, the space-related expressions used, such as "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "vertical," "top," "bottom," "inner," "outer," "clockwise," "counterclockwise," "axial," "radial," and "circumferential," indicate the orientation or positional relationship based on the orientation or positional relationship shown in the specific embodiments or drawings. They are only for the purpose of describing the specific embodiments of this application or for the reader's understanding, and do not indicate or imply that the device or component referred to must have a specific position, a specific orientation, or be constructed or operated in a specific orientation. Therefore, they should not be construed as limitations on the embodiments of this application.
[0076] Please see Figure 1 In a first aspect, this embodiment provides a method for drawing personnel trajectories based on AI algorithms, including:
[0077] S101. Collect positioning device signals and monitoring video data. The positioning device signals include the real-time location information of the current user, and the monitoring video data includes multi-view spatiotemporal continuous images.
[0078] S102. Perform a first preprocessing on the positioning device signal to obtain first trajectory data. The first preprocessing is configured as signal denoising processing based on dynamic Kalman filtering.
[0079] In addition, the surveillance video data is subjected to a second preprocessing to obtain the first video data, and the second preprocessing is configured as multi-target tracking enhancement processing based on adaptive illumination compensation;
[0080] S103. Extract the spatiotemporal continuity trajectory features from the first trajectory data, and perform trajectory prediction compensation for the signal interruption segment based on the LSTM neural network to obtain the compensated trajectory data.
[0081] In addition, human motion posture features and scene semantic features are extracted from the first video data;
[0082] S104. A spatiotemporal correlation graph is constructed by using a graph convolutional neural network to combine the compensated trajectory data, personnel motion posture features and scene semantic features to obtain enhanced trajectory data. The spatiotemporal correlation graph is configured as a trajectory-scene joint optimization model.
[0083] S105. Multi-scale trajectory reconstruction is performed on the enhanced trajectory data based on the improved Transformer architecture to obtain the final personnel trajectory information. The multi-scale trajectory reconstruction is configured as an end-to-end trajectory generation that integrates positioning signals, visual observations and environmental topology.
[0084] In addition, based on the final personnel trajectory information, a visual trajectory report is generated that includes trajectory accuracy assessment, abnormal behavior marking, and electronic fence triggering events. The visual trajectory report is configured to support spatiotemporal data analysis that supports real-time monitoring and historical backtracking.
[0085] In step S101, the collected positioning device signal refers to the two-dimensional coordinate sequence acquired in real time through short-term positioning devices such as personnel badges and wristbands, used in conjunction with positioning base stations. This is particularly suitable for precise positioning in indoor environments such as the bidding area. The monitoring video data refers to the multi-view video streams collected by security cameras deployed in key areas such as the bidding area, overnight bidding rest area, and sample room. Its spatiotemporal continuous image characteristics are not only used for trajectory drawing but also support the realization of video behavior analysis and early warning capabilities, such as mobile phone call behavior analysis and rest area departure behavior analysis. The positioning device signal and monitoring video data are synchronized through a unified timestamp, providing a data foundation for subsequent electronic fence triggering and abnormal behavior marking.
[0086] In step S102, the dynamic Kalman filter used in the first preprocessing is specifically optimized for signal interference in the complex electromagnetic environment of the evaluation area, effectively suppressing positioning drift caused by the dense distribution of electronic devices. The adaptive illumination compensation algorithm in the second preprocessing focuses on solving typical illumination problems such as reflections from glass curtain walls and insufficient nighttime lighting in the evaluation area, ensuring the image quality required for video behavior analysis and early warning. The first trajectory data obtained after preprocessing retains the electronic fence triggering capability of the original positioning signal, while the first video data meets the detection accuracy requirements for behavior analysis scenarios such as device non-return warnings.
[0087] In step S103, the trajectory prediction compensation module of the LSTM neural network is specifically trained for signal interruption scenarios such as experts briefly leaving the evaluation area. Its prediction results will serve as a reference for handling suspected risk information. The personnel movement posture features extracted from the video data include upper limb movement patterns unique to answering and making phone calls. Scene semantic features clearly distinguish the spatial boundary features of electronic fence-controlled areas such as the evaluation area and sample rooms. Furthermore, the scene semantic features include a road network topology extracted from architectural CAD drawings, which includes corridor connectivity and access control point location information, used to constrain the trajectory prediction output of the LSTM neural network.
[0088] In step S104, the constructed spatiotemporal correlation map specifically enhances the semantic annotation of nodes in the electronic fence area. When the spatial relationship between the compensated trajectory data and the electronic fence is abnormal (such as non-project personnel entering the bidding area), the map will generate corresponding early warning information signals. The trajectory-scene joint optimization model, by fusing positioning signals and video behavior analysis results, can effectively distinguish between real risks and false alarms that need to be excluded by AI automatic false alarm prevention.
[0089] In step S105, the multi-scale trajectory reconstruction process pays special attention to the trajectory accuracy of key areas such as the overnight bidding area. The final personnel trajectory information output includes complete fields such as timestamps and location coordinates required for risk item archiving. The generated visualized trajectory report not only displays the basic movement trajectory but also integrates various early warning and handling process information, such as mobile phone call alerts and area unauthorized entry alerts, allowing platform administrators to view and handle these processes according to their permissions. Abnormal behavior markers in the report strictly adhere to the standard manual review process, and confirmed risk items are automatically linked to the corresponding project information and expert information databases.
[0090] This embodiment significantly improves the data quality of positioning device signals and monitoring video data in complex environments through the synergistic effect of the first preprocessing of dynamic Kalman filtering and the second preprocessing of adaptive illumination compensation. The trajectory prediction compensation based on LSTM neural network effectively solves the problem of trajectory loss during signal interruption. Combined with the spatiotemporal correlation map constructed by graph convolutional neural network, it realizes multimodal deep fusion of compensated trajectory data, personnel motion posture features and scene semantic features. The improved Transformer architecture multi-scale trajectory reconstruction technology ensures that the final personnel trajectory information simultaneously meets the positioning accuracy requirements of electronic fence trigger events and the temporal continuity requirements of abnormal behavior marking. The generated visualized trajectory report provides full-dimensional spatiotemporal data analysis support for regional control scenarios such as bidding area control and overnight bidding area control by uniformly integrating trajectory accuracy assessment, abnormal behavior marking and electronic fence trigger events.
[0091] Please see Figure 2 In some embodiments, the positioning device signal undergoes a first preprocessing step to obtain first trajectory data. This first preprocessing is configured as a signal denoising process based on dynamic Kalman filtering, including:
[0092] S201. Construct an adaptive Kalman filter model based on the motion state of the positioning device. The state vector of the adaptive Kalman filter model includes the device position coordinates, the device motion speed, and the device acceleration. The observation noise covariance matrix of the adaptive Kalman filter model is dynamically adjusted according to the current signal strength of the positioning device. The process noise covariance matrix of the adaptive Kalman filter model is used to represent the error of motion state prediction.
[0093] S202. Perform coarse filtering on the original positioning signal to remove abnormal points that exceed the physical motion constraints. The physical motion constraints include the maximum velocity threshold and the maximum acceleration threshold.
[0094] S203. A sliding window mechanism is used to calculate motion trend parameters. The window size of the sliding window mechanism is adaptively adjusted according to the positioning update frequency.
[0095] S204. Update the process noise covariance matrix based on motion trend parameters. When a turning feature is detected, increase the weight of the angle component in the process noise covariance matrix.
[0096] S205. Perform Kalman prediction-update iteration and output the smoothed first trajectory data. The first trajectory data includes the denoised three-dimensional coordinate sequence, the motion state confidence at each time step, the signal quality evaluation index, and the outlier marking information.
[0097] In step S201, an adaptive Kalman filter model based on the motion state of the positioning device is constructed. The state vector of the adaptive Kalman filter model includes the device position coordinates (x, y, z) and the device motion velocity (v). x ,v y ,v z ) and equipment acceleration (a x ,a y ,a z ). Observation noise covariance matrix R t The statistical characteristics of measurement noise are characterized, satisfying R0. t =σ 2 I, where,σ 2 = α / (SNR+β), where α is the first preset coefficient, β is the second preset coefficient, SNR is the real-time signal-to-noise ratio, and Q is the process noise covariance matrix. t The error used to describe the prediction of motion state is mainly reflected in the adaptive adjustment to sudden motion (such as sudden stop, turn).
[0098] In step S202, the original positioning signal is an unprocessed coordinate sequence directly output by the positioning device, and the maximum velocity threshold v in the physical motion constraints. max The maximum acceleration threshold α is determined based on the statistical distribution of normal walking speed. max The inertia of human motion is then considered and set as the physiological limit. The coarse filtering process, by establishing joint velocity-acceleration constraints, can effectively identify and eliminate coordinate jump points caused by multipath effects or equipment failures.
[0099] In step S203, the sliding window mechanism refers to a finite-length data buffer that slides along the time axis. Preferably, the window size T of the sliding window mechanism is adaptively adjusted according to the positioning update frequency to satisfy... Where f is the current sampling frequency. The calculated motion trend parameters include characteristic quantities reflecting the motion law, such as the average velocity vector and the rate of change of acceleration.
[0100] In step S204, the turning characteristic refers to a motion state where the rate of change of the motion direction angle exceeds a set threshold. Increasing the weight of the angle component in the process noise covariance matrix allows the filter to respond more quickly to changes in motion direction. The adjustment of the angle component weights satisfies ΔQ. θ = k·||ω||, where ω is the angular velocity and k is the adjustment coefficient. This adjustment strategy is particularly suitable for the walking pattern of personnel who frequently turn within the evaluation area, avoiding trajectory distortion caused by filtering lag.
[0101] In step S205, the confidence level of the motion state at each moment in the output first trajectory data reflects the reliability of the current estimate. Signal quality assessment index Q t =ω1SNR+ω2(1-σ 2 This comprehensively considers multiple factors such as observation noise intensity and prediction residuals; outlier marker information. The system records the locations of the original data points rejected by the coarse filter and the types of physical constraints they violated, where τ is the dynamic threshold. The entire processing effectively suppresses the influence of various measurement noises while preserving the true motion characteristics.
[0102] This embodiment significantly improves the denoising effect of the original positioning signal by constructing an adaptive Kalman filter model that includes the device's position coordinates, speed, and acceleration, and dynamically adjusting the observation noise covariance matrix based on signal strength. It effectively identifies and eliminates outliers by combining coarse filtering with physical motion constraints with motion trend parameter calculation using a sliding window mechanism. By dynamically updating the angular component weights in the process noise covariance matrix, the first trajectory data accurately reflects complex motion patterns such as turning characteristics while maintaining the continuity of the motion state. The final output first trajectory data includes multi-dimensional information such as motion state confidence and signal quality evaluation indicators, providing high-precision foundational data for subsequent trajectory compensation and reconstruction.
[0103] In some embodiments, the surveillance video data undergoes a second preprocessing step to obtain first video data. The second preprocessing step is configured as multi-target tracking enhancement processing based on adaptive illumination compensation, including:
[0104] An adaptive compensation model based on illumination perception is constructed. The adaptive compensation model includes an illumination component estimation module, a reflection component calculation module, and a dynamic compensation module. The dynamic compensation module is used to adjust the compensation intensity parameters according to changes in scene illumination.
[0105] Multi-target tracking processing is performed on the compensated video data to output the first video data. The first video data includes a video frame sequence with illumination compensation, multi-target tracking results with identification tags, target motion state data, and tracking quality evaluation results. The multi-target tracking processing includes:
[0106] The video frames are processed by a people detection network, and the detection results are output, including the coordinates of the people detection box and the confidence level.
[0107] A multi-feature fusion algorithm is used to perform target association on the detection results. Target association includes motion state matching and appearance feature association.
[0108] In this embodiment, the constructed adaptive compensation model based on illumination perception refers to a dynamic illumination adjustment system specifically designed for surveillance video scenes. The illumination component estimation module separates the global illumination components in the video frame, using guided filtering to extract the low-frequency illumination component L(X,y,t) from the original video frame. The reflection component calculation module extracts the inherent visual features of the target object, obtaining the reflection component through R(x,y,t) = V(x,y,t) / L(x,y,t), where V(x,y,t) is the original video frame. The dynamic compensation module adjusts the scene illumination according to the scene illumination change rate ΔL. t =||L t -L t-1 ||2 Adaptive adjustment of compensation strength parameters
[0109] The multi-target tracking enhancement process includes: building a person detection network based on the YOLOv5 architecture, outputting the coordinates of the detection boxes and their confidence scores, using the DeepSORT algorithm for target association, where motion state estimation uses an 8-dimensional state vector, appearance feature extraction uses a lightweight CNN network, and a tracking quality evaluation index is established.
[0110] In the first output video data, the video frame sequence that has completed illumination compensation is a set of continuous video images that have been processed by an adaptive compensation model to eliminate the effects of sudden changes in illumination; the multi-target tracking results with identification refer to cross-frame target association data marked with unique IDs; the target motion state data includes the target's position, velocity, and direction of motion information in the image coordinate system; and the tracking quality evaluation results reflect tracking reliability indicators such as the degree of target occlusion and feature consistency.
[0111] The personnel detection network processing refers to the process of analyzing video frames using deep learning object detection algorithms. The output detection results use pixel coordinates for the personnel detection bounding boxes, and the confidence score reflects the probability of a person being present within the bounding box. The multi-feature fusion algorithm is a dual discrimination mechanism that integrates motion state matching and appearance feature association. Motion state matching is based on the consistency of target displacement between adjacent frames, while appearance feature association is achieved through visual feature similarity extracted by a convolutional neural network. The entire processing workflow maintains stable multi-target tracking performance even under complex lighting conditions.
[0112] This embodiment effectively overcomes the problem of sudden changes in illumination in monitoring scenes by constructing an adaptive compensation model that includes an illumination component estimation module, a reflection component calculation module, and a dynamic compensation module. It combines personnel detection network processing with a multi-feature fusion algorithm to ensure the accuracy of multi-target tracking results with identification tags. The output first video data simultaneously includes target motion state data and tracking quality assessment results, providing complete and reliable multi-target tracking information for subsequent analysis. The entire processing flow significantly improves the video surveillance quality and tracking stability under complex lighting conditions.
[0113] In some embodiments, spatiotemporal continuity trajectory features are extracted from the first trajectory data, and trajectory prediction compensation is performed on the signal interruption segment based on an LSTM neural network to obtain compensated trajectory data, including:
[0114] A trajectory prediction model based on LSTM is constructed. The trajectory prediction model includes a trajectory feature extraction module and a trajectory prediction compensation module. The trajectory prediction compensation module is used to predict the trajectory of the interrupted segment based on motion features.
[0115] Trajectory prediction compensation is performed on the signal interruption segment, and the compensated trajectory data is output. The compensated trajectory data includes a complete three-dimensional coordinate sequence, position compensation markers, prediction confidence scores, and correction records. The trajectory prediction compensation includes:
[0116] Trajectory features are extracted from the first trajectory data, including position coordinate sequence, velocity change trend, and motion direction angle;
[0117] The trajectory features are input into the LSTM neural network, which outputs a sequence of predicted trajectory coordinates.
[0118] When the signal is restored, the predicted trajectory is verified and corrected against the actual positioning data.
[0119] In this embodiment, the constructed LSTM-based trajectory prediction model is a deep learning model specifically designed to handle the interruption of positioning signals. The trajectory feature extraction module is used to obtain key features such as position coordinate sequence, velocity change trend and motion direction angle from the first trajectory data. The trajectory prediction compensation module uses these features to predict the trajectory coordinate sequence during the interruption period through an LSTM neural network.
[0120] Compensated trajectory data refers to the complete trajectory information after prediction compensation processing. The complete three-dimensional coordinate sequence includes the spatial positions of the original positioning points and the predicted compensation points. The position compensation marker is used to distinguish between the actual measurement points and the predicted compensation points. The prediction confidence score reflects the reliability assessment of the prediction results by the LSTM neural network. The correction record saves the process information of verifying and correcting the predicted trajectory and the actual positioning data after signal recovery.
[0121] The specific implementation process of trajectory prediction compensation includes: first, extracting multi-dimensional motion features containing position, velocity, and direction from the first trajectory data; then, inputting these temporal features into an LSTM neural network for modeling and prediction; finally, during signal recovery, completing trajectory correction by comparing the predicted trajectory with the actual positioning data. This entire process effectively solves the problem of trajectory discontinuity caused by positioning signal interruption.
[0122] This embodiment constructs an LSTM-based trajectory prediction model that includes a trajectory feature extraction module and a trajectory prediction compensation module, which can accurately predict the trajectory coordinate sequence of signal interruption segments. The output compensated trajectory data not only contains a complete three-dimensional coordinate sequence, but also provides position compensation markers, prediction confidence scores, and correction records, ensuring the integrity and reliability of the trajectory data. By inputting trajectory features into the LSTM neural network and implementing verification and correction after signal recovery, the problem of trajectory discontinuity caused by positioning signal interruption is effectively solved, and the spatiotemporal continuity of trajectory data is significantly improved.
[0123] In some embodiments, extracting human motion posture features and scene semantic features from the first video data includes:
[0124] A deep learning-based 3D pose estimation model is constructed. The 3D pose estimation model includes a 2D keypoint detection network and a 3D pose reconstruction network. The 2D keypoint detection network is used to detect the 2D coordinates of each joint of the human body from video frames. The 3D pose reconstruction network is used to map the 2D joint coordinates to 3D space and calculate the relative motion relationship between the joints.
[0125] The first video data is processed by a three-dimensional pose estimation model. Human body detection and joint point localization are performed on the input video frames of the first video data, and the two-dimensional joint point coordinates including the main body parts such as the head and limbs are output.
[0126] Based on multi-view geometric constraints and prior knowledge of human skeleton, the coordinates of two-dimensional joint points are reconstructed into three-dimensional spatial coordinates.
[0127] Calculate the motion changes of joints between adjacent video frames and output the human motion posture features, which include a three-dimensional spatial coordinate sequence, joint angle changes, and overall motion velocity vector.
[0128] A scene semantic understanding model is constructed, which includes a scene classification network and a semantic segmentation network. The scene classification network is used to identify the scene category of the monitored area, and the semantic segmentation network is used to label the pixel-level boundaries of each functional area in the scene.
[0129] The first video data is processed by a scene semantic understanding model. The input video frames of the first video data are classified into scene partitions, which include corridors, rooms and stairwells.
[0130] Perform pixel-level segmentation on the rooms in the scene partition, and output a segmentation mask containing entrances / exits, rest areas, and work areas to obtain the segmentation result;
[0131] A scene topology graph is constructed based on the segmentation results, and scene semantic features are generated.
[0132] In this embodiment, the deep learning-based 3D pose estimation model is a neural network model used to extract 3D motion information of the human body from video data. The 2D keypoint detection network uses a convolutional neural network architecture to detect the 2D coordinates of each joint of the human body, while the 3D pose reconstruction network uses a multilayer perceptron structure to map the 2D coordinates to 3D space and establish a motion relationship model between the joints. This model learns the mapping relationship from 2D images to 3D poses through end-to-end training.
[0133] Human motion posture features are quantitative descriptions of human motion obtained through processing with a 3D pose estimation model. The 3D spatial coordinate sequence represents the spatial position changes of each joint point in consecutive video frames, the joint angle changes reflect the relative amplitude of limb movement, and the overall motion velocity vector describes the human's movement speed and direction within the scene. These features are obtained by calculating the inter-frame joint displacements of adjacent video frames.
[0134] The scene semantic understanding model is a computer vision model used to analyze the spatial structure of a surveillance scene. The scene classification network identifies the overall scene category of the monitored area using image classification algorithms, while the semantic segmentation network preferably employs an encoder-decoder architecture to achieve pixel-level segmentation of each functional area. The classification results for areas such as corridors, rooms, and stairwells within the scene partition are obtained through the fully connected layer output of the scene classification network. Furthermore, a road network parsing layer is added to the scene semantic understanding model, incorporating key path elements such as corridors and access control points as special nodes into the spatiotemporal relational graph.
[0135] Scene semantic features are an abstract representation of the spatial structure of a monitoring scene, preferably modeled through a scene topology map to depict the spatial relationships between functional areas. This feature first identifies functional areas such as entrances, rest areas, and work areas based on the segmentation mask output by the semantic segmentation network. Then, it constructs a scene topology map based on the adjacency relationships and connectivity of these areas, ultimately forming a feature representation that includes scene structure and functional zoning information.
[0136] This embodiment collaboratively processes video data by constructing a 3D pose estimation model and a scene semantic understanding model. The 3D pose estimation model first detects the 2D joints of the human body, then reconstructs the 3D motion trajectory through multi-view geometric constraints, extracting motion pose features including joint angles and velocities. Simultaneously, the scene semantic understanding model performs scene classification and semantic segmentation on video frames, identifies functional regions, and constructs a scene topology map. Finally, the human motion pose features are combined with scene semantic features to achieve correlation analysis between human behavior and the scene environment, providing multi-dimensional feature support for subsequent behavior understanding.
[0137] In some embodiments, a graph convolutional neural network is used to construct a spatiotemporal correlation graph from the compensated trajectory data, personnel motion posture features, and scene semantic features to obtain enhanced trajectory data, including:
[0138] A spatiotemporal correlation model based on graph neural networks is constructed. The spatiotemporal correlation model includes a node feature encoder and a graph convolution operation module. The node feature encoder is used to uniformly encode the compensated trajectory data, personnel motion posture features and scene semantic features into node feature vectors. The graph convolution operation module is used to model the spatiotemporal correlation relationship between nodes.
[0139] Using trajectory points, human body joints, and scene partitions as graph nodes, an initial graph structure containing spatially adjacent edges and temporally continuous edges is established.
[0140] Based on the scene topology, establish personnel-scene interaction edges, which include spatial inclusion relationships and movement direction relationships, and update personnel-scene interaction edges to spatial adjacency edges;
[0141] Additionally, establish trajectory continuity edges along the time dimension, which include velocity consistency and motion smoothness constraints, and update the trajectory continuity edges to the time continuity edges;
[0142] Generate the final graph structure;
[0143] The features of neighboring nodes in the final graph structure are aggregated by multi-layer graph convolution operation. The first weight coefficient of spatially adjacent edges and the second weight coefficient of temporally continuous edges are dynamically adjusted based on the attention mechanism, and the output is enhanced trajectory data containing spatial correction and temporal smoothing.
[0144] Enhanced trajectory data includes refined trajectory coordinates that integrate multi-source features, annotations of personnel-scene interaction states, trajectory reliability scores, and abnormal behavior detection markers.
[0145] In this embodiment, the spatiotemporal correlation model is a graph neural network architecture used to fuse multi-source motion data. The node feature encoder uses fully connected layers to uniformly map the 3D coordinate sequence of the compensated trajectory data, the joint motion parameters of the personnel's motion posture features, and the partitioned topological information of the scene semantic features to the feature space. The graph convolution operation module achieves cross-modal feature interaction and fusion through a message passing mechanism. This model learns the spatiotemporal correlation patterns between different features through end-to-end training.
[0146] The initial graph structure is a graph data structure used to represent the relationships between multi-source data. Trajectory point nodes represent the spatial location of the compensated trajectory data, human joint node nodes reflect the three-dimensional coordinates of the person's movement posture characteristics, and scene partition nodes correspond to functional areas of the scene's semantic features. Spatial adjacency edges determine the connection relationship between adjacent nodes using a Euclidean distance threshold, while temporally continuous edges establish node associations based on the timestamp order of video frames.
[0147] Person-scene interaction edges are a special type of edge used to model the spatial relationship between the human body and the scene. The spatial containment relationship is determined by whether the coordinates of the joint point are located inside the scene partition, and the motion direction relationship is calculated based on the angle between the joint point's motion velocity vector and the scene's entrance / exit normal vector. Trajectory continuity edges are constraint edges used to maintain the consistency of motion sequence. Their velocity consistency is calculated by the velocity difference between adjacent trajectory points, and motion smoothness is evaluated by the change in acceleration.
[0148] The final graph structure achieves the propagation and updating of node features through iterative graph convolution operations. The first weight coefficient is dynamically adjusted based on the feature similarity of spatially adjacent nodes, while the second weight coefficient is adaptively allocated based on the motion consistency of temporally continuous nodes.
[0149] The enhanced trajectory data includes refined trajectory coordinates that integrate multi-source features, personnel-scene interaction status annotations, trajectory reliability scores, and abnormal behavior detection markers. Preferably, the refined trajectory coordinates are obtained by aggregating the positional features of spatially adjacent nodes, the personnel-scene interaction status annotations reflect the degree of matching between the motion trajectory and the scene function, the trajectory reliability score is calculated based on the attention weight coefficient, and the abnormal behavior detection markers are determined by the abnormal score output by graph convolution.
[0150] This embodiment fuses multi-source motion data by constructing a spatiotemporal correlation graph. Using trajectory points, keypoints, and scene partitions as graph nodes, a graph structure is established, incorporating three edge types: spatial adjacency, temporal continuity, and personnel-scene interaction. Neighborhood features are aggregated using a graph convolutional neural network. An attention mechanism is used to dynamically adjust the weight coefficients of different edges, enhancing spatial accuracy while maintaining motion smoothness. The final output is enhanced trajectory data that integrates scene semantics and motion posture, including refined coordinates, interaction states, and anomaly markers. This embodiment constructs a spatiotemporal correlation graph containing compensated trajectory data, personnel motion posture features, and scene semantic features. By dynamically fusing multi-source information using a graph convolutional neural network, the generated enhanced trajectory data possesses both spatial correction and temporal smoothing characteristics, significantly improving the refinement of trajectory coordinates. It also includes personnel-scene interaction state annotations, trajectory reliability scores, and abnormal behavior detection markers, providing a more comprehensive and reliable trajectory representation for behavior analysis.
[0151] In some embodiments, multi-scale trajectory reconstruction is performed on the enhanced trajectory data based on an improved Transformer architecture to obtain the final personnel trajectory information, including:
[0152] A multi-scale trajectory reconstruction model based on an improved Transformer is constructed. The multi-scale trajectory reconstruction model includes a spatiotemporal feature encoding module and a trajectory generation module. The spatiotemporal feature encoding module is used to extract trajectory features at different time and spatial scales, and the trajectory generation module is used to output the optimized trajectory sequence.
[0153] In the time dimension, long-time windows, medium-time windows, and short-time windows are used to extract motion features respectively;
[0154] In the spatial dimension, location features are extracted by establishing local motion neighborhoods and global scene ranges respectively;
[0155] The correlation weights between motion features and position features are calculated using a spatiotemporal attention mechanism to generate correlation weight information.
[0156] Based on the association weight information, a feature fusion network is used to perform feature fusion, and the final personnel trajectory information is output.
[0157] In this embodiment, the multi-scale trajectory reconstruction model includes a spatiotemporal feature encoding module and a trajectory generation module. The spatiotemporal feature encoding module extracts spatiotemporal patterns from the enhanced trajectory data through a multi-head self-attention mechanism, while the trajectory generation module outputs an optimized continuous trajectory sequence through a decoder structure. This architecture improves the accuracy of trajectory reconstruction by processing motion features at different scales in parallel.
[0158] Long-term, medium-term, and short-term windows are different time spans used to capture motion features. Long-term windows focus on overall motion trends, medium-term windows analyze motion phase characteristics, and short-term windows capture subtle changes in motion. In the spatial dimension, the local motion neighborhood represents fine motion patterns through the spatial distribution of neighboring trajectory points, while the global scene scope considers the constraints of the entire scene space on the motion trajectory.
[0159] Preferably, the spatiotemporal attention mechanism dynamically generates associated weight information through a query-key-value matching mechanism; the feature fusion network weights and combines spatiotemporal features according to the associated weight information, ultimately outputting final personnel trajectory information that includes global consistency and local details. Furthermore, a road network attention mechanism can be set up, using the CAD road network topology as a key parameter in the calculation of spatiotemporal attention weights, ensuring that the reconstructed trajectory strictly follows the physical constraints of building passageways, improving trajectory accuracy at key nodes such as access control points and turns.
[0160] This embodiment improves the multi-scale trajectory reconstruction model of the Transformer architecture by extracting temporal features using long, medium, and short time windows, and combining spatial feature analysis of local motion neighborhoods and global scene ranges. The associated weight information generated by the spatiotemporal attention mechanism guides the feature fusion network to perform accurate weighting. The final output optimized trajectory sequence retains both the integrity of the motion trend and the accuracy of action details, significantly improving the completeness and reliability of personnel trajectory reconstruction.
[0161] In some embodiments, a visual trajectory report is generated based on the final personnel trajectory information, including trajectory accuracy assessment, abnormal behavior markers, and electronic fence trigger events, comprising:
[0162] A visualization report generation model is constructed, which includes a trajectory analysis unit and an event marking unit. The trajectory analysis unit is used to calculate trajectory accuracy indicators, and the event marking unit is used to identify abnormal behavior and electronic fence events.
[0163] The trajectory accuracy assessment is calculated by the trajectory analysis unit, and the trajectory accuracy assessment result is obtained. The trajectory accuracy assessment result includes the positioning error index and the trajectory smoothness index.
[0164] Generate accuracy visualization markers based on trajectory accuracy assessment results;
[0165] Additionally, abnormal behavior events and electronic fence triggering events are detected through event tagging units to obtain detection results;
[0166] Generate behavioral anomaly markers and regional violation markers based on the detection results;
[0167] A visual trajectory report is generated based on accuracy visualization markers, anomaly markers, and regional violation markers.
[0168] In this embodiment, the visualization report generation model is a system architecture that transforms the final personnel trajectory information into a structured analysis report. The trajectory analysis unit quantifies trajectory quality by calculating positioning error and trajectory smoothness indices. The positioning error index reflects the degree of deviation between trajectory points and their actual locations, calculated by comparing the coordinate differences between the reconstructed trajectory and the reference trajectory. The trajectory smoothness index characterizes the continuity of the trajectory, determined by analyzing the rate of change of motion vectors between adjacent trajectory points. This model is particularly suitable for real-time personnel positioning scenarios in bidding evaluation areas, accurately displaying the movement trajectories of personnel wearing positioning devices.
[0169] The event tagging unit is used to identify special trajectory events. Abnormal behavior events are detected by analyzing features such as sudden changes in movement speed and deviations from the normal path pattern, specifically including violations such as making or receiving phone calls and entering unauthorized areas. Electronic fence trigger events are identified by judging the spatial relationship between trajectory coordinates and the boundaries of preset restricted areas, covering various early warning scenarios such as bidding area control, overnight bidding area control, and sample room control. The detection results are converted into behavior anomaly tags and area violation tags. Behavior anomaly tags use specific icons to mark abnormal trajectory segments, while area violation tags are visualized by highlighting the illegally entered electronic fence area.
[0170] Precision visualization markers transform numerical trajectory accuracy assessment results into graphical elements, using color gradients or symbol sizes to reflect quality differences across different trajectory segments. The resulting visualized trajectory report integrates precision visualization markers, behavioral anomaly markers, and area violation markers, forming a visual output containing multi-dimensional analysis results. This report supports viewing video footage along the trajectory path and can be linked to early warning and response processes, including three stages: early warning information generation, handling of suspected risk information, and risk item archiving. It also features AI-powered automatic false alarm removal, allowing managers to intuitively grasp the quality status of personnel trajectories and the distribution of abnormal events.
[0171] This embodiment achieves accurate assessment of personnel trajectory quality and reliable identification of abnormal behavior by generating trajectory accuracy evaluation results and event tagging unit detection results through a visualization report generation model. The comprehensive display of accuracy visualization tags, abnormal behavior tags, and area violation tags enables managers to intuitively grasp the trajectory quality status and the distribution of violation events. Combined with the automatic detection of electronic fence triggered events and abnormal behavior events, it effectively improves the accuracy and timeliness of bidding area control and provides visualized decision support for the early warning and handling process.
[0172] Unlike existing technologies, the above technical solution has the following advantages: The synergistic effect of the first preprocessing using dynamic Kalman filtering and the second preprocessing using adaptive illumination compensation significantly improves the data quality of positioning device signals and surveillance video data in complex environments; trajectory prediction compensation based on LSTM neural networks effectively solves the problem of missing trajectories during signal interruptions, and the spatiotemporal correlation map constructed by graph convolutional neural networks achieves multimodal deep fusion of compensated trajectory data, personnel motion posture features, and scene semantic features; the improved Transformer architecture's multi-scale trajectory reconstruction technology ensures that the final personnel trajectory information simultaneously meets the positioning accuracy requirements of electronic fence trigger events and the temporal continuity requirements of abnormal behavior marking; the generated visualized trajectory report provides full-dimensional spatiotemporal data analysis support for regional control scenarios such as bidding area control and overnight bidding area control by uniformly integrating trajectory accuracy assessment, abnormal behavior marking, and electronic fence trigger events.
[0173] In a second aspect, the present invention also provides a personnel trajectory mapping system based on an AI algorithm, which is applicable to the method described in the first aspect.
[0174] The above technical solution constructs a spatiotemporal correlation map that includes compensated trajectory data, personnel motion posture features, and scene semantic features. It uses graph convolutional neural networks to dynamically fuse multi-source information, generating enhanced trajectory data that simultaneously possesses spatial correction and temporal smoothing characteristics. This significantly improves the refinement of trajectory coordinates and includes personnel-scene interaction state annotations, trajectory reliability scores, and abnormal behavior detection markers, providing a more comprehensive and reliable trajectory representation for behavior analysis.
[0175] Finally, it should be noted that although the above embodiments have been described in the text and drawings of this application, this should not limit the scope of patent protection of this application. Any technical solutions that are based on the essential concept of this application and utilize the content described in the text and drawings of this application, resulting in equivalent structural or procedural substitutions or modifications, as well as the direct or indirect application of the technical solutions of the above embodiments to other related technical fields, are all included within the scope of patent protection of this application.
Claims
1. A method for drawing personnel trajectories based on AI algorithms, characterized in that, include: The system collects positioning device signals and monitoring video data. The positioning device signals include the real-time location information of the current user, and the monitoring video data includes multi-view spatiotemporal continuous images. The positioning device signal is subjected to a first preprocessing to obtain first trajectory data, wherein the first preprocessing is configured as signal denoising processing based on dynamic Kalman filtering; In addition, the surveillance video data is subjected to a second preprocessing to obtain the first video data, wherein the second preprocessing is configured as multi-target tracking enhancement processing based on adaptive illumination compensation; The spatiotemporal continuity trajectory features are extracted from the first trajectory data, and trajectory prediction compensation is performed on the signal interruption segment based on the LSTM neural network to obtain the compensated trajectory data. In addition, personnel motion posture features and scene semantic features are extracted from the first video data; A spatiotemporal correlation graph is constructed by using a graph convolutional neural network to combine the compensated trajectory data, personnel motion posture features and scene semantic features to obtain enhanced trajectory data. The spatiotemporal correlation graph is configured as a trajectory-scene joint optimization model. The enhanced trajectory data is reconstructed using an improved Transformer architecture to obtain the final personnel trajectory information. The multi-scale trajectory reconstruction is configured as an end-to-end trajectory generation that integrates positioning signals, visual observations, and environmental topology. Furthermore, based on the final personnel trajectory information, a visual trajectory report is generated that includes trajectory accuracy assessment, abnormal behavior markers, and electronic fence trigger events. The visual trajectory report is configured to support spatiotemporal data analysis for real-time monitoring and historical backtracking. Specifically, multi-scale trajectory reconstruction is performed on the enhanced trajectory data based on the improved Transformer architecture to obtain the final personnel trajectory information, including: A multi-scale trajectory reconstruction model based on an improved Transformer is constructed. The multi-scale trajectory reconstruction model includes a spatiotemporal feature encoding module and a trajectory generation module. The spatiotemporal feature encoding module is used to extract trajectory features at different time and spatial scales, and the trajectory generation module is used to output the optimized trajectory sequence. In the time dimension, long-time windows, medium-time windows, and short-time windows are used to extract motion features respectively; In the spatial dimension, location features are extracted by establishing local motion neighborhoods and global scene ranges respectively; The correlation weights between motion features and location features are calculated through a spatiotemporal attention mechanism to generate correlation weight information. In this process, a road network attention mechanism is set up, and the CAD road network topology is used as a key parameter to participate in the spatiotemporal attention weight calculation. This ensures that the reconstructed trajectory strictly follows the physical constraints of the building passage, thereby improving the trajectory accuracy at key nodes such as access points and turns. Based on the associated weight information, a feature fusion network is used to perform feature fusion, and the final personnel trajectory information is output.
2. The method for drawing personnel trajectories based on AI algorithms according to claim 1, characterized in that, The positioning device signal undergoes a first preprocessing step to obtain first trajectory data. This first preprocessing step is configured as a signal denoising process based on dynamic Kalman filtering, including: An adaptive Kalman filter model based on the motion state of a positioning device is constructed. The state vector of the adaptive Kalman filter model includes the device's position coordinates, device's motion speed, and device's acceleration. The observation noise covariance matrix of the adaptive Kalman filter model is dynamically adjusted according to the current signal strength of the positioning device. The process noise covariance matrix of the adaptive Kalman filter model is used to represent the error of motion state prediction. The original positioning signal is coarsely filtered to remove outliers that exceed the physical motion constraints, which include a maximum velocity threshold and a maximum acceleration threshold. Motion trend parameters are calculated using a sliding window mechanism, and the window size of the sliding window mechanism is adaptively adjusted according to the positioning update frequency. The process noise covariance matrix is updated based on the motion trend parameters, and the weight of the angle component in the process noise covariance matrix is increased when a turning feature is detected. Perform Kalman prediction-update iteration to output smoothed first trajectory data, which includes a denoised three-dimensional coordinate sequence, motion state confidence at each time step, signal quality evaluation index, and outlier marker information.
3. The method for drawing personnel trajectories based on AI algorithms according to claim 1, characterized in that, The surveillance video data undergoes a second preprocessing step to obtain first video data. This second preprocessing is configured as a multi-target tracking enhancement process based on adaptive illumination compensation, including: An adaptive compensation model based on illumination perception is constructed. The adaptive compensation model includes an illumination component estimation module, a reflection component calculation module, and a dynamic compensation module. The dynamic compensation module is used to adjust the compensation intensity parameters according to changes in scene illumination. The compensated video data is subjected to multi-target tracking processing to output first video data. The first video data includes a video frame sequence with illumination compensation, multi-target tracking results with identification tags, target motion state data, and tracking quality evaluation results. The multi-target tracking processing includes: The video frames are processed by a people detection network, and the detection results are output, including the coordinates of the people detection box and the confidence level. A multi-feature fusion algorithm is used to perform target association on the detection results. The target association includes motion state matching and appearance feature association.
4. The method for drawing personnel trajectories based on AI algorithms according to claim 1, characterized in that, The spatiotemporal continuity trajectory features are extracted from the first trajectory data, and trajectory prediction compensation is performed on the signal interruption segment based on the LSTM neural network to obtain the compensated trajectory data, including: A trajectory prediction model based on LSTM is constructed. The trajectory prediction model includes a trajectory feature extraction module and a trajectory prediction compensation module. The trajectory prediction compensation module is used to predict the trajectory of the interrupted segment based on motion features. Trajectory prediction compensation is performed on the signal interruption segment, and compensated trajectory data is output. The compensated trajectory data includes a complete three-dimensional coordinate sequence, position compensation markers, prediction confidence scores, and correction records. The trajectory prediction compensation includes: Trajectory features are extracted from the first trajectory data, including position coordinate sequence, velocity change trend, and motion direction angle; The trajectory features are input into an LSTM neural network, which outputs a sequence of predicted trajectory coordinates. When the signal is restored, the predicted trajectory is verified and corrected against the actual positioning data.
5. The method for drawing personnel trajectories based on AI algorithms according to claim 1, characterized in that, Extracting human motion posture features and scene semantic features from the first video data, including: A deep learning-based 3D pose estimation model is constructed. The 3D pose estimation model includes a 2D keypoint detection network and a 3D pose reconstruction network. The 2D keypoint detection network is used to detect the 2D coordinates of each joint of the human body from video frames. The 3D pose reconstruction network is used to map the 2D joint coordinates to 3D space and calculate the relative motion relationship between the joints. The first video data is processed by the three-dimensional pose estimation model. Human body detection and joint point localization are performed on the input video frames of the first video data, and the two-dimensional joint point coordinates including the head, limbs and other main body parts are output. Based on multi-view geometric constraints and prior knowledge of human skeleton, the coordinates of two-dimensional joint points are reconstructed into three-dimensional spatial coordinates. Calculate the motion changes of joints between adjacent video frames and output the human motion posture features, which include a three-dimensional spatial coordinate sequence, joint angle changes, and overall motion velocity vector. A scene semantic understanding model is constructed, which includes a scene classification network and a semantic segmentation network. The scene classification network is used to identify the scene category of the monitored area, and the semantic segmentation network is used to annotate the pixel-level boundaries of each functional area in the scene. The first video data is processed by the scene semantic understanding model. The input video frames of the first video data are classified into scenes to obtain scene partitions, including corridors, rooms, and stairwells. Perform pixel-level segmentation on the rooms in the scene partition, and output a segmentation mask containing entrances / exits, rest areas, and work areas to obtain the segmentation result; Based on the segmentation results, a scene topology graph is constructed, and the scene semantic features are generated.
6. The method for drawing personnel trajectories based on AI algorithms according to claim 1, characterized in that, A spatiotemporal correlation graph is constructed using a graph convolutional neural network to combine the compensated trajectory data, personnel motion posture features, and scene semantic features, resulting in enhanced trajectory data, including: A spatiotemporal correlation model based on graph neural networks is constructed. The spatiotemporal correlation model includes a node feature encoder and a graph convolution operation module. The node feature encoder is used to uniformly encode the compensated trajectory data, personnel motion posture features and scene semantic features into node feature vectors. The graph convolution operation module is used to model the spatiotemporal correlation relationship between nodes. Using trajectory points, human body joints, and scene partitions as graph nodes, an initial graph structure containing spatially adjacent edges and temporally continuous edges is established. Based on the scene topology, establish personnel-scene interaction edges, which include spatial inclusion relationships and movement direction relationships, and update personnel-scene interaction edges to spatial adjacency edges; In addition, a trajectory continuity edge is established along the time dimension, the trajectory continuity edge including velocity consistency and motion smoothness constraints, and the trajectory continuity edge is updated to the time continuity edge; Generate the final graph structure; The features of neighboring nodes in the final graph structure are aggregated by multi-layer graph convolution operation. The first weight coefficient of spatially adjacent edges and the second weight coefficient of temporally continuous edges are dynamically adjusted based on the attention mechanism, and the output is enhanced trajectory data containing spatial correction and temporal smoothing. The enhanced trajectory data includes refined trajectory coordinates that fuse multi-source features, personnel-scene interaction status annotations, trajectory reliability scores, and abnormal behavior detection markers.
7. The method for drawing personnel trajectories based on AI algorithms according to claim 1, characterized in that, Based on the final personnel trajectory information, a visual trajectory report is generated that includes trajectory accuracy assessment, abnormal behavior markers, and electronic fence trigger events, including: A visualization report generation model is constructed, which includes a trajectory analysis unit and an event marking unit. The trajectory analysis unit is used to calculate trajectory accuracy indicators, and the event marking unit is used to identify abnormal behavior and electronic fence events. The trajectory accuracy assessment is calculated by the trajectory analysis unit to obtain the trajectory accuracy assessment result, which includes a positioning error index and a trajectory smoothness index. Generate accuracy visualization markers based on trajectory accuracy assessment results; Furthermore, abnormal behavior events and electronic fence triggering events are detected through the event marking unit to obtain detection results; Based on the detection results, generate behavioral anomaly markers and regional violation markers; A visual trajectory report is generated based on the accuracy visualization markers, anomaly markers, and regional violation markers.
8. A personnel trajectory mapping system based on AI algorithms, characterized in that, Applicable to the method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Pedestrian trajectory prediction method based on window attention and space diagram interaction network
CN120318273A
Techniques to perform trajectory predictions
US20230394823A1