Intelligent human shape trajectory prediction and alarm system and method based on multi-modal video analysis
Through the intelligent humanoid trajectory prediction and alarm system based on multimodal video analysis, the problem of insufficient accurate extraction of target motion state information and trajectory prediction accuracy in the prior art is solved, and high-precision target behavior monitoring and dynamic environment adaptation are achieved, which significantly improves the system's response ability and prediction accuracy.
Patent Information
- Application Number
- CN202510140790.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-05-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art cannot accurately extract target motion state information, insufficient trajectory prediction accuracy and poor adaptability to dynamic environments, making it difficult to meet the needs of high-precision and multi-dimensional behavior monitoring.
The intelligent humanoid trajectory prediction and alarm system based on multimodal video analysis is adopted. Multimodal video data is collected through the video acquisition module, the data processing module extracts the motion state information of the target humanoid, the trajectory modeling module builds a trajectory model and generates a predicted trajectory through a dynamic optimization algorithm, the risk assessment module calculates the behavioral risk value, and triggers the alarm and optimization module for feedback optimization through the alarm module.
The refined modeling and dynamic prediction of the target motion state are achieved, the accuracy of trajectory prediction and the system's response ability to target behavior are improved, the adaptability to the dynamic environment is enhanced, and the incidence of missed and false alarms is reduced.
Smart Images

Figure CN120047897A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of human motion monitoring and recognition, and in particular to an intelligent human trajectory prediction and alarm system and method based on multimodal video analysis. Background Art
[0002] With the increasing demand for social security, intelligent humanoid recognition and alarm systems based on surveillance cameras have gained widespread application in security, public space management, industrial park monitoring, and other fields. These systems collect video data from surveillance scenes to detect, identify, and analyze the behavior of target individuals, providing automatic early warning of abnormal behavior. Their core approach lies in accurately extracting the target's motion characteristics and dynamically predicting their trajectory. This allows for timely alarm triggering when abnormal behavior occurs, reducing oversight by human monitoring and improving security efficiency.
[0003] In the existing technology, target recognition and alarm systems mostly rely on the simple capture and judgment of the target position in the video frame. The system usually determines whether the target has crossed the boundary or entered a specific risk area by setting fixed area or static threshold rules. However, this type of method lacks in-depth analysis of the target's motion state (such as speed, direction, etc.) and cannot cope with the complex changes in the target's behavior. For example, when the target person changes speed, makes quick turns or other abnormal actions in a dynamic scene, the traditional method cannot accurately capture these behavioral characteristics due to the single data dimension, resulting in a large number of false alarms or missed alarms.
[0004] Furthermore, existing trajectory modeling techniques typically rely on static rules or simple linear prediction algorithms. These methods often only consider the target's historical trajectory, making rough inferences about its future trajectory and lacking support for dynamic optimization. Especially in situations where environmental constraints (such as obstacles and risk areas) are complex, traditional techniques are unable to fully model the correlation between the environment and the target, resulting in large deviations in predicted trajectories and reduced system reliability and adaptability. Therefore, existing technologies are particularly inadequate in multi-target scenarios, making it difficult to meet the actual needs of high-precision, multi-dimensional behavior monitoring. Summary of the Invention
[0005] In response to the shortcomings of the existing technology, the present invention provides an intelligent humanoid trajectory prediction and alarm system and method based on multimodal video analysis, which solves the problems in the existing technology such as the inability to accurately extract target motion state information, insufficient trajectory prediction accuracy, and poor adaptability to dynamic environments.
[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: an intelligent humanoid trajectory prediction and alarm system based on multimodal video analysis, comprising: Video acquisition module, used to collect video stream data of monitoring scenes; A data processing module extracts motion state information of the target humanoid based on the video stream data, wherein the motion state information includes the joint position, speed and direction of the target humanoid; The trajectory modeling module builds the target person's trajectory model based on the target person's motion state and generates the target person's predicted trajectory through a dynamic optimization algorithm; The risk assessment module calculates the behavioral risk value based on the target person's predicted trajectory and scene environment characteristics; The alarm module compares the behavior risk value with the alarm threshold and triggers the alarm based on the comparison result; The optimization module performs feedback optimization on the trajectory model based on the deviation between the actual trajectory and the predicted trajectory after the alarm is triggered.
[0007] Preferably, the video acquisition module includes an infrared camera and a visible light camera, and solves the problem of capturing human targets under low light or complex lighting conditions through multimodal data fusion.
[0008] Preferably, the data processing module extracts key point data of the target human figure through a posture estimation algorithm, and the key points include joints of the head, shoulders, elbows, knees and feet.
[0009] Preferably, the trajectory modeling module implements trajectory modeling through an optimal control algorithm, and the optimal control algorithm predicts the future trajectory of the target person based on the historical trajectory of the target person in combination with scene environment constraints.
[0010] Preferably, the optimal control algorithm for trajectory prediction in the trajectory modeling module is solved based on a Lagrangian function, wherein the Lagrangian function includes the kinetic energy of the target person and the potential energy in the scene environment, wherein the potential energy is described by risk area characteristics and behavioral resistance characteristics.
[0011] Preferably, the risk assessment module fits the historical behavior distribution of the target person based on a Gaussian mixture model, and calculates the behavior risk value according to the matching degree between the predicted trajectory and the behavior distribution.
[0012] Preferably, the risk assessment module compares the calculated behavior risk value with a preset threshold value and triggers different levels of alarms through a graded alarm mechanism, wherein the alarm levels include high-risk alarm, medium-risk alarm and low-risk record.
[0013] Preferably, the optimization module optimizes and adjusts the potential energy function parameters of the trajectory modeling module based on a feedback learning mechanism.
[0014] Preferably, the optimization module adopts a federated learning method to optimize the local trajectory model through each surveillance camera node, and upload the optimization results to the cloud. The global model aggregated in the cloud is used to update the trajectory modeling module of each node.
[0015] An intelligent humanoid trajectory prediction and alarm method based on multimodal video analysis includes the following steps: Collect video stream data of the monitoring scene and obtain the movement information of people in the target area; Analyze the video stream data to extract the target person's motion state information, including the target person's joint position, speed, and direction; A trajectory model is constructed based on the target person's motion state, and a predicted trajectory of the target person is generated through a dynamic optimization algorithm; a behavioral risk value is calculated based on the predicted trajectory and scene environmental characteristics, wherein the scene environmental characteristics include risk area characteristics and behavioral resistance characteristics; Compare the behavior risk value with the preset alarm threshold. If the behavior risk value exceeds the alarm threshold, an alarm is triggered through the alarm module; The trajectory model is optimized based on the deviation between the actual trajectory and the predicted trajectory after the alarm is triggered, so as to improve the accuracy of trajectory prediction and the overall adaptability of the system.
[0016] The present invention provides an intelligent humanoid trajectory prediction and alarm system and method based on multimodal video analysis. It has the following beneficial effects: 1. The present invention uses a data processing module to extract the motion state information of the target humanoid, providing high-precision input data for the subsequent trajectory modeling module. The trajectory modeling module constructs a trajectory model and, combined with a dynamic optimization algorithm, generates a predicted trajectory of the target person. Compared with existing technical solutions that rely solely on coarse-grained target detection or simple position tracking, the present invention achieves refined modeling and dynamic prediction of the target's motion state, solving the problem of existing methods' inability to accurately reflect complex behavioral changes. At the same time, it significantly improves the accuracy of the predicted trajectory and the system's responsiveness to the target's behavior.
[0017] 2. This invention utilizes multimodal data fusion technology, combining infrared and visible light cameras within a video acquisition module, to overcome the shortcomings of single-modality devices in low-light or complex lighting conditions. Specifically, the infrared camera provides high-contrast target outline data, while the visible light camera supplements detail and color information. A fusion algorithm enables complementary optimization of this data. Compared to existing approaches that rely solely on visible light cameras for target capture, this invention significantly improves target capture accuracy and environmental adaptability, demonstrating significant advantages in low-light and dynamic background scenarios.
[0018] 3. This invention utilizes trajectory modeling and risk assessment based on an optimal control algorithm, combining the target individual's historical trajectory with real-time status information to dynamically generate predicted trajectories and calculate behavioral risk values. By incorporating the Lagrangian function modeling system's changing kinetic and potential energy characteristics, this approach addresses the inadequate ability of traditional trajectory modeling schemes to predict complex behaviors. Compared to existing approaches that rely on static threshold judgments, this invention accurately assesses potential abnormal behaviors of target individuals, effectively reducing the incidence of missed alerts and false positives.
[0019] 4. The present invention utilizes the deviation between the actual trajectory after the alarm and the predicted trajectory through an optimization module to dynamically adjust the parameters in the trajectory model, and combines it with federated learning to achieve distributed optimization among multiple nodes. Through the dynamic learning rate mechanism, the system can accelerate the convergence of model parameters and improve optimization efficiency. Compared with the static parameter model in the prior art, which is difficult to adapt to scene changes, the present invention significantly improves the adaptability and prediction accuracy of the system through a feedback learning mechanism, enabling it to continuously adapt to complex and changing monitoring environments, thereby ensuring the long-term stability of the overall performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 Schematic diagram of the system architecture of the present invention; Figure 2 Schematic diagram of the method of the present invention. DETAILED DESCRIPTION
[0021] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the specification of the present invention. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0022] In order to better understand the present invention, the above contents are described in detail below in conjunction with specific embodiments.
[0023] Please see the attached Figure 1 The embodiment of the present invention provides an intelligent humanoid trajectory prediction and alarm system based on multimodal video analysis, comprising: Video acquisition module, used to collect video stream data of monitoring scenes; In this embodiment, the role of the video acquisition module is to provide basic data input for the entire system, specifically including collecting the number of video streams in the monitoring scene, and its output provides support for the subsequent data processing module and trajectory modeling module.
[0024] Generally speaking, the video acquisition module must ensure efficient capture of occupant movement within the target area under a variety of lighting conditions, including complex lighting, dynamic backgrounds, and low-light environments. Specifically, the video acquisition module utilizes multimodal data fusion technology to process images captured by both visible and infrared light sources, addressing the limitations of single-modality cameras in certain scenarios.
[0025] Specifically, the video acquisition module includes an infrared camera and a visible light camera, which are used to collect video stream data from the monitored area. The infrared camera can provide high-contrast human silhouette data in low-light environments, while the visible light camera is used to capture clear color image data under normal light conditions.
[0026] In some embodiments, the video acquisition module aligns the data collected by the two cameras using a fusion algorithm. The frames collected by the visible light and infrared cameras are first synchronized according to the timestamps, and then the multimodal images are spatially aligned based on spatial geometric transformations.
[0027] As an option, the video capture module also supports multi-resolution capture mode: Specifically, the module dynamically adjusts acquisition resolution based on scene requirements. For example, in low-light environments, the system prioritizes increasing the infrared camera's frame rate while reducing the visible light camera's sampling resolution to alleviate computational pressure. In bright-light environments, the visible light camera's high-resolution mode is switched, deprioritizing infrared data.
[0028] In one possible implementation, the video stream after multimodal data fusion can be further processed to improve the subsequent object detection accuracy. The fused video stream has the following characteristics: The infrared video stream provides thermal sensing information of the target human figure in the background; Visible light video stream provides shape and color information of the target.
[0029] Expanded technical content: The multimodal fusion algorithm of the present invention adopts a weight allocation strategy, and the weights of different modal data are dynamically adjusted according to the ambient lighting conditions and the difficulty of target detection. The weight factor is calculated using the following formula: Among them, ω r : weight of infrared camera data; ω v : weight of visible light camera data; L r : brightness value of infrared image; L v : Brightness value of visible light image.
[0030] Through this weight adjustment method, the system can prioritize the use of infrared data in low-light conditions, while under normal lighting conditions, the system relies more on visible light data to ensure the quality and stability of the output video stream.
[0031] In one possible implementation, the video acquisition module also supports a region of interest (ROI) mode: Specifically, this module can crop the video capture range based on a preset monitoring area. ROI mode allows the camera to focus on target acquisition in a specific area, such as a doorway, staircase, or fence, while ignoring irrelevant background scenes. This effectively reduces the computational burden on subsequent data processing modules.
[0032] Generally, the range of the ROI can be set by preset parameters or dynamically adjusted based on historical alarm data. For example, in an area where alarms are frequently triggered, the system will prioritize increasing the sampling frequency of that area.
[0033] To enhance system robustness, the video acquisition module of the present invention also features adaptive environmental sensing capabilities. This module uses built-in sensors to monitor ambient light intensity and temperature changes in real time, dynamically adjusting camera parameters based on these sensing results, including but not limited to exposure time, white balance, and gain. In extreme environments (such as heavy rain, sandstorms, and strong backlight), the video acquisition module can also switch to infrared priority mode, directly ignoring data from the visible light camera, thereby improving the detection stability of humanoid targets.
[0034] For example, in nighttime surveillance scenarios, the video acquisition module primarily relies on infrared cameras to collect real-time thermal imaging data of the human body. Meanwhile, the visible light camera, while operating at a low resolution, is still able to capture some environmental features for subsequent data enhancement. In daytime outdoor environments, the infrared camera serves only as a supplement, with the high-resolution visible light camera providing the primary data stream.
[0035] Therefore, through the video acquisition module, the system can achieve high-precision capture of humanoid targets under various lighting conditions. Multimodal data fusion and adaptive parameter adjustment mechanisms make the acquired data more robust, ensuring that subsequent modules can maintain consistent performance in different scenarios.
[0036] The video acquisition module mentioned above is the input source of the entire intelligent human recognition alarm system, and its output data stream provides a guarantee for the accuracy of the subsequent data processing module and trajectory modeling module.
[0037] A data processing module extracts motion state information of the target humanoid based on the video stream data, wherein the motion state information includes the joint position, speed and direction of the target humanoid; In this embodiment, the data processing module is used to extract the target humanoid's motion state information from the video stream captured by the video acquisition module. This primarily involves extracting and calculating the target humanoid's joint positions, velocity, and motion direction. The processing results provide the necessary input data for the trajectory modeling module, enabling dynamic modeling and prediction of the target's behavior.
[0038] Generally speaking, the data processing module needs to be able to detect target human figures in real time within a video stream and comprehensively analyze the motion information from the detection results. Specifically, the module uses pose estimation to obtain key point data of the target and combines it with time series analysis to calculate its speed and direction of motion. Through this module's processing, the system can quickly and accurately extract the dynamic characteristics of the target human figure in complex scenes.
[0039] In the present invention, the data processing module not only supports single target detection, but also can separate and track multiple targets in a multi-target scenario, ensuring the integrity and accuracy of the output data.
[0040] Specifically, the data processing module receives continuous video stream frames from the video acquisition module and detects and analyzes the humanoid target in each frame. The motion state of the target humanoid consists of three main components: joint position, joint velocity, and joint motion direction.
[0041] Specifically, joint locations are extracted from each frame of the video stream using a pose estimation algorithm. This pose estimation algorithm uses a deep learning model (such as OpenPose or HRNet) that can identify multiple key parts of the human body, including the head, shoulders, elbows, wrists, hips, knees, and ankles.
[0042] In some embodiments, the spatial coordinates of the key points are expressed in two dimensions as: Among them, P t represents the target joint point set at time t; (x i ,y i ) represents the two-dimensional coordinate of the i-th joint point; N represents the total number of joint points.
[0043] The joint velocity is obtained by calculating the rate of change of the joint coordinates over time. The calculation formula for the joint velocity is: Among them, v i is the velocity of joint point i; Δx i and Δy i is the coordinate difference of the joint point at two consecutive time points; Δt is the time interval between consecutive time points.
[0044] The joint point direction is calculated based on the components of the velocity. The joint point direction can be described by the following formula: Among them, θ i is the movement direction of joint point i; Δy i and Δx i are the vertical and horizontal displacements of the joint point at two consecutive time points, respectively.
[0045] As an option, the data processing module also supports multi-frame accumulation analysis of the target's motion state: Specifically, this module can perform cumulative calculations on the position, velocity, and direction of each joint point in consecutive frames to improve the robustness and accuracy of detection. The result of the cumulative calculation can be expressed by the following formula: in, is the average velocity of joint point i in n frames; v i (t) is the velocity of joint point i at time t, n represents the number of frames in the time window, that is, the total number of consecutive frames selected when calculating the average velocity, and represents the time point corresponding to the end frame of the time window. The time window starts from t1+n and lasts for n frames.
[0046] In one possible implementation, the data processing module also integrates a target tracking algorithm: The target tracking algorithm combines motion state information with target features (such as color and outline) to track multiple targets. In a multi-target scenario, the system distinguishes multiple targets by assigning unique target IDs and records the motion state of each target.
[0047] In general, the target tracking algorithm is implemented based on the Kalman filter, and its update equation is as follows: P t+1|t =FP t F T +Q Among them? represents the predicted value of the state at time t+1 based on the information at time t (the predicted system state vector), and F is the state transition matrix, which describes how the system state transfers from time t to time t+1, for example, how the position of an object changes over time. is the state estimate at time t (estimated system state vector), B is the control input matrix, describing the external control input U t Impact on system state (e.g. acceleration, velocity, etc.); P t+1|tis the state covariance matrix of the predicted state at time t+1 based on the information at time t, which represents the uncertainty of the predicted state. t is the state covariance matrix at time t, representing the uncertainty of the current state. T is the transpose of the state transfer matrix, which is used to maintain the symmetry of matrix calculation. Q is the process noise covariance matrix, which represents the influence of uncertainty or random disturbance within the system.
[0048] To further improve the performance of the data processing module, some embodiments of the present invention incorporate anomaly detection methods based on time series analysis. By analyzing the patterns of changes in the target's motion state over time, the system can identify possible abnormal behaviors. For example, a target with significant changes in speed or direction between consecutive frames may indicate rapid running or sudden changes in direction.
[0049] In addition, the data processing module of the present invention supports dynamic modeling of the scene background to filter static objects and background interference. In one possible implementation, the module extracts dynamic targets through a background subtraction algorithm and optimizes it in combination with the target's motion state information.
[0050] The trajectory modeling module builds the target person's trajectory model based on the target person's motion state and generates the target person's predicted trajectory through a dynamic optimization algorithm; In this embodiment, the trajectory modeling module constructs a trajectory model of the target individual based on the target individual's motion state information extracted by the data processing module and generates a predicted trajectory using a dynamic optimization algorithm. This module plays a bridging role in the overall system, and its output, the predicted trajectory, directly provides data support for the risk assessment module.
[0051] Typically, the trajectory modeling module performs dynamic modeling based on the target person's historical motion state. Specifically, this module utilizes the target person's joint position, velocity, and direction, and employs optimal control methods and environmental constraints to construct a trajectory model that matches the target's motion characteristics. Based on this, a dynamic optimization algorithm is then used to generate a predicted trajectory for the target person. This approach effectively addresses trajectory changes in complex environments, thereby improving the system's prediction accuracy.
[0052] Specifically, the trajectory modeling module first receives the target person's motion state information extracted by the data processing module, including joint position, speed, and direction. Based on this motion state information, the trajectory modeling module constructs a trajectory model of the target person, describing the target's motion characteristics, and optimizes it using a dynamic optimization algorithm.
[0053] Specifically, the target person's trajectory can be expressed as a function x(t) of time t. In a two-dimensional plane, the target person's trajectory is modeled as follows: Among them, x i (t) and y i (t) represents the horizontal and vertical positions of joint point i at time t; N represents the number of joint points of the target person.
[0054] To predict the target person's trajectory, the system needs to combine its historical trajectory and current motion state information to build a model. Generally, this modeling process is described by the optimal control method.
[0055] As an option, the present invention uses a Lagrangian optimization model for trajectory modeling: Specifically, the trajectory model is generated by solving the optimal solution of the Lagrangian function. The dynamic evolution of the target trajectory can be expressed as the following optimization problem: in, is the Lagrangian function used to describe the trajectory characteristics of the target; t0 and t1 represent the start time and end time of the trajectory respectively; x(t) is the trajectory state at time t; is the velocity of the trajectory, which indicates the motion intensity of the target.
[0056] In one possible implementation, the Lagrangian function consists of kinetic energy and potential energy: Specifically, the Lagrangian function has the following form: in, is the kinetic energy of the target person, describing the intensity of his motion; U(x, t) is the potential energy function, representing the constraint of the scene environment on the target motion; m is the parameter of the target motion mass, which can be approximately taken as a unit value.
[0057] The potential energy function U(x,t) is used to describe the restrictions imposed by the scene environment on the target trajectory, including the risk area characteristics and behavioral resistance characteristics. The formula is as follows: U(x,t)=w1R(x)+w2B(x) Among them, R(x) is the risk area function, which is used to describe the risk of the target's location; w2B(x) is the behavioral resistance function, which is used to describe the obstruction of the external environment to the target's movement; w1 and w2 are the weight parameters of the environmental constraints, which are dynamically adjusted to adapt to different scenarios.
[0058] The trajectory is optimized using the Euler-Lagrange equations: In order to obtain the predicted value of the target trajectory, the Euler-Lagrange equation is used in this embodiment to optimize the target trajectory. The specific calculation formula is: in, Describes the contribution of a system's velocity at a given moment to its motion, such as its kinetic energy during travel. Describes contributions to the motion from the system's position, such as potential energy due to external forces or obstacles in the scene. Used to reflect the impact of speed changes (i.e. acceleration) on the trajectory.
[0059] Substituting the above Lagrangian function, we get the dynamic equation of the target trajectory: in, is the acceleration of the target trajectory; is the gradient of the potential energy with respect to the trajectory.
[0060] By discretizing the above equation, the predicted value of the target trajectory can be calculated within the time step Δt: Among them, x t and x t+1 are the trajectory positions at time t and t+1 respectively; v t is the trajectory speed at time t; a t is the trajectory acceleration at time t.
[0061] In some embodiments, the trajectory modeling module can incorporate real-time feedback data to update the model. For example, after an alarm is triggered, the system will record the deviation between the actual trajectory and the predicted trajectory and dynamically adjust the trajectory model by optimizing the weight parameters w1 and w2 in the potential energy function.
[0062] Furthermore, the trajectory modeling module can adapt to multi-target scenarios, generating multiple predicted trajectories by simultaneously processing the trajectory information of multiple targets. During this process, the system assigns a unique ID to each target to ensure the accuracy and traceability of trajectory predictions.
[0063] For example, in an industrial plant monitoring scenario, the trajectory modeling module analyzes the movement of a target person and constructs a historical trajectory model. Combining risk areas (such as high-risk areas around hazardous equipment) and resistance characteristics (such as movement obstacles in crowded areas) in the environment, the module generates a predicted trajectory for the target and identifies potential proximity to high-risk areas. This information is then passed to the risk assessment module.
[0064] Therefore, through the above approach, the trajectory modeling module of the present invention can accurately construct the target individual's motion trajectory and generate a predicted trajectory through a dynamic optimization algorithm. Combining the Lagrangian optimization model and the Euler-Lagrange equations, the module achieves high-precision modeling and optimization of the target trajectory. The module's output not only enhances the system's predictive capabilities but also provides critical support for the risk assessment module. The completeness and detailed nature of the formula ensures the system's mathematical logic rigor and operational feasibility.
[0065] The risk assessment module calculates the behavioral risk value based on the predicted trajectory of the target person and the characteristics of the scene environment; the alarm module compares the behavioral risk value with the alarm threshold and triggers the alarm based on the comparison result; In this embodiment, the risk assessment module quantifies the risk of the target behavior based on the predicted trajectory of the target person generated by the trajectory modeling module and the characteristics of the scene environment, outputting a behavioral risk value. This module provides decision-making basis for the alarm module, and its accuracy and real-time performance directly affect the system's response to abnormal behavior.
[0066] Generally, the risk assessment module calculates the probability of a target person entering a high-risk area or performing potentially abnormal behavior by analyzing the correlation between the predicted trajectory and scene characteristics. Specifically, this module needs to model the target person's motion characteristics and the risk distribution in the scene, and quantify the risk value through probabilistic methods. In addition, the risk assessment module can incorporate dynamic environmental factors (such as changes in the target's motion speed and the density of risk areas) to improve the accuracy of the assessment without increasing computational complexity.
[0067] Specifically, the risk assessment module receives the target person's predicted trajectory generated by the trajectory modeling module And scenario environment characteristic data. Scenario environment characteristics typically include risk area characteristics and behavioral resistance characteristics. This data is quantitatively described through scenario modeling, specifically the degree of match between the target person's trajectory and the risk area.
[0068] Calculation method of risk value Specifically, the calculation formula of the behavioral risk value R is: Among them, R represents the degree of behavioral risk of the target person on the predicted trajectory; is the conditional probability, indicating the predicted trajectory Corresponding abnormal behavior A t The probability of an event occurring; W is the risk weight, which is related to the characteristics of the scenario environment.
[0069] In the present invention, the conditional probability The calculation of is based on the target behavior distribution fitted by Gaussian mixture model (GMM): Among them, μ is the mean of the behavior distribution, representing the central position of the normal trajectory distribution; Σ is the covariance matrix, representing the degree of dispersion of the trajectory distribution; |Σ| is the determinant value of the covariance matrix; n is the dimension of the trajectory data, represents the transpose of the deviation vector.
[0070] By comparing the predicted trajectory with the trajectory distribution of normal behavior, the system can evaluate whether the target person has abnormal behavior.
[0071] Modeling of scene environment characteristics The scene environment characteristics mainly consist of risk area characteristics and behavior resistance characteristics, representing the behavior risk level of the target person in different areas and the constraints of the external environment on movement respectively.
[0072] Generally, the risk area characteristic R(x) represents the risk value for the target person to enter a specific area, and its calculation formula is: Among them, R(x) is the risk value at position x; ρ i is the risk coefficient of the risk area Ω i ; 1(x∈Ω i ) is an indicator function, which takes the value of 1 when the position x is in the risk area Ω i and 0 otherwise; k is the number of risk areas in the scene.
[0073] As an option, the behavior resistance characteristic B(x) describes the external resistance to the target movement, and the formula is as follows: B(x) = k·f(x) Among them, B(x) is the resistance value at position x; k is the resistance coefficient; f(x) is the distance function between the current position of the target and the resistance source.
[0074] In a possible implementation, the system will dynamically adjust the risk area characteristics and behavior resistance characteristics. For example, in a monitoring scenario, when the target approaches the boundary area, the system will increase the weight value ρ i .
[0075] Hierarchical evaluation of behavior risk value When evaluating the behavior risk of the target, the system determines the risk level by comparing the calculated risk value R with the preset alarm threshold. Generally, the system divides the risks into the following three categories: High risk R>T1 The target behavior is highly abnormal and the alarm is triggered immediately; Medium risk T2<R≤T1 The target behavior is somewhat abnormal and the alarm is triggered with a delay; Low risk R≤T2: The target behavior is normal and only relevant information is recorded.
[0076] As a possible implementation method, the risk thresholds T1 and T2 can be dynamically adjusted according to different scenarios, for example, raising the medium-risk and high-risk thresholds in high-density crowd scenarios.
[0077] To further improve the performance of the risk assessment module, the present invention supports risk trend analysis based on time series. For example, the system can analyze the changing trend of the target behavior risk value to determine whether the target is a potential threat. The calculation formula is: ΔR=R t+1 -R t Among them, ΔR is the change in risk value; R t and R t+1 are the behavioral risk values at time t and t+1 respectively.
[0078] If ΔR is greater than the positive threshold for multiple consecutive times, the system considers the target behavior to be abnormal and triggers an alarm in advance.
[0079] For example, in a surveillance scenario, a target person's predicted trajectory approaches a high-risk area (such as a fence). The risk assessment module calculates the degree of match between their trajectory and the distribution of the area and assesses the target's behavioral risk. If the risk value exceeds a preset threshold, the system determines that the target may have attempted to climb over the fence and sends an alarm signal to the alarm module.
[0080] Therefore, the risk assessment module calculates the target individual's behavioral risk by combining predicted trajectories with the characteristics of the scene environment. Its core technologies include conditional probability modeling, dynamic adjustment of risk areas and behavioral resistance, and a risk grading assessment mechanism. This module enables the system to accurately quantify behavioral risk, providing a reliable basis for the alarm module, enabling early detection and real-time response to abnormal behavior. The integrity of the formula ensures the scientific and rigorous nature of the assessment process.
[0081] The optimization module performs feedback optimization on the trajectory model based on the deviation between the actual trajectory and the predicted trajectory after the alarm is triggered.
[0082] In this embodiment, after an alarm is triggered, the optimization module performs feedback optimization on the trajectory model based on the deviation between the actual trajectory and the predicted trajectory, thereby improving trajectory prediction accuracy and overall system adaptability. Connected between the risk assessment module and the trajectory modeling module, the optimization module uses real-time trajectory data feedback to dynamically adjust model parameters, enabling the system to continuously adapt to different scenarios and changes in target behavior.
[0083] Typically, the optimization module compares the actual trajectory data after an alarm is triggered with the corresponding predicted trajectory and calculates the deviation. Subsequently, a feedback learning mechanism is used to adjust the core parameters of the trajectory model (such as dynamic optimization weights and environmental characteristic parameters) to optimize the trajectory modeling process. In this way, the optimization module can reduce model prediction errors and improve the system's performance in complex scenarios.
[0084] Specifically, the optimization module receives the actual trajectory X of the target after the alarm is triggered from the alarm module. act (t), and combined with the predicted trajectory generated by the trajectory modeling module Calculate the deviation between the two. The quantitative formula of the deviation is as follows: Where ΔX(t) represents the trajectory deviation vector at time t; X act (t) represents the actual trajectory at time t; represents the predicted trajectory at time t.
[0085] Feedback optimization mechanism Specifically, the optimization module adjusts the parameters of the trajectory model according to the deviation ΔX(t). Generally, the optimization of trajectory model parameters is achieved through the gradient descent method, and its update formula is: Among them, Θ new is the updated trajectory model parameter; Θ old are the trajectory model parameters before optimization; η is the learning rate, which controls the optimization step size; J is the objective function, which is used to quantify the model error; is the gradient of the trajectory model error with respect to the parameters.
[0086] Alternatively, the objective function J can use the squared norm of the trajectory deviation: Among them, ∥ΔX(t)∥ 2 is the sum of squares of trajectory deviations, representing the error between the predicted trajectory and the actual trajectory.
[0087] Through the above optimization mechanism, the system can dynamically adjust the weight parameters in the trajectory model to make it more consistent with the actual motion characteristics of the target.
[0088] Dynamically adjust parameters In one possible implementation, the optimization module will focus on optimizing the core parameters in the trajectory modeling module, such as environmental constraint parameters (risk area weight w1, resistance weight w2) and dynamic optimization parameters (such as the acceleration term in the state transition matrix F). The specific adjustment method is as follows: Environmental constraint parameter adjustment Based on the distribution characteristics of the deviation, the optimization module dynamically adjusts the risk area weight w1 and the resistance weight w2 to make the trajectory model more suitable for the current scenario. For example, when the target trajectory approaches a high-risk area but no abnormal behavior occurs, the optimization module reduces the risk area weight w1 to reduce the error amplification of the prediction results.
[0089] Dynamic optimization parameter adjustment The optimization module modifies the dynamic parameters in the state transfer matrix F. For example, in low-speed motion scenarios, the weight of the acceleration term is reduced to reduce the prediction error caused by speed fluctuations.
[0090] In the present invention, the optimization module also supports a distributed optimization method based on federated learning. In a multi-node monitoring scenario, each monitoring node independently performs the feedback optimization process and uploads the optimized model parameters to the cloud server. The cloud server aggregates the optimization results of all nodes and generates global model parameters. The formula is as follows: Among them, Θ global : global model parameters; M: number of monitoring nodes; Θ i : Model parameters of the i-th node.
[0091] The aggregated global model parameters will be distributed to each monitoring node to achieve unified optimization in distributed scenarios.
[0092] To further improve the optimization effect, the optimization module also combines time series analysis methods to model the changing trend of trajectory deviation. For example, by calculating the rate of change of trajectory deviation, the system can identify the convergence speed of model error and dynamically adjust the learning rate η. The formula is as follows: η=η0·exp(-α·∥ΔX(t)∥) Where η0 is the current learning rate, η is the current learning rate, which is used to control the step size of the model parameter update. The larger the learning rate, the larger the optimization step size; the smaller the learning rate, the smaller the optimization step size. -α is the learning rate attenuation coefficient, ∥ΔX(t)∥ is the second norm of the trajectory deviation at the current time t, and exp(·) is the natural exponential function.
[0093] Therefore, the dynamic learning rate mechanism can accelerate the convergence of the model and improve the optimization efficiency.
[0094] For example, in a campus monitoring scenario, when the system detects a target person approaching a fence but exhibiting no unusual behavior, a certain deviation occurs between the actual trajectory and the predicted trajectory. The optimization module records the deviation and adjusts the weight w1 of the risk area in the trajectory model based on the deviation, thereby reducing errors in subsequent prediction results. Furthermore, through a federated learning mechanism, the optimized parameters are synchronized to other monitoring nodes, improving the overall system's predictive capabilities.
[0095] Therefore, the optimization module uses feedback to optimize the deviation between the actual trajectory and the predicted trajectory after an alarm is triggered, dynamically adjusting the core parameters of the trajectory model, improving trajectory prediction accuracy and system adaptability. Combining feedback optimization mechanisms with federated learning technology, the module not only adapts to complex environments with multiple scenarios and targets, but also enables distributed, unified model optimization.
[0096] Please see the attached Figure 2 The present invention provides a method for intelligent humanoid trajectory prediction and alarm system based on multimodal video analysis, comprising the following steps: Collect video stream data of the monitoring scene and obtain the movement information of people in the target area; Analyze the video stream data to extract the target person's motion state information, including the target person's joint position, speed, and direction; A trajectory model is constructed based on the target person's motion state, and a predicted trajectory of the target person is generated through a dynamic optimization algorithm; a behavioral risk value is calculated based on the predicted trajectory and scene environmental characteristics, wherein the scene environmental characteristics include risk area characteristics and behavioral resistance characteristics; Compare the behavior risk value with the preset alarm threshold. If the behavior risk value exceeds the alarm threshold, an alarm is triggered through the alarm module; The trajectory model is optimized based on the deviation between the actual trajectory and the predicted trajectory after the alarm is triggered, so as to improve the accuracy of trajectory prediction and the overall adaptability of the system.
[0097] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. An intelligent humanoid trajectory prediction and alarm system based on multimodal video analysis, characterized in that: include: Video acquisition module, used to collect video stream data of monitoring scenes; A data processing module extracts motion state information of a target humanoid based on video stream data, wherein the motion state information includes joint position, speed and direction of the target humanoid; The trajectory modeling module builds the trajectory model of the target person based on the target person's motion state and generates the predicted trajectory of the target person through a dynamic optimization algorithm; The risk assessment module calculates the behavior risk value based on the predicted trajectory of the target person and the characteristics of the scene environment; The alarm module compares the behavior risk value with the alarm threshold and triggers an alarm based on the comparison result; The optimization module performs feedback optimization on the trajectory model based on the deviation between the actual trajectory and the predicted trajectory after the alarm is triggered.
2. According to claim 1, the intelligent humanoid trajectory prediction and alarm system based on multimodal video analysis is characterized in that: The video acquisition module includes an infrared camera and a visible light camera, and solves the problem of capturing human targets under low light or complex lighting conditions through multimodal data fusion.
3. The intelligent humanoid trajectory prediction and alarm system based on multimodal video analysis according to claim 1 is characterized in that: The data processing module extracts key point data of the target human figure through a posture estimation algorithm, and the key points include joint points of the head, shoulders, elbows, knees and feet.
4. The intelligent humanoid trajectory prediction and alarm system based on multimodal video analysis according to claim 3 is characterized in that: The trajectory modeling module implements trajectory modeling through an optimal control algorithm, and the optimal control algorithm predicts the future trajectory of the target person based on the historical trajectory of the target person in combination with scene environmental constraints.
5. The intelligent humanoid trajectory prediction and alarm system based on multimodal video analysis according to claim 4 is characterized in that: The optimal control algorithm for trajectory prediction in the trajectory modeling module is solved based on a Lagrangian function, which includes the kinetic energy of the target person and the potential energy in the scene environment, wherein the potential energy is described by risk area characteristics and behavioral resistance characteristics.
6. The intelligent humanoid trajectory prediction and alarm system based on multimodal video analysis according to claim 1 is characterized in that: The risk assessment module fits the historical behavior distribution of the target person based on the Gaussian mixture model, and calculates the behavior risk value according to the matching degree between the predicted trajectory and the behavior distribution.
7. The intelligent humanoid trajectory prediction and alarm system based on multimodal video analysis according to claim 6 is characterized in that: The risk assessment module compares the calculated behavior risk value with a preset threshold value, and triggers different levels of alarms through a graded alarm mechanism, wherein the alarm levels include high risk alarm, medium risk alarm and low risk record.
8. The intelligent humanoid trajectory prediction and alarm system based on multimodal video analysis according to claim 1 is characterized in that: The optimization module optimizes and adjusts the potential energy function parameters of the trajectory modeling module based on a feedback learning mechanism.
9. The intelligent humanoid trajectory prediction and alarm system based on multimodal video analysis according to claim 8, characterized in that: The optimization module adopts a federated learning method to optimize the local trajectory model through each surveillance camera node, and uploads the optimization result to the cloud. The global model aggregated in the cloud is used to update the trajectory modeling module of each node.
10. An intelligent humanoid trajectory prediction and alarm method based on multimodal video analysis, based on an intelligent humanoid trajectory prediction and alarm system based on multimodal video analysis according to any one of claims 1 to 9, characterized in that: The following steps are involved: Collect video stream data of monitoring scenes and obtain movement information of people in the target area; Analyze the video stream data to extract the motion state information of the target person, wherein the motion state information includes the joint position, speed and direction of the target person; A trajectory model is constructed based on the target person's motion state, and a predicted trajectory of the target person is generated through a dynamic optimization algorithm; Calculating a behavior risk value according to the predicted trajectory and scene environment characteristics, wherein the scene environment characteristics include risk area characteristics and behavior resistance characteristics; Compare the behavior risk value with the preset alarm threshold. If the behavior risk value exceeds the alarm threshold, an alarm is triggered through the alarm module; The trajectory model is optimized based on the deviation between the actual trajectory and the predicted trajectory after the alarm is triggered to improve the accuracy of trajectory prediction and the overall adaptability of the system.
Citation Information
Cited By
Intelligent security intercom system based on PIR and AI behavior analysis
CN120343209A
Intelligent security intercom system based on PIR and AI behavior analysis
CN120343209B
Wafer level test-oriented multi-probe cooperative contact control and adjustment method and system
CN120539571A
Multi-probe collaborative contact control and adjustment method and system for wafer-level testing
CN120539571B
Smart home control method and system based on multi-modal fusion and deep learning
CN120630745A