Real-time video stream behavior identification and early warning system

By constructing a dynamic behavior recognition model and cross-modal difference analysis, combined with the physical space constraints of the monitoring area, the video surveillance system has achieved accurate behavior recognition and anomaly warning in complex scenarios. This solves the problems of incomplete feature extraction and resource waste in existing technologies, and improves recognition accuracy and management efficiency.

CN120894831AActive Publication Date: 2025-11-04GAOZI TECHNOLOGY (SHENZHEN) CO LTD

Patent Information

Application Number
CN202511227289.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-29
Publication Date
2025-11-04
Estimated Expiration
2045-08-29

AI Technical Summary

Technical Problem

Existing video surveillance systems struggle to achieve accurate behavior recognition and anomaly warning in complex scenarios. In particular, feature extraction is incomplete when there is occlusion, changes in lighting, or rapid target movement. Furthermore, they cannot combine the physical space constraints of the monitored area with the target's movement trajectory to effectively predict abnormal behavior, resulting in a lack of targeted warnings and a waste of computing resources.

Method used

A dynamic behavior recognition model is constructed through a spatiotemporal feature modeling module, which combines skeletal key points, 3D coordinates, motion optical flow vector fields, and micro-expression muscle activity intensity spectra to generate theoretical behavior pattern vectors. A behavior fragment extraction module performs cross-modal difference analysis to generate multimodal difference feature tensors. An anomaly propagation modeling module simulates the propagation path of abnormal behavior in combination with the physical space constraints of the monitored area, and a risk area positioning module marks high-risk areas. An early warning strategy generation module dynamically configures monitoring parameters, including high frame rate micro-expression capture and motion trajectory disturbance testing.

Benefits of technology

It achieves accurate identification and efficient early warning of target object behavior, reduces behavior identification bias, can promptly detect potential abnormal behavior, clearly present the propagation path and impact range of abnormal behavior, optimize the utilization of computing resources, and improve management efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120894831A_ABST
    Figure CN120894831A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of video behavior recognition, and discloses a behavior recognition and early warning system for a real-time video stream. The system comprises a spatio-temporal feature modeling module, a behavior fragment extraction module, an anomaly propagation modeling module, a risk area positioning module and an early warning strategy generation module. The spatial-temporal feature modeling module builds a dynamic model based on historical data, captures a skeleton key point three-dimensional coordinate sequence, a motion optical flow vector field and a micro-expression intensity spectrum, and outputs a theoretical behavior mode vector; the behavior fragment extraction module generates a multi-modal difference feature tensor through cross-modal difference analysis; the exception propagation modeling module generates an exception propagation path risk probability distribution cloud picture in combination with spatial constraint and trajectory information; the risk area positioning module identifies a high-risk area and marks a boundary; and the early warning strategy generation module dynamically configures monitoring parameters, starts high-frame-rate micro-expression capture for a high-risk area, and performs a track disturbance test on an adjacent area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video behavior recognition technology, specifically to a real-time video stream behavior recognition and early warning system. Background Technology

[0002] With the widespread application of video surveillance technology in public safety, traffic management, and industrial production, the demand for accurate identification and anomaly warning of target object behavior in real-time video streams is becoming increasingly urgent. Current mainstream behavior recognition systems mostly rely on single-modal data analysis, such as judging human movement solely through skeletal keypoint coordinates or identifying object movement trajectories based solely on motion optical flow vector fields. This single-modal analysis method has significant limitations. When there are occlusions, changes in lighting, or rapid target movement in the monitored scene, incomplete feature extraction can easily occur, leading to behavior recognition errors.

[0003] Existing systems lack effective propagation modeling mechanisms for handling abnormal behavior. Most systems can only issue warnings after detecting a single abnormal behavior, failing to combine the physical spatial constraints of the monitored area (such as passage width and obstacle distribution) with the target object's movement trajectory to analyze the possible spread paths and impact range of the abnormal behavior. For example, in densely populated areas, when an abnormal running behavior of a person is detected, existing systems struggle to predict whether this behavior will trigger a chain reaction among surrounding people, and cannot accurately identify high-risk gathering areas. This results in a lack of targeted warning information, making it difficult to assist management personnel in taking timely and effective intervention measures.

[0004] Traditional monitoring systems often use fixed parameter configurations, making it impossible to dynamically adjust them according to the risk level of a scenario. Maintaining the same frame rate and feature extraction frequency for monitoring different areas not only wastes computational resources in low-risk areas but may also miss key behavioral features in high-risk areas due to insufficient sampling frequency. For example, in station waiting areas and entrance / exit passages—areas with different risk levels—fixed monitoring parameters cannot meet the need for rapid anomaly response in the passage area and also consume unnecessary computational resources in the waiting area, affecting the overall system efficiency. Summary of the Invention

[0005] The purpose of this invention is to provide a real-time video stream behavior recognition and early warning system to solve the problems mentioned in the background art.

[0006] To achieve the above objectives, the present invention provides a real-time video stream behavior recognition and early warning system, the system comprising:

[0007] Spatiotemporal feature modeling module: Based on the historical behavior video data of the target object, a dynamic behavior recognition model is constructed. The model captures the three-dimensional coordinate sequence of the skeletal key points of the target, the motion optical flow vector field and the intensity spectrum of micro-expression muscle activity in the current video stream in real time. The theoretical behavior pattern vector is output through the dynamic behavior recognition model.

[0008] Behavior segment extraction module: Perform cross-modal difference analysis on the theoretical behavior pattern vector and the measured behavior vector of the video analysis engine. The cross-modal difference analysis includes temporal continuity deviation, spatial trajectory offset and micro-expression spectrum similarity, and generates a multimodal difference feature tensor at the target object level.

[0009] Anomaly propagation modeling module: Input the multimodal difference feature tensor into the scene topology inference network, and combine the physical space constraints of the monitoring area with the target object's movement trajectory information to generate a risk probability distribution cloud map of the abnormal behavior propagation path;

[0010] Risk area location module: Identifies high-risk cluster areas based on the risk probability distribution cloud map, and marks the physical boundaries of abnormal behavior areas where the probability value exceeds a preset risk probability threshold;

[0011] Early warning strategy generation module: Configures monitoring parameters based on the risk probability distribution cloud map, including:

[0012] High frame rate micro-expression capture mode is enabled for high-risk gathering areas;

[0013] Test the motion trajectory perturbation of targets in the vicinity.

[0014] Preferably, the spatiotemporal feature modeling module specifically includes:

[0015] Historical behavior feature mining: Multi-scale spatiotemporal decomposition processing is performed on the historical behavior video data of the target object, including: using an adaptive spatiotemporal pyramid to extract the energy distribution of steady-state and transient motion of the skeletal key point sequence, establishing the correlation mapping between motion vector field and behavioral intention through optical flow coupling analysis, and using sequence alignment algorithm to calibrate the intensity pattern of micro-expression activity in different scenarios;

[0016] Dynamic behavior recognition model construction: The processed historical behavior data is input into the fusion prediction architecture, which includes:

[0017] Temporal modeling units based on behavior evolution curves are used to generate basic behavior prediction vectors;

[0018] A 3D convolutional network with embedded spatial attention mechanism corrects prediction bias caused by changes in viewpoint;

[0019] Based on the optical flow feature compensator, the prediction weights are dynamically adjusted according to the real-time captured motion optical flow vector field.

[0020] The edge computing unit synchronously collects the position drift rate and velocity fluctuation parameters of key skeletal points, the amplitude gradient distribution and directional consistency coefficient of motion optical flow vectors, and the frequency band energy entropy value of micro-expression intensity spectrum.

[0021] Theoretical value generation: Input the real-time captured data into the dynamic behavior recognition model to obtain the theoretical behavior pattern vector.

[0022] Preferably, the behavior fragment extraction module specifically includes:

[0023] Time continuity deviation calculation: The theoretical behavior pattern vector and the measured behavior vector are compared by sliding comparison with a preset time segment. The non-uniformly sampled behavior sequence is aligned by the sequence dynamic warping algorithm. The cumulative trajectory deviation in each segment is calculated to generate a time domain deviation feature matrix.

[0024] Spatial trajectory offset detection: Spatial domain transformation decomposition is performed on the three-dimensional coordinates of skeletal key points of theoretical and measured values, the trajectory curvature similarity ratio is calculated, the spatial offset index of each joint node is extracted, and a spatial offset feature matrix is ​​constructed.

[0025] Micro-expression spectral similarity assessment: Based on the pattern matching algorithm, the distance distribution between theoretical values ​​and measured micro-expression intensity spectra is measured, the phase synchronization error of activation intensity of specific muscle groups is calculated, the divergence difference of spectral energy distribution is quantified, and micro-expression similarity vectors are generated.

[0026] Multimodal difference feature tensor generation: The temporal deviation feature matrix, spatial offset feature matrix, and micro-expression similarity vector are fused with high-order features. Through modal importance weighted standardization, scale differences are eliminated, and a third-order multimodal difference feature tensor with dimensions of [target identifier x time segment x modal type] is output.

[0027] Preferably, the anomaly propagation modeling module specifically includes:

[0028] Scene topology modeling: Construct a spatial topology map of the monitoring area based on the target object's movement trajectory information, mark the physical connectivity attributes between each sub-area, overlay the field of view coverage constraints of the fixed monitoring equipment on the topology map, and generate a physical topology model containing a spatial association matrix and an area reachability matrix;

[0029] Anomaly propagation deduction: Mapping the multimodal differential feature tensor to the corresponding region of the physical topology model;

[0030] Anomaly propagation simulation is performed based on graph convolutional networks. The computation of the anomaly propagation simulation includes:

[0031] Calculate the attenuation factor of anomalous behavior based on the region's physical connectivity attributes;

[0032] Capture anomalous correlation features through cross-regional attention mechanisms;

[0033] A random walk method is used to simulate the propagation path of anomalous behavior in a topological network.

[0034] Risk distribution generation: Statistically count the frequency of abnormal behavior in each sub-region during simulated propagation, combine regional accessibility parameters to calculate the probability value of abnormal behavior residence, and generate a risk probability distribution cloud map covering the entire region;

[0035] Physical boundary delineation: Perform density clustering analysis on the risk probability distribution cloud map to identify high-risk probability clustering areas;

[0036] Based on the location of the target object and its physical topology connection, mark the physical boundaries of the abnormal behavior area.

[0037] Preferably, the early warning strategy generation module specifically includes:

[0038] When the anomaly probability value of a certain area in the risk probability distribution cloud map exceeds a preset risk probability threshold, a monitoring command is sent to the video acquisition terminal belonging to that area, and the following is executed:

[0039] Increase the frame rate for micro-expression capture to 3-5 times the original frame rate;

[0040] Simultaneously enable real-time muscle activity intensity tracking mode to capture frequency band energy mutations in specific facial regions;

[0041] Deploy transient behavior detectors at the edge to record abnormal micro-expression waveform fragments;

[0042] Motion trajectory perturbation test execution: Apply multi-dimensional motion perturbation to the target in the neighborhood area with the largest change in risk probability gradient in the risk probability distribution cloud map.

[0043] Preferably, the multidimensional motion perturbation includes:

[0044] Inject a preset motion trajectory interference signal through a virtual reality interface;

[0045] Theoretical behavioral response calculation: Based on the physical topology model and spatial correlation matrix, calculate the theoretical behavioral response spectrum;

[0046] Acquisition of measured behavioral response: Record the behavioral patterns of each target after the injection of interference signal to obtain the measured behavioral response spectrum;

[0047] Anomaly offset determination: Calculate the characteristic distance between the theoretical behavioral response spectrum and the measured behavioral response spectrum;

[0048] By comparing the characteristic distance between the measured behavioral response spectrum and the theoretical behavioral response spectrum, the degree of deviation of abnormal targets is calculated. When the degree of deviation of abnormal targets exceeds the preset deviation threshold, they are marked as abnormal behavior associated targets.

[0049] Preferably, the spatiotemporal feature modeling module further includes:

[0050] Adaptive normalization processing based on ambient lighting conditions eliminates characteristic noise caused by changes in light intensity;

[0051] The correlation characteristics of skeleton, optical flow, and micro-expression are integrated through a multi-source feature fusion algorithm;

[0052] The output contains a theoretical behavior pattern vector that includes the normal behavior fluctuation range. The theoretical behavior pattern vector is dynamically updated as the behavior pattern evolves.

[0053] Preferably, the anomaly propagation modeling module further includes outputting a set of suspicious target identifiers and an anomaly propagation main path sequence, wherein:

[0054] The suspicious target identifier set is based on the monitoring area covered by the physical boundary of the abnormal behavior area, and the monitoring area is bound to the actual target object to form the suspicious target identifier set;

[0055] The anomaly propagation main path sequence is obtained by counting the occurrence frequency of each path during the anomaly diffusion process simulated by random walk, selecting the path sequence with the highest cumulative occurrence frequency, and outputting an ordered list of nodes reflecting the main propagation trajectory of the abnormal behavior in the scene.

[0056] Preferably, in the construction of the dynamic behavior recognition model:

[0057] The optical flow feature compensator uses a motion trajectory reconstruction algorithm to decompose the motion pattern of the real-time captured motion optical flow vector field and identify the basic motion component and abnormal fluctuation component.

[0058] The prediction weight coefficients of the time series modeling unit are dynamically adjusted based on the energy proportion of the abnormal fluctuation components.

[0059] Preferably, in the behavior fragment extraction module:

[0060] The micro-expression spectrum similarity assessment uses a key frequency band energy comparison algorithm to extract the energy distribution features of preset emotion-related frequency bands;

[0061] The energy density of theoretical and measured data in the emotionally relevant frequency bands is quantitatively compared, and a micro-expression similarity vector is constructed based on the degree of relative difference.

[0062] Compared with the prior art, the beneficial effects of the present invention are:

[0063] Through multi-module collaborative operation, accurate identification and efficient early warning of target object behavior are achieved. At the feature extraction level, the spatiotemporal feature modeling module simultaneously captures three types of multimodal data: 3D coordinate sequences of skeletal key points, motion optical flow vector fields, and micro-expression muscle activity intensity spectra. Based on historical behavioral video data, a dynamic behavior recognition model is constructed to generate theoretical behavior pattern vectors. This multimodal data fusion approach can compensate for the shortcomings of single-modal data in complex scenarios. Even in monitoring scenarios with occlusion, changes in lighting, or rapid target movement, the complementary nature of different modal data can fully extract target behavioral features, reduce biases in the behavior recognition process, and improve the accuracy of identifying various behavior patterns.

[0064] The behavior fragment extraction module compares theoretical behavior pattern vectors with measured behavior vectors through cross-modal difference analysis, generating multimodal difference feature tensors from three dimensions: temporal continuity deviation, spatial trajectory offset, and micro-expression spectrum similarity. This multi-dimensional difference analysis approach can more comprehensively capture the subtle differences between theoretical and actual behaviors, avoiding missed or misjudged abnormal behaviors caused by single-dimensional analysis. This ensures that the system can promptly detect potential abnormal behavior signs, providing detailed feature evidence for subsequent abnormal propagation modeling.

[0065] The anomaly propagation modeling module inputs multimodal differential feature tensors into the scene topology inference network and combines this with the physical space constraints of the monitored area and the target object's movement trajectory information to generate a risk probability distribution cloud map of the anomaly behavior propagation path. This module overcomes the limitation of traditional systems that can only detect single anomalies. By constructing scene topology relationships, it simulates the diffusion patterns of anomalies in different physical spaces, clearly presenting the possible propagation paths and impact ranges of anomalies. Managers can intuitively understand the diffusion trend of anomalies through the risk probability distribution cloud map, predict potential chain reactions in advance, and provide a clear direction for developing intervention strategies.

[0066] The risk area location module identifies high-risk clusters based on a risk probability distribution cloud map and marks the physical boundaries of abnormal behavior areas where the probability value exceeds a preset threshold. This function enables the system to accurately pinpoint areas requiring key attention, avoiding the need for managers to manually check through massive amounts of monitoring footage, reducing manpower costs, and ensuring that intervention measures are precisely applied to high-risk areas, thereby improving management efficiency.

[0067] The early warning strategy generation module dynamically configures monitoring parameters based on the risk probability distribution cloud map. For high-risk cluster areas, a high-frame-rate micro-expression capture mode is activated, enabling more detailed capture of micro-expression changes in target objects within that area, and timely detection of hidden abnormal emotions and behavioral tendencies. Applying motion trajectory perturbation tests to targets in adjacent areas further verifies whether there are potential abnormal behavioral trends, allowing for early risk assessment. This dynamic parameter configuration method rationally allocates computing resources according to the risk level of different areas, ensuring monitoring accuracy in high-risk areas and reducing unnecessary computational consumption in low-risk areas, thus optimizing system resource utilization and improving overall operational efficiency. Attached Figure Description

[0068] Figure 1 This is a timing diagram of the real-time video stream behavior recognition and early warning system described in this invention;

[0069] Figure 2 A flowchart illustrating the operation of the spatiotemporal feature modeling module;

[0070] Figure 3 A flowchart illustrating the operation of the behavior fragment extraction module;

[0071] Figure 4 A flowchart illustrating the workflow of the anomaly propagation modeling module;

[0072] Figure 5 A flowchart for multi-dimensional motion perturbation execution. Detailed Implementation

[0073] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0074] Please see Figure 1 This invention provides a real-time video stream behavior recognition and early warning system, the system comprising:

[0075] This system achieves real-time identification and early warning of target object behavior in video streams through the collaborative work of multiple modules. The spatiotemporal feature modeling module constructs a dynamic behavior recognition model based on historical video data of the target object's behavior. It captures in real-time the 3D coordinate sequence of the target's skeletal key points, the motion optical flow vector field, and the intensity spectrum of micro-expression muscle activity in the current video stream, and outputs a theoretical behavior pattern vector through the model. The behavior segment extraction module performs cross-modal difference analysis between the theoretical behavior pattern vector and the measured behavior vector from the video analysis engine, including temporal continuity deviation, spatial trajectory offset, and micro-expression spectral similarity, generating a multimodal difference feature tensor at the target object level. The anomaly propagation modeling module inputs the multimodal difference feature tensor into the scene topology inference network, combining the physical spatial constraints of the monitored area and the target object's movement trajectory information to generate a risk probability distribution cloud map of the abnormal behavior propagation path. The risk area localization module identifies high-risk cluster areas based on the risk probability distribution cloud map and marks the physical boundaries of abnormal behavior areas with probability values ​​exceeding a preset risk probability threshold. The early warning strategy generation module configures monitoring parameters based on the risk probability distribution cloud map, including enabling a high frame rate micro-expression capture mode for high-risk cluster areas and applying motion trajectory perturbation tests to targets in adjacent areas. The system achieves a complete processing chain from behavioral feature extraction to risk early warning through the above process.

[0076] Example 1: See Figure 2 The implementation of the spatiotemporal feature modeling module involves multi-scale spatiotemporal decomposition processing of historical behavior video data of the target object. An adaptive spatiotemporal pyramid is used to extract the energy distribution of steady-state and transient motions in the skeletal keypoint sequence. This processing decomposes the video sequence into multiple spatiotemporal scales, calculating the distribution characteristics of motion energy at each scale. Steady-state motion components reflect continuous behavioral patterns, while transient motion components capture sudden action features. This decomposition allows for a more accurate description of the spatiotemporal characteristics of behavior. Optical flow coupling analysis establishes a mapping between the motion vector field and behavioral intent. By analyzing the changes in the direction and amplitude of the optical flow vector, the probability of behavioral intent is inferred. This mapping is based on the statistical correlation between the optical flow field and historical behavioral patterns. A sequence alignment algorithm calibrates the intensity patterns of micro-expression activities in different scenarios. This algorithm uses dynamic time warping or similar methods to adjust the time alignment and intensity normalization of micro-expression data obtained under different shooting conditions, ensuring data consistency.

[0077] The construction of the dynamic behavior recognition model inputs processed historical behavior data into the fusion prediction architecture. A temporal modeling unit based on the behavior evolution curve generates a basic behavior prediction vector. This unit uses a recurrent neural network or a temporal convolutional network to learn the temporal evolution of the behavior sequence and outputs a vector representing the expected behavior pattern. A 3D convolutional network with embedded spatial attention mechanism corrects prediction bias caused by viewpoint changes. This network introduces attention weights in the 3D convolution operation, focusing on key spatial regions and reducing feature distortion caused by viewpoint differences. An optical flow feature compensator dynamically adjusts the prediction weights based on the real-time captured motion optical flow vector field. This component analyzes the characteristic changes of the optical flow field and adjusts the contribution of each part of the model accordingly. Multiple real-time data parameters are simultaneously collected through an edge computing unit. The positional drift rate of skeletal keypoints is obtained by calculating the rate of change of keypoint coordinates between consecutive frames, using the formula:

[0078]

[0079] Where: D t The average positional drift rate over time t is represented by N, where N is the number of keypoints and p is the number of keypoints. i (t) represents the three-dimensional coordinates of the i-th keypoint at time t, and ||·|| represents the Euclidean distance. The velocity fluctuation parameter is calculated based on the second-order difference of the position sequence, reflecting the changes in motion acceleration. The amplitude gradient distribution of the motion optical flow vector is obtained by calculating the spatial derivative of the optical flow amplitude, while the directional consistency coefficient is calculated by statistically analyzing the variance of the optical flow direction within a local region. The frequency band energy entropy value of the micro-expression intensity spectrum is calculated by dividing the spectrum of the micro-expression signal and then calculating the energy distribution entropy value of each sub-band, using the following formula:

[0080]

[0081] Wherein: H f This represents the frequency band energy entropy value, where K is the number of frequency bands, and E is the frequency band energy entropy k This is the normalized energy value for the k-th frequency band. This entropy value quantifies the distributional complexity of the micro-expression spectrum energy. In the theoretical value generation stage, the real-time captured data is input into the dynamic behavior recognition model to obtain a theoretical behavior pattern vector. The model processing includes feature encoding, temporal modeling, and feature decoding, ultimately outputting a vector representation of the expected behavior pattern. This vector contains multimodal feature information of the behavior for differential analysis.

[0082] An adaptive normalization process based on ambient lighting conditions is applied to the feature extraction process to eliminate feature noise caused by changes in lighting. This process estimates scene lighting conditions, normalizes the intensity of visual features, and enhances contrast, reducing the impact of lighting changes on feature stability. A multi-source feature fusion algorithm integrates the correlation characteristics of skeleton, optical flow, and micro-expressions, employing weighted fusion or learning-based fusion strategies to combine the feature advantages of each modality, improving the robustness and discriminative power of behavioral representation. The output theoretical behavioral pattern vector contains the normal behavioral fluctuation range, which is determined by statistically analyzing the feature distribution range of historical data, used to define the boundaries of normal behavior. The theoretical behavioral pattern vector is dynamically updated as behavioral patterns evolve. The update mechanism is based on online learning or sliding window statistical methods, enabling the model to adapt to gradual changes in behavioral patterns. The entire implementation emphasizes the fine processing and analysis of multimodal features, achieving comprehensive modeling of behavioral features through the combination of various algorithms and models. The components collaborate with each other through data flow and parameter transfer, forming a complete behavioral feature extraction and modeling pipeline. The processing emphasizes a balance between computational efficiency and accuracy, employing edge computing to reduce data transmission latency and ensure the system can meet real-time processing requirements. The feature representation design considers the spatiotemporal dynamics of behavior, capturing changes in both short-term actions and long-term behavioral patterns. The model's adaptability is achieved through a dynamic update mechanism, enabling the system to cope with changes in environmental conditions and behavioral patterns, maintaining stable recognition performance.

[0083] Example 2: See Figure 3 The implementation of the behavior segment extraction module involves multi-dimensional comparison and analysis of theoretical behavior pattern vectors and measured behavior vectors. The temporal continuity deviation calculation employs a sliding window mechanism, dividing the continuous video stream into fixed-length time segments and aligning theoretical predictions with actual observations within each segment. A sequence dynamic warping algorithm addresses temporal misalignment caused by differences in sampling rates or action execution speeds, dynamically adjusting the correspondence on the time axis to achieve optimal matching between the two sequences in the temporal dimension. The cumulative trajectory deviation calculation considers the differences in feature vectors at each aligned time point, summarizing the overall deviation within the entire time segment through integration to form a feature matrix reflecting temporal dimension differences. This temporal analysis method can capture rhythmic anomalies or disordered action sequences during behavior execution.

[0084] Spatial trajectory offset detection focuses on the differences in motion trajectories of skeletal keypoints in three-dimensional space. A spatial domain transformation is performed on the theoretically predicted skeletal point coordinate sequence and the actual detection results, converting the original coordinate data into a more discriminative feature space. The trajectory curvature similarity ratio calculation quantifies the degree of deviation in the geometric characteristics of the motion path by comparing the curvature differences between the theoretical and actual trajectories at each keypoint. The spatial offset index of each joint node comprehensively considers multiple factors such as positional deviation, velocity differences, and acceleration changes. By weighted fusion of these spatial features, a feature matrix reflecting the overall spatial motion differences is constructed. This spatial analysis method can identify behavioral characteristics such as abnormal body postures, unnatural limb movements, or abnormal interactions with environmental objects.

[0085] Micro-expression spectral similarity assessment focuses on the analysis of subtle facial muscle activity. Pattern matching algorithms compare the theoretical micro-expression intensity spectrum with the actual observed spectrum in multiple dimensions, calculating differences in spectral shape, energy distribution, and phase characteristics. Phase synchronization error analysis of specific muscle group activation intensity focuses on the coordination changes of muscle activity in different facial regions. By comparing the temporal relationship between theoretical predictions and actual observations of muscle activation, the degree of incoordination during expression execution is quantified. The divergence difference calculation of spectral energy distribution uses information theory methods to measure the overall difference between the two spectra, generating a comprehensive index vector reflecting micro-expression anomalies. This micro-expression analysis method can capture subtle flaws in deliberately feigned expressions or abnormal features of emotional fluctuations.

[0086] The generation of multimodal differential feature tensors requires the effective fusion of features from different analytical dimensions. High-order feature fusion techniques employ tensor operations to combine the temporal bias matrix, spatial offset matrix, and micro-expression similarity vectors in a unified high-dimensional space, preserving the original structural relationships of each modality's features. Modality importance-weighted standardization automatically adjusts the contribution weight of each feature modality in the final feature representation based on its discriminative power in behavior recognition, while eliminating scale differences caused by different feature dimensions and numerical ranges. The final output third-order tensor structure is organized along three dimensions: target identifier, time segment, and modality type, forming a multimodal feature representation that comprehensively reflects behavioral anomalies. This tensor structure facilitates the processing module's analysis of behavioral anomaly features from different perspectives.

[0087] Key frequency band energy comparison algorithms have specific application value in micro-expression analysis. This algorithm first identifies several frequency bands closely related to emotional expression; these bands typically correspond to the activity characteristics of specific facial muscle groups. A refined comparison is then performed on the energy distribution of theoretically predicted and actually observed micro-expression data within these key frequency bands, calculating the relative difference in energy density within each band. The selection of emotion-related frequency bands is based on prior knowledge from the facial action coding system, focusing on muscle activity patterns highly correlated with basic emotional expressions. The quantification and comparison of energy density considers not only amplitude differences but also analyzes the changes in energy distribution patterns in the frequency domain, thereby constructing a more comprehensive feature vector of micro-expression anomalies. This frequency band-specific analysis method improves the detection sensitivity of deliberately controlled expressions.

[0088] The overall processing flow of the behavior fragment extraction module emphasizes multi-angle and multi-level anomaly feature mining. Temporal continuity analysis reveals temporal anomalies in the behavior execution process, spatial trajectory detection discovers geometric deviations in motion paths and postures, and micro-expression analysis captures abnormal patterns in subtle facial muscle activities. These three analytical dimensions complement each other to jointly construct a comprehensive description of abnormal behavior. The module adopts a phased processing strategy, first performing refined feature extraction and difference calculation within each modality, then performing cross-modal feature fusion and standardization, ultimately forming a structured multimodal difference feature representation. This processing approach preserves the independence of features in each modality while establishing intermodal relationships, providing rich input features for anomaly propagation analysis. The entire implementation process emphasizes the balance between computational efficiency and feature discriminative power, employing a parallel computing framework to accelerate multimodal feature extraction while maintaining the discriminativeness and interpretability of the features.

[0089] Example 3: See Figure 4The implementation of the anomaly propagation modeling module begins with the physical space modeling of the monitoring scene. Based on the target object's movement trajectory data (derived from the 3D coordinate sequence of skeletal key points output by the spatiotemporal feature modeling module: by calculating the positional changes of skeletal key points within 10 consecutive frames, the system fits the target object's movement path, eliminates single-point coordinate deviations caused by occlusion, and ensures trajectory continuity), the system constructs a spatial topology map containing the connection relationships of each sub-region. Nodes in the map represent functional zones within the monitoring area, and edges represent physical connectivity paths between regions (such as doorways and corridors). Movement trajectory data is obtained through the aforementioned skeletal key point derivation process, synchronously recording the target's location information and dwell time in each region at different time periods, providing dynamic data support for the association of nodes in the topology map. Physical connectivity attribute annotations include the connection methods of actual building structures such as doorways and corridors, while also considering obstacles and passage restrictions. The field of view coverage of fixed monitoring equipment is superimposed on the topology map in the form of polygonal regions, marking the spatial range that each camera can effectively monitor. By integrating this information, the system generates a physical topology model that includes a spatial correlation matrix and a regional reachability matrix. The spatial correlation matrix quantifies the connection strength between regions, while the regional reachability matrix calculates the probability of transition from one region to another.

[0090] The mapping process from multimodal differential feature tensors to the physical topology model employs a region binding mechanism. Each behavioral anomaly feature tensor is associated with a corresponding region node in the topology graph based on its target location information at the time of its generation. This mapping establishes a correspondence between anomaly features and spatial locations, providing a data foundation for propagation analysis. Graph convolutional networks play a core role in anomaly propagation simulation. The network structure design considers the spatial characteristics of the monitoring scenario, employing multi-layer graph convolution operations to capture anomaly propagation patterns between regions. The attenuation factor calculation for anomaly behavior is based on the physical distance and connection method between regions; anomaly signals propagating over long distances or passing through obstacle areas are appropriately attenuated. A cross-regional attention mechanism automatically learns the correlation strength of anomaly features between different regions, focusing on spatially separated but similarly patterned region pairs. A random walk method simulates the diffusion process of anomaly behavior in the topology network, generating multiple possible propagation paths, each representing a potential anomaly propagation mode.

[0091] The generation process of the risk probability distribution cloud map integrates multiple simulation results. The system statistically analyzes the frequency of occurrence of each sub-region in all random walk paths; a higher frequency indicates a greater likelihood that the area will be affected by abnormal behavior. The regional accessibility parameter adjusts the frequency statistics to consider the impact of the ease of actual movement on the propagation of anomalies. The dwell probability value calculation combines occurrence frequency and dwell time factors to reflect the continued likelihood of abnormal behavior in a specific area. The final generated risk probability distribution cloud map is presented in the form of a heatmap, using different color depths to represent the anomaly risk level of each area. This visualization method intuitively displays high-risk clusters in the monitored scenario, facilitating security personnel to quickly locate areas of interest.

[0092] Density clustering algorithms play a crucial role in identifying high-risk areas. The system performs multi-scale clustering analysis on the risk probability distribution cloud map, identifying clusters with significantly higher probability values ​​than surrounding areas. During clustering, spatial continuity is considered, merging physically adjacent areas with similar risk levels into a single high-risk zone. Target location information is used to further pinpoint the potential source of abnormal behavior, and topological connectivity is combined to infer the possible direction of anomaly propagation. The physical boundaries of abnormal behavior areas are marked using a polygonal fence approach, clearly marking areas requiring close monitoring on the electronic map of the monitoring system. Boundary delineation considers both the spatial distribution of clustering results and actual monitoring needs and management convenience. The generation of the suspicious target identification set is based on the correlation analysis between abnormal areas and actual monitored targets. The system marks target objects appearing within high-risk areas, recording their appearance characteristics, behavioral patterns, and movement trajectories. These targets constitute a preliminary suspicious set, which is further filtered through behavioral verification. The extraction process of the main path sequence of anomaly propagation analyzes all random walk simulation results and statistically analyzes the frequency of occurrence of each path. The most frequent path sequences reflect the most likely propagation trajectories of abnormal behavior within a scenario. The system represents these as an ordered list of nodes, with each node corresponding to a monitored area. This master path information helps predict the direction of abnormal behavior's spread and provides a reference for deploying prevention and control measures.

[0093] The entire anomaly propagation modeling process emphasizes the organic integration of spatial factors and behavioral characteristics. The physical topology model accurately reflects the actual layout and connectivity of the monitoring scene, providing realistic spatial constraints for propagation simulation. Graph convolutional networks and random walk methods effectively capture the diffusion characteristics of abnormal behavior in spatial networks, generating propagation paths that conform to reality. The calculation of risk probability distribution comprehensively considers multiple factors, and the generated cloud map can reliably indicate potential risk areas. The identification of suspicious targets and main paths is based on rigorous data analysis, providing valuable reference information for security decision-making. The system implementation focuses on balancing computational efficiency and result reliability, employing a distributed computing framework to accelerate the analysis process of large-scale topology networks while ensuring the accuracy of risk assessment. The output of the anomaly propagation modeling module provides spatial dimension analysis results for the entire behavior recognition and early warning system, compensating for the limitations of simple behavioral feature analysis.

[0094] Example 4: See Figure 5 The implementation process of the early warning strategy generation module is based on the real-time analysis results of the risk probability distribution cloud map. When the abnormal probability value of a specific area in the cloud map exceeds the preset risk threshold, the system automatically sends a set of monitoring instructions to the video acquisition terminal belonging to that area. The execution of these instructions includes increasing the micro-expression capture frame rate from the base frequency to 3-5 times the original frame rate, with the specific multiple dynamically adjusted according to the risk level. The high frame rate capture mode is achieved by adjusting the reading timing and signal processing parameters of the image sensor, ensuring more intensive time sampling while maintaining image quality. The simultaneously enabled real-time muscle activity intensity tracking mode focuses on monitoring specific areas of the face, which typically include areas with active micro-expressions such as the area between the eyebrows, the corners of the eyes, and the corners of the mouth. The frequency band energy mutation detection algorithm analyzes the spectral characteristics of these areas in real time, and immediately triggers the recording mechanism when a significant change in energy distribution is detected. The transient behavior detector deployed on the edge side continuously analyzes micro-expression waveform data using a sliding window approach. When an abnormal waveform pattern is identified, it automatically saves waveform segment data for that time period, which contains complete information such as waveform amplitude, frequency, and phase.

[0095] The motion trajectory perturbation test is implemented on targets in the vicinity of the risk probability distribution cloud map where the gradient change is most significant. The system first calculates the risk probability gradient field at each location in the cloud map and identifies the boundaries of regions with the largest probability change rate. Multi-dimensional motion perturbations are applied to target objects within these boundary regions. The perturbation method injects preset motion trajectory interference signals through a virtual reality interface. These interference signals simulate visual stimuli under various sudden conditions, such as suddenly appearing virtual obstacles, changes in path guidance, or abrupt changes in ambient lighting. The calculation of the theoretical behavioral response spectrum is based on the spatial correlation characteristics of the physical topology model, and its calculation formula is as follows:

[0096]

[0097] Where: R theory This represents the theoretical behavioral response spectrum, where M is the number of associated regions, and w i r is the weight coefficient of the i-th region. i Represents the spatial coordinate vector of the region, v i Φ is the regional feature vector, and Φ is the response calculation function. This function comprehensively considers the spatial correlation strength and regional feature similarity to predict the expected behavior pattern of the target under interference conditions.

[0098] The measured behavioral response spectrum is acquired by recording the actual behavioral changes of the target after the injection of interference signals. The system employs a multimodal data acquisition method, simultaneously recording multi-dimensional response data such as the target's motion trajectory, body posture, and facial expressions. These data undergo feature extraction and fusion processing to form a complete measured response spectrum, containing behavioral change features over time. In the abnormal offset determination stage, the feature distance between the theoretical response spectrum and the measured response spectrum is calculated using a distance metric method in a multi-dimensional feature space. The feature distance calculation considers the overall similarity of behavioral patterns and the differences in local features, obtaining a comprehensive offset by weighted combination of multiple distance indices. The calculation of the degree of abnormal target offset compares the comprehensive offset with a preset threshold. When the offset continuously exceeds the threshold range, the system marks the target as an abnormal behavior-related target. The marking process employs a progressive confirmation mechanism, requiring consistent judgments across multiple consecutive time segments to ultimately confirm the abnormal state.

[0099] The entire early warning strategy generation process emphasizes a combination of real-time performance and adaptability. Adjustments to monitoring parameters are based on dynamic changes in risk distribution, ensuring optimal allocation of monitoring resources. Upgrades to the micro-expression monitoring mode not only increase sampling frequency but also enhance the ability to capture subtle muscle movements. Motion trajectory perturbation testing reveals potential abnormal behaviors through active stimulation, providing a proactive detection method. The calculation of the theoretical behavioral response spectrum fully utilizes the spatial information of the physical topology model, making the prediction results more consistent with the characteristics of real-world scenarios. Multimodal acquisition of measured responses ensures the comprehensiveness and accuracy of behavioral assessment. The anomaly judgment mechanism employs a multi-indicator comprehensive judgment to reduce the possibility of false alarms. The system implementation emphasizes the collaborative work between components, ensuring rapid transition from risk detection to early warning response. All processing processes consider real-time requirements, employing a streaming processing framework to ensure the analysis and decision-making cycle is completed in the shortest possible time. The generation of early warning strategies is based not only on the current risk state but also on historical behavioral patterns and environmental context information, making decision-making more comprehensive and reliable.

[0100] Example 5: Implementation of the Optical Flow Feature Compensator employs a motion trajectory reconstruction algorithm to perform in-depth analysis of the real-time captured motion optical flow vector field. This algorithm first decomposes the continuous optical flow vector sequence into different motion pattern components, distinguishing between fundamental motion components and anomalous fluctuation components through frequency domain analysis and pattern recognition techniques. Fundamental motion components represent the target object's regular behavior patterns, which typically exhibit certain regularity and predictability, reflecting normal movement characteristics and behavioral habits. Anomalous fluctuation components capture motion characteristics that deviate from the regular patterns, potentially manifesting as sudden speed changes, unusual motion directions, or abnormal trajectory shapes. The motion pattern decomposition process employs a multi-scale analysis method, considering both short-term motion details and long-term motion trends to ensure effective identification of various types of anomalous fluctuations.

[0101] The energy proportion calculation focuses on the relative importance of anomalous fluctuation components within the overall motion pattern. The system quantifies the significance of anomalous motion by calculating the ratio of the amplitude energy of the anomalous component to the total motion energy. This ratio reflects the contribution of anomalous fluctuations to the overall motion pattern; a higher ratio indicates more pronounced and prominent anomalous motion characteristics. Based on this energy proportion value, the system dynamically adjusts the prediction weight coefficients of the temporal modeling unit. When the energy proportion of the anomalous fluctuation component is high, the weight of the temporal prediction model is correspondingly reduced, minimizing its impact on the final behavior prediction. This dynamic adjustment mechanism ensures that the behavior recognition model can adapt to various motion conditions, promptly adjusting the prediction strategy when anomalous motion occurs, thus improving the system's robustness and adaptability.

[0102] The prediction weight coefficient adjustment process of the temporal modeling unit adopts a smooth transition mechanism to avoid fluctuations in prediction results caused by sudden changes in weights. The system gradually adjusts the weight coefficients according to the changing trend of energy proportions, while also considering historical weight changes to ensure the stability and continuity of the adjustment process. The adjustment range of the weight coefficients is standardized and kept within a reasonable numerical range, ensuring timely response to anomalies while avoiding prediction instability caused by over-adjustment. This refined weight management mechanism enables the behavior recognition model to effectively cope with various abnormal motion situations while maintaining prediction accuracy.

[0103] The motion trajectory reconstruction algorithm involves multiple processing stages. The initial stage preprocesses the optical flow vector field, including noise filtering, vector smoothing, and missing data compensation, to improve the quality of the motion data. Motion feature extraction is then performed, extracting various feature parameters such as velocity, acceleration, and direction of motion from the optical flow field. The pattern decomposition stage employs clustering and principal component analysis to decompose complex motion patterns into several basic motion components. The anomaly detection algorithm, based on statistical learning principles, identifies motion components that significantly differ from conventional patterns and marks them as anomalous fluctuation components. The entire process utilizes incremental learning, continuously optimizing the accuracy of motion pattern decomposition as data accumulates.

[0104] The collaborative work between the optical flow feature compensator and other modules of the system is reflected in the integration of data exchange and processing flows. The compensator receives real-time optical flow data from the video analytics engine, processes it, and outputs motion pattern analysis results and weight adjustment suggestions. This output data is passed to the temporal modeling unit and other relevant components, influencing the behavior prediction and analysis process. Simultaneously, the compensator also obtains environmental context information and historical behavior data from other modules; this information is used to optimize the accuracy of motion pattern decomposition and the precision of anomaly detection. This bidirectional data exchange mechanism ensures coordinated operation among all components of the system, improving overall behavior recognition performance.

[0105] During implementation, special attention was paid to balancing computational efficiency and real-time requirements. The motion trajectory reconstruction algorithm employed an optimized computational strategy to minimize computational complexity while ensuring analytical accuracy. A parallel processing architecture was used to accelerate the analysis of large-scale optical flow data, and a distributed computing framework ensured that the system could handle the massive amounts of optical flow data generated by high-resolution video streams. Memory management mechanisms optimized data storage and access patterns, reducing latency during data processing. These technical measures ensured that the optical flow feature compensator could operate stably in real-time video stream processing scenarios, meeting the system's requirements for response speed and processing efficiency.

[0106] The performance optimization of the optical flow feature compensator is ongoing, with algorithm parameters continuously adjusted by monitoring its operational status and analyzing processing results. The system records performance metrics such as the accuracy of motion pattern decomposition and the recall rate of anomaly detection, using this data to optimize algorithm performance. An adaptive learning mechanism enables the compensator to adapt to different monitoring scenarios and target types, maintaining stable performance under various environmental conditions. The recognition threshold for abnormal fluctuation components is dynamically adjusted according to actual conditions, avoiding false positives or false negatives caused by environmental changes. This continuous optimization mechanism ensures that the optical flow feature compensator can operate stably and reliably over the long term, providing accurate motion analysis results for the behavior recognition system.

[0107] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0108] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A real-time video stream behavior recognition and early warning system, characterized in that, include: Spatiotemporal feature modeling module: Based on the historical behavior video data of the target object, a dynamic behavior recognition model is constructed. The model captures the three-dimensional coordinate sequence of the skeletal key points of the target, the motion optical flow vector field and the intensity spectrum of micro-expression muscle activity in the current video stream in real time. The theoretical behavior pattern vector is output through the dynamic behavior recognition model. Behavior segment extraction module: Perform cross-modal difference analysis on the theoretical behavior pattern vector and the measured behavior vector of the video analysis engine. The cross-modal difference analysis includes temporal continuity deviation, spatial trajectory offset and micro-expression spectrum similarity, and generates a multimodal difference feature tensor at the target object level. Anomaly propagation modeling module: Input the multimodal difference feature tensor into the scene topology inference network, and combine the physical space constraints of the monitoring area with the target object's movement trajectory information to generate a risk probability distribution cloud map of the abnormal behavior propagation path; Risk area location module: Identifies high-risk cluster areas based on the risk probability distribution cloud map, and marks the physical boundaries of abnormal behavior areas where the probability value exceeds a preset risk probability threshold; Early warning strategy generation module: Configures monitoring parameters based on the risk probability distribution cloud map, including: High frame rate micro-expression capture mode is enabled for high-risk gathering areas; Test the motion trajectory perturbation of targets in the vicinity.

2. The real-time video stream behavior recognition and early warning system according to claim 1, characterized in that, The spatiotemporal feature modeling module specifically includes: Historical behavior feature mining: Multi-scale spatiotemporal decomposition processing is performed on the historical behavior video data of the target object, including: using an adaptive spatiotemporal pyramid to extract the energy distribution of steady-state and transient motion of the skeletal key point sequence, establishing the correlation mapping between motion vector field and behavioral intention through optical flow coupling analysis, and using sequence alignment algorithm to calibrate the intensity pattern of micro-expression activity in different scenarios; Dynamic behavior recognition model construction: The processed historical behavior data is input into the fusion prediction architecture, which includes: Temporal modeling units based on behavior evolution curves are used to generate basic behavior prediction vectors; A 3D convolutional network with embedded spatial attention mechanism corrects prediction bias caused by changes in viewpoint; Based on the optical flow feature compensator, the prediction weights are dynamically adjusted according to the real-time captured motion optical flow vector field. The edge computing unit synchronously collects the position drift rate and velocity fluctuation parameters of key skeletal points, the amplitude gradient distribution and directional consistency coefficient of motion optical flow vectors, and the frequency band energy entropy value of micro-expression intensity spectrum. Theoretical value generation: Input the real-time captured data into the dynamic behavior recognition model to obtain the theoretical behavior pattern vector.

3. The real-time video stream behavior recognition and early warning system according to claim 2, characterized in that, The behavior fragment extraction module specifically includes: Time continuity deviation calculation: The theoretical behavior pattern vector and the measured behavior vector are compared by sliding comparison with a preset time segment. The non-uniformly sampled behavior sequence is aligned by the sequence dynamic warping algorithm. The cumulative trajectory deviation in each segment is calculated to generate a time domain deviation feature matrix. Spatial trajectory offset detection: Spatial domain transformation decomposition is performed on the three-dimensional coordinates of skeletal key points of theoretical and measured values, the trajectory curvature similarity ratio is calculated, the spatial offset index of each joint node is extracted, and a spatial offset feature matrix is ​​constructed. Micro-expression spectral similarity assessment: Based on the pattern matching algorithm, the distance distribution between theoretical values ​​and measured micro-expression intensity spectra is measured, the phase synchronization error of activation intensity of specific muscle groups is calculated, the divergence difference of spectral energy distribution is quantified, and micro-expression similarity vectors are generated. Multimodal difference feature tensor generation: The temporal deviation feature matrix, spatial offset feature matrix, and micro-expression similarity vector are fused with high-order features. Through modal importance weighted standardization, scale differences are eliminated, and a third-order multimodal difference feature tensor with dimensions of [target identifier x time segment x modal type] is output.

4. The real-time video stream behavior recognition and early warning system according to claim 3, characterized in that, The anomaly propagation modeling module specifically includes: Scene topology modeling: Construct a spatial topology map of the monitoring area based on the target object's movement trajectory information, mark the physical connectivity attributes between each sub-area, overlay the field of view coverage constraints of the fixed monitoring equipment on the topology map, and generate a physical topology model containing a spatial association matrix and an area reachability matrix; Anomaly propagation deduction: Mapping the multimodal differential feature tensor to the corresponding region of the physical topology model; Anomaly propagation simulation is performed based on graph convolutional networks. The computation of the anomaly propagation simulation includes: Calculate the attenuation factor of anomalous behavior based on the region's physical connectivity attributes; Capture anomalous correlation features through a cross-regional attention mechanism; A random walk method is used to simulate the propagation path of anomalous behavior in a topological network. Risk distribution generation: Statistically analyze the frequency of abnormal behavior in each sub-region during simulated propagation, combine regional accessibility parameters to calculate the probability value of abnormal behavior residence, and generate a risk probability distribution cloud map covering the entire region; Physical boundary delineation: Perform density clustering analysis on the risk probability distribution cloud map to identify high-risk probability clustering areas; Based on the location of the target object and its physical topology connection, mark the physical boundaries of the abnormal behavior area.

5. The real-time video stream behavior recognition and early warning system according to claim 4, characterized in that, The early warning strategy generation module specifically includes: When the anomaly probability value of a certain area in the risk probability distribution cloud map exceeds a preset risk probability threshold, a monitoring command is sent to the video acquisition terminal belonging to that area, and the following is executed: Increase the frame rate for micro-expression capture to 3-5 times the original frame rate; Simultaneously enable real-time muscle activity intensity tracking mode to capture frequency band energy mutations in specific facial regions; Deploy transient behavior detectors at the edge to record abnormal micro-expression waveform fragments; Motion trajectory perturbation test execution: Apply multi-dimensional motion perturbation to the target in the neighborhood area with the largest change in risk probability gradient in the risk probability distribution cloud map.

6. The real-time video stream behavior recognition and early warning system according to claim 5, characterized in that, The multidimensional motion perturbation includes: Inject a preset motion trajectory interference signal through a virtual reality interface; Theoretical behavioral response calculation: Based on the physical topology model and spatial correlation matrix, calculate the theoretical behavioral response spectrum; Acquisition of measured behavioral response: Record the behavioral patterns of each target after the injection of interference signal to obtain the measured behavioral response spectrum; Anomaly offset determination: Calculate the characteristic distance between the theoretical behavioral response spectrum and the measured behavioral response spectrum; By comparing the characteristic distance between the measured behavioral response spectrum and the theoretical behavioral response spectrum, the degree of deviation of abnormal targets is calculated. When the degree of deviation of abnormal targets exceeds the preset deviation threshold, they are marked as abnormal behavior associated targets.

7. The real-time video stream behavior recognition and early warning system according to claim 3, characterized in that, The spatiotemporal feature modeling module also includes: Adaptive normalization processing based on ambient lighting conditions eliminates characteristic noise caused by changes in light intensity; The correlation characteristics of skeleton, optical flow, and micro-expression are integrated through a multi-source feature fusion algorithm; The output contains a theoretical behavior pattern vector that includes the normal behavior fluctuation range. The theoretical behavior pattern vector is dynamically updated as the behavior pattern evolves.

8. The real-time video stream behavior recognition and early warning system according to claim 4, characterized in that, The anomaly propagation modeling module also includes an output containing a set of suspicious target identifiers and an anomaly propagation main path sequence, wherein: The suspicious target identifier set is based on the monitoring area covered by the physical boundary of the abnormal behavior area, and the monitoring area is bound to the actual target object to form the suspicious target identifier set; The anomaly propagation main path sequence is obtained by counting the occurrence frequency of each path during the anomaly diffusion process simulated by random walk, selecting the path sequence with the highest cumulative occurrence frequency, and outputting an ordered list of nodes reflecting the main propagation trajectory of the abnormal behavior in the scene.

9. The real-time video stream behavior recognition and early warning system according to claim 2, characterized in that, In the construction of the dynamic behavior recognition model: The optical flow feature compensator uses a motion trajectory reconstruction algorithm to decompose the motion pattern of the real-time captured motion optical flow vector field and identify the basic motion component and abnormal fluctuation component. The prediction weight coefficients of the time series modeling unit are dynamically adjusted based on the energy proportion of the abnormal fluctuation components.

10. The real-time video stream behavior recognition and early warning system according to claim 3, characterized in that, In the behavior fragment extraction module: The micro-expression spectrum similarity assessment uses a key frequency band energy comparison algorithm to extract the energy distribution features of preset emotion-related frequency bands; The energy density of theoretical and measured data in the emotionally relevant frequency bands is quantitatively compared, and a micro-expression similarity vector is constructed based on the degree of relative difference.

Citation Information

Patent Citations

  • Micro-expression recognition method and system based on optical flow and RGB modal contrast learning

    CN113139479A

  • Behavior recognition method for cross-modal three-dimensional point cloud sequence spatial-temporal feature network

    CN114973418A

  • Identification method and system based on multi-modal data fusion

    CN116934926A

  • Target generation method and device for radar system software testing and electronic equipment

    CN117407319A

  • Intelligent human shape trajectory prediction and alarm system and method based on multi-modal video analysis

    CN120047897A

Cited By

  • Space-time sequence data processing method and system for pet abnormal behavior recognition

    CN121071757A

  • Personnel abnormal behavior supervision method and system for important places

    CN121353990A

  • A method and system for monitoring abnormal behavior of personnel in important places

    CN121353990B

  • AI-based motion posture recognition system

    CN121564446A

  • Multi-dimensional evaluation method and system for thirst degree of hemodialysis patient

    CN121726035A