Behavior recognition and warning system for real-time video streams
By using spatiotemporal feature modeling and cross-modal difference analysis, combined with scene topology inference network, a risk probability distribution cloud map of abnormal behavior propagation path is generated. This solves the problems of behavior recognition bias and resource waste in existing video surveillance systems in complex scenarios, and achieves accurate anomaly early warning and efficient resource utilization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GAOZI TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2025-08-29
- Publication Date
- 2026-04-10
AI Technical Summary
Existing video surveillance systems struggle to achieve accurate behavior recognition and anomaly warning in complex scenarios. In particular, feature extraction is incomplete when there is occlusion, changes in lighting, or rapid target movement, making it impossible to effectively predict the spread path and impact range of abnormal behavior. Furthermore, the monitoring parameter configuration cannot be dynamically adjusted, resulting in wasted computing resources or missed key behavioral features.
By employing a spatiotemporal feature modeling module that combines skeletal key points, 3D coordinates, motion optical flow vector fields, and micro-expression muscle activity intensity spectra, a theoretical behavior pattern vector is generated through a dynamic behavior recognition model. Combined with cross-modal difference analysis and scene topology inference network, a risk probability distribution cloud map of abnormal behavior propagation paths is generated, and monitoring parameters are dynamically configured to achieve efficient early warning.
It achieves accurate behavior recognition in complex scenarios, reduces recognition bias, can promptly detect potential abnormal behaviors, clearly present abnormal propagation paths, optimize the utilization of computing resources, and improve management efficiency.
Smart Images

Figure CN120894831B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of video behavior recognition, in particular to a behavior recognition and early warning system for real-time video stream. BACKGROUND
[0002] With the wide application of video monitoring technology in public safety, traffic management, industrial production and other fields, the demand for accurate identification and abnormal early warning of target object behavior in real-time video stream is increasingly urgent. The current mainstream behavior recognition system relies on single modal data for analysis, such as determining human motion only through skeletal key point coordinates or identifying object moving trajectory only according to motion optical flow vector field. This single modal analysis method has obvious limitations. When there are occlusions, changes in lighting or rapid movement of targets in the monitoring scene, incomplete feature extraction is easy to occur, which in turn leads to behavior recognition deviation.
[0003] Existing systems lack effective propagation modeling mechanisms in abnormal behavior processing. Most systems can only issue an early warning after detecting a single abnormal behavior, and cannot analyze the possible diffusion path and impact range of the abnormal behavior in combination with the physical space constraints (such as channel width, obstacle distribution) of the monitoring area and the moving trajectory of the target object. For example, in a crowded place, when an abnormal running behavior of a person is detected, the existing system is difficult to predict whether this behavior will trigger a chain reaction of surrounding people, and also difficult to accurately label high-risk gathering areas, resulting in lack of pertinence of the early warning information, and making it difficult to assist management personnel to take timely and effective intervention measures.
[0004] The monitoring parameter configuration of traditional systems mostly adopts a fixed mode, which cannot be dynamically adjusted according to the scene risk level. For monitoring of different areas, the same frame rate and feature extraction frequency are maintained, which not only wastes computing resources in low-risk areas, but also may miss key behavior features due to insufficient sampling frequency in high-risk areas. For example, in the station waiting area and the entrance and exit channel, which are two areas of different risk levels, the fixed monitoring parameters cannot meet the demand for rapid abnormal response in the channel area, and also consume unnecessary computing resources in the waiting area, affecting the overall operation efficiency of the system. SUMMARY
[0005] The present application aims to provide a behavior recognition and early warning system for real-time video stream to solve the problems raised in the background.
[0006] To achieve the above-mentioned purpose, the present application provides a behavior recognition and early warning system for real-time video stream, which comprises:
[0007] a spatio-temporal feature modeling module: based on the target object historical behavior video data to build a dynamic behavior recognition model, real-time capture of the current video stream of the target object of the three-dimensional coordinate sequence of the skeleton key point, the motion optical flow vector field and the micro-expression muscle activity intensity spectrum, output the theoretical behavior mode vector through the dynamic behavior recognition model;
[0008] a behavior segment extraction module: cross-modal difference analysis of the theoretical behavior mode vector and the video analysis engine measured behavior vector, the cross-modal difference analysis includes time continuity deviation, spatial trajectory offset degree and micro-expression frequency spectrum similarity, generating a multi-modal difference feature tensor of the target object level;
[0009] an abnormal propagation modeling module: input the multi-modal difference feature tensor into the scene topology reasoning network, combine the monitoring area physical space constraint condition and the target object moving track information, generate the risk probability distribution cloud chart of the abnormal behavior propagation path;
[0010] a risk area positioning module: according to the risk probability distribution cloud chart, identify the high-risk aggregation area, and label the abnormal behavior physical area boundary with a probability value exceeding a preset risk probability threshold;
[0011] a warning strategy generation module: based on the risk probability distribution cloud chart, configure the monitoring parameters, including:
[0012] enable high frame rate micro-expression capture mode for high-risk aggregation area;
[0013] apply motion trajectory disturbance test to adjacent area target.
[0014] Preferably, the spatio-temporal feature modeling module specifically includes:
[0015] historical behavior feature mining: multi-scale spatio-temporal decomposition processing of target object historical behavior video data, including: using adaptive spatio-temporal pyramid to extract energy distribution of steady-state motion and transient motion of skeleton key point sequence, establishing the correlation mapping of motion vector field and behavior intention through optical flow coupling analysis, and using sequence alignment algorithm to calibrate micro-expression activity intensity mode in different scenes;
[0016] dynamic behavior recognition model construction: input the processed historical behavior data into the fusion prediction architecture, the fusion prediction architecture includes:
[0017] a time series modeling unit based on behavior evolution curve, used to generate a basic behavior prediction vector;
[0018] a three-dimensional convolution network embedded with spatial attention mechanism, which corrects the prediction deviation caused by the change of viewing angle;
[0019] a light flow feature compensator, which dynamically adjusts the prediction weight according to the real-time captured motion optical flow vector field;
[0020] synchronously collecting, by the edge computing unit, a position drift rate and a speed fluctuation parameter of the skeletal key points, an amplitude gradient distribution and a direction consistency coefficient of a motion optical flow vector, and a frequency band energy entropy value of a micro-expression intensity spectrum;
[0021] Theoretical value generation: input real-time captured data into the dynamic behavior recognition model to obtain a theoretical behavior pattern vector.
[0022] Preferably, the behavior segment extraction module specifically comprises:
[0023] Time continuity deviation calculation: slidingly compare the theoretical behavior pattern vector and the measured behavior vector in a preset time segment, align the non-uniformly sampled behavior sequence using a sequence dynamic regularization algorithm, calculate the cumulative trajectory deviation amount in each segment, and generate a time domain deviation feature matrix;
[0024] Spatial trajectory deviation detection: spatial domain transformation and decomposition are performed on the three-dimensional coordinates of the skeletal key points of the theoretical value and the measured value, the trajectory curvature similarity ratio is calculated, the spatial deviation index of each joint node is extracted, and a spatial deviation feature matrix is constructed;
[0025] Micro-expression frequency spectrum similarity evaluation: based on a pattern matching algorithm, the distance distribution of the theoretical value and the measured micro-expression intensity spectrum is measured, the phase synchronization error of the activation intensity of a specific muscle group is calculated, the divergence difference of the frequency spectrum energy distribution is quantified, and a micro-expression similarity vector is generated.
[0026] Multi-modal difference feature tensor generation: high-order feature fusion is performed on the time domain deviation feature matrix, the spatial deviation feature matrix, and the micro-expression similarity vector, and through standardized processing of modal importance weighting, scale differences are eliminated, and a three-order multi-modal difference feature tensor with dimensions [target identifier x time segment x modal type] is output.
[0027] Preferably, the abnormality propagation modeling module specifically comprises:
[0028] Scene topology modeling: constructing a spatial topology graph of the monitoring area according to the target object movement trajectory information, labeling the physical connectivity attributes between each sub-region, superimposing the field of view coverage constraint conditions of fixed monitoring equipment in the topology graph, and generating a physical topology model containing a spatial correlation matrix and a region reachability matrix;
[0029] Abnormality propagation deduction: mapping the multi-modal difference feature tensor to the corresponding region of the physical topology model;
[0030] Performing abnormality propagation simulation based on a graph convolution network, and the calculation of the abnormality propagation simulation includes:
[0031] Calculating the attenuation factor of the abnormal behavior according to the region physical connectivity attribute;
[0032] Capture abnormal correlation features through cross-region attention mechanism;
[0033] Simulate the diffusion path of abnormal behavior in the topological network by random walk method;
[0034] Risk distribution generation: count the frequency of abnormal behavior in the simulation propagation of each sub-region, calculate the abnormal behavior residence probability value combined with the region accessibility parameter, and generate a risk probability distribution cloud map covering the whole region;
[0035] Physical boundary demarcation: perform density clustering analysis on the risk probability distribution cloud map to identify high-risk probability aggregation area;
[0036] According to the position of the target object and the connection relationship of the physical topology, the physical region boundary of the abnormal behavior is marked.
[0037] Preferably, the early warning strategy generation module specifically comprises:
[0038] When the abnormal probability value of a certain region in the risk probability distribution cloud map exceeds the preset risk probability threshold, a monitoring instruction is issued to the video acquisition terminal to which the region belongs, and the following is executed:
[0039] The micro-expression capture frame rate is increased to 3-5 times the original frame rate;
[0040] Synchronously enable real-time tracking mode of muscle activity intensity to capture energy mutation of specific facial region frequency band;
[0041] Deploy a transient behavior detector on the edge side to record abnormal micro-expression waveform fragments;
[0042] Motion trajectory disturbance test execution: apply multi-dimensional motion disturbance to the target in the adjacent region with the largest risk probability gradient change in the risk probability distribution cloud map.
[0043] Preferably, the multi-dimensional motion disturbance comprises:
[0044] Inject a preset motion trajectory interference signal through a virtual reality interface;
[0045] Theoretical behavior response calculation: calculate the theoretical behavior response spectrum based on the physical topology model and the spatial correlation matrix;
[0046] Measured behavior response acquisition: record each target behavior mode after the interference signal injection to obtain the measured behavior response spectrum;
[0047] Abnormal offset determination: calculate the feature distance between the theoretical behavior response spectrum and the measured behavior response spectrum;
[0048] The characteristic distance of the measured behavior response spectrum and the theoretical behavior response spectrum is compared, the abnormal target deviation degree is calculated, and when the abnormal target deviation degree exceeds the preset deviation threshold, the abnormal behavior correlation target is marked.
[0049] Preferably, the space-time feature modeling module further comprises:
[0050] Adaptive normalization processing based on environmental light conditions eliminates feature noise caused by light changes;
[0051] The correlation characteristics of bones, optical flow and micro-expression are integrated through a multi-source feature fusion algorithm.
[0052] The theoretical behavior mode vector containing the normal behavior fluctuation range is output, and the theoretical behavior mode vector is dynamically updated with the evolution of the behavior mode.
[0053] Preferably, the abnormal propagation modeling module further comprises an output containing a suspicious target identification set and an abnormal propagation main path sequence, wherein:
[0054] The suspicious target identification set is based on the monitoring area covered by the abnormal behavior physical area boundary, binds the monitoring area with the actual target object, and constitutes the suspicious target identification set;
[0055] The abnormal propagation main path sequence selects the path sequence with the highest cumulative appearance frequency by counting the appearance frequency of each path in the random walk simulation abnormal diffusion process, and outputs an ordered node list reflecting the main propagation trajectory of the abnormal behavior in the scene.
[0056] Preferably, in the dynamic behavior recognition model construction:
[0057] The optical flow feature compensator adopts a motion trajectory reconstruction algorithm to perform motion pattern decomposition on the real-time captured motion optical flow vector field, and identifies the basic motion component and the abnormal fluctuation component.
[0058] The prediction weight coefficient of the time series modeling unit is dynamically adjusted based on the energy proportion of the abnormal fluctuation component.
[0059] Preferably, in the behavior segment extraction module:
[0060] The micro-expression frequency spectrum similarity evaluation adopts a key frequency band energy comparison algorithm to extract the energy distribution characteristics of the preset emotion-related frequency band.
[0061] The energy density of the theoretical and measured data in the emotion-related frequency band is quantitatively compared, and a micro-expression similarity vector is constructed based on the relative difference degree.
[0062] Compared with the prior art, the present application has the following advantages:
[0063] Through the collaborative work of multiple modules, accurate identification and efficient early warning of target object behavior are achieved. At the feature extraction level, the spatio-temporal feature modeling module simultaneously captures three types of multi-modal data: 3D coordinate sequences of skeletal key points, motion optical flow vector fields, and micro-expression muscle activity intensity spectra. Based on historical behavior video data, a dynamic behavior recognition model is constructed to generate a theoretical behavior pattern vector. This multi-modal data fusion approach can compensate for the shortcomings of single modal data in complex scenarios. Even in the presence of occlusions, lighting changes, or target rapid movement in the monitoring scene, different modal data can complement each other to fully extract target behavior features, reduce bias in the behavior recognition process, and improve the recognition accuracy of various behavior patterns.
[0064] The behavior segment extraction module compares the theoretical behavior pattern vector with the measured behavior vector through cross-modal difference analysis, generating a multi-modal difference feature tensor from three dimensions: temporal continuity deviation, spatial trajectory offset, and micro-expression frequency spectrum similarity. This multi-dimensional difference analysis approach can more comprehensively capture the subtle differences between theoretical behavior and actual behavior, avoiding false negatives or false positives due to single-dimensional analysis, and ensuring that the system can timely detect potential signs of abnormal behavior, providing detailed feature basis for subsequent abnormal behavior propagation modeling.
[0065] The abnormal behavior propagation modeling module inputs the multi-modal difference feature tensor into the scene topology reasoning network, and combines the physical space constraints and target object movement trajectory information in the monitoring area to generate a risk probability distribution cloud map of the abnormal behavior propagation path. This module breaks through the limitations of traditional systems that can only detect a single abnormal behavior. By constructing the scene topology relationship, it simulates the diffusion law of abnormal behavior in different physical spaces, clearly presenting the possible propagation path and impact range of abnormal behavior. Management personnel can intuitively understand the diffusion trend of abnormal behavior through the risk probability distribution cloud map, and predict possible chain reactions in advance to provide a clear direction for developing intervention strategies.
[0066] The risk area positioning module identifies high-risk aggregation areas based on the risk probability distribution cloud map and labels the physical area boundaries of abnormal behavior with probability values exceeding the preset threshold. This function enables the system to accurately lock onto areas that require close attention, avoiding the need for management personnel to manually search through a large number of monitoring screens, reducing labor costs, and ensuring that intervention measures can accurately target high-risk areas, improving management efficiency.
[0067] The early warning strategy generation module dynamically configures monitoring parameters based on the risk probability distribution cloud map, enables a high frame rate micro-expression capture mode for a high-risk aggregation area, can capture micro-expression changes of target objects in the area in more detail, and discovers hidden abnormal emotions and behavior tendencies in time; motion trajectory disturbance tests are applied to adjacent areas, which can further verify whether the target in the area has a potential abnormal behavior trend, and risks can be checked in advance. This dynamic parameter configuration method can reasonably allocate computing resources according to the risk levels of different areas, ensure monitoring accuracy in high-risk areas, reduce unnecessary computing consumption in low-risk areas, optimize the use of system resources, and improve overall operation efficiency. BRIEF DESCRIPTION OF DRAWINGS
[0068] Figure 1 A timing diagram of the real-time video stream behavior recognition and early warning system according to the present application;
[0069] Figure 2 A flowchart of the spatiotemporal feature modeling module;
[0070] Figure 3 A flowchart of the behavior segment extraction module;
[0071] Figure 4 A flowchart of the abnormal propagation modeling module;
[0072] Figure 5 A flowchart of the multi-dimensional motion disturbance execution. DETAILED DESCRIPTION
[0073] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0074] Please refer to Figure 1 The present application provides a real-time video stream behavior recognition and early warning system, which comprises:
[0075] The system realizes real-time recognition and early warning of target object behavior in video stream through the cooperative work of multiple modules. The spatio-temporal feature modeling module constructs a dynamic behavior recognition model based on historical behavior video data of the target object, captures the three-dimensional coordinate sequence of the target's skeletal key points, the motion optical flow vector field, and the micro-expression muscle activity intensity spectrum in real time in the current video stream, and outputs a theoretical behavior mode vector through the model. The behavior segment extraction module performs cross-modal difference analysis on the theoretical behavior mode vector and the measured behavior vector of the video analysis engine, including time continuity deviation, spatial trajectory offset degree, and micro-expression frequency spectrum similarity, to generate a multi-modal difference feature tensor at the target object level. The abnormal propagation modeling module inputs the multi-modal difference feature tensor into a scene topology reasoning network, combines the physical space constraint conditions and the target object movement trajectory information of the monitoring area, and generates a risk probability distribution cloud map of the abnormal behavior propagation path. The risk area positioning module identifies high-risk aggregation areas according to the risk probability distribution cloud map and labels the abnormal behavior physical area boundary with a probability value exceeding a preset risk probability threshold. The early warning strategy generation module configures monitoring parameters based on the risk probability distribution cloud map, including enabling a high-frame-rate micro-expression capture mode for high-risk aggregation areas and applying a movement trajectory disturbance test to adjacent area targets. The system realizes the whole-chain processing from behavior feature extraction to risk early warning through the above process.
[0076] Embodiment 1: refer to Figure 2 The implementation process of the spatio-temporal feature modeling module involves multi-scale spatio-temporal decomposition processing of historical behavior video data of the target object. The adaptive spatio-temporal pyramid is used to extract the energy distribution of the steady-state motion and transient motion of the skeletal key point sequence. This processing decomposes the video sequence into multiple spatio-temporal scales, and calculates the distribution features of the motion energy at each scale. The steady-state motion component reflects the persistent behavior pattern, and the transient motion component captures the burst action features. Through this decomposition, the spatio-temporal characteristics of the behavior can be described more accurately. The optical flow coupling analysis establishes the association mapping between the motion vector field and the behavior intention, and infers the possibility of behavior intention by analyzing the direction and amplitude changes of the optical flow vector. This mapping is based on the statistical association relationship between the optical flow field and the historical behavior mode. The sequence alignment algorithm calibrates the micro-expression activity intensity mode in different scenes. This algorithm uses dynamic time warping or similar methods to adjust the time alignment and intensity normalization of micro-expression data obtained under different shooting conditions, ensuring the consistency of the data.
[0077] The construction of the dynamic behavior recognition model inputs the processed historical behavior data into the fusion prediction architecture, and generates a basic behavior prediction vector based on the time sequence modeling unit of the behavior evolution curve. The unit adopts a recurrent neural network or a time convolution network to learn the time evolution law of the behavior sequence and outputs a vector representing the expected behavior pattern. The three-dimensional convolution network with embedded space attention mechanism corrects the prediction deviation caused by the change of the perspective, and the network introduces attention weights in the three-dimensional convolution operation, focusing on the key spatial area and reducing the feature distortion caused by the difference in perspective. The optical flow feature compensator dynamically adjusts the prediction weight according to the real-time captured motion optical flow vector field. The component analyzes the feature changes of the optical flow field and adjusts the contribution of each part of the model accordingly. Through the edge computing unit, multiple real-time data parameters are synchronously collected. The position drift rate of the skeletal key points is obtained by calculating the change rate of the key point coordinates between consecutive frames, and the formula is as follows:
[0078]
[0079] where D t represents the average position drift rate at time t, N is the number of key points, p i (t) is the three-dimensional coordinates of the i-th key point at time t, and ||·|| represents the Euclidean distance. The velocity fluctuation parameter is calculated based on the second-order difference of the position sequence, reflecting the change of the motion acceleration. The amplitude gradient distribution of the motion optical flow vector is obtained by calculating the spatial derivative of the optical flow amplitude, and the direction consistency coefficient is calculated by counting the variance of the optical flow direction in the local area. The frequency band energy entropy value of the micro-expression intensity spectrum is calculated by dividing the micro-expression signal spectrum and calculating the energy distribution entropy value of each sub-band, and the formula is as follows:
[0080]
[0081] where H f represents the frequency band energy entropy value, K is the number of frequency bands, and E k is the normalized energy value of the k-th frequency band. The entropy value quantifies the complexity of the micro-expression spectrum energy distribution. In the theoretical value generation stage, the above real-time captured data is input into the dynamic behavior recognition model to obtain the theoretical behavior pattern vector. The model processing process includes feature encoding, time sequence modeling and feature decoding, and finally outputs a vector representing the expected behavior pattern. The vector contains multi-modal feature information of the behavior, which is used for difference analysis.
[0082] Adaptive normalization based on ambient lighting conditions is applied to the feature extraction process to eliminate feature noise caused by light changes. This process reduces the impact of lighting changes on feature stability by estimating scene lighting conditions, normalizing intensity and enhancing contrast of visual features. Multi-source feature fusion algorithm integrates the correlation characteristics of skeleton, optical flow and micro-expression, uses weighted fusion or learning-based fusion strategy, combines the feature advantages of each modality, and improves the robustness and discriminability of behavior representation. The output theoretical behavior pattern vector contains the normal behavior fluctuation interval, which is determined by the statistical feature distribution range of historical data, and is used to define the boundary of normal behavior. The theoretical behavior pattern vector is updated dynamically with the evolution of behavior pattern, and the update mechanism is based on online learning or sliding window statistical method, so that the model can adapt to the gradual changes of behavior pattern. The whole implementation process emphasizes the fine processing and analysis of multi-modal features, and realizes the comprehensive modeling of behavior features through the combination of various algorithms and models. The components cooperate with each other through data flow and parameter transmission, forming a complete behavior feature extraction and modeling pipeline. The processing process pays attention to the balance between calculation efficiency and accuracy, uses edge computing to reduce data transmission delay, and ensures that the system can meet the real-time processing requirements. The design of feature representation considers the spatiotemporal dynamic characteristics of behavior, which can capture the changes of short-term action and long-term behavior pattern. The adaptive ability of the model is realized through the dynamic update mechanism, so that the system can respond to changes in environmental conditions and behavior patterns, and maintain the stability of recognition performance.
[0083] Example 2: refer to Figure 3 The implementation process of the behavior segment extraction module involves multi-dimensional comparison and analysis of the theoretical behavior pattern vector and the measured behavior vector. The time continuity deviation calculation uses a sliding window mechanism to divide the continuous video stream into fixed-length time segments, and aligns the theoretical prediction value and the actual observation value in each segment. The sequence dynamic adjustment algorithm processes the time misalignment problem caused by sampling rate difference or action execution speed difference, and adjusts the corresponding relationship on the time axis dynamically to achieve the best matching state of the two sequences in the time dimension. The cumulative trajectory deviation calculation considers the difference between feature vectors at each aligned time point, and integrates the overall deviation degree in the whole time segment through integration to form a feature matrix reflecting the time dimension difference. This time series analysis method can capture rhythm abnormalities or action sequence disorders in the behavior execution process.
[0084] The spatial trajectory deviation detection focuses on the motion trajectory difference of the skeletal key points in three-dimensional space. The spatial domain transformation is performed on the theoretically predicted skeletal point coordinate sequence and the actual detection result, and the original coordinate data is converted to a more discriminative feature space. The trajectory curvature similarity ratio calculation compares the bending degree difference of the theoretical trajectory and the actual trajectory at each key point, quantifying the deviation of the two in the motion path geometry. The spatial deviation index of each joint node considers multiple factors such as position deviation, velocity difference and acceleration change, and by weighted fusion of these spatial features, a feature matrix reflecting the overall spatial motion difference is constructed. This spatial analysis method can identify abnormal body posture, unnatural limb movement or abnormal interaction with environmental objects and other behavior characteristics.
[0085] The micro-expression frequency spectrum similarity evaluation focuses on the analysis of facial subtle muscle activity. The pattern matching algorithm compares the theoretical micro-expression intensity spectrum with the actual observed spectrum in multiple dimensions, calculating the differences in spectrum shape, energy distribution and phase characteristics. The phase synchronization error analysis of specific muscle group activation intensity focuses on the coordinated changes of muscle activity in different facial regions, and by comparing the theoretical expectation and the actual observed muscle activation timing relationship, the degree of incoordination in the expression execution process is quantified. The divergence difference calculation of spectral energy distribution uses information theory to measure the overall difference between the two spectra, generating a comprehensive index vector reflecting micro-expression abnormalities. This micro-expression analysis method can capture subtle flaws in deliberate fake expressions or abnormal features of emotional fluctuations.
[0086] The generation process of the multi-modal difference feature tensor needs to effectively fuse features from different analysis dimensions. High-order feature fusion technology uses tensor operation method to combine time deviation matrix, spatial deviation matrix and micro-expression similarity vector in a unified high-dimensional space, preserving the original structural relationship of each modal feature. The standardization processing of modal importance weighting automatically adjusts the contribution weight of each feature modal in the final feature representation according to its discriminability difference in behavior recognition, while eliminating the scale difference problem caused by different feature dimensions and numerical ranges. The final output of the three-order tensor structure is organized according to the three dimensions of target identification, time segment and modal type, forming a multi-modal feature representation that can fully reflect behavior abnormalities. This tensor structure facilitates the analysis of behavior abnormal features from different angles by the processing module.
[0087] The key band energy ratio algorithm has specific application value in micro-expression analysis. The algorithm first determines several frequency band ranges closely related to emotional expression, which usually correspond to the activity characteristics of specific facial muscle groups. The energy distribution of theoretically predicted and actually observed micro-expression data in these key bands is finely compared, and the relative difference degree of energy density in each band is calculated. The selection of emotion-related frequency bands is based on the prior knowledge of the facial action coding system, focusing on muscle activity patterns with high correlation to basic emotional expression. The quantitative comparison of energy density not only considers the amplitude difference, but also analyzes the distribution form change of energy in the frequency domain, thereby constructing a more comprehensive micro-expression abnormal feature vector. This band-specific analysis method improves the detection sensitivity of deliberately controlled expressions.
[0088] The overall processing flow of the behavior segment extraction module emphasizes multi-angle and multi-level abnormal feature mining. Time continuity analysis reveals the timing abnormalities in behavior execution, spatial trajectory detection finds the geometric deviations of movement paths and postures, and micro-expression analysis captures abnormal patterns of facial subtle muscle activity. The three analysis dimensions complement each other and together build a comprehensive description of abnormal behavior. The module adopts a phased processing strategy, first performing fine feature extraction and difference calculation within each modality, then performing cross-modality feature fusion and standardization, and finally forming a structured multi-modal difference feature representation. This processing method not only preserves the independence of each modality feature, but also establishes the correlation between modalities, providing rich input features for abnormal propagation analysis. The entire implementation process focuses on the balance between computational efficiency and feature discriminability, using a parallel computing framework to accelerate multi-modal feature extraction while maintaining feature discriminability and interpretability.
[0089] Example 3: see Figure 4The implementation process of the anomaly propagation modeling module starts from the modeling of the physical space of the monitoring scene. According to the moving trajectory data of the target object (which is derived from the sequence of three-dimensional coordinates of the skeletal key points output by the spatio-temporal feature modeling module: by calculating the position changes of the skeletal key points within 10 consecutive frames, fitting the moving path of the target object, eliminating single-point coordinate deviations caused by occlusion, and ensuring the continuity of the trajectory), the system constructs a spatial topology graph containing the connection relationship of each sub-region. The nodes in the graph represent the functional partitions in the monitoring area, and the edges represent the physical connection paths (such as doorways, corridors) between regions. The moving trajectory data is obtained through the above skeletal key point derivation process, and the position information and the duration of the target appearing in each region at different time periods are recorded synchronously, providing dynamic data support for the association of the topology graph nodes. The physical connection attribute labeling includes the connection mode of the actual building structure such as doorways and corridors, while considering obstacles and traffic restrictions. The field of view coverage range of the fixed monitoring equipment is superimposed on the topology graph in the form of a polygonal region, marking the space range that each camera can effectively monitor. By integrating these information, the system generates a physical topology model containing a spatial correlation matrix and a region accessibility matrix, where the spatial correlation matrix quantifies the connection strength between regions, and the region accessibility matrix calculates the transition probability from one region to another.
[0090] The mapping process of the multi-modal difference feature tensor to the physical topology model adopts a region binding mechanism. Each behavior anomaly feature tensor is associated with the corresponding region node in the topology graph according to the target position information at the time of its generation. This mapping establishes a correspondence between the abnormal features and the spatial position, providing a data basis for propagation analysis. Graph convolution networks play a core role in anomaly propagation simulation. The network structure design takes into account the spatial characteristics of the monitoring scene, and adopts multi-layer graph convolution operations to capture the anomaly propagation patterns between regions. The decay factor calculation of abnormal behavior is based on the physical distance and connection mode between regions. Abnormal signals that propagate over long distances or pass through obstacle regions will be appropriately attenuated. The cross-region attention mechanism automatically learns the correlation strength between different regions, focusing on those region pairs that are spatially separated but have similar behavior patterns. The random walk method simulates the diffusion process of abnormal behavior in the topology network, generating multiple possible propagation paths, each representing a potential anomaly propagation mode.
[0091] The generation process of the risk probability distribution cloud map integrates multiple simulation results, and the system statistically counts the frequency of each sub-region in all random walk paths. The higher the frequency, the greater the possibility of the region being affected by abnormal behavior. The region accessibility parameter adjusts the frequency statistics result, considering the influence of the actual movement difficulty on the abnormal propagation. The residence probability value calculation combines the appearance frequency and residence time factors to reflect the possibility of abnormal behavior in a specific area. The final generated risk probability distribution cloud map is presented in the form of a heat map, with different color depths representing the abnormal risk level of each region. This visualization method intuitively displays the high-risk aggregation area in the monitoring scene, facilitating the security personnel to quickly locate the area of concern.
[0092] The density clustering algorithm plays a key role in high-risk area identification. The system performs multi-scale clustering analysis on the risk probability distribution cloud map to identify aggregation areas with probability values significantly higher than surrounding areas. The spatial continuity of the region is considered during the clustering process, and regions that are physically adjacent and have similar risk levels are merged into a high-risk area. The target object position information is used to further accurately locate the potential source of abnormal behavior, combined with the topological connection relationship to infer the direction of abnormal behavior propagation. The annotation of the physical region boundary of abnormal behavior uses a polygon fence method to clearly mark the area range that needs to be focused on on the electronic map of the monitoring system. The boundary demarcation considers both the spatial distribution of the clustering result and the actual monitoring needs and management convenience. The generation of the suspicious target identification set is based on the association analysis of abnormal areas and actual monitoring targets. The system marks the target objects appearing in the high-risk area and records their appearance characteristics, behavior patterns, and movement trajectories. These targets constitute the initial suspicious set, which will be further filtered through behavior verification. The extraction process of the abnormal behavior propagation main path sequence analyzes all random walk simulation results and counts the frequency of each path. The path sequence with the highest frequency reflects the most likely propagation trajectory of abnormal behavior in the scene, which is represented as an ordered node list by the system, with each node corresponding to a monitoring region. This main path information helps to predict the diffusion direction of abnormal behavior and provides a reference for the deployment of prevention and control measures.
[0093] The whole abnormal propagation modeling process emphasizes the organic combination of spatial factors and behavioral characteristics. The physical topology model accurately reflects the actual layout and connection relationship of the monitoring scene, providing real spatial constraints for propagation simulation. The graph convolution network and random walk method effectively capture the diffusion characteristics of abnormal behavior in the spatial network, producing propagation paths that conform to actual situations. The calculation of risk probability distribution takes into account various factors, and the generated cloud map can reliably indicate potential risk areas. The identification of suspicious targets and main paths is based on rigorous data analysis, providing valuable reference information for security decisions. In system implementation, attention is paid to the balance between computational efficiency and result reliability, and a distributed computing framework is used to accelerate the analysis process of large-scale topology networks while ensuring the accuracy of risk assessment. The output of the abnormal propagation modeling module provides spatial dimension analysis results for the whole behavior recognition and early warning system, making up for the limitations of pure behavior characteristic analysis.
[0094] Embodiment 4: refer to Figure 5 The implementation process of the early warning strategy generation module is based on real-time analysis results of the risk probability distribution cloud map. When the abnormal probability value of a specific area in the cloud map exceeds the preset risk threshold, the system automatically issues a set of monitoring instructions to the video acquisition terminal to which the area belongs. The execution of these instructions includes increasing the micro-expression capture frame rate from the basic frequency to 3-5 times the original frame rate, and the specific multiple value is dynamically adjusted according to the risk level. The high-frame-rate capture mode is realized by adjusting the reading timing and signal processing parameters of the image sensor, ensuring that more intensive time sampling is obtained while maintaining image quality. The muscle activity intensity real-time tracking mode enabled simultaneously focuses on monitoring specific areas of the face, which usually include the inter-brow, eye corner, and mouth corner, etc. The frequency band energy mutation detection algorithm analyzes the spectral characteristics of these areas in real time, and triggers the recording mechanism immediately when a significant change in energy distribution is detected. The transient behavior detector deployed on the edge continuously analyzes the micro-expression waveform data using a sliding window method, and automatically saves the waveform segment data of the time period when an abnormal waveform pattern is identified. These segments contain complete information such as waveform amplitude, frequency, and phase.
[0095] The motion trajectory disturbance test is implemented for the adjacent area targets with the most significant gradient change in the risk probability distribution cloud map. The system first calculates the risk probability gradient field of each position in the cloud map and identifies the region boundaries with the largest probability change rate. Multi-dimensional motion disturbances are applied to the target objects within these boundary regions, and the disturbance mode injects preset motion trajectory interference signals through a virtual reality interface. These interference signals simulate various visual stimuli in sudden situations, such as sudden appearance of virtual obstacles, changes in path guidance, or sudden changes in environmental lighting, etc. The calculation of the theoretical behavior response spectrum is based on the spatial correlation characteristics of the physical topology model, and its calculation formula is:
[0096]
[0097] where: R theory represents the theoretical behavior response spectrum, M is the number of associated regions, w i is the weight coefficient of the i-th region, r i represents the region spatial coordinate vector, v i is the region feature vector, and Φ is the response calculation function. This function comprehensively considers the spatial correlation strength and region feature similarity to predict the expected behavior pattern of the target under interference conditions.
[0098] The acquisition of the actual behavior response spectrum is achieved by recording the actual behavior changes of the target after the injection of interference signals. The system uses a multi-modal data acquisition method to synchronously record the multi-dimensional reaction data such as the motion trajectory, body posture, and facial expression of the target. These data are processed through feature extraction and fusion to form a complete actual response spectrum, containing the behavior change characteristics in time series. The feature distance between the theoretical response spectrum and the actual response spectrum is calculated in the abnormal deviation determination stage, and the distance measurement method in the multi-dimensional feature space is adopted. The feature distance calculation considers the overall similarity and local feature difference of the behavior pattern, and obtains the comprehensive deviation by weighted combination of multiple distance indicators. The calculation of the abnormal target deviation degree compares the comprehensive deviation with the preset threshold value, and when the deviation continuously exceeds the threshold range, the system marks the target as an abnormal behavior associated target. The marking process adopts a gradual confirmation mechanism, which requires consistent judgment of multiple time segments to finally confirm the abnormal state.
[0099] The entire early warning strategy generation process emphasizes the combination of real-time and adaptability, and the adjustment of monitoring parameters is based on the dynamic changes of risk distribution, ensuring the optimal allocation of monitoring resources. The upgrade of micro-expression monitoring mode not only improves the sampling frequency, but also enhances the ability to capture subtle muscle activities. The motion trajectory disturbance test reveals potential abnormal behavior through active stimulation, providing an active detection means. The calculation of the theoretical behavior response spectrum fully utilizes the spatial information of the physical topology model, making the prediction results more consistent with the characteristics of the actual scene. The multi-modal acquisition of the actual response ensures the comprehensiveness and accuracy of behavior evaluation. The abnormality determination mechanism adopts multi-index comprehensive judgment to reduce the possibility of false positives. In the system implementation, attention is paid to the cooperative work between various components to ensure the rapid conversion from risk detection to early warning response. All processing processes consider real-time requirements, and a streaming processing framework is adopted to ensure that the analysis and decision-making cycle is completed in the shortest time. The generation of early warning strategies not only considers the current risk state, but also considers historical behavior patterns and environmental context information, making the decision more comprehensive and reliable.
[0100] Example 5: Implementation of the optical flow feature compensator employs a motion trajectory reconstruction algorithm for in-depth analysis of real-time captured motion optical flow vector fields. This algorithm first decomposes the continuous optical flow vector sequence into different motion pattern components, distinguishing between the basic motion components and abnormal fluctuation components through frequency domain analysis and pattern recognition techniques. The basic motion components represent the regular behavior patterns of the target object, which usually exhibit certain regularity and predictability, reflecting normal movement characteristics and behavioral habits. The abnormal fluctuation components capture those motion features that deviate from the regular patterns, which may manifest as sudden changes in speed, unusual motion directions, or abnormal motion trajectory morphologies. The motion pattern decomposition process employs a multi-scale analysis method, considering both short-term motion details and long-term motion trends, ensuring effective identification of various types of abnormal fluctuations.
[0101] The energy proportion calculation focuses on the relative importance of abnormal fluctuation components in the overall motion pattern. The system quantifies the significance of abnormal motion by calculating the ratio of the amplitude energy of abnormal components to the total motion energy. This ratio reflects the contribution proportion of abnormal fluctuations in the overall motion pattern, and a higher ratio indicates that the abnormal motion features are more obvious and prominent. Based on this energy proportion value, the system dynamically adjusts the prediction weight coefficient of the time series modeling unit. When the energy proportion of abnormal fluctuation components is high, the weight of the time series prediction model is correspondingly reduced, reducing its influence on the final behavior prediction. This dynamic adjustment mechanism ensures that the behavior recognition model can adapt to various motion conditions, adjusting the prediction strategy in a timely manner when abnormal motion occurs, improving the robustness and adaptability of the system.
[0102] The prediction weight coefficient adjustment process of the time series modeling unit adopts a smooth transition mechanism to avoid fluctuations in the prediction results caused by sudden changes in the weight. The system gradually adjusts the weight coefficient according to the change trend of the energy proportion, while considering the historical weight change situation to ensure the stability and continuity of the adjustment process. The adjustment range of the weight coefficient is normalized to maintain it within a reasonable numerical interval, ensuring timely response to abnormal situations while avoiding prediction instability caused by excessive adjustment. This refined weight management mechanism enables the behavior recognition model to maintain prediction accuracy while effectively dealing with various abnormal motion situations.
[0103] The implementation of the motion trajectory reconstruction algorithm includes multiple processing stages. The initial stage preprocesses the optical flow vector field, including noise filtering, vector smoothing, and missing data compensation, to improve the quality of the motion data. The motion feature extraction stage extracts various characteristic parameters such as velocity, acceleration, and motion direction from the optical flow field. The pattern decomposition stage uses clustering analysis and principal component analysis to decompose complex motion patterns into several basic motion components. The anomaly detection algorithm is based on statistical learning principles and identifies motion components that are significantly different from the regular pattern, marking them as abnormal fluctuation components. The entire processing process uses incremental learning, which can continuously optimize the accuracy of motion pattern decomposition as data accumulates.
[0104] The cooperation of the optical flow feature compensator with other modules of the system is reflected in data exchange and the integration of processing flow. The compensator receives real-time optical flow data from the video analysis engine, processes it, and outputs motion pattern analysis results and weight adjustment suggestions. These output data are passed to the time series modeling unit and other related components, affecting the behavior prediction and analysis process. At the same time, the compensator also obtains environmental context information and historical behavior data from other modules, which are used to optimize the accuracy of motion pattern decomposition and the precision of anomaly detection. This two-way data exchange mechanism ensures the coordinated operation of each component in the system, improving the overall behavior recognition performance.
[0105] Particular attention is paid to balancing the requirements of computational efficiency and real-time performance during implementation. The motion trajectory reconstruction algorithm uses an optimized calculation strategy to minimize computational complexity while ensuring analysis accuracy. Parallel processing architecture is used to accelerate the analysis process of large-scale optical flow data, and a distributed computing framework ensures that the system can handle massive optical flow data generated by high-resolution video streams. Memory management mechanisms optimize data storage and access patterns, reducing latency during data processing. These technical measures ensure that the optical flow feature compensator can operate stably in real-time video stream processing scenarios, meeting the system's requirements for response speed and processing efficiency.
[0106] Performance optimization of the optical flow feature compensator is ongoing, with algorithm parameters being adjusted continuously based on monitoring of running status and analysis of processing results. The system records performance indicators such as the accuracy of motion pattern decomposition and the recall rate of anomaly detection, and optimizes algorithm performance based on these indicator data. The adaptive learning mechanism enables the compensator to adapt to different monitoring scenarios and target types, maintaining stable performance in various environmental conditions. The identification threshold of abnormal fluctuation components is dynamically adjusted according to actual conditions, avoiding false positives or false negatives due to environmental changes. This continuous optimization mechanism ensures that the optical flow feature compensator can work stably and reliably for a long time, providing accurate motion analysis results for the behavior recognition system.
[0107] It is to be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting; it is not intended to exclude myriad other embodiments of the present application that other inventors can develop based on the same general inventive concepts embodied by the described embodiments. That is, although the present application is described in terms of particular embodiments and illustrative figures, it should be apparent that the scope of the present application is not limited to these specific embodiments.
[0108] While the embodiments of the application have been shown and described herein, it will be understood by those skilled in the art that many changes, modifications, substitutions and alterations to these embodiments can be made without departing from the principles and spirits of the application, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A real-time video stream behavior recognition and early warning system, characterized in that, include: Spatiotemporal feature modeling module: Based on the historical behavior video data of the target object, a dynamic behavior recognition model is constructed. The model captures the three-dimensional coordinate sequence of the skeletal key points of the target, the motion optical flow vector field and the intensity spectrum of micro-expression muscle activity in the current video stream in real time. The theoretical behavior pattern vector is output through the dynamic behavior recognition model. Behavior segment extraction module: Perform cross-modal difference analysis on the theoretical behavior pattern vector and the measured behavior vector of the video analysis engine. The cross-modal difference analysis includes temporal continuity deviation, spatial trajectory offset and micro-expression spectrum similarity, and generates a multimodal difference feature tensor at the target object level. Anomaly propagation modeling module: Input the multimodal difference feature tensor into the scene topology inference network, and combine the physical space constraints of the monitoring area with the target object's movement trajectory information to generate a risk probability distribution cloud map of the abnormal behavior propagation path; Risk area location module: Identifies high-risk cluster areas based on the risk probability distribution cloud map, and marks the physical boundaries of abnormal behavior areas where the probability value exceeds a preset risk probability threshold; Early warning strategy generation module: Configures monitoring parameters based on the risk probability distribution cloud map, including: High frame rate micro-expression capture mode is enabled for high-risk gathering areas; Tests were conducted by applying motion trajectory perturbations to targets in the vicinity. The spatiotemporal feature modeling module specifically includes: Historical behavior feature mining: Multi-scale spatiotemporal decomposition processing is performed on the historical behavior video data of the target object, including: using an adaptive spatiotemporal pyramid to extract the energy distribution of steady-state and transient motion of the skeletal key point sequence, establishing the correlation mapping between motion vector field and behavioral intention through optical flow coupling analysis, and using sequence alignment algorithm to calibrate the intensity pattern of micro-expression activity in different scenarios; Dynamic behavior recognition model construction: The processed historical behavior data is input into the fusion prediction architecture, which includes: Temporal modeling units based on behavior evolution curves are used to generate basic behavior prediction vectors; A 3D convolutional network with embedded spatial attention mechanism corrects prediction bias caused by changes in viewpoint; Based on the optical flow feature compensator, the prediction weights are dynamically adjusted according to the real-time captured motion optical flow vector field. The edge computing unit synchronously collects the position drift rate and velocity fluctuation parameters of key skeletal points, the amplitude gradient distribution and directional consistency coefficient of motion optical flow vectors, and the frequency band energy entropy value of micro-expression intensity spectrum. Theoretical value generation: Input the real-time captured data into the dynamic behavior recognition model to obtain the theoretical behavior pattern vector.
2. The real-time video stream behavior recognition and early warning system according to claim 1, characterized in that, The behavior fragment extraction module specifically includes: Time continuity deviation calculation: The theoretical behavior pattern vector and the measured behavior vector are compared by sliding comparison with a preset time segment. The non-uniformly sampled behavior sequence is aligned by the sequence dynamic warping algorithm. The cumulative trajectory deviation in each segment is calculated to generate a time domain deviation feature matrix. Spatial trajectory offset detection: The spatial domain transformation decomposition of the three-dimensional coordinates of the skeletal key points of theoretical and measured values is performed, the trajectory curvature similarity ratio is calculated, the spatial offset index of each joint node is extracted, and a spatial offset feature matrix is constructed. Micro-expression spectral similarity assessment: Based on the pattern matching algorithm, the distance distribution between theoretical values and measured micro-expression intensity spectra is measured, the phase synchronization error of muscle group activation intensity is calculated, the divergence difference of spectral energy distribution is quantified, and micro-expression similarity vectors are generated. Multimodal difference feature tensor generation: The temporal deviation feature matrix, spatial offset feature matrix, and micro-expression similarity vector are fused with high-order features. Through modal importance weighted standardization, scale differences are eliminated, and a third-order multimodal difference feature tensor with dimensions of [target identifier x time segment x modal type] is output.
3. The real-time video stream behavior recognition and early warning system according to claim 2, characterized in that, The anomaly propagation modeling module specifically includes: Scene topology modeling: Construct a spatial topology map of the monitoring area based on the target object's movement trajectory information, mark the physical connectivity attributes between each sub-area, overlay the field of view coverage constraints of the fixed monitoring equipment on the topology map, and generate a physical topology model containing a spatial association matrix and an area reachability matrix; Anomaly propagation deduction: Mapping the multimodal differential feature tensor to the corresponding region of the physical topology model; Anomaly propagation simulation is performed based on graph convolutional networks. The computation of the anomaly propagation simulation includes: Calculate the attenuation factor of anomalous behavior based on the region's physical connectivity attributes; Capture anomalous correlation features through cross-regional attention mechanisms; A random walk method is used to simulate the propagation path of anomalous behavior in a topological network. Risk distribution generation: Statistically analyze the frequency of abnormal behavior in each sub-region during simulated propagation, combine regional accessibility parameters to calculate the probability value of abnormal behavior residence, and generate a risk probability distribution cloud map covering the entire region; Physical boundary delineation: Perform density clustering analysis on the risk probability distribution cloud map to identify high-risk probability clustering areas; Based on the location of the target object and its physical topology connection, mark the physical boundaries of the abnormal behavior area.
4. The real-time video stream behavior recognition and early warning system according to claim 3, characterized in that, The early warning strategy generation module specifically includes: When the anomaly probability value of a certain area in the risk probability distribution cloud map exceeds a preset risk probability threshold, a monitoring command is sent to the video acquisition terminal belonging to that area, and the following is executed: Increase the frame rate for micro-expression capture to 3-5 times the original frame rate; Simultaneously enable real-time muscle activity intensity tracking mode to capture frequency band energy mutations in the facial region; Deploy transient behavior detectors at the edge to record abnormal micro-expression waveform fragments; Motion trajectory perturbation test execution: Apply multi-dimensional motion perturbation to the target in the neighborhood area with the largest change in risk probability gradient in the risk probability distribution cloud map.
5. The real-time video stream behavior recognition and early warning system according to claim 4, characterized in that, The multidimensional motion perturbation includes: Inject a preset motion trajectory interference signal through a virtual reality interface; Theoretical behavioral response calculation: Based on the physical topology model and spatial correlation matrix, calculate the theoretical behavioral response spectrum; Acquisition of measured behavioral response: Record the behavioral patterns of each target after the injection of interference signal to obtain the measured behavioral response spectrum; Anomaly offset determination: Calculate the characteristic distance between the theoretical behavioral response spectrum and the measured behavioral response spectrum; By comparing the characteristic distance between the measured behavioral response spectrum and the theoretical behavioral response spectrum, the degree of deviation of abnormal targets is calculated. When the degree of deviation of abnormal targets exceeds the preset deviation threshold, they are marked as abnormal behavior associated targets.
6. The real-time video stream behavior recognition and early warning system according to claim 5, characterized in that, The spatiotemporal feature modeling module also includes: Adaptive normalization processing based on ambient lighting conditions eliminates characteristic noise caused by changes in light intensity; The correlation characteristics of skeleton, optical flow, and micro-expression are integrated through a multi-source feature fusion algorithm; The output contains a theoretical behavior pattern vector that includes the normal behavior fluctuation range. The theoretical behavior pattern vector is dynamically updated as the behavior pattern evolves.
7. The real-time video stream behavior recognition and early warning system according to claim 6, characterized in that, The anomaly propagation modeling module also includes an output containing a set of suspicious target identifiers and an anomaly propagation main path sequence, wherein: The suspicious target identifier set is based on the monitoring area covered by the physical boundary of the abnormal behavior area, and the monitoring area is bound to the actual target object to form the suspicious target identifier set; The anomaly propagation main path sequence is obtained by counting the occurrence frequency of each path during the anomaly diffusion process simulated by random walk, selecting the path sequence with the highest cumulative occurrence frequency, and outputting an ordered list of nodes reflecting the propagation trajectory of the abnormal behavior in the scene.
8. The real-time video stream behavior recognition and early warning system according to claim 7, characterized in that, In the construction of the dynamic behavior recognition model: The optical flow feature compensator uses a motion trajectory reconstruction algorithm to decompose the motion pattern of the real-time captured motion optical flow vector field and identify the basic motion component and abnormal fluctuation component. The prediction weight coefficients of the time series modeling unit are dynamically adjusted based on the energy proportion of the abnormal fluctuation components.
9. The real-time video stream behavior recognition and early warning system according to claim 8, characterized in that, In the behavior fragment extraction module: The micro-expression spectrum similarity assessment uses a key frequency band energy comparison algorithm to extract the energy distribution features of preset emotion-related frequency bands; The energy density of theoretical and measured data in the emotionally relevant frequency bands is quantitatively compared, and a micro-expression similarity vector is constructed based on the degree of relative difference.
Citation Information
Patent Citations
Intelligent human shape trajectory prediction and alarm system and method based on multi-modal video analysis
CN120047897A