A method and system for identifying human behavior in video surveillance
By fusion of the monitoring video sequence and positioning the bone joints, building action connection diagrams and timing correlation edges, learning behavior patterns and abnormal measurements, the insufficient identification of personnel behavior in the existing technology in complex scenarios is solved, and the accurate identification and risk assessment of personnel behavior is achieved, and the accuracy and timeliness of abnormal detection are improved.
Patent Information
- Application Number
- CN202411622814.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-13
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2044-11-13
AI Technical Summary
When the prior art deals with the identification of personnel behavior in complex scenarios, it lacks the overall modeling of human behavior in the time and space dimensions, and it is difficult to accurately capture the logical correlation and evolution laws between behaviors. The lack of in-depth analysis and risk assessment of abnormality detection, resulting in insufficient accuracy and timeliness of early warning information.
By extracting the human body's movement texture features and motion trajectory features from the monitoring video sequence, performing feature fusion, combining bone joint positioning and action node extraction, building action connection diagrams and timing correlation edges, conducting behavior pattern learning, performing action semantic mapping and abnormal measurements, evaluating behavior risk levels, and generating behavior recognition results and early warning information.
It realizes accurate identification of personnel behavior and dynamic assessment of risk levels, improves the sensitivity of abnormal behavior detection and timely warning, and ensures the safety of monitoring scenarios.
Smart Images

Figure CN119580352B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing, and in particular to a method and system for identifying human behavior in video surveillance. Background Art
[0002] In recent years, with the widespread application of video surveillance systems in public security, intelligent video analysis technology has garnered increasing attention. Existing methods for human behavior recognition primarily utilize techniques such as image processing-based feature extraction, deep learning-based behavior classification, and rule-based anomaly detection. These methods extract human appearance, motion, and posture features from video sequences, combining them with behavioral pattern analysis and temporal relationship modeling to automatically identify human behavior and provide early warning of abnormal behavior. Current technical solutions typically decompose the behavior recognition problem into multiple subtasks, including feature extraction, behavior classification, and anomaly detection. They employ methods such as multimodal feature fusion and deep learning to improve recognition accuracy.
[0003] However, existing technical solutions still have some shortcomings when dealing with human behavior recognition in complex scenarios: first, traditional feature extraction methods often only focus on features of a single dimension and lack holistic modeling of human behavior in the time and space dimensions; second, existing behavior recognition algorithms have a relatively shallow understanding of action semantics and find it difficult to accurately capture the logical connections and evolution patterns between behaviors; third, abnormal behavior detection is mostly based on simple threshold judgments, lacking in-depth analysis of behavior patterns and refined assessment of risk levels, resulting in insufficient accuracy and timeliness of warning information. Summary of the Invention
[0004] The present application provides a method and system for identifying human behavior in video surveillance, which is used to achieve accurate identification of human behavior and dynamic assessment of risk levels based on multimodal feature fusion and spatiotemporal behavior modeling.
[0005] In the first aspect, the present application provides a method for identifying human behavior through video surveillance, which includes: extracting human action texture features and motion trajectory features from an input monitoring video sequence, performing feature fusion processing through behavior saliency analysis, and obtaining a human behavior feature vector; performing skeletal joint positioning processing on the human behavior feature vector, and performing action node and joint angle extraction processing on the positioning area to obtain human action skeleton sequence data; constructing an action connection graph and temporal correlation edges for the human action skeleton sequence data, performing behavior pattern learning processing through motion paradigm analysis, and obtaining spatiotemporal behavior features; performing action semantic mapping processing on the spatiotemporal behavior features and human behavior feature vectors, and obtaining behavior semantic description features through behavior temporal chain construction processing; performing behavior pattern matching and anomaly measurement processing on the behavior semantic description features, and performing feature discrimination processing through behavior similarity comparison to obtain a behavior type identifier and anomaly degree index; performing behavior continuity analysis processing on the behavior type identifier and anomaly degree index, and obtaining behavior recognition results and warning information through behavior risk level assessment processing.
[0006] In a second aspect, the present application provides a system for identifying human behavior through video surveillance, the system comprising:
[0007] The acquisition module is used to extract the texture features and motion trajectory features of human actions from the input surveillance video sequence, perform feature fusion processing through behavior saliency analysis, and obtain the human behavior feature vector;
[0008] A positioning module is used to perform skeleton joint positioning processing on the human behavior feature vector, and to extract action nodes and joint angles from the positioning area to obtain human action skeleton sequence data;
[0009] A construction module is used to construct an action connection graph and a temporal correlation edge for the human action skeleton sequence data, and perform behavior pattern learning processing through motion paradigm analysis to obtain spatiotemporal behavior characteristics;
[0010] A mapping module is used to perform action semantic mapping processing on the spatiotemporal behavior features and the human behavior feature vectors, and obtain behavior semantic description features through behavior time sequence chain construction processing;
[0011] A measurement module is used to perform behavior pattern matching and anomaly measurement processing on the behavior semantic description features, perform feature discrimination processing by behavior similarity comparison, and obtain a behavior type identifier and an anomaly degree index;
[0012] The analysis module is used to perform behavior persistence analysis on the behavior type identification and abnormality degree index, and obtain behavior recognition results and early warning information through behavior risk level assessment.
[0013] The technical solution provided in this application extracts human action texture features and motion trajectory features from surveillance video sequences and integrates them with behavioral saliency analysis to achieve feature fusion, effectively capturing the appearance and motion information of human behavior and improving the integrity of feature representation. By locating skeletal joints and extracting action nodes from human action feature vectors, key human body parts are accurately located, achieving a precise description of human posture. Action connection graphs and temporal correlation edges are constructed based on human action skeleton sequence data. Behavioral pattern learning is combined with motion paradigm analysis to establish a spatiotemporal structural representation of behavior, enhancing the ability to model complex behavioral patterns. Action semantics are mapped between spatiotemporal behavioral features and human action feature vectors. By constructing a behavioral temporal chain, an effective mapping from low-level features to high-level semantics is achieved, improving the accuracy of behavior understanding. Behavioral pattern matching and anomaly measurement are performed based on behavioral semantic description features. Feature discrimination is performed through behavioral similarity comparison, effectively identifying abnormal behavior patterns and improving the sensitivity of abnormal behavior detection. Finally, continuous analysis and risk assessment of behavior type identification and anomaly severity indicators are performed, enabling timely early warning of abnormal behavior and ensuring the safety of the surveillance scene. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0015] Figure 1 This is a schematic diagram of an embodiment of a method for identifying human behavior in video surveillance in an embodiment of the present application;
[0016] Figure 2 This is a schematic diagram of an embodiment of a human behavior recognition system for video surveillance in an embodiment of the present application. DETAILED DESCRIPTION
[0017] The embodiments of the present application provide a method and system for identifying human behavior in video surveillance. The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or that are inherent to these processes, methods, products or devices.
[0018] For ease of understanding, the specific process of the embodiment of the present application is described below. Figure 1 In one embodiment of the present application, a method for identifying human behavior through video surveillance includes:
[0019] Step S101: extracting human action texture features and motion trajectory features from the input surveillance video sequence, performing feature fusion processing through behavior saliency analysis, and obtaining a human action feature vector;
[0020] Step S102: performing skeleton joint positioning processing on the human behavior feature vector, and extracting action nodes and joint angles from the positioning area to obtain human action skeleton sequence data;
[0021] Step S103: constructing an action connection graph and temporal correlation edges for the human action skeleton sequence data, performing behavior pattern learning processing through motion paradigm analysis, and obtaining spatiotemporal behavior features;
[0022] Step S104: performing action semantic mapping processing on the spatiotemporal behavior features and the human behavior feature vectors, and constructing and processing the behavior time sequence chain to obtain behavior semantic description features;
[0023] Step S105: Performing behavior pattern matching and anomaly measurement processing on the behavior semantic description features, performing feature discrimination processing by behavior similarity comparison, and obtaining a behavior type identifier and an anomaly degree index;
[0024] Step S106: Conduct behavior continuity analysis on the behavior type identifier and abnormality index, and obtain behavior recognition results and warning information through behavior risk level assessment.
[0025] It is understandable that the execution subject of this application can be a video surveillance personnel behavior recognition system, or a terminal or a server, which is not limited here. The embodiment of this application is described by taking the server as the execution subject as an example.
[0026] Specifically, feature extraction is performed on the input surveillance video sequence. Texture features are extracted from video frames using Gabor filters. This process involves convolution operations on the image using a set of two-dimensional Gabor filters of varying orientations and scales. This process extracts local texture information related to human motion, including features such as clothing texture and limb contours. Optical flow is also used to calculate the motion vector field between adjacent frames to obtain human motion trajectory features. Optical flow calculates the displacement vectors of image pixels between consecutive frames to obtain speed and direction information of the moving target. A significance analysis is performed on the extracted texture and motion features, and weight coefficients are calculated for the feature regions, highlighting the feature components that contribute most to action recognition. Finally, a comprehensive human action feature vector is constructed through weighted fusion. Based on the obtained human action feature vector, precise localization of human skeletal joints is performed. First, a region segmentation algorithm is used to determine the human body region and generate a human contour mask. Then, geometric constraints are applied within this region to locate key skeletal nodes, including 17 major joints: the head, neck, shoulder, elbow, wrist, hip, knee, and ankle. The spatial distribution of these nodes is verified by introducing anatomical constraints, and the connection relationships and three-dimensional angle information between each joint are calculated. For each joint point, its spatial coordinates and confidence score are recorded, and jitter is eliminated through temporal smoothing, ultimately obtaining skeleton sequence data that describes changes in human posture.
[0027] An action connection graph is constructed based on skeleton sequence data. Nodes in the graph represent skeletal joints, and edges represent the physical connections between joints. Edges are weighted by calculating the Euclidean distance and angular change between nodes to characterize the strength of joint motion. A temporal sliding window is used to analyze the correlation between consecutive frames, with a window size of 25 frames and a step size of 5 frames. Temporal correlation edges are established in the temporal dimension to characterize the temporal evolution of actions. Pattern learning is performed on the constructed spatiotemporal graph structure to extract node and edge features. Graph convolution is then used to fuse spatial and temporal information to obtain a unified spatiotemporal behavior representation. Semantic mapping is then performed between the spatiotemporal behavior features and human behavior feature vectors to establish a correspondence between actions and semantic concepts. Semantic layering of action sequences is performed to decompose complex behaviors into basic action units, analyzing the temporal dependencies and causal relationships between these units. A behavior rule library is constructed based on scene semantic constraints, and semantic reasoning is performed on action sequences to identify action intent. Through action rule matching and semantic association analysis, a behavior temporal chain is constructed to describe the semantic evolution of behaviors and generate behavior semantic description features with temporal structure.
[0028] Based on the semantic description features of behavior, behavioral pattern matching and anomaly detection are performed. First, the behavioral features are encoded into standardized feature vectors and compared with a predefined behavioral pattern library for similarity. Cosine similarity is used to measure the distance between feature vectors, and an adaptive threshold is set for behavior classification. At the same time, the abnormality measurement value of the behavior is calculated, and the degree of abnormality of the behavior is evaluated by analyzing the degree of deviation between the behavioral features and normal behavior patterns. Based on the similarity calculation and anomaly measurement results, the behavior type identification and abnormality degree index are obtained. Finally, the behavior type and abnormality degree are continuously analyzed to evaluate the duration and evolution trend of the behavior. The changing patterns of behavioral features are statistically analyzed through time series windows to predict the development trend of the behavior. Combined with the scenario risk level, a risk propagation model is established to comprehensively evaluate the dangerousness of the behavior and generate graded warning information.
[0029] For example, in a shopping mall surveillance scenario, a camera captured a person lingering near a jewelry counter. The system first extracted the person's motion texture features and trajectory. Using a Gabor filter bank, it obtained a 64-dimensional texture feature vector. Optical flow was then used to calculate the velocity field, showing that the person repeatedly moved back and forth within a 3-meter area at a speed of 0.3-0.5 m / s. Skeleton extraction was used to analyze the person's posture, locating 17 key nodes. The system found frequent changes in head and torso posture, with bending angles fluctuating between 30 and 45 degrees, and head rotations exceeding 120 degrees when looking around. Combining spatiotemporal behavioral features with semantic analysis, the system identified this pattern of repeated wandering and frequent looking as suspicious. Comparing the system with a feature library of normal shopping behavior yielded a behavioral similarity of 0.35 and an anomaly score of 0.85, significantly exceeding the warning threshold of 0.7. Further analysis by the system revealed that the behavior lasted for 16 minutes and occurred in the high-risk jewelry area. The risk level assessment value reached 8.5 (out of 10 points). After comprehensive assessment, a second-level warning information was generated, prompting security personnel to pay close attention and conduct on-site verification.
[0030] In this embodiment, by extracting human action texture features and motion trajectory features from surveillance video sequences and integrating them with behavioral saliency analysis to achieve feature fusion, the appearance and motion information of human behavior are effectively captured, improving the integrity of feature representation. By locating skeletal joints and extracting action nodes from human action feature vectors, key human body parts are accurately located, achieving a precise description of human posture. Based on human action skeleton sequence data, an action connection graph and temporal correlation edges are constructed. Combined with motion paradigm analysis, behavioral pattern learning is performed to establish a spatiotemporal structural representation of behavior, enhancing the ability to model complex behavioral patterns. Action semantics are mapped between spatiotemporal behavioral features and human action feature vectors. By constructing a behavioral temporal chain, an effective mapping from low-level features to high-level semantics is achieved, improving the accuracy of behavior understanding. Based on the behavioral semantic description features, behavioral pattern matching and anomaly measurement are performed. Feature discrimination is performed by comparing behavioral similarities, effectively identifying abnormal behavior patterns and improving the sensitivity of abnormal behavior detection. Finally, continuous analysis and risk assessment of behavior type identification and abnormality severity indicators are performed, enabling timely early warning of abnormal behavior and ensuring the safety of the monitoring scenario.
[0031] In a specific embodiment, the process of executing step S101 may specifically include the following steps:
[0032] (1) Performing frame sampling processing on the input monitoring video sequence to obtain video frame sequence data, and performing grayscale and normalization processing on the video frame sequence data to obtain preprocessed image data;
[0033] (2) The pre-processed image data is subjected to texture feature extraction processing by Gabor filter to obtain the human motion texture feature matrix, and the pre-processed image data is subjected to motion vector calculation processing by optical flow method to obtain the human motion trajectory feature matrix;
[0034] (3) performing feature dimension alignment processing on the human motion texture feature matrix and the human motion trajectory feature matrix to obtain an aligned feature matrix, and performing principal component analysis processing on the aligned feature matrix to obtain dimension-reduced feature data;
[0035] (4) Performing regional weight calculation on the reduced-dimensional feature data through saliency calculation to obtain a feature weight matrix, and performing weighted combination processing on the reduced-dimensional feature data and the feature weight matrix to obtain the initial fusion feature;
[0036] (5) Perform feature correlation analysis on the initial fusion features through the covariance matrix to obtain a feature correlation matrix, and perform feature selection on the feature correlation matrix to obtain a core feature set;
[0037] (6) Performing frequency domain feature extraction processing on the core feature set through Fourier transform to obtain a frequency domain feature vector, and performing frequency domain feature screening processing on the frequency domain feature vector to obtain a frequency domain feature set;
[0038] (7) performing feature concatenation processing on the frequency domain feature set and the core feature set to obtain a combined feature vector, and performing feature normalization processing on the combined feature vector to obtain a normalized feature vector;
[0039] (8) The standardized feature vector is subjected to time-frequency feature analysis through wavelet transform to obtain time-frequency feature data, and the time-frequency feature data is subjected to multi-scale feature extraction to obtain the human behavior feature vector.
[0040] Specifically, a frame sampling operation is performed, the sampling frequency is set to 25 frames per second, and a video frame sequence is obtained by uniform sampling. The sampled video frames are grayscaled, and the RGB three-channel image is converted into a single-channel grayscale image. At the same time, the pixel values are normalized, and the pixel value range is uniformly mapped to the [0,1] interval to obtain standardized preprocessed image data. Based on the preprocessed image, the Gabor filter is used to extract texture features. The Gabor filter is a bandpass filter that can effectively capture local texture information. A Gabor filter group with 8 directions and 4 scales is designed to perform convolution operations on the image. The response result of each filter forms a feature map, and finally a human motion texture feature matrix consisting of 32 feature maps is obtained. At the same time, the Lucas-Kanade optical flow method is used to calculate the pixel displacement field between adjacent frames by solving the optical flow equation:
[0041]
[0042] in, Represents the spatial gradient of the image in the x and y directions,
[0043] represents the time gradient, v x 、v y Represent the optical flow velocity component to be determined, obtain the motion vector field, and construct the human motion trajectory feature matrix.
[0044] The texture and motion feature matrices were dimensionally aligned, and spatial interpolation was used to bring the two features to the same spatial resolution, resulting in an aligned feature matrix. Principal component analysis (PCA) was then performed on the aligned feature matrix for dimensionality reduction. The eigenvalues and eigenvectors of the feature covariance matrix were calculated, and the principal components with a cumulative contribution rate of 95% were selected as the reduced-dimensional feature data.
[0045] Based on the feature data after dimensionality reduction, the saliency weight of the feature region is calculated. The saliency calculation is based on the local contrast and spatial distribution of the feature, using the following formula:
[0046]
[0047] Among them, W(x, y) is the significance weight, (x, y) represents the coordinates of the feature point, (μ x , μ y ) represents the center position, (σ x ,σ y_ ) represents the spatial distribution parameter, and C(x, y) represents the local contrast. After obtaining the feature weight matrix, it is weightedly combined with the dimensionality reduction feature data to generate the initial fusion feature. Feature correlation analysis is performed by calculating the covariance matrix of the initial fusion feature, the correlation between different feature components is evaluated, and features with lower correlation are selected to form the core feature set. The core feature set is subjected to a one-dimensional Fourier transform to obtain the frequency domain feature vector, and the main frequency components are retained through energy threshold screening to form a frequency domain feature set. The frequency domain feature set and the core feature set are feature spliced to construct a combined feature vector, and feature normalization is performed to obtain a standardized feature vector. Finally, the standardized feature vector is subjected to a wavelet transform, and the Daubechies wavelet basis function is used for multi-scale decomposition to extract the time-frequency features at different scales, and finally the human behavior feature vector is obtained.
[0048] Taking supermarket surveillance as an example, identifying customer shopping behavior, surveillance video was sampled at 25 frames per second with a resolution of 1920 × 1080 pixels. Gabor filters were used to extract the customer's clothing texture and body features, generating 32 feature maps, each 240 × 135 pixels. Optical flow was used to calculate the velocity field, showing that the customer moved between shelves at a speed of 0.5 m / s. The aligned feature matrix had dimensions of 240 × 135 × 64. PCA dimensionality reduction was performed to retain the first 20 principal components, reducing the feature dimensions to 240 × 135 × 20. Saliency calculations revealed that the customer's hands and shopping cart areas had high weights, exceeding 0.8. Feature correlation analysis identified 12 core features, and Fourier transforms retained frequency components with 80% energy. This yielded a 256-dimensional human behavior feature vector for subsequent behavior recognition analysis.
[0049] In a specific embodiment, the process of executing step S102 may specifically include the following steps:
[0050] (1) performing human body region segmentation processing on the human behavior feature vector to obtain human body region mask data, and performing bounding box positioning processing on the human body region mask data to obtain human body region bounding box data;
[0051] (2) Generate candidate skeleton nodes for the human body region bounding box data to obtain a skeleton node candidate set, and perform node screening on the skeleton node candidate set through geometric constraints to obtain the coordinates of key skeleton nodes;
[0052] (3) The coordinates of key skeletal nodes are processed according to human anatomical constraints to obtain an anatomical node relationship diagram, and the anatomical node relationship diagram is processed to construct joint connection relationships to obtain skeletal joint connection data;
[0053] (4) performing joint angle calculation processing on the skeletal joint connection data to obtain a joint angle sequence, and performing angle range verification processing on the joint angle sequence to obtain valid joint angle data;
[0054] (5) performing motion trajectory construction processing on the effective joint angle data to obtain a joint motion trajectory set, and performing action node extraction processing on the joint motion trajectory set to obtain action key node data;
[0055] (6) Performing temporal alignment processing on the key node data of the action to obtain an aligned skeleton sequence, and performing skeleton feature extraction processing on the aligned skeleton sequence to obtain human body action skeleton sequence data.
[0056] Specifically, human body region segmentation is performed, and the image is classified at the pixel level using the semantic segmentation method. The image is divided into the human body region and the background region, and binary human body region mask data is generated. In the mask data, the pixel value of the human body region is 1, and the pixel value of the background region is 0. Subsequently, the mask data is subjected to connected region analysis and boundary extraction, and the minimum circumscribed rectangle is calculated to obtain the bounding box data containing the human body, and the upper left corner coordinates and lower right corner coordinates of the bounding box are recorded. Within the obtained human body region bounding box, a hierarchical feature point detection method is used to generate skeleton node candidate points. First, grid points are uniformly sampled and generated within the bounding box area with a grid spacing of 8 pixels. Then, the local feature response value of each grid point is calculated, and the points with higher response values are selected as candidate points. The candidate point set is screened by geometric constraints. The geometric constraints include distance constraints and angle constraints between point pairs. The key node coordinates that conform to the human skeletal structure are obtained through threshold screening.
[0057] Anatomical constraints are applied to the selected skeletal node coordinates. Based on the physiological structural characteristics of the human skeleton, the spatial positional relationships and motion limits of 17 key nodes are defined. Anatomical constraint rules are used to construct a node relationship graph, where nodes represent key skeletal points and edges represent the anatomical connections between nodes. Based on this relationship graph, skeletal joint connection data is further constructed, recording the connection properties and degrees of freedom of motion of each pair of connected joints.
[0058] Calculate the joint angles based on the skeletal joint connection data. The angle formed by three adjacent key points is calculated as follows:
[0059]
[0060] Among them, P1, P2, and P3 represent the coordinates of three adjacent key points. and Represents a vector, and θ is the joint angle. Calculate the three-dimensional Euler angle of each joint:
[0061]
[0062] Among them, a, β, and γ represent the rotation angles around the x, y, and z axes respectively. (n x , n y , n z ) represents the normal vector of the joint coordinate system, (a x , a y ) represents the projection component of the axis vector of the joint coordinate system onto the xy plane. The angle validity is verified based on the physiological range of motion of the human joint, and valid joint angle data is screened.
[0063] Joint motion trajectories are constructed based on effective joint angle data. A cubic spline interpolation method is used to fit the motion paths of joint points in a time series, generating a continuous trajectory curve. Key points are extracted from the trajectory curve, and points with significant curvature changes are selected as key motion nodes. The spatiotemporal coordinates and motion state parameters of these key nodes are recorded. Finally, the key motion node data is time-series aligned, and a dynamic time warping algorithm is used to standardize motion sequences of varying lengths, uniformly sampling them to the same time length to obtain an aligned skeleton sequence. By extracting the spatial structural features and temporal variation characteristics of the aligned skeleton sequence, skeleton sequence data describing human motion is ultimately generated.
[0064] Taking the analysis of an athlete's long jump as an example, the athlete's image is first segmented to obtain a human region mask with bounding box coordinates of [(150, 100), (450, 500)]. 85 candidate points are detected within the bounding box, and 17 key skeletal points are retained after geometric constraint screening. Joint connections are constructed based on anatomical constraints, and 16 major joint connections are recorded. The calculated hip joint angle sequence varies from 120 to 75 degrees during the takeoff phase, and the knee joint angle varies from 165 to 90 degrees, both within the physiological range. Spline interpolation is used to generate a continuous trajectory of 25 frames over 0.5 seconds, extracting five key action nodes: the end of the run-up, takeoff, mid-air posture, landing preparation, and landing completion. After temporal alignment, a standardized 80-frame skeleton sequence is generated, fully capturing the changes in human posture during the long jump.
[0065] In a specific embodiment, the process of executing step S103 may specifically include the following steps:
[0066] (1) Performing skeleton node connection processing on the human body motion skeleton sequence data to obtain an initial motion connection graph, and performing edge weight calculation processing on the initial motion connection graph to obtain an motion connection graph;
[0067] (2) Performing action feature extraction on the nodes in the action connection graph to obtain a node feature matrix, and performing node similarity calculation on the node feature matrix to obtain a node similarity matrix;
[0068] (3) Performing time series sliding window processing on the node similarity matrix to obtain time series window data, and performing time series edge construction processing on the time series window data to obtain a time series correlation edge set;
[0069] (4) Perform edge weight distribution processing on the temporal correlation edge set to obtain the temporal correlation edge, and perform graph fusion processing on the temporal correlation edge and the action connection graph to obtain the spatiotemporal action graph;
[0070] (5) Performing action pattern clustering processing on the spatiotemporal action graph to obtain action pattern clusters, and performing paradigm feature extraction processing on the action pattern clusters to obtain action paradigm features;
[0071] (6) The action paradigm features are processed into behavioral pattern encoding to obtain a behavioral coding sequence, and the behavioral coding sequence is processed into feature extraction through paradigm analysis to obtain spatiotemporal behavioral features.
[0072] Specifically, the connection relationships between skeleton nodes are established. Based on the human anatomy, 17 key bone points are connected through 16 edges to form the initial action connection diagram. The edge connection relationships include the main bone connections such as the head to the neck, the neck to the torso, the torso to the left and right shoulders, the shoulder to the elbow, and the elbow to the wrist. The weight value of each edge is calculated using the following formula:
[0073]
[0074] Among them, w ij represents the weight of the edge between node i and node j, d ij represents the Euclidean distance between two nodes, v i and v j Representing the movement speed of the two nodes, a weighted action connection graph is obtained. Then, the action features of each node in the action connection graph are extracted. The node features include spatial position, movement speed, acceleration and other information. For node k, its feature vector is calculated as follows:
[0075]
[0076] Among them, (x k ,y k ) is the node space coordinate, Δxk and Δy k The node feature matrix is obtained by calculating the displacement increment and Δt the time interval. The feature similarity between nodes is calculated using the cosine similarity metric to construct the node similarity matrix. A temporal analysis is performed on the node similarity matrix, using a sliding window size of 25 frames and a step size of 5 frames. Local temporal features are extracted by sliding along the time dimension. Within each temporal window, the correlation of node motion patterns is analyzed and temporal correlation edges are established. Temporal correlation edges represent the dynamic relationship between nodes at different times and include attributes such as time span and motion correlation.
[0077] The resulting temporal correlation edge sets are weighted, with the weight values reflecting the strength of the temporal correlation. The temporal correlation edges are fused with the original action connection graph to construct a spatiotemporal action graph that encompasses both spatial structure and temporal evolution. Each node in the spatiotemporal action graph contains both spatial location information and temporal variation characteristics. Action pattern clustering is performed based on the spatiotemporal action graph, using spectral clustering to group similar action patterns into the same category. Clustering considers the spatial distribution and temporal variation characteristics of nodes to obtain action pattern clusters with similar motion patterns. Typical features are extracted from each pattern cluster, including the spatial configuration of key nodes and motion trajectories, to form action paradigm features. Action paradigm features describe the standard pattern of a certain type of action and are highly distinctive and representative.
[0078] The action paradigm features are encoded into a fixed-dimensional feature sequence, taking into account both spatial structure and temporal variation. Through paradigm analysis, key patterns in the sequence are extracted, ultimately yielding spatiotemporal behavioral features that characterize the complete behavior.
[0079] Taking basketball shooting action analysis as an example, we first construct a skeleton connection graph containing 17 nodes. The calculated edge weight from the shoulder joint to the elbow joint is 0.85, and the edge weight from the elbow joint to the wrist is 0.78. When extracting node features, the displacement of the right wrist node in 0.5 seconds is (0.8m, 1.2m), the velocity is (1.6m / s, 2.4m / s), and the acceleration is (3.2m / s 2 ,4.8m / s 2 ). Using a 25-frame sliding window to analyze the temporal features, it was found that the movements of the three arm joints (shoulder, elbow, and wrist) were highly correlated, with a similarity of 0.92. A complete action graph was obtained through spatiotemporal fusion, containing 80 spatial edges and 120 temporal edges. Cluster analysis revealed three main action pattern clusters: the preparation phase, the force phase, and the follow-up phase. The paradigm features of each phase were extracted, such as the arm joint angle change range of 45°-170° in the force phase and the movement duration of 0.3 seconds. Finally, a 384-dimensional spatiotemporal behavior feature vector was generated for subsequent behavior recognition analysis.
[0080] In a specific embodiment, the process of executing step S104 may specifically include the following steps:
[0081] (1) Performing action semantic layering processing on spatiotemporal behavior features and human behavior feature vectors to obtain layered action semantic elements, and then performing basic action deconstruction processing on the layered action semantic elements to obtain a basic action library;
[0082] (2) Perform scene-action matching on the basic action library to obtain scene-related action sequences, and perform action correlation deduction on the scene-related action sequences to obtain an action correlation feature matrix;
[0083] (3) Perform action sequence modeling on the action association feature matrix to obtain an action dependency network, and perform semantic reasoning on the action dependency network to obtain action intention features;
[0084] (4) Mining the temporal regularity of action intention features to obtain an action rule set, and performing behavioral semantic mapping on the action rule set to obtain a behavioral semantic mapping graph;
[0085] (5) Performing hierarchical semantic abstraction processing on the behavior semantic mapping graph to obtain a hierarchical behavior description, and then performing action constraint construction processing on the hierarchical behavior description to obtain a behavior constraint network;
[0086] (6) Perform spatiotemporal correlation verification on the behavior constraint network to obtain a verification behavior sequence, and perform semantic feature aggregation on the verification behavior sequence to obtain behavior semantic description features.
[0087] Specifically, a three-layer semantic structure is constructed: the bottom layer contains basic action units, such as raising a hand, turning around, and bending over; the middle layer contains combined action sequences, such as picking up an item and examining an item; and the top layer contains complete behavioral semantics, such as shopping and patrolling. The input features are decomposed according to the semantic hierarchy to obtain hierarchical action semantic elements. Then, the basic action deconstruction is performed on each layer of semantic elements, extracting the core components and temporal relationships of the actions, and constructing a standardized basic action library. Based on the constructed basic action library, scene action matching is performed, and the standard action units in the action library are compared with the action sequences in the actual scene. By calculating the similarity and temporal consistency of the action features, the action sequences related to the current scene are identified. The matched scene-related action sequences are subjected to correlation analysis, and the temporal dependency and semantic correlation between the actions are calculated to generate an action correlation feature matrix. Each element in the matrix represents the strength of the correlation between two actions.
[0088] Sequence modeling is performed on the action-association feature matrix to construct an action dependency network in the form of a directed acyclic graph. Nodes in the network represent basic action units, edges represent dependencies between actions, and edge weights reflect the strength of dependencies. The network structure is used to analyze the combinatorial patterns and semantic dependencies of actions. Semantic reasoning is performed on the dependency network, combined with scene context information, to infer the underlying intent of the action sequence and generate feature vectors describing the purpose and intent of the action. Temporal patterns are mined on the action intention features to analyze the temporal variation of the action sequence. By identifying recurring action combinations, action transition probabilities, and other information, a set of rules describing the temporal patterns of actions is derived. This set of rules is mapped into a semantic space, establishing a correspondence between action rules and behavioral semantics, and constructing a behavioral semantic mapping graph.
[0089] The resulting behavioral semantic mapping graph is hierarchically abstracted, elevating specific action rules to a higher-level behavioral semantic description. Through semantic aggregation and generalization, a multi-level behavioral description system is formed, including semantic expressions at the action, intention, and behavior levels. An action constraint network is constructed based on the hierarchical behavioral description, defining the temporal, spatial, and semantic constraints between actions. Finally, the behavior constraint network is verified for spatiotemporal correlation to check whether the action sequence satisfies the various constraints. This verification screens out valid behavior sequences that meet the constraints, and these sequences are subjected to semantic feature aggregation to extract common and distinguishing features, ultimately generating a complete behavioral semantic description feature.
[0090] Taking abnormal behavior identification during surveillance as an example, the input action sequence consists of basic actions such as "walking quickly - looking around frequently - pausing to observe - quickly leaving." Semantic layering deconstructs this into underlying action units: walking speed of 2.5 m / s, head turning frequency of 2 times / second, and pausing duration of 15 seconds. During the scene matching phase, comparison with a database of normal shopping behaviors revealed an abnormality level of 0.85 for this action combination. Action correlation analysis revealed a strong correlation of 0.92 between frequent looking and pausing to observe. The dependency network revealed that this sequence exhibited clear surveillance characteristics, suggesting the action intent was based on scene exploration. Temporal pattern analysis revealed that a similar action pattern recurred four times within three minutes, demonstrating a periodic pattern. After semantic mapping and constraint verification, this behavior sequence was labeled "suspicious site reconnaissance" with an abnormality level of 0.88, triggering a security alert.
[0091] In a specific embodiment, the process of executing step S105 may specifically include the following steps:
[0092] (1) Performing behavior type encoding on the behavior semantic description features to obtain a behavior encoding vector, and then performing classification standard matching on the behavior encoding vector to obtain behavior category features;
[0093] (2) Perform standard behavior pattern extraction on the behavior category features to obtain a behavior pattern library, and perform similarity calculation on the behavior category features and the behavior pattern library to obtain a behavior similarity matrix;
[0094] (3) Perform threshold segmentation on the behavior similarity matrix to obtain similar behavior groups, and calculate the intra-group differences of similar behavior groups to obtain the behavior deviation index;
[0095] (4) performing abnormality evaluation on the behavioral deviation indicators to obtain a behavioral abnormality measurement value, and performing grade classification on the behavioral abnormality measurement value to obtain an abnormality degree grade;
[0096] (5) performing feature combination processing on the abnormality level and the behavior category features to obtain a behavior feature set, and performing category determination processing on the behavior feature set to obtain a behavior type identifier;
[0097] (6) Statistical analysis is performed on the behavior type identification and abnormality level to obtain the behavior statistical characteristics, and the behavior statistical characteristics are quantitatively processed to obtain the behavior type identification and abnormality level index.
[0098] Specifically, the input behavior semantic description features are first encoded into behavior types. The encoding process uses a one-hot encoding method to map different behavior types into binary vectors of fixed dimensions, with each position representing a basic behavior pattern. For example, in a supermarket monitoring scenario, the encoding dimension is 32, where bits 1-8 represent movement behaviors (such as walking, running, and staying), bits 9-16 represent interaction behaviors (such as picking up and putting back items), bits 17-24 represent posture behaviors (such as bending, reaching, and looking), and bits 25-32 represent social behaviors (such as talking, following, and gathering). The generated behavior encoding vector is matched with predefined classification criteria to obtain behavior category features. Subsequently, standard behavior patterns are extracted from the behavior category features to construct a behavior library containing typical behavior patterns. Each pattern in the behavior pattern library contains descriptions of three dimensions: temporal features, spatial features, and semantic features. Calculate the similarity between the behavior category features to be analyzed and the standard patterns in the pattern library, use the cosine distance to measure the similarity between the feature vectors, and generate an NxM-dimensional behavior similarity matrix, where N is the number of behaviors to be analyzed and M is the number of standard patterns.
[0099] The behavioral similarity matrix is segmented by an adaptive threshold, with the threshold set to the interval of the similarity mean plus or minus the standard deviation. Behaviors with similarity above the threshold are clustered into the same group to form similar behavior groups. For each similar behavior group, the differences between behaviors within the group are calculated, including temporal differences, spatial differences, and semantic differences, and a behavioral deviation index is obtained. The deviation index reflects the degree of abnormality of the behavior relative to the standard pattern. The abnormality is assessed based on the behavioral deviation index. The deviation index is mapped to the range of 0-1 to obtain a standardized behavioral abnormality measurement value. According to the distribution characteristics of the abnormality measurement value, the abnormality degree is divided into five levels: normal (0-0.2), mild abnormality (0.2-0.4), moderate abnormality (0.4-0.6), severe abnormality (0.6-0.8), and extreme abnormality (0.8-1.0).
[0100] The anomaly level and behavior category features are combined to construct a multidimensional behavioral feature set. This feature set includes behavior type information, anomaly level information, temporal variation information, and scenario context information. The feature set is analyzed using pre-set judgment rules to determine the final behavior type identification. The behavior type identification includes the specific behavior category and anomaly attribute tags. Finally, a statistical analysis is performed on the behavior type identification and anomaly level, calculating statistical characteristics such as the frequency, duration, and transition probability of different behavior types. These statistical characteristics are quantified to generate a standardized indicator system, ultimately outputting the behavior type identification and anomaly level indicator.
[0101] Taking the identification of abnormal behavior in a bank branch as an example, a surveillance system captures a sequence of individual behaviors. The behavior is first encoded, generating a 32-dimensional binary vector in which the presence of 1s in the resting position (bit 3), frequent looking around (bit 23), and solitude (bit 28) indicates the presence of these basic behaviors. Comparing this with a standard behavioral pattern database, a similarity matrix is calculated. The similarity between this behavior and normal queuing patterns is 0.35, and the similarity with suspicious loitering patterns is 0.82. Using a similarity threshold of 0.65, the individual is segmented and classified as a suspicious behavior group. Computing intra-group differences reveals that this behavior deviates significantly from the normal pattern in dwell time (over 30 minutes) and frequency of looking around (an average of 6 times per minute), resulting in a behavioral deviation index of 0.78. An abnormality assessment identifies this behavior as severe. Combined with behavioral feature analysis, this behavior is identified as a "suspicious scouting" type. The resulting abnormality index is 0.78, triggering a level 2 alert for the security system.
[0102] In a specific embodiment, the process of executing step S106 may specifically include the following steps:
[0103] (1) Perform time series window segmentation processing on the behavior type identification and abnormality degree index to obtain a time series analysis sequence, and perform behavior duration statistics processing on the time series analysis sequence to obtain the behavior duration feature;
[0104] (2) Analyze and process the behavior evolution trend of the behavior persistence characteristics to obtain behavior trend data, and perform event development prediction processing on the behavior trend data to obtain the behavior development sequence;
[0105] (3) Conduct risk assessment on the behavior development sequence to obtain risk assessment indicators, and then classify the risk assessment indicators into risk levels to obtain risk level data;
[0106] (4) Perform scenario-related analysis on the risk level data to obtain scenario risk characteristics, and perform risk propagation path analysis on the scenario risk characteristics to obtain a risk propagation map;
[0107] (5) Performing early warning level determination processing on the risk propagation map to obtain early warning level data, and performing early warning information generation processing on the early warning level data to obtain an early warning message set;
[0108] (6) Perform information aggregation and processing on the warning message set and behavior development sequence to obtain behavior recognition results and warning information.
[0109] Specifically, in the process of processing behavior type identification and abnormality level indicators, a time series window segmentation method is first used, setting the window size to 30 seconds and the sliding step size to 10 seconds to segment the continuous behavior sequence. Statistical analysis is performed on the behavioral data within each time window, recording the distribution of behavior types and the changing trend of abnormality levels. By cumulatively counting the duration, frequency of occurrence, and transition patterns of each behavior type, a feature vector describing the persistence characteristics of the behavior is obtained. Trend analysis is performed based on the behavioral persistence characteristics, and the changing patterns of the feature values in the time series are statistically analyzed. The analysis includes the evolution trend of the behavior type, the progressive relationship of the abnormality level, and the transition characteristics of the behavior pattern. A time series forecasting method is used for the trend data to predict the development trend of behavior in the future. The forecast content includes the probability of change of the behavior type, the growth rate of the abnormality level, and the potential behavior transition points.
[0110] The predicted behavior development sequence is evaluated for its riskiness. The evaluation metrics include three dimensions: harmfulness, scope of impact, and urgency. Harmfulness is quantified based on the destructive power and threat level of the behavior. Scope of impact considers the spatial scale and number of people affected by the behavior. Urgency is based on the time urgency of the situation. Scoring across these three dimensions is combined to generate a standardized risk assessment index. Based on the numerical distribution of the assessment indicators, the risk level is categorized into five levels: slight risk (0-0.2), low risk (0.2-0.4), moderate risk (0.4-0.6), high risk (0.6-0.8), and extreme risk (0.8-1.0). Risk level data is correlated with specific scenarios for analysis. Considering characteristics such as scenario importance, vulnerability, and protection capabilities, the actual impact of risks in different scenarios is assessed. For example, the same risk level in a densely populated area poses a higher threat than in a remote area. Based on the results of the scenario correlation analysis, a risk propagation path map is constructed to depict the risk diffusion process across space and time. Nodes in the map represent risk impact points, edges represent risk propagation channels, and weights reflect the probability of propagation.
[0111] Based on the risk propagation map, the warning level is determined and set into four levels: blue warning (minor threat), yellow warning (moderate threat), orange warning (severe threat), and red warning (extremely severe threat). Corresponding measures and response processes are defined for each warning level, generating a warning message collection containing a description of the threat, the scope of impact, and recommended measures. Finally, a comprehensive analysis and information aggregation of the warning message collection and the behavioral development sequence is conducted to form a complete behavior identification result and warning information. The summary includes the determination of the behavior type, the assessment of the degree of abnormality, the determination of the risk level, and the recommended warning measures.
[0112] Taking the security monitoring system at a large shopping mall as an example, the surveillance system detected a sequence of suspicious individuals' behaviors. Within a 30-second window, three main behaviors were recorded: repeated patrolling (45%), scattered wandering (35%), and whispered conversations (20%). Behavioral persistence revealed that the patrolling behavior was repeated six times within 15 minutes, each lasting 2-3 minutes. Trend analysis predicted that the group's activities would expand into the valuables area, with the level of abnormality increasing by 0.3 levels within 15 minutes. A risk assessment indicated a current risk index of 0.75, placing it at a high risk level. Scenario analysis identified the jewelry section of the mall, which has a high customer traffic density and a high level of security. Risk propagation analysis indicated that the risk would extend throughout the jewelry area within 5 minutes. Based on this analysis, an orange alert was triggered, generating the warning message: "A suspicious group of multiple individuals has been detected in the jewelry area, displaying clear scouting patterns. Security personnel are advised to increase patrol frequency and prioritize deployment in key areas."
[0113] The above describes the method for identifying the behavior of a person under video surveillance in the embodiment of the present application. The following describes the system for identifying the behavior of a person under video surveillance in the embodiment of the present application. Figure 2 In one embodiment of the present application, a system for identifying human behavior using video surveillance includes:
[0114] The acquisition module 201 is used to extract the human action texture features and motion trajectory features from the input surveillance video sequence, perform feature fusion processing through behavior saliency analysis, and obtain the human action feature vector;
[0115] A positioning module 202 is used to perform skeleton joint positioning processing on the human behavior feature vector, and extract action nodes and joint angles from the positioning area to obtain human action skeleton sequence data;
[0116] A construction module 203 is used to construct an action connection graph and temporal correlation edges for the human action skeleton sequence data, and perform behavior pattern learning processing through motion paradigm analysis to obtain spatiotemporal behavior features;
[0117] A mapping module 204 is configured to perform action semantic mapping processing on the spatiotemporal behavior features and the human behavior feature vectors, and obtain behavior semantic description features through behavior time sequence chain construction processing;
[0118] The measurement module 205 is used to perform behavior pattern matching and anomaly measurement processing on the behavior semantic description features, perform feature discrimination processing by behavior similarity comparison, and obtain a behavior type identifier and an anomaly degree index;
[0119] The analysis module 206 is used to perform behavior persistence analysis on the behavior type identifier and abnormality index, and obtain behavior recognition results and warning information through behavior risk level assessment.
[0120] Through the collaborative efforts of the aforementioned components, the system extracts texture features and motion trajectory features from surveillance video sequences and integrates them with behavioral saliency analysis to achieve feature fusion, effectively capturing the appearance and motion information of human behavior and improving the integrity of feature representation. By locating skeletal joints and extracting action nodes from human behavior feature vectors, key human body parts are accurately located, enabling a precise description of human posture. Based on human motion skeleton sequence data, an action connection graph and temporal correlation edges are constructed. Combined with motion paradigm analysis, behavioral pattern learning is performed to establish a spatiotemporal structural representation of behavior, enhancing the ability to model complex behavioral patterns. Action semantics are mapped between spatiotemporal behavioral features and human behavior feature vectors. By constructing a behavioral temporal chain, an effective mapping from low-level features to high-level semantics is achieved, improving the accuracy of behavior understanding. Based on the behavioral semantic description features, behavioral pattern matching and anomaly measurement are performed. Feature discrimination is performed through behavioral similarity comparison, effectively identifying abnormal behavior patterns and improving the sensitivity of abnormal behavior detection. Finally, continuous analysis and risk assessment of behavior type identification and anomaly severity indicators are performed, enabling timely early warning of abnormal behavior and ensuring the safety of the surveillance scene.
[0121] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for identifying human behavior in video surveillance, characterized in that: The method for identifying human behavior through video surveillance includes: Extract human action texture features and motion trajectory features from the input surveillance video sequence, perform feature fusion processing through behavior saliency analysis, and obtain human action feature vectors; Performing skeleton joint positioning processing on the human behavior feature vector, and extracting action nodes and joint angles from the positioning area to obtain human action skeleton sequence data; Constructing an action connection graph and temporal correlation edges for the human action skeleton sequence data, performing behavior pattern learning processing through motion paradigm analysis, and obtaining spatiotemporal behavior characteristics; Performing action semantic mapping processing on the spatiotemporal behavior features and the human behavior feature vectors, and constructing and processing the behavior time sequence chain to obtain behavior semantic description features; Performing behavior pattern matching and anomaly measurement processing on the behavior semantic description features, performing feature discrimination processing by behavior similarity comparison, and obtaining a behavior type identifier and an anomaly degree index; Conducting behavioral persistence analysis on the behavior type identifier and abnormality index, and performing behavioral risk level assessment to obtain behavior recognition results and warning information; The step of constructing an action connection graph and temporal correlation edges for the human action skeleton sequence data, performing behavior pattern learning processing through motion paradigm analysis, and obtaining spatiotemporal behavior features includes: Performing skeleton node connection processing on the human body motion skeleton sequence data to obtain an initial motion connection graph, and performing edge weight calculation processing on the initial motion connection graph to obtain an motion connection graph; Performing action feature extraction processing on the nodes in the action connection graph to obtain a node feature matrix, and performing node similarity calculation processing on the node feature matrix to obtain a node similarity matrix; Performing time series sliding window processing on the node similarity matrix to obtain time series window data, and performing time series edge construction processing on the time series window data to obtain a time series associated edge set; Performing edge weight distribution processing on the time series associated edge set to obtain time series associated edges, and performing graph fusion processing on the time series associated edges and the action connection graph to obtain a spatiotemporal action graph; Performing action pattern clustering processing on the spatiotemporal action graph to obtain action pattern clusters, and performing paradigm feature extraction processing on the action pattern clusters to obtain action paradigm features; Performing behavior pattern coding processing on the action paradigm features to obtain a behavior coding sequence, and performing feature extraction processing on the behavior coding sequence through paradigm analysis to obtain spatiotemporal behavior features; The step of performing action semantic mapping on the spatiotemporal behavior features and the human behavior feature vectors and constructing a behavior time sequence chain to obtain behavior semantic description features includes: Performing action semantic layering processing on the spatiotemporal behavior features and the human behavior feature vectors to obtain layered action semantic elements, and performing basic action deconstruction processing on the layered action semantic elements to obtain a basic action library; Performing scene-action matching processing on the basic action library to obtain a scene-related action sequence, and performing action correlation deduction processing on the scene-related action sequence to obtain an action correlation feature matrix; Performing action sequence modeling processing on the action association feature matrix to obtain an action dependency network, and performing semantic reasoning processing on the action dependency network to obtain action intention features; Performing temporal regularity mining on the action intention features to obtain an action rule set, and performing behavior semantic mapping on the action rule set to obtain a behavior semantic mapping graph; Performing hierarchical semantic abstraction processing on the behavior semantic mapping graph to obtain a hierarchical behavior description, and performing action constraint construction processing on the hierarchical behavior description to obtain a behavior constraint network; A spatiotemporal correlation verification process is performed on the behavior constraint network to obtain a verification behavior sequence, and a semantic feature aggregation process is performed on the verification behavior sequence to obtain a behavior semantic description feature.
2. The method for identifying human behavior in video surveillance according to claim 1, characterized in that: The human action texture features and motion trajectory features are extracted from the input surveillance video sequence, and feature fusion processing is performed through behavior saliency analysis to obtain a human action feature vector, including: Performing frame sampling processing on the input surveillance video sequence to obtain video frame sequence data, and performing grayscale and normalization processing on the video frame sequence data to obtain preprocessed image data; Performing texture feature extraction processing on the pre-processed image data through a Gabor filter to obtain a human motion texture feature matrix, and performing motion vector calculation processing on the pre-processed image data through an optical flow method to obtain a human motion trajectory feature matrix; Performing feature dimension alignment processing on the human motion texture feature matrix and the human motion trajectory feature matrix to obtain an aligned feature matrix, and performing principal component analysis processing on the aligned feature matrix to obtain dimension-reduced feature data; Performing regional weight calculation processing on the reduced-dimensionality feature data through saliency calculation to obtain a feature weight matrix, and performing weighted combination processing on the reduced-dimensionality feature data and the feature weight matrix to obtain an initial fusion feature; Performing feature correlation analysis on the initial fusion features through a covariance matrix to obtain a feature correlation matrix, and performing feature selection on the feature correlation matrix to obtain a core feature set; Performing frequency domain feature extraction processing on the core feature set through Fourier transform to obtain a frequency domain feature vector, and performing frequency domain feature screening processing on the frequency domain feature vector to obtain a frequency domain feature set; Performing feature splicing processing on the frequency domain feature set and the core feature set to obtain a combined feature vector, and performing feature normalization processing on the combined feature vector to obtain a normalized feature vector; The standardized feature vector is subjected to time-frequency feature analysis processing by wavelet transform to obtain time-frequency feature data, and the time-frequency feature data is subjected to multi-scale feature extraction processing to obtain a human behavior feature vector.
3. The method for identifying human behavior in video surveillance according to claim 1, characterized in that: The method of performing skeleton joint positioning processing on the human behavior feature vector and extracting action nodes and joint angles from the positioning area to obtain human action skeleton sequence data includes: Performing human body region segmentation processing on the human behavior feature vector to obtain human body region mask data, and performing bounding box positioning processing on the human body region mask data to obtain human body region bounding box data; Performing skeleton node candidate point generation processing on the human body region bounding box data to obtain a skeleton node candidate set, and performing node screening processing on the skeleton node candidate set through geometric constraints to obtain key skeleton node coordinates; Performing human anatomical constraint processing on the key skeletal node coordinates to obtain an anatomical node relationship diagram, and performing joint connection relationship construction processing on the anatomical node relationship diagram to obtain skeletal joint connection data; Performing joint angle calculation processing on the skeletal joint connection data to obtain a joint angle sequence, and performing angle range verification processing on the joint angle sequence to obtain valid joint angle data; Performing motion trajectory construction processing on the effective joint angle data to obtain a joint motion trajectory set, and performing action node extraction processing on the joint motion trajectory set to obtain action key node data; The action key node data is subjected to time sequence alignment processing to obtain an aligned skeleton sequence, and the aligned skeleton sequence is subjected to skeleton feature extraction processing to obtain human body action skeleton sequence data.
4. The method for identifying human behavior in video surveillance according to claim 1, characterized in that: The behavioral semantic description features are subjected to behavioral pattern matching and anomaly measurement processing, and feature discrimination processing is performed through behavioral similarity comparison to obtain a behavior type identifier and an anomaly degree index, including: Performing behavior type coding processing on the behavior semantic description feature to obtain a behavior coding vector, and performing classification standard matching processing on the behavior coding vector to obtain a behavior category feature; Performing standard behavior pattern extraction processing on the behavior category features to obtain a behavior pattern library, and performing similarity calculation processing on the behavior category features and the behavior pattern library to obtain a behavior similarity matrix; Performing threshold segmentation processing on the behavior similarity matrix to obtain similar behavior groups, and performing intra-group difference calculation processing on the similar behavior groups to obtain a behavior deviation index; Performing abnormality evaluation processing on the behavioral deviation indicator to obtain a behavioral abnormality measurement value, and performing grade classification processing on the behavioral abnormality measurement value to obtain an abnormality degree grade; Performing feature combination processing on the abnormality level and the behavior category feature to obtain a behavior feature set, and performing category determination processing on the behavior feature set to obtain a behavior type identifier; The behavior type identifier and the abnormality level are statistically analyzed to obtain behavior statistical characteristics, and the behavior statistical characteristics are quantitatively analyzed to obtain behavior type identifier and abnormality level index.
5. The method for identifying human behavior in video surveillance according to claim 1, characterized in that: The behavior type identification and abnormality degree index are subjected to behavior persistence analysis and processing, and the behavior risk level assessment is performed to obtain behavior recognition results and warning information, including: Performing time series window segmentation processing on the behavior type identifier and abnormality degree index to obtain a time series analysis sequence, and performing behavior duration statistics processing on the time series analysis sequence to obtain a behavior duration feature; Performing behavior evolution trend analysis on the behavior persistence characteristics to obtain behavior trend data, and performing event development prediction processing on the behavior trend data to obtain a behavior development sequence; Performing a risk assessment process on the behavior development sequence to obtain a risk assessment index, and performing a risk level classification process on the risk assessment index to obtain risk level data; Performing scenario association analysis on the risk level data to obtain scenario risk characteristics, and performing risk propagation path analysis on the scenario risk characteristics to obtain a risk propagation map; Performing warning level determination processing on the risk propagation map to obtain warning level data, and performing warning information generation processing on the warning level data to obtain a warning message set; The warning message set and behavior development sequence are aggregated and processed to obtain behavior recognition results and warning information.
6. A system for identifying human behavior in video surveillance, for implementing the method for identifying human behavior in video surveillance according to any one of claims 1 to 5, characterized in that: The personnel behavior recognition system of the video surveillance includes: The acquisition module is used to extract the texture features and motion trajectory features of human actions from the input surveillance video sequence, perform feature fusion processing through behavior saliency analysis, and obtain the human behavior feature vector; A positioning module is used to perform skeleton joint positioning processing on the human behavior feature vector, and to extract action nodes and joint angles from the positioning area to obtain human action skeleton sequence data; A construction module is used to construct an action connection graph and a temporal correlation edge for the human action skeleton sequence data, and perform behavior pattern learning processing through motion paradigm analysis to obtain spatiotemporal behavior characteristics; A mapping module is used to perform action semantic mapping processing on the spatiotemporal behavior features and the human behavior feature vectors, and obtain behavior semantic description features through behavior time sequence chain construction processing; A measurement module is used to perform behavior pattern matching and anomaly measurement processing on the behavior semantic description features, perform feature discrimination processing by behavior similarity comparison, and obtain a behavior type identifier and an anomaly degree index; The analysis module is used to perform behavior persistence analysis on the behavior type identification and abnormality degree index, and obtain behavior recognition results and early warning information through behavior risk level assessment.
Citation Information
Patent Citations
Human body abnormal behavior recognition method under transformer substation video monitoring based on posture estimation
CN116912930A
Human body action recognition method, human body action recognition system, and device
WO2022000420A1