Two-stage limb violent behavior recognition method based on improved OpenPose and space-time diagram convolution ST-GCN

By improving the two-stage method of OpenPose and spatiotemporal graph convolution ST-GCN, the static limitations and real-time problems in physical violence behavior recognition are solved, and high-precision and efficient physical violence behavior recognition is achieved.

CN120833632APending Publication Date: 2025-10-24ZHEJIANG NORMAL UNIV +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510923204.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-10-24

AI Technical Summary

Technical Problem

Existing deep learning technology has the problems of static framed images, large fluctuations in human joints, and high real-time requirements in the identification of physical violence, making it difficult to achieve accurate monitoring around the clock.

Method used

The improved OpenPose algorithm is used to extract key points of the human skeleton, and the spatiotemporal graph convolution ST-GCN structure is combined for spatiotemporal feature analysis. The lightweight MobileNet-V3 network and the Hungarian algorithm are used to optimize key point matching. The five-partition strategy is introduced to optimize the graph convolution operation and enhance the modeling ability of the motion patterns of human body parts.

Benefits of technology

The accuracy and robustness of physical violence behavior recognition are improved, the semantic ambiguity of actions and joint positioning errors are successfully alleviated, and real-time and recognition capabilities are guaranteed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120833632A_ABST
    Figure CN120833632A_ABST
Patent Text Reader

Abstract

The invention discloses a two-stage limb violent behavior recognition method based on improved OpenPose and space-time diagram convolution ST-GCN, which is applied to the technical field of limb violent behavior recognition and comprises the following steps: extracting human skeleton key point time sequence data based on an improved OpenPose algorithm; wherein the improvement comprises the steps of extracting human skeleton key points based on MobileNet-V3 and carrying out global optimal matching on the human skeleton key points based on a Hungary algorithm; based on a space-time diagram convolution ST-GCN structure, analyzing and extracting space-time features of the time sequence data of the key points of the human skeleton, and outputting a limb violent behavior recognition result; wherein the space-time diagram convolution ST-GCN structure is formed by stacking a plurality of ST-GCN modules, and each ST-GCN module is composed of an attention layer, a space diagram convolution and a time diagram convolution. According to the method, the limb violent behavior identification precision is effectively improved, and the real-time performance is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of limb violence behavior recognition, and particularly relates to a two-stage limb violence behavior recognition method based on improved OpenPose and spatio-temporal graph convolution ST-GCN. BACKGROUND

[0002] Public security guarantee mechanism has an important strategic position in maintaining social order and citizen welfare. As an abnormal behavior, limb violence not only infringes on individual rights, but also may trigger regional security crisis, especially in the value cognition of juvenile groups in the scene. At present, video monitoring is widely used in the field of security and protection, but traditional manual monitoring is limited by visual fatigue and scene coverage, and it is difficult to achieve all-weather accurate monitoring. With the help of deep learning technology for intelligent analysis, automatic recognition and real-time warning of limb violence behavior can be realized, so as to improve the response efficiency, reduce the dependence on artificial, and enhance the intelligent level of security and protection system.

[0003] Although deep learning has made progress in behavior recognition, limb violence behavior recognition still faces problems such as static limitations of frame images, large variation range of human joint points, and high real-time requirements of actual deployment.

[0004] Therefore, how to provide a two-stage limb violence behavior recognition method based on improved OpenPose and spatio-temporal graph convolution ST-GCN to effectively solve the above problems of deep learning in the field of limb violence behavior recognition is a problem that those skilled in the art need to solve. SUMMARY

[0005] Therefore, the present application provides a two-stage limb violence behavior recognition method based on improved OpenPose and spatio-temporal graph convolution ST-GCN.

[0006] In order to achieve the above purpose, the present application adopts the following technical solutions:

[0007] A two-stage limb violence behavior recognition method based on improved OpenPose and spatio-temporal graph convolution ST-GCN, comprising:

[0008] Step 1: based on the improved OpenPose algorithm, extracting human skeleton key point time series data; wherein the improvement of the OpenPose algorithm includes: extracting human skeleton key points based on MobileNet-V3 and globally optimally matching human skeleton key points based on the Hungarian algorithm;

[0009] Step 2: Based on the spatio-temporal graph convolution ST-GCN structure, the spatio-temporal features of the human body skeleton key point time series data are analyzed and extracted, and the limb violence behavior recognition result is output; wherein the spatio-temporal graph convolution ST-GCN structure is stacked by multiple ST-GCN modules, and each ST-GCN module is composed of an attention layer, a spatial graph convolution and a temporal graph convolution.

[0010] Optionally, in step 1, MobileNet-V3 adopts a depth separable convolution technology, combines the use of an inverted residual structure and a linear bottleneck design, and introduces an SE-Net module for adaptive adjustment of model parameters, and uses an h-swish activation function instead of a traditional ReLU activation function.

[0011] Optionally, the depth separable convolution technology is specifically:

[0012] Three single-layer convolution kernels are used to perform convolution operations on the three channels of the input to generate feature maps consistent with the number of input layer channels;

[0013] A point convolution kernel with a size of 1x1xM is used to weight and combine the feature maps output in the previous step in the channel dimension to finally generate the output result; wherein M is the number of channels of the input image;

[0014] The convolution parameter scale of each layer is as follows:

[0015] S layer =S input ×S dwc ×S pwc ×N filter ;

[0016] Wherein, S layer is the parameter scale of the current layer, S input is the input feature map size, S dwc is the channel convolution kernel size, S pwc is the point convolution kernel size, and N filter is the set output feature quantity.

[0017] Optionally, the inverted residual structure and the linear bottleneck design are specifically:

[0018] First, dimensionality is increased through 1x1 convolution; then, 3x3 depth separable convolution is used for feature extraction; finally, 1x1 convolution is used again to reduce dimensionality and compress the number of channels, and a linear activation function is used instead of ReLU in the last dimension reduction.

[0019] Optionally, the h-swish activation function is as follows:

[0020]

[0021] Optionally, in step 1, the global optimal matching of human skeleton key points is performed based on the Hungarian algorithm, specifically:

[0022] Step1: Obtain two groups of key points to be matched by using the PCM algorithm, and calculate the affinity between all point pairs;

[0023] Step2: Screen the possible matching key point pairs according to the affinity threshold;

[0024] Step3: Select a key point from set 1 in turn, if it is taken, jump to Step5, otherwise continue to select the key point with the highest affinity from set 2 which has not been matched; if the key point has not been matched, directly connect and repeat Step3; if the key point has been matched, execute Step4; if there is no matching key point, skip the point and continue Step3;

[0025] Step4: Traverse all possible augmented paths and calculate their affinity sums Ci, select the augmented path corresponding to the maximum Ci as the optimal path, and return to Step3 after path adjustment;

[0026] Step5: The global optimal matching of key points is completed, and the matching process is ended.

[0027] Optionally, in step 2, before analyzing and extracting the space-time features of the human skeleton key point time series data based on the space-time graph convolution ST-GCN structure, it further includes: constructing a space-time graph, specifically:

[0028] For a skeleton sequence with N joints and T frames, a undirected space-time graph G=(V,E) is constructed; wherein V is the node set, V={v ti |t=1,…,T; i=1,…,N}, representing the positions of all joints in each time frame, as the input of ST-GCN, the feature vector on each node F(vti) is composed of the joint coordinate vector and its confidence estimate in frame t; the process of constructing the space-time graph on the skeleton sequence can be divided into two steps: one is spatial connection, within each frame, the joints are connected according to the physical structure of the human body; the second is time connection, the same position of each joint in consecutive frames is connected to form the connection in time sequence;

[0029] E is the edge set; in the graph structure, the edge set E is divided into two subsets, subset E s : E s =(v ti ,v tj ) | (i,j) ∈ H is used to describe the connection of the same frame, wherein H represents the set of joints that meet the natural connectivity of the human body; subset E F : E F =(vti ,v (t+1)i ) record the timing relationship between frames, that is, connect the nodes of the same joint in adjacent frames, and for each joint, the timing edge represents its trajectory over time;

[0030] Set the convolution kernel size as KxK, the input feature map as f in , and the number of channels as c, then the single-channel output value at the spatial position is as follows:

[0031]

[0032] wherein p is a sampling function for determining the feature points in the neighborhood of x participating in the convolution operation, for image convolution, the sampling function is expressed as: p(x, h, w)=x+p'(h, w); x is the current feature point; p'(h, w) is the offset relative to the current position (h, w) and represents the offset of the field position; w is a weight function as the core parameter of the convolution operation, which is used for inner product calculation with the c-dimensional input feature vector to complete feature extraction.

[0033] Optionally, in step 2, the calculation expression of the spatial graph convolution is as follows:

[0034]

[0035]

[0036] w(v ti ,v tj )=w'(l ti (v tj ));

[0037] Z ti (v tj )=|{v tk |l ti (v tk )=l ti (v ti )}|;

[0038] wherein f out (v ti ) is the feature output after the spatial graph convolution operation of the node v ti ; v ti is the i-th joint node of the t-th frame in the graph structure; B(v ti ) is the adjacent point set of the node v ti ; f in is an input feature function; p(v ti ,v tj ) is a sampling function; D is a coefficient for defining the neighborhood range; d(v tj ,v ti) is the shortest path from node v tj to v ti ; the first equation in the gathering function indicates that v tj will satisfy the condition |d(v ti ,v tj ) ≤ D as the adjacent node of v ti ; the second equation in the gathering function indicates that the adjacent node of v ti is itself when D = 1; w(v ti ,v tj ) is a weight function based on the neighborhood division; l ti is a mapping function, i.e., l ti = B(v ti ) → 0, …, K-1, where K is the number of divided neighborhoods, and B(v ti ) is mapped to the divided neighborhood based on the mapping relationship; w' is a weight distribution function for distributing corresponding weights to the adjacent nodes mapped to different neighborhoods; Z ti (v tj ) is a standardization item; v tk is a node that is mapped to the same neighborhood as v tj after v ti is mapped by the mapping function l tj .

[0039] Optionally, the adjacent node set B(v ti ) of node v ti also includes the same relevant nodes in the continuous frames, and is calculated as follows:

[0040] B(v ti ) = v qj |d(v qj ,v ti ) ≤ K, |q-t| ≤ |Γ / 2|

[0041] where v qj is the same relevant node of node v ti in the continuous frames; d(v qj ,v ti ) is the shortest path from node v qj to v ti ; K is a spatial distance threshold; q and t are the time frame serial numbers of v qj and v ti , respectively; and Γ is a time kernel control time range in the adjacent graph in time.

[0042] The mapping function is also redefined with the change of the adjacent node set of the node, as follows:

[0043] l ST (v ti ) = Lti (v tj )+(q-t+[Γ / 2])×K;

[0044] wherein, l ST (v ti ) is a new mapping function.

[0045] Optionally, the partition strategy of the neighborhood adopts five-partition, including: root node, near-center adjacent group, far-center adjacent group, near-center non-adjacent group and far-center non-adjacent group.

[0046] Through the above technical solutions, compared with the prior art, the application proposes a two-stage limb violence behavior recognition method based on improved OpenPose and space-time graph convolution ST-GCN. First, the lightweight improved algorithm MobileNet-V3-OpenPose is used to extract human skeleton key points, providing input for subsequent space-time feature analysis. Then, a skeleton space-time graph based on ST-GCN is constructed, and ST-GCN is used to analyze the space-time features of the key point sequence, and a five-partition strategy is innovatively introduced to optimize the graph convolution operation, to deeply mine the time sequence change and spatial dependence of fighting actions, and enhance the modeling ability of different human body part motion patterns. This method can effectively capture the rapid displacement of key nodes and their mutual relationship changes when violence occurs, improve the recognition accuracy, and successfully alleviate the problems of action semantic ambiguity and joint positioning error in the monitoring scene. The experimental results show that the model has high accuracy and robustness in the limb violence behavior recognition task, and can effectively improve the recognition ability while ensuring real-time performance. BRIEF DESCRIPTION OF DRAWINGS

[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only embodiments of the present application, and those skilled in the art can obtain other drawings according to the provided drawings without creative labor.

[0048] Figure 1 It is a method flowchart of the present application.

[0049] Figure 2 It is a human key point detection flowchart based on the improved OpenPose algorithm of the present application.

[0050] Figure 3 It is a deep separable convolution operation diagram of the present application.

[0051] Figure 4 It is an inverted residual structure diagram of the present application.

[0052] Figure 5 A schematic diagram of the Bneck Block structure of the present application.

[0053] Figure 6 A schematic diagram of an example of global optimal matching of human skeleton key points based on the Hungarian algorithm of the present application.

[0054] Figure 7 A schematic diagram of the human spatiotemporal skeleton construction of the present application.

[0055] Figure 8 A schematic diagram of the five-partitioning strategy of the present application.

[0056] Figure 9 A schematic diagram of the ST-GCN structure of the present application. DETAILED DESCRIPTION

[0057] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.

[0058] Embodiment 1

[0059] Embodiment 1 of the present application discloses a two-stage limb violent behavior recognition method based on improved OpenPose and spatiotemporal graph convolution ST-GCN, as shown in Figure 1 , which comprises:

[0060] Step 1: Based on the improved OpenPose algorithm, extract human skeleton key point time series data; wherein the improvement of the OpenPose algorithm comprises: extracting human skeleton key points based on MobileNet-V3 and globally optimally matching human skeleton key points based on the Hungarian algorithm.

[0061] In the improvement based on the OpenPose algorithm, in order to improve the detection efficiency and realize the global optimal matching of key points in a multi-person scene, the present application selects a lightweight MobileNet-V3 network to replace VGG-19 to reduce the parameter amount, reduce the calculation cost and improve the processing speed. Subsequently, the Hungarian algorithm is introduced to solve the bipartite graph matching problem to ensure the global optimal distribution of key points. The specific process is shown in Figure 2 .

[0062] In MobileNet, V1 reduces the amount of calculation by using deep separable convolution, V2 introduces an inverted residual structure and a linear bottleneck design, and V3 further optimizes the network architecture to improve efficiency and accuracy.

[0063] As Figure 3 shown, under the DSC (deep separable convolution) strategy, three single-layer convolution kernels are used to perform convolution operations on the three channels of the input, generating feature maps consistent with the number of input layer channels. To further fuse the feature information of different channels at the same spatial position, a point convolution kernel with a size of 1x1xM is used to perform weighted combination of the feature maps output in the previous step in the channel dimension, finally generating the output result; where M is the number of channels of the input image; this method effectively reduces the number of parameters and reduces the calculation cost.

[0064] The convolution parameter scale of each layer of the DSC structure is as follows:

[0065] S layer = S input x S dwc x S pwc x N filter ;

[0066] Where S layer is the parameter scale of the current layer, S input is the input feature map size, S dwc is the channel convolution kernel size, S pwc is the point convolution kernel size, and N filter is the set output feature quantity.

[0067] By introducing the inverted residual structure and linear bottleneck design, not only the feature transmission ability is improved, but also the information loss of low-dimensional input is avoided, so as to maintain the feature integrity. The inverted residual structure is as shown in Figure 4 . The inverted residual structure and linear bottleneck design are as follows:

[0068] First, dimensionality is increased by 1x1 convolution; then, 3x3 deep separable convolution is used for feature extraction; finally, 1x1 convolution is used again to reduce dimension and compress the number of channels, and linear activation function is used instead of ReLU (linear bottleneck design) when reducing dimension in the last layer. The reason for using this strategy is that deep convolution itself cannot directly adjust the number of channels, thereby limiting the information interaction between channels and affecting the effective capture of some features. To improve the feature expression ability, the number of channels needs to be expanded before feature extraction, and then subsequent calculations are performed. This structure effectively improves the information loss problem.

[0069] MobileNet-V3 not only continues to use the deep separable convolution technology, combines the use of inverted residual structure and linear bottleneck design, and additionally introduces the SENet module for adaptive adjustment of model parameters, and uses the h-swish activation function instead of the traditional ReLU activation function. The present application selects MobileNet-V3 as the replacement scheme of the feature network part of the OpenPose algorithm to improve the device inference speed while maintaining the accuracy, and has better efficiency compared with the original VGG-19 network architecture.

[0070] The MobileNet-V3 network structure is shown in Table 1.

[0071] Table 1 MobileNet-V3 network structure

[0072] Input Operator Expsize S #out NL SE 224 2 x 3 Conv2d, 3x3 - 2 16 HS - 112 2 x 16 Bneck, 3x3 16 2 16 RE √ 56 2 x 16 Bneck, 5x5 72 2 24 HS - 28 2 x 24 Bneck, 5x5 96 2 40 HS √ 14 2 x 40 Bneck, 5x5 240 1 40 HS √ 14 2 x 40 Bneck, 5x5 120 1 48 HS √ 14 2 x 48 Bneck, 5x5 144 1 48 HS √ 14 2 x 48 Bneck, 5x5 288 2 48 HS √ 7 2 x 96 Bneck, 5x5 576 1 96 HS √ 7 2 x 96 Conv2d, 5x5 576 1 96 HS √ 7 2 x 96 Pool, 7x7 - 1 576 Conv2d, 1x1, NBN - 7 2 x 576 HS - 1 - - - 1 2 ×576]]> Conv2d, 1x1, NBN - 1 1024 Figure 5 - 1 2 x 1024 Figure 6 - 1 k - -

[0073] In Table 1, “Input” represents the shape change of each feature layer of MobileNet-V3; “Operator” explains the calculation module used by each feature layer; “Bneck” refers to the bottleneck module; “exp size” represents the channel number after the dimension increase of B inverse residual structure; “#out” represents the feature channel number when input to Bneck; “SE” identifies whether the attention mechanism is introduced in this layer; “NL” represents the type of activation function used, where “RE” is the ReLU activation function and “HS” is the h-swish activation function; “NBN” is used to represent that batch normalization is not used; and “stride” represents the stride used in each block structure.

[0074] In MobileNet-V3, the optimized h-swish activation function replaces the traditional ReLU activation function. Compared with ReLU, h-swish has smoother characteristics, which can reduce the impact of ReLU and improve the performance of the model on low-dimensional data. In addition, with the increase of network layers, h-swish can also reduce the number of hyperparameters, thereby effectively reducing the computational cost. The h-swish activation function is as follows:

[0075]

[0076] The improved version of OpenPose optimizes the feature extraction part by introducing the Bneck (bottleneck) module. The structure of the Bneck module is shown in Figure 7 , which first expands the channel number through a 1×1 convolution layer in the processing flow, then performs batch normalization (BN) and uses the h-swish function for activation. Finally, the channel compression is completed through a 1×1 convolution layer or average pooling. In addition, in order to prevent information loss in deep network and promote feature flow, the module introduces a residual connection mechanism.

[0077] MobileNet-V3 improves performance in low-dimensional feature spaces by using the h-swish activation function and enhances the weight adjustment capabilities of feature maps with the SENet module. These optimizations enable MobileNet-V3 to achieve a better balance between computational efficiency and performance.

[0078] To avoid the potential failure to achieve optimal results due to discarding partial matches, the OpenPose algorithm proposes a bipartite graph matching strategy based on the Hungarian algorithm to achieve maximum matching of key points. Its core principle is to use a maximum affinity matching mechanism. If a matching target is already occupied, the algorithm uses augmenting paths to traverse and select the path with the highest overall affinity for matching. The following uses the key point matching of two people as an example to explain the algorithm's workflow in detail.

[0079] The global optimal matching of key points of human skeleton is performed based on the Hungarian algorithm, specifically:

[0080] Step 1: Use the PCM algorithm to obtain two sets of key points to be matched and calculate the affinity between all point pairs;

[0081] Step 2: Screen possible matching key point pairs based on affinity threshold;

[0082] Step 3: Select key points from set 1 in sequence. If all key points are selected, jump to Step 5. Otherwise, continue to select the key point with the highest affinity that is not currently matched from set 2. If the key point has not been matched, directly connect the line and repeat Step 3. If the key point has been matched, execute Step 4. If there is no matching key point, skip the point and continue to Step 3.

[0083] Step 4: Traverse all possible augmenting paths and calculate their affinity sum Ci. Select the augmenting path corresponding to the maximum Ci as the optimal path. After completing the path adjustment, return to Step 3.

[0084] Step 5: The global optimal matching of key points is completed and the matching process ends.

[0085] like Figure 8 The method is further illustrated by an example. The dashed lines in the figure represent possible keypoint pairs that can be matched after threshold filtering. It can be observed that when the Hungarian algorithm fails to match, the augmenting path traversal optimization successfully adjusts the matching object for left ear 2. This not only corrects the matching failure of left ear 3 but also optimizes the matching relationship for left ear 2. By traversing all possible augmenting paths, this strategy achieves global optimal matching of keypoints under the principle of maximum affinity, while also improving matching efficiency.

[0086] Step 2: based on the spatio-temporal graph convolution ST-GCN structure, the spatio-temporal features of the human body skeleton key point time series data are analyzed and extracted, and the limb violence behavior recognition result is output; wherein the spatio-temporal graph convolution ST-GCN structure is stacked by multiple ST-GCN modules, and each ST-GCN module is composed of an attention layer, a spatial graph convolution and a temporal graph convolution.

[0087] The graph convolution network (GCN) and the convolutional neural network (CNN) are both essentially aimed at extracting features. The main difference between the two is that the CNN is mainly suitable for regular Euclidean data such as images, videos and texts, which have a clear grid structure in space. The GCN is suitable for data with non-Euclidean structure, such as graph data, in which the connection relationship between data points (nodes) does not follow a regular grid structure, but is determined by the topology of the graph. The task of the present application involves violence action recognition based on human body skeleton key points, and this data type belongs to graph data in non-Euclidean data. Therefore, the present application adopts a graph convolution method for behavior recognition. As shown in Figure 9 The ST-GCN model combines the GCN and the temporal convolutional network (TCN), and its architecture is stacked by multiple spatio-temporal graph convolution modules. Each module includes two main operations: spatial graph convolution and temporal convolution. In the spatial dimension, the ST-GCN regards the human body joint as a graph node and aggregates neighborhood information through graph convolution to capture the spatial topological relationship between the joints. In the time dimension, the model uses a one-dimensional convolution kernel to model the motion trajectory of the joints in consecutive frames along the time axis, thereby capturing the dynamic evolution of the human skeleton. This dual-path design enables the ST-GCN to capture both spatial connectivity and temporal continuity, thereby more comprehensively understanding the human motion features. In addition, the key points in the skeleton sequence are connected by spatial edges and temporal edges. The spatial edges represent the topological relationship between the joints, and the temporal edges capture the joint trajectory at different time steps. This design enables the ST-GCN to achieve efficient motion recognition in the spatio-temporal dimension.

[0088] Before the spatio-temporal features of the human body skeleton key point time series data are analyzed and extracted based on the spatio-temporal graph convolution ST-GCN structure, it further includes: constructing a spatio-temporal graph to represent the hierarchical structure of the skeleton sequence, specifically:

[0089] These skeleton sequences not only contain spatial connections between nodes, but also have temporal connections across frames. For a skeleton sequence with N joints and T frames, a undirected spatio-temporal graph G=(V,E) is constructed; wherein V is the node set (Vertex, also called Node), V={v tit = 1, ..., T; i = 1, ..., N}, represents the position of all joints in each time frame. As the input of ST-GCN, the feature vector of each node F(vti) consists of the joint coordinate vector of the node in frame t and its confidence estimate. The process of constructing the spatiotemporal graph on the skeleton sequence can be divided into two steps: the first is spatial connection, in which the joints are connected according to the physical structure of the human body in each frame; the second is temporal connection, in which each joint at the same position in consecutive frames is connected to form a temporal connection.

[0090] E is the edge set (Edge); in the graph structure, the edge set E is divided into two, the subset E s :E s =(v ti ,v tj )|(i,j)∈H is used to describe the skeleton connection in the same frame, where H represents the set of joints that conform to the natural connectivity of the human body; the subset E F :E F =(v ti ,v (t+1)i ) records the temporal relationship between frames, that is, connecting the nodes of the same joint in adjacent frames. For each joint, the temporal edge represents its trajectory over time;

[0091] Set the convolution kernel size to K×K and the input feature map to f in , the number of channels is c, then the single-channel output value at the spatial position is as follows:

[0092]

[0093] Among them, p is the sampling function, which is used to determine the feature points participating in the convolution operation in the neighborhood of x. For image convolution, the sampling function can be expressed as: p(x,h,w)=x+p′(h,w); x is the current feature point; p′(h,w) is the offset relative to the current position (h,w), indicating the offset of the domain position; w is the weight function as the core parameter of the convolution operation, and the inner product calculation is performed with the c-dimensional input feature vector to complete the feature extraction.

[0094] In the field of images, the sampling function p(h,w) is usually defined as the neighboring pixels relative to the center position x. Similarly, in the graph structure, the sampling function can be used to define the neighboring pixels relative to the node v. ti The mathematical expression of the adjacent point set is as follows:

[0095] B(v ti )=v tj |d(v tj ,v ti )≤D;

[0096] Among them, d(v tj ,vti ) is the shortest path from node v tj to v ti ; when D = 1, the acquisition function can be simplified as: p(v ti , v tj ) = v ti .

[0097] In the graph convolution operation, due to the large difference in the number of neighbors of each node, the method of neighborhood division is usually used to construct the weight function. Specifically, the neighbor set B(v ti ) of node v ti can be divided into K subsets of fixed size, and each subset corresponds to a label. In order to classify adjacent nodes into the corresponding subset, a mapping function l ti : B(v ti )→0,…,K-1 is defined. Based on the mapping relationship, the weight function can be represented by a tensor (c, K), where c is the number of channels and K represents the total number of subsets. When independent weights are assigned to different subsets, as follows:

[0098] w(v ti , v tj ) = w'(l ti (v tj )).

[0099] Combining the above sampling function and weight function, the calculation expression of spatial graph convolution is as follows:

[0100]

[0101]

[0102] w(v ti , v tj ) = w'(l ti (v tj ));

[0103] Z ti (v tj ) = |{v tk |l ti (v tk ) = l ti (v tj )}|;

[0104] where f out (v ti ) is the feature output by the spatial graph convolution operation of node v ti ; v ti is the i-th key node of the t-th frame in the graph structure; B(v ti ) is the node v tiThe set of adjacent points of in is the input characteristic function; p(v ti ,v tj ) is the acquisition function; D is the coefficient that defines the neighborhood range; d(v tj ,v ti ) is the node v tj to v ti The first equation in the acquisition function indicates that |d(v tj ,v ti )≤D condition v tj As v ti The second equation in the acquisition function indicates that when D = 1, v ti The adjacent point of is itself; w(v ti ,v tj ) is the weight function based on neighborhood division; l ti is the mapping function, that is: l ti =B(v ti )→0,…,K-1, where K is the number of divided neighborhoods. Based on this mapping relationship, B(v ti ) is mapped to the divided neighborhood; w' is a weight distribution function used to assign corresponding weights to adjacent points mapped to different neighborhoods; Z ti (v tj ) is a standardized item; v tk For v tj After mapping function l ti After mapping, the divided area is consistent with v tj Same nodes.

[0105] In the spatiotemporal graph, in order to extend the spatial graph to the spatiotemporal graph, the connection of the same joint point in consecutive frames needs to be taken into account. Therefore, the node v ti The adjacent point set B(v ti ) includes not only the adjacent nodes of the current frame, but also the same joint points in consecutive frames, and is calculated as follows:

[0106] B(v ti )=v qj |d(v qj ,v ti )≤K,|qt|≤|Γ / 2;

[0107] Among them, v qj For node v ti The same joint point in consecutive frames; d(v qj ,v ti ) is the node v qj to v ti The shortest path; K is the spatial distance threshold; q, t are v qjand v ti the time frame number; Γ is the time kernel control the time range in the adjacent graph;

[0108] The mapping function is also redefined as the adjacent point set of the node changes as follows:

[0109] l ST (v ti ) = l ti (v tj ) + (q-t + [Γ / 2]) x K;

[0110] Wherein, l ST (v ti ) is the new mapping function.

[0111] The adjacent point set B(v ti ) of the center node v ti ) is divided into several fixed subsets, and this division method directly affects the construction of the weight function. In order to more effectively capture the complex spatial relationship in the skeleton data, in the processing mode of single frame image D=1, the ST-GCN algorithm adopts three different neighborhood division strategies, including single division strategy (Uni-labeling), distance partitioning strategy (Distance partitioning) and spatial configuration partitioning strategy (Spatial configuration partitioning), which are designed to enhance the modeling ability of the model to the spatial structure in the skeleton data. However, the traditional ST-GCN partition strategy only focuses on the weight allocation between the root node and its adjacent nodes, and the interaction span between the nodes is small, although it is more advantageous in detecting small amplitude actions of the human body, but when the human body is in violent motion (such as the occurrence of violent behavior), the amplitude of the human joint node change will be very dramatic in both spatial and temporal dimensions. The above three partitioning methods obviously ignore the interaction with nodes at a greater distance, which will inevitably lead to a decline in detection performance. Therefore, in order to better capture the large amplitude action features of the human body, the present application proposes an improved partition method: on the basis of the existing spatial division, two new regions are added, and the domain span D is expanded from 1 to 2, and the adjacent region of the root node is expanded to five sub-regions, as shown in ​ These regions include: the root node itself; the near-center adjacent group: the nodes with a distance of 1 from the root node and closer to the skeleton center than the root node; the far-center adjacent group: the nodes with a distance of 1 from the root node and farther away from the skeleton center than the root node; the near-center non-adjacent group: the nodes with a distance of 2 from the root node and closer to the skeleton center than the root node; the far-center non-adjacent group: the nodes with a distance of 2 from the root node but farther away from the skeleton center than the root node. The specific partitioning is as follows:

[0112]

[0113] where r i is the distance from the root node to the center of gravity, and d(v tj ,v ti ) is the distance from the root node v ti to the node v tj .

[0114] The newly proposed five-partition division strategy not only optimizes the weight distribution of the whole body nodes when targeting limb violence, but also pays more attention to the correlation between the local limb movements of the violent person, thereby enhancing the model's perception of the overall human violence and improving the accuracy of limb violence recognition.

[0115] The overall process of the ST-GCN algorithm includes four key steps: data input, feature extraction, spatio-temporal modeling, and final classification. First, the original spatial-temporal data of the joints is preprocessed, converting the five-dimensional tensor [N, C, T, V, M] to a four-dimensional tensor [N x M, C, T, V], where N is the batch dimension, C is the channel number, T is the time dimension, V is the number of joints, and M is the number of human instances. This conversion optimizes the data storage form and simplifies subsequent operations. The data is normalized by Batch Normalization (BN) to standardize the feature distribution, stabilize the training process, avoid gradient vanishing or explosion, and accelerate convergence. The normalized data is fed into the network main body composed of multiple ST-GCN modules stacked together. Each ST-GCN module consists of an attention layer, a spatial graph convolution, and a temporal graph convolution, which are alternately stacked to gradually extract spatio-temporal features.

[0116] In the attention layer, the network generates an attention matrix that adjusts the edge weights of the graph structure, dynamically adjusting the weights of the edges in the graph. This allows the network to better capture the spatial and temporal relationships between joints, enhancing its sensitivity to action details.

[0117] Next, the GCN spatial convolution aggregates the spatial relationship features between joints through convolution operations. The initial input joint feature dimension is 3, representing the three-dimensional coordinates and confidence information of the joints. As the ST-GCN modules are stacked, the expressive power of the spatial features gradually increases.

[0118] The TCN temporal convolution performs convolution operations in the time dimension, modeling the dynamic changes of actions over time. By compressing the time dimension, the network effectively extracts fine-grained joint features in each frame, improving the expressive power of the temporal features. Through this process, the ST-GCN can handle complex features in both time and space dimensions.

[0119] As ​As shown, in the stacking process of multiple ST-GCN modules, the time dimension is gradually compressed from 150 frames to 38, and the joint feature dimension is expanded from 3 to 256. This design of feature dimension expansion and time dimension compression enables the network to capture the details and stage characteristics of the action. Finally, the network integrates the spatio-temporal features through global average pooling, and then classifies the extracted high-level semantic information through a fully connected layer to complete the final limb violence action recognition task. This complete processing flow enables ST-GCN to effectively capture and express the spatio-temporal patterns in the skeleton data, providing strong support for complex action classification tasks.

[0120] The COCO Keypoints dataset is respectively trained and learned by the original OpenPose network and the MobileNet-V3-OpenPose network, and the performance of the VGG-19 and MobileNet series feature extraction network models on the COCO dataset is compared. As shown in Table 2, although the VGG-19 has certain feature extraction capability, its calculation complexity is high, resulting in a low frame rate. In contrast, the replaced MobileNet-V3-OpenPose network is superior to the original VGG-19 network in terms of precision, recall and F1 score, proving that the MobileNet architecture has better performance while improving efficiency. Among them, the recognition accuracy of MobileNet-V3 reaches 91.68%, the precision is stable above 87%, and the processing speed is 29 to 33 FPS, far exceeding VGG-19, with a frame rate increase of more than three times. In terms of parameter quantity, the MobileNet-V3-OpenPose network is reduced by about 8.57 times compared with VGG-19, greatly reducing the model complexity and storage overhead. Overall, MobileNet-V3 not only performs best in terms of precision and recall, but also has higher real-time processing capability, balancing detection accuracy and efficiency, so the MobileNet-V3 is finally selected as the backbone network of the OpenPose feature extraction.

[0121] Table 2 Comparison of feature network algorithm performance

[0122]

[0123] The different partition methods are used on the ST-GCN algorithm to perform experiments on the NTU RGB+D dataset and Fight-NTU RGB+D dataset, and performance comparison and evaluation are performed, as shown in Table 3. Compared with other partition strategies, the five-partition method proposed in the application performs better in recognition accuracy. On the NTU RGB+D dataset, the Top-1 accuracy is increased by 42.6%, 4.11%, and 1.87%, respectively, and the Top-5 accuracy is increased by 29.24%, 1.19%, and 0.18%, respectively; on the Fight-NTU RGB+D dataset, the Top-1 accuracy is increased by 42.27%, 2.9%, and 1.29%, respectively, and the Top-5 accuracy is increased by 23.22%, 1.9%, and 0.14%, respectively. The experimental results fully verify the effectiveness and superiority of the five-partition method in feature region division, and provide strong support for improving the accuracy of limb violence behavior recognition.

[0124] Table 3 Model performance comparison of different partition methods

[0125]

[0126] To verify the superiority of the improved ST-GCN model after the partition strategy, comparison experiments are performed with existing classical algorithms on the NTU RGB+D dataset and Fight-NTU RGB+D dataset, and the experimental results are shown in Table 4. As can be seen from the data, the GCN method is better than the LSTM method as a whole, indicating that the feature modeling based on graph convolution has stronger expression ability in behavior recognition tasks. In addition, the improved ST-GCN of the application has a Top-1 and Top-5 accuracy of 80.52% and 96.21% on the NTU RGB+D dataset, and a Top-1 and Top-5 accuracy of 87.23% and 98.56% on the Fight-NTU RGB+D dataset, which shows better recognition effect compared with other algorithms. This further verifies the significant advantages of the model proposed in the application in violence action feature modeling and recognition accuracy.

[0127] Table 4 Performance comparison of different classical algorithms

[0128]

[0129] To verify the effectiveness of the lightweight MobileNet-V3-OpenPose algorithm and the ST-GCN structure under the improved partitioning strategy, we designed four sets of experiments, testing the original ST-GCN, MobileNet-V3-OpenPose + ST-GCN, the original OpenPose + improved ST-GCN, and MobileNet-V3-OpenPose + improved ST-GCN. As shown in Table 5, on the NTURGB+D dataset, this method achieved a Precision of 83.12%, a Recall of 79.75%, an F1 score of 85.31%, and an FPS of 24.46. The original ST-GCN uses OpenPose to extract human skeleton sequences. As a bottom-up pose estimation algorithm, OpenPose has a fast detection speed, but there is still room for improvement in accuracy and parameter count. After the introduction of MobileNet-V3-OpenPose, the lightweight network reduces parameters, reduces computational complexity, and improves processing speed. At the same time, the Hungarian algorithm is combined for bipartite graph matching. Compared with the original ST-GCN, while maintaining precision, recall, and F1 score, the operation speed is greatly improved, and the number of parameters is reduced by about 7.5 times. In addition, the spatial configuration strategy adopted by the original ST-GCN does not consider the weight distribution of distant nodes, resulting in limited recognition accuracy. The improved ST-GCN of the present invention introduces a five-partition method, which increases Precision to 88.65%, Recall to 83.65%, and F1 score to 86.23%. Further integrating the improved OpenPose with ST-GCN, the Precision, Recall, and F1 of MobileNetV3-OpenPose + improved ST-GCN increase by 2.91, 3.31, and 2.03 percentage points, respectively, and the number of parameters is reduced by about 4.8 times, fully verifying the effectiveness and superiority of the proposed method.

[0130] Table 5 Ablation experiment

[0131]

[0132] The embodiment of the application discloses a two-stage limb violence behavior recognition method based on improved OpenPose and space-time graph convolution ST-GCN. First, the human skeleton key points are extracted by using a lightweight improved algorithm MobileNet-V3-OpenPose, so as to provide input for subsequent space-time feature analysis. Then, a skeleton space-time graph based on ST-GCN is constructed, the key point sequence is analyzed by using ST-GCN, and a five-partition strategy is innovatively introduced to optimize the graph convolution operation, so that the time sequence change and spatial dependency of fighting actions are deeply mined, and the modeling ability of different human body part motion patterns is enhanced. The method can effectively capture the rapid displacement of the key points and the change of the mutual relationship when the violence behavior occurs, improve the recognition accuracy, and successfully relieve the problems of action semantic ambiguity and joint positioning error in the monitoring scene. The experimental results show that the model has high accuracy and robustness in the limb violence behavior recognition task, and can effectively improve the recognition ability while ensuring the real-time performance.

[0133] The embodiments in the specification are described in a progressive manner, and each embodiment focuses on the difference from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed by the embodiments, since it corresponds to the method disclosed by the embodiments, the description is relatively simple, and the related parts can be referred to the method part.

[0134] The above description of the disclosed embodiments enables a person skilled in the art to implement or use the application. Various modifications to the embodiments will be apparent to those skilled in the art, and the general principles defined in the application can be implemented in other embodiments without departing from the spirit or scope of the application. Therefore, the application will not be limited to the embodiments shown in the application, but will conform to the widest scope consistent with the principles and novel features disclosed in the application.

Claims

1. A two-stage limb violence behavior recognition method based on improved OpenPose and space-time graph convolution ST-GCN, characterized in that, Comprise: Step 1: based on improved OpenPose algorithm, extract human skeleton key point time series data;Wherein, the improvement of the OpenPose algorithm, including: based on MobileNet-V3 extract human skeleton key points and based on the Hungarian algorithm for the global optimal matching of the human skeleton key points; Step 2: based on the space-time graph convolution ST-GCN structure, the space-time feature analysis and extraction of the human skeleton key point time series data are carried out, and the limb violence behavior recognition result is output;Wherein, the space-time graph convolution ST-GCN structure is stacked by a plurality of ST-GCN modules, and each ST-GCN module is composed of an attention layer, a spatial graph convolution and a time graph convolution.

2. The two-stage limb violence behavior recognition method based on improved OpenPose and space-time graph convolution ST-GCN according to claim 1, wherein, In step 1, the MobileNet-V3 uses a depth separable convolution technology, combines the use of an inverted residual structure and a linear bottleneck design, and introduces an SENet module for adaptive adjustment of model parameters, and uses an h-swish activation function instead of a traditional ReLU activation function.

3. The two-stage limb violence behavior recognition method based on improved OpenPose and space-time graph convolution ST-GCN according to claim 2, characterized in that, The depth separable convolution technology is as follows: Three single-layer convolution kernels are used to perform convolution operation on the three channels of the input, respectively, to generate feature maps consistent with the channel number of the input layer; A point convolution kernel with a size of 1x1xM is used to weight and combine the feature maps output in the last step in the channel dimension, and finally generate the output result;Wherein, M is the channel number of the input image; The convolution parameter scale of each layer is as follows: S layer = S input x S dwc x S pwc x N filter ; where S layer is the parameter size of the current layer, S input is the input feature map size, S dwc is the channel convolution kernel size, S pwc is the point convolution kernel size, N filter is the set output feature quantity.

4. The two-stage limb violence behavior recognition method based on improved OpenPose and space-time graph convolution ST-GCN according to claim 2, characterized in that, The inverted residual structure and linear bottleneck design are as follows: First, the dimension is increased by 1x1 convolution;Then, 3x3 depth separable convolution is used for feature extraction;Finally, 1x1 convolution is used again to reduce the dimension and compress the channel number, and linear activation function is used instead of ReLU in the last dimension reduction.

5. The two-stage limb violence behavior recognition method based on improved OpenPose and space-time graph convolution ST-GCN according to claim 2, characterized in that, The h-swish activation function is as follows:

6. The two-stage limb violence behavior recognition method based on improved OpenPose and space-time graph convolution ST-GCN according to claim 1, characterized in that, In step 1, based on the Hungarian algorithm for global optimal matching of the human skeleton key points, specifically: Step1: use PCM algorithm to obtain two groups of key points to be matched, and calculate the affinity between all point pairs; Step2: select the possible matching key point pair according to the affinity threshold; Step3: select the key point from set 1 in turn, if it is taken, jump to Step5, otherwise continue to select the current unmatched key point with the highest affinity from set 2;If the key point has not been matched, directly connect the line and repeat Step3;If the key point has been matched, execute Step4;If there is no matching key point, skip this point and continue Step3; Step4: traverse all possible augmented paths and calculate the affinity sum Ci, select the augmented path corresponding to the maximum Ci as the optimal path, and return to Step3 after adjusting the path; Step5: the global optimal matching of the key points is completed, and the matching process is ended.

7. The two-stage limb violence behavior recognition method based on improved OpenPose and space-time graph convolution ST-GCN according to claim 1, characterized in that, In step 2, before the space-time graph convolution ST-GCN structure is used to analyze and extract the space-time features of the human skeleton key point time series data, it also includes: constructing a space-time graph, specifically: For a skeleton sequence with N joints and T frames, a time-space graph G=(V, E) is constructed; wherein V is the node set, V={v ti |t=1,…,T; i=1,…,N} represents the positions of all joints in each time frame, as the input of ST-GCN, the feature vector on each node F(vti) is composed of the joint coordinate vector of the node in frame t and its confidence estimate; the process of constructing the time-space graph on the skeleton sequence can be divided into two steps: one is spatial connection, within each frame, the joints are connected according to the physical structure of the human body; the other is time connection, the same position of each joint in consecutive frames is connected to form the connection in time sequence; E is an edge set; the edge set E is divided into two subsets E s : E s = (v ti , v tj ) | (i, j) e H is used to describe the skeleton connection within the same frame, where H represents a set of joints that meet the natural connectivity of the human body; the subset E F : E F = (v ti , v (t+1)i ) records the timing relationship between frames, that is, the nodes connecting the same joints in adjacent frames, and for each joint, the timing edge represents its trajectory over time; Set the convolution kernel size to KxK, the input feature map to f in , and the number of channels to c. The single-channel output value at a spatial position is as follows: Wherein, p is a sampling function for determining the feature points in the neighborhood of x participating in the convolution operation, for image convolution, the sampling function can be expressed as: p(x, h, w) = x + p'(h, w); x is the current feature point; p'(h, w) is the offset relative to the current position (h, w), indicating the offset of the field position; w is the weight function as the core parameter of the convolution operation, and the c-dimensional input feature vector is calculated by inner product to complete feature extraction.

8. The two-stage limb violence behavior recognition method based on improved OpenPose and space-time graph convolution ST-GCN according to claim 1, characterized in that, In step 2, the calculation expression of the spatial graph convolution is as follows: w(v ti ,v tj ) = w'(l ti (v tj )) ; Z ti (v tj ) = |{v tk |l ti (v tk ) = l ti (v tj )}|; wherein f out (v ti ) is the feature of node v ti output after the spatial graph convolution operation; v ti is the i-th key node of the t-th frame in the graph structure; B(v ti ) is the set of adjacent nodes of node v ti ; f in is the input feature function; p(v ti ,v tj ) is the collection function; D is the coefficient defining the neighborhood range; d(v tj ,v ti ) is the shortest path from node v tj to v ti ; the first equation in the collection function indicates that v tj that satisfies the condition |d(v ti ,v tj )≤D will be the adjacent node of v ti ; the second equation in the collection function indicates that when D=1, the adjacent node of v ti is itself; w(v ti ,v tj ) is the weight function based on neighborhood division; l ti is the mapping function, i.e., l ti =B(v ti )→0,…,K-1, wherein K is the number of divided neighborhoods, and based on the mapping relationship, B(v ti ) is mapped to the divided neighborhood; w' is the weight allocation function, which is used to allocate corresponding weights to the adjacent nodes mapped to different neighborhoods; Z ti (v tj ) is the standardization item; v tk is the node that is mapped to the same region as v tj after v ti is mapped by the mapping function l tj .

9. The two-stage limb violence behavior recognition method based on improved OpenPose and space-time graph convolution ST-GCN according to claim 8, characterized in that, Node v ti The set of adjacent nodes B(v ti ) also includes the same adjacent nodes in consecutive frames, calculated as follows: B(v ti ) = v qj |d(v qj ,v ti ) ≤ K, |q - t| ≤ |Γ / 2; wherein v qj is a node ti in the same graph in consecutive frames; d(v qj ,v ti ) is the shortest path from node v qj to v ti ; K is a spatial distance threshold; q, t are the time frame numbers of v qj and v ti , respectively; and Γ is a time kernel controlling the time range in the adjacent graph in time. The mapping function is also redefined with the change of the adjacent point set of the node, as follows: l ST (v ti )=l ti (v tj )+(q-t+[Γ / 2])×K; where l ST (v ti ) is a new mapping function.

10. The two-stage physical violence behavior recognition method based on improved OpenPose and spatiotemporal graph convolution ST-GCN according to claim 8 is characterized in that: The division strategy of the neighborhood adopts five-partition, including: root node self, near-center adjacent group, far-center adjacent group, near-center non-adjacent group and far-center non-adjacent group.

Citation Information

Cited By

  • Method and system for identifying behavior pattern of red-crowned crane based on artificial intelligence visual technology

    CN121330723A