Open set action recognition method based on high-order space-time self-attention mechanism

By combining the advanced space-time self-attention mechanism and open set recognition algorithm, the improved ST-GCN model can effectively identify unknown categories of actions in complex environments, solving the recognition limitations of traditional methods, improving the accuracy and robustness of action recognition, and is suitable for human-computer collaboration, intelligent monitoring and other fields.

CN120259935APending Publication Date: 2025-07-04NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510234651.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Existing action recognition methods are difficult to identify unknown categories of actions in complex environments, especially in industrial and complex environments. Traditional closed-set recognition methods cannot effectively deal with unknown categories of actions, resulting in insufficient recognition capabilities and poor robustness.

Method used

Combining the advanced space-time self-attention mechanism and open set recognition algorithm, through the improved ST-GCN model, the spatial multi-attention layer, the temporal multi-attention layer and the OpenMax layer are introduced to optimize the bone topology structure, enhance the model's ability to capture long-term dependencies and global context information, and realize the recognition of unknown categories of actions.

Benefits of technology

It significantly improves the accuracy and robustness of action recognition in complex environments, can effectively distinguish known and unknown categories, and is suitable for open set scenarios, improving the applicability and real-timeness of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259935A_ABST
    Figure CN120259935A_ABST
Patent Text Reader

Abstract

The invention discloses an open set action recognition method based on a high-order space-time self-attention mechanism, and the method comprises the steps: introducing the high-order space-time self-attention mechanism, enabling a model to effectively capture the long-time dependence relation and global context information in a human body action sequence, and improving the recognition precision of complex actions, shielding and background changes. In an industrial scene, the action of an operator is often influenced by environmental factors, the recognition effect of a traditional method is poor under the condition, and the accuracy and robustness of action recognition are remarkably improved under the complex conditions. Through combination of the OpenMax layer, actions of known categories and unknown categories can be distinguished, and the limitation that only known categories in a training set can be identified in the prior art is solved. According to the technology, the model can make correct recognition and judgment when facing unknown actions in an open set environment, the recognition precision in an open set scene is improved, and the applicability of the system is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and particularly to an open-set action recognition method based on a high-order spatio-temporal self-attention mechanism. Background Art

[0002] With the rapid development of artificial intelligence technology, human action recognition has achieved remarkable application results in fields such as human-machine collaboration, intelligent monitoring, and medical assistance. The core of human action recognition lies in accurately extracting and interpreting the dynamic features of the human body through effective action representation and classification models. However, with the rapid development of human-machine collaboration technology, human action recognition faces many challenges in complex environments. Existing action recognition methods usually rely on training data of known categories and belong to the closed-set recognition problem. In this method, the model can only recognize the categories existing in the training set and has weak recognition ability for newly emerging action categories that do not appear in the training set. This problem is particularly prominent in practical applications, especially in industrial and complex environments where there are often a large number of action categories of unknown types, making it difficult for traditional closed-set recognition methods to handle, and open-set recognition has become a research hotspot.

[0003] Different from closed-set recognition, open-set recognition methods can recognize and distinguish samples of known and unknown categories in the face of unknown categories. This ability enables the model to better adapt to unknown situations in a dynamically changing environment and improve the robustness and accuracy of recognition. Open-set action recognition methods are of great significance in practical applications, especially in scenarios such as industrial operations, robot control, and intelligent monitoring, where operators may perform some actions not included in the training set, and these actions still need to be recognized.

[0004] In recent years, the spatio-temporal graph convolutional network (ST-GCN), as an effective skeletal action recognition method developed in recent years, uses graph convolutional operations to capture the spatial relationships and temporal dependencies between human joints and can better model the spatio-temporal features of human skeleton data. Although the ST-GCN model has achieved remarkable results in action recognition, it still faces some challenges, such as insufficient ability to capture long-term dependence relationships, difficulty in handling occlusion problems, and lack of open-set recognition ability.

[0005] Therefore, to solve the above problems, this patent proposes an IST-GCN action recognition model framework that combines a high-order spatio-temporal self-attention mechanism and an open-set recognition algorithm, aiming to enhance the adaptability and robustness of existing action recognition models in complex environments, especially the recognition ability when facing unknown action categories. Summary of the Invention

[0006] To solve the above technical problems, the present invention aims to provide an open-set action recognition method based on a high-order spatio-temporal self-attention mechanism, aiming to solve the technical problems faced by existing action recognition algorithms in complex industrial environments and open-set scenarios.

[0007] This application is implemented based on the following technical solutions: An open-set action recognition method based on a high-order spatio-temporal self-attention mechanism, characterized by comprising the following steps:

[0008] S1 Spatio-temporal modeling based on skeletons:

[0009] Perform pose estimation on the input video, extract the skeleton topology HRCS-11, where the HRCS-11 contains 11 key skeleton nodes, including upper limb and neck nodes; connect the skeleton nodes along the time series to construct a spatio-temporal topology graph G=(V,E), where the node set V={v ti |t = 1,…,T, i = 1,2,…,N} represents a skeleton sequence containing T frames, each frame contains N joint points, and the edge set E includes a spatial connection Es and a time connection ET;

[0010] S2 Optimize the spatio-temporal graph convolutional network ST-GCN based on a high-order spatio-temporal self-attention mechanism:

[0011] Embed the STTA module in the spatio-temporal graph convolutional network ST-GCN. The STTA module includes a spatial multi-head self-attention layer S-MHSA, a temporal multi-head self-attention layer T-MHSA, a feed-forward network FFN with residual connections, and a layer normalization LN to obtain an improved ST-GCN model;

[0012] Extract the relationship between joints in the spatial domain through the spatial multi-head self-attention layer S-MHSA, and capture the long-term temporal dependencies through the T-MHSA;

[0013] S3 Optimize the improved ST-GCN model based on an open-set recognition algorithm:

[0014] Add an OpenMax layer after the fully connected layer of the improved ST-GCN model, and use the activation vectors AV of the samples in the training set to fit the cumulative distribution function CDF of each category;

[0015] In the test phase, according to the distance between the activation vector AV of the test sample and the mean activation vector MAV of each category, calculate the probability that it belongs to a known category through the cumulative distribution function CDF, and generate an N+1-dimensional probability vector to identify unknown category actions.

[0016] Furthermore, the construction of the spatio-temporal topology graph in S1 specifically includes:

[0017] Spatial edge set E s ={(vti , v tj ) | 1 ≤ t ≤ T, (i, h) ∈ P}, where P is the preset bone node connection pair in the optimized bone topology HRCS - 11; v ti , v tj are the nodes in the ti - th and tj - th bone topologies respectively, and T represents the number of frames in the time dimension;

[0018] The time - edge set E T = {(v ti , v ki ) | 1 ≤ k, t ≤ T, 1 ≤ i ≤ N, |k - t| = 1};

[0019] The edge set E of the human body's topology in the spatio - temporal dimension is represented as E = E s ∪ E T .

[0020] Furthermore, the calculation of the multi - head self - attention mechanism in the STTA module includes:

[0021] Project the input spatio - temporal features into query, key, and value matrices, and calculate the attention weights of each individual attention head through the multi - head attention formula .

[0022] After the outputs of all attention heads are concatenated, they are projected back to the original dimensional representation, forming the following expression:

[0023] MHSA(Q, K, V) = Concat(Attn1, Attn2, … Attn h )W o

[0024] The operations of the two - layer feed - forward network FFN and layer normalization in the STTA module are as follows:

[0025] H′ = LayerNorm(MHSA(X)+X)

[0026] FFN(H′) = σ(H′W1 + b1)W2 + b2

[0027] H = LayerNorm(FFN(H′)+H′)

[0028] where X ∈ R d represents the input data of the self - attention layer, d f is the transformation dimension in the feed - forward network, and are weight matrices, b1 and b2 are bias vectors, σ represents the GELU activation function, H is the final output of the STTA module, Attn1, Attn2, … Attnh They are the attention of each head, W o is the weight matrix used to connect each attention head.

[0029] Furthermore, in step S3, adding an OpenMax layer after the fully connected layer of the improved ST-GCN model specifically means:

[0030] For each category of samples in the training set, calculate the mean MAV of its activation vector AV and the set of Euclidean distances from MAV;

[0031] Fit a Weibull distribution based on the distance set to generate the cumulative distribution function CDF for each category;

[0032] In the test phase, correct its score vector and generate the probability of the unknown category through the distance between the activation vector AV of the test sample and the mean MAV of the activation vector and the cumulative distribution function CDF.

[0033] Furthermore, in step S2, the specific implementation process of the temporal multi-head self-attention layer T-MHSA is as follows:

[0034] First, project the input features into query, key, and value vectors, defined as follows:

[0035] For the spatio-temporal block of the l-th layer, the query, key, and value vectors at each position (p, t) are respectively expressed as:

[0036]

[0037] Among them, is the feature encoding from the previous layer module, and l represents the current layer number of the ST-GCN module; and are respectively the learnable linear transformation matrices that map the input feature to the query, key, and value spaces;

[0038] Next, in order to obtain the attention weights across time frames, for each pair of spatio-temporal blocks where p = 1, …, n and t, j = 1, …, T, calculate the attention weights through the dot product of the query and key vectors. Specifically, for each time step (p, t), calculate the attention weights through the following formula:

[0039]

[0040] Among them, represents the attention scores between frames in the time dimension, which are normalized by the softmax function to ensure the attention distribution between different time frames;

[0041] Subsequently, for each value vector weight it according to the corresponding attention weights and calculate its weighted sum to generate a new temporal representation

[0042]

[0043] Furthermore, in S2, the specific implementation process of the spatial multi-head self-attention layer S-MHSA is as follows:

[0044] Use three linear transformation matrices and to generate query, key, and value vectors, denoted as and Calculate the attention scores in the spatial dimension as follows:

[0045]

[0046] where represents the attention weight of the node in the spatial dimension;

[0047] Then, perform a weighted sum on the attention score vector in the spatial dimension to generate a new representation in the spatial dimension

[0048]

[0049] Furthermore, the method further includes real-time action recognition optimization:

[0050] Compress the spatio-temporal feature dimension through a pooling layer to reduce the computational complexity;

[0051] In the temporal convolutional layer, use a convolutional operation with a kernel size of 9×1 and perform pooling processing at a specified layer.

[0052] Furthermore, the node connections pair P of the skeletal topology HRCS-11 is: P =

[0053] {(0,1),(1,2),(2,3),(3,4),(1,5),(5,6),(6,7),(1,8),(8,9),(8,10)};

[0054] where 0 is the nose key point, 1 is the neck key point, 2 is the left shoulder key point, 3 is the left elbow key point, 4 is the left wrist key point, 5 is the right shoulder key point, 6 is the right elbow key point, 7 is the right wrist key point, 8 is the mid-hip key point, 9 is the left hip key point, and 10 is the right hip key point.

[0055] Furthermore, the method is applied to the open-set action recognition task in scenarios of human-robot collaboration, intelligent monitoring, or robot control.

[0056] Beneficial effects:

[0057] By introducing a high-order spatio-temporal self-attention mechanism and an open-set recognition algorithm, the present invention has significant beneficial effects in solving the limitations faced by existing action recognition technologies in complex industrial environments. Specifically, the technical solution provided by the present invention can:

[0058] 1. By introducing a high-order spatio-temporal self-attention mechanism, the model can effectively capture the long-term dependence relationships and global context information in the human action sequence, thereby improving the recognition accuracy for complex actions, occlusions, and background changes. In industrial scenarios, the actions of operators are often affected by environmental factors, and the recognition effect of traditional methods is poor in such cases. However, the present invention significantly improves the accuracy and robustness of action recognition under these complex conditions.

[0059] 2. By combining the OpenMax layer, the present invention can distinguish between known and unknown action categories, solving the limitation in the prior art that only known categories in the training set can be recognized. This technology enables the model to make correct recognition judgments when facing unknown actions in an open-set environment, improving the recognition accuracy in the open-set scenario and enhancing the applicability of the system.

[0060] 3. Through the optimized skeletal topology HRCS-11, while retaining the key upper limb and neck nodes, the present invention reduces the annotation of lower limbs and unnecessary joints, effectively reducing data redundancy and occlusion interference, and improving the efficiency and accuracy of action recognition. This structure is particularly suitable for skeleton data modeling in human-robot collaboration scenarios, further improving the application effect of the model in industrial environments.

[0061] 4. By combining a real-time action recognition optimization algorithm, the present invention significantly reduces the computational overhead and improves the real-time performance. This enables the technical solution of the present invention to be applicable not only to static data sets but also to function in a dynamic real-time environment, meeting the real-time response requirements in industrial applications. Description of the Drawings

[0062] Figure 1 It is the overall framework diagram of the IST-GCN action recognition model in the embodiment of the present application;

[0063] Figure 2 It is the skeletal structure diagram of HRCS-11 in the embodiment of the present application;

[0064] Figure 3 It is the model framework diagram of STTA-GCN in the embodiment of the present application;

[0065] Figure 4 This is the MHSA network structure diagram in the embodiment of the present application;

[0066] Figure 5 This is the structure diagram of the open-set action recognition algorithm training module in the embodiment of the present application;

[0067] Figure 6 This is the structure diagram of the open-set action recognition algorithm testing module in the embodiment of the present application. Detailed implementation manners

[0068] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings:

[0069] The present invention can be implemented in many different forms and should not be considered limited to the examples described herein. On the contrary, these embodiments are provided so that this disclosure is thorough and complete, and will fully convey the scope of the present invention to those skilled in the art.

[0070] The present invention proposes an open-set action recognition method based on a high-order spatio-temporal self-attention mechanism, which is applied to a complex open-set environment to solve the problem that existing closed-set action recognition methods cannot effectively recognize actions of unknown categories. This method combines a high-order spatio-temporal self-attention mechanism and improves the model's recognition ability for unknown action categories through an open-set recognition algorithm. This method can be widely applied to fields such as human-robot collaboration, intelligent monitoring, and robot control, and has high practical application value.

[0071] This patent focuses on the action recognition method based on skeletal features and proposes an IST-GCN action recognition model framework that combines a high-order spatio-temporal self-attention mechanism and an open-set recognition algorithm to address the challenges of action recognition in complex human-robot collaboration scenarios. First, based on the optimized skeletal topology HRCS-11, a human body spatio-temporal topology graph is constructed, providing a basis for the spatio-temporal modeling of skeletal action data; second, on the basis of the original ST-GCN model, a high-order spatio-temporal self-attention mechanism is introduced to enhance the model's ability to capture long-term dependency relationships and global context information, thereby improving its accuracy in short-term occluded action recognition; finally, combined with an improved open-set recognition algorithm, the accuracy and robustness of the model in an open environment are improved. The overall framework of the IST-GCN action recognition model is as Figure 1 shown.

[0072] Development environment: The development environment of the present invention uses the Python 2.7 language and is developed in combination with the PyTorch framework. The development platform is the Ubuntu 18.04 operating system, equipped with an NVIDIA GPU (such as GTX 1080Ti) to support deep learning tasks.

[0073] First, based on the spatial topology of the skeleton, this patent connects the skeleton nodes along the time series to form a spatio-temporal topology. For the human-robot collaboration scenario centered around the workbench, this patent proposes an optimized skeleton topology HRCS-11, which contains 11 key skeleton points, omits unnecessary annotations such as the lower limbs, eyes, and ears, and adds a neck key point to more accurately represent the position of the head and body, as Figure 2 shown. The node connection pairs P of the skeleton topology HRCS-11 are: P = {(0,1),(1,2),(2,3),(3,4),(1,5),(5,6),(6,7),(1,8),(8,9),(8,10)}; where 0 is the nose key point, 1 is the neck key point, 2 is the left shoulder key point, 3 is the left elbow key point, 4 is the left wrist key point, 5 is the right shoulder key point, 6 is the right elbow key point, 7 is the right wrist key point, 8 is the mid-hip key point, 9 is the left hip key point, and 10 is the right hip key point.

[0074] Through this simplified skeleton representation method, the influence of data redundancy and occlusion problems is effectively reduced, and the efficiency and accuracy of action recognition are improved. By connecting the skeleton nodes along the time series, a spatio-temporal topology structure is formed, which can be represented as an undirected graph G=(V,E), where V is the set of nodes and E is the set of edges connecting the nodes. The node set is defined as V = {v ti |1≤t≤T,1≤i≤N}, where T represents the number of frames in the time dimension and N represents the number of skeleton nodes in each frame. The edge set E includes connections in the spatial dimension and the time dimension. In the spatial dimension, the edge set E s is defined as E s ={(v ti ,v tj )|1≤t≤T,(i,j)∈P}, where P is the set of pairs of connected nodes in the skeleton structure. Taking HRCS-11 as an example, P = {(0,1),…,(8,10)}. In the time dimension, the edge set E T is defined as E T ={(v ti ,v ki )|1≤k,t≤T,1≤i≤N,|k - t| = 1}. Therefore, the edge set E of the topological structure of the human body in the spatio-temporal dimension can be expressed as E = E s ∪E T .

[0075] On this basis, the ST-GCN model is used to perform pose estimation on the input video, extract the skeleton sequence, and construct a spatio-temporal graph. Specifically, for the input skeleton sequence, the ST-GCN model first normalizes the data through a batch normalization layer. Then, the model uses an ST-GCN module to expand the channel dimension from 3 to 64, and further extracts action features through 9 ST-GCN modules with residual connections. These 9 modules are divided into three groups, and the output channel numbers are 64, 128, and 256 respectively, forming a multi-scale action feature representation, so as to realize human action recognition in the video. Each ST-GCN module contains a GCN graph convolutional unit and a TCN temporal convolutional unit. First, spatial features are extracted through GCN operations, and then TCN is used to model temporal dynamics. To reduce the risk of overfitting, the model randomly discards some human spatio-temporal features with a probability of 50%. At the same time, through the residual mechanism, the original input information is fused into the output. The convolutional kernel size of the temporal convolutional layer is set to 9×1, and pooling is performed on the temporal convolutional layers of the 4th and 7th layers to achieve channel upsampling of human action features. Finally, the features processed by 9 ST-GCN modules pass through the global pooling layer, and the classification probability is output by the SoftMax classifier.

[0076] However, although the ST-GCN model shows high efficiency in learning the spatial and temporal dependencies of non-Euclidean space data (such as skeleton graphs), it lacks flexibility in the feature extraction process and fails to explicitly focus on the spatial topological connections between joint points and their importance at the high-order spatio-temporal level. In complex scenarios of human-robot collaboration, the human action recognition method based on a single camera often loses the action information of the operator due to the temporary occlusion of devices such as robotic arms. The traditional ST-GCN model is difficult to effectively capture global context information due to the lack of a high-order spatio-temporal self-attention module, which may lead to action recognition failure. Therefore, this patent proposes a new module combining ST-GCN and Transformer - the STTA module. This module enhances the ability of ST-GCN to capture long-range spatio-temporal dependencies between joint points by introducing a high-order spatio-temporal self-attention mechanism while retaining the GCN topological feature extraction function.

[0077] The core of the STTA module is to use the multi-head self-attention mechanism to effectively capture human action features, especially high-order spatio-temporal dependencies, in local and global contexts. It mainly includes three main parts: the spatial multi-head self-attention layer (S-MHSA), the temporal multi-head self-attention layer (T-MHSA), and a two-layer feed-forward network (FFN) with residual connections, followed by layer normalization (LN), and finally the skeleton data is embedded through a multi-layer perceptron (MLP). The spatio-temporal features extracted in the ST-GCN block are first projected into three different matrices: and where D q , D k and D v represent the dimensions of queries, keys, and values respectively. Through the multi-head self-attention mechanism (MHSA), these D-dimensional representations are projected into multiple subspaces and H different learned projections are used. For the queries, keys, and values of each set of projections, the output of a single attention head is calculated by the following formula:

[0078]

[0079] After the outputs of all attention heads are concatenated, they are projected back to the original D-dimensional representation, finally forming the following expression:

[0080] MHSA(Q, K, V) = Concat(Attn1, Attn2, … Attn h )W o

[0081] The two-layer feed-forward network (FFN) and layer normalization operations of the STTA module are as follows:

[0082] H′ = LayerNorm(MHSA(X) + X)

[0083] FFN(H′) = σ(H′W1 + b1)W2 + b2

[0084] H = LayerNorm(FFN(H′) + H′)

[0085] where X ∈ R d represents the input data of the self-attention layer, d f is the transformation dimension in the feed-forward network, and are weight matrices, b1 and b2 are bias vectors, σ represents the GELU activation function, and H is the final output of the STTA module.

[0086] Figure 3 Shows the structural improvement of ST-GCN after introducing the STTA module, mainly enhancing the model's feature capture ability in the spatial and temporal dimensions through S-MHSA and T-MHSA.

[0087] Among them, T-MHSA focuses on the temporal dimension, captures the dynamic features of joint points changing over time to obtain global temporal information, thereby enhancing the model's representation of the temporal dependence of actions. The specific structure of MHSA is as Figure 4 shown. The T-MHSA module first projects the input features into query, key, and value vectors, defined as follows:

[0088] For the spatio-temporal block of the l-th layer, the query, key, and value vectors at each position (p, t) are respectively represented as:

[0089]

[0090] where, is the feature encoding from the previous layer module, and l represents the current layer number of the ST-GCN module. and are learnable linear transformation matrices that map the input feature to the query, key, and value spaces respectively.

[0091] Next, in order to obtain the attention weights across time frames, for each pair of spatio-temporal blocks where p = 1, …, n and t, j = 1, …, T, the attention weights are calculated by the dot product of the query and key vectors. Specifically, for each time step (p, t), the attention weights are calculated by the following formula:

[0092]

[0093] where, represents the attention scores between frames in the time dimension, which are normalized by the softmax function to ensure the attention distribution between different time frames.

[0094] Subsequently, each value vector is weighted according to the corresponding attention weight and its weighted sum is calculated to generate a new temporal representation

[0095]

[0096] To capture feature information more comprehensively, the multi-head mechanism is implemented through different self-attention module groups. The self-attention representation of each group where h = 1, …, H are concatenated together and passed through a learnable linear transformation W o to obtain the final result. The multi-head self-attention mechanism allows the model to simultaneously focus on different information patterns and dependencies in different subspaces, thereby enhancing the richness and diversity of feature representation.

[0097] The introduction of the T-MHSA module enables the model to effectively capture long-term dependencies and adapt to complex spatio-temporal patterns. This method is particularly suitable for processing skeleton data containing temporal dynamics, enabling the model to more accurately capture the temporal variation characteristics of actions in action recognition tasks, further improving the accuracy and robustness of the model.

[0098] Similar to T-MHSA, this patent uses three linear transformation matrices in the spatial dimension and to generate query, key, and value vectors, denoted as and These vectors are used to calculate the spatial attention scores as follows:

[0099]

[0100] where represents the attention weights of nodes in the spatial dimension, which are calculated by normalizing the similarity between different nodes to capture the spatial relationships between different parts of the human body. Then, the obtained attention scores are used to perform a weighted sum on the value vectors to generate a new representation in the spatial dimension

[0101]

[0102] S-MHSA captures the structural dependencies of the human skeleton in space by aggregating relevant spatial information that is far from the target node. This method enables the model to focus on the relative position relationships of distant joints in an action, thereby enhancing the model's ability to express action features in the spatial dimension.

[0103] Finally, in the action recognition task, the traditional ST-GCN model calculates the estimated scores for each category through a fully connected layer and processes these scores through the SoftMax function to obtain probability values, ensuring that the sum of the probabilities is 1. However, due to this normalization property, the model has limitations in the open-set recognition problem and cannot identify whether the input object belongs to an unknown category. To solve this problem, this patent improves the ST-GCN model by adding an OpenMax layer to its structure. This layer uses the activation vectors (AVs) generated by samples in the training set in the fully connected layer to fit the cumulative distribution functions (CDFs) for each category. In the test phase, the N-dimensional scores of the test samples are adjusted according to the fitted CDFs and finally converted into an N+1-dimensional probability vector, where N represents the number of known categories, and the additional dimension represents the probability that the input is an unknown category. The improved model will be elaborated in detail from two aspects: the training module and the test module below.

[0104] It is observed that for different samples {a1, a2, …, a N} of the same category A, the scores {a 1A , a 2A , …, a NA} of class A in its AV usually fall within an interval; while for unknown samples {x1, x2, …, x N} that are highly similar to class A samples, their scores {x 1A , x 2A , …, x NA} often deviate from this interval.

[0105] Based on this discovery, the original action recognition model was improved. The improved training module process is as Figure 5 shown, and the specific steps are as follows:

[0106] Step 1: Extract skeletal features. Perform pose estimation on the input video set and extract the corresponding skeleton sequences. The skeleton sequence of each frame is used as input and fed into the IST - GCN model for processing.

[0107] Step 2: Calculate the activation vector and the mean activation vector. Taking the i - th class of actions as an example, calculate the AV of all training samples of this class. If the i - th class contains N i training samples, then N i AVs will be generated. Calculate the mean of the AVs to obtain the mean activation vector (MAV) of this class.

[0108] Step 3: Calculate the distance set. For all correctly classified AVs in the i - th class, calculate their Euclidean distances from the MAV to generate the distance set of the training samples of the i - th class.

[0109] Step 4: Fit the cumulative distribution function. Use this distance set to fit the Weibull distribution to obtain the cumulative distribution function (CDF) of the i - th class, so as to be able to calculate the probability that the distance is less than a certain specific value. This process establishes a distance - based probability model for each class, and these CDFs will be used in the score correction process during the test phase.

[0110] Step 5: Repeat steps 2 to 5 until the corresponding CDFs are generated for all classes in the training set.

[0111] In the training module, by processing the training samples of each class, the corresponding cumulative distribution function (CDF) is fitted, and the mean activation vector (MAV) of each class is calculated. In the test phase, relying on the obtained CDF and MAV, combined with the distance calculation and correction strategy, the scores are further adjusted to improve the model's ability to distinguish known and unknown classes. The improved test module process is as Figure 6 shown, and the specific steps are as follows:

[0112] After inputting a test sample, first calculate its activation vector AV, and then calculate the distances from this AV to the MAVs of each category, thus obtaining N distances. Substitute these distances into the CDF obtained by fitting in the data preprocessing stage respectively to get N probabilities. These probabilities reflect the possibility of the occurrence of the distance between the activation vector of the test sample and the mean activation vectors of each category in the distance distribution of each category. In short, by substituting the distance values into the CDF, the probabilities of the test sample belonging to each known category can be obtained.

[0113] Based on these probabilities, next, the score vector of the test sample is corrected. The specific correction process is as follows: calculate 1 - CDF i (the distance from AV to MAV), that is, the probability corresponding to the distance from AV to the MAV of the i-th category, and this value represents the probability that the test sample does not belong to the i-th category, and use this probability as the correction weight. Among them, CDF i represents the cumulative distribution function of the i-th category, which is obtained by fitting the Weibull distribution in the training stage.

[0114] Then, multiply these N correction weights by the score vector obtained after the test sample is processed by the SoftMax function respectively. Through this process, the original score vector is corrected to the corrected score vector. Finally, the deducted parts in the predicted scores of each category are accumulated to generate the score of this sample belonging to the unknown category, thus obtaining an N + 1-dimensional score vector to achieve the open-set recognition function.

[0115] By introducing a high-order spatio-temporal self-attention mechanism, the present invention enhances the model's ability to capture long-term dependencies and global context information in human action sequences, thereby improving the recognition accuracy of occluded and complex actions. At the same time, combined with the open-set recognition algorithm, the model can effectively distinguish between known and unknown category actions, solve the challenges in open-set action recognition, and enhance the application value and robustness of the model in real industrial scenarios.

[0116] Through the technical solution of the present invention, it is possible to effectively process action recognition tasks with various dynamic changes, occlusions, and complex backgrounds in complex industrial environments. Compared with traditional closed-set recognition methods, the present invention can not only improve the recognition accuracy of known categories, but also accurately identify actions of unknown categories in open-set scenarios, greatly enhancing the applicability and robustness of the model. The open-set action recognition method based on the high-order spatio-temporal self-attention mechanism proposed by the present invention has strong long-term dependence capture ability and global context modeling ability, can effectively handle long-term dependence, occlusion, and complex scene problems, thereby improving the overall performance of action recognition, and providing a more reliable solution for applications in fields such as human-robot collaboration and intelligent manufacturing.

[0117] In summary, this patent proposes an action recognition method for complex human-machine collaboration scenarios, which combines a high-order spatio-temporal self-attention mechanism with an open-set recognition algorithm to form an IST-GCN model framework. Based on the original ST-GCN model, a high-order spatio-temporal self-attention mechanism is fused, including two core modules: spatial multi-head self-attention (S-MHSA) and temporal multi-head self-attention (T-MHSA). This significantly enhances the model's ability to capture long-term dependencies and global context information, and shows a significant performance improvement in dealing with short-term occluded actions. In addition, an open-set recognition algorithm based on the OpenMax function is introduced. By calculating the distance between samples and the mean activation vector and performing distribution fitting, it effectively distinguishes known classes from unknown classes and improves the recognition accuracy in an open environment. Finally, the superiority and applicability of the IST-GCN model are verified through experiments.

[0118] There are many specific application ways of the present invention. The above description is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle of the present invention, several improvements can be made, and these improvements should also be regarded as the protection scope of the present invention.

Claims

1. An open-set action recognition method based on a high-order spatio-temporal self-attention mechanism, characterized in that It includes the following steps: S1 Spatio-temporal modeling based on skeletons: Perform pose estimation on the input video to extract the skeletal topology HRCS-11, where the HRCS-11 contains 11 key skeletal nodes, including upper limb and neck nodes; connect the skeletal nodes along the time series to construct a spatio-temporal topology graph G=(V,E), where the node set V={v ti |t = 1,…,T, i = 1,2,…,N} represents a skeleton sequence containing T frames, each frame contains N joint points, and the edge set E includes spatial connections Es and temporal connections ET; S2 Optimizing the spatio-temporal graph convolutional network ST-GCN based on the high-order spatio-temporal self-attention mechanism: Embedding the STTA module in the spatio-temporal graph convolutional network ST-GCN, where the STTA module includes a spatial multi-head self-attention layer S-MHSA, a temporal multi-head self-attention layer T-MHSA, a feed-forward network FFN with residual connections, and layer normalization LN, to obtain an improved ST-GCN model; Extracting the inter-joint relationship in the spatial domain through the spatial multi-head self-attention layer S-MHSA, and through the T- MHSA capturing the long-term dependence relationship in the temporal domain; S3 Optimizing the improved ST-GCN model based on the open-set recognition algorithm: Adding an OpenMax layer after the fully-connected layer of the improved ST-GCN model, and fitting the cumulative distribution function CDF of each category using the activation vector AV of the samples in the training set; In the test phase, according to the distance between the activation vector AV of the test sample and the mean activation vector MAV of each category, calculate the probability that it belongs to a known category through the cumulative distribution function CDF, and generate an N+1-dimensional probability vector to identify unknown category actions.

2. The method according to claim 1, wherein The construction of the spatio-temporal topology graph in S1 specifically includes: Spatial edge set E s ={(v ti , v tj ) | 1 ≤ t ≤ T, (i, j) ∈ P}, where P is the preset bone node connection pair in the optimized bone topology HRCS-11; v ti , v tj are the nodes in the ti-th and tj-th bone topologies respectively, and T represents the number of frames in the time dimension; Time edge set $E$ T = {(v ti , v ki ) | 1 ≤ k, t ≤ T, 1 ≤ i ≤ N, |k - t| = 1}; The edge set E of the topological structure of the human body in the spatio-temporal dimension is represented as E = E s ∪ E T .

3. The method according to claim 1, wherein The calculation of the multi-head self-attention mechanism in the STTA module includes: Project the input spatio-temporal features into query, key, and value matrices, and calculate the attention weights of each individual attention head through the multi-head attention formula Calculate the attention weights of each individual attention head; After the outputs of all attention heads are concatenated, they are projected back to the original dimensional representation, forming the following expression: MHSA(Q, K, V) = Concat(Attn1, Attn2, … Attn h )W o The operations of the two-layer feed-forward network FFN and layer normalization of the STTA module are as follows: H′ = LayerNorm(MHSA(X)+X) FFN(H′) = σ(H′W1 + b1)W2 + b2 H = LayerNorm(FFN(H ′ ) + H′) where X ∈ R d represents the input data of the self-attention layer, and d f is the transformation dimension in the feed-forward network, and are weight matrices, b1 and b2 are bias vectors, σ represents the GELU activation function, H is the final output of the STTA module, Attn1, Attn2, … Attn h are the attention heads respectively, and W o is the weight matrix used to connect each attention head.

4. The method according to claim 1, characterized in that, In S3, adding an OpenMax layer after the fully-connected layer of the improved ST-GCN model specifically means: For each category of samples in the training set, calculate the mean MAV of its activation vector AV and the set of Euclidean distances from MAV; Fitting the Weibull distribution based on the distance set to generate the cumulative distribution function CDF of each category; In the test phase, through the distance between the activation vector AV of the test sample and the mean MAV of the activation vector and the cumulative distribution function CDF, correct its score vector and generate the probability of the unknown category.

5. The method according to claim 3, characterized in that, In S2, the temporal multi-head self-attention layer The specific implementation process of T-MHSA is: First, project the input features into queries, keys, and values vectors, defined as follows: For the spatio-temporal block of the l-th layer, the queries, keys, and values vectors at each position (p,t) are respectively represented as: Among them, is the feature encoding from the previous layer module, where l represents the current layer number of the ST-GCN module; and are respectively learnable linear transformation matrices that map the input feature to the query, key, and value spaces; Next, in order to obtain attention weights across time frames, for each pair of spatio-temporal blocks where p = 1, …, n and t, j = 1, …, T, the attention weights are calculated by the dot product of the queries and keys vectors; specifically, for each time step (p, t), the attention weights are calculated by the following formula: Among them, represents the attention scores between frames in the time dimension, which are normalized by the softmax function to ensure the attention distribution among different time frames. Subsequently, each value vector is weighted according to the corresponding attention weight to calculate its weighted sum to generate a new temporal representation 6. The method according to claim 5, wherein In S2, the specific implementation process of the spatial multi-head self-attention layer S-MHSA is: Use three linear transformation matrices and to generate query, key, and value vectors, denoted as and respectively. Calculate the attention scores in the computational space as follows: Among them, represents the attention weight of the node in the spatial dimension; Next, a weighted sum is performed on the spatial attention score vectors to generate a new representation of the spatial dimension 7. The method according to claim 1, wherein The method further includes real-time action recognition optimization: Compressing the spatio-temporal feature dimension through a pooling layer to reduce the computational complexity; Adopting a convolution operation with a kernel size of 9×1 in the temporal convolution layer and performing pooling processing at the specified layer.

8. The method according to any one of claims 1 to 5, characterized in that, The node connection pair P of the bone topology structure HRCS-11 is: P={(0,1),(1,2),(2,3),(3,4),(1,5),(5,6),(6,7),(1,8),(8,9),(8,10)}; Among them, 0 is the key point of the nose, 1 is the key point of the neck, 2 is the key point of the left shoulder, 3 is the key point of the left elbow, 4 is the key point of the left wrist, 5 is the key point of the right shoulder, 6 is the key point of the right elbow, 7 is the key point of the right wrist, 8 is the key point of the mid-hip, 9 is the key point of the left hip, and 10 is the key point of the right hip.

9. The method according to claim 1, characterized in that, The method is applied to the open-set action recognition task in the scenarios of human-robot collaboration, intelligent monitoring, or robot control.