Human-machine marshalling intention recognition method and system based on skeleton trajectory and graph convolution
By optimizing the skeleton trajectory extraction algorithm and the improved ST-GCN network model, the delay and robustness issues of skeleton data recognition in human-machine teaming scenarios are solved, fast and accurate recognition of personnel behavior intentions is achieved, and the efficiency of human-machine collaboration is improved.
Patent Information
- Application Number
- CN202510235303.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-02-28
AI Technical Summary
Existing graph convolutional networks based on skeleton data face response delays, insufficient multi-view action recognition accuracy, and robustness defects under dynamic interference in human-machine teaming scenarios, which limits their application in real-time collaborative tasks.
By optimizing the skeleton trajectory extraction algorithm, designing a multi-node information aggregation mechanism, combining the channel attention module and multi-angle data enhancement strategy, the feature extraction effect and recognition speed are improved. The lightweight MobileNet V1 and improved ST-GCN network models are used to enhance the ability to capture spatial and temporal features.
The robot can quickly and accurately identify the behavioral intentions of people, adapt to changes in perspective at different angles, improve the efficiency of human-machine collaboration, and is suitable for real-time collaborative tasks in complex dynamic environments.
Smart Images

Figure CN120708269A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of personnel behavior intention recognition, and in particular to a method and system for human-machine grouping intention recognition based on skeleton trajectory and graph convolution. Background Art
[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.
[0003] With the widespread application of human-robot teaming in collaborative operations, urban search and rescue, target capture, and environmental monitoring, robots' real-time response and collaborative capabilities in complex and dynamic tasks have become key technical requirements. Especially in high-pressure, complex scenarios involving emergency missions, robotic systems must quickly respond to the needs of interacting humans and flexibly adjust their strategies, placing higher demands on robots to autonomously identify human activity intentions.
[0004] Vision-based intent recognition is one of the core technologies for achieving efficient collaboration in human-robot teaming. Existing research constructs recognition models using multimodal features (such as RGB sequences, optical flow, audio signals, and skeleton data). Skeleton-based methods have become mainstream due to their lightweight nature and robustness to background interference. Graph Convolutional Networks (GCNs) have become a leading solution by modeling skeletal topological relationships to extract spatiotemporal dynamic features. However, existing methods face response delays, insufficient multi-view action recognition accuracy, and robustness to dynamic interference in human-robot teaming scenarios, limiting their practical application in real-time collaborative tasks. Summary of the Invention
[0005] To overcome the shortcomings of the aforementioned existing technologies, this paper proposes a method for identifying behavioral intentions in human-robot groupings based on spatiotemporal skeletal trajectory modeling and enhanced graph convolutional networks. By optimizing the skeletal trajectory extraction algorithm, designing a multi-node information aggregation mechanism, integrating a channel attention module, and combining a multi-angle data enhancement strategy, this method improves feature extraction and recognition speed, enabling rapid and accurate recognition of human behavioral intentions by the robot. The method can also adapt to changes in perspective at different angles, improving collaborative efficiency.
[0006] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions:
[0007] First, a method for identifying human-machine grouping intentions based on skeleton trajectories and graph convolution is disclosed, including:
[0008] Obtain real-time video stream data of personnel activities in human-machine teaming scenarios;
[0009] Extract human skeleton sequence information from real-time personnel activity video stream data and generate a 2D skeleton sequence graph of the personnel;
[0010] The graph convolutional network model is used to process the 2D skeleton sequence graph of the personnel and output the predicted personnel behavior intention.
[0011] As a further technical solution, extracting human skeleton sequence information from real-time personnel activity video stream data specifically includes:
[0012] A lightweight MobileNet V1 is used as the backbone network to extract feature maps from real-time personnel activity video stream data;
[0013] Predict skeleton key points based on the extracted feature map;
[0014] Generate partial correlation fields based on the extracted feature maps to describe the spatial connection relationship between various joint points;
[0015] The predicted skeletal key points and partial association fields are assembled into a 2D skeleton sequence graph of the person.
[0016] As a further technical solution, when extracting feature maps, the following steps are specifically included:
[0017] The lightweight MobileNet V1 performs convolution operations on each input channel of the real-time people activity video stream data independently, and then convolutions the output of all channels;
[0018] After being processed by MobileNet V1, the generated feature map contains high-level features related to the human posture in the image.
[0019] As a further technical solution, a graph convolutional network model is used to process the 2D skeleton sequence graph of the person and output the predicted person activity intention, specifically including:
[0020] Perform batch normalization on the 2D skeleton sequence images of people to obtain normalized standard data;
[0021] Use spatiotemporal graph convolution operations on the normalized standard data to extract high-level features in the skeletal posture space and time dimensions;
[0022] After extracting high-level features by calling the global average pooling layer and the fully connected layer, the Softmax classifier is used to obtain the corresponding action classification.
[0023] As a further technical solution, we use spatiotemporal graph convolution operations on the normalized standard data to extract high-level features in the skeletal posture space and time dimensions, including:
[0024] The spatiotemporal graph convolutional network model includes: ST-GCN unit; ST-GCN unit alternately uses GCN graph convolutional network, spatial expansion module SEM, TCN temporal convolutional network and channel attention module SENet to extract features from both spatial and temporal dimensions;
[0025] GCN is used to process the spatial features of skeleton data, that is, the spatial relationship between each bone point. Through graph convolution operation, GCN can fuse the spatial information of each bone point with the information of its adjacent points to obtain richer spatial features;
[0026] The convolution operation in the spatial extension module (SEM) not only processes the local relationship between adjacent nodes, but also captures the interaction between the first non-adjacent nodes by expanding the skeleton graph, which is used to enhance the model's ability to capture the long-range dependency relationship between skeleton points.
[0027] TCN effectively captures dynamic dependencies in time series through convolution operations and is used to process the temporal dimension of skeleton data, that is, how the model understands the changes in the skeleton in different time steps.
[0028] As a further technical solution, the graph convolutional network model introduces SENet as a channel attention mechanism to optimize the network's attention to key information;
[0029] SENet adaptively learns the importance of each channel, strengthening the model's focus on key information and ignoring irrelevant features;
[0030] Specifically, SENet compresses each channel through global average pooling (GAP) and generates a global description of the feature map of each channel through pooling operations.
[0031] As a further technical solution, after extracting high-level features by calling the global average pooling layer and the fully connected layer, the Softmax classifier is used to obtain the corresponding action classification, specifically including:
[0032] After the spatiotemporal graph convolution layer, the generated feature map is input to the global average pooling (GAP) layer; the GAP layer operates by averaging the features of each time step and compressing the feature vector into a fixed-length representation;
[0033] Next, the features output by GAP are passed to the fully connected layer, which performs a linear transformation through the weight matrix and bias term to map the input features to a higher-dimensional space and prepare appropriate outputs for the classification task;
[0034] Finally, the Softmax classifier is used to convert the feature vector output by the fully connected layer into a probability distribution of the category. Softmax returns the category with the highest probability as the predicted intent category.
[0035] Secondly, a human-machine grouping intention recognition system based on skeleton trajectory and graph convolution is disclosed, including:
[0036] The data acquisition module is configured to: acquire real-time personnel activity video stream data in a human-machine teaming scenario;
[0037] The person 2D skeleton sequence diagram generation module is configured to: extract human skeleton sequence information from real-time person activity video stream data to generate a person 2D skeleton sequence diagram;
[0038] The personnel activity intention prediction module is configured to: use the graph convolutional network model to perform calculations on the personnel 2D skeleton sequence graph and output the predicted personnel activity intention.
[0039] One or more of the above technical solutions have the following beneficial effects:
[0040] This paper optimizes a lightweight human pose estimation model, significantly improving its inference speed while maintaining accuracy. This optimization enables the model to run in real time in resource-constrained, dynamic environments, meeting rapidly changing behavior recognition requirements. Especially on edge devices, the optimized model can efficiently handle complex action recognition tasks, thereby enhancing the practical application value of human pose estimation.
[0041] The Spatial Extension Module (SEM) designed in this paper enhances the feature extraction capabilities of spatial skeleton sequences, especially in complex and rapidly changing dynamic environments. By incorporating information from non-adjacent nodes, this module enables the network to not only focus on skeleton points in the local neighborhood but also effectively capture the spatial dependencies between distant nodes. This design improves the recognition accuracy of complex activities and effectively avoids the information loss problem in traditional methods, thereby greatly enhancing the accurate recognition of complex activity patterns.
[0042] This paper introduces the SENet module as a channel attention mechanism, significantly improving the model's focus on key information. This mechanism adaptively adjusts the weights of each channel, enabling the model to more accurately identify and extract important skeletal features, especially when processing complex dynamic activities. By optimizing the transfer of channel information, the channel attention mechanism improves the model's precision and robustness, significantly enhancing the accuracy and reliability of behavioral intent recognition.
[0043] To address the issue of perspective sensitivity in human-machine collaboration scenarios, this paper proposes a robust training framework based on a multi-view data augmentation strategy. By constructing a multi-view data space that includes geometric rotation (±45°), dynamic lighting, and complex background perturbations, this paper systematically addresses the model's generalization shortcomings caused by traditional single-view training data. By integrating limb motion saliency feature extraction with cross-view feature alignment constraints, this strategy enables the model to maintain consistency in action representation under multiple observation conditions, such as frontal and side views, thereby reducing recognition errors caused by angle changes.
[0044] Advantages of additional aspects of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0046] Figure 1 is a flow chart of the method of the present invention;
[0047] Figure 2 This is a schematic diagram of the network structure of the lightweight OpenPose of the present invention;
[0048] Figure 3 This is a schematic diagram of the network structure of the optimized ST-GCN of the present invention;
[0049] Figure 4 This is a schematic diagram of the structural design principle of the space expansion module SEM of the present invention;
[0050] Figure 5 This is a schematic diagram of the network structure of the present invention based on the channel attention mechanism SENet;
[0051] Figure 6 This is the robot experimental platform used in the present invention. DETAILED DESCRIPTION
[0052] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.
[0053] It should be noted that the terms used herein are for describing particular embodiments only and are not intended to limit the exemplary embodiments according to the present invention.
[0054] In the absence of conflict, the embodiments of the present invention and the features thereof may be combined with each other.
[0055] Since human-robot teaming scenarios in complex dynamic environments are highly uncertain, the actual interactions between humans and robots are often extremely complex. Therefore, it is necessary to improve the adaptability of algorithm models in such scenarios so that the identification of human behavioral intentions is more accurate and practical. Based on this, this embodiment proposes a method for identifying human-robot teaming behavioral intentions suitable for complex environments. Through this method, the behavioral intentions of humans can be accurately identified in real time, and accurate operating instructions can be provided to robots, thereby achieving efficient human-robot collaboration. The present invention is described in detail below with reference to specific embodiments.
[0056] Example 1
[0057] This embodiment discloses a method for identifying human-machine grouping intention based on skeleton trajectory and graph convolution, including:
[0058] S1: RealSense D435i camera is used to capture real-time video stream data of people's activities in complex environments.
[0059] Step 101: Use a RealSense D435i camera to capture a color video stream of the scene where the device is located at a resolution of 640×480 and 30 frames per second.
[0060] Step 102: The collected video stream data is transmitted in real time to a processing system based on the NVIDIA Jetson AGX Orin platform and equipped with the Ubuntu 20.04 operating system for further use in subsequent posture estimation and behavior recognition.
[0061] Step 103: Determine the pixel coordinate system of the video stream according to the resolution of the video stream, which provides a basis for the subsequent generation and standardization of the 2D skeleton sequence diagram and ensures the accuracy of the skeleton point positioning.
[0062] S2: This embodiment uses the optimized lightweight OpenPose human pose estimation model to process the video stream data captured in step S1, predicts the skeleton points of the person from the bottom up, extracts the two-dimensional coordinates and confidence levels of the skeleton points, and generates a 2D skeleton sequence diagram of the person.
[0063] In this step, the lightweight OpenPose model is trained using the COCO dataset released by the Microsoft team. This dataset contains over 200,000 annotated images, covering multi-angle human poses, object positioning, and semantic segmentation information. The human key point annotation data is specifically used to improve the accuracy of skeletal joint detection. After training, the model outputs a neural network weight file (.pth) within the PyTorch framework. After structured quantization, the PyTorch model is converted to the TensorRT inference engine (.trt) using the ONNX intermediate format, ensuring high real-time performance for subsequent deployment on the NVIDIA Jetson AGX Orin embedded platform.
[0064] The network is based on MobileNet v1 as the backbone and uses dilated convolution to improve feature extraction efficiency while reducing computational complexity. During training, a preliminary two-stage network (including initialization and refinement) was employed. By continuously optimizing the network architecture and training process, a model capable of running efficiently on edge devices was ultimately obtained.
[0065] The obtained TensorRT model file is used for real-time human posture estimation, and the inference process is accelerated by the TensorRT engine, significantly improving the inference speed and efficiency. Specifically, the model first generates preliminary estimated keypoint heatmaps through the keypoint prediction branch: this branch uses the convolution layer to optimize the feature map so that the probability value of each bone point is concentrated in its corresponding area, and forms the probability distribution of the key points through the Softmax activation function, thereby locating key parts such as the nose, eyes, ears, shoulders, elbows, wrists, hips, knees and ankles. At the same time, the model synchronously generates a vector field that describes the spatial connection relationship between joint points through the part affinity fields (PAFs) prediction branch: this branch uses the convolution layer to efficiently capture the direction and correlation strength between joint points, providing a spatial dependency basis for subsequent bone point connections.
[0066] During the refinement stage, the model iteratively optimizes the initially generated keypoint heatmaps and PAFs. Through multi-scale feature fusion and confidence weighting, it eliminates redundant detections and corrects mismatched joint connections. Finally, combining the optimized heatmaps and PAFs, the model uses the Hungarian algorithm to achieve global optimal matching of joints, assembling a complete and accurate 2D human skeleton sequence diagram, providing high-precision data support for subsequent behavioral intention recognition.
[0067] Specifically, such as Figure 2As shown in the figure, the steps for extracting skeleton points of a person using the lightweight OpenPose human pose estimation model include:
[0068] S2.1: Lightweight MobileNet V1 is used as the backbone network for feature map extraction. This embodiment significantly reduces the complexity of the model by introducing depthwise separable convolution to replace traditional convolution. Specifically, depthwise separable convolution consists of two parts: depthwise convolution and pointwise convolution: first, each channel of the input feature map (multi-channel image features after preprocessing such as image normalization and size adjustment) is independently subjected to depth convolution to extract spatial dimension features. Subsequently, the features between channels are linearly combined through point-by-point convolution to further fuse semantic information. Through this structure, MobileNetV1 greatly reduces the amount of computation and parameters while ensuring feature extraction capabilities. After processing by the backbone network, the generated high-level feature map contains spatial and semantic information related to human posture, providing a robust data basis for subsequent key point prediction.
[0069] S2.2: Predict skeletal key points. In the key point prediction branch, the convolution layer inside the lightweight OpenPose model further extracts and optimizes the feature map output by the backbone network, and enhances the response characteristics of the key point area through multi-level convolution operations, so that the probability value of each skeletal point is concentrated at the spatial position of its corresponding node. Specifically, the network outputs the confidence of the existence of the key point for each pixel position of the feature map, and normalizes the confidence within the entire image through the Softmax activation function to generate a two-dimensional probability distribution map. This probability distribution map can characterize the predicted probability density of key points at different positions. For example, the probability peak area is the predicted joint point coordinates (such as elbows, knees, etc.). In subsequent steps, the model extracts the coordinate point with the highest confidence from the probability distribution map as the final detection result through non-maximum suppression (NMS) or threshold screening, thereby accurately locating the key points of the human skeleton.
[0070] S2.3: Generate partial association fields (PAFs). In parallel with the key point prediction branch, the model generates partial association fields (PAFs) that describe the spatial connection relationship between joint points through the PAFs prediction branch. Specifically, this branch is based on the backbone network feature map and uses depth-wise separable convolution to efficiently extract directional association features between joint points (such as the spatial orientation of the shoulder-elbow connection line), and outputs multi-channel PAFs. Each PAF channel corresponds to a predefined pair of joint points, and the two-dimensional vector and confidence of its pixel position characterize the geometric connection strength and direction between adjacent joint points. For example, the PAF channel of the elbow and wrist will generate a high-confidence direction vector on the path connecting the two, providing spatial constraints for subsequent skeletal point connections and ensuring the consistency of the human body's topological structure.
[0071] S2.4: Connect the skeleton points and assemble the skeleton graph. Based on the predicted key point heat map and partial association fields (PAFs), the model first performs association matching on predefined joint point pairs (such as shoulder-elbow, hip-knee): by integrating the direction vectors in the PAFs along the joint point connection path, calculating the geometric consistency score, and combining the confidence of the joint points in the heat map to screen effective connections. Subsequently, the Hungarian algorithm is used to globally optimize the matching results to eliminate contradictory connections such as limb crossing, ensuring that the skeleton topology conforms to the human anatomical structure. Finally, the model outputs a 2D skeleton sequence graph containing the two-dimensional coordinates and connection relationships of each joint point. Its spatiotemporal information provides high-precision structured data support for subsequent recognition of personnel behavior intentions.
[0072] S3: Convert the skeleton sequence graph (original joint point coordinate sequence) extracted in step S2 into a spatiotemporal graph structure and input it into the improved spatiotemporal graph convolutional network (ST-GCN) model. The specific model structure is shown in Figure 3 、 Figure 4 and Figure 5 Based on the above optimization features, the model can accurately predict the activity intentions of personnel in the context of human-robot teaming, thereby enabling the robot to effectively perceive and predict personnel activities.
[0073] It should be noted that the skeleton sequence graph refers to the raw skeletal motion data, such as the time series of joint coordinates (x, y). The human skeleton graph is a structured model of skeletal data, representing the connection relationship between nodes (joints) and edges in the form of a graph, which is used to adapt to the input requirements of the graph convolutional network (ST-GCN).
[0074] Therefore, we construct a human skeleton graph: the skeleton sequence is usually represented by the coordinates of each joint in each frame (such as 2D (x, y) or 3D (x, y, z)). Through the spatiotemporal graph structure, the skeleton sequence can form a hierarchical topological representation. The specific method is as follows:
[0075] For a skeleton sequence containing N joints and T frames, an undirected spatiotemporal graph G = (V, E) is constructed.
[0076] Where: V is a node set, representing the coordinates of the joint points in all time frames, defined as v ti is the coordinate of the i-th joint point in the t-th frame, and d is the coordinate dimension (2D or 3D).
[0077] E is an edge set, which is divided into spatial edges (joints connected within the same frame) and temporal edges (joints connected across frames). S is the spatial connection between joints in the same frame, which is determined by the predefined anatomical relationship H (e.g. ); Time side E T For the temporal connection of the same joint point between consecutive time frames (such as ). Specifically expressed as:
[0078]
[0079] Among them, v ti represents the i-th joint point in time frame t, v tj represents the jth joint point in time frame t, H represents the anatomical connection relationship between human joints, and v (t+1)i represents the i-th joint point in time frame t+1.
[0080] The key points of the human skeleton are represented as (x, y) coordinates in the spatial coordinate system. In the graph convolution process, the human skeleton information is processed through spatiotemporal graph convolution, and the formula is as follows:
[0081]
[0082] Among them, f ω (v ti ) represents node v ti The output features, N(v ti ) represents node v ti Neighborhood (including spatial domain E S and time domain E T ), m(v tj ) represents the feature information of the neighborhood nodes, and ω is a learnable weight function used to adjust the importance of different connections.
[0083] About the spatial expansion module: In traditional graph convolutional networks (GCNs), spatial graph convolution models usually only focus on the information of the node itself and its adjacent nodes, which may lead to the loss of skeleton information, thereby affecting the accuracy of behavior recognition. In order to solve this problem, the present invention designs a spatial expansion module (SEM), which not only focuses on the skeleton information of the node itself and adjacent nodes when calculating the spatial features of the node, but also introduces the skeleton information of the first non-adjacent node at the same time. In this way, the spatial expansion module can capture longer-distance dependencies between skeleton points, thereby enhancing the spatial modeling capabilities between nodes in the skeleton graph, avoiding information waste, and improving the accuracy of behavior recognition.
[0084] In the specific implementation, the spatial expansion module extends the traditional graph convolution calculation method to consider the information of non-adjacent nodes. By introducing non-adjacent node information, the spatial expansion module expands the calculation formula to:
[0085]
[0086] Where N(v i ) is the node v i The set of adjacent nodes, f in (v j ) is the node v j Input features, W ij is the weight between nodes. The second term represents the first non-adjacent node v k Information, W ik For node v i and non-adjacent nodes v k In this way, the module can not only focus on the information of adjacent nodes, but also effectively transmit the information of distant nodes, enhancing the modeling ability of spatial features.
[0087] In addition, to further optimize the feature extraction capabilities of the spatial expansion module, a skeleton point weighting mechanism is designed to adjust the influence of each node in the calculation process. This weighting mechanism dynamically assigns weights by calculating the distance between nodes. The formula is:
[0088]
[0089] Among them, α ik Represents node v i and non-adjacent nodes v k The weight between ik For node v i and v k The distance between them, μ is the center value of the distance, and β is a hyperparameter that controls the weight distribution. Through this weighting mechanism, the model can adjust the transmission of information according to the distance, thereby extracting spatial features more accurately.
[0090] About channel attention mechanism (SENet module): SENet module is introduced as a channel attention mechanism to enhance the network's attention to important features and ignore the influence of irrelevant channels. SENet module uses the "compression-excitation" mechanism to first compress the global information of each channel through global average pooling and calculate the global description z of the cth channel. c :
[0091]
[0092] Among them, H and W are the height and width of the feature map, c represents the number of channels, and f c (n) (i, j) represents the eigenvalue of the cth channel at position (i, j) in the nth layer.
[0093] Then, through the first fully connected layer (weight matrix W1∈R C×C / r , bias term b1) for the global descriptor z c Perform dimensionality reduction and apply the ReLU activation function:
[0094] u c =ReLU(W1·z c +b1)
[0095] Where r is the reduction ratio (usually set to 16) and C / r is the number of channels after dimensionality reduction.
[0096] Through the second fully connected layer (weight matrix W2∈R C / r×C , the bias term b2) restores the channel dimension and applies the Sigmoid function to generate the attention weights:
[0097] s c =σ(W2·u c +b2)
[0098] Where σ is the Sigmoid function, s c ∈[0,1] represents the importance weight of the c-th channel.
[0099] Finally, the attention weight s c Multiply the original feature map channel by channel to obtain the optimized feature map:
[0100]
[0101] In this way, the SENet module enables the network to adaptively adjust the weights of the channels, thereby more effectively capturing key information and improving the accuracy of behavioral intention recognition.
[0102] The improved ST-GCN model combines a spatial expansion module and a channel attention mechanism to improve spatial modeling capabilities while also enhancing the model's dynamic processing capabilities in time series. The improved network structure includes 9 main spatiotemporal units, with output channels of 64, 64, 64, 128, 128, 128, 256, 256, and 256, respectively. The network performs convolution operations at each layer to extract more refined features. The improvement of the network structure not only improves the computational efficiency of the model, but also improves the prediction accuracy of activity intentions. The formula is as follows:
[0103]
[0104] Among them, f out represents the output feature, W k is the weight of the convolution operation, A k Represents the spatial feature map obtained by convolution operation, M k Indicates the characteristics of the activity, Represents a convolution operation.
[0105] Optimization and Prediction: The optimized ST-GCN model described above enables the network to more accurately capture the spatial and temporal relationships between different skeleton points. During the refinement phase, the network further refines the initially estimated skeleton point positions to ensure accurate identification of each skeleton point. Ultimately, the model outputs an accurate 2D skeleton sequence graph and uses a spatiotemporal graph convolutional network to predict activity intent.
[0106] Specifically, such as Figure 3 As shown in the figure, the steps for predicting personnel activity intentions by the optimized GCN model include:
[0107] S3.1: Batch normalize the skeleton input data. Before the data is input to the model, the skeleton data is first batch normalized. It is divided into the following stages:
[0108] During the training phase, the mean μ is calculated for each batch of skeleton data B and variance in m is the batch size, x i is the i-th sample in the batch.
[0109] The "batch" in batch normalization refers to the data processing method during training. During training, the model updates the normalization parameters by calculating the mean and variance of small batches of examples. During prediction, the average parameters accumulated during training are directly used, regardless of the number of batches of input data.
[0110] Subsequently, the data is normalized and scaled to restore expressive power, as follows:
[0111]
[0112] where ò is a numerical stability constant, and γ (scaling factor) and β (bias factor) are parameters learned during training.
[0113] In the prediction phase, the global mean μ accumulated during the training process is used global and variance (updated by moving average), rather than the current batch statistics, the formula is:
[0114]
[0115] Through the above operations, batch normalization ensures the stability of the input distribution of the model during training and prediction, thereby improving the feature extraction efficiency and generalization ability of the spatiotemporal graph convolutional network.
[0116] S3.2: Use a series of spatiotemporal graph convolution operations to extract high-level features in the skeletal pose space and time dimensions.
[0117] The ST-GCN unit with the space expansion module is as follows Figure 4 As shown in the figure, after batch normalization, the input data is fed into the ST-GCN unit with the spatial expansion module. This unit alternates between GCN (graph convolutional network), spatial expansion module (SEM), and TCN (temporal convolutional network) to extract features from both spatial and temporal dimensions.
[0118] Specifically, GCN is used to process the spatial features of skeleton data, that is, the spatial relationship between each bone point. The specific graph convolution operation can be expressed as:
[0119]
[0120] in, is the normalized adjacency matrix, H (l) is the input feature matrix of the lth layer, W (l) is the learned weight matrix, and σ is the activation function. Through graph convolution operations, GCN can fuse the spatial information of each skeleton point with the information of its neighboring points to obtain richer spatial features.
[0121] The spatial extension module (SEM) is used to enhance the model’s ability to capture long-range dependencies between skeleton points. In the spatial extension module, the graph convolution operation not only processes the local relationship between nodes, but also captures the interaction between remote nodes by expanding the skeleton graph. Specifically, for each skeleton node v i , whose characteristics are determined by the adjacent nodes v j and the first non-adjacent node v k The formula can be expressed as:
[0122]
[0123] Among them, f in (v j ) and f in (v k ) is the node v j and v k Input features, W ij and W ik is the connection weight, N (vi) Represents node v i In this way, the spatial expansion module can capture the global information between skeleton points in the spatial dimension, thereby improving the limitation of traditional methods that only rely on local information.
[0124] TCN is used to process the time dimension of skeleton data, that is, how the model understands the changes in the skeleton at different time steps. TCN effectively captures the dynamic dependencies in the time series through one-dimensional convolution operations. The specific convolution operation formula is as follows:
[0125]
[0126] Among them, x t-k+1 is the input feature at time step t-k+1, w k is the weight of the convolution kernel, y t is the output feature. TCN captures the dynamic evolution of skeleton points through this sliding window method, identifies the temporal pattern of actions, and ensures that the model can handle long-range dependencies in time series data.
[0127] The ST-GCN unit that introduces the spatial expansion module and channel attention mechanism is as follows Figure 5 As shown. Based on the spatial expansion module, this embodiment introduces SENet as a channel attention mechanism to further optimize the network's focus on key information. SENet adaptively learns the importance of each channel, strengthens the model's focus on key information, and ignores irrelevant features. Specifically, SENet compresses each channel through global average pooling (GAP), converting the feature map F of each channel into (c) Generate a global description z through pooling operation c :
[0128]
[0129] Among them, F (c) (i, j) is the feature map on channel c, H and W are the height and width of the image. Then, the channel attention coefficient z is calculated through two fully connected layers and activation functions (ReLU and Sigmoid):
[0130] z=σ(W2·ReLU(W1·z c ))
[0131] Finally, the network weights the features of each channel by the generated attention coefficients, which improves the robustness and accuracy of the network in action recognition.
[0132] S3.3: Map the features extracted by the spatiotemporal graph convolutional network into action category probabilities through a global average pooling layer, a fully connected layer, and a Softmax classifier. The specific steps are as follows:
[0133] First, the feature map generated by the spatiotemporal graph convolution layer is input to the global average pooling (GAP) layer. The GAP layer operates by slicing the features at each time step. The averaging process is performed and the calculation formula is:
[0134]
[0135] Where T is the time dimension, representing the number of time steps. This layer operation eliminates the complexity of the time dimension T.
[0136] The above spatiotemporal graph convolution layer refers to the series of spatiotemporal graph convolution operations in S3.2, specifically Figure 3 Medium GCN+SEM+TCN block and GCN+SENet+SEM+TCN block.
[0137] Next, the feature F output by GAP is GAP Through the learnable weight matrix and bias Perform a linear transformation to map the input features into a higher-dimensional space and prepare appropriate outputs for the classification task:
[0138] F E =W·F GAP +b
[0139] Where K is the number of behavior categories in the classification task.
[0140] Finally, the Softmax classifier is used to transform the feature vector F output by the fully connected layer E Converted to the probability distribution of the category. The formula of the Softmax function is as follows:
[0141]
[0142] Finally, Softmax outputs the category with the highest probability as the predicted intent category.
[0143] Example 2
[0144] The purpose of this embodiment is to provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the program.
[0145] Example 3
[0146] The purpose of this embodiment is to provide a computer-readable storage medium.
[0147] A computer-readable storage medium stores a computer program, which, when executed by a processor, performs the steps of the above method.
[0148] Example 4
[0149] The purpose of this embodiment is to provide a human-machine grouping intention recognition system based on skeleton trajectory and graph convolution, including:
[0150] The data acquisition module is configured to: acquire real-time personnel activity video stream data in a human-machine teaming scenario;
[0151] The person 2D skeleton sequence diagram generation module is configured to: extract human skeleton sequence information from real-time person activity video stream data to generate a person 2D skeleton sequence diagram;
[0152] The personnel activity intention prediction module is configured to: use the graph convolutional network model to perform calculations on the personnel 2D skeleton sequence graph and output the predicted personnel activity intention.
[0153] The sub-technical solution of this embodiment proposes an innovative solution to the problems existing in existing behavioral intention recognition methods, such as slow recognition speed, low accuracy, and being affected by the relative angle between the person and the robot. First, the present invention optimizes the human posture estimation model, significantly improving the estimation speed while ensuring the estimation accuracy; secondly, the introduction of the space expansion module and the channel attention mechanism effectively improves the intention recognition algorithm's ability to extract the spatial features of skeletal information, and optimizes the weight distribution of important features, thereby ignoring irrelevant information and improving the recognition accuracy; in addition, the optimized data set composition supports the recognition of multi-angle personnel activity intentions, and improves the robustness of multi-view recognition. The present invention is suitable for the recognition of personnel behavioral intentions in human-machine grouping scenarios, and has significant practical value, especially in the fields of military and industrial robot collaboration. It has important application prospects.
[0154] Example 5
[0155] The purpose of this embodiment is to provide a computer program product containing instructions, which, when running on a computer, enables the computer to execute the methods and functions involved in any of the above embodiments.
[0156] The steps involved in the apparatus of the above embodiment correspond to those of the method embodiment 1. For detailed implementation, please refer to the relevant description of embodiment 1. The term "computer-readable storage medium" should be understood to mean a single medium or multiple media containing one or more instruction sets; it should also be understood to include any medium capable of storing, encoding, or carrying an instruction set for execution by a processor and causing the processor to perform any method of the present invention.
[0157] Those skilled in the art will appreciate that the modules or steps of the present invention described above can be implemented using a general-purpose computer device. Alternatively, they can be implemented using program code executable by a computing device, which can then be stored in a storage device and executed by the computing device. Alternatively, they can be fabricated into separate integrated circuit modules, or multiple modules or steps can be fabricated into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.
[0158] Although the above describes the specific embodiments of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without any creative work are still within the scope of protection of the present invention.
Claims
1. A human-machine grouping intention recognition method based on skeleton trajectory and graph convolution is characterized by: include: Obtain real-time video stream data of personnel activities in human-machine teaming scenarios; Extract human skeleton sequence information from real-time personnel activity video stream data and generate a 2D skeleton sequence graph of the personnel; The graph convolutional network model is used to process the 2D skeleton sequence graph of the personnel and output the predicted personnel activity intention.
2. The method for identifying human-machine grouping intention based on skeleton trajectory and graph convolution as claimed in claim 1 is characterized in that: Extracting human skeleton sequence information from real-time personnel activity video stream data specifically includes: A lightweight MobileNet V1 is used as the backbone network to extract feature maps from real-time personnel activity video stream data; Predict skeleton key points based on the extracted feature map; Generate partial correlation fields based on the extracted feature maps to describe the spatial connection relationship between various joint points; The predicted skeletal key points and partial association fields are assembled into a 2D skeleton sequence graph of the person.
3. The method for identifying human-machine grouping intention based on skeleton trajectory and graph convolution as claimed in claim 2 is characterized in that: When extracting feature maps, it specifically includes: The lightweight MobileNet V1 performs convolution operations on each input channel of the real-time people activity video stream data independently, and then convolutions the output of all channels; After being processed by MobileNet V1, the generated feature map contains high-level features related to the human posture in the image; Prioritize using a graph convolutional network model to process the 2D skeleton sequence graph of the person and output the predicted person activity intention, specifically including: Perform batch normalization on the 2D skeleton sequence images of people to obtain normalized standard data; Use spatiotemporal graph convolution operations on the normalized standard data to extract high-level features in the skeletal posture space and time dimensions; After extracting high-level features by calling the global average pooling layer and the fully connected layer, the Softmax classifier is used to obtain the corresponding action classification.
4. The method for identifying human-machine grouping intention based on skeleton trajectory and graph convolution as claimed in claim 3 is characterized in that: The normalized standard data is subjected to spatiotemporal graph convolution operations to extract high-level features in the skeletal posture space and time dimensions, including: The spatiotemporal graph convolutional network model includes: ST-GCN unit; ST-GCN unit alternately uses GCN graph convolutional network, spatial expansion module SEM, TCN temporal convolutional network and channel attention module SENet to extract features from both spatial and temporal dimensions; GCN is used to process the spatial features of skeleton data, that is, the spatial relationship between each bone point. Through graph convolution operation, GCN can fuse the spatial information of each bone point with the information of its adjacent points to obtain richer spatial features; The convolution operation in the spatial extension module (SEM) not only processes the local relationship between nodes, but also captures the interaction of remote nodes by expanding the skeleton graph, which is used to enhance the model's ability to capture the long-range dependency relationship between skeleton points. TCN effectively captures dynamic dependencies in time series through convolution operations and is used to process the temporal dimension of skeleton data, that is, how the model understands the changes in the skeleton in different time steps.
5. The method for identifying human-machine grouping intention based on skeleton trajectory and graph convolution as claimed in claim 1 is characterized in that: The graph convolutional network model introduces SENet as a channel attention mechanism to optimize the network's attention to key information; SENet adaptively learns the importance of each channel, strengthening the model's focus on key information and ignoring irrelevant features; Specifically, SENet compresses each channel through global average pooling (GAP) and generates a global description of the feature map of each channel through pooling operations.
6. The method for identifying human-machine grouping intention based on skeleton trajectory and graph convolution as claimed in claim 1, characterized in that: After extracting high-level features by calling the global average pooling layer and the fully connected layer, the Softmax classifier is used to obtain the corresponding action classification, including: After the spatiotemporal graph convolution layer, the generated feature map is input to the global average pooling (GAP) layer; the GAP layer operates by averaging the features of each time step and compressing the feature vector into a fixed-length representation; Next, the features output by GAP are passed to the fully connected layer, which performs a linear transformation through the weight matrix and bias term to map the input features to a higher-dimensional space and prepare appropriate outputs for the classification task; Finally, the Softmax classifier is used to convert the feature vector output by the fully connected layer into a probability distribution of the category. Softmax returns the category with the highest probability as the predicted intent category.
7. A human-machine grouping intention recognition system based on skeleton trajectory and graph convolution is characterized by: include: The data acquisition module is configured to: acquire real-time personnel activity video stream data in a human-machine teaming scenario; The person 2D skeleton sequence diagram generation module is configured to: extract human skeleton sequence information from real-time person activity video stream data to generate a person 2D skeleton sequence diagram; The personnel activity intention prediction module is configured to: use the graph convolutional network model to perform calculations on the personnel 2D skeleton sequence graph and output the predicted personnel activity intention.
8. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the method described in any one of claims 1 to 6 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method described in any one of claims 1 to 6 are performed.
Citation Information
Patent Citations
Graph convolution action recognition method, device and equipment based on 2S-AGCN
CN113642400A
Behavior recognition model training method and device, equipment and storage medium
CN114792401A
Behavior detection method based on lightweight OpenPose space-time diagram network
CN115546894A
Expansion action recognition method based on graph convolutional network
CN115937965A
Man-machine cooperation method based on multi-scale image convolutional neural network
CN116665312A
Cited By
Orthopedic surgery key point positioning method and system based on anatomical structure feature recognition
CN121600073A
Behavior recognition method and system based on artificial intelligence
CN122049981A
An artificial intelligence-based behavior recognition method and system
CN122049981B