Human-centered visual perception method facing human-machine cooperation scene

Through the enhanced human mesh recovery algorithm and the combination of uncertainty estimation with the accumulation trigger, the problem of inaccurate human body shape and posture estimation in the occlusion environment is solved, and real-time and stable human body state and intention recognition is achieved in human-computer collaboration.

CN120299081APending Publication Date: 2025-07-11NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510235723.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Existing human-machine collaboration technologies are difficult to accurately estimate the shape and posture of human bodies in occlusion environments, and real-time motion recognition is unstable, resulting in challenges to understanding human intentions and security.

Method used

The human mesh recovery algorithm with prior knowledge enhancement is used to combine image data enhancement to learn pose prior knowledge through 2D key point regression branches, and real-time action recognition is achieved using uncertainty estimation and accumulation triggers to reduce errors in the occlusion environment.

Benefits of technology

It improves the accuracy and real-time estimation of human body shape and posture in occlusion environments, reduces the error rate in action recognition, and enhances the robustness and stability of human-computer collaboration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299081A_ABST
    Figure CN120299081A_ABST
Patent Text Reader

Abstract

The invention discloses a human-centered visual perception method for a human-machine cooperation scene, and the method comprises the following steps: S1, predicting the shape and posture parameters of a human body in a shielding scene through a priori knowledge enhanced human body grid recovery algorithm; s2, recombining attitude parameters predicted by the priori knowledge enhanced human body grid recovery algorithm into a skeleton spatio-temporal topological graph in an HRCA-11 format, performing data enhancement processing, and inputting the skeleton spatio-temporal topological graph into action recognition models with different weights for joint reasoning to obtain an action type and uncertainty of single recognition; and S3, sending the uncertainty and the action category into an accumulation trigger, and outputting a smooth and stable action category in combination with a historical recognition record. According to the method, priori knowledge of human body postures is learned through a 2D key point regression branch, errors of the model in a shielding environment are reduced, real-time action recognition is achieved through a trigger accumulation mode, and the accuracy and stability of the model during real-time recognition are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of human-machine collaboration, in particular to the human body shape and pose recognition technology in the presence of occlusion in human-machine collaboration. Specifically, it is a method for accurately obtaining the state and intention of humans in an occluded environment. Background Art

[0002] Human-machine collaboration is a human-centered intelligent manufacturing mode, which requires full perception of humans and advocates the subjective initiative of humans in the manufacturing process. Vision perception methods centered on humans, such as human body shape and pose estimation, and action recognition, are the basis for machines to understand humans. The perception of human state can be multi-dimensional, including body posture, gestures, and facial expressions. Existing human state perception mainly has two methods: pose estimation and shape and pose estimation. Pose estimation only provides a coarse-grained visual description, while shape and pose estimation models humans as a parameterized triangular mesh, such as the SMPL model (Skinned Multi-Person Linear model), which represents the human body as two sets of parameters: shape and pose. However, existing vision perception methods are insufficient in the face of the requirements of human-machine collaboration with occlusion, high real-time performance, and accuracy. Specifically, it is impossible to correctly estimate the shape and pose of humans in the presence of occlusion, and there will also be a negative transfer phenomenon in the unoccluded area. Existing human action recognition methods mainly rely on skeletal features, which are independent of the environment, so they have strong transfer ability in different industrial scenarios. However, existing methods can only process video fragments or skeletal sequences containing only a single action. For the real-time and coherent action recognition required by the HRC task, it is not accurate enough, often misidentifies in the area of action transition, and often has label jitter in real-time recognition, and cannot achieve stable recognition. This poses a great challenge to correctly understanding human intentions and personnel safety.

[0003] Patent Application Number: CN202380031168.6, titled "A 2D Human Pose Estimation Method for Dual-Branch Human Structure Modeling", which can restore the key points of human poses even in severely occluded situations. This method constructs a convolutional neural network consisting of a backbone network, a dual-branch inference module, and a feature fusion decoding module. The dual-branch inference module is composed of a high-confidence key point inference branch and a component inference branch, which process key point and component information respectively. The feature fusion decoding module includes a multi-layer perceptron and a decoding stage, enabling the network to fuse key point and component information to generate the final predicted key point heatmap. This method trains the network using a training set and a validation set, and finally achieves high-precision 2D human pose estimation on the test set, performing excellently in dealing with occlusion and complex backgrounds, significantly improving the accuracy and robustness of 2D human pose estimation. This patent uses 17 key points, and compared with this patent, it reduces the lower limb key points, which is in line with the fact that the lower limbs of the human body do not participate in cooperation in the workbench scenario, effectively reducing recognition errors. Secondly, the inference branch of this patent relies on visual feature maps, resulting in greater interference for the model in complex backgrounds. Summary of the Invention

[0004] Aiming at the deficiencies of the above-mentioned existing technologies, the present invention proposes a human-centered visual perception method for human-machine collaboration scenarios. By means of a 2D key point regression branch, it learns the prior knowledge of human poses, and through an image data augmentation method, it reduces the error of the model in occlusion environments. Through an uncertainty-based action recognition method, it filters the data between actions and realizes real-time action recognition by means of an accumulative trigger, improving the accuracy and stability of the model in real-time recognition.

[0005] To achieve the above objectives, the technical solutions adopted by the present invention are as follows:

[0006] A human-centered visual perception method for human-machine collaboration scenarios of the present invention includes the following steps:

[0007] S1: Predict the human body shape and pose parameters in an occlusion scenario through a human mesh recovery algorithm enhanced by prior knowledge;

[0008] S2: Recombine the pose parameters predicted by the human mesh recovery algorithm enhanced by prior knowledge into a skeletal spatio-temporal topology map in the HRCA-11 format, and after data augmentation processing, input it into action recognition models with different weights for joint inference to obtain the action type and uncertainty of a single recognition;

[0009] S3: Send the uncertainty and action category into an accumulative trigger, and combine the historical recognition records to output a smooth and stable action category.

[0010] Further, step S1 specifically includes:

[0011] Use HRNet as the backbone network to extract multi-resolution features of the input image;

[0012] Divide the multi-resolution features into an SMPL parameter regression branch and a 2D keypoint regression branch through Neck operations;

[0013] Iteratively optimize the shape parameters, pose parameters, and camera parameters through the SMPL parameter regression branch, and combine an adversarial training strategy;

[0014] Generate a heatmap of the multi-resolution features through the 2D keypoint regression branch to enhance the pose prior knowledge in the occluded areas;

[0015] Adopt a lower limb cropping data augmentation strategy during the training phase and enable it after the network converges.

[0016] Further, step S2 specifically includes:

[0017] Construct the predicted pose parameters into a skeletal spatio-temporal topology graph;

[0018] Use a deep ensemble model to jointly predict the action category and uncertainty. The deep ensemble model is trained through random weight initialization, data rearrangement, and random data augmentation;

[0019] Output a smooth action category through an accumulative trigger mechanism that combines historical recognition records and uncertainty thresholds.

[0020] Further, the SMPL parameter regression branch adopts the error feedback iteration IEF method to optimize the parameters through the following steps:

[0021] Input the average shape, pose, and camera parameters of the SMPL model;

[0022] Iteratively predict the parameter differences through a multi-layer perceptron and accumulate and update them with the previous results;

[0023] Use 6D pose to represent the pose parameters and generate 2D skeletal point supervision signals through weak perspective camera projection.

[0024] Further, the adversarial training strategy includes:

[0025] 25 discriminators, where 23 discriminators respectively correspond to the relative rotation angles of 23 joints of the human body, 1 discriminates the authenticity of the overall pose, and 1 discriminates the shape parameters;

[0026] The generator and discriminator are alternately optimized. The loss function of the generator is the L2 loss between the discriminator output and the true label.

[0027] Further, the lower limb cropping data augmentation strategy is specifically:

[0028] During the training phase, the lower half of the image is cropped to retain the skeletal point supervision signal in the occluded area;

[0029] After the model converges, the cropping strategy is enabled, and data diversity is enhanced by combining random rotation and horizontal flipping.

[0030] Furthermore, the uncertainty calculation method of the deep ensemble model is as follows:

[0031] Apply random data augmentation multiple times to the same input data to generate multiple prediction results;

[0032] Calculate the action category probability through the expectation of the prediction results, and measure the uncertainty through the standard deviation or divergence.

[0033] Furthermore, the cumulative trigger mechanism specifically includes:

[0034] Set an accumulator for each action category, and accumulate the current category count value when the uncertainty is lower than the threshold;

[0035] When the accumulated value exceeds the lower threshold, trigger action determination, and stop accumulating when it exceeds the upper threshold;

[0036] The sliding window covers 20 frames of skeletal data and dynamically updates the historical recognition records.

[0037] Furthermore, the construction method of the skeletal spatio-temporal topology graph is as follows:

[0038] Process the skeletal sequence through the ST-GCN network, where the graph convolution module divides the neighboring nodes into central points, near-central points, and far-central points;

[0039] Adopt a spatio-temporal separated convolution strategy, using graph convolution in the spatial dimension and temporal convolution in the time dimension.

[0040] Furthermore, it also includes interactive action recognition based on environmental perception:

[0041] Determine the operation target at the moment when the distance between the object and the human body is the farthest;

[0042] Combine the template: Human[action] represents the human action template, a represents the action that the human is performing, and o represents the object that the human is operating on.

[0043] Advantages of the present invention:

[0044] The present invention is used in the field of human-machine collaboration, especially for human body shape and pose recognition technology in the presence of occlusion in human-machine collaboration. Specifically, it is a method for accurately obtaining the state and intention of humans in an occluded environment. Its advantages are as follows:

[0045] 1. Human Mesh Recovery Algorithm Enhanced by Prior Knowledge: This application proposes a human shape and pose estimation algorithm enhanced by prior knowledge. It learns the prior knowledge of human poses through a 2D key point regression branch, and through image data augmentation methods, enables the network to learn how to predict occluded humans from both data and knowledge aspects. This algorithm has the ability to predict human poses in real-time and can accurately predict in the face of occlusion, solving the problem that the human shape and pose estimation algorithm is not good at estimating occluded humans in a human-machine collaboration working environment. It greatly improves the real-time and accuracy of prediction, significantly enhances the robust prediction ability for local occlusion scenarios, and expands the application boundary of human-machine collaboration.

[0046] Action Recognition Algorithm Based on Uncertainty Estimation and Accumulation Trigger: Filters the data between actions through uncertainty and realizes real-time action recognition through the accumulation trigger method. Reduces the problem that the real-time action recognition in human-machine collaboration often predicts incorrect transition data between two actions. Brief Description of the Drawings

[0047] Figure 1 It is the logic structure diagram of the human mesh recovery algorithm enhanced by prior knowledge in the embodiment of this application;

[0048] Figure 2 It is the network structure diagram of the Neck operation on HRNet in the embodiment of this application;

[0049] Figure 3 It is the heat map of human key points in the embodiment of this application;

[0050] Figure 4 It is the effect diagram of cropping the human body for data augmentation in the embodiment of this application;

[0051] Figure 5 It is the network structure diagram of action recognition in the embodiment of this application;

[0052] Figure 6 It is the network structure diagram of ST-GCN in the embodiment of this application;

[0053] Figure 7 It is the effect diagram of the trigger and accumulator of actions in the embodiment of this application. Detailed Embodiment

[0054] For the convenience of those skilled in the art to understand, the present invention will be further described below in conjunction with the embodiments and the drawings. The content mentioned in the embodiments does not limit the present invention.

[0055] Refer to Figures 1 to 7 As shown, a human-centered visual perception method for a human-machine collaboration scenario of the present invention is as follows:

[0056] Figure 1 Shows the network structure of the prior knowledge enhanced human body mesh recovery algorithm. Since HRC usually occurs in single-person scenarios, the network adopts a top-down structure. The top-down method eliminates the need for complex post-processing of the bottom-up method and provides better real-time performance and accuracy. Specifically, the network is divided into three parts: the backbone network, the SMPL parameter regression module, and the 2D keypoint regression module. Among them, the SMPL parameter regression module is the main task of the network. The 2D keypoint regression module is an auxiliary task of this application. It enhances the feature extraction ability of the backbone network by learning the features of 2D keypoints, obtains a differentiated feature representation, and improves the network's recovery performance for occluded humans. Specifically, after the image passes through the backbone network, a multi-resolution visual representation is obtained. The feature representations of different resolutions will pass through two different branches. After the features of all resolutions are aggregated, they will be sent to the regression branch of the SMPL parameters, and the features of the highest resolution will be sent to the branch of the 2D keypoint regression. The SMPL parameters, 2D keypoints, etc. predicted by each branch are respectively compared with the ground truth in the dataset to calculate the loss and reverse-optimize the neural network. In addition, data augmentation is also performed during the training process. By cropping the lower limbs of the image and sending it into the neural network, the occluded situation is simulated to enhance the network's learning ability for occluded humans.

[0057] Backbone network:

[0058] In this application, HRNet is selected as the backbone network to extract the features of the image because HRNet can obtain multi-resolution visual representations and is naturally suitable for the task of human keypoint detection. At the same time, this application believes that although the network for predicting SMPL parameters has a different structure from the heatmap-based keypoint detection network, they belong to similar downstream tasks and should also have similar visual representations. First, after the image is cropped according to the detection box in the dataset, it is resized to 224×224, and the pixels in the blank area are filled with zeros. The processed image is sent into HRNet to obtain visual representations of multiple resolutions. Among them These visual representations are divided into two categories after passing through the Neck and are used as the inputs for the SMPL parameter regression module and the 2D keypoint regression module, denoted as and The specific Neck operation is as Figure 2 shown. The representations of different resolutions expand the channel length to 4 times through the Bottleneck module, that is Among them Starting from the feature with the highest resolution, through 2x downsampling, a feature with the same dimension as the feature of the second-highest resolution is output and added to the feature of the second-highest resolution to obtain a new feature. Repeat the above operations for the new features and the features with lower resolutions until the features with the lowest resolution. The unified feature that aggregates all resolutions is This feature undergoes operations of 1×1 convolution and average pooling to obtain the feature input R1 for the SMPL branch. The feature input R2 for the 2D keypoint branch is the same as the feature F1 with the highest resolution output by the original HRNet.

[0059] SMPL Parameter Regression Module:

[0060] The SMPL parameter regression module takes the feature R1 as input and regresses the SMPL shape parameters pose parameters and camera parameters Since it is difficult to directly regress the pose parameters, this application follows the regression method of error feedback iteration (IEF) proposed by HMR. The pseudocode is shown in Algorithm 1. The input of IEF is the average shape parameters provided by the SMPL model average pose parameters average camera parameters They are all serialized into 1D, and after concatenation, the predicted differences, Δβ, Δθ, ΔI, are iteratively predicted. These predicted differences are accumulated with the original input to obtain the input for the next IEF, that is, β1 = β0 + Δβ, θ1 = θ0 + Δθ, I1 = I0 + ΔI. After a certain number of iterations, the final SMPL parameters and camera parameters are obtained. For the pose parameters of SMPL, this application uses 6D pose, which is a continuous rotation representation and is more friendly to the prediction of neural networks. After that, the 6D pose will be converted into other representation forms supported by SMPL, such as Euler angles, quaternions, rotation matrices, etc.

[0061] The obtained shape and pose parameters are fed into the SMPL model to obtain a more fine-grained human representation, that is, the mesh in SMPL The process can be expressed as V = M(β, θ). In addition, because the bone points in different datasets are different, the mesh of SMPL can also be mapped to different bone points to utilize more data. Specifically, for the 3D bone representation form in a dataset Through a mapping matrix the mesh of SMPL can be mapped to the dimension that conforms to the bone representation form in the dataset, that is, to achieve multi-dataset supervision.

[0062] Furthermore, project the 3D bone points to 2D through the predicted camera parameters and supervise with more extensive 2D labels. Specifically, the weak perspective camera used in this application which respectively represents a scaling parameter This to some extent reflects the degree of proximity of people, and two translation parameters respectively include the normalized translation t along the x-axis x and the normalized translation t along the y-axis y . This process can be described as where Π(·) represents the projection operation of the weak perspective camera.

[0063] Furthermore, because the relevant dataset does not open the labels of SMPL parameters, direct supervision of SMPL parameters cannot be obtained. Only using the supervision method of 3D or 2D bone points will have an adverse impact on the results. The relevant dataset does not open the SMPL parameters corresponding to the images, but the SMPL parameters of the real human body without images are available. Therefore, using adversarial loss can alleviate the deficiency of the lack of real SMPL parameter supervision to a certain extent. Specifically, this application follows the setting of HMR, and the discriminator D = {D1, D2,.., D 25}}, a total of 25. 23 of them are used to judge the authenticity of the relative rotation angles of 23 joints of a person, 1 discriminator is used to judge the authenticity of the overall pose of a person, and the remaining 1 discriminator is used to judge the shape parameters of a person, as shown in Algorithm 2. The prediction SMPL backbone acts as the generator part in the generative adversarial network. The discriminator and the generator are updated sequentially. The discriminator is used to distinguish real SMPL parameters from the SMPL parameters predicted by the neural network, and the generator is used to make the SMPL parameters predicted by the neural network closer to the real values.

[0064] 2D keypoint regression module:

[0065] In addition to the main branch that regresses SMPL parameters, this application also designs a branch to handle the auxiliary task of 2D keypoint detection to implicitly learn pose-related prior knowledge and enhance the network's representation of occluded human bodies. It is a heatmap-based detection method that converts the estimation of the keypoint positions of the human body into the corresponding pixel coordinates in the heatmap. Different channels of the heatmap predict different keypoints, as Figure 3 shown. Specifically, the high-resolution representation obtained by the backbone network is sent to a CNN to directly output the heatmap of human keypoints where c = 54, representing the union of the keypoint categories included in the dataset, and the coordinates of the maximum point in each channel are the positions of the predicted keypoints.

[0066] Training strategy for data augmentation:

[0067] In addition to the auxiliary task of using 2D key point detection, this application also introduces data augmentation, aiming to enable the neural network to learn possible situations of occluded regions from the data aspect. Specifically, this application crops the lower half of the image and feeds it into the neural network. For the 3D bone points regressed by SMPL parameters and their projected 2D bone points, the cropped parts are still supervised to enable it to learn the situations of occluded regions, such as Figure 4 As shown, the red area is the detection box after cropping, and the green area is the original detection box. For the heatmap branch, the occluded key points are not supervised because the heatmap method cannot represent points outside the image. In addition to cropping the image, this application also uses random rotation and horizontal flipping of the image, which can further expand the data. It should be noted that cropping the image is not enabled at the beginning of training, but only after the network has basically converged and can infer reasonable human SMPL parameters. Because enabling it at the beginning of training will greatly increase the learning difficulty of the network and make it difficult to converge.

[0068] Design of the Loss function:

[0069] In view of the diversity of annotations in each dataset, this application proposes a set of weighted losses to make full use of the information contained in each dataset. The Loss function of this application has a total of 5 categories. 3D bone point loss As shown in formula (1), where is the ground truth of the 3D bone points in the dataset, is the predicted value of the neural network. This application uses the L2 loss to calculate their losses. For the bone point annotations that do not exist in the dataset, the losses of these bone points are not calculated. 2D bone point loss As shown in formula (2), its definition method is the same as that of formula (1). Loss of the heatmap for 2D key point regression As shown in formula (3), it is also defined by the L2 norm, is the ground truth of the heatmap, and p i,j,k is the heatmap predicted by the network. Their losses are calculated pixel by pixel and then summed as the heatmap loss. In addition, the loss of the generator As shown in formula (4), Φ = {β, θ}, which are the SMPL parameters predicted by the network. This application hopes that through the discriminator D i (·), it can be close to the real SMPL data, that is, close to 1. The loss of each discriminator As shown in formula (5), they are summed as the final discriminative loss. Θ comes from the true data distribution with its corresponding label being 1, and Φ comes from the fake data (i.e., the data generated by the network) with its label being 0. This application hopes that it can distinguish the true data and the fake data as much as possible, providing a better benchmark for the optimization of the generator. The discriminator and the generator are iteratively optimized in each round. Finally, this application combines the above losses into a weighted loss As shown in formula (6). λ1 to λ4 are hyperparameters used to balance each loss, and they are 100, 10, 300, and 1 respectively.

[0070]

[0071] The action recognition algorithm based on uncertainty estimation and accumulative trigger identifies the action category of the skeletal spatio-temporal topology map. This method inputs the skeleton regenerated from the human pose regenerated by the human mesh recovery algorithm enhanced by the above prior knowledge, and obtains the action type and uncertainty. The action type and uncertainty are input into the accumulative trigger, and combined with the historical recognition records, a smooth and stable action category is output. It reduces the problems of misrecognition at the action transition and label jitter in real-time recognition. Thus, it realizes robust human intention perception, providing a basis for subsequent robot task reasoning.

[0072] Implementation method:

[0073] Figure 5 Show the structural diagram of the action recognition algorithm based on uncertainty estimation and accumulative trigger. First, the human pose predicted by the human shape and pose estimation algorithm is recombined into a skeletal spatio-temporal topology map in the form of HRCA-11. Subsequently, the topology map is sent into action recognition models with different weights after data augmentation for joint prediction, obtaining the action type and uncertainty of a single recognition. Finally, the real-time uncertainty and action category are sent into the accumulative trigger, and combined with the historical recognition records, a smooth and stable action category is output.

[0074] Skeleton-based action recognition module:

[0075] This application uses a graph convolutional neural network to process the skeletal spatio-temporal graph to achieve action recognition. ST-GCN is the benchmark method of this application. Each layer of graph convolution can be abstracted as consisting of a sampling function and a weight function. Similar to CNN, the sampling function of GCN is used to sample the elements within the neighborhood of an element, and the weight function is used to assign weights to each sampled element.

[0076] For a root node v located at the sampling center ti , similar to the center point of the convolutional kernel in CNN, its neighborhood can be expressed as B(v ti ) = {v tj|d(v tj ,v ti )≤D}, d(v tj ,v ti ) represents the shortest path length between two nodes in the spatio-temporal topological graph of the skeleton defined above, where D = 1. For a root node, only the nodes with a length of 1 around it are recorded as its neighborhood. Let the sampling function of the GCN be p(v ti ,v tj ), where v tj ∈ B(v ti ), and p(v ti ,v tj ) only samples the nodes in its 1-hop neighborhood around it, as shown in Equation (8).

[0077] p(v ti ,v tj ) = v tj (8)

[0078] Furthermore, similar to CNN, the GCN assigns weights to the nodes in the 1-hop neighborhood B(v ti ) of the root node v ti . The difference is that 2D convolution processes a sequence of pixels with a fixed spatial order, while the structure of the graph has no specific spatial order. Therefore, it is necessary to divide B(v ti ) into different subsets and assign different weights to different subsets. Specifically, through a mapping function l ti (·), the nodes in the spatial neighborhood B(v ti ) of the root node v ti ) are mapped into K subsets, as shown in Equation (9). For each node v ti , there is a weight w(l ti (v tj )) corresponding one-to-one to the subset it belongs to, which maps the nodes in the subset into features, as shown in Equation (10).

[0079] l ti ∶ B(v ti ) → {1, 2, …, K - 1} (9)

[0080]

[0081] This application adopts a similar method to that in the ST-GCN paper and uses the mapping function l ti (·) to divide the neighborhood B(v ti ). The ST-GCN method sets the center of gravity position g t of a person, and g t is artificially set to the position of the neck. By comparing the nodes v tj in the neighborhood ∈ B(vti ) The shortest distance r to the centroid g j , and the root node v ti Divide the nodes according to the distance from the centroid g to the shortest distance, r j = d(v tj , g t ). The nodes in the neighborhood are mapped into 3 subsets, as shown in formula (11), namely the center point r j = r i , the near-center point r j < r i , and the far-center point r j > r i .

[0082]

[0083] For the spatio-temporal graph convolution process without subsets and with all nodes sharing weights, it can be expressed as formula (12). Where f k (·) represents mapping the nodes to the output features of the k-th layer, |B(·)| represents the cardinality of the neighborhood of the nodes, w 0 represents the weights of the neural network, represents the convolution along the time dimension. Further, it is extended to a graph convolution with subsets and different weights assigned to different subsets, and its process is as shown in formula (13). Where, Z ti (v tj ) represents the cardinality of the subset where different nodes are located. Further, formula (13) is extended from a single node v ti to the set of nodes V in a single frame, so formula (14) is derived. Where, is the set of features output by the k-th layer, A j is the adjacency matrix A divided according to subsets, A + I = ∑ j A j . D j is the degree of the divided adjacency matrix A j , It should be noted that the graph convolution here only acts on the spatial dimension, that is, it convolves the spatial topology graph V t of each frame separately, but this is a parallel calculation. The convolution along the time dimension is completed by , and they are carried out iteratively in sequence.

[0084]

[0085] Specifically, the above formula (14) is a basic module of ST-GCN, and the ST-GCN network is stacked by 9 such modules. The original skeletal time series data is first fed into the network to extract features After that, the feature F1 are fed into 9 ST-GCN modules for operation to obtain feature F 10 will perform average pooling along the temporal and spatial dimensions The pooled features will be fed into the head layer to project scores for each category cls is the number of action categories. Finally, through SoftMax, the probability y of each category is obtained, and the category with the highest probability is the action category i, that is, i = argmax i (y (i) ), as Figure 6 shown. Among them, GCN performs convolution on the spatial graph, and TCN performs convolution along the temporal dimension. They iteratively operate on the input features in turn.

[0086] Deep ensemble uncertainty estimation module:

[0087] The above method can directly classify actions. However, this neural network for single action recognition often makes incorrect classifications for out-of-domain (outside the training set) data because the neural network can only output the category probabilities within the training set, and the probabilities it predicts for actions that do not belong to the existing categories are meaningless. Therefore, this application aims to identify such situations by estimating uncertainty, enabling the neural network to output its uncertainty about this prediction while predicting the action category, as a measure of whether to accept this prediction. When the uncertainty is higher than a certain threshold, this prediction is considered untrustworthy; otherwise, this prediction result is adopted.

[0088] The neurons in each layer of the neural network are fixed values, resulting in the fact that a single forward process for the same data can only output fixed values. A natural idea to make the neural network output uncertainty is to hope that the output result of the neural network is not a value but a distribution, so that the uncertainty of the network can be measured according to the degree of dispersion of the distribution. Bayesian neural network is a classic method. It makes each neuron follow a distribution rather than a fixed value, and multiple random samples are taken from the distribution during forward inference, so that the predicted results also form a distribution, and the standard deviation of this distribution can be used as the uncertainty of the network. However, the number of parameters in the neural network is very large, and it is almost unrealistic to strictly calculate the posterior distribution of the output according to the Bayesian formula.

[0089] This application uses the deep ensemble method to estimate the prediction uncertainty of a neural network. The literature has shown that deep ensemble neural networks have an effect similar to Bayesian networks, thanks to the randomness in the training process. The deep ensemble method trains multiple models with a large degree of randomness and also uses the weights of multiple models for inference in the inference stage. The expected value of the distribution is considered the prediction result of the neural network, and the standard deviation of the distribution is the uncertainty of the neural network. Specifically, this application trains M groups of models with a large degree of randomness and introduces three types of randomness: (1) randomly initialized weights of the neural network; (2) randomly shuffled data input; (3) random data augmentation. For (3) data augmentation, this application randomly determines the starting frame of the skeletal sequence, divides and reassembles the skeletal sequence based on this, and repeats it to a specific length. This setting remains unchanged in the inference stage to increase randomness. The pseudo-codes for training and inference are shown in Algorithms 3 and 4. In the inference stage, the weights {θ1,…,θ M} of the trained ensemble model are loaded. For the data x fed into the neural network, data augmentation is applied N times on the neural network with weights θ m to obtain x m = {x m1 ,…,x mN}. These data are fed into the neural network to obtain the prediction results y m = {y m1 ,…,y mN}. This process is repeated M times on the ensemble model with weights . The set of prediction results obtained can be considered a uniform mixture model, and its prediction result p E (y|x) and uncertainty u E (y|x) are as shown in Formulas (15) and (16), where KL(·||·) represents the KL divergence between two distributions.

[0090]

[0091] Real-time action recognition based on an accumulative trigger:

[0092] In the real-time action recognition stage, the camera captures the video stream. The object detection and human mesh shape and pose estimation algorithms predict the skeletal sequences of the human body, and these skeletal sequences are stored. A sliding window covers the skeletal data within a region ending with the current frame. These data are fed into the uncertainty-based behavior recognition model to obtain the predicted class i = argmax i (p E (y|x) (i) ) and the uncertainty u E(y|x). Although the uncertainty can well identify data outside the distribution, since the method outputs a predicted category for each frame, adjacent frames will be predicted as different categories, showing a jitter phenomenon. This application adopts a discriminant method based on a trigger, which is used to smooth the real-time classification results of the neural network and avoid accidental errors caused by category jitter. Specifically, each action category i has an accumulator A(i), which records the action category of the current frame in real time. When the uncertainty u E (y|x) is lower than the threshold Θ E , the accumulator A(i) of this action category is incremented by 1, and the accumulators A(j) of other action categories are decremented by 1, where j ≠ i. When the value of the action accumulator exceeds the lower threshold Θ L , the trigger TR(i) of this action category is triggered, and TR(i) = 1, which is the final action recognition result. When the value of the action accumulator is lower than the lower threshold Θ L , the trigger of this action category, TR(i) = 0, indicating that this action ends. At the same time, this application also sets an upper threshold Θ U , which ensures sensitivity and avoids interference from long actions to subsequent recognition. When the value of the accumulator is higher than Θ U , no further accumulation is performed, and A(i) = min(A(i) + 1, Θ U ). In this application, the sliding window is 20 frames, and the hyperparameters are determined by experience and statistical results. Θ E = 0.1, Θ L = 8, Θ U = 12. For the specific process, see Figure 7 .

[0093] Interactive action recognition based on environmental perception:

[0094] The above method can achieve stable and accurate action recognition, but it only considers human actions and lacks perception of the surrounding environment. For actions such as picking up an object, it is more appropriate to describe them in the form of human-object interaction. Specifically, through a manually set template, "Human[action]:I{a}the{o *}.", which is used to describe fine-grained action intentions. Among them, o ∈ O represents the object operated by the human, and a ∈ A represents the action being performed by the human. During the time period T when the action a occurs, the object o is determined in the following way * , where M(·) maps the object to pixel coordinates, and t max = argmax t∈T ‖M(hand t ) - M(root t) ‖, indicating the moment when the distance between the hand and the body is the farthest. At this moment, the object closest to the human hand is the object operated by humans.

[0095] The specific application ways of the present invention are numerous. The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art of this technology, several improvements can be made without departing from the principle of the present invention, and these improvements should also be regarded as the protection scope of the present invention.

Claims

1. A human-centered visual perception method for human-robot collaboration scenarios, characterized in that, It includes the following steps: S1: Use a human mesh recovery algorithm enhanced by prior knowledge to predict the human shape and pose parameters in an occluded scenario; S2: Reconstruct the pose parameters predicted by the prior knowledge enhanced human mesh recovery algorithm into a skeletal spatio-temporal topology map in the HRCA-11 format, and after data augmentation processing, input it into action recognition models with different weights for joint inference to obtain the action types and uncertainties recognized in a single instance; S3: Send the uncertainty and action category into an accumulative trigger, and combine with historical recognition records to output a smooth and stable action category.

2. The human-centered visual perception method for a human-robot collaboration scenario according to claim 1, wherein Step S1 specifically includes: Use HRNet as the backbone network to extract multi-resolution features of the input image; Divide the multi-resolution features into an SMPL parameter regression branch and a 2D keypoint regression branch through Neck operations; Iteratively optimize the shape parameters, pose parameters, and camera parameters through the SMPL parameter regression branch, and combine with an adversarial training strategy; Generate a heat map of the multi-resolution features through the 2D keypoint regression branch to enhance the pose prior knowledge in the occluded area; Adopt a lower limb cropping data augmentation strategy in the training stage and enable it after the network converges.

3. A human-centered visual perception method for a human-robot collaboration scenario according to claim 1, characterized in that, Step S2 specifically includes: Construct the predicted pose parameters into a skeletal spatio-temporal topology map; Use a deep ensemble model to jointly predict the action category and uncertainty. The deep ensemble model is trained by randomly initializing weights, data rearrangement, and random data augmentation; Output a smooth action category through an accumulative trigger mechanism in combination with historical recognition records and an uncertainty threshold.

4. The human - centered visual perception method for a human - machine collaboration scenario according to claim 2, wherein, The SMPL parameter regression branch adopts an error feedback iteration IEF method to optimize the parameters through the following steps: Input the average shape, pose, and camera parameters of the SMPL model; Iteratively predict the parameter difference through a multi-layer perceptron and accumulate and update it with the previous result; Use 6D pose to represent the pose parameters and generate a 2D bone point supervision signal through weak perspective camera projection.

5. The human-centered visual perception method for a human-robot collaboration scenario according to claim 2, wherein The adversarial training strategy includes: 25 discriminators, among which 23 discriminators respectively correspond to the relative rotation angles of 23 joints of the human body, 1 discriminates the authenticity of the overall pose, and 1 discriminates the shape parameters; The generator and discriminator are alternately optimized, and the loss function of the generator is the L2 loss between the discriminator output and the true label.

6. The human-centered vision perception method for a human-robot collaboration scenario according to claim 2, wherein The lower limb cropping data augmentation strategy specifically is: In the training stage, crop the lower half of the image and retain the bone point supervision signal in the occluded area; Enable the cropping strategy after the model converges and combine with random rotation and horizontal flipping to enhance data diversity.

7. The human - centered visual perception method for a human - machine collaboration scenario according to claim 3, characterized in that, The uncertainty calculation method of the deep ensemble model is: Apply random data augmentation to the same input data multiple times to generate multiple prediction results; Calculate the action category probability through the expectation of the prediction results, and measure the uncertainty through the standard deviation or divergence.

8. The human-centered visual perception method for a human-machine collaboration scenario according to claim 3, wherein The accumulative trigger mechanism specifically includes: Set an accumulator for each action category, and accumulate the current category count value when the uncertainty is lower than the threshold; Trigger action determination when the accumulated value exceeds the lower threshold and stop accumulating when it exceeds the upper threshold; The sliding window covers 20 frames of bone data and dynamically updates the historical recognition records.

9. The human - centered visual perception method for a human - robot collaboration scenario according to claim 3, wherein, The construction method of the skeletal spatio-temporal topology map is: Process the skeleton sequence through the ST-GCN network, where the graph convolution module divides the neighborhood nodes into central points, near-center points, and far-center points; Adopt a spatio-temporal separated convolution strategy, using graph convolution in the spatial dimension and temporal convolution in the temporal dimension.

10. The human-centered visual perception method for a human-robot collaboration scenario according to claim 2, characterized in that It also includes interactive action recognition based on environmental perception: Determine the operation target at the moment when the distance between the object and the human body is the farthest; Combined with the template: Human[action] represents the human action template, a represents the action that the human is performing, and o represents the object that the human is operating on.

Citation Information

Patent Citations

  • Natural language control of robot

    CN118900751A