A Robot Redirection Method Based on Graph Neural Networks
By combining graph neural network encoder and potential neural network, an efficient and accurate robot redirection method is realized, solving the problems of complex computing and poor interpretation in the existing technology, and improving mapping accuracy and real-time performance.
Patent Information
- Application Number
- CN202411084199.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-08
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-08-08
AI Technical Summary
The existing robot redirection method has complex calculations, is prone to local optimality, poor interpretability and controllability, and is poor in adaptability to new postures.
Using a graph neural network-based method, the graph neural network encoder and potential neural networks capture the difference in human-machine motion, and the graph decoder and differentiable positive kinematics layer are used to achieve end-to-end vision-motion mapping.
The mapping process is simplified, the mapping accuracy and generalization capabilities are improved, the computing complexity is reduced, and efficient real-time performance and accurate motion redirection is achieved.
Smart Images

Figure CN119068148B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of robot computing, and in particular to a robot redirection method based on graph neural network. Background Art
[0002] With the rapid development of robot technology, vision-guided robot teleoperation technology has received extensive attention. This technology captures the operator's actions through various devices and then maps them to the robot platform in real time to achieve remote control, which has important applications in fields such as special environment operations and surgical assistance. The key to vision teleoperation is to achieve high-precision motion redirection mapping.
[0003] Traditional methods are mainly based on optimization algorithms such as inverse kinematics and motion planning, which require manual design of objective functions and constraint conditions, with complex calculations and prone to falling into local optima. In recent years, data-driven methods based on deep learning have emerged, attempting to directly learn the mapping model from visual data end-to-end, overcoming the defects of traditional methods. However, due to being black-box operations, they have poor interpretability and controllability, are difficult to debug and optimize, and also have poor adaptability to new poses not covered.
[0004] Secondly, existing technologies such as methods based on reinforcement learning model motion redirection as a Markov decision process and solve the mapping strategy from state to action through deep reinforcement algorithms. The advantage of such methods is strong generalization, but the disadvantages are that they require building an accurate environmental model, with high computational cost and lack of interpretability, which are restricted in practical applications. At the same time, they also have the problem of insufficient interpretability. Summary of the Invention
[0005] Aiming at the deficiencies of the existing technology, the present invention provides a robot redirection method based on graph neural network.
[0006] The technical solution of the present invention is as follows: A robot redirection method based on graph neural network, comprising the following steps:
[0007] S1), collect the action video of the operator; and extract the human skeletal structure from the action video;
[0008] S2), model the human skeletal structure and the robot structure in the form of a tree-shaped topological graph according to the graph;
[0009] S3), use the corresponding graph neural network encoder to encode and optimize the tree-shaped topological graphs of the human body and the robot respectively to obtain the latent space representations of the tree-shaped topological graphs of the human body and the robot;
[0010] S4), Input the optimized latent space representation of the human-machine skeleton diagram obtained in step S3) into the latent neural network, and capture the latent features of the human-machine motion differences through the latent neural network, so as to obtain the joint angles and link poses of the robot;
[0011] S5), Through the differentiable forward kinematics layer, convert the joint angles and link poses output by the graph decoder into the desired position of the robot end effector; that is, complete the final motion redirection mapping.
[0012] Preferably, in step S1), after obtaining the action video, perform denoising, color correction, and frame rate adjustment processing on it.
[0013] Preferably, in step S1), the visual detection algorithm OpenPose and the estimation algorithm Frankmocap extract the human body three-dimensional joint point coordinate data and attitude quaternion information from the action video as the human body bone structure.
[0014] Preferably, in step S2), use the joints of the robot as nodes and the links connecting two joints as undirected edges; therefore, model the structure of the robot as an acyclic tree topology graph Gr=(Vr, Er); where, Vr is the set of robot joint points; Er is the set of robot link edges.
[0015] Preferably, in step S2), the human body bone structure is modeled as an acyclic tree topology graph Gh=(Vh, Eh); where, Vh is the set of human body joint points; Er is the set of human body bone connection edges.
[0016] Preferably, in step S3), the graph neural network encoder adopts a hierarchical convolution structure, and the graph neural network encoder includes multiple stages, each stage includes multiple convolution layers of the same type, and each stage corresponds to processing features of different scales.
[0017] Preferably, in step S3), the graph neural network encoders are respectively a hand encoder and a body encoder; respectively use the hand encoder and the body encoder to encode and optimize the tree topology graphs of the human body and the robot, and capture the subtle changes in the hand movements of the human body and the robot through the convolution operation of the hand encoder; the body encoder uses the backbone path encoding operation to learn the macroscopic features of the overall body movements from the full-body skeleton graphs of the human body and the robot.
[0018] Preferably, in step S3), using the graph neural network encoder to encode and optimize the tree topology graphs of the human body and the robot specifically includes the following steps:
[0019] S31) For the human body tree topology, at each node, the input feature is the rotation angle of the corresponding joint, and at each edge, the input features are the initial offset and rotation between adjacent edges;
[0020] For the robot's tree topology, at each node, the output feature will represent the rotation angle of the corresponding joint, and at each edge, the output features will represent the initial offset and rotation between links;
[0021] S32) Use a graph neural network encoder to encode the tree topologies of the human body and the robot respectively to generate a latent space representation;
[0022] S33) Perform message passing and update on node features and edge features in the graph neural network encoder, and repeatedly perform multi-level graph convolution operations;
[0023] S34) Through the multi-level feature extraction and update of the graph neural network encoder, optimize the initial input features; and output a concise and efficient latent space representation.
[0024] Preferably, in step S4), the latent neural network includes a graph encoder and a graph decoder; wherein, the graph encoder performs node and edge update operations on the optimized encoding of the input human-robot skeleton graph through a graph convolutional network GCN. The graph convolutional network GCN is stacked by multiple layers of graph convolutional layers, and each layer of graph convolutional layer includes two operations: node update and edge update. The node update fuses the current node feature and the aggregated feature from adjacent nodes, and the edge update fuses the current edge feature and the aggregated feature from adjacent edges; high-level node features and edge features are extracted through the graph convolutional network GCN to capture higher-level motion patterns and differences layer by layer;
[0025] The graph decoder decodes the latent features through a fully connected layer or a transposed convolutional layer to output the robot joint angle and link pose information.
[0026] Preferably, in step S4), using the latent neural network to extract the robot's joint angle and link pose information specifically includes the following steps:
[0027] S41) Map the optimized encodings of the human body and robot skeleton graphs into the latent neural network respectively, and then perform continuous constraint optimization through the graph encoder in the latent space to obtain latent features that can minimize the representation of the human-robot difference;
[0028] S42) Input the optimized and aligned latent encoding into the graph decoder to decode the corresponding joint angle and link pose information, and ensure that the decoded output meets the kinematic constraints of the robot through the prior structure information of the robot during the decoding process.
[0029] Preferably, in step S4), the expression for the graph encoder to update the fused node features and edge features is:
[0030] x_v t+1 =Φ(x_v t ,∑_{u∈N(v)}ψ(x_v t ,x_u t ,e_attr(v,u)));
[0031] In the formula, x_v t+1 represents the feature vector of node v updated at time t + 1; x_v t represents the feature vector of node v at time t; N(v) is the set of neighbor nodes of node v, u represents a neighbor node of v, and x_u t represents the feature vector of node u at time t; e_attr(v,u) is the feature vector of the edge connecting v and u, ψ is the message function for calculating the information vector based on node and edge information; ∑_{u∈N(v)} is used to aggregate the influences of all neighbor nodes; Φ is the update function for integrating its own information and the influences of neighbor nodes to update the node features.
[0032] Preferably, in step S4), the objective function for training the graph decoder is:
[0033]
[0034] In the formula; z is the latent encoding; are the parameters of the graph decoder; represents the reconstruction loss function, which is used to measure the difference between the output of the graph decoder and the input X; Ω(z) is the regularization term imposed on the latent encoding z; λ is the weight coefficient of the regularization term.
[0035] Preferably, in step S5), the differentiable forward kinematics layer is used to convert the joint angles and link poses information output by the graph decoder into the desired position of the robot end effector, which specifically includes the following steps:
[0036] S51), Take the robot joint angles and link poses information output by the graph encoder as the input and pass it to the differentiable forward kinematics layer;
[0037] S52), The forward kinematics layer calculates the transformation matrix of each joint according to the kinematic model and DOF settings of the robot;
[0038] S53), Calculate the overall motion transformation from the base coordinate system to the end effector through the transformation matrices of each joint to obtain the desired position of the end effector;
[0039] S54), Compare the desired position of the end effector with the true target position to obtain a position error;
[0040] S55), Based on the position error, backpropagate to the graph decoder and graph encoder to optimize the latent representation of the graph encoder and the output of the graph decoder;
[0041] S56), Through multiple iterations of optimization, minimize the position error of the end effector to obtain an accurate motion redirection mapping; thereby obtaining a sequence of robot joint angles that can accurately reach the desired position and completing the action redirection mapping task.
[0042] Preferably, in step S5), the goal of the motion redirection is to determine a set of robot joint angles r* such that the motion posture R(r*) of the robot can maximally imitate and fit a given human posture H under given constraint conditions; at the same time, it cannot exceed the range of the robot's own joint angles. By optimizing the objective function and satisfying the constraints, obtain the motion trajectory that allows the robot to best imitate human actions, that is:
[0043] r * = argminL(H, R(r));
[0044] s.t. r lower ≤ r ≤ r upper ;
[0045] In the formula; r* is the robot joint angle; L represents the target mapping function from the joint angle to the final posture of the robot; H is the human posture information; R(r) represents the motion of the robot, determined by the joint angle r; r lower 、r upper are the lower and upper limits of the robot motion respectively, reflecting the constraints of the robot motion range.
[0046] Preferably, in step S5), the joint angles and link poses of the robot output by the graph decoder are mapped to the motion trajectory of the robot end effector that can best imitate human actions by optimizing the target mapping function L; the expression of the target mapping function L of the motion redirection is:
[0047] L = λ ees L ees + λ ee L ee + λ ori L ori + λ fin L fin ;
[0048] In the formula, L ees is the velocity loss function; L ee is the end effector loss function; Lori is the velocity loss function; L fin is the finger loss function; λ ees 、λ ee 、λ ori 、λ fin are the weights of the respective loss functions.
[0049] Preferably, in step S5), the velocity loss function L ees has the following expression:
[0050]
[0051] where s i and r i are the velocity and normalization coefficient of the i-th frame of the robot; and are the velocity and normalization coefficient of the i-th frame of the human body.
[0052] Preferably, in step S5), the direction loss function L ori has the following expression:
[0053]
[0054] where; D i is the rotation matrix of the end effector of the i-th frame of the robot; is the rotation matrix of the end effector of the i-th frame of the human body.
[0055] Preferably, in step S5), the finger loss function L fin has the following expression:
[0056]
[0057] where; f i tip is the fingertip position of the i-th frame of the robot; f i meta is the finger tip position of the i-th frame of the robot; r tm is the normalization coefficient of the i-th robot arm; is the fingertip position of the i-th frame of the human body; is the finger tip position of the i-th frame of the human body; is the normalization coefficient of the i-th human arm.
[0058] Preferably, in step S5), the end effector loss function L ee has the following expression:
[0059]
[0060] where: ei The sum is r i The position of the end effector and the normalization coefficient of the i-th frame of the robot, and are the end effector and the normalization coefficient of the i-th frame of the human.
[0061] The beneficial effects of the present invention are as follows:
[0062] 1. The present invention combines a graph neural network encoder with a latent neural network, thereby retaining the robot structure prior while making full use of the powerful representation ability of the latent space; automatically learning the structural mapping relationship between the human body and the robot by using the graph neural network encoder; efficiently capturing the topological features in the graph-structured data through the graph neural network encoder; adaptively extracting the association patterns between humans and robots; greatly simplifying the mapping process, reducing the labor cost, and improving the mapping accuracy and generalization ability at the same time;
[0063] 2. Through the latent coding optimization strategy, the present invention maps the original joint space to a compact latent space, transforming the complex motion mapping problem into an optimization problem in the latent space; searching for the optimal human-robot motion mapping in the latent space through the gradient descent algorithm, greatly reducing the computational complexity of the mapping process and improving the real-time performance;
[0064] 3. The present invention captures human joint information through an RGB camera, performs graph modeling on the human skeleton and the robot structure; learns the structural mapping relationship between humans and robots through the graph encoder; avoiding complex manual design and parameter debugging;
[0065] 4. Through the composite optimization objective function including velocity loss, acceleration loss, end effector loss, direction loss, and finger loss, the present invention comprehensively considers the spatio-temporal characteristics of human-robot motion, making the mapping result more accurate and natural;
[0066] 5. The present invention uses a graph decoder to recover the robot joint angle information from the optimized latent coding, and converts it into the robot end position through a differentiable forward kinematics layer, realizing end-to-end vision-motion mapping. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] Figure 1 is a schematic flow chart of the method of the present invention;
[0068] Figure 2 is a framework diagram of the method of the present invention;
[0069] Figure 3 is a schematic framework diagram of the graph encoder of the present invention;
[0070] Figure 4 is a schematic framework diagram of the graph decoder of the present invention. Detailed implementation manners
[0071] The following further describes the detailed implementation manners of the present invention with reference to the accompanying drawings:
[0072] As Figure 1 and 2 shown, this embodiment provides a robot redirection method based on a graph neural network, including the following steps:
[0073] S1), Use an RGB camera to collect the action video of the operator; and use a visual detection algorithm and an estimation algorithm to extract the three-dimensional joint point coordinates and attitude quaternion information of the human body from the collected operator image or video as the human body bone structure.
[0074] In this embodiment, the visual detection algorithm and the estimation algorithm respectively adopt the OpenPose and Frankmocap algorithms.
[0075] And after obtaining the action video of the operator in this embodiment, it is necessary to perform denoising, color correction, and frame rate adjustment processing on it to improve the accuracy of subsequent pose extraction.
[0076] S2), Model the human body bone structure and the robot structure in the form of a graph as a tree topology graph.
[0077] In this embodiment, for constructing the robot structure into a tree topology graph, specifically: use the joints of the robot as nodes, and the connecting rod connecting two joints as an undirected edge; therefore, model the structure of the robot as an acyclic tree topology graph Gr=(Vr,Er); where, Vr is the set of robot joint points; Er is the set of robot connecting rod edges.
[0078] Correspondingly, model the human body bone structure as an acyclic tree topology graph Gh=(Vh,Eh); where, Vh is the set of human body joint points; Er is the set of human body bone connection edges.
[0079] S3), Use the corresponding graph neural network encoder to respectively encode and optimize the tree topology graphs of the human body and the robot to obtain the latent space representations of the tree topology graphs of the human body and the robot.
[0080] In this embodiment, the graph neural network encoder adopts a hierarchical convolution structure, and the graph neural network encoder includes multiple stages, each stage includes multiple convolution layers of the same type, and each stage corresponds to processing features of different scales.
[0081] Moreover, the graph neural network encoders are respectively a hand encoder and a body encoder; the hand encoder and the body encoder are respectively used to encode and optimize the tree-shaped topological graphs of the human body and the robot. The convolutional operation of the hand encoder is used to capture the subtle changes in the hand movements of the human body and the robot; the body encoder uses the backbone path encoding operation to learn the macroscopic features of the overall body movements from the full-body skeleton graphs of the human body and the robot.
[0082] Encoding and optimizing the tree-shaped topological graphs of the human body and the robot by using the graph neural network encoders specifically includes the following steps:
[0083] S31): For the tree-shaped topological graph of the human body, at each node, the input feature is the rotation angle of the corresponding joint, and at each edge, the input feature is the initial offset and rotation between adjacent edges;
[0084] For the tree-shaped topological graph of the robot, at each node, the output feature will represent the rotation angle of the corresponding joint, and at each edge, the output feature will represent the initial offset and rotation between the connecting rods;
[0085] S32): Use the graph neural network encoders to respectively encode the tree-shaped topological graphs of the human body and the robot to generate latent space representations;
[0086] S33): Perform message passing and updating on the node features and edge features in the graph neural network encoder, and repeatedly perform multi-level graph convolution operations;
[0087] S34): Optimize the initial features of the input through the multi-level feature extraction and updating of the graph neural network encoder; and output a concise and efficient latent space representation.
[0088] S4): Input the optimized latent space representations of the human and robot skeleton graphs obtained in step S3) into the latent neural network, and capture the latent features of the human-robot motion differences through the latent neural network, so as to obtain the joint angles and link poses of the robot.
[0089] In this embodiment, preferably, in step S4), the latent neural network includes a graph encoder and a graph decoder; wherein, the graph encoder performs node and edge update operations on the optimized encoding of the input human-robot skeleton graph through the graph convolutional network GCN. The graph convolutional network GCN is stacked by multiple layers of graph convolutional layers, and each layer of graph convolutional layer includes two operations: node update and edge update. The node update fuses the current node feature and the aggregated features from adjacent nodes, and the edge update fuses the current edge feature and the aggregated features from adjacent edges; high-level node features and edge features are extracted through the graph convolutional network GCN, and higher-level motion patterns and differences are captured layer by layer;
[0090] The described graph decoder decodes the latent features through a fully connected layer or a transposed convolutional layer, and outputs the robot joint angles and link poses. See Figure 3 and 4 shown below.
[0091] S5), through a differentiable forward kinematics layer, convert the joint angles and link poses output by the graph decoder into the desired position of the robot end effector; that is, complete the final motion redirection mapping.
[0092] Preferably in this embodiment, in step S4), the latent neural network is used to extract the robot joint angles and link poses, which specifically includes the following steps:
[0093] S41), respectively map the optimized encodings of the human body and robot skeleton graphs into the latent neural network. The graph encoder of the latent neural network uses the graph convolutional network GCN to perform constrained encoding on the optimized encodings of the input human-robot skeleton graphs; alternately update the human and robot node / edge features through the graph convolutional network GCN, and capture higher-level motion patterns and differences layer by layer; thus obtaining latent features that can minimize the representation of human-robot differences;
[0094] S42), input the optimized and aligned latent encoding into the graph decoder to decode the corresponding joint angles and link poses. During the decoding process, use the prior structure information of the robot to ensure that the decoding output meets the kinematic constraints of the robot.
[0095] Preferably in this embodiment, in step S41), the expression for the graph encoder to update and fuse node features and edge features is:
[0096] x_v t+1 =Φ(x_v t ,∑_{u∈N(v)}ψ(x_v t ,x_u t ,e_attr(v,u)));
[0097] In the formula, x_v t+1 represents the updated feature vector of node v at time t + 1; x_v t represents the feature vector of node v at time t; N(v) is the set of neighbor nodes of node v, u represents a neighbor node of v, and x_u t represents the feature vector of node u at time t; e_attr(v,u) is the feature vector of the edge connecting v and u, ψ is the message function for calculating the information vector based on node and edge information; ∑_{u∈N(v)} is used to summarize the influences of all neighbor nodes; Φ is the update function for integrating its own information and the influences of neighbor nodes to update node features.
[0098] Preferably, in step S42), the objective function for training the graph decoder is as follows:
[0099]
[0100] In the formula, z is the latent encoding; are the graph decoder parameters; represents the reconstruction loss function, which is used to measure the difference between the output of the graph decoder and the input X; Ω(z) is a regularization term imposed on the latent encoding z; λ is the weight coefficient of the regularization term.
[0101] Preferably, in step S5), the differentiable forward kinematics layer is used to convert the joint angles and link poses output by the graph decoder into the desired position of the robot end effector, which specifically includes the following steps:
[0102] S51), Transfer the robot joint angles and link poses output by the graph encoder as inputs to the differentiable forward kinematics layer;
[0103] S52), The forward kinematics layer calculates the transformation matrix of each joint according to the kinematic model and DOF settings of the robot;
[0104] S53), Calculate the overall motion transformation from the base coordinate system to the end effector through the transformation matrices of each joint to obtain the desired position of the end effector;
[0105] S54), Compare the desired position of the end effector with the real target position to obtain the position error;
[0106] S55), Based on the position error, backpropagate to the graph decoder and the graph encoder to optimize the latent representation of the graph encoder and the output of the graph decoder;
[0107] S56), Through multiple iterative optimizations, minimize the position error of the end effector to obtain an accurate motion redirection mapping; thereby obtaining a sequence of robot joint angles that can accurately reach the desired position and completing the action redirection mapping task.
[0108] Preferably, in step S5), the goal of the motion redirection is to determine a set of robot joint angles r* such that the motion pose R(r*) of the robot can best imitate and fit the given human pose H under the given constraints; at the same time, it cannot exceed the range of the robot's own joint angle limits. By optimizing the objective function and satisfying the constraints, obtain the motion trajectory that allows the robot to best imitate human actions, that is:
[0109] r * = argminL(H, R(r));
[0110] s.t.r lower ≤r≤r upper ;
[0111] Where; r* is the robot joint angle; L represents the target mapping function from the joint angle to the final pose of the robot; H is the human pose information; R(r) represents the movement of the robot, which is determined by the joint angle r; r lower and r upper are the lower and upper limits of the robot movement respectively, reflecting the constraints of the robot movement range.
[0112] Preferably, in step S5), the joint angle and link pose information of the robot output by the graph decoder are mapped through the optimized target mapping function L into the movement trajectory of the robot end effector that can best imitate human actions; the expression of the target mapping function L for motion redirection is:
[0113] L = λ ees L ees + λ ee L ee + λ ori L ori + λ fin L fin ;
[0114] Where, L ees is the velocity loss function; L ee is the end effector loss function; L ori is the direction loss function; L fin is the finger loss function; λ ees and λ ee and λ ori and λ fin are the weights of the corresponding loss functions respectively.
[0115] Preferably, in step S5), the expression of the velocity loss function L ees is:
[0116]
[0117] Where, s i and r i are the velocity of the i-th frame of the robot and the normalization coefficient; and are the velocity of the i-th frame of the human body and the normalization coefficient.
[0118] Preferably, in step S5), the expression of the direction loss function L ori is:
[0119]
[0120] where; D i is the rotation matrix of the end effector of the i-th frame of the robot; is the rotation matrix of the end effector of the i-th frame of the human body.
[0121] Preferably, in step S5), the finger loss function L fin has the following expression:
[0122]
[0123] where; f i tip is the fingertip position of the i-th frame of the robot; f i meta is the finger tip position of the i-th frame of the robot; r tm is the normalization coefficient of the i-th robot arm; is the fingertip position of the i-th frame of the human body; is the finger tip position of the i-th frame of the human body; is the normalization coefficient of the i-th human arm.
[0124] Preferably, in step S5), the end effector loss function L ee has the following expression:
[0125]
[0126] where: e i and are r i the position and normalization coefficient of the end effector of the i-th frame of the robot, and are the end effector and normalization coefficient of the i-th frame of the human.
[0127] The above embodiments and descriptions in the specification only illustrate the principles and the best embodiments of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed.
Claims
1. A robot redirection method based on graph neural network, characterized in that, The steps are as follows: S1), Collect the action video of the operator; and extract the human skeletal structure from the action video; S2), Construct a corresponding tree topology map according to the human skeletal structure and the robot structure; S3), Use the graph neural network encoder to encode and optimize the tree topology maps of the human body and the robot respectively, and obtain the latent space representations of the tree topology maps of the human body and the robot; Among them, using the graph neural network encoder to encode and optimize the tree topology maps of the human body and the robot specifically includes the following steps: S31), For the human body tree topology map, at each node, the input feature is the rotation angle of the corresponding joint, and at each edge, the input feature is the initial offset and rotation between adjacent edges; For the robot tree topology map, at each node, the output feature will represent the rotation angle of the corresponding joint, and at each edge, the output feature will represent the initial offset and rotation between the connecting rods; S32), Use the graph neural network encoder to encode the tree topology maps of the human body and the robot respectively to generate latent space representations; S33), Perform message passing and update on the node features and edge features in the graph neural network encoder, and repeatedly perform multi-level graph convolution operations; S34), Through the multi-level feature extraction and update of the graph neural network encoder, optimize the initial features of the input; and output a concise and efficient latent space representation; S4), Input the optimized latent space representations of the human and robot skeleton graphs obtained in step S3) into the latent neural network, and capture the latent features of the human-robot motion difference through the latent neural network to obtain the joint angle and link pose information of the robot; S5), Through the differentiable forward kinematics layer, convert the joint angle and link pose information output by the graph decoder into the desired position of the robot end effector; that is, complete the final motion redirection mapping.
2. The robot redirection method based on graph neural network according to claim 1, wherein: In step S1), after obtaining the action video, perform denoising, color correction, and frame rate adjustment processing on it; and use the visual detection algorithm OpenPose and the estimation algorithm Frankmocap algorithm to extract the human three-dimensional joint point coordinate data and attitude quaternion information.
3. The robot redirection method based on a graph neural network according to claim 2, wherein: In step S2), take the joints of the robot as nodes, and the connecting rods connecting two joints as undirected edges; model the structure of the robot as an acyclic tree topology graph Gr=(Vr,Er); where, Vr is the set of robot joint points; Er is the set of robot connecting rod edges; the acyclic tree topology graph of the human skeletal structure is Gh=(Vh,Eh); where, Vh is the set of human joint points; Er is the set of human skeletal connection edges.
4. A robot redirection method based on a graph neural network according to claim 1, characterized in that: In step S3), the graph neural network encoder adopts a hierarchical convolution structure, and the graph neural network encoder includes multiple stages, each stage includes multiple convolution layers of the same type, and each stage corresponds to processing features of different scales.
5. A robot redirection method based on a graph neural network according to claim 1, characterized in that: In step S4), the potential neural network includes a graph encoder and a graph decoder; wherein, the graph encoder performs node and edge update operations on the optimized encoding of the input human-robot skeleton graph through a graph convolutional network GCN. The graph convolutional network GCN is stacked by multiple graph convolutional layers, and each graph convolutional layer includes two operations: node update and edge update. The node update fuses the current node feature and the aggregated feature from adjacent nodes, and the edge update fuses the current edge feature and the aggregated feature from adjacent edges; the graph convolutional network GCN extracts high-level node features and edge features, and captures higher-level motion patterns and differences layer by layer. The graph decoder decodes the latent features through a fully connected layer or a transposed convolutional layer, and outputs the robot joint angles and link poses information.
6. The robot redirection method based on a graph neural network according to claim 5, characterized in that: In step S4), using the potential neural network to extract the robot's joint angles and link poses information, specifically includes the following steps: S41), Map the optimized encodings of the human and robot skeleton graphs into the potential neural network respectively, and then perform continuous constraint optimization through the graph encoder in the latent space to obtain the latent features that can minimize the representation of the human-robot difference; the expression for the graph encoder to update and fuse node features and edge features is: x_v t+1 = Φ(x_v t , ∑_{u∈N(v)} ψ(x_v t , x_u t , e_attr(v, u))); where \(x_v\) t+1 represents the updated feature vector of node \(v\) at time \(t + 1\); \(x_v\) t represents the feature vector of node \(v\) at time \(t\); \(N(v)\) is the set of neighbor nodes of node \(v\), \(u\) represents a neighbor node of \(v\), and \(x_u\) t represents the feature vector of node \(u\) at time \(t\); \(e_{attr}(v, u)\) is the feature vector of the edge connecting \(v\) and \(u\), and \(\psi\) is the message function for calculating the information vector based on node and edge information; \(\sum\) - \(\{u\in N(v)\}\) is used to aggregate the influence of all neighbor nodes; Φ is an update function that integrates its own information and the influence of neighboring nodes to update node features. S42), Input the optimized and aligned latent encoding into the graph decoder to decode the corresponding joint angles and link poses information. During the decoding process, use the prior structure information of the robot to ensure that the decoding output satisfies the kinematic constraints of the robot; the objective function for training the graph decoder is: where; z is the latent code; are the parameters of the graph decoder; represents the reconstruction loss function, which is used to measure the difference between the output of the graph decoder and the input X; Ω(z) is the regularization term imposed on the latent code z; λ is the weight coefficient of the regularization term.
7. The method for robot redirection based on a graph neural network according to claim 6, wherein: In step S5), use the differentiable forward kinematics layer to convert the joint angles and link poses information output by the graph decoder into the desired position of the robot end effector, specifically includes the following steps: S51), Take the robot joint angles and link poses information output by the graph encoder as input and pass it to the differentiable forward kinematics layer. S52), The forward kinematics layer calculates the transformation matrix of each joint according to the kinematic model and DOF settings of the robot. S53), Calculate the overall motion transformation from the base coordinate system to the end effector through the transformation matrices of each joint to obtain the desired position of the end effector. S54), Compare the desired position of the end effector with the real target position to obtain the position error. S55), Based on the position error, backpropagate to the graph decoder and the graph encoder to optimize the latent representation of the graph encoder and the output of the graph decoder. S56), Through multiple iterations of optimization, minimize the position error of the end effector to obtain an accurate motion redirection mapping; thus obtain a sequence of robot joint angles that can accurately reach the desired position and complete the action redirection mapping task.
8. The robot redirection method based on a graph neural network according to claim 7, characterized in that: In step S5), the target of the motion redirection is to determine a set of robot joint angles r* such that the motion posture R(r*) of the robot can imitate and fit the given human posture H to the greatest extent under the given constraints; at the same time, it cannot exceed the range of the robot's own joint angle limits. By optimizing the objective function and satisfying the constraints, the motion trajectory that enables the robot to best imitate human actions is obtained, that is: r * = argmin L(H, R(r)); s.t.r lower ≤r≤r upper ; where; r* is the robot joint angle; L represents the target mapping function from the joint angle to the final pose of the robot; H is the human pose information; R(r) represents the motion of the robot, which is determined by the joint angle r; r lower , r upper are the lower and upper limits of the robot motion respectively, reflecting the constraints of the robot motion range.
9. The robot redirection method based on a graph neural network according to claim 8, characterized in that: In step S5), the joint angle and link pose information of the robot output by the graph decoder are mapped to the motion trajectory of the robot end effector that can best imitate human actions through optimizing the target mapping function L; the expression of the target mapping function L of the motion redirection is: L = λ ees L ees + λ ee L ee + λ ori L ori + λ fin L fin ; Where, L ees is the velocity loss function; L ee is the end effector loss function; L ori is the direction loss function; L fin is the finger loss function; λ ees , λ ee , λ ori , λ fin are the weights of the respective loss functions.
Citation Information
Patent Citations
Motion capture missing data recovery method combining graph neural network and bone length constraint
CN116563952A
Sign language recognition method and device, electronic equipment and storage medium
CN117437694A