Robot control method based on heterogeneity modeling
By modeling the robot structure as a heterogeneous graph and constructing a heterogeneous graph Transformer model, the performance degradation problem of the general control strategy model in the existing technology on complex morphological robots is solved, and more efficient control performance and generalization capabilities are achieved.
Patent Information
- Application Number
- CN202510090251.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-30
AI Technical Summary
In the prior art, the general control strategy model has a performance degradation due to oversmoothing problems when dealing with complex robot morphology, and homogeneous graph modeling does not fully utilize different functions of structural units, resulting in poor control effects.
The robot control method based on heterogeneity modeling is adopted to model the robot structure as a heterogeneous graph, and a general control strategy model based on heterogeneous graph Transformer is constructed, and the model is trained through reinforcement learning to be suitable for robots of different forms.
Through heterogeneous graph modeling and Transformer model, the diversity and interaction relationship of robot structural units can be more accurately represented, control performance, reduce learning costs, and enhance the generalization ability of the model.
Smart Images

Figure CN120065815A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of voxel robots, and particularly to a robot control method based on heterogeneous modeling. Background Art
[0002] In the field of voxel robot control, the practice of combining reinforcement learning with robot control tasks to enable robots to automatically learn control strategies has achieved extensive success. However, this method has limitations, that is, the strategies of each robot need to be independently trained from scratch. As the number of robots increases, the learning and time costs brought by independent training strategies also increase significantly. To solve this problem, the concept of a general control strategy has been proposed. By regarding the training processes of robots with different morphologies as multi-task reinforcement learning problems, a single strategy model that can be applied to different morphologies is trained, thereby saving training costs. The general control strategy model also has the ability to quickly transfer to unknown morphologies.
[0003] The method of using a graph neural network as a control strategy model was the first to achieve general control. However, due to the inherent over-smoothing defect of the graph neural network, the more complex the robot morphology and the larger the number of structural units, the more significantly the performance of the general controller will decline. The method of using a Transformer to achieve general control avoids this problem and discovers the important role of the robot's morphological information in the performance of the general control strategy. However, due to the compromise on the structural integrity of the robot, the robot structure is defaultly simply modeled as a homogeneous graph structure, and all structural units and connection relationships of the robot are regarded as homogeneous nodes and edges. The homogenization processing of structural units may hinder the control model from utilizing their different functions, and the effect of the general control strategy is not good. Summary of the Invention
[0004] To solve some or all of the above technical problems existing in the prior art, the present invention provides a robot control method based on heterogeneous modeling.
[0005] The technical solution of the present invention is as follows:
[0006] There is provided a robot control method based on heterogeneous modeling, the method comprising:
[0007] Modeling the robot structure as a heterogeneous graph, wherein the robot structural units are used as graph nodes, the structural unit types are used as node types, the connection relationships between the robot structural units are used as graph edges, and the connection relationship types are used as edge types;
[0008] Constructing a general control strategy model based on a heterogeneous graph Transformer, the general control strategy model generating actions and obtaining rewards based on the heterogeneous graph structure of the robot and environmental observation information;
[0009] The general control policy model is trained by using reinforcement learning to maximize the average reward of the general control policy on robots of various forms;
[0010] The general control of robots of various forms is realized by using the trained general control policy model.
[0011] In an embodiment of the present invention, the general control policy model includes a feature extraction module, a node feature update module, and an action generation module. The general control policy model generates actions and obtains rewards based on the heterogeneous graph structure of the robot and environmental observation information, including:
[0012] The feature extraction module receives environmental observation information and robot form information, and generates initial node features of all nodes;
[0013] The node feature update module updates the initial node features to obtain the final node features;
[0014] The action generation module generates action signals according to the final node features.
[0015] In an embodiment of the present invention, the feature extraction module receives environmental observation information and robot form information, and generates initial node features of all nodes, further including:
[0016] The feature extraction module receives the form information and observation information of the robot. The observation information includes local observation information and global observation information. The local observation information comes from each structural unit of the robot, and the global observation information comes from the task scenario;
[0017] Determine the corresponding linear mapping according to the node type;
[0018] Project the node local observation information into the embedding vector of the corresponding feature space according to the determined linear mapping;
[0019] Sum the one-dimensional position of the node and the embedding vector to obtain the initial node features.
[0020] In an embodiment of the present invention, the feature extraction module is an Encoder module composed of linear layers.
[0021] In an embodiment of the present invention, the node feature update module updates the initial node features to obtain the final node features, specifically including:
[0022] Heterogeneous attention calculation: Calculate the attention score of the target node for each adjacent node, and consider the corresponding linear mapping relationship determined by the node type in the calculation process;
[0023] Heterogeneous information transmission: Calculate the transmitted heterogeneous information to obtain the heterogeneous information that the source node finally transmits to the target node after passing through a specific type of edge relationship;
[0024] Target-specific information fusion: Obtain the final output feature of the target node according to the attention score of the target node and the transmitted heterogeneous information.
[0025] In an embodiment of the present invention, the feature update module is a heterogeneous graph Transformer module.
[0026] In an embodiment of the present invention, during heterogeneous attention calculation, it passes through multiple stacked layers, and the output of each layer is the input of the next layer, so as to continuously update the node features.
[0027] In an embodiment of the present invention, obtaining the final output feature of the target node according to the attention score of the target node and the transmitted heterogeneous information specifically includes:
[0028] Multiply the attention matrix as the weight matrix with the information matrix, and then apply an activation function to map it back to a specific type of distribution belonging to the target node, and apply a residual connection to avoid the target node forgetting its own features.
[0029] In an embodiment of the present invention, the action generation module generates an action signal according to the final feature of the node, specifically including:
[0030] Input the final feature of the node into the decoder;
[0031] Input the global observation information into the decoder;
[0032] Multiply the output of the decoder with a covariance matrix of fixed parameters to generate an action distribution;
[0033] Sample the action distribution to obtain the action signal of each structural unit.
[0034] In an embodiment of the present invention, the decoder is a Decoder module with a linear layer.
[0035] The main advantages of the technical solution of the present invention are as follows:
[0036] The robot control method based on heterogeneous modeling of the present invention models the robot structure as a heterogeneous graph. The heterogeneous graph structure can more accurately represent the diversity and interaction relationships of the robot structure units, providing more sufficient morphological information for the control model. The heterogeneous graph Transformer can effectively understand the functional differences of different types of structure units and differentially utilize the information from different structure units according to the heterogeneity of the connection relationships, thereby improving the control performance. The trained general control strategy model can be applied to robots of different morphologies, thus reducing the learning cost and enhancing the generalization ability of the model. Brief Description of the Drawings
[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0038] Figure 1 It is a flowchart of the robot control method based on heterogeneous modeling according to an embodiment of the present invention;
[0039] Figure 2 It is a model diagram of the robot control method based on heterogeneous modeling provided by an embodiment of the present invention. Detailed Embodiments
[0040] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions of the present invention in conjunction with the specific embodiments and corresponding drawings of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0041] The following will detail the technical solutions provided by the embodiments of the present invention in conjunction with the drawings.
[0042] The embodiments of the present invention provide a robot control method based on heterogeneous modeling, as shown in the Figure 1 accompanying drawings, including:
[0043] S1, model the robot structure as a heterogeneous graph, where the robot structure units are used as graph nodes, the structure unit types are used as node types, the connection relationships between the robot structure units are used as graph edges, and the connection relationship types are used as edge types.
[0044] In the prior art, due to the compromise on the structural integrity of the robot, the robot structure is defaultly simply modeled as a homogeneous graph structure, and all structural units and connection relationships of the robot are regarded as homogeneous nodes and edges. The homogenization of structural units may hinder the control model from utilizing their different functions. And the assumption that all connection relationships between structural units have equal relevance has also been verified to be unreasonable. This excessive simplification of the robot's morphological information limits the model's ability to design the best general control strategy.
[0045] In the embodiments of the present invention, the robot structure is modeled as a heterogeneous graph, where the robot structural units are used as graph nodes, and the connection relationships between the robot structural units are used as graph edges, and the types of nodes and edges correspond to the types of structural units and connection relationships. For example, for a robot composed of wheels, joints, and limbs, the wheels, joints, and limbs can be used as different types of nodes respectively, and the connection relationships between them (such as the wheel is connected to the joint, and the joint is connected to the limb) are used as different types of edges.
[0046] The heterogeneous graph can more accurately represent the diversity and interaction relationships of the robot structural units. For example, the wheel node can have a "driving" type, the joint node can have a "connecting" type, and the limb node can have an "executing" type. Different types of nodes can have different feature dimensions and feature representation methods. For example, the wheel node can contain information such as radius and mass, the joint node can contain information such as angle and torque, and the limb node can contain information such as length and stiffness. Through heterogeneous graph modeling, the robot's morphology is accurately characterized, providing more sufficient morphological information for the general control strategy model.
[0047] S2. Construct a general control strategy model based on the heterogeneous graph Transformer. The general control strategy model generates actions and obtains rewards based on the heterogeneous graph structure of the robot and environmental observation information.
[0048] Using the heterogeneous graph Transformer to process the heterogeneous graph can more reasonably process the information from the heterogeneous graph, effectively understand the functional differences of different types of structural units, and differentially utilize the information from different structural units according to the heterogeneity of the connection relationships. The general strategy model trained in this way fully considers the differences between different robot morphologies and has superior control performance and morphological adaptability.
[0049] S3. Train the general control strategy model by means of reinforcement learning to maximize the average reward of the general control strategy on robots of various morphologies.
[0050] Train the general control policy model using reinforcement learning to maximize the average reward of the general control policy on robots of various forms. For example, the deep Q-learning algorithm or the policy gradient algorithm can be used for training. During the training process, the model needs to learn how to generate optimal actions based on the heterogeneous graph structure of the robot and the environmental observation information to obtain the maximum reward.
[0051] S4. Use the trained general control policy model to achieve the general control of robots of various forms.
[0052] With the trained general control policy model, the general control of robots of different forms can be achieved. For example, the trained model can be applied to robots of different forms, and control actions can be generated according to the current observation information and form information of the robot. Since the model has learned the common characteristics of robots of different forms, it can effectively control robots of different forms.
[0053] In summary, for the robot control method based on heterogeneous modeling provided by the embodiments of the present invention, by modeling the robot structure as a heterogeneous graph, the heterogeneous graph structure can more accurately represent the diversity and interaction relationships of the robot structure units, providing more sufficient form information for the control model; the heterogeneous graph Transformer can effectively understand the functional differences of different types of structure units and differentially utilize the information from different structure units according to the heterogeneity of the connection relationships, thereby improving the control performance. The trained general control policy model can be applied to robots of different forms, thus reducing the learning cost and enhancing the generalization ability of the model.
[0054] The following details each step and the related principles in the robot control method based on heterogeneous modeling provided by the embodiments of the present invention.
[0055] The training process of the general control policy can be regarded as a multi-task reinforcement learning problem. First, generate a set of K robots with different forms through the modular robot design space. All robots share the same control policy model, and use the reinforcement learning algorithm to train this control policy model. The training objective is that the learned general control policy can make the average reward of each form reach the maximum value.
[0056] Please refer to the appendix Figure 2, the model of the robot control method based on heterogeneity modeling provided by the embodiments of the present invention includes a feature extraction module, a node feature update module, and an action generation module. The feature extraction module is a linear Encoder layer that completes the preprocessing and embedding of the observation information. The node feature update module is a heterogeneous graph Transformer layer. In this layer, heterogeneous attention calculation and heterogeneous information transfer are first performed in parallel, and then the information output by both is fused with target-specific information. The action generation module is a linear Decoder layer that completes the generation of action signals.
[0057] We take a robot k randomly generated from the robot design space as an example to illustrate the specific method flow.
[0058] Step 1: Robot heterogeneity modeling
[0059] As shown in the appendix Figure 2 , for the robot k, it can be represented as a heterogeneous directed graph with distinguishable node and edge types: G=(V, E, U, P). Each node v i ∈V for i∈1…n represents a structural unit that makes up the robot, where n is the total number of structural units of the robot. The types of structural units correspond to the types of nodes. The directed edge e i,j ∈E represents the connection relationship from node v i to node v j . Each node and each edge are respectively associated with their type mapping functions τ(v): V→U and φ(e): E→P.
[0060] Step 2: Preprocessing and embedding of observation information
[0061] The general control strategy model generates actions and obtains rewards by receiving the observation information returned by the environment and the morphological information of the robot itself at each time step. Please refer to the appendix Figure 2 . The starting part of the model is an Encoder composed of linear layers. First, it is necessary to preprocess the observation information returned by the environment according to the heterogeneous graph structure of the robot in the Encoder.
[0062] The Encoder receives the morphological information M k (heterogeneous graph structure) of the kth robot and the observation information S k (which can be divided into local observation information and global observation information . The local observation information comes from each structural unit of the robot, and the global observation information comes from the task scenario). Relying on M k, when processing the feature information of each node, the model will also consider the type of the node. Specifically, it will learn an independent linear mapping for each node type, so as to be able to differentiate the distributions of different types of node features. The calculation formula is as follows:
[0063]
[0064] Here, Encoder τ(x) is a linear network related to the node type, N k represents all nodes in robot k, is the feature information corresponding to node x, τ(x) represents the type corresponding to node x, d is the feature dimension of the node, and Encoder τ(x) projects the τ(x)-type node into the corresponding embedding vector. In addition, we will also add the learnable one-dimensional position embedding W pos to the embedded node features to automatically learn the position information, and finally obtain the output H 0 (x) ∈ R dim , as the initial node feature of node x.
[0065] Step 3. Heterogeneous attention calculation
[0066] After obtaining the initial features of all nodes of the robot, it will enter the heterogeneous graph Transformer part. The heterogeneous attention module in the heterogeneous graph Transformer will be stacked for a total of L layers, and the output of each layer will be used as the input of the next layer to continuously update the node features. In each layer, heterogeneous attention calculation will only be performed between adjacent nodes.
[0067] For a given target node t and all its source nodes (neighbor nodes) s ∈ N(t), the multi-head attention mechanism is adopted, and the number of attention heads is h. The calculation formula is as follows:
[0068]
[0069]
[0070]
[0071]
[0072] Similarly, a corresponding linear mapping is selected according to the type of the node for calculation. For the i-th attention head the target node feature is used to generate the query vector Q i (t), and the source node feature is used to generate the key vector K i(s), the attention score of the target node t to the source node s at the i-th head can be obtained by taking the dot product of the query and the key. After obtaining the results of all attention heads, they are concatenated and softmax is applied to obtain the attention scores of the target node for each neighbor node.
[0073] Step 4: Heterogeneous information propagation
[0074] The heterogeneous information propagation process is carried out synchronously with the process of calculating heterogeneous attention, and the information is only transmitted from the source node to the corresponding target node. For a pair of nodes e = (s, t), the multi-head information is calculated through the following formula:
[0075]
[0076]
[0077] First, through linear projection The features of the source node s of type τ(s) are projected into the i-th information head Subsequently, all a total of h information heads are concatenated, and according to the edge type The corresponding edge relation weight matrix is selected And multiplied by the concatenated To obtain the information Message(s, t, e) that the source node s finally transmits to the target node t after passing through a specific type of edge relation.
[0078] Step 5: Target-specific information aggregation
[0079] During the target-specific information aggregation process, it is necessary to aggregate and transmit the source node information from different feature distributions to the target node. The importance of the information from each node is measured by the attention scores. The attention matrix is used as the weight matrix to multiply the information matrix, and then the activation function σ is applied to this vector and mapped back into the specific type of distribution belonging to the target node. At the same time, to avoid the target node forgetting its own feature information, we also use the technique of residual connection. The specific formula is as follows:
[0080]
[0081]
[0082] In this way, the output H l (t) of the target node t at the l-th layer is obtained, which is used to be delivered to the next layer. The output of the last layer is the final output feature of the target node t. The final output features of all nodes will be transmitted to the Decoder to generate actions.
[0083] Step 6: Generation of action signals
[0084] Each of the obtained final node features is input into a Decoder consisting of linear layers, which also receives the global observation information g learned by the linear network MLP and the state representation Gloabl k . The output μ of the Decoder and a diagonal covariance matrix Σ with fixed parameters are used to model the action distribution as a Gaussian distribution, thereby obtaining the control policy π θ . The formula is as follows:
[0085]
[0086] μ(S k ) = Decoder(H L (t), Gloabl k )
[0087] π θ (a k |s k ) = N(μ(S k ), Σ)
[0088] After obtaining the action distribution, sampling it can obtain the action signals corresponding to each structural unit of the robot.
[0089] Step 7. Train the general control policy model
[0090] Apply various reinforcement learning algorithms to train the general control policy model, with the goal of maximizing the average reward of the general control policy on robots of various forms. The finally trained general control policy model can achieve good generalization control for robots of any form.
[0091] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. In addition, in this article, "front", "rear", "left", "right", "up" and "down" are all referenced to the placement state shown in the drawings.
[0092] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A robot control method based on heterogeneity modeling, characterized in that: include: The robot structure is modeled as a heterogeneous graph, where the robot structural units are graph nodes, the structural unit types are node types, the connection relationships between the robot structural units are graph edges, and the connection relationship types are edge types. Construct a general control strategy model based on heterogeneous graph Transformer, which generates actions and obtains rewards based on the robot's heterogeneous graph structure and environmental observation information; The universal control strategy model is trained by reinforcement learning to maximize the average reward of the universal control strategy on robots of various forms; The trained general control strategy model is used to realize the general control of robots of various forms.
2. The robot control method based on heterogeneous modeling according to claim 1 is characterized in that: The general control strategy model includes a feature extraction module, a node feature update module, and an action generation module. The general control strategy model generates actions and obtains rewards based on the robot's heterogeneous graph structure and environmental observation information, including: The feature extraction module receives the environmental observation information and the robot morphology information, and generates the initial node features of all nodes; The node feature updating module updates the initial node features to obtain the final node features; The action generation module generates an action signal according to the final feature of the node.
3. The robot control method based on heterogeneous modeling according to claim 2 is characterized in that: The feature extraction module receives the environment observation information and the robot morphology information, generates the initial node features of all nodes, and further includes: The feature extraction module receives the morphological information and observation information of the robot, wherein the observation information includes local observation information and global observation information, wherein the local observation information comes from each structural unit of the robot, and the global observation information comes from the task scene; Determine the corresponding linear mapping according to the node type; According to the determined linear mapping, the local observation information of the node is projected into an embedding vector of the corresponding feature space; The one-dimensional position of the node is summed with the embedding vector to obtain the initial node features.
4. The robot control method based on heterogeneous modeling according to claim 3 is characterized in that: The feature extraction module is an Encoder module composed of linear layers.
5. The robot control method based on heterogeneous modeling according to claim 2 is characterized in that: The node feature updating module updates the initial node feature to obtain the final node feature, specifically including: Heterogeneous attention calculation: Calculate the attention score of the target node for each adjacent node. The calculation process considers the node type to determine the corresponding linear mapping relationship; Heterogeneous information transmission: Calculate the transmitted heterogeneous information and obtain the heterogeneous information that the source node finally transmits to the target node after passing through a specific type of edge relationship; Target-specific information fusion: The final output features of the target node are obtained according to the attention score of the target node and the transferred heterogeneous information.
6. The robot control method based on heterogeneous modeling according to claim 5 is characterized in that: The feature updating module is a heterogeneous graph Transformer module.
7. The robot control method based on heterogeneous modeling according to claim 6 is characterized in that: Heterogeneous attention calculations go through multiple stacked layers, and the output of each layer is the input of the next layer, thereby continuously updating the node features.
8. The robot control method based on heterogeneous modeling according to claim 7 is characterized in that: The final output feature of the target node is obtained according to the attention score of the target node and the transmitted heterogeneous information, specifically including: The attention matrix is multiplied with the information matrix as a weight matrix, and then an activation function is applied to it to map it back to the specific type of distribution belonging to the target node. A residual connection is applied to prevent the target node from forgetting its own features.
9. The robot control method based on heterogeneous modeling according to claim 2, characterized in that: The action generation module generates an action signal according to the final feature of the node, specifically including: Input the final node features into the decoder; Input global observation information into the decoder; Multiply the decoder output with the covariance matrix of fixed parameters to generate the action distribution; The action distribution is sampled to obtain the action signal of each structural unit.
10. The robot control method based on heterogeneous modeling according to claim 9, characterized in that: The decoder is a linear layer Decoder module.