Control method and device for coordinated motion of multiple robotic arms
Through the prediction model, the position information of the robot arm and workpiece is processed, and the expected stress information and action operation instructions are generated, which solves the problem of low stability in the coordinated movement of multiple robot arms and achieves more stable workpiece handling.
Patent Information
- Application Number
- CN202510533638.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2045-04-27
AI Technical Summary
There is a complex force/position coupling effect in the coordinated movement of multiple robotic arms, resulting in low stability and affecting the stability of workpiece handling.
The first predictive sub-model in the prediction model is used to process the position information of the robot arm and the workpiece, obtain the expected stress information of the workpiece, and generate action operation instructions matching the expected stress information through the second predictive sub-model, and control the robot arm to perform the handling task.
Through the application of the prediction model, the force/position coupling effect during the movement of the multi-robot arm is reduced, and the stability of the coordinated movement of the multi-robot arm and the stability of workpiece handling are improved.
Smart Images

Figure CN120056135B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of mechanical control and intelligent manufacturing, and more particularly to a control method and device for the coordinated motion of multiple robotic arms. Background Art
[0002] In industrial automation, intelligent manufacturing, and precision assembly, multiple robotic arms coordinate motion to move workpieces to target locations for assembly. Related technologies consider the motion of multiple robotic arms, but the complex force / position coupling effects they generate during movement can lead to low stability in their coordinated motion. Summary of the Invention
[0003] In view of the above problems, the present invention provides a control method and device for the coordinated motion of multiple robotic arms.
[0004] According to a first aspect of the present invention, there is provided a control method for the coordinated motion of multiple robotic arms, comprising: utilizing a first prediction sub-model in a prediction model to process robotic arm posture information of multiple robotic arms and workpiece posture information of a workpiece to obtain expected force information of the workpiece, wherein the multiple robotic arms are used to collaboratively carry the workpiece, the expected force information represents the force required for the workpiece to be in the expected workpiece posture, and the workpiece posture information includes the expected workpiece posture; utilizing a second prediction sub-model in the prediction model corresponding to each of the multiple robotic arms to respectively process the expected force information and the workpiece posture information to obtain action operation instructions for each of the multiple robotic arms that match the expected force information, wherein the prediction model is obtained based on training of an excitation function; and utilizing multiple action operation instructions to respectively control the corresponding robotic arms to perform the handling task on the workpiece.
[0005] According to an embodiment of the present invention, the first prediction sub-model in the prediction model is used to process the manipulator posture information of multiple manipulators and the workpiece posture information of the workpiece to obtain the expected force information of the workpiece, including: using the convolutional layer of the first prediction sub-model to process the multiple manipulator posture information and the workpiece posture information to obtain a directed graph, the directed graph including a sub-graph representing the manipulator posture information and a global node representing the workpiece posture information, and the global node is connected to the sub-graph; and inputting the directed graph into the attention layer of the first prediction sub-model to obtain the expected force information.
[0006] According to an embodiment of the present invention, inputting a directed graph into the attention layer of the first prediction sub-model to obtain expected force information includes: inputting the directed graph into the attention layer of the first prediction sub-model to obtain local node features corresponding to multiple robotic arms; aggregating information of multiple local node features based on global nodes to obtain global features; and predicting the expected force information based on the global features.
[0007] According to an embodiment of the present invention, the subgraph includes local nodes and subgraph directed edges, the local nodes represent the joint posture information of the joints in the robotic arm, the robotic arm includes multiple joints, and the subgraph directed edges represent the association relationship between the multiple joint posture information.
[0008] According to an embodiment of the present invention, the directed graph also includes a cross-subgraph directed edge, the multiple robotic arms include a first robotic arm and a second robotic arm, and the cross-subgraph directed edge represents the association relationship between the first target joint pose information of the first robotic arm and the second target joint pose information of the second robotic arm.
[0009] According to an embodiment of the present invention, the robot arm posture information includes the current robot arm posture and the expected robot arm posture, and the workpiece posture information also includes the current workpiece posture.
[0010] According to an embodiment of the present invention, the incentive function includes a team incentive sub-function, and the above method also includes: when the workpiece training posture of the workpiece meets the target workpiece posture condition, the preset team reward value is determined as the team incentive sub-function value corresponding to the team incentive sub-function, and the workpiece training posture is the workpiece posture after the robotic arm executes the training action operation instruction.
[0011] According to an embodiment of the present invention, the excitation function also includes an individual excitation sub-function, and the above method also includes: when the manipulator training posture of the manipulator meets the target manipulator posture condition, the preset arrival excitation value is determined as the first individual excitation sub-function value, and the manipulator training posture is the manipulator posture after the manipulator executes the training action operation instruction; when the manipulator training posture represents a collision between multiple manipulators, the preset collision excitation value is determined as the second individual excitation sub-function value; and when the training running time meets the time threshold condition, the preset timeout excitation value is determined as the third individual excitation sub-function value, and the training running time represents the time for multiple manipulators to execute their respective training action operation instructions. The individual excitation sub-function values corresponding to the individual excitation sub-function include at least one of the following: the first individual excitation sub-function value, the second individual excitation sub-function value and the third individual excitation sub-function value.
[0012] According to an embodiment of the present invention, the expected force information includes at least one of the following: force and moment acting on the center of mass of the workpiece.
[0013] The second aspect of the present invention provides a control device for the coordinated movement of multiple robotic arms, including: a first processing module, used to use a first prediction sub-model in a prediction model to process the robotic arm posture information of multiple robotic arms and the workpiece posture information of a workpiece to obtain expected force information of the workpiece, wherein multiple robotic arms are used to collaboratively carry workpieces, the expected force information represents the force required for the workpiece to be in the expected workpiece posture, and the workpiece posture information includes the expected workpiece posture; a second processing module, used to use a second prediction sub-model in the prediction model corresponding to each of the multiple robotic arms to respectively process the expected force information and the workpiece posture information to obtain action operation instructions for each of the multiple robotic arms that match the expected force information, wherein the prediction model is obtained based on training of an excitation function; and a control module, used to use multiple action operation instructions to respectively control the corresponding robotic arms to perform the handling task on the workpiece.
[0014] A third aspect of the present invention provides an electronic device, comprising: one or more processors; a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above method.
[0015] The fourth aspect of the present invention further provides a computer-readable storage medium having a computer program or instructions stored thereon, which implements the steps of the above method when the computer program or instructions are executed by a processor.
[0016] The fifth aspect of the present invention further provides a computer program product, comprising a computer program or instructions, which implement the steps of the above method when executed by a processor.
[0017] According to an embodiment of the present invention, by using the first prediction sub-model in the prediction model to process the manipulator arm posture information of multiple manipulators and the workpiece posture information of the workpiece, the expected force information of the workpiece is obtained; then, using the second prediction sub-model in the prediction model to process the expected force information and workpiece posture information, the motion operation instructions of each of the multiple manipulator arms that match the expected force information of the workpiece are accurately predicted, so that the forces acting on the workpiece by the multiple manipulator arms are balanced and the interaction between the multiple manipulator arms is reduced. Therefore, using multiple motion operation instructions to control the corresponding manipulator arms respectively to perform the handling task on the workpiece reduces the force / position coupling effect during the movement of the multiple manipulator arms, which is beneficial to the stability of the coordinated movement of the multiple manipulator arms. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The above contents and other objects, features and advantages of the present invention will become more apparent through the following description of the embodiments of the present invention with reference to the accompanying drawings, in which:
[0019] Figure 1A diagram showing an application scenario of a method for controlling coordinated motion of multiple robotic arms according to an embodiment of the present invention;
[0020] Figure 2 A flowchart of a method for controlling coordinated motion of multiple robotic arms according to an embodiment of the present invention is shown;
[0021] Figure 3 A schematic diagram of a prediction model according to an embodiment of the present invention is shown;
[0022] Figure 4 shows the iterative return value of the prediction model training according to an embodiment of the present invention;
[0023] Figure 5 shows the success rate of prediction model training according to an embodiment of the present invention;
[0024] Figure 6 A schematic diagram showing a method for controlling coordinated motion of multiple robotic arms according to an embodiment of the present invention is shown;
[0025] Figure 7 A structural block diagram of a control device for coordinated motion of multiple robotic arms according to an embodiment of the present invention is shown;
[0026] Figure 8 A block diagram of an electronic device that can be used to implement the method for controlling the coordinated motion of multiple robotic arms according to an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0027] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present invention. In the following detailed description, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of embodiments of the present invention. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessary confusion of the concept of the present invention.
[0028] The terms used herein are only for describing specific embodiments and are not intended to limit the present invention. The terms "comprise", "include", etc. used herein indicate the presence of the features, steps, operations and / or components, but do not exclude the presence or addition of one or more other features, steps, operations or components.
[0029] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0030] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).
[0031] In industrial automation, intelligent manufacturing, and precision assembly, multiple robotic arms coordinate motion to move workpieces to target locations for assembly. Related technologies consider the motion of multiple robotic arms, but the complex force / position coupling effects they generate during movement can lead to low stability in their coordinated motion.
[0032] In view of this, an embodiment of the present invention provides a control method for the coordinated movement of multiple robotic arms, including: using a first prediction sub-model in a prediction model to process the robotic arm posture information of multiple robotic arms and the workpiece posture information of a workpiece to obtain expected force information of the workpiece, wherein multiple robotic arms are used to collaboratively carry workpieces, the expected force information represents the force required for the workpiece to be in the expected workpiece posture, and the workpiece posture information includes the expected workpiece posture; using a second prediction sub-model in the prediction model corresponding to each of the multiple robotic arms to respectively process the expected force information and the workpiece posture information to obtain action operation instructions for each of the multiple robotic arms that match the expected force information, wherein the prediction model is obtained based on training of an excitation function; and using multiple action operation instructions to respectively control the corresponding robotic arms to perform the handling task on the workpiece.
[0033] Figure 1 An application scenario diagram of a method for controlling coordinated motion of multiple robotic arms according to an embodiment of the present invention is shown.
[0034] like Figure 1 As shown, the application scenario diagram of the control method for coordinated motion of multiple robotic arms includes a first robotic arm 110 , a second robotic arm 120 and a workpiece 130 .
[0035] The first robot arm 110 and the second robot arm 120 move in coordination to move the workpiece 130 from the ground to a target location.
[0036] According to embodiments of the present invention, the next motion of each of the multiple robotic arms can be predicted based on their current motion sequences. However, when multiple robotic arms coordinately transport a workpiece 130, force-position coupling effects often exist. The predicted motion of the robotic arms may also cause deviations in the smoothness of the transport of the workpiece 130.
[0037] The present invention reversely infers the respective motion operation instructions of the first robotic arm 110 and the second robotic arm 120 based on the force information of the workpiece 130, thereby effectively reducing the force-position coupling effect and improving the stability of transporting the workpiece 130.
[0038] The position of the workpiece 130 on the ground is the current position. The expected force information of the workpiece 130 at the target position can be the force F and the torque M acting on the center of mass of the workpiece. Based on the expected force information and the workpiece position information, the action operation instructions of the multiple robot arms that match the expected force information are obtained.
[0039] The first and second robotic arms 110 and 120 use the same control algorithm and can be controlled by the same server. The server can also execute the control method for coordinated multi-robotic arm motion. For example, the robotic arm control algorithm can include fuzzy control, sliding mode control, adaptive control, or PID control (Proportional-Integral-Derivative).
[0040] By inputting the motion operation instructions into the robotic arm control algorithm, smooth control of the first robotic arm 110 and / or the second robotic arm 120 can be achieved.
[0041] There is no limitation on the number of joints, motor types, end effectors, etc. of the first robotic arm 110 and the second robotic arm 120 .
[0042] It should be understood that Figure 1 The number of robotic arms in is only exemplary and any number of robotic arms may be provided.
[0043] In some embodiments, Figure 2 A flow chart of a method for controlling coordinated motion of multiple robotic arms according to an embodiment of the present invention is shown.
[0044] like Figure 2 As shown, the method for controlling the coordinated motion of multiple robotic arms in this embodiment includes operations S210 to S230.
[0045] In operation S210 , the robot arm posture information of the plurality of robot arms and the workpiece posture information of the workpiece are processed using a first prediction sub-model in the prediction model to obtain expected force information of the workpiece.
[0046] In operation S220, the expected force information and the workpiece posture information are processed respectively using the second prediction sub-models corresponding to the multiple robotic arms in the prediction model to obtain the motion operation instructions of the multiple robotic arms that match the expected force information.
[0047] In operation S230 , a plurality of motion operation instructions are used to control the corresponding robotic arms respectively to perform a transport task on the workpiece.
[0048] According to an embodiment of the present invention, a plurality of robotic arms are used to collaboratively carry a workpiece, and the plurality of robotic arms can carry the workpiece to a target position of a target workpiece to complete assembly or fitting between the workpiece and the target workpiece.
[0049] For example, the workpiece could be a vehicle engine, and the target workpiece could be a vehicle frame, with a target location for the engine on the frame. The engine is collaboratively moved to the target location on the frame using the end effectors (e.g., grippers) of multiple robotic arms. The end effectors are then replaced with welding guns, and the engine is then mounted on the frame using the welding guns of each of the robotic arms.
[0050] According to an embodiment of the present invention, the manipulator pose information may include a reference manipulator pose, a joint pose, a link pose, an actuator pose, an expected manipulator pose, and the like.
[0051] The baseline manipulator pose can also be called the initial pose in the initial coordinate system, which can be used as a reference position and pose of the manipulator in the workspace.
[0052] The joint pose can be the position and posture of each joint in the robotic arm.
[0053] The link pose can be the relative position and posture of the structural parts between each joint in the robotic arm relative to the initial coordinate system.
[0054] An actuator can be the end of a robotic arm. For example, an actuator can be a clamp, welding gun, etc. The actuator pose can be the position and orientation of the actuator relative to the initial coordinate system.
[0055] The desired robot arm posture may be a position and posture that the robot arm needs to achieve in order to place the workpiece in the desired workpiece posture.
[0056] According to an embodiment of the present invention, the expected force information represents the force required for the workpiece to be in the expected workpiece posture. The workpiece posture information includes the expected workpiece posture.
[0057] For example, the expected force information may be an upward support force of 5 Newtons to lift the workpiece by 5 centimeters.
[0058] According to an embodiment of the present invention, the prediction model is obtained through training based on an activation function.
[0059] According to an embodiment of the present invention, the prediction model may include a first prediction sub-model and multiple second prediction sub-models. The first prediction sub-model and the second prediction sub-model may be the same or different machine learning models.
[0060] For example, the first prediction sub-model can be a deep learning model, and the second prediction sub-model can be a random forest model. The deep learning model can learn the relationship between multiple robotic arms from their manipulator pose information, and then use this relationship and the workpiece pose information to determine the expected force information for the workpiece. The random forest model can then determine the manipulator's motion instructions from multiple candidate motion instructions based on the expected force information and the workpiece pose information.
[0061] For example, the first and second prediction sub-models can be neural network models. However, the second prediction sub-models for multiple robotic arms can form a distributed policy network. These multiple second prediction sub-models can collaboratively learn by sharing information about the robotic arm's position and posture, taking into account not only the movement of that particular robotic arm but also the movement of other robotic arms, enabling coordinated movement of multiple robotic arms and reducing collisions between them.
[0062] According to an embodiment of the present invention, by using the first prediction sub-model in the prediction model to process the manipulator arm posture information of multiple manipulators and the workpiece posture information of the workpiece, the expected force information of the workpiece is obtained; then, using the second prediction sub-model in the prediction model to process the expected force information and workpiece posture information, the motion operation instructions of each of the multiple manipulator arms that match the expected force information of the workpiece are accurately predicted, so that the forces acting on the workpiece by the multiple manipulator arms are balanced and the interaction between the multiple manipulator arms is reduced. Therefore, using multiple motion operation instructions to control the corresponding manipulator arms respectively to perform the handling task on the workpiece reduces the force / position coupling effect during the movement of the multiple manipulator arms, which is beneficial to the stability of the coordinated movement of the multiple manipulator arms.
[0063] According to an embodiment of the present invention, the first prediction sub-model in the prediction model is used to process the manipulator posture information of multiple manipulators and the workpiece posture information of the workpiece to obtain the expected force information of the workpiece, including: using the convolutional layer of the first prediction sub-model to process the multiple manipulator posture information and the workpiece posture information to obtain a directed graph, the directed graph including a sub-graph representing the manipulator posture information and a global node representing the workpiece posture information, and the global node is connected to the sub-graph; and inputting the directed graph into the attention layer of the first prediction sub-model to obtain the expected force information.
[0064] According to an embodiment of the present invention, the first prediction sub-model may be a Graph Attention Network (GAT). For example, the first prediction sub-model may include a convolutional layer, an attention layer, a normalization layer, a fully connected layer, and the like.
[0065] According to an embodiment of the present invention, the subgraph may include multiple nodes, and the nodes represent joint pose information of the robotic arm.
[0066] The first prediction sub-model can enhance the expressive power of the model by introducing an attention mechanism to weight the neighbors of nodes in the directed graph.
[0067] According to an embodiment of the present invention, the role of the global node may be to globally model the relationship with the nodes in the directed graph, optimize the information flow in the global scope of the directed graph, and avoid deviation of local information of the subgraph.
[0068] According to an embodiment of the present invention, inputting a directed graph into the attention layer of the first prediction sub-model to obtain expected force information includes: inputting the directed graph into the attention layer of the first prediction sub-model to obtain local node features corresponding to multiple robotic arms; aggregating information of multiple local node features based on global nodes to obtain global features; and predicting the expected force information based on the global features.
[0069] The first prediction sub-model can include three attention layers. These layers extract features from the directed graph, obtaining local node features corresponding to each of the multiple robotic arms. Each node aggregates information from the input robotic arm and workpiece pose information to obtain node aggregate information. This node aggregate information is used to update global node features.
[0070] The node update equation of the graph attention network is:
[0071] (1)
[0072] Represents the feature vector of node i in the l+1th attention layer. K represents the number of linear projection heads, and the value of K can be 3. Concat represents concatenating the outputs of K independent attention heads. Each attention head generates a feature vector, and the final feature representation is obtained after concatenation. ELU represents Exponential Linear Unit, an activation function. represents the learnable weight matrix in the kth attention head, which is used to linearly transform the node features of the lth attention layer; represents the feature vector of node j in the lth attention layer; represents the normalized attention weight of node i to neighbor node j in the lth attention layer and the kth attention head, which is specifically defined as:
[0073] (2)
[0074] (3)
[0075] exp represents the exponential function, which is used to convert the attention score into a positive number; represents the unnormalized attention weight, calculated by formula (3); represents the sum of the unnormalized attention weights of all neighbors j′ of node i (including node i itself) to achieve normalization; Represents the Rectified LinearUnit, an activation function; represents the learnable weight vector used to calculate the attention weight in the k-th attention head of the l-th attention layer; Represents the feature vector of node i in the lth attention layer. Through formulas (1), (2) and (3), GAT learns an average of the neighborhood features of each node and sparsely weights them by the importance of each neighborhood.
[0076] The second prediction sub-model can be a Long Short-Term Memory (LSTM) network. For example, the distributed policy network consists of n LSTM networks, where n represents the number of robotic arms. The LSTM network can extract a fixed-size sequence of robotic arm poses from the current and historical poses of the robotic arm.
[0077] According to an embodiment of the present invention, the expected force information includes at least one of the following: force and moment acting on the center of mass of the workpiece.
[0078] Each second prediction sub-model extracts information from the expected force information and the workpiece posture information, and outputs an action operation instruction that matches the expected force information.
[0079] For example, the robot pose information of multiple robot arms and the workpiece pose information of each workpiece can be a time-sequential sequence of machine poses and workpiece poses, respectively. These sequences can form a 256-dimensional vector. After state extraction, the first prediction sub-model outputs a 12-dimensional vector through three fully connected layers with a tanh activation function (hyperbolic tangent function). This 12-dimensional vector serves as the input to the second prediction sub-model, namely the expected force information and workpiece pose information.
[0080] The second prediction sub-model can first predict the action value of each robot arm, and then determine the action operation instruction according to the action value of each robot arm. Each second prediction sub-model has the same structure but different network parameters.
[0081] According to an embodiment of the present invention, performing a workpiece assembly task requires that each of n robotic arms has a corresponding sequence of motion instructions. These sequences are motion instructions corresponding to action values, arranged in a sequence of motions. Expected force information can be used to pre-calibrate the action values, enabling the predicted motion instructions to achieve more stable coordinated motion among the multiple robotic arms.
[0082] According to an embodiment of the present invention, the subgraph includes local nodes and subgraph directed edges, the local nodes represent the joint posture information of the joints in the robotic arm, the robotic arm includes multiple joints, and the subgraph directed edges represent the association relationship between the multiple joint posture information.
[0083] According to an embodiment of the present invention, the association relationship between multiple joint posture information may be a spatial position relationship, a dependency relationship of handling tasks, etc. For example, the movement directions of joint postures A and B in a robotic arm are different.
[0084] According to an embodiment of the present invention, the directed graph also includes a cross-subgraph directed edge, the multiple robotic arms include a first robotic arm and a second robotic arm, and the cross-subgraph directed edge represents the association relationship between the first target joint pose information of the first robotic arm and the second target joint pose information of the second robotic arm.
[0085] According to an embodiment of the present invention, the association between the first target joint pose information of the first robotic arm and the second target joint pose information of the second robotic arm can be a spatial positional relationship or a dependency relationship of the handling task. Cross-subgraph directed edges can also represent communication relationships between multiple robotic arms, relative motion between robotic arms, collision constraints, and so on.
[0086] According to an embodiment of the present invention, the robot arm posture information includes the current robot arm posture and the expected robot arm posture, and the workpiece posture information also includes the current workpiece posture.
[0087] Figure 3 A schematic diagram of a prediction model according to an embodiment of the present invention is shown.
[0088] like Figure 3 As shown, the prediction model includes a first prediction sub-model and multiple second prediction sub-models. The first prediction sub-model can be a graph attention network, which inputs the workpiece posture information and the robot arm posture information into the first prediction sub-model to obtain the expected force information.
[0089] The workpiece pose information may include the current workpiece pose and the expected workpiece pose. The robot arm pose information may include the current robot arm pose and the expected robot arm pose.
[0090] The second prediction sub-model can be LSTM. n robotic arms correspond to n LSTMs. For example, n robotic arms correspond to LSTM1...LSTM n .
[0091] The workpiece posture information and expected force information are input into LSTM1 to obtain the first action operation instruction of the first robot arm. The workpiece posture information and expected force information are input into LSTMn to obtain the nth action operation instruction of the nth robot arm.
[0092] According to an embodiment of the present invention, the incentive function includes a team incentive sub-function, and the above method also includes: when the workpiece training posture of the workpiece meets the target workpiece posture condition, the preset team reward value is determined as the team incentive sub-function value corresponding to the team incentive sub-function, and the workpiece training posture is the workpiece posture after the robotic arm executes the training action operation instruction.
[0093] According to an embodiment of the present invention, in order to promote cooperative behavior, all robotic arms are rewarded only when the workpiece successfully reaches the desired workpiece position. The team incentive sub-function is as follows:
[0094] (4)
[0095] p am , p bm Represent the current workpiece pose and the expected workpiece pose at both ends of the workpiece, respectively, where , δ w is the threshold of the team incentive sub-function.
[0096] When m=1, p a1 and p b1 They correspond to the current workpiece pose and the expected workpiece pose at one end of the workpiece respectively.
[0097] When m=2, p a2 and p b2 They correspond to the current workpiece pose and the expected workpiece pose at the other end of the workpiece, respectively.
[0098] According to an embodiment of the present invention, the team incentive function focuses on multiple robotic arms carrying the workpiece to the desired workpiece posture, which can promote collaboration between multiple second prediction sub-models and reduce conflicts. According to an embodiment of the present invention, the incentive function also includes an individual incentive sub-function, and the above method also includes: when the robotic arm training posture of the robotic arm meets the target robotic arm posture condition, the preset arrival incentive value is determined as the first individual incentive sub-function value, and the robotic arm training posture is the robotic arm posture after the robotic arm executes the training action operation instruction; when the robotic arm training posture represents a collision between multiple robotic arms, the preset collision incentive value is determined as the second individual incentive sub-function value; and when the training run time meets the time threshold condition, the preset timeout incentive value is determined as the third individual incentive sub-function value, and the training run time represents the time for multiple robotic arms to execute their respective training action operation instructions. The individual incentive sub-function values corresponding to the individual incentive sub-function include at least one of the following: the first individual incentive sub-function value, the second individual incentive sub-function value, and the third individual incentive function value.
[0099] According to an embodiment of the present invention, the target robotic arm posture conditions may include that the distance between the position of the robotic arm training posture and the expected position of the expected robotic arm posture is less than a distance threshold, the posture of the robotic arm training posture reaches the expected robotic arm posture, and when the running direction of the robotic arm is consistent with the direction of the expected robotic arm posture.
[0100] According to an embodiment of the present invention, the time threshold condition may be that the training run time exceeds the time threshold.
[0101] Individual excitation subfunction The formula is as follows:
[0102] (5)
[0103] Indicates the distance between the position of the robot arm training pose and the expected position of the desired robot arm pose Less than the distance threshold δ L , the obtained position reaches the incentive value; Indicates the preset collision excitation value obtained when a collision occurs between multiple robotic arms; Indicates when a collision occurs between multiple robotic arms. Indicates when the training runs When the time threshold T is exceeded, the preset timeout incentive value is obtained; It means that when the posture of the training posture of the robot arm reaches the desired posture of the robot arm, the posture reaches the incentive value; approaching the goal means that when the posture of the training posture of the robot arm reaches the desired posture of the robot arm, the posture reaches the incentive value. Indicates the direction excitation value obtained when the robot arm's running direction matches the direction of the expected robot arm posture. Indicates the direction of the desired robot arm pose.
[0104] According to an embodiment of the present invention, the preset arrival incentive value may include at least one of a position arrival incentive value, a posture arrival incentive value, and a direction incentive value.
[0105] According to an embodiment of the present invention, the individual incentive sub-function may be an optimization objective function for multiple second prediction sub-models, and the team incentive sub-function may be a collaboration between the first prediction sub-model and multiple second prediction sub-models in the prediction model to achieve global optimization.
[0106] According to an embodiment of the present invention, by measuring the individual performance of the second prediction sub-model through the individual incentive sub-function, its own behavioral performance can be optimized. Combined with the team incentive sub-function, it can not only realize the individual learning of the second prediction sub-model, but also promote the collaboration between the first prediction sub-model and multiple second prediction sub-models, so that the action operation instructions of multiple robotic arms can be predicted more accurately.
[0107] Figure 4 It shows the iterative return value of the prediction model training according to an embodiment of the present invention.
[0108] like Figure 4 As shown in the figure, "Our method" can refer to the results obtained by training the prediction model using the excitation function of the present invention, and "Baseline method" refers to the results obtained by training the prediction model using the baseline method of directly controlling each robotic arm. As the number of training steps increases, the iterative return value of the training of the present invention tends to converge and reach a stable value, and the result is better than the baseline method.
[0109] Figure 5 The success rate of the prediction model training according to an embodiment of the present invention is shown.
[0110] like Figure 5 As shown in the figure, "Our method" refers to the results obtained by training the prediction model using the excitation function of the present invention, and "Baseline method" refers to the results obtained by training the prediction model using the baseline method of directly controlling each robotic arm. As the number of training steps increases, the success rate of the training of the present invention increases and converges to around 0.95, which is better than the baseline method.
[0111] In some embodiments, Figure 6 A schematic diagram of a method for controlling coordinated motion of multiple robotic arms according to an embodiment of the present invention is shown.
[0112] like Figure 6 As shown, during the offline training phase, the prediction model 620 is trained based on the activation function using the simulated manipulator pose information of each of the multiple simulated manipulators and the simulated workpiece pose information of the simulated workpiece in the simulation environment 610. The activation function can optimize the parameters of the prediction model 620.
[0113] The online control phase involves deploying the prediction model 620 to the server 630 . The robot arm pose information can be obtained from sensors mounted on the robot arm, or from an image acquisition device or sensors mounted on the workpiece. The desired workpiece pose and the desired robot arm pose can be obtained from the client 650 .
[0114] The server 630 is used to use the first prediction sub-model in the prediction model 620 to process the robot arm posture information of multiple robot arms and the workpiece posture information of the workpiece to obtain the expected force information of the workpiece; and use the second prediction sub-model in the prediction model 620 corresponding to each of the multiple robot arms to respectively process the expected force information and the workpiece posture information to obtain the action operation instructions of each of the multiple robot arms that match the expected force information.
[0115] The server 630 is further configured to send motion operation instructions to the position controller 640 of the robotic arm.
[0116] The position controller 640 is used to control the movement of the end effector of the robot arm and send the robot arm posture information and task feedback results of the robot arm to the client 650. The task feedback results may include workpiece posture information and workpiece posture error values.
[0117] The client 650 (eg, a control interface) is used to send a target task (eg, specify a grabbing location) or a cancellation instruction.
[0118] In some embodiments, Figure 7 A structural block diagram of a control device for coordinated motion of multiple robotic arms according to an embodiment of the present invention is shown.
[0119] like Figure 7 As shown, the control device 700 for coordinated motion of multiple robotic arms in this embodiment includes a first processing module 710 , a second processing module 720 and a control module 730 .
[0120] The first processing module 710 is configured to process the manipulator pose information of the multiple manipulators and the workpiece pose information of the workpiece using the first prediction sub-model in the prediction model to obtain expected force information for the workpiece, wherein the multiple manipulators are configured to collaboratively carry the workpiece, the expected force information represents the force required for the workpiece to be in the expected pose, and the workpiece pose information includes the expected pose. In one embodiment, the first processing module 710 can be configured to perform operation S210 described above, which will not be further described here.
[0121] The second processing module 720 is configured to use the second prediction sub-models corresponding to each of the multiple robotic arms in the prediction model to respectively process the expected force information and workpiece pose information, thereby obtaining motion instructions for each of the multiple robotic arms that match the expected force information. The prediction model is trained based on an incentive function, which includes an individual incentive sub-function and a team incentive sub-function. In one embodiment, the second processing module 720 can be configured to perform operation S220 described above, which will not be further described here.
[0122] The control module 730 is used to control the corresponding robotic arms using a plurality of motion operation instructions to perform a handling task on the workpiece. In one embodiment, the control module 730 can be used to perform the operation S230 described above, which will not be repeated here.
[0123] According to an embodiment of the present invention, the first processing module 710 includes a first processing submodule and an input submodule. The first processing submodule is configured to process multiple robot arm pose information and workpiece pose information using the convolutional layer of the first prediction submodel to obtain a directed graph. The directed graph includes a subgraph representing the robot arm pose information and a global node representing the workpiece pose information, with the global node connected to the subgraph. The input submodule is configured to input the directed graph into the attention layer of the first prediction submodel to obtain expected force information.
[0124] According to an embodiment of the present invention, the input submodule includes an input unit, an information aggregation unit, and a prediction unit. The input unit is used to input the directed graph into the attention layer of the first prediction submodel to obtain local node features corresponding to each of the multiple robotic arms; the information aggregation unit is used to aggregate the multiple local node features based on the global node to obtain global features; and the prediction unit is used to predict the expected force information based on the global features.
[0125] According to an embodiment of the present invention, the subgraph includes local nodes and subgraph directed edges, the local nodes represent the joint posture information of the joints in the robotic arm, the robotic arm includes multiple joints, and the subgraph directed edges represent the association relationship between the multiple joint posture information.
[0126] According to an embodiment of the present invention, the directed graph also includes a cross-subgraph directed edge, the multiple robotic arms include a first robotic arm and a second robotic arm, and the cross-subgraph directed edge represents the association relationship between the first target joint pose information of the first robotic arm and the second target joint pose information of the second robotic arm.
[0127] According to an embodiment of the present invention, the robot arm posture information includes the current robot arm posture and the expected robot arm posture, and the workpiece posture information also includes the current workpiece posture.
[0128] According to an embodiment of the present invention, the incentive function includes a team incentive sub-function, and the apparatus further includes a first determination module. The first determination module is configured to determine a preset team reward value as a team incentive sub-function value corresponding to the team incentive sub-function when a workpiece training pose of the workpiece satisfies a target workpiece pose condition, the workpiece training pose being the pose of the workpiece after the robotic arm executes a training action operation instruction.
[0129] According to an embodiment of the present invention, the excitation function further includes an individual excitation sub-function, and the above-mentioned device further includes a second determination module, a third determination module, and a fourth determination module. The second determination module is used to determine the preset arrival excitation value as the first individual excitation sub-function value when the manipulator arm training posture of the manipulator meets the target manipulator arm posture condition, and the manipulator arm training posture is the manipulator arm posture after the manipulator executes the training action operation instruction; the third determination module is used to determine the preset collision excitation value as the second individual excitation sub-function value when the manipulator arm training posture represents a collision between multiple manipulators; and the fourth determination module is used to determine the preset timeout excitation value as the third individual excitation sub-function value when the training run time meets the time threshold, and the training run time represents the time for multiple manipulators to respectively execute their respective training action operation instructions. The individual excitation sub-function values corresponding to the individual excitation sub-function include at least one of the following: the first individual excitation sub-function value, the second individual excitation sub-function value, and the third individual excitation sub-function value.
[0130] According to an embodiment of the present invention, the expected force information includes at least one of the following: force and moment acting on the center of mass of the workpiece.
[0131] According to embodiments of the present invention, any multiple modules among the first processing module 710, the second processing module 720, and the control module 730 may be combined into a single module, or any one of these modules may be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules may be combined with at least part of the functionality of other modules and implemented in a single module. According to embodiments of the present invention, at least one of the first processing module 710, the second processing module 720, and the control module 730 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or may be implemented in hardware or firmware through any other reasonable means of circuit integration or packaging, or may be implemented in any one of software, hardware, and firmware, or any suitable combination of these. Alternatively, at least one of the first processing module 710, the second processing module 720, and the control module 730 may be at least partially implemented as a computer program module that, when executed, performs the corresponding functionality.
[0132] According to an embodiment of the present invention, the present invention further provides an electronic device, a readable storage medium and a computer program product.
[0133] Figure 8 A block diagram of an electronic device that can be used to implement the method for controlling the coordinated motion of multiple robotic arms according to an embodiment of the present invention is shown.
[0134] Electronic device is intended to refer to various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Electronic device may also refer to various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are intended to be examples only and are not intended to limit the implementation of the inventions described and / or claimed herein.
[0135] like Figure 8 As shown, the device includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. Various programs and data required for the operation of the device 800 can also be stored in the RAM 803. The computing unit 801, ROM 802, and RAM 803 are connected to each other via a bus 804. An I / O interface 805 is also connected to the bus 804.
[0136] Multiple components in the electronic device are connected to the I / O interface 805, including an input unit 806, such as a keyboard, mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, optical disk, etc.; and a communication unit 809, such as a network card, modem, wireless communication transceiver, etc. The communication unit 809 allows the device to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0137] The computing unit 801 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as the avatar driving method. For example, in some embodiments, the avatar driving method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the avatar driving method described above can be performed. Alternatively, in other embodiments, the computing unit 801 can be configured to perform the multi-manipulator coordinated motion control method through any other suitable means (e.g., via firmware).
[0138] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-a-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0139] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. Such program code can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0140] In the context of the present invention, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0141] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0142] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0143] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. This client-server relationship is established by computer programs running on the respective computers, establishing a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a host product within the cloud computing service ecosystem, addressing the management difficulties and limited scalability of traditional physical hosts and VPS ("Virtual Private Server") services. The server may also be a server in a distributed system or a server integrated with blockchain.
[0144] Those skilled in the art will appreciate that various combinations and / or combinations of features described in the various embodiments and / or claims of the present invention may be made, even if such combinations and / or combinations are not explicitly described in the present invention. In particular, various combinations and / or combinations of features described in the various embodiments and / or claims of the present invention may be made, without departing from the spirit and teachings of the present invention. All such combinations and / or combinations fall within the scope of the present invention.
[0145] Although the present invention has been shown and described with reference to specific exemplary embodiments thereof, it will be understood by those skilled in the art that various changes in form and details may be made to the present invention without departing from the spirit and scope of the invention as defined by the appended claims and their equivalents. Therefore, the scope of the present invention should not be limited to the above-described embodiments, but should be determined not only by the appended claims but also by the equivalents of the appended claims.
Claims
1. A method for controlling the coordinated motion of multiple robotic arms, characterized in that: The method comprises: Processing robot arm posture information of a plurality of robot arms and workpiece posture information of a workpiece using a first prediction sub-model in a prediction model to obtain expected force information of the workpiece, wherein the plurality of robot arms are used to collaboratively carry the workpiece, the expected force information represents the force required for the workpiece to be in the expected workpiece posture, and the workpiece posture information includes the expected workpiece posture; using a second prediction sub-model in the prediction model corresponding to each of the plurality of robotic arms to respectively process the expected force information and the workpiece posture information to obtain motion operation instructions for each of the plurality of robotic arms that match the expected force information, wherein the prediction model is obtained based on activation function training; and Using the plurality of motion operation instructions to respectively control the corresponding robotic arms to perform a transport task on the workpiece; The method of using a first prediction sub-model in the prediction model to process the manipulator pose information of the plurality of manipulators and the workpiece pose information of the workpiece to obtain expected force information of the workpiece includes: Processing the plurality of the robotic arm pose information and the workpiece pose information using the convolutional layer of the first prediction sub-model to obtain a directed graph, wherein the directed graph includes a subgraph representing the robotic arm pose information and a global node representing the workpiece pose information, wherein the global node is connected to the subgraph; and The directed graph is input into the attention layer of the first prediction sub-model to obtain the expected force information.
2. The method according to claim 1, characterized in that Inputting the directed graph into the attention layer of the first prediction sub-model to obtain the expected force information includes: Inputting the directed graph into the attention layer of the first prediction sub-model to obtain local node features corresponding to each of the plurality of robotic arms; Aggregating information of the plurality of local node features based on the global node to obtain a global feature; and The expected force information is obtained based on the global feature prediction.
3. The method according to claim 1, characterized in that The subgraph includes local nodes and subgraph directed edges, the local nodes represent joint pose information of the joints in the robotic arm, the robotic arm includes multiple joints, and the subgraph directed edges represent the association relationship between the multiple joint pose information.
4. The method according to claim 1, wherein The directed graph also includes a cross-subgraph directed edge, the multiple robotic arms include a first robotic arm and a second robotic arm, and the cross-subgraph directed edge represents the association relationship between the first target joint pose information of the first robotic arm and the second target joint pose information of the second robotic arm.
5. The method according to claim 1, wherein The robot arm posture information includes the current robot arm posture and the expected robot arm posture, and the workpiece posture information also includes the current workpiece posture.
6. The method according to claim 1, characterized in that The incentive function includes a team incentive sub-function, and the method further includes: When the workpiece training posture of the workpiece meets the target workpiece posture condition, the preset team reward value is determined as the team incentive sub-function value corresponding to the team incentive sub-function, and the workpiece training posture is the workpiece posture after the robot arm executes the training action operation instruction.
7. The method according to claim 6, characterized in that The excitation function further includes an individual excitation sub-function, and the method further includes: When the manipulator arm training posture of the manipulator arm satisfies the target manipulator arm posture condition, determining the preset arrival excitation value as the first body excitation sub-function value, wherein the manipulator arm training posture is the manipulator arm posture after the manipulator arm executes the training action operation instruction; In the case where the training posture of the manipulator represents a collision between a plurality of the manipulators, determining a preset collision excitation value as a second individual excitation sub-function value; and When the training running time meets the time threshold condition, the preset timeout incentive value is determined as the third individual incentive sub-function value. The training running time represents the time for the multiple robotic arms to respectively execute their respective training action operation instructions. The individual incentive sub-function value corresponding to the individual incentive sub-function includes at least one of the following: the first individual incentive sub-function value, the second individual incentive sub-function value and the third individual incentive sub-function value.
8. The method according to claim 1, characterized in that The expected force information includes at least one of the following: force and moment acting on the center of mass of the workpiece.
9. A control device for coordinated motion of multiple robotic arms, characterized in that: The device comprises: a first processing module, configured to process, using a first prediction sub-model in the prediction model, robot arm posture information of a plurality of robot arms and workpiece posture information of a workpiece to obtain expected force information of the workpiece, wherein the plurality of robot arms are used to collaboratively carry the workpiece, the expected force information represents a force required for the workpiece to be in the expected workpiece posture, and the workpiece posture information includes the expected workpiece posture; a second processing module, configured to respectively process the expected force information and the workpiece posture information using second prediction sub-models in the prediction model corresponding to each of the plurality of robotic arms, to obtain motion operation instructions for each of the plurality of robotic arms that match the expected force information, wherein the prediction model is obtained based on training of an incentive function, and the incentive function includes an individual incentive sub-function and a team incentive sub-function; and A control module, configured to use the plurality of motion operation instructions to respectively control the corresponding robotic arms to perform a transport task on the workpiece; The first processing module includes: A first processing submodule is configured to process the plurality of robot arm pose information and the workpiece pose information using a convolutional layer of the first prediction submodel to obtain a directed graph, wherein the directed graph includes a subgraph representing the robot arm pose information and a global node representing the workpiece pose information, wherein the global node is connected to the subgraph; An input submodule is used to input the directed graph into the attention layer of the first prediction submodel to obtain the expected force information.
Citation Information
Patent Citations
Multi-mechanical-arm cooperation method
CN113305845A
Compliant control method and device for on-orbit assembly of double-arm robot
CN117182929A