Robot control method, electronic device, readable storage medium and program product
By establishing robot and workpiece models in a simulation environment and training the robot control model based on actual pose information, the safety risks and generalization failures of the robot control model in the real environment are solved, and safe and efficient training and control are achieved.
Patent Information
- Application Number
- CN202511284849.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-09
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2045-09-09
AI Technical Summary
Existing robot control models have safety risks such as collisions and overloads during training, making it difficult to conduct high-frequency trial and error training. Traditional simulation environments cannot accurately reproduce the motion model errors of real robots, causing the trained models to fail to generalize when deployed to real-world scenarios.
Establish a simulation robot model and a simulation workpiece model corresponding to the robot, train the preset learning model based on pose information, collect the actual pose of the robot and determine the actual deviation information of the simulation model, input it to the robot control model to output predicted action commands, avoid collision and overload risks in real environment, and integrate real environment data to improve the model's generalization ability.
By training robot models in a virtual simulation environment, safety risks in real-world environments are avoided, the domain gap between virtual and real environments is bridged, and the generalization ability and training efficiency of the models are improved.
Smart Images

Figure CN120773066B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of robotics, and more specifically to a robot control method, electronic device, readable storage medium, and program product. Background Technology
[0002] Assembly robots are the core equipment of flexible automated assembly systems, consisting of a manipulator, controller, end effector, and sensing system. The manipulator can be categorized into horizontal articulated, Cartesian coordinate, multi-joint, and cylindrical coordinate types. The controller typically employs a multi-CPU or multi-level computer system. The end effector is designed with various grippers and wrists to accommodate different assembly workpieces.
[0003] When using robot control models to control robots, this method may pose risks such as collisions and overloads because the robot control models are trained using data from the real environment. Summary of the Invention
[0004] In view of the above problems, this application provides a robot control method, a robot control device, an electronic device, a computer-readable storage medium, and a computer program product.
[0005] According to the first aspect of this application, a method for training a robot control model is provided, comprising: establishing a simulated robot model corresponding to the robot and a simulated workpiece model corresponding to the workpiece to be operated by the robot; training a preset learning model based on the pose information of the robot, the simulated robot model, and the simulated workpiece model to obtain a trained robot control model, wherein the pose information represents the position and posture of the object; acquiring the actual robot pose of the robot, and determining the actual simulated pose of the simulated robot model and the actual deviation information between the actual simulated workpiece pose of the simulated workpiece model and a template based on the actual robot pose of the robot, wherein the template indicates the target pose of the simulated workpiece model; inputting the actual simulated pose and the actual deviation information into the robot control model and outputting predicted motion commands; and transmitting the predicted motion commands to the robot to control the robot.
[0006] A second aspect of this application provides a robot control device, comprising:
[0007] The module is used to create a simulation robot model corresponding to the robot and a simulation workpiece model corresponding to the workpiece that the robot will operate.
[0008] The training module is used to train a preset learning model based on the pose information of the robot, the simulated robot model, and the simulated workpiece model to obtain a trained robot control model. The pose information represents the position and posture of the object.
[0009] The determination module is used to collect the actual robot pose and, based on the actual robot pose, determine the actual simulation pose of the simulated robot model and the actual deviation information between the actual simulation workpiece pose and the template. The template indicates the target pose of the simulated workpiece model.
[0010] The prediction module is used to input the actual simulated pose and actual deviation information into the robot control model and output predicted action commands.
[0011] The control module is used to transmit predicted motion commands to the robot in order to control the robot.
[0012] A third aspect of this application provides an electronic device comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the method described above.
[0013] A fourth aspect of this application also provides a computer-readable storage medium having a computer program or instructions stored thereon, which, when executed by a processor, implement the steps of the above-described method.
[0014] The fifth aspect of this application also provides a computer program product, including a computer program or instructions that, when executed by a processor, implement the steps of the above-described method.
[0015] According to embodiments of this application, a simulated robot model and a simulated workpiece model corresponding to a robot in a real environment are established. Based on the pose information of the robot, the simulated robot model, and the simulated workpiece model, a preset learning model is trained to obtain a trained robot control model. The actual robot pose is collected, and the actual simulated pose of the simulated robot model and the actual deviation information between the actual simulated workpiece pose and the template are determined. The actual simulated pose and actual deviation information are input into the robot control model, and predicted action commands are output, thereby controlling the robot using the predicted action commands. Since the real robot is embedded in the virtual simulation environment for simulation, the safety risks of collisions and overloads in the real environment caused by using data generated by the real robot grasping and placing workpieces for model training are avoided. At the same time, robot data from the real environment is integrated during the training process, solving the domain gap problem between the virtual environment and the real environment, and improving the generalization ability of the model. Attached Figure Description
[0016] The above-mentioned contents, other objects, features and advantages of this application will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:
[0017] Figure 1 An application scenario diagram of the robot control method according to an embodiment of this application is shown.
[0018] Figure 2 A flowchart of a robot control method according to an embodiment of this application is shown.
[0019] Figure 3 A schematic diagram of a simulation environment according to an embodiment of this application is shown.
[0020] Figure 4 A flowchart of a method for training a robot control model according to an embodiment of this application is shown.
[0021] Figure 5 Schematic diagrams of different coordinate systems in a simulation environment according to embodiments of this application are shown.
[0022] Figure 6 A schematic diagram of a robot with marked points according to an embodiment of this application is shown.
[0023] Figure 7 A flowchart of a method for training a robot control model according to another embodiment of this application is shown.
[0024] Figure 8 A structural block diagram of a training device for a robot control model according to an embodiment of this application is shown.
[0025] Figure 9 A block diagram of an electronic device suitable for implementing the above-described method according to an embodiment of this application is shown. Detailed Implementation
[0026] The embodiments of this application will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of this application. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of this application for ease of explanation. However, it will be apparent that one or more embodiments may be implemented without these specific details. Furthermore, descriptions of well-known structures and technologies are omitted in the following description to avoid unnecessarily obscuring the concepts of this application.
[0027] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0028] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.
[0029] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).
[0030] The robot control models currently used in related technologies have the following technical problems during the training process:
[0031] (1) In the pure physical training scenario, safety risks such as collision and overload of the robotic arm make it impossible to carry out high-frequency trial and error training. Traditional solutions lack a hardware-in-the-loop safety buffer mechanism, making it difficult to complete strategy iteration under the premise of ensuring equipment safety.
[0032] (2) Reconstructing training conditions in a pure physical environment (such as different workpiece postures and environmental disturbances) requires manual resetting of hardware configuration, which is time-consuming, labor-intensive and inefficient, and cannot meet the needs of reinforcement learning for large-scale and diverse training data.
[0033] (3) In the traditional simulation environment training mode, the simulation environment is difficult to accurately reproduce the motion model error (such as joint encoder error, joint flexible deformation) and mechanical characteristics (such as friction torque) of the real robot, which leads to the problem of generalization failure of the training model when it is deployed to the physical scene due to the "domain gap".
[0034] In view of this, embodiments of this application provide a robot control method, electronic device, readable storage medium, and program product, which can be applied to the field of robot technology. The method includes establishing a simulated robot model corresponding to the robot and a simulated workpiece model corresponding to the workpiece to be operated by the robot; training a preset learning model based on the pose information of the robot, the simulated robot model, and the simulated workpiece model to obtain a trained robot control model; acquiring the actual robot pose of the robot, and determining the actual simulated pose of the simulated robot model and the actual deviation information between the actual simulated workpiece pose of the simulated workpiece model and the template based on the actual robot pose of the robot, wherein the template indicates the target pose of the simulated workpiece model; inputting the actual simulated pose and the actual deviation information into the robot control model and outputting predicted action commands; and transmitting the predicted action commands to the robot to control the robot.
[0035] In the technical solution of this application, the user information (including but not limited to user personal information, user image information, user device information, such as location information) and data (including but not limited to data used for analysis, stored data, and displayed data) involved are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of related data all comply with relevant laws, regulations, and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entry points for users to choose to authorize or refuse.
[0036] Figure 1 An application scenario diagram of the robot control method according to an embodiment of this application is shown.
[0037] like Figure 1 As shown, application scenario 100 according to this embodiment may include an electronic device assembly scenario. Communication unit 104 is used to provide a communication link between robot 101, positioning devices represented by multiple monocular cameras 102, and server 103. Communication unit 104 may include various connection types, such as wired or wireless communication links or fiber optic cables, etc.
[0038] Robot 101 can be any type of robot, such as an assembly robot that can install devices such as memory and hard drives in specified locations in electronic devices.
[0039] The positioning device can locate the robot 101. In one example, the positioning device uses at least two monocular cameras 102 to locate the robot 101. The monocular camera 102 can be any type of camera or video camera. Other positioning techniques are also applicable.
[0040] Server 103 can be a server that provides various services, can perform virtual simulation of a real environment, and can train a preset learning model to obtain a robot control model.
[0041] It should be noted that the robot control method provided in this application embodiment can generally be executed by the server 103. Accordingly, the robot control device provided in this application embodiment can generally be set in the server 103.
[0042] Figure 2 A flowchart of a robot control method according to an embodiment of this application is shown. Figure 3 A schematic diagram of a simulation environment according to an embodiment of this application is shown.
[0043] like Figure 2 As shown, the robot control method includes operations S201 to S205.
[0044] In operation S201, a simulation robot model corresponding to the robot and a simulation workpiece model corresponding to the workpiece to be operated by the robot are established.
[0045] In operation S202, based on the pose information of the robot, the simulated robot model, and the simulated workpiece model, a preset learning model is trained to obtain a trained robot control model, wherein the pose information represents the position and posture of the object.
[0046] In operation S203, the actual robot pose of the robot is acquired, and based on the actual robot pose, the actual simulation pose of the simulated robot model and the actual deviation information between the actual simulation pose of the simulated workpiece model and the template are determined. The template indicates the target pose of the simulated workpiece model.
[0047] When operating S204, the actual simulated pose and actual deviation information are input into the robot control model, and the predicted motion command is output.
[0048] In operation S205, the predicted motion commands are transmitted to the robot to control the robot.
[0049] A robot can refer to a workpiece in the field of mechanical manufacturing that installs parts into designated locations, such as robots used in the production of mobile phones, computers, and automobiles. A workpiece can refer to any component, such as an engine, memory, or hard drive. The robot holds the workpiece with its grippers, thereby installing it into the designated location. In this application, when referring to "robot" or "workpiece," it refers to the physical entity of the robot or workpiece in a real-world environment.
[0050] In robot control, a pre-set learning model needs to be trained first. During training, a simulation model of the robot can be established first. This involves constructing a simulated robot model and a simulated workpiece model corresponding to the workpiece the robot will operate within a simulation environment. Initially, the workpiece and template in the simulation environment are identical to those in the real environment. In this application, when referring to "simulated robot model" or "simulated workpiece model," it refers to a virtual model simulating the physical "robot" or physical "workpiece" in a simulation environment; it can also be considered a digital twin of the physical "robot" or physical "workpiece" in the simulation environment. Figure 3 As shown, in the simulation environment, in addition to the simulated robot model and the simulated workpiece model, other environmental models such as templates, measurement unit models, and workbench models can also be simulated. The template can indicate the target pose of the simulated workpiece model; that is, in the simulation environment, the simulated workpiece model should be assembled to a specified position in what posture. Although in Figure 3The diagram illustrates a simulation environment displayed via an electronic device's monitor. However, it's important to note that the simulation environment can be a digital space, and its existence is independent of whether it's displayed. In a real-world environment, considering that workpieces are typically expensive or easily damaged parts, real workpieces are not configured during training. Furthermore, real workbenches or similar equipment may not be required in the real-world environment, allowing the robot to be trained in a relatively open area. This avoids safety risks such as collisions and overloads, and reduces the cost of trial and error with the physical hardware.
[0051] During model training, the robot's pose information is collected and synchronized to the simulation environment to obtain the pose information of the simulated robot model and the simulated workpiece model. This pose information is then used to train a pre-defined learning model, resulting in a trained robot control model. In this application, "pose information" includes position and orientation. The robot's pose information or actual robot pose is collected, such as the robot's specific location and the joint angles of different joints in the robot arm. This collection is performed in a real environment. For example, the robot's specific position can be determined using a positioning device, and the information can be directly read from the robot's Application Programming Interface (API).
[0052] During the robot control process after training is completed, the actual robot pose can also be read through the robot's API interface. This actual robot pose can be synchronized to the simulation environment to determine the actual simulation pose of the simulated robot model and the actual deviation information between the actual simulation workpiece pose and the template based on the actual robot pose.
[0053] The actual simulated pose of the robot model and the actual deviation information between the simulated workpiece model and the template are input into the trained robot control model, and the predicted motion commands are output. The predicted motion commands are then transmitted to the robot to control the robot.
[0054] After the robot installs the workpiece at the designated location, the simulation environment and the robot can be initialized to prepare for the installation of the next workpiece.
[0055] According to embodiments of this application, a simulated robot model and a simulated workpiece model corresponding to a robot in a real environment are established. Based on the pose information of the robot, the simulated robot model, and the simulated workpiece model, a preset learning model is trained to obtain a trained robot control model. The actual robot pose is collected, and the actual simulated pose of the simulated robot model and the actual deviation information between the actual simulated workpiece pose and the template are determined. The actual simulated pose and actual deviation information are input into the robot control model, and predicted action commands are output, thereby controlling the robot using the predicted action commands. Since the real robot is embedded in the virtual simulation environment for simulation, the safety risks of collisions and overloads in the real environment caused by using data generated by the real robot grasping and placing workpieces for model training are avoided. At the same time, robot data from the real environment is integrated during the training process, solving the domain gap problem between the virtual environment and the real environment, and improving the generalization ability of the model.
[0056] Figure 4 A flowchart of a method for training a robot control model according to an embodiment of this application is shown.
[0057] According to embodiments of this application, pose information includes the actual training pose of the simulated robot model and the actual training pose of the simulated workpiece model, determined based on the robot's actual training pose.
[0058] like Figure 4 As shown, the training method for the robot control model in this embodiment may include operations S410 to S440.
[0059] During operation of S410, the preset learning model outputs action commands based on the input data set. The input data set includes the actual training pose of the simulated robot model and the actual training deviation between the actual training pose of the simulated workpiece model and the template.
[0060] When operating the S420, the robot is used to collect training poses after taking the actions indicated by the action commands.
[0061] During operation of S430, the motion training poses of the simulated robot model and the simulated workpiece model are determined based on the robot's motion training pose.
[0062] When operating S440, the preset learning model is updated based on the motion training deviation between the motion training pose of the simulated workpiece model and the template, thus obtaining the trained robot control model.
[0063] In a simulation environment, the actual training poses of the robot model and the simulated workpiece model can be determined based on the robot's actual training pose in the real environment. Furthermore, the actual training pose of the simulated workpiece model and the template can be used to determine the actual training deviation between their respective poses. The following section will describe in detail how to determine the poses of the simulated robot model and the simulated workpiece model in the simulation environment based on the robot's pose in the real environment, and how to determine the deviation between the simulated workpiece model's pose and the template.
[0064] The input data set, including the actual training pose of the simulated robot model and the actual training deviation between the actual training pose of the simulated workpiece model and the template, is input into a preset learning model. The preset learning model can output motion commands, which can instruct the joint movements of different joints on the robot arm and / or the opening and closing movements of the gripper at the end of the mechanical part. The preset learning model can be any type of neural network, including but not limited to continuous motion space reinforcement learning models, such as the Advantage Actor-Critic (A2C) model and the Deep Deterministic Policy Gradient (DDPG) model.
[0065] After the motion command is transmitted to the robot in the real environment through the communication unit, the robot can respond to the motion command and take the action indicated by the motion command. After the action is completed, the robot's motion execution result, i.e. motion training pose, can be collected. Thus, the motion training pose of the simulated workpiece model and the motion training deviation between the motion training pose of the simulated workpiece model and the template can be determined again in the simulation environment as described above. Based on the motion training deviation, the model parameters of the preset learning model can be adjusted to obtain the trained robot control model.
[0066] According to embodiments of this application, a simulated robot model and a simulated workpiece model corresponding to a robot in a real environment are established. The actual training poses of the simulated robot model and the simulated workpiece model are determined using the robot's actual training poses and input into a learning model to predict motion commands. After the robot executes the motion command, the motion training deviation between the motion training pose of the simulated workpiece model in the simulated environment and the template is determined based on the robot's motion training pose after the motion. The model is updated based on the motion training deviation to obtain a robot control model. By embedding a real robot into a virtual simulation environment to simulate the robot's operation on the workpiece, the safety risks of collisions and overloads in the real environment caused by using data generated from the real robot's operation of the workpiece for model training are avoided. Simultaneously, the integration of robot data from the real environment during training solves the domain gap problem between the virtual and real environments, improving the model's generalization ability.
[0067] Figure 5 Schematic diagrams of different coordinate systems in a simulation environment according to embodiments of this application are shown.
[0068] According to an embodiment of this application, establishing a simulated robot model corresponding to the robot and a simulated workpiece model corresponding to the workpiece to be operated by the robot includes: establishing overlapping environmental coordinate systems in the real environment and the simulated environment respectively; determining the coordinate system transformation matrix from the coordinate system of the positioning device in the real environment to the environmental coordinate system, wherein the positioning device is used to position the robot.
[0069] The positioning device can be a multi-camera array, such as at least two monocular cameras deployed in a non-diagonal position. In one specific embodiment, the multi-camera array can capture images of multiple marker points set in the robot's base, thereby completing the robot's positioning.
[0070] like Figure 5 As shown, in a simulation environment, different simulation models can have their own coordinate systems, such as... The base coordinate system of the simulated robot model has three degrees of freedom: Planar movement and rotation Rotation of the axis; The attitude coordinate system representing the simulated workpiece model in a gripped state has four degrees of freedom: along... Translation and rotation of the three axes Rotation of the axis; The template coordinate system, which is defined in the coordinate system of the measurement unit model (hereinafter referred to as the "measurement coordinate system"), has six degrees of freedom: translation along the three axes and rotation about the three axes. This represents a measurement coordinate system with six degrees of freedom: translation along the three axes and rotation about the three axes.
[0071] Before training, both the simulation and real environments can be initialized. In the real environment, the robot's initial pose, including the position of the moving chassis, can be set in the environmental coordinate system. and orientation And the joint angles of the robotic arm. Embodiments of this application are illustrated by way of a robot with seven joints, thereby giving the robot seven-axis joint angles. Similarly, the simulated robot model also has seven-axis joint angles, and the environmental coordinate system can be a coordinate system constructed with a certain position in the real environment as the origin.
[0072] In a simulation environment, the coordinate systems described above can be randomized, for example, by applying specific position errors and three-axis orientation angle errors. (This is used to define the measurement coordinate system.) For example, the observation position of the measurement model during system initialization. As shown in formula (1):
[0073] (1)
[0074] in, Let the three-dimensional rotation matrix generated by Euler angles satisfy: , , These represent rotations about the z-axis, y-axis, and x-axis, respectively. , ,× represents the Cartesian product, A hypercube region in three-dimensional space. This indicates the center position of the random field during initialization. , , These are the orientation angles along three axes, i.e., the three-axis orientation angles. The value can be set according to actual needs; for example, the minimum value can be -5° and the maximum value can be +5°. , Based on the hypercube region The position within is determined.
[0075] Similarly, the attitude uncertainty of other coordinate systems can be set in the same way as above, with missing degrees of freedom set to 0.
[0076] Define an environmental coordinate system that coincides with the environmental coordinate system of the simulated environment. A marker point can be placed at the origin of the real-world environmental coordinate system (which can be any location within the real-world environment) to establish the environmental coordinate system. Based on coordinate system calibration methods (such as Zhang Zhengyou's calibration method), the coordinate system transformation matrix from the positioning device to the environmental coordinate system is obtained. Among them, the coordinate system transformation matrix The conditions for satisfying formula (2) are:
[0077] (2)
[0078] in, , represents the homogeneous coordinates of the same point in the environmental coordinate system and the positioning device coordinate system, respectively, and i represents the serial number of any monocular camera in the positioning device.
[0079] According to embodiments of this application, by establishing overlapping environmental coordinate systems in the real and simulated environments and determining the coordinate transformation matrix from the positioning device coordinate system to the environmental coordinate system, it is convenient to map poses from the real environment to the virtual environment. This facilitates the use of robot operation data from the real environment in model training, improving the model's generalization ability. Furthermore, the simulation environment can simulate different postures and environmental noise conditions. The reconfigurability of the simulation environment overcomes the problems of time-consuming scene switching and low efficiency due to repetitive hardware configuration in purely realistic scenarios.
[0080] Figure 6 A schematic diagram of a robot with marked points according to an embodiment of this application is shown.
[0081] According to an embodiment of this application, the robot includes a base and a robotic arm disposed on the base.
[0082] According to an embodiment of this application, collecting the actual robot pose of the robot includes: locating the spatial orientation information of the base using a positioning device and obtaining the joint angle information of the robotic arm; and determining the actual robot pose of the robot based on the spatial orientation information and the joint angle information.
[0083] During model training and robot control, the robot's base can be positioned using a positioning device, thereby determining its spatial orientation. Specifically, at least three marker points can be set on the base, such as... Figure 6 Four marker points are set on the robot's base. By taking pictures of multiple marker points with the positioning device, the spatial orientation information of the base can be determined, which in turn determines the robot's spatial position.
[0084] Simultaneously, the joint angles of different joints on the robot's robotic arm can be directly read from the API interface in a real-world environment. Based on the robot's spatial orientation information, combined with multiple joint angle information, the robot's pose can be determined, such as the actual robot pose during robot control.
[0085] According to the embodiments of this application, since the spatial orientation information of the robot is determined by the positioning device and the joint angle information is directly read during the model training and control process, the real data is mapped to the simulation environment, thereby solving the domain gap problem between the simulation environment and the real environment and improving the generalization ability of the model.
[0086] Figure 7 A flowchart of a method for training a robot control model according to another embodiment of this application is shown.
[0087] Reference Figure 7 In operation S701, the simulation environment is first initialized. Then, in operation S702, the input data set is constructed. Next, in operation S703, the input data set is input into the preset learning model to predict the action command. Then, in operation S704, the robot executes the action command. In operation S705, the action training pose after the robot executes the action is synchronized to the simulation environment. Then, in operation S706, the reward function is calculated based on the action training deviation in the simulation environment. In operation S707, it is determined whether the termination condition is met. If it is met, the simulation environment is initialized again. If it is not met, the model is trained iteratively.
[0088] According to an embodiment of this application, determining the actual simulation pose of a simulated robot model and the actual simulation pose of a simulated workpiece model based on the actual robot pose includes: transforming the actual robot pose of the robot based on a coordinate system transformation matrix to obtain a transformed pose; synchronizing the transformed pose to the simulated robot model to obtain the actual simulation pose of the simulated robot model; and determining the actual simulation pose of the simulated workpiece model based on the actual simulation pose of the simulated robot model.
[0089] In a real environment, the actual robot pose can be determined by positioning devices and other means. Based on the coordinate system transformation matrix determined above, the actual robot pose can be transformed into the environmental coordinate system to obtain the transformed pose. This transformed pose is then synchronized to the simulation robot model to update the model's pose, thus obtaining the actual simulation pose of the simulation robot model.
[0090] Since the workpiece is simulated simultaneously in the simulation environment, the actual simulation pose of the simulated workpiece model can be directly determined from the simulation environment after the actual simulation pose of the simulated robot model is determined. Specifically, the actual simulation pose of the simulated workpiece model can be determined based on the actual simulation pose of the simulated robot model (e.g., the spatial position of the simulated robot model and the posture of the robotic arm) and the gripping relationship between the simulated robot model and the simulated workpiece model.
[0091] According to an embodiment of this application, a preset learning model outputs action commands based on an input data set, including: acquiring the opening and closing state information of the gripper in the robot and the action command previously output by the preset learning model, wherein the input data set further includes the opening and closing state information and the previously output action command; inputting the input data set into the preset learning model and outputting the action command.
[0092] Since the model aims to accurately install the workpiece in a designated position, the opening and closing state information of the robot's gripper also needs to be considered. Therefore, the input data set, combining the motion command output from the previous step of the pre-learning model, the actual training pose of the simulated robot model, and the actual training deviation between the actual training pose of the simulated workpiece model and the template, is fed into the pre-learning model for prediction, thus obtaining the motion command. The input data set can include a one-dimensional vector concatenated from multiple source state features. The motion command can indicate joint movements and / or gripper movements, such as joint increments and / or gripper changes.
[0093] It should be noted that during the first training or prediction in the control of the model, since the preset learning model did not have any output in the previous time, the action instruction that does not need to be executed can be used as the prediction action instruction of the previous output.
[0094] According to the embodiments of this application, since the opening and closing state information of the gripper, the previously output action command, the actual training pose of the simulated robot model, and the actual training deviation between the simulated workpiece model and the template are comprehensively considered during the training process of the preset learning model, the action command predicted by the preset learning model is more in line with the action in the real environment, which effectively improves the accuracy of the robot in installing the workpiece at the designated position, reduces the possibility of the workpiece colliding in the real environment, and thus improves the safety of the robot during training.
[0095] According to an embodiment of this application, the actual simulated pose and actual deviation information are input into the robot control model, and predicted motion commands are output. This includes: inputting the actual simulated pose and actual deviation information into the robot control model, and outputting the mean motion characteristics and standard deviation motion characteristics of different joints of the robot; for any joint, establishing a Gaussian distribution model based on the mean motion characteristics and standard deviation motion characteristics; and sampling the Gaussian distribution model to obtain the joint motion commands, wherein the predicted motion commands include joint motion commands corresponding to different joints.
[0096] When processing input data in the preset learning model, for multi-source state features such as actual simulated pose and actual deviation information, the multi-source state features can be concatenated into a one-dimensional vector, and this one-dimensional vector can be input into the preset learning model. The actual simulated pose of the simulated robot model can include the joint angles of multiple joints, and the input data can also include opening and closing state information, which represents the degree to which the gripper is open or closed.
[0097] It should be noted that the preset learning model can include multiple hidden layers, such as a hidden layer with three neurons having dimensions of 256, 128, and 64 respectively. Furthermore, to improve the stability during the training phase, a batch normalization layer can be set between any two hidden layers.
[0098] Based on the mean and standard deviation features of the motion output by the pre-defined learning model, a Gaussian distribution model of the robotic arm's motion at time i is established. As shown in formula (3):
[0099] (3)
[0100] Where j represents a joint. and These are the action mean feature and the action standard deviation feature, respectively.
[0101] By performing random sampling on the Gaussian distribution model, the joint motion commands can be obtained. These commands can then be sent to the robot in the real environment via the robot API interface or communication unit to execute the corresponding actions.
[0102] According to an embodiment of this application, locating the spatial orientation information of the robot's base using a positioning device includes: collecting the coordinates of at least three marker points on the robot's base in a real environment using the positioning device; and determining the spatial orientation information of the robot's base based on the camera coordinate system transformation matrix and the coordinates of the at least three marker points.
[0103] The positioning device can be a camera array composed of multiple monocular cameras. For example, any two of four monocular cameras can be combined to form a binocular system, and at least two cameras at non-diagonal positions can completely capture at least three marker points. The number of marker points can be set according to actual needs, but should not be less than three.
[0104] For each marker point, based on the coordinate information of multiple points collected by at least two positioning cameras (i.e., monocular cameras), the focal lengths of the multiple positioning cameras, and the camera coordinate system transformation matrix... Determine the marker coordinates of the marker point in the camera coordinate system. According to the marked coordinates The camera coordinate system transformation matrix is used to generate the target marker coordinates. Based on the target marker coordinates of multiple marker points, determine the location of the chassis in the spatial orientation information. and orientation .
[0105] In a specific embodiment, taking a total of 4 marker points as an example, if both camera 1 and camera 2 can capture 3 marker points, the coordinates of a marker point in each of the two cameras are... , By transforming the marker point into the coordinate system of camera 1, the marker coordinates can be obtained. Specifically, the calculation is performed using formula (4):
[0106] (4)
[0107] in and These represent the focal lengths of camera 1 and camera 2, respectively. , Camera coordinate system transformation matrix The elements satisfy formula (5):
[0108] (5)
[0109] Based on the coordinate system transformation matrix specified above The coordinates of the marker point in the environment coordinate system can be obtained. As shown in formula (6):
[0110] (6)
[0111] based on The marker layout template shown allows you to determine the chassis position based on the target marker coordinates of multiple marker points. and orientation .
[0112] According to embodiments of this application, by using multiple monocular cameras to photograph pre-marked points, the points are then transformed into the camera coordinate system. A camera coordinate system transformation matrix is then used to convert the points into the environment coordinate system. This allows for accurate determination of the robot's position based on the coordinates of multiple points in the environment coordinate system, enabling precise adjustment of the robot's posture in the simulation environment. This forms a closed-loop reinforcement learning training mechanism of "state observation - policy reasoning - action execution," retaining the convenience of the simulation environment while introducing the randomness of the real robot's state, thus improving the generalization ability of reinforcement learning.
[0113] According to an embodiment of this application, the actual deviation information is generated as follows: the actual robot pose of the robot is synchronized to the simulation robot model to obtain an updated simulation robot model; the actual simulation workpiece pose and the actual simulation template pose of the template in the environmental coordinate system are determined based on the updated simulation robot model; and the actual deviation information is calculated based on the actual simulation workpiece pose and the actual simulation template pose.
[0114] For actual deviation information (or motion training deviation), the corresponding robot pose (such as the robot's actual training pose or the robot's motion training pose) can be synchronized to the simulation environment, thereby obtaining an updated simulation robot model. From this, the actual simulated workpiece pose and the actual simulated template pose in the environmental coordinate system of the simulated workpiece model can be determined. Based on the actual simulated workpiece pose and the actual simulated template pose, the actual deviation information is calculated.
[0115] In one specific embodiment, the actual simulated workpiece pose of workpiece c in the environmental coordinate system of the simulation environment. The actual simulated template pose at template position t ,in, Indicates transpose. and middle Indicates spatial location, A quaternion represents a rotation in three-dimensional space. A quaternion is a mathematical tool used to represent rotations in three-dimensional space, consisting of a real part... and three imaginary parts composition.
[0116] Based on the actual simulated workpiece pose and actual simulation template pose Calculate the actual deviation information.
[0117] According to embodiments of this application, the actual deviation information includes target distance and orientation angle deviation.
[0118] According to an embodiment of this application, the actual deviation information is calculated based on the actual simulated workpiece pose and the actual simulated template pose, including: transforming the actual simulated workpiece pose and the actual simulated template pose based on a coordinate system transformation matrix to obtain workpiece pose data and template pose data in the measurement coordinate system; and calculating the target distance and orientation angle deviation based on the workpiece pose data and the template pose data.
[0119] The attitude of the coordinate system of the measurement unit model d in the environment coordinate system Based on the coordinate system transformation matrix, the actual simulated workpiece pose and the actual simulated template pose in the environmental coordinate system are transformed to the measurement coordinate system d, thereby obtaining the workpiece pose data. ,in, Including spatial location and attitude quaternions The workpiece pose data are obtained by calculating using formulas (7) and (8) respectively. :
[0120] (7)
[0121] (8)
[0122] Similarly, template pose data Spatial position and attitude quaternions The template pose data are obtained by calculating using formulas (9) and (10) respectively. :
[0123] (9)
[0124] (10)
[0125] Based on workpiece pose data and template pose data Calculate the target distance As shown in formula (11):
[0126] (11)
[0127] Finally, based on the workpiece pose data and template pose data, the orientation angle deviation is calculated. As shown in formula (12).
[0128] (12)
[0129] in, This represents the arctangent function in the four quadrants.
[0130] According to embodiments of this application, by converting the actual simulated workpiece pose and the actual simulated template pose in the environmental coordinate system into workpiece pose data and template pose data in the measurement coordinate system, and then calculating the target distance and orientation angle deviation based on the workpiece pose data and template pose data, the prediction of motion commands can be generated based on the target distance and orientation angle deviation. This can improve the prediction accuracy of the motion commands, ensure that the workpiece can be accurately installed at the designated position, and reduce the physical hardware trial and error costs caused by workpiece collisions.
[0131] According to an embodiment of this application, based on the motion training deviation between the motion training pose of the simulated workpiece model and the template, the preset learning model is updated to obtain a trained robot control model, including: calculating a reward function based on the target distance and orientation angle deviation in the motion training deviation and the multiple joint angle poses of the simulated robot model; updating the model parameters of the preset learning model based on the reward function to obtain the robot control model.
[0132] In each iteration of training, a reward function is calculated based on the target distance, orientation angle deviation, and multiple joint angle poses of the simulated robot model in the current iteration of action training deviation. Then, the parameters of the preset learning model are adjusted based on the reward function to obtain the robot control model.
[0133] According to an embodiment of this application, the reward function is calculated as follows: for any joint angle pose, the joint angle pose is differentiated to obtain a derivative value; multiple derivative values are summed to obtain a smoothing index, wherein the smoothing index characterizes the smoothness of the action when the robot executes the action command; and a reward function is generated based on the orientation angle deviation, target distance, smoothing index, and preset time penalty information of the action training deviation in the current iteration.
[0134] For any joint angle pose, the derivative of that pose can be calculated to obtain the corresponding derivative value. The derivative values of multiple joint angles can then be summed to obtain the smoothness index. As shown in formula (13):
[0135] (13)
[0136] Where n represents the number of joint angles, and when the number of joint angles is 7, n=7; This indicates the angle of the u-th joint. The derivative value obtained by performing gradient differentiation.
[0137] Based on the orientation angle deviation, target distance, smoothing index, and preset time penalty information of the action training bias in the current iteration, a reward function is generated. As shown in formula (14):
[0138] (14)
[0139] Among them, A and Indicates the weighting coefficient. This indicates the preset time penalty information, which can be set according to actual needs, for example... .
[0140] According to embodiments of this application, by integrating a multi-dimensional reward function that incorporates orientation angle deviation, smoothness index, and preset time penalty information for safety constraints, and adjusting model parameters based on the reward function, the resulting robot control model enables the robot to execute action commands with smoother movements in practical use. At the same time, using preset time penalty information can improve the training speed of the model and enhance its generalization ability under complex working conditions.
[0141] According to an embodiment of this application, when the simulated robot model collides with any simulated object model, the simulated workpiece model falls from the simulated robot model, or the execution step size (i.e., the number of iterations) exceeds a first preset threshold, the robot and the simulated robot model are initialized; when the reward function satisfies the reward threshold or the posture error between the simulated workpiece model and the template is less than a second preset threshold, the currently trained learning model is determined as the robot control model.
[0142] The first preset threshold, the reward threshold, and the second preset threshold can all be set according to actual needs. For example, the first preset threshold can be 500 times, the reward threshold can be any value in [-1,1], such as 0.8, and the second preset threshold can be an angle error of 0.1°, which can also include position deviation error.
[0143] During any iteration of training, if the simulated robot model collides with other simulated object models in the simulation environment, or the simulated workpiece model falls off the simulated robot model, or the execution step exceeds the first preset threshold, the robot and the simulated robot model can be initialized, such as resetting their poses, so that iterative training can be performed again. In this initialization operation, the model parameters of the learning model are not initialized.
[0144] If the number of iterations reaches the preset threshold, or the reward function in the current training satisfies the preset reward threshold, or the posture error between the simulated workpiece model and the template position in the current training is less than the preset angle threshold, then the current learning model can be determined as the trained robot control model.
[0145] Figure 8 A structural block diagram of a training device for a robot control model according to an embodiment of this application is shown.
[0146] like Figure 8 As shown, the robot control device 800 in this embodiment includes a setup module 810, a training module 820, a determination module 830, a prediction module 840, and a control module 850.
[0147] The module 810 is used to create a simulation robot model corresponding to the robot and a simulation workpiece model corresponding to the workpiece to be operated by the robot.
[0148] The training module 820 is used to train a preset learning model based on the pose information of the robot, the simulated robot model, and the simulated workpiece model to obtain a trained robot control model, wherein the pose information represents the position and posture of the object.
[0149] The determination module 830 is used to collect the actual robot pose of the robot, and based on the actual robot pose, to determine the actual simulation pose of the simulated robot model and the actual deviation information between the actual simulation workpiece pose of the simulated workpiece model and the template. The template indicates the target pose of the simulated workpiece model.
[0150] The prediction module 840 is used to input the actual simulated pose and actual deviation information into the robot control model and output predicted action commands.
[0151] The control module 850 is used to transmit predicted motion commands to the robot for robot control.
[0152] According to embodiments of this application, a simulated robot model and a simulated workpiece model corresponding to a robot in a real environment are established. Based on the pose information of the robot, the simulated robot model, and the simulated workpiece model, a preset learning model is trained to obtain a trained robot control model. The actual robot pose is collected, and the actual simulated pose of the simulated robot model and the actual deviation information between the actual simulated workpiece pose and the template are determined. The actual simulated pose and actual deviation information are input into the robot control model, and predicted action commands are output, thereby controlling the robot using the predicted action commands. Since the real robot is embedded in the virtual simulation environment for simulation, the safety risks of collisions and overloads in the real environment caused by using data generated by the real robot grasping and placing workpieces for model training are avoided. At the same time, robot data from the real environment is integrated during the training process, solving the domain gap problem between the virtual environment and the real environment, and improving the generalization ability of the model.
[0153] According to embodiments of this application, any multiple modules among the establishment module 810, training module 820, determination module 830, prediction module 840, and control module 850 can be combined into one module, or any one of these modules can be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules can be combined with at least part of the functionality of other modules and implemented in one module. According to embodiments of this application, at least one of the establishment module 810, training module 820, determination module 830, prediction module 840, and control module 850 can be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or implemented in hardware or firmware by any other reasonable means of integrating or packaging the circuitry, or implemented in software, hardware, or firmware, or in any suitable combination of any of these three implementation methods. Alternatively, at least one of the establishment module 810, training module 820, determination module 830, prediction module 840, and control module 850 can be at least partially implemented as a computer program module, which can perform corresponding functions when the computer program module is run.
[0154] Figure 9 A block diagram of an electronic device suitable for implementing the above-described method according to an embodiment of this application is shown.
[0155] like Figure 9 As shown, an electronic device 900 according to an embodiment of this application includes a processor 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage portion 908 into a random access memory (RAM) 903. The processor 901 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 901 may also include onboard memory for caching purposes. The processor 901 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of this application.
[0156] RAM 903 stores various programs and data required for the operation of electronic device 900. Processor 901, ROM 902, and RAM 903 are interconnected via bus 904. Processor 901 executes various operations of the method flow according to embodiments of this application by executing programs in ROM 902 and / or RAM 903. It should be noted that the programs may also be stored in one or more memories other than ROM 902 and RAM 903. Processor 901 may also execute various operations of the method flow according to embodiments of this application by executing programs stored in said one or more memories.
[0157] According to embodiments of this application, the electronic device 900 may further include an input / output (I / O) interface 905, which is also connected to a bus 904. The electronic device 900 may also include one or more of the following components connected to the input / output (I / O) interface 905: an input section 906 including a keyboard, mouse, etc.; an output section 907 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 908 including a hard disk, etc.; and a communication section 909 including a network interface card such as a LAN card, modem, etc. The communication section 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the input / output (I / O) interface 905 as needed. A removable medium 911, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 910 as needed so that computer programs read from it can be installed into the storage section 908 as needed.
[0158] This application also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the method according to the embodiments of this application.
[0159] According to embodiments of this application, the computer-readable storage medium can be a non-volatile computer-readable storage medium, such as including but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this application, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to embodiments of this application, the computer-readable storage medium may include ROM 902 and / or RAM 903 and / or one or more memories other than ROM 902 and RAM 903 described above.
[0160] Embodiments of this application also include a computer program product comprising a computer program containing program code for performing the methods shown in the flowchart. When the computer program product is run on a computer system, the program code is used to cause the computer system to implement the methods provided in the embodiments of this application.
[0161] When the computer program is executed by the processor 901, it performs the functions defined in the system / apparatus of this application embodiment. According to the embodiments of this application, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0162] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and downloaded and installed via the communication section 909, and / or installed from a removable medium 911. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.
[0163] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 909, and / or installed from the removable medium 911. When the computer program is executed by the processor 901, it performs the functions defined in the system of this application embodiment. According to the embodiments of this application, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0164] According to embodiments of this application, program code for executing the computer programs provided in the embodiments of this application can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C", or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0165] The embodiments of this application have been described above. However, these embodiments are merely illustrative and not intended to limit the scope of this application. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Without departing from the scope of this application, those skilled in the art can make various substitutions and modifications, all of which should fall within the scope of this application.
Claims
1. A robot control method, characterized in that, The control method includes: Establish a simulation robot model corresponding to the robot and a simulation workpiece model corresponding to the workpiece to be operated by the robot; Based on the pose information of the robot, the simulated robot model, and the simulated workpiece model, a preset learning model is trained to obtain a trained robot control model. The pose information represents the position and posture of the object, and includes the actual training pose of the simulated robot model and the actual training pose of the simulated workpiece model determined based on the actual training pose of the robot. The actual robot pose is acquired, and based on the actual robot pose, the actual simulation pose of the simulated robot model and the actual deviation information between the actual simulation pose of the simulated workpiece model and the template are determined. The template indicates the target pose of the simulated workpiece model. The actual simulated pose and the actual deviation information are input into the robot control model, and the predicted action command is output. The predicted action command is transmitted to the robot to control the robot; Specifically, based on the pose information of the robot, the simulated robot model, and the simulated workpiece model, a preset learning model is trained to obtain a trained robot control model, including: Based on the input data set, the preset learning model outputs action commands, wherein the input data set includes the actual training pose of the simulated robot model and the actual training deviation between the actual training pose of the simulated workpiece model and the template. Collect the robot's training pose after it performs the action indicated by the action command; Based on the robot's motion training pose, determine the motion training pose of the simulated robot model and the motion training pose of the simulated workpiece model. Based on the motion training deviation between the simulated workpiece model and the template, the preset learning model is updated to obtain a trained robot control model, including: For any joint angle pose among multiple joint angle poses of the simulated robot model, the joint angle pose is differentiated to obtain the derivative value; The smoothing index is obtained by summing multiple derivative values, where the smoothing index characterizes the smoothness of the robot's actions when executing motion commands. Based on the orientation angle deviation, target distance, smoothing index, and preset time penalty information of the action training bias in the current iteration, a reward function is generated. As shown in formula (1): (1) Among them, A and Indicates the weighting coefficient. This indicates the preset time penalty information. For smoothing exponent, For the orientation angle deviation, The target distance is n, where n represents the number of joint angles. This indicates the angle of the u-th joint. The derivative value obtained by performing gradient differentiation; The model parameters of the preset learning model are updated according to the reward function to obtain the robot control model.
2. The control method according to claim 1, characterized in that, Establishing a simulation robot model corresponding to the robot and a simulation workpiece model corresponding to the workpiece to be operated by the robot, including: Establish overlapping environmental coordinate systems in both the real and simulated environments; Determine the coordinate transformation matrix from the coordinate system of the positioning device in the real environment to the environmental coordinate system, wherein the positioning device is used to locate the robot.
3. The control method according to claim 1, characterized in that, The robot includes a base and a robotic arm mounted on the base. The actual robot pose of the data acquisition robot includes: The spatial orientation information of the base is located using a positioning device, and the joint angle information of the robotic arm is obtained. The actual robot pose is determined based on the spatial orientation information and the joint angle information.
4. The control method according to claim 2, characterized in that, Based on the actual robot pose of the robot, the actual simulation pose of the simulated robot model and the actual simulation workpiece pose of the simulated workpiece model are determined, including: The actual robot pose of the robot is transformed based on the coordinate system transformation matrix to obtain the transformed pose; The transformed pose is synchronized to the simulated robot model to obtain the actual simulated pose of the simulated robot model; The actual simulated workpiece pose of the simulated workpiece model is determined based on the actual simulated pose of the simulated robot model.
5. The control method according to claim 1, characterized in that, Based on the input data set, the pre-defined learning model outputs action instructions, including: The robot acquires the opening and closing state information of the gripper and the action command output by the preset learning model last time. The input data set further includes the opening and closing state information and the action command output last time. The input data set is input into the preset learning model, and the action command is output.
6. The control method according to claim 1, characterized in that, The actual simulated pose and the actual deviation information are input into the robot control model, and predicted action commands are output, including: The actual simulated pose and the actual deviation information are input into the robot control model, and the mean motion characteristics and standard deviation motion characteristics of different joints of the robot are output. For any of the joints, a Gaussian distribution model is established based on the mean and standard deviation characteristics of the movements. The Gaussian distribution model is sampled to obtain the joint motion commands of the joint, wherein the predicted motion commands include joint motion commands corresponding to different joints.
7. The control method according to claim 3, characterized in that, Locating the spatial orientation information of the base using the positioning device includes: In a real environment, the positioning device is used to collect the coordinates of at least three marker points on the base of the robot; Based on the camera coordinate system transformation matrix and the coordinates of the at least three marker points, the spatial orientation information of the robot's base is determined.
8. The control method according to claim 1, characterized in that, The actual deviation information is generated in the following way: The actual robot pose of the robot is synchronized to the simulation robot model to obtain an updated simulation robot model; Based on the updated simulation robot model, determine the actual simulation workpiece pose of the simulation workpiece model and the actual simulation template pose of the template in the environmental coordinate system; The actual deviation information is calculated based on the actual simulated workpiece pose and the actual simulated template pose.
9. The control method according to claim 8, characterized in that, The actual deviation information includes target distance and orientation angle deviation; The calculation of the actual deviation information based on the actual simulated workpiece pose and the actual simulated template pose includes: Based on the coordinate system transformation matrix, the actual simulated workpiece pose and the actual simulated template pose are transformed to obtain workpiece pose data and template pose data in the measurement coordinate system. Based on the workpiece pose data and template pose data, the target distance and the orientation angle deviation are calculated.
10. The control method according to claim 1, characterized in that, In response to a collision between the simulated robot model and any simulated object model, the simulated workpiece model falling from the simulated robot model, or the execution step length exceeding a first preset threshold, the robot and the simulated robot model are initialized. In response to the reward function satisfying the reward threshold or the posture error between the simulated workpiece model and the template being less than a second preset threshold, the currently trained preset learning model is determined as the robot control model.
11. An electronic device, comprising: One or more processors; Memory, used to store one or more computer programs. The characteristic feature is that the one or more processors execute the one or more computer programs to implement the steps of the control method according to any one of claims 1 to 10.
12. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, When the computer program or instructions are executed by the processor, they implement the steps of the control method according to any one of claims 1 to 10.
13. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the control method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Collaborative robot multi-shaft-hole assembling method and system, electronic equipment and medium
CN117067209A
Mechanical arm control method, device and equipment of composite robot and storage medium
CN120422256A