Robot control method, device, electronic device and computer program product
Patent Information
- Application Number
- CN202510464920.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2045-04-14
AI Technical Summary
因此,现有的模型训练方法需要花费大量的时间在确定有效动作策略上,从而导致所需要的训练时间较长,模型收敛速度较慢,进而使得模型训练的效率较低
[0048] In this embodiment, the electronic device can obtain correction parameters fed back by the researchers during model training and update the training model by combining the correction parameters with the expected actions. Specifically, during model training, when researchers determine that the robot's current actions deviate significantly, they can input correction parameters into the electronic device to prevent the robot from continuing to attempt failed combinations. Therefore, the electronic device can determine an effective action strategy based on the interaction state corresponding to the correction parameters. Thus, the method provided in this embodiment can reduce the time required for the electronic device to determine an effective action strategy, thereby reducing the model's training time and improving its training efficiency and convergence speed.
Smart Images

Figure CN120395816B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of robotics technology, and in particular to a robot control method, device, electronic device, and computer program product. Background Technology
[0002] A humanoid robot is a robot designed to mimic human appearance and behavior, possessing human form and function. Due to its anthropomorphic mechanical structure, humanoid robots typically have a certain degree of mobility and operational capabilities, enabling them to replace humans in performing dangerous, repetitive, and high-precision tasks. A humanoid robot mainly consists of a mechanical structure, a perception system, and a control system. The perception system primarily comprises various sensors, used to collect environmental information and the robot's own state information. The control system manipulates the mechanical structure based on the environmental and state information collected by the perception system, as well as the motion strategy model, to complete the operational tasks. To enable the humanoid robot to understand the goals and requirements of the operational tasks, researchers typically need to train the motion strategy model within the humanoid robot.
[0003] In existing model training methods, the state information collected by humanoid robots and the feedback action parameters are randomly combined. Therefore, humanoid robots typically need to try a large number of unsuccessful combinations before they can determine the action parameters that match a particular state information, i.e., determine the effective action policy corresponding to a given state information, and then train the model based on the effective action policy. As a result, existing model training methods require a significant amount of time to determine the effective action policy, leading to long training times, slow model convergence, and ultimately low model training efficiency. Summary of the Invention
[0004] In view of this, embodiments of this application provide a robot control method, device, electronic device, and computer program product to reduce model training time, thereby improving model convergence speed and training efficiency.
[0005] The first aspect of this application provides a robot control method, including:
[0006] In response to the first interaction state between the robot's actuators and the corresponding target object;
[0007] The initial action parameters corresponding to the first interaction state are determined by training the model, and the execution component is controlled to interact with the target object according to the initial action parameters;
[0008] If the interaction result obtained based on the initial action parameters does not meet the task completion conditions of the target task, then the correction parameters corresponding to the first interaction state are obtained.
[0009] The training model is updated based on the correction parameters, the initial action parameters, and at least one expected action parameter; the expected action parameter is the action parameter of the execution component during the process of manipulating the execution component to perform the target task on the target object.
[0010] If the updated training model does not meet the preset training completion conditions, then return to the operation of obtaining the first interaction state between the robot's execution component and the target object and its subsequent operations until the training model meets the training completion conditions. The training model that meets the training completion conditions will then be used as the action strategy model corresponding to the robot.
[0011] In one possible implementation of the first aspect, updating the training model based on the correction parameters, the initial action parameters, and at least one desired action parameter includes:
[0012] The first dataset is determined based on the initial action parameters and the correction parameters;
[0013] The second dataset is determined based on the expected action parameters and the correction parameters;
[0014] Based on a preset period, a number of first training data corresponding to a first proportional coefficient are obtained from the first dataset, and a number of second training data corresponding to a second proportional coefficient are obtained from the second dataset.
[0015] A first update operation is performed on the training model based on all the first training data and the second training data.
[0016] In one possible implementation of the first aspect, updating the trained model based on the correction parameters, the initial action parameters, and at least one desired action parameter further includes:
[0017] The initial action parameters are input into the reward function to generate the feedback parameters corresponding to the initial action parameters;
[0018] A second update operation is performed on the trained model based on the feedback parameters.
[0019] In one possible implementation of the first aspect, updating the training model based on the correction parameters, the initial action parameters, and at least one desired action parameter includes:
[0020] While performing parameter iteration operations, the training model is updated based on the correction parameters, the initial action parameters, and at least one desired action parameter; the parameter iteration operations include: performing an operation to obtain the first interaction state between the robot's execution component and the target object, and performing an operation to determine the initial action parameters corresponding to the first interaction state through the training model, and controlling the execution component to interact with the target object based on the initial action parameters.
[0021] In one possible implementation of the first aspect, the first interaction state includes a first posture corresponding to the target object, a second posture corresponding to the execution component, and tactile information of the execution component and the target object;
[0022] The first interaction state in response to the robot's actuator and the corresponding target includes:
[0023] During the process of controlling the robot's execution components to perform the target task, the tactile information between the execution components and the target object is determined based on the deformation information of the contact area on the execution components;
[0024] Obtain a partial image containing the execution component and an overall image containing all components of the robot;
[0025] The local image and the overall image are input into a preset image processing algorithm to obtain the first pose and the second pose;
[0026] Accordingly, determining the initial action parameters corresponding to the first interaction state through training the model includes:
[0027] The initial motion parameters are generated by processing the first pose, the second pose, and the tactile information through a trained model.
[0028] In one possible implementation of the first aspect, the tactile information includes the central axis of the actual contact area between the actuating component and the target object; the deformation information of the area to be contacted includes the displacement vectors of each contact point in the area to be contacted in a three-dimensional spatial coordinate system.
[0029] The determination of the tactile information between the actuator and the target object based on the deformation information of the contact area on the actuator includes:
[0030] For any point to be contacted in the area to be contacted, if there is a displacement vector in any coordinate system in the displacement vector of any point to be contacted that is greater than or equal to the displacement threshold corresponding to any coordinate system, then the point to be contacted is determined as the actual contact point.
[0031] The actual contact area between the actuating component and the target object is determined based on all the actual contact points.
[0032] The central axis of the actual contact area is determined based on the shape of the actual contact area.
[0033] In one possible implementation of the first aspect, controlling the robot to perform actions according to the action strategy model includes:
[0034] During the process of controlling the robot's execution components to perform the target task, at any given control moment, the second interaction state between the robot's execution components and the target object is acquired;
[0035] The second interaction state is input into the action strategy model to generate the target action parameters corresponding to the second interaction state;
[0036] The execution component is controlled to interact with the target object according to the target motion parameters.
[0037] A second aspect of this application provides a robot control device, including:
[0038] The state acquisition module is used to respond to the first interaction state between the robot's execution parts and the corresponding target object;
[0039] An interaction module is used to determine the initial action parameters corresponding to the first interaction state through a trained model, and to control the execution component to interact with the target object according to the initial action parameters;
[0040] The parameter acquisition module is used to acquire the correction parameters corresponding to the first interaction state if the interaction result obtained based on the initial action parameters does not meet the task completion conditions of the target task.
[0041] The update module is used to update the training model based on the correction parameters, the initial action parameters, and at least one expected action parameter; the expected action parameter is the action parameter of the execution component during the process of manipulating the execution component to perform the target task on the target object;
[0042] The model determination module is used to return to the operation of obtaining the first interaction state between the robot's execution component and the target object and subsequent operations if the updated training model does not meet the preset training completion conditions, until the training model meets the training completion conditions, and the training model that meets the training completion conditions is used as the action strategy model corresponding to the robot.
[0043] The control module is used to control the robot to perform actions according to the action strategy model.
[0044] A third aspect of this application provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the robot control method described in the first aspect above.
[0045] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the robot control method described in the first aspect above.
[0046] A fifth aspect of this application provides a computer program product that, when run on a computer, causes the computer to execute the robot control method described in the first aspect.
[0047] Compared with the prior art, the embodiments of this application have the following advantages:
[0048] In this embodiment, the electronic device can obtain correction parameters fed back by the researchers during model training and update the training model by combining the correction parameters with the expected actions. Specifically, during model training, when researchers determine that the robot's current actions deviate significantly, they can input correction parameters into the electronic device to prevent the robot from continuing to attempt failed combinations. Therefore, the electronic device can determine an effective action strategy based on the interaction state corresponding to the correction parameters. Thus, the method provided in this embodiment can reduce the time required for the electronic device to determine an effective action strategy, thereby reducing the model's training time and improving its training efficiency and convergence speed. Attached Figure Description
[0049] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0050] Figure 1 This is a schematic diagram of a robot model training method provided in an embodiment of this application;
[0051] Figure 2 This is a schematic diagram of a model training process provided in an embodiment of this application;
[0052] Figure 3 This is a schematic diagram of another robot model training method provided in an embodiment of this application;
[0053] Figure 4This is a schematic diagram of another robot model training method provided in an embodiment of this application;
[0054] Figure 5 This is a schematic diagram of a robot control method provided in an embodiment of this application;
[0055] Figure 6 This is a schematic diagram of another robot control method provided in an embodiment of this application;
[0056] Figure 7 This is a schematic diagram of another robot control method provided in an embodiment of this application;
[0057] Figure 8 This is a schematic diagram of another robot control method provided in an embodiment of this application;
[0058] Figure 9 This is a schematic diagram of a central axis provided in an embodiment of this application;
[0059] Figure 10 This is a schematic diagram of a robot control device provided in an embodiment of this application;
[0060] Figure 11 This is a schematic diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0061] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0062] With the rapid development of robotics technology, humanoid robots, with their unique advantage of being able to mimic human form and behavior, are finding increasingly widespread applications in numerous fields. From precision assembly tasks in industrial production to service assistance in everyday life scenarios, such as family companionship and logistics delivery, humanoid robots have demonstrated enormous potential and value.
[0063] In existing humanoid robot control technologies, vision-based control schemes are quite common. Typically, a camera module is installed in the robot's head to capture state images of the target object and its actuators. The initial intention of this approach was to leverage the intuitiveness and richness of visual information to provide comprehensive data support for robot control. By using the state images captured by the camera module as input data to the control module, the module can generate motion parameters for controlling the actuators based on the relative position and posture of the target object and actuators presented in the images, using a motion strategy model. This achieves precise control of the robot's movements. Therefore, the accuracy of the motion strategy model directly affects the accuracy of the robot's movements. Consequently, how to establish highly accurate motion strategy models has always been a key research focus for researchers.
[0064] In existing technologies, researchers can construct motion strategy models based on mechanical principles, such as mass-spring models, finite element models, and pull-sliding dynamics models. However, motion strategy models constructed based on mechanical principles not only require significant computational time in practical applications, but also necessitate researchers to construct corresponding motion strategy models for each target object based on its size, material, and other information, further consuming considerable modeling time. Therefore, establishing motion strategy models based on neural network models is currently a major development direction in robot control technology. However, researchers need to spend a significant amount of time training models when building motion strategy models based on neural network models. In view of this, embodiments of this application provide a method for training motion strategy models, offering a method for training motion strategy models with shorter training times.
[0065] The technical solution of this application will be described below through specific embodiments.
[0066] This application provides a method for managing motion strategy models. Specifically, the motion strategy model is the model used when a robot interacts with a target object.
[0067] The management method in this application includes two stages: a model training stage for generating an applicable action strategy model, and a model application stage for applying the action strategy model.
[0068] Phase 1: Model Training Phase
[0069] Reference Figure 1This diagram illustrates a robot model training method according to an embodiment of this application. The method can be applied to the robot's control module, which can be an electronic device such as a microcontroller unit (MCU), microprocessor unit (MPU), digital signal processor (DSP), field-programmable gate array (FPGA), or system-on-chip (SoC). The robot model training method described above may specifically include the following steps:
[0070] S101, the first interaction state in response to the robot's actuator and the corresponding target object.
[0071] In this embodiment, the target task can be a task that researchers need the robot to perform on a target object, such as a cable manipulation task. The execution component can be an action on the robot used to perform the target task with the target object. For example, when the target task is a cable manipulation task, the target object can be a cable, and the execution component can be the robot's robotic arm; when the target task is a walking task, the target object can be the ground, and the execution component can be the robot's robotic leg. The interaction state can include at least the tactile information of the robot and the target object, the first posture of the target object, and the second posture of the execution component. The specific method for the electronic device to obtain the interaction state is the same as the method for obtaining the interaction state in the model application stage; please refer to the content of embodiments four to six of this application.
[0072] Before controlling a robot to perform a target task, researchers need to train a training model to obtain the corresponding action strategy model. At this stage, researchers can manipulate the robot's actuators to perform the target task (e.g., by controlling the robot to interact with the target object via a teleoperation device to complete the task), and manipulate the robot's desired action parameters and the corresponding interaction states during the process. This allows for model training based on the collected desired action parameters. After collecting the desired action parameters, researchers can issue a model training command to the robot. The electronic device can respond to the user-initiated model training command and train the training model to determine the action strategy model. During model training, the electronic device can continuously collect the interaction states between the actuators and the target object, and input these interaction states into the training model to generate initial action parameters, enabling the electronic device to train the model based on these real-time generated initial action parameters.
[0073] S102. Determine the initial action parameters corresponding to the first interaction state through the training model, and control the execution component to interact with the target object according to the initial action parameters.
[0074] In this embodiment, after acquiring the first interaction state, the electronic device can input the first interaction state into a training model to generate initial action parameters corresponding to the first interaction state. These initial action parameters can be parameters used to control the execution component to perform the corresponding action. Then, the electronic device can control the execution component to interact with the target object based on the initial action parameters.
[0075] For example, when the target task is to grip a cable, the electronic device can control the robotic arm according to the initial motion parameters so that the robotic arm can achieve the purpose of gripping the cable.
[0076] S103. If the interaction result obtained based on the initial action parameters does not meet the task completion conditions, then obtain the correction parameters corresponding to the first interaction state.
[0077] In this embodiment, the interaction result can be the result of interacting with the target object by executing the action corresponding to the initial action parameters through the execution component. The correction parameter can be the expected action parameter corresponding to the current tactile information input by the developer when the interaction result between the execution component and the target object does not meet the task completion conditions. Continuing with the above example, for instance, the initial action parameters are used to achieve the purpose of inserting the cable end into a specific interface, but the actual interaction result is that the cable end is not inserted into the specific interface, but there is still a certain distance between the cable end and the interface. In this case, the interaction result is identified as not meeting the task completion conditions.
[0078] After controlling the interaction between the actuator and the target object based on the initial motion parameters, the electronic device can determine whether the interaction result between the actuator and the target object meets the task completion conditions. If the electronic device determines that the interaction result between the actuator and the target object does not meet the task completion conditions, the electronic device can obtain the correction parameters input by the R&D personnel.
[0079] Specifically, if the electronic device determines that the interaction between the execution component and the target object does not meet the task completion conditions, the electronic device can obtain the current first posture of the execution component and the corresponding second posture of the target object, and determine whether the current interaction between the execution component and the target object meets the correction conditions based on the first and second postures. For details on how the electronic device obtains the first and second postures, please refer to the content of the fifth embodiment of this application. For example, when the target task is a cable manipulation task, if the electronic device determines, based on the first and second postures, that the cable is about to detach from the gripper's fingers, and / or, based on the first posture, that the current movement amplitude of the robotic arm is too large, the electronic device can determine that the execution component currently meets the correction conditions. If the electronic device determines that the first and / or second postures meet the correction conditions, the electronic device can generate correction prompt information and display the correction prompt information through a preset display device. The correction prompt information is used to prompt the R&D personnel to input correction parameters. The electronic device can then continuously receive correction parameters input by the R&D personnel. If the electronic device determines that neither the first nor the second posture meets the correction conditions, the electronic device may not generate correction prompt information.
[0080] If the electronic device determines that the interaction result between the execution component and the target object meets the task completion conditions, the electronic device may not acquire correction parameters. At this time, the electronic device can further determine whether the training model meets the training completion conditions. If the electronic device determines that the training model meets the training completion conditions, it can stop training the training model and designate the current training model as the action policy model. If the electronic device determines that the training model does not meet the training completion conditions, it can control the execution component to return to the initial state and re-execute S101 to S105 until the training model meets the training completion conditions. The method by which the electronic device determines whether the training model meets the training completion conditions is the same as in S105 of this embodiment; please refer to the content of S105 of this embodiment for understanding.
[0081] In one possible implementation, after the electronic device controls the actuator to interact with the target object based on initial motion parameters, it can acquire an overall image containing all parts of the robot via a camera mounted on the robot's head, and / or acquire local images containing the actuator and the target object via a camera mounted on the actuator. The electronic device can then input the acquired overall and local images into a discriminative model, which generates the interaction result between the actuator and the target object. The discriminative model can be a binary classification model, such as a decision tree model or a random forest model. The discriminative model can be trained by researchers using the overall and local images acquired during offline demonstrations.
[0082] Specifically, during training, when the discriminant model determines that the task is completed based on the overall and local images, it can generate a first interactive result to indicate task completion. When the discriminant model determines that the task is not completed based on the overall and local images, it cannot generate an interactive result. When the discriminant model determines that the robot has engaged in dangerous actions during training, such as the joint angular velocity of the actuator exceeding a threshold or the end effector experiencing excessive force, it can generate a second interactive result to indicate task failure and control the robot to stop moving. Therefore, when the discriminant model does not output an interactive result, the electronic device can determine that the current interactive result obtained based on the initial motion parameters does not meet the task completion conditions, and the electronic device can control the actuator to continue moving. When the discriminant model outputs an interactive result, the electronic device can determine that the current interactive result obtained based on the initial motion parameters meets the task completion conditions, and the electronic device can control the actuator to stop moving.
[0083] S104. Update the trained model based on the correction parameters, initial action parameters, and at least one desired action parameter.
[0084] In this embodiment, after acquiring the correction parameters, the electronic device can update the trained model based on the correction parameters, initial action parameters, and at least one expected action parameter. The expected action parameter refers to the action parameters of the execution component during the process of the developer manipulating the execution component to perform a target task on a target object before model training. For example, when the motion strategy model is a model for performing a cable manipulation task, the target task can be a cable manipulation task, and the expected action parameter can be the action parameters of the robotic arm during the process of the developer manipulating the robotic arm to perform cable manipulation to insert the cable's insertion end into a specific interface before model training.
[0085] In one possible implementation, see Figure 2 This illustration shows a schematic diagram of a model training process provided in an embodiment of this application. Figure 2 As shown, the electronic device may include a first process and a second process. The first process may be used to acquire initial action parameters and correction parameters; that is, the first process may be used to execute steps S101 to S103. The second process may be used to perform update operations on the trained model; that is, the second process may be used to execute steps S104 to S105.
[0086] Specifically, in response to a user-initiated model training command, the electronic device can invoke a first process. This first process acquires the initial interaction state between the robot's actuators and the target object, and inputs this state into the training model to determine the initial action parameters corresponding to that state. Then, the first process controls the actuators to interact with the target object based on these initial action parameters. After the interaction is complete, the first process can continue to acquire the initial interaction state between the actuators and the target object, and continue to determine the initial action parameters based on this state and the training model. This process continuously generates the initial action parameters corresponding to the training model based on the real-time interaction state.
[0087] During this process, the first process can also continuously receive correction parameters input by the R&D personnel. Figure 2 The slashed portion in the text can indicate the moment when the first process acquires the correction parameters. For example... Figure 2 As shown, the frequency with which researchers input correction parameters gradually decreases as the training model is updated. After each generation of initial action parameters and acquisition of correction parameters, the first process can transmit these parameters to the second process, enabling the second process to update the training model based on them.
[0088] During the process of the first process controlling the robot's execution components to perform the aforementioned operations, the electronic device can periodically call the second process according to a preset cycle set by the developers. This allows the training model to be updated via the second process during the execution of the first process. Specifically, during the execution of the first process—that is, while the first process acquires the first interaction state between the robot's execution components and the target object, and / or while the first process processes the first interaction state using the training model, generates initial motion parameters, and controls the execution components to interact with the target object based on the initial motion parameters—if the preset cycle is met, the electronic device can update the training model via the second process based on correction parameters, initial motion parameters, and at least one desired motion parameter. After updating the training model, the electronic device can transmit the updated training model to the first process, enabling the first process to generate initial motion parameters based on the updated training model.
[0089] In existing technologies, data acquisition and model updating are typically performed sequentially. Therefore, when training a model using existing technologies, the electronic device must wait for data acquisition while the model is updating, and vice versa. However, in this embodiment, the electronic device can execute data acquisition and model updating in parallel through two processes. Thus, the method provided in this embodiment allows the electronic device to continuously collect new data in the first process while the second process promptly uses the acquired data to update the model. Therefore, the method provided in this embodiment can effectively reduce model training time and significantly improve model training efficiency.
[0090] S105. If the updated training model does not meet the preset training completion conditions, then return to the operation of obtaining the first interaction state between the robot's execution parts and the target object and its subsequent operations until the training model meets the training completion conditions. The training model that meets the training completion conditions will be used as the robot's corresponding action strategy model.
[0091] In this embodiment, after updating the training model, the electronic device can determine whether the updated training model meets the training completion conditions. If the electronic device determines that the training model does not meet the training completion conditions, the electronic device can return to execute S101-S105 to continue training the training model. If the electronic device determines that the training model does not meet the training completion conditions, the electronic device can stop training the training model and use the currently qualified training model as the robot's corresponding action strategy model.
[0092] In one possible implementation, the electronic device can update the update count of the training model after each update operation and determine whether the training model meets the training completion condition based on the update count. Specifically, when the update count of the training model is greater than or equal to a preset threshold, the electronic device can determine that the current training model meets the training completion condition. When the update count of the training model is less than the threshold, the electronic device can determine that the current training model does not meet the training completion condition.
[0093] In this embodiment, after generating the motion strategy model, the electronic device can respond to user-initiated control commands and control the execution component to perform interactive actions with the target object. Specifically, during the process of controlling the robot's execution component, the electronic device can continuously acquire the second interaction state between the robot's execution component and the target object. The specific method for acquiring the second interaction state is the same as the method for acquiring the first interaction state; please refer to the content in the fourth embodiment of this application. After acquiring the second interaction state, the electronic device can input the second interaction state into the trained motion strategy model to generate target motion parameters corresponding to the second interaction state, and control the execution component to interact with the target object based on the target motion parameters. The electronic device can continuously collect the second interaction state between the execution component and the target object and continuously generate target motion parameters until the execution component completes the target task with the target object.
[0094] The method provided in this embodiment allows the electronic device to obtain correction parameters fed back by researchers based on the current interaction state during model training, and update the training model according to the correction parameters and expected action parameters. Therefore, the method provided in this embodiment enables the electronic device to directly determine the effective action strategy corresponding to the interaction state based on the correction parameters and expected action parameters, thereby reducing the robot's trial time, reducing the model training time, and improving the model's convergence speed and training efficiency.
[0095] Figure 3 A flowchart illustrating the specific implementation of a robot model training method S104 according to the second embodiment of this application is shown. See also... Figure 3 Compared to Figure 1 In the embodiment described above, S104 of the robot model training method provided in this embodiment includes: S301 to S304, which are detailed below:
[0096] S301. Determine the first dataset based on the initial motion parameters and correction parameters.
[0097] In this embodiment, as Figure 2 As shown, after obtaining the correction parameters, the first process in the electronic device can transmit the current initial action parameters and correction parameters to the second process. After receiving the initial action parameters and correction parameters, the second process can write the received initial action parameters and correction parameters into the first data buffer of the second process to determine the first dataset based on the initial action parameters and correction parameters.
[0098] S302. Determine the second dataset based on the expected motion parameters and correction parameters.
[0099] In this embodiment, as Figure 2As shown, after receiving the initial action parameters and correction parameters, the second process in the electronic device can also write the received correction parameters and expected action parameters into the second data buffer of the second process to determine the second dataset based on the expected action parameters and correction parameters.
[0100] S303. Based on a preset period, obtain several first training data corresponding to the first proportional coefficient from the first dataset, and obtain several second training data corresponding to the second proportional coefficient from the second dataset.
[0101] In this embodiment, the second process in the electronic device can randomly obtain a plurality of first training data corresponding to a first scaling factor from the first dataset, and randomly obtain a plurality of second training data corresponding to a second scaling factor from the second dataset. Specifically, the second process can sample the first dataset and the second dataset equally to obtain the same amount of first training data and second training data. At the same time, the first process in the electronic device can continue to execute the operations of S101 to S103 in the first embodiment of this application.
[0102] S304. Perform a first update operation on the trained model based on all the first training data and the second training data.
[0103] In this embodiment, as Figure 2 As shown, after the second process in the electronic device obtains the first training data and the second training data, it can calculate the error value based on all the first training data and the second training data, and perform a first update operation on the training model based on the currently determined error value to update the network parameters in the training model.
[0104] The method provided in this embodiment allows the electronic device to determine the first and second datasets based on the initial and expected action parameters, respectively, in conjunction with correction parameters. This enables the electronic device to fully utilize data from different sources to update the trained model. Furthermore, since the correction parameters are data input by researchers during training, they can compensate for any potential errors or deficiencies. In summary, the method provided in this embodiment provides a higher-quality data foundation for model training, thereby improving the accuracy of the trained action policy model.
[0105] Figure 4 A flowchart illustrating the specific implementation of step S104 in a robot model training method according to the third embodiment of this application is shown. See also... Figure 4 Compared to Figure 1 In the embodiment described above, S104 of the robot model training method provided in this embodiment includes: S401 to S402, which are detailed below:
[0106] S401. Input the initial action parameters into the reward function to generate the feedback parameters corresponding to the initial action parameters.
[0107] In this embodiment, after obtaining the initial action parameters, the second process in the electronic device can input the initial action parameters into a pre-set reward function to generate feedback parameters corresponding to the initial action parameters. The reward function can determine the feedback parameters corresponding to the initial action parameters based on the degree of deviation between the input initial action parameters and the desired action parameters. For example, when the target task is a cable operation task, the reward function can generate positive feedback parameters to provide a positive reward when the deviation between the initial action parameters and the desired action parameters is less than a deviation threshold, or when the deviation is gradually decreasing; conversely, the reward function can generate negative feedback parameters to provide a negative reward when the deviation between the initial action parameters and the desired action parameters is greater than or equal to the deviation threshold, or when the deviation is gradually increasing.
[0108] S402. Perform a second update operation on the trained model based on the feedback parameters.
[0109] In this embodiment, after the second process in the electronic device obtains the feedback parameters corresponding to the initial action parameters, it can perform a second update operation on the training model according to the feedback parameters to update the network parameters in the training model.
[0110] The method provided in this embodiment can further improve the training efficiency of the model because the second process can perform a second update operation on the training model by combining the feedback parameters obtained from the reward function while performing the first update operation on the training model.
[0111] Phase 2: Model Application Phase
[0112] Reference Figure 5 This diagram illustrates a robot control method according to a fourth embodiment of this application. Specifically, it may include the following steps:
[0113] S501, the first interaction state in response to the robot's actuator and the corresponding target object.
[0114] S502. Determine the initial action parameters corresponding to the first interaction state through the training model, and control the execution component to interact with the target object according to the initial action parameters.
[0115] S503. If the interaction result obtained based on the initial action parameters does not meet the task completion conditions of the target task, then obtain the correction parameters corresponding to the first interaction state.
[0116] S504. Update the trained model based on the correction parameters, initial action parameters, and at least one desired action parameter.
[0117] S505. If the updated training model does not meet the preset training completion conditions, then return to the operation of obtaining the first interaction state between the robot's execution parts and the target object and its subsequent operations until the training model meets the training completion conditions. The training model that meets the training completion conditions will be used as the robot's corresponding action strategy model.
[0118] The specific implementation methods of S501 to S505 in this embodiment are the same as those of S101 to S105 in the first embodiment of this application. Please refer to the content of S101 to S105 in the first embodiment for understanding, and they will not be repeated here.
[0119] S506. Control the robot to perform actions according to the motion strategy model.
[0120] In this embodiment, after determining the motion strategy model, the electronic device can control the robot to perform actions according to the trained motion strategy model. The specific method for controlling the robot to perform actions according to the motion strategy model can be found in the model application phase of this application, and will not be repeated here.
[0121] Figure 6 A flowchart illustrating the specific implementation of step S506 in a robot control method according to the fifth embodiment of this application is shown. See also... Figure 6 Compared to Figure 5 In the embodiment provided, the robot control method S506 includes: S5061 to S5065, which are detailed below:
[0122] S5061. During the process of controlling the robot's execution component to perform the target task, for any control moment, based on the deformation information of the area to be contacted on the execution component, determine the tactile information between the execution component and the target object.
[0123] In this embodiment, the area to be contacted can be a region on the execution component that can contact the target object. For example, when the target task is a hand operation task, the execution component can be the robot's robotic arm, and the area to be contacted can be the fingertip area of the robot's finger; when the target task is a walking task, the execution component can be the robot's robotic leg, and the area to be contacted can be the robot's foot. The interaction state between the execution component and the target object can include a first posture corresponding to the target object, a second posture corresponding to the execution component, and tactile information between the execution component and the target object.
[0124] When a user needs the robot to interact with a target object to complete a target task, the user can send control commands corresponding to the target task to the robot. After receiving the control commands from the user, the robot can determine the execution component and motion strategy model corresponding to the target task based on the control commands. Then, the electronic device can acquire the second interaction state between the target object and the execution component. The second interaction state can include the first posture of the target object, the second posture of the execution component, and tactile information between the execution component and the target object. Specifically, the electronic device can continuously acquire deformation information corresponding to the area to be contacted through tactile sensors installed on the execution component. Then, the electronic device can determine the tactile information between the execution component and the target object based on the currently acquired deformation information. For example, when the target task is a cable handling task, the tactile information acquired by the electronic device based on the deformation information can include the central axis, friction force, and contact state between the execution component and the target object; when the target task initiated by the user is a task other than cable handling, the tactile information acquired by the electronic device can include the friction force and contact state between the execution component and the target object. The specific method by which the electronic device determines tactile information based on deformation information is the same as the method in embodiments five to seven of this application. Please refer to the content in embodiments five to seven of this application for understanding.
[0125] In one possible implementation, the surface of the contact area of the actuating component can be covered with an elastic material. Therefore, when the contact area comes into contact with the target object, the elastic material on the surface of the contact area can deform. Specifically, the deformation information acquired by the electronic device may include, but is not limited to, a deformation image of the elastic material on the surface of the contact area and the displacement vectors of each contact point in the contact area in a three-dimensional spatial coordinate system. For example, a camera can be installed under the elastic material of the contact area, and the electronic device can acquire the deformation image of the elastic material through the camera under the elastic material. After acquiring the deformation image, the electronic device can reconstruct a three-dimensional shape model of the elastic material based on the deformation image, and determine the displacement vectors of each contact point in the contact area in a three-dimensional spatial coordinate system based on the three-dimensional model.
[0126] For example, multiple tiny markers can be embedded inside the elastic material on the surface of the area to be contacted. When the target object comes into contact with the area to be contacted, the elastic material will deform, and the markers inside the elastic material will also move with the deformation of the elastic body. After the electronic device acquires the deformation image, it can determine the position information of each marker in the deformation image through image processing algorithms, and then determine the displacement vector of each contact point in the area to be contacted in the three-dimensional spatial coordinate system.
[0127] S5062. Obtain a partial image containing the actuator and an overall image containing all parts of the robot.
[0128] In this embodiment, the electronic device can also acquire partial images of the execution component through a camera mounted on the execution component, and acquire an overall image containing all components of the robot through a camera mounted on the robot's head.
[0129] For example, when the target task is a cable handling task, the electronic device can acquire a partial image through a camera mounted on the wrist of the robot's robotic arm, and obtain an overall image based on a camera mounted on the top of the robot's head.
[0130] S5063. Input the local image and the overall image into a preset image processing algorithm to obtain the first pose and the second pose.
[0131] In this embodiment, after acquiring the local and overall images, the electronic device can first perform image preprocessing on the acquired local and overall images. For example, the electronic device can first crop the local and overall images, adjusting them to the same size to focus the image's emphasis on the region containing the execution component and the target object. Then, the electronic device can input the preprocessed local and overall images into a preset image processing algorithm to obtain a first pose corresponding to the execution component and a second pose corresponding to the target object. It should be noted that the image processing algorithm in this embodiment can be any image processing algorithm known to those skilled in the art, such as a Residual Network (ResNet), and this embodiment is not intended to specifically limit the image processing algorithm.
[0132] S5064. Input the first posture, the second posture, and tactile information into the action strategy model corresponding to the target task to generate target action parameters.
[0133] In this embodiment, after obtaining the first posture, the second posture, and tactile information, the electronic device can input these information into the motion strategy model corresponding to the target task to generate target motion parameters. For example, when the target task is a cable handling task, the electronic device can input the first posture, the second posture, and tactile information into the motion strategy model trained based on the expected motion parameters corresponding to the cable handling task.
[0134] S5065. Control the interaction between the execution component and the target object according to the target motion parameters.
[0135] In this embodiment, after determining the target motion parameters, the electronic device can control the execution component to interact with the target object based on the target motion parameters. Specifically, when the execution component is a robot's robotic arm, the electronic device can control the posture of the robotic arm's hands and the bending angle of the joints of each finger on the robotic arm based on the target motion parameters, and control the robotic arm to move in a certain direction with that posture.
[0136] The method provided in this embodiment enables the robot to accurately perceive the actual contact state between the actuator and the target object based on real-time deformation information of the area to be contacted, since the electronic device can determine the tactile information of the actuator relative to the target object based on the deformation information of the area to be contacted. Furthermore, since the tactile information is determined directly by the electronic device based on the deformation information of the area to be contacted, the tactile information determined in this embodiment can accurately reflect the actual contact between the actuator and the target object in real time, without relying on visual imaging. Therefore, using accurate tactile information as input data for the motion strategy model allows the first motion parameters generated by the motion strategy model to better match the actual relationship between the actuator and the target object, thereby improving control accuracy.
[0137] Figure 7 A flowchart illustrating a specific implementation of a robot control method S5061 according to the sixth embodiment of this application is shown. See also... Figure 7 Compared to Figure 6 In the embodiment provided, the robot control method S5061 includes: S701 to S703, which are detailed below:
[0138] S701. Based on the deformation information of the area to be contacted, determine the actual contact area between the actuator and the target object from the area to be contacted.
[0139] In this embodiment, the actual contact area can be the region in the area to be contacted where the target object and the execution component interact. The point to be contacted can be a point on the area to be contacted that can make contact with the target object. The deformation information acquired by the electronic device can include the displacement vectors of each point to be contacted in the three-dimensional spatial coordinate system within the area to be contacted.
[0140] When the user-initiated target task is any task other than cable handling, the tactile information acquired by the electronic device can include the frictional force and contact state between the actuator and the target object. After acquiring the deformation information of the area to be contacted, the electronic device can determine the actual contact area between the actuator and the target object based on the deformation information.
[0141] Specifically, the robot's contact area may include multiple contact points. After acquiring the displacement vectors of each contact point, the electronic device can determine whether any displacement vector in any coordinate system is greater than or equal to a displacement threshold for that coordinate system. The displacement thresholds for different coordinate systems can be different, and each coordinate system's threshold can be preset by the developers. For any contact point in the contact area, if any displacement vector in a coordinate system is greater than or equal to the corresponding displacement threshold, the electronic device can designate that contact point as an actual contact point. If all displacement vectors in the contact point's displacement vector are less than the displacement threshold of their corresponding coordinate system, the electronic device may not designate that contact point as an actual contact point. After determining all actual contact points in the contact area, the electronic device can define the area formed by all actual contact points as the actual contact area between the actuator and the target object.
[0142] S702. Determine the frictional force between the actuator and the target object based on the actual contact area.
[0143] In this embodiment, after determining the actual contact area, the electronic device can determine the frictional force between the actuator and the target object based on the actual contact area.
[0144] In one possible implementation, after determining the actual contact area, the electronic device can acquire historical deformation information corresponding to each actual contact point within that area. This historical deformation information can be data collected by the electronic device during the process of controlling the robot's actuators to perform the current target task, with the acquisition time preceding the control time corresponding to the current deformation information. After determining the historical deformation information corresponding to each actual contact point, for any given actual contact point, the electronic device can determine the deformation change information corresponding to that actual contact point based on both the historical deformation information and the current deformation information. After determining the deformation change information corresponding to all actual contact points, the electronic device can calculate the frictional force between the actuator and the target object based on the deformation change information corresponding to all actual contact points.
[0145] In one possible implementation, the electronic device can acquire multiple historical deformation images with acquisition times earlier than the current control time, based on a preset statistical duration. Then, the electronic device can input all historical deformation images and the current deformation image into a preset matching algorithm to determine the historical displacement vector of each actual contact point within the statistical duration, and determine the displacement change value of the center point based on the historical displacement vector and the current displacement vector. To improve matching accuracy, the matching algorithm can be configured with filtering algorithms, motion constraints, and optimized objective functions to ensure the smoothness of displacement changes at each actual contact point during matching. After determining the displacement change value, the electronic device can calculate the frictional force between the cable and the fingertip area based on the mean of all displacement change values, the statistical duration, and a preset friction coefficient.
[0146] S703. The area of the actual contact area determines the contact state between the actuator and the target object.
[0147] In this embodiment, the contact state between the execution component and the target object can include a first contact state and a second contact state. The first contact state can be a state in which there is effective contact between the execution component and the target object. The second contact state can be a state in which there is no effective contact between the execution component and the target object, or a state in which there is no contact between the execution component and the target object.
[0148] After determining the actual contact area, the electronic device can calculate the area of the actual contact area and determine the contact state between the actuator and the target object based on the area. Specifically, after determining the area of the actual contact area, the electronic device can query a database based on the object identifier corresponding to the target object to obtain the area threshold corresponding to the target object. Then, the electronic device can determine whether the area of the current actual contact area is greater than the area threshold corresponding to the target object. If the electronic device determines that the area of the actual contact area is greater than or equal to the area threshold corresponding to the target object, the electronic device can determine the contact state between the actuator and the target object as the first contact state. If the electronic device determines that the area of the actual contact area is less than the area threshold corresponding to the target object, the electronic device can determine the contact state between the actuator and the target object as the second contact state.
[0149] After determining the friction force and contact state corresponding to the deformation information, the electronic device can input the friction force and contact state into a training model to generate initial motion parameters for the actuator. For example, when the friction force between the actuator and the target object is less than a preset friction force threshold, and the contact state between the actuator and the target object is a second contact state, the motion parameters generated by the electronic device can be parameters that control the actuator to increase the friction force and actual contact area with the target object. For instance, the motion parameters generated by the electronic device can include larger force parameters and bending angles.
[0150] When a robot performs interactive actions with a target object, the target objects of different materials have different surface characteristics and mechanical properties. Friction and contact states can provide a wealth of detailed feedback to the robot's control system, helping the robot to distinguish between target objects of different materials. This, in turn, enables the robot to generate more accurate motion parameters that better match the characteristics of the target object. Therefore, the method provided in this embodiment can improve the accuracy of motion parameters generated by the electronic device.
[0151] Figure 8 A flowchart illustrating a specific implementation of a robot control method S5061 according to the sixth embodiment of this application is shown. See also... Figure 8 Compared to Figure 6 In the embodiment provided, the robot control method S5061 includes: S801 to S804, which are detailed below:
[0152] S801. Based on the deformation information of the area to be contacted, determine the actual contact area between the robotic arm and the cable from the area to be contacted.
[0153] In this embodiment, when the target task initiated by the user is a cable operation task, the target object can be a cable-type target object; the execution component can be the robotic arm of a robot; the area to be contacted can be the fingertip area of the robot's finger; the tactile information obtained by the electronic device can include the central axis, friction force, and contact state between the execution component and the target object.
[0154] During the cable handling task, the electronic device can acquire deformation information of the fingertip area of each target finger through sensors installed on the fingertip area of the robot's fingers, and determine the actual contact area between the fingertip area of each target finger and the cable based on the deformation information. The target fingers can be the fingers used by the robot to grip the cable. For example, the target fingers can be the robot's thumb and index finger. The specific method by which the electronic device determines the contact state is the same as in S701 of the fifth embodiment of this application. Readers can refer to S701 of the sixth embodiment of this application for understanding, and it will not be repeated here.
[0155] S802. Determine the central axis of the actual contact area between the robotic arm and the target object based on the shape of the actual contact area.
[0156] In this embodiment, the central axis of the actual contact area can be used to represent the current clamping posture of the cable in the finger.
[0157] In cable manipulation tasks, cables are often made of flexible materials, lack a fixed shape, and are typically long and thin. Robotic fingers can grip the cables in various ways, so the electronic device needs to determine the next direction of the robotic arm's movement based on the cable's current gripping posture within the fingers. After determining the actual contact area, the electronic device can determine the central axis of that area based on its shape. Specifically, after acquiring the actual contact area, the electronic device can input it into a principal component analysis (PCA) algorithm to determine the corresponding central axis. The PCA algorithm can determine the central axis based on the shape of the actual contact area.
[0158] See Figure 9 This diagram illustrates a central axis provided in an embodiment of this application. Figure 9 In (a), the area filled with diagonal lines can represent the actual contact area between the cable and the fingertip of a finger. Specifically, the shape of this actual contact area can be elliptical. Figure 9 In (b), a schematic structure of a finger and a cable is shown. The figure shows a finger-shaped component with a cable diagonally passing through the finger. The dotted line on the finger can represent the central axis of the cable as determined by the actual contact area.
[0159] S803. Determine the friction force between the robotic arm and the cable based on the center point on the central axis.
[0160] In this embodiment, after determining the central axis corresponding to the actual contact area, the electronic device can determine the frictional force between the robotic arm and the cable based on the center point on the central axis. The specific method by which the electronic device calculates the frictional force based on the center point is the same as the method for calculating frictional force in the sixth embodiment S702 of this application. Readers can refer to the content of the sixth embodiment S702 of this application and replace "actual contact point" with "center point" to understand the specific method for calculating frictional force in this embodiment.
[0161] S804. The area of the actual contact area determines the contact state between the robotic arm and the cable.
[0162] In this embodiment, after determining the actual contact area, the electronic device can further determine the contact state between the execution component and the target object based on the area of the actual contact area. The specific method by which the electronic device determines the contact state is the same as that in S703 of the sixth embodiment of this application; please refer to the specific method in S703 of the sixth embodiment of this application, which will not be repeated here. Specifically, when the robot performs a cable manipulation task, the first contact state can indicate that there is effective contact between the robotic arm's fingers and the cable, meaning that the robotic arm's fingers are currently gripping the cable with good quality; the second contact state can indicate that there is no effective contact between the robotic arm's fingers and the cable, meaning that the robotic arm's fingers are currently gripping the cable with poor quality. For example, when the electronic device determines that the contact state between the robotic arm and the cable is the second contact state, the initial motion parameters generated by the electronic device can include the joint bending angles corresponding to each joint in the thumb and index finger, to control the robotic arm to increase the actual contact area with the cable.
[0163] The method provided in this embodiment allows the electronic device to obtain the frictional force, contact state, and central axis between the cable and the robotic arm based on deformation information during cable handling tasks. The central axis provides a clear positional reference for the electronic device, enabling it to understand the spatial location and direction of the cable and thus accurately plan the robotic arm's movement path. Considering frictional force and contact state allows the electronic device to adjust the robotic arm's gripping force and method based on the cable's surface characteristics (e.g., smooth or rough) and the actual contact conditions. For example, when friction is low, the robotic arm automatically increases the gripping force; when the contact state is uneven, the bending angles of the joints in the index finger and thumb can be adjusted to ensure the cable is stably gripped and prevent slippage or wobbling. Therefore, the method provided in this embodiment can improve the accuracy and stability of the robot when performing cable handling.
[0164] It should be noted that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0165] Reference Figure 10 The diagram illustrates a robot control device according to an embodiment of this application, which may include a state acquisition module 1001, an interaction module 1002, a parameter acquisition module 1003, an update module 1004, a model determination module 1005, and a control module 1006, wherein:
[0166] The state acquisition module 1001 is used to respond to the first interaction state between the robot's execution component and the corresponding target object;
[0167] The interaction module 1002 is used to determine the initial action parameters corresponding to the first interaction state through a training model, and to control the execution component to interact with the target object according to the initial action parameters;
[0168] The parameter acquisition module 1003 is used to acquire the correction parameters corresponding to the first interaction state if the interaction result obtained based on the initial action parameters does not meet the task completion conditions of the target task.
[0169] Update module 1004 is used to update the training model based on the correction parameters, the initial action parameters, and at least one expected action parameter; the expected action parameter is the action parameter of the execution component during the process of manipulating the execution component to perform the target task on the target object;
[0170] The model determination module 1005 is used to return to the operation of obtaining the first interaction state between the robot's execution component and the target object and its subsequent operations if the updated training model does not meet the preset training completion conditions, until the training model meets the training completion conditions, and the training model that meets the training completion conditions is used as the action strategy model corresponding to the robot.
[0171] The control module 1006 is used to control the robot to perform actions according to the action strategy model.
[0172] The update module 1004 can also be used to determine a first dataset based on initial action parameters and correction parameters; determine a second dataset based on expected action parameters and correction parameters; obtain several first training data corresponding to a first proportional coefficient from the first dataset based on a preset period, and obtain several second training data corresponding to a second proportional coefficient from the second dataset; and perform a first update operation on the training model based on all the first and second training data.
[0173] The update module 1004 can also be used to input the initial action parameters into the reward function to generate the feedback parameters corresponding to the initial action parameters; and to perform a second update operation on the trained model based on the feedback parameters.
[0174] The update module 1004 can also be used to update the training model based on the correction parameters, initial action parameters, and at least one expected action parameter while performing parameter iteration operations. The parameter iteration operations include: performing the operation of obtaining the first interaction state between the robot's execution component and the target object, and performing the operation of processing the first interaction state through the training model to generate initial action parameters, and controlling the execution component to interact with the target object based on the initial action parameters.
[0175] The state acquisition module 1001 can also be used to determine the tactile information of the execution component and the target object based on the deformation information of the contact area on the execution component during the process of controlling the execution component of the robot to perform the target task; acquire a local image containing the execution component and an overall image containing all components of the robot; input the local image and the overall image into a preset image processing algorithm to obtain the first pose and the second pose;
[0176] The interaction module 1002 can also be used to process the first pose, the second pose, and tactile information through the training model to generate initial motion parameters.
[0177] The status acquisition module 1001 can also be used to determine any contact point in the contact area as an actual contact point if there is a displacement vector in any coordinate system that is greater than or equal to the displacement threshold corresponding to any coordinate system. The actual contact area between the execution component and the target object is determined based on all actual contact points. The central axis of the actual contact area is determined based on the shape of the actual contact area.
[0178] The control device 1006 can also be used to, during the process of controlling the robot's execution component to perform the target task, acquire the second interaction state between the robot's execution component and the target object at any control moment; input the second interaction state into the action strategy model to generate the target action parameters corresponding to the second interaction state; and control the execution component to interact with the target object according to the target action parameters.
[0179] As the apparatus embodiments are basically similar to the method embodiments, they are described in a relatively simple manner. For relevant details, please refer to the description in the method embodiment section.
[0180] Reference Figure 11 The diagram illustrates an electronic device according to an embodiment of this application. Figure 11 As shown, the electronic device 1100 in this embodiment includes: a processor 1110, a memory 1120, and a computer program 1121 stored in the memory 1120 and executable on the processor 1110. When the processor 1110 executes the computer program 1121, it implements the steps of the various embodiments of the robot control method described above, for example... Figure 5 The steps S501 to S506 are shown. Alternatively, when the processor 1110 executes the computer program 1121, it implements the functions of each module / unit in the above-described device embodiments, for example... Figure 9 The functions of modules 901 to 906 are shown.
[0181] For example, the computer program 1121 can be divided into one or more modules / units, which are stored in the memory 1120 and executed by the processor 1110 to complete this application. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, which can be used to describe the execution process of the computer program 1121 in the electronic device 1100. For example, the computer program 1121 can be divided into a status acquisition module, an interaction module, a parameter acquisition module, an update module, a model determination module, and a control module, with the specific functions of each module as follows:
[0182] The state acquisition module is used to respond to the first interaction state between the robot's execution parts and the corresponding target object;
[0183] An interaction module is used to determine the initial action parameters corresponding to the first interaction state through a trained model, and to control the execution component to interact with the target object according to the initial action parameters;
[0184] The parameter acquisition module is used to acquire the correction parameters corresponding to the first interaction state if the interaction result obtained based on the initial action parameters does not meet the task completion conditions of the target task.
[0185] The update module is used to update the training model based on the correction parameters, the initial action parameters, and at least one expected action parameter; the expected action parameter is the action parameter of the execution component during the process of manipulating the execution component to perform the target task on the target object;
[0186] The model determination module is used to return to the operation of obtaining the first interaction state between the robot's execution component and the target object and subsequent operations if the updated training model does not meet the preset training completion conditions, until the training model meets the training completion conditions, and the training model that meets the training completion conditions is used as the action strategy model corresponding to the robot.
[0187] The control module is used to control the robot to perform actions according to the action strategy model.
[0188] The electronic device 1100 may be a desktop computer, a cloud server, or other computing device. The electronic device 1100 may include, but is not limited to, a processor 1110 and a memory 1120. Those skilled in the art will understand that... Figure 11This is merely one example of electronic device 1100 and does not constitute a limitation on electronic device 1100. It may include more or fewer components than shown, or combine certain components, or different components. For example, electronic device 1100 may also include input / output devices, network access devices, buses, etc.
[0189] The processor 1110 can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0190] The memory 1120 can be an internal storage unit of the electronic device 1100, such as a hard disk or memory of the electronic device 1100. The memory 1120 can also be an external storage device of the electronic device 1100, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc., equipped on the electronic device 1100. Furthermore, the memory 1120 can include both internal and external storage units of the electronic device 1100. The memory 1120 is used to store the computer program 1121 and other programs and data required by the electronic device 1100. The memory 1120 can also be used to temporarily store data that has been output or will be output.
[0191] This application also discloses an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the robot model training method and the robot control method as described in the foregoing embodiments.
[0192] This application also discloses a computer-readable storage medium storing a computer program that, when executed by a processor, implements the robot model training method and robot control method as described in the foregoing embodiments.
[0193] This application also discloses a computer program product that, when run on a computer, causes the computer to execute the robot model training method and robot control method described in the foregoing embodiments.
[0194] The embodiments described above are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A method for controlling a robot, characterized in that, include: In response to a first interaction state between the robot's actuator and a corresponding target object; wherein the first interaction state includes a first posture corresponding to the target object, a second posture corresponding to the actuator, and tactile information of the actuator and the target object; The initial action parameters corresponding to the first interaction state are determined by training the model, and the execution component is controlled to interact with the target object according to the initial action parameters; If the interaction result obtained based on the initial action parameters does not meet the task completion conditions of the target task, then the correction parameters corresponding to the first interaction state are obtained. The training model is updated based on the correction parameters, the initial action parameters, and at least one expected action parameter; the expected action parameter is the action parameter of the execution component during the process of manipulating the execution component to perform the target task on the target object. If the updated training model does not meet the preset training completion conditions, then return to the operation of the first interaction state of the robot's execution component and the corresponding target object and its subsequent operations until the training model meets the training completion conditions. The training model that meets the training completion conditions will be used as the action strategy model corresponding to the robot. The robot is controlled to perform actions according to the action strategy model; The step of updating the training model based on the correction parameters, the initial action parameters, and at least one expected action parameter includes: determining a first dataset based on the initial action parameters and the correction parameters; determining a second dataset based on the expected action parameters and the correction parameters; obtaining several first training data corresponding to a first proportional coefficient from the first dataset and several second training data corresponding to a second proportional coefficient from the second dataset based on a preset period; performing a first update operation on the training model based on all the first training data and the second training data; and inputting the initial action parameters into a reward function to generate feedback parameters corresponding to the initial action parameters, and performing a second update operation on the training model based on the feedback parameters. The first interaction state in response to the robot's execution component and the corresponding target object includes: during the process of controlling the robot's execution component to perform the target task, determining the tactile information of the execution component and the target object based on the deformation information of the area to be contacted on the execution component; acquiring a local image containing the execution component and an overall image containing all components of the robot; inputting the local image and the overall image into a preset image processing algorithm to obtain the first pose and the second pose.
2. The method according to claim 1, characterized in that, The step of updating the trained model based on the correction parameters, the initial action parameters, and at least one desired action parameter includes: While performing parameter iteration operations, the training model is updated based on the correction parameters, the initial action parameters, and at least one desired action parameter; the parameter iteration operations include: performing an operation to obtain the first interaction state between the robot's execution component and the target object, and performing an operation to determine the initial action parameters corresponding to the first interaction state through the training model, and controlling the execution component to interact with the target object based on the initial action parameters.
3. The method according to any one of claims 1-2, characterized in that, The step of determining the initial action parameters corresponding to the first interaction state by training a model includes: The initial motion parameters are generated by processing the first pose, the second pose, and the tactile information through a trained model.
4. The method according to claim 3, characterized in that, The tactile information includes the central axis of the actual contact area between the actuator and the target object; the deformation information of the area to be contacted includes the displacement vector of each point to be contacted in the three-dimensional spatial coordinate system. The determination of the tactile information between the actuator and the target object based on the deformation information of the contact area on the actuator includes: For any point to be contacted in the area to be contacted, if there is a displacement vector in any coordinate system in the displacement vector of any point to be contacted that is greater than or equal to the displacement threshold corresponding to any coordinate system, then the point to be contacted is determined as the actual contact point. The actual contact area between the actuating component and the target object is determined based on all the actual contact points. The central axis of the actual contact area is determined based on the shape of the actual contact area.
5. The method according to any one of claims 1-2, characterized in that, The step of controlling the robot to perform actions according to the action strategy model includes: During the process of controlling the robot's execution components to perform the target task, at any given control moment, the second interaction state between the robot's execution components and the target object is acquired; The second interaction state is input into the action strategy model to generate the target action parameters corresponding to the second interaction state; The execution component is controlled to interact with the target object according to the target motion parameters.
6. A robot model training device, characterized in that, include: A state acquisition module is used to respond to a first interaction state between the robot's execution component and the corresponding target object; wherein, the first interaction state includes a first posture corresponding to the target object, a second posture corresponding to the execution component, and tactile information of the execution component and the target object; An interaction module is used to determine the initial action parameters corresponding to the first interaction state through a trained model, and to control the execution component to interact with the target object according to the initial action parameters; The parameter acquisition module is used to acquire the correction parameters corresponding to the first interaction state if the interaction result obtained based on the initial action parameters does not meet the task completion conditions of the target task. The update module is used to update the training model based on the correction parameters, the initial action parameters, and at least one expected action parameter; the expected action parameter is the action parameter of the execution component during the process of manipulating the execution component to perform the target task on the target object; The model determination module is used to return to the operation of executing the first interaction state of the robot's execution component and the corresponding target object and its subsequent operations if the updated training model does not meet the preset training completion conditions, until the training model meets the training completion conditions, and the training model that meets the training completion conditions is used as the action strategy model corresponding to the robot. The control module is used to control the robot to perform actions according to the action strategy model; The update module is further configured to: determine a first dataset based on the initial action parameters and the correction parameters; determine a second dataset based on the expected action parameters and the correction parameters; obtain several first training data corresponding to a first proportional coefficient from the first dataset and several second training data corresponding to a second proportional coefficient from the second dataset based on a preset period; perform a first update operation on the training model based on all the first training data and the second training data; and input the initial action parameters into a reward function to generate feedback parameters corresponding to the initial action parameters, and perform a second update operation on the training model based on the feedback parameters. The state acquisition module is further configured to: determine the tactile information of the execution component and the target object based on the deformation information of the contact area on the execution component during the process of controlling the execution component of the robot to perform the target task; acquire a local image containing the execution component and an overall image containing all components of the robot; input the local image and the overall image into a preset image processing algorithm to obtain the first posture and the second posture.
7. An electronic device, characterized in that, The device includes a processor, a memory, and a computer program stored in the memory and executable on the processor, characterized in that, when the processor executes the computer program, the electronic device implements the model training method for the robot as described in any one of claims 1-5.
8. A computer program product, characterized in that, Includes a computer program, which, when run, causes the model training method for the robot as described in any one of claims 1-5 to be executed.
Citation Information
Patent Citations
Mechanical arm motion track planning method and system, storage medium and electronic equipment
CN115070764A
Robot motion strategy model optimization method and related device
CN119292077A