Object control method and device, electronic equipment and storage medium

By deploying a policy model trained on simulation data onto real objects, the problem of policy transfer caused by the differences between simulation and real robots is solved, and the goal of completing the task is achieved efficiently and safely.

CN121742292APending Publication Date: 2026-03-27CHINA AUTOMOTIVE INNOVATION CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-17
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

In existing technologies, there are differences between simulated robots and real robots, such as dynamic model errors, sensor noise, and actuator response delays. This causes strategies that perform well on simulated robots to perform poorly or even fail on real robots. Furthermore, training and adjusting directly on real robots is costly and risky.

Method used

By acquiring the current state data of the real object in the target simulation environment, the target policy model is trained and deployed in the real object. The simulation and real control parameters are updated using preset demonstration data, so that the simulated object and the real object are highly consistent at the behavioral level, and correction signals are generated to control the real object to complete the target task.

Benefits of technology

It effectively narrows the gap between simulation and real robots, improves the success rate of strategy transfer, reduces the risk and cost of adjusting real objects, and enhances the adaptability, stability and task completion efficiency of the control system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121742292A_ABST
    Figure CN121742292A_ABST
Patent Text Reader

Abstract

The invention relates to an object control method and apparatus, an electronic device and a storage medium. The method comprises the steps of obtaining first current state data of a simulation object corresponding to a real object in a target simulation environment corresponding to a target task; based on the first current state data and a target task instruction corresponding to the target task, training an initial strategy model in the simulation object to obtain a target strategy model corresponding to the target task; the target strategy model is deployed in a real object, so that the target strategy model generates a real control instruction for the real object based on the target task instruction, and a control parameter of a real controller corresponding to the real object is a target real control parameter; and the real control instruction is sent to a real controller, so that the real controller generates a correction signal based on the target real control parameter and the real control instruction to control a real object, and the real object completes the target task. According to the embodiment of the invention, the real object can be controlled to reliably complete the target task.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of robot control, and particularly relates to an object control method and device, electronic equipment and storage medium. BACKGROUND

[0002] With the rapid development of artificial intelligence and robot technology, reinforcement learning has become an important technical means in the field of robot motion control. It can autonomously learn the optimal motion control strategy through trial and error of an intelligent agent in an environment. However, directly adjusting reinforcement learning on a real robot has a long adjustment cycle, high trial and error cost, and risks of equipment damage. Therefore, the industry generally adopts a simulation-first strategy, that is, training a control strategy model on a simulation robot first, and then migrating it to a real robot for execution. However, there are real differences between the simulation robot and the real robot, such as dynamic model errors, sensor noise, and actuator response delays, which result in poor performance or even complete failure of the strategy that performs well on the simulation robot on the real robot. Therefore, a method is needed to effectively reduce the differences between the simulation robot and the real robot and improve the success rate of strategy migration to control the real robot to reliably complete the target task. SUMMARY

[0003] The present disclosure provides an object control method, device, electronic equipment and storage medium to reduce the differences between the simulation robot and the real robot and improve the success rate of strategy migration to control the real robot to reliably complete the target task.

[0004] According to a first aspect of an embodiment of the present disclosure, an object control method is provided, comprising: obtaining first current state data of a simulation object corresponding to a real object in a target simulation environment corresponding to a target task, a simulation parameter corresponding to the simulation object being a target simulation parameter, the target simulation parameter being obtained by updating an initial simulation parameter based on preset demonstration data corresponding to a preset demonstration action, the first current state data representing a current motion state of the simulation object; training an initial strategy model in the simulation object based on the first current state data and a target task instruction corresponding to the target task to obtain a target strategy model corresponding to the target task; the target strategy model is used to generate a control instruction for an application object to make the application object complete the target task, the application object being an object to which the target strategy model is applied; deploy the target policy model in the real object, so that the target policy model generates a real control instruction for the real object based on the target task instruction, a control parameter of a real controller corresponding to the real object is a target real control parameter, and the target real control parameter is obtained by updating an initial real control parameter based on the preset demonstration data; and the target real control parameter represents a response characteristic of the real controller. send the real control instruction to the real controller, so that the real controller generates a correction signal to control the real object based on the target real control parameter and the real control instruction, and makes the real object complete the target task.

[0005] According to a second aspect of the embodiments of the present disclosure, an object control apparatus is provided, including: a first current state data acquisition module configured to acquire first current state data of a simulation object corresponding to a real object in a target simulation environment corresponding to a target task, a simulation parameter corresponding to the simulation object is a target simulation parameter, the target simulation parameter is obtained by updating an initial simulation parameter based on preset demonstration data corresponding to a preset demonstration action, and the first current state data represents a current motion state of the simulation object; a policy model training module configured to train an initial policy model in the simulation object based on the first current state data and a target task instruction corresponding to the target task, to obtain a target policy model corresponding to the target task; the target policy model is used to generate a control instruction for an application object, so that the application object completes the target task, and the application object is an object to which the target policy model is applied; a model deployment module configured to deploy the target policy model in the real object, so that the target policy model generates a real control instruction for the real object based on the target task instruction, a control parameter of a real controller corresponding to the real object is a target real control parameter, and the target real control parameter is obtained by updating an initial real control parameter based on the preset demonstration data; and the target real control parameter represents a response characteristic of the real controller. a real object control module configured to send the real control instruction to the real controller, so that the real controller generates a correction signal to control the real object based on the target real control parameter and the real control instruction, and makes the real object complete the target task.

[0006] According to a third aspect of the embodiments of the present disclosure, an electronic device is provided, including a processor, a memory for storing instructions executable by the processor, and the processor is configured to execute the instructions to implement the method according to any one of the first aspect.

[0007] According to a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, which, when instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to perform the method in any one of the first aspect of the embodiments of the present disclosure. According to a fifth aspect of the embodiments of the present disclosure, a computer program product containing instructions, which, when running on a computer, enables the computer to perform the method in any one of the first aspect of the embodiments of the present disclosure.

[0008] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects: The first current state data of the simulation object corresponding to the real object in the target simulation environment corresponding to the target task is acquired, wherein the simulation parameter corresponding to the simulation object is the target simulation parameter obtained by updating the initial simulation parameter based on the preset demonstration data corresponding to the preset demonstration action, and a high-fidelity basis is provided for policy model training; the target policy model corresponding to the target task is trained by using the first current state data and the target task instruction; then the target policy model is deployed into the real object, so that the target policy model generates real control instructions for the real object based on the target task instruction, wherein the control parameter of the real controller corresponding to the real object is the target real control parameter obtained by updating the initial real control parameter based on the preset demonstration data, and the initial simulation parameter and the initial real control parameter are updated based on the same preset demonstration data, so as to ensure that the simulation object and the real object are highly consistent in behavior level; finally, the real controller generates a correction signal by combining the target real control parameter updated based on the preset demonstration data and the real control instruction, so as to control the real object to stably complete the target task, achieve efficient and safe alignment of the simulation object and the real object in the motion control level, reduce the risk and cost of adjusting the real object, control the real object to reliably complete the target task, and thus improve the adaptability, stability and task completion efficiency of the control system as a whole.

[0009] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF DRAWINGS

[0010] The accompanying drawings, which are incorporated into and form part of the specification, illustrate an embodiment consistent with the present disclosure and, together with the specification, serve to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.

[0011] Figure 1 is a flowchart of an object control method according to an exemplary embodiment; Figure 2 is a flowchart of determining a target simulation parameter according to an exemplary embodiment; Figure 3 is a flowchart of a process of constructing a simulation object according to an example embodiment; Figure 4 is a flowchart of a process of training an initial policy model in a simulation object based on first current state data and target task instructions corresponding to a target task according to an example embodiment; Figure 5 is a flowchart of a process of deploying a target policy model in a real object to enable the target policy model to generate real control instructions for the real object based on target task instructions according to an example embodiment; Figure 6 is a flowchart of a process of determining a target real control parameter according to an example embodiment; Figure 7 is a block diagram of an object control device according to an example embodiment; Figure 8 is a block diagram of an electronic device for object control according to an example embodiment. DETAILED DESCRIPTION

[0012] In order to make the ordinary person skilled in the art better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be described clearly and completely below in conjunction with the drawings.

[0013] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar different contents, and do not necessarily describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The implementation described in the following example embodiments does not represent all implementations consistent with the present disclosure. Rather, they are only examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0014] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, analyzed data, etc.) involved in the present disclosure are all information and data authorized by the user or authorized by all parties.

[0015] Figure 1 is a flowchart of an object control method according to an example embodiment, which is applied to an electronic device such as a server, as shown in Figure 1 includes the following steps: In step S101, first current state data of a simulation object corresponding to a real object in a target simulation environment corresponding to a target task is acquired.

[0016] In one specific embodiment, the real object is a real system existing in the physical world and having physical structure and dynamics, and the real object may be a real robot, for example. The simulation object is a virtual model simulating the real object and running in a simulation environment, and the running of the simulation object is completely based on the computer simulation environment.

[0017] In one specific embodiment, the target task is a preset operation task (such as walking, jumping, and precise grasping) to be performed by the real object in the real world. However, since the parameter debugging, strategy verification, or action optimization related to the task are directly carried out on the real object, there is a risk of long adjustment cycle, high trial and error cost, and easy damage to the equipment. Therefore, a simulation object matching the characteristics of the real object is constructed, and training of a target strategy model corresponding to the target task is completed based on the simulation object in the target simulation environment corresponding to the target task.

[0018] In one specific embodiment, the target simulation environment is a virtual running scene constructed based on the characteristics of the real scene to adapt to the requirements of the target task, and is used to provide simulation running conditions matching the target task for the simulation object and support the simulation object to complete virtual execution of the target task.

[0019] In one specific embodiment, the first current state data represents the current motion state of the simulation object. For example, the first current state data may include at least one of current joint angle data, current joint angular velocity data, and current joint torque data. Specifically, the current joint angle data represents the current spatial position state of each motion joint of the simulation object, the current joint angular velocity data represents the current rotation speed state of each motion joint of the simulation object, and the current joint torque data represents the current force or driving force state of each motion joint of the simulation object.

[0020] In one specific embodiment, the simulation parameters corresponding to the simulation object are target simulation parameters, and the target simulation parameters are obtained by updating initial simulation parameters based on preset demonstration data corresponding to a preset demonstration action. The preset demonstration data is motion state data of a preset reference object recorded in a process in which the preset reference object performs the preset demonstration action.

[0021] Exemplarily, the preset demonstration data can include at least one of preset joint angle data, preset joint angular velocity data, and preset joint torque data of the preset reference object. Specifically, the preset joint angle data is the angle value of each joint of the preset reference object at different time when performing the preset demonstration action, and is used to reflect the position state of the joint; the preset joint angular velocity data is the rate value of the change of the joint angle with time, and is used to reflect the speed of the joint movement; and the preset joint torque data is the torque value required to drive each joint to complete the preset demonstration action, and is used to reflect the power demand of the joint movement.

[0022] In one specific embodiment, the target simulation parameters include target physical parameters of the simulation object and target simulation control parameters of the simulation controller corresponding to the simulation object; the target physical parameters represent inherent physical properties of the simulation object, i.e., physical characteristics inherent to the simulation object, which do not change with the control process and determine the basic dynamic behavior of the simulation object; and the target simulation control parameters represent the response characteristics of the simulation controller, i.e., the feedback mode of the simulation controller to the received control instructions, specifically, the adjustment logic of the simulation controller for adjusting the movement of the simulation object after receiving the control instructions. Specifically, the target physical parameters are obtained by updating the initial physical parameters, and the target simulation control parameters are obtained by updating the initial simulation control parameters.

[0023] Exemplarily, the target physical parameters can include at least one of joint mass data, joint inertia data, and joint friction coefficient data of the simulation object. Specifically, the joint mass data is the mass value of the components (such as motors and connecting rods) to which the joints of the simulation object belong; the joint inertia data is the value of the inertia of the movement of the joints of the simulation object around the rotation axis; and the joint friction coefficient data is the value of the resistance of the internal mechanical contact (such as bearings and gears) of the joints of the simulation object.

[0024] Exemplarily, the target simulation control parameters can include at least one of proportional-differential gain parameters and impedance control parameters. Specifically, the proportional-differential gain parameters can include at least one of proportional gain and differential gain, and the impedance control parameters can include at least one of virtual stiffness and virtual damping. The proportional gain is used to adjust the output response strength of the simulation controller according to the deviation between the actual state of the simulation object and the control instruction, and to improve the action adjustment sensitivity; the differential gain is used to predict the rate of change of the state deviation of the simulation object, to suppress the oscillation phenomenon in the action response process, and to ensure the stability of the movement; the virtual stiffness is used to simulate the rigidity characteristics of the simulation object when interacting with the environment or the target, to control the posture holding ability in the action execution, and to realize the output force adaptability; and the virtual damping is used to buffer the speed change of the simulation object movement, to reduce the impact feeling when the action starts and stops, and to optimize the movement stability and the action accuracy.

[0025] In one specific embodiment, asFigure 2 As shown in the above, the target physical parameter and the target simulation control parameter are determined in the following manner: In step S201, based on preset demonstration data, the simulation object is controlled to perform a preset demonstration action in a preset simulation environment to obtain current simulation action data.

[0026] In one specific embodiment, the current simulation action data is information reflecting the real-time motion state of the simulation object when the simulation object performs the preset demonstration action in the preset simulation environment. The data included in the preset demonstration data corresponds one-to-one to the data included in the current simulation action data.

[0027] In the case where the preset demonstration data includes preset joint angle data, preset joint angular velocity data and preset joint torque data of a preset reference object, the current simulation action data includes current joint angle data, current joint angular velocity data and current joint torque data of the simulation object. The specific details of the current joint angle data, the current joint angular velocity data and the current joint torque data can be referred to the preset joint angle data, the preset joint angular velocity data and the preset joint torque data respectively, which will not be described here.

[0028] In step S203, based on the current simulation action data and the preset demonstration data, the initial physical parameter and the initial simulation control parameter are updated.

[0029] In one specific embodiment, the updating of the initial physical parameter and the initial simulation control parameter based on the current simulation action data and the preset demonstration data can include determining the difference between each data included in the current simulation action data and the corresponding data in the preset demonstration data, and updating the initial physical parameter and the initial simulation control parameter.

[0030] In step S205, based on the simulation object after the initial physical parameter and the initial simulation control parameter are updated, the loop iteration operation of repeating step S201 of controlling the simulation object to perform a preset demonstration action in a preset simulation environment based on preset demonstration data to obtain current simulation action data, and step S203 of updating the initial physical parameter and the initial simulation control parameter based on the current simulation action data and the preset demonstration data is repeated until the difference between the current simulation action data and the preset demonstration data is less than a second preset threshold, and the corresponding current physical parameter is taken as the target physical parameter and the corresponding current simulation control parameter is taken as the target simulation control parameter.

[0031] In the above embodiment, the current simulation action data is generated as a comparison benchmark by driving the simulation object to perform the demonstration action in the preset environment using the preset demonstration data; and then the initial physical parameters and the initial simulation control parameters are iteratively optimized based on the difference between the current data and the demonstration data, and when the difference between the current simulation action data and the preset demonstration data is less than a second preset threshold, the current parameters are confirmed as the target physical parameters and the target simulation control parameters, thereby providing a simulation carrier that fits the real object for subsequent control strategy training.

[0032] In one specific embodiment, as shown in Figure 3 The simulation object is constructed in the following manner: In step S301, the initial physical parameters and the initial simulation control parameters are determined based on the real object.

[0033] In one specific embodiment, the initial physical parameters and the initial simulation control parameters are determined based on the real object, which can include determining the initial physical parameters of the simulation object based on the inherent physical characteristics of the real object, and determining the initial simulation control parameters of the simulation object corresponding to the simulation controller based on the control characteristics of the real controller corresponding to the real object.

[0034] In step S303, the initial physical parameters and the initial simulation control parameters are configured in the preset simulation environment to construct the simulation object.

[0035] In the above embodiment, the initial physical parameters and the initial simulation control parameters are determined based on the real object to ensure that the dynamics of the simulation object are consistent with those of the real object; and then these parameters are configured in the preset simulation environment to construct a high-fidelity simulation object that can reproduce the motion characteristics of the real object in the real world, thereby providing a reliable foundation for strategy training and parameter optimization and effectively improving the efficiency and success rate of simulation to reality migration.

[0036] In step S103, the initial strategy model in the simulation object is trained based on the first current state data and the target task instruction corresponding to the target task to obtain a target strategy model corresponding to the target task.

[0037] In one specific embodiment, the target strategy model is used to generate control instructions for the application object to complete the target task. The application object is the object to which the target strategy model is applied.

[0038] In one specific embodiment, the target task instruction is an instruction that explicitly indicates at least one of the action execution requirements, the expected effect, and the constraint conditions of the target task.

[0039] With the simulation object being a robot, the target task is an example of making the robot's mechanical arm grasp a target object, and the target task instruction can include at least one of an end effector target position instruction, a joint motion parameter instruction, an end effector interaction parameter instruction, and a motion constraint instruction.

[0040] Specifically, the end effector target position instruction is used to explicitly indicate the three-dimensional coordinate range of the target object and the position accuracy requirement that the end effector needs to reach; the joint motion parameter instruction can include at least one of the motion range constraint and the target motion accuracy threshold of each joint of the mechanical arm, representing the allowed boundary and performance standard of joint action; the end effector interaction parameter instruction can include at least one of the clamping force constraint range, the opening and closing amplitude limit, and the contact response sensitivity requirement of the end effector in the grasping process, to specify the interaction criteria with the target object; and the motion constraint instruction can include at least one of the maximum speed upper limit, the acceleration limit, and the spatial range of the prohibited collision area of the mechanical arm motion, to guarantee the safety and compliance of task execution.

[0041] In actual applications, the correspondence between the target task and the target strategy model can be flexibly designed according to the similarity between tasks, resource constraints, and system optimization objectives. One target task can correspond to one target strategy model, for example, when the differences between tasks are significant or highly specialized strategies are needed, a dedicated target strategy model for each target task is trained, such as walking and jumping tasks. Multiple target tasks can also share one target strategy model, for example, when tasks have common dynamic characteristics or the strategy generalization ability needs to be improved, such as moving tasks on different terrains.

[0042] In one specific embodiment, as shown in Figure 4 Based on the first current state data and the target task instruction corresponding to the target task, training an initial strategy model in the simulation object includes: In step S401, the first current state data and the target task instruction are input into the initial strategy model to obtain a simulation control instruction for the simulation object.

[0043] In one specific embodiment, the simulation control instruction is a specific operation instruction that adapts to the characteristics of the simulation object, used to explicitly indicate the action execution parameters of the simulation object.

[0044] With the simulation object being a robot, the target task is an example of making the robot's mechanical arm grasp a target object, and the simulation control instruction can include at least one of a joint motion control instruction, an end effector control instruction, and a motion trajectory planning instruction.

[0045] Specifically, the joint motion control instruction includes target joint angles, joint angular velocities and angular accelerations of each joint of the robot arm, is used to directly drive each joint to move according to a preset rule, realize spatial position movement of the end effector, and ensure that the joint action matches the posture requirement of the target task; the end effector control instruction includes a target clamping force, an opening and closing amplitude and a motion triggering time of the end effector, is used to control the end effector to complete core operations such as grabbing and releasing, and avoid target objects from falling off or being damaged due to improper clamping force; and the motion trajectory planning instruction includes motion path parameters (such as path point coordinates and motion timing) of the robot arm from an initial position to a target position, is used to plan an optimal motion trajectory, take into account action efficiency and safety, and avoid collision with obstacles in the simulation environment.

[0046] In step S403, the simulation object is controlled to execute the simulation control instruction, and second current state data is obtained.

[0047] In one specific embodiment, the second current state data is new state data generated after the simulation object executes the simulation control instruction. For specific details of the second current state data, reference can be made to the first current state data described above, which will not be repeated here.

[0048] In step S405, based on the second current state data, current reward data obtained after the simulation object executes the simulation control instruction is determined.

[0049] In one specific embodiment, the determination of the current reward data obtained after the simulation object executes the simulation control instruction based on the second current state data can include: based on the second current state data reflecting the actual state of the simulation object after executing the simulation control instruction, combining a preset reward function of the target task, judging the degree of fit of the action execution effect of the simulation object to the requirement of the target task, and then determining the corresponding current reward data.

[0050] In one specific embodiment, the preset reward function is a function that is set in advance for the target task and is used to quantitatively map the degree of fit of the action execution effect of the simulation object reflected by the second current state data to the requirement of the target task.

[0051] Taking the simulation object as a robot dog and the target task as straight-line walking as an example, the preset reward function can be a weighted sum of a speed reward term and a gait regularity reward. Specifically, the speed reward term obtains a positive reward when the actual forward speed of the robot dog approaches a target speed (such as 0.5 m / s), and the reward value is lower when the speed deviation is larger; the gait regularity reward gives an additional reward through the regularity of the foot end contact timing, and promotes the generation of a coordinated gait. Optionally, the weights corresponding to the speed reward term and the gait regularity reward can be set in advance.

[0052] With the robot as the simulation object, the target task is to make the robot's mechanical arm grasp the target object, and the preset reward function can be a weighted sum of the distance reward, the grasping success reward, and the time efficiency reward. Specifically, the distance reward is a reward given in proportion to the distance between the end effector and the target object; the grasping success reward is a reward given when the force sensor detects successful grasping; and the time efficiency reward is a time reward obtained when the task is completed within a limited time. Optionally, the weights corresponding to the distance reward, the grasping success reward, and the time efficiency reward can be preset.

[0053] In step S407, the parameters of the initial policy model are updated based on the first current state data, the simulation control instruction, the second current state data, and the current reward data.

[0054] In step S409, the second current state data is taken as the first current state data, and the steps S401 and S407 are repeated, i.e., the first current state data and the target task instruction are input into the initial policy model to obtain the simulation control instruction for the simulation object, and the parameters of the initial policy model are updated based on the first current state data, the simulation control instruction, the second current state data, and the current reward data, until the performance index of the current policy model meets the preset convergence condition, and the current policy model that meets the preset convergence condition is taken as the target policy model.

[0055] In one specific embodiment, the preset convergence condition is used to determine whether the policy model training is up to standard. For example, the preset convergence condition can be an average reward threshold or a performance stability criterion.

[0056] Specifically, the average reward threshold is that the average reward value of the policy model in a continuous preset number of training periods continuously exceeds a preset target; and the performance stability criterion is that the standard deviation of the reward value in the last preset number of iterations is less than a preset threshold.

[0057] In the above embodiment, the first current state data and the target task instruction are input into the initial policy model to generate the simulation control instruction, drive the simulation object to perform an action, and obtain the second current state data reflecting the actual state change of the simulation object after executing the instruction; the current reward data is determined based on the second current state data, and the model parameters are updated in combination with the initial state, the control instruction, and the reward signal; the adaptability of the policy is gradually improved through iterative optimization until the model performance meets the convergence condition, and finally the target policy model capable of accurately guiding the simulation object to complete the target task is obtained, thereby improving the decision-making accuracy of the target policy model and the adaptability to the target task.

[0058] In step S105, the target policy model is deployed in the real object, so that the target policy model generates a real control instruction for the real object based on the target task instruction.

[0059] In one specific embodiment, the control parameter of the real object corresponding to the real controller is a target real control parameter, and the target real control parameter is obtained by updating the initial real control parameter based on preset demonstration data; the target real control parameter represents the response characteristic of the real controller, i.e., the feedback mode of the real controller to the received control instruction, specifically, the adjustment logic of the real controller for adjusting the motion of the real object after receiving the control instruction. For specific refinement of the target real control parameter, refer to the above-mentioned target simulation control parameter, which will not be described here.

[0060] In one specific embodiment, as shown in Figure 5 The above-mentioned deployment of the target strategy model in the real object to make the target strategy model generate a real control instruction for the real object based on the target task instruction includes: In step S501, the target strategy model is deployed in the real object.

[0061] In one specific embodiment, the above-mentioned deployment of the target strategy model in the real object can include: integrating the target strategy model adapted to the target task and the characteristics of the real object after simulation training into the control unit or hardware system of the real object.

[0062] In step S503, the current real state data of the real object is obtained.

[0063] In one specific embodiment, the current real state data represents the current motion state of the real object. For specific refinement of the current real state data, refer to the above-mentioned first current state data, which will not be described here.

[0064] In step S505, the current real state data and the target task instruction are input into the target strategy model to obtain a real control instruction.

[0065] Specifically, for specific refinement of the real control instruction, refer to the above-mentioned simulation control instruction, which will not be described here.

[0066] In the above-mentioned embodiment, the target strategy model is deployed in the real object, and then the current real state data of the real object is obtained, and then the current real state data and the target task instruction are input into the target strategy model to generate a corresponding real control instruction, thereby realizing accurate task analysis and action planning and improving the reliability of task completion.

[0067] In one specific embodiment, as shown in Figure 6 The above-mentioned target real control parameter is determined in the following manner: In step S601, based on preset demonstration data, the real object is controlled to perform a preset demonstration action to obtain current real action data.

[0068] In one specific embodiment, the current real action data is information reflecting the real-time motion state of the real object when the real object is performing the preset demonstration action. The preset demonstration data includes data corresponding to the data included in the current real action data. For specific refinement of the current real action data, please refer to the above-mentioned current simulation action data, which will not be repeated here.

[0069] In step S603, the initial real control parameter is updated based on the current real action data and the preset demonstration data.

[0070] In one specific embodiment, the above-mentioned updating of the initial real control parameter based on the current real action data and the preset demonstration data can include updating the initial real control parameter based on the difference between the current real action data and the preset demonstration data.

[0071] In step S605, based on the real object after updating the initial real control parameter, the step S601 of controlling the real object to perform the corresponding preset demonstration action based on the preset demonstration data is repeated to obtain the current real action data, and the step S603 of updating the initial real control parameter based on the current real action data and the preset demonstration data is iterated until the difference between the current real action data and the preset demonstration data is less than a first preset threshold, and the corresponding current real control parameter is taken as the target real control parameter.

[0072] In the above-mentioned embodiment, the real object is controlled to perform the corresponding preset demonstration action based on the preset demonstration data, and the current real action data is obtained, the initial real control parameter is updated based on the difference between the current real action data and the preset demonstration data, the response characteristics of the real controller are adjusted in a targeted manner, and the loop iteration operation of repeatedly performing the demonstration action to obtain data and updating the parameter based on the difference is performed to continuously reduce the deviation between the real action and the preset demonstration action until the difference between the two meets the preset requirement, and the target real control parameter finally determined can make the response characteristics of the real controller accurately adapt to the preset demonstration action, thereby improving the accuracy of the target real control parameter and ensuring the accuracy of the action performed by the real object.

[0073] In step S107, the real control instruction is sent to the real controller, so that the real controller generates a correction signal based on the target real control parameter and the real control instruction to control the real object to complete the target task.

[0074] In one specific embodiment, the correction signal is used to eliminate the execution deviation between the current motion state of the real object and the requirement of the target task.

[0075] In one specific embodiment, the sending of the real control instruction to the real controller, so that the real controller generates a correction signal based on the target real control parameter and the real control instruction to control the real object to complete the target task, comprises: sending the real control instruction to the real controller, so that the real controller generates a correction signal based on the current real state data, the target real control parameter and the real control instruction to control the real object to complete the target task.

[0076] Taking the real object as a robot, the target task as making the robot's mechanical arm grasp a target object, the real control instruction as containing target angles of each joint of the mechanical arm, target force of an end effector and target motion speed, the current real state data as containing real-time angles of each joint of the mechanical arm, real-time force of the end effector and real-time motion speed, and the target real control parameter as including proportional gain, differential gain, virtual stiffness and virtual damping, the real controller generates a correction signal by combining each target real control parameter in the following way: The real controller first calculates the deviation of the joint target angle in the real control instruction from the real-time angle of the joint in the current real state data, multiplies the angle deviation by the proportional gain to obtain a proportional correction component, which is used to quickly adjust the driving force of the joint and reduce the deviation of the real-time angle from the target angle. Based on the joint angle deviation at the current and previous time, the real controller calculates the rate of change of the angle deviation, multiplies the rate of change by the differential gain to obtain a differential correction component, which is used to suppress the oscillation phenomenon in the joint motion process and avoid unstable motion caused by too fast angle adjustment. The real controller calculates the deviation of the target force of the end effector in the real control instruction from the real-time force in the current real state data, multiplies the force deviation by the virtual stiffness to obtain a stiffness correction component, which is used to simulate the stiffness characteristics when the mechanical arm interacts with the target object, ensure that the end effector maintains a preset posture when contacting the object, and avoid excessive deformation. The real controller calculates the deviation of the target motion speed of the end effector in the real control instruction from the real-time motion speed in the current real state data, multiplies the speed deviation by the virtual damping to obtain a damping correction component, which is used to buffer the speed change and reduce the impact when the mechanical arm starts or stops or turns, making the motion process smoother. The real controller weights and sums the proportional correction component, the differential correction component, the stiffness correction component and the damping correction component to generate a correction signal.

[0077] Optionally, the preset weights corresponding to the proportional correction component, the differential correction component, the stiffness correction component and the damping correction component can be preset.

[0078] In actual application, the real controller can judge the motion scene in combination with the current real state data, and dynamically adjust the weight. For example, when the joint angle deviation is large, the weight of the proportional correction component is increased to accelerate the convergence of the deviation; when it is detected that the joint motion is oscillating, the weight of the differential correction component is increased to enhance stability; and when the end effector contacts the target object, the weight of the stiffness correction component is increased to ensure the interaction posture.

[0079] In a specific embodiment, the real controller can also limit the amplitude of the correction signal to ensure that it does not exceed the rated working range of the real object actuator (such as a motor or a cylinder), thereby avoiding damage to the equipment due to signal overload, and finally outputting a correction signal that can directly drive the real object to move.

[0080] In the above embodiments, the real control instruction is sent to the real controller, so that the controller can generate a correction signal that can correct the motion deviation by comprehensively considering the current real state data, the target real control parameter and the real control instruction, dynamically adjust the motion execution process of the real object through the correction signal, eliminate the execution deviation, and make the motion of the real object always meet the target task requirements, thereby ensuring that the real object efficiently completes the target task.

[0081] Figure 7 is a block diagram of an object control device according to an example embodiment. Referring to Figure 7 The device includes: A first current state data acquisition module 710 is configured to acquire first current state data of a simulation object corresponding to a real object in a target simulation environment corresponding to a target task, the simulation parameter corresponding to the simulation object being a target simulation parameter, the target simulation parameter being obtained by updating an initial simulation parameter based on preset demonstration data corresponding to a preset demonstration motion, and the first current state data representing a current motion state of the simulation object. A strategy model training module 720 is configured to train an initial strategy model in the simulation object based on the first current state data and a target task instruction corresponding to the target task, to obtain a target strategy model corresponding to the target task; the target strategy model is configured to generate a control instruction for an application object, so that the application object completes the target task, and the application object is an object to which the target strategy model is applied. A model deployment module 730 is configured to deploy the target strategy model in the real object, so that the target strategy model generates a real control instruction for the real object based on the target task instruction, and the control parameter of the real controller corresponding to the real object is a target real control parameter, the target real control parameter being obtained by updating an initial real control parameter based on preset demonstration data; the target real control parameter represents the response characteristic of the real controller. The real object control module 740 is configured to send the real control instruction to the real controller, so that the real controller generates a correction signal based on the target real control parameter and the real control instruction to control the real object, and the real object completes the target task.

[0082] In an optional embodiment, the model deployment module 730 includes: a model deployment unit configured to deploy the target policy model in the real object; a current real state data acquisition unit configured to acquire current real state data of the real object, the current real state data representing a current motion state of the real object; a real control instruction determination unit configured to input the current real state data and the target task instruction into the target policy model to obtain the real control instruction.

[0083] In an optional embodiment, the real object control module 740 includes: a real object control unit configured to send the real control instruction to the real controller, so that the real controller generates a correction signal based on the current real state data, the target real control parameter and the real control instruction to control the real object, and the real object completes the target task.

[0084] In an optional embodiment, the target real control parameter is determined by the following modules: a current real action data determination module configured to control the real object to perform a preset demonstration action corresponding to preset demonstration data based on the preset demonstration data to obtain current real action data; an initial real control parameter update module configured to update the initial real control parameter based on the current real action data and the preset demonstration data; a target real control parameter determination module configured to repeatedly control the real object to perform a preset demonstration action corresponding to preset demonstration data based on the real object after the initial real control parameter is updated to obtain current real action data, until a loop iteration operation based on the current real action data and the preset demonstration data to update the initial real control parameter is performed until a difference between the current real action data and the preset demonstration data is less than a first preset threshold, and a current real control parameter corresponding to the target real control parameter.

[0085] In an optional embodiment, the target simulation parameter includes a target physical parameter of the simulation object and a target simulation control parameter of a simulation controller corresponding to the simulation object; the target physical parameter represents an inherent physical property of the simulation object, and the target simulation control parameter represents a response characteristic of the simulation controller; the target physical parameter is obtained by updating an initial physical parameter, and the target simulation control parameter is obtained by updating an initial simulation control parameter; the target physical parameter and the target simulation control parameter are determined by the following modules: The current simulation action data determination module is configured to control the simulation object to perform the preset demonstration action corresponding to the preset demonstration data in the preset simulation environment based on the preset demonstration data, to obtain current simulation action data. The initial simulation parameter updating module is configured to update the initial physical parameter and the initial simulation control parameter based on the current simulation action data and the preset demonstration data. The target simulation parameter determination module is configured to repeat the operation of controlling the simulation object to perform the preset demonstration action corresponding to the preset demonstration data in the preset simulation environment based on the preset demonstration data, to obtain the current simulation action data, until the difference between the current simulation action data and the preset demonstration data is less than a second preset threshold, and to take the corresponding current physical parameter as the target physical parameter and the corresponding current simulation control parameter as the target simulation control parameter, based on the updated initial physical parameter and initial simulation control parameter.

[0086] In an optional embodiment, the simulation object is constructed by using the following modules: The initial simulation parameter determination module is configured to determine the initial physical parameter and the initial simulation control parameter based on the real object. The simulation object construction module is configured to configure the initial physical parameter and the initial simulation control parameter in the preset simulation environment, to construct the simulation object.

[0087] In an optional embodiment, the strategy model training module 720 includes: The simulation control instruction determination unit is configured to input the first current state data and the target task instruction into the initial strategy model, to obtain the simulation control instruction for the simulation object. The second current state data acquisition unit is configured to control the simulation object to execute the simulation control instruction, and to acquire the second current state data, which is new state data generated after the simulation object executes the simulation control instruction. The current reward data determination unit is configured to determine the current reward data obtained by the simulation object after executing the simulation control instruction, based on the second current state data. The strategy model parameter updating unit is configured to update the parameters of the initial strategy model, based on the first current state data, the simulation control instruction, the second current state data, and the current reward data. The target policy model determination unit is configured to input the second current state data as the first current state data into the initial policy model repeatedly with the first current state data and the target task instruction, to obtain a simulation control instruction for the simulation object, and to update the parameters of the initial policy model based on the first current state data, the simulation control instruction, the second current state data, and the current reward data, until the performance index of the current policy model meets the preset convergence condition, and to take the current policy model that meets the preset convergence condition as the target policy model.

[0088] Figure 8 is a block diagram of an electronic device for object control according to an example embodiment. The electronic device can be a server, and its internal structure can be as shown in Figure 8 The electronic device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. The processor of the electronic device is configured to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The network interface of the electronic device is configured to communicate with external terminals through network connections. The computer program is executed by the processor to implement an object control method. The display screen of the electronic device can be a liquid crystal display screen or an electronic ink display screen. The input device of the electronic device can be a touch layer overlaid on the display screen, or a key, trackball, or touchpad provided on the housing of the electronic device. The input device can also be an external keyboard, touchpad, or mouse, etc.

[0089] Those skilled in the art can understand that Figure 8 the structure shown in the above description is only a block diagram of part of the structure related to the present disclosure, and does not constitute a limitation on the electronic device to which the present disclosure is applied. A specific electronic device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components. In an example embodiment, an electronic device is also provided, including a processor, and a memory for storing instructions executable by the processor. The processor is configured to execute the instructions to implement an object control method as in the embodiments of the present disclosure.

[0090] In an example embodiment, a computer-readable storage medium is also provided, and when the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform an object control method as in the embodiments of the present disclosure.

[0091] In an example embodiment, a computer program product containing instructions which, when the program is executed by a computer, causes the computer to carry out the object control method in the embodiments of the present disclosure is also provided.

[0092] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a nonvolatile computer readable storage medium, and when the computer program is executed, the computer program can include the processes of the above-mentioned embodiments of the methods. Any reference to memory, storage, databases, or other media in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM), or external cache memory. As an illustration but not limitation, RAM is available in many forms such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0093] Other embodiments of the present disclosure will be apparent to those skilled in the art with the consideration of the specification and practice of the disclosure disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present disclosure following the general principles thereof and including modifications and equivalents that are obvious to those skilled in the art. The specification and examples are to be regarded as exemplary only, and the true scope and spirit of the present disclosure are indicated by the following claims.

[0094] It should be understood that the present disclosure is not limited to the precise structures as herein described and illustrated in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the claims that follow.

Claims

1. An object control method characterized by, The method comprises: obtaining first current state data of a simulation object corresponding to a real object in a target simulation environment corresponding to a target task, a simulation parameter corresponding to the simulation object being a target simulation parameter, the target simulation parameter being obtained by updating an initial simulation parameter based on preset demonstration data corresponding to preset demonstration actions, the first current state data representing a current motion state of the simulation object; training an initial strategy model in the simulation object based on the first current state data and target task instructions corresponding to the target task to obtain a target strategy model corresponding to the target task; the target strategy model being used to generate control instructions for an application object to enable the application object to complete the target task, the application object being an object to which the target strategy model is applied; deploying the target strategy model in the real object to enable the target strategy model to generate real control instructions for the real object based on the target task instructions, a control parameter of a real controller corresponding to the real object being a target real control parameter, the target real control parameter being obtained by updating an initial real control parameter based on the preset demonstration data; the target real control parameter representing a response characteristic of the real controller; sending the real control instructions to the real controller to enable the real controller to generate a correction signal based on the target real control parameter and the real control instructions to control the real object, so that the real object completes the target task.

2. The method of claim 1, wherein, The deploying of the target strategy model in the real object to enable the target strategy model to generate real control instructions for the real object based on the target task instructions comprises: deploying the target strategy model in the real object; obtaining current real state data of the real object, the current real state data representing a current motion state of the real object; inputting the current real state data and the target task instructions into the target strategy model to obtain the real control instructions.

3. The method of claim 2, wherein, The sending of the real control instructions to the real controller to enable the real controller to generate a correction signal based on the target real control parameter and the real control instructions to control the real object, so that the real object completes the target task comprises: sending the real control instructions to the real controller to enable the real controller to generate the correction signal based on the current real state data, the target real control parameter and the real control instructions to control the real object, so that the real object completes the target task.

4. The method of claim 1, wherein, The target real control parameter is determined in the following manner: controlling the real object to perform corresponding preset demonstration actions based on the preset demonstration data to obtain current real action data; updating the initial real control parameter based on the current real action data and the preset demonstration data; The operation of repeating the operation of controlling the real object to perform the corresponding preset demonstration action based on the preset demonstration data until the current real action data and the preset demonstration data are updated based on the current real action data and the preset demonstration data is repeated until a difference between the current real action data and the preset demonstration data is less than a first preset threshold, and a corresponding current real control parameter is taken as the target real control parameter.

5. The method of claim 1, wherein, The target simulation parameter includes a target physical parameter of the simulation object and a target simulation control parameter of a simulation controller corresponding to the simulation object; the target physical parameter represents an inherent physical property of the simulation object, and the target simulation control parameter represents a response characteristic of the simulation controller; The target physical parameter is obtained by updating an initial physical parameter, and the target simulation control parameter is obtained by updating an initial simulation control parameter; The target physical parameter and the target simulation control parameter are determined in the following manner: The simulation object is controlled to perform a corresponding preset demonstration action in a preset simulation environment based on the preset demonstration data to obtain current simulation action data; The initial physical parameter and the initial simulation control parameter are updated based on the current simulation action data and the preset demonstration data; The operation of repeating the operation of controlling the simulation object to perform the corresponding preset demonstration action in the preset simulation environment based on the preset demonstration data until the current simulation action data and the preset demonstration data are updated based on the current simulation action data and the preset demonstration data is repeated until a difference between the current simulation action data and the preset demonstration data is less than a second preset threshold, and a corresponding current physical parameter is taken as the target physical parameter, and a corresponding current simulation control parameter is taken as the target simulation control parameter.

6. The method of claim 5, wherein, The simulation object is constructed in the following manner: The initial physical parameter and the initial simulation control parameter are determined based on the real object; The initial physical parameter and the initial simulation control parameter are configured in the preset simulation environment to construct the simulation object.

7. The method of claim 1, wherein, The operation of training an initial policy model in the simulation object based on the first current state data and a target task instruction corresponding to the target task includes: The first current state data and the target task instruction are input into the initial policy model to obtain a simulation control instruction for the simulation object; The simulation object is controlled to execute the simulation control instruction, and second current state data is obtained, the second current state data being new state data generated after the simulation object executes the simulation control instruction; Current reward data obtained after the simulation object executes the simulation control instruction is determined based on the second current state data; and The initial policy model is trained based on the current reward data and the target task instruction. update parameters of the initial policy model based on the first current state data, the simulation control instruction, the second current state data, and the current reward data; repeat the inputting the first current state data and the target task instruction into the initial policy model to obtain a simulation control instruction for the simulation object, until a performance index of a current policy model meets a preset convergence condition, and the current policy model that meets the preset convergence condition is taken as the target policy model, based on the first current state data, the simulation control instruction, the second current state data, and the current reward data, and updating the parameters of the initial policy model.

8. An object control device, characterized by comprising: comprise: a first current state data acquisition module configured to acquire first current state data of a simulation object corresponding to a real object in a target simulation environment corresponding to a target task, a simulation parameter corresponding to the simulation object being a target simulation parameter, the target simulation parameter being obtained by updating an initial simulation parameter based on preset demonstration data corresponding to preset demonstration actions, the first current state data representing a current motion state of the simulation object; a policy model training module configured to train an initial policy model in the simulation object based on the first current state data and a target task instruction corresponding to the target task, to obtain a target policy model corresponding to the target task; the target policy model being configured to generate a control instruction for an application object, so that the application object completes the target task, the application object being an object to which the target policy model is applied; a model deployment module configured to deploy the target policy model in the real object, so that the target policy model generates a real control instruction for the real object based on the target task instruction, a control parameter of a real controller corresponding to the real object being a target real control parameter, the target real control parameter being obtained by updating an initial real control parameter based on the preset demonstration data; the target real control parameter representing a response characteristic of the real controller; a real object control module configured to send the real control instruction to the real controller, so that the real controller generates a correction signal based on the target real control parameter and the real control instruction to control the real object, so that the real object completes the target task.

9. An electronic device, comprising: comprise: a processor; a memory for storing instructions executable by the processor; wherein the processor is configured to execute the instructions to implement the object control method of any one of claims 1 to 7.

10. A computer readable storage medium characterized by, When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device can perform the object control method of any one of claims 1 to 7.