Manipulator control method and device and electronic equipment
Through neural network reconstruction of robotic tasks, generating control strategies and parameters, the flexible and changeable self-learning control of intelligent robotic robots is realized, and the problem that traditional control methods cannot meet the flexibility and self-learning needs of the improved intelligence level are solved.
Patent Information
- Application Number
- CN202411594216.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-08
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-11-08
AI Technical Summary
The existing technology is difficult to realize the flexible and changeable self-learning control of intelligent robot robots, and traditional control methods cannot meet the needs of action flexibility and self-learning after the degree of intelligence is improved.
By obtaining the task of the robot and determining the neural network, using the neural network to reconstruct the task, obtaining control strategies, including basic control instructions and scene control instructions, based on these instructions, obtaining control parameters and controlling the robot to perform actions.
It realizes flexible and changeable self-learning control of the robot, meets the robot's accurate control needs in complex scenarios, and reduces the requirements for the robot's computing power, and is suitable for a wider range of scenarios.
Smart Images

Figure CN120038738A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application belong to the technical field of manipulators, and in particular, relate to a manipulator control method, device and electronic device. Background Art
[0002] There are various control methods for manipulators, such as traditional PID control (proportional, integral, derivative control), fuzzy control and other methods. When a system and a controlled object are not fully understood, or system parameters cannot be obtained through effective measurement means, PID control technology is most suitable. In PID control, there are also PI and PD controls in practice. A PID controller controls by calculating a control amount based on the error of the system using proportion, integral, and derivative. In fuzzy control, the input quantity is fuzzily quantified into a fuzzy variable, the fuzzy variable is inferred through fuzzy rules to obtain a fuzzy output, and a clear output quantity is obtained through defuzzification for control.
[0003] However, there are many external factors that affect the manipulator of an intelligent robot to obtain better control force. With the improvement of the intelligence level of robots and the development of artificial intelligence, the actions of intelligent robots through traditional control methods can no longer meet the flexible and self-learning operation types, and corresponding control cannot be accurately performed. Summary of the Invention
[0004] The present application aims to at least solve one of the technical problems existing in the related art. For this purpose, the present application provides a manipulator control method, device and electronic device, which can realize the flexible and self-learning control of the manipulator.
[0005] In a first aspect, the present application provides a manipulator control method, which includes:
[0006] Obtain a first task of the manipulator, and determine a first neural network according to the first task;
[0007] Reconstruct the first task by using the first neural network to obtain a control strategy for the reconstructed first task, where the control strategy includes a series of basic control instructions;
[0008] Obtain control parameters for controlling the manipulator based on the control strategy, and control the manipulator to execute a control action according to the control parameters, where the control parameters include parameters corresponding to the basic control instructions.
[0009] In some embodiments, the obtaining the first task of the manipulator includes:
[0010] Obtain the first task and environment state information corresponding to the first task, where the environment state information includes the surrounding environment information and self-state information of the manipulator in a simulation environment;
[0011] Reconstructing the first task by using a pre-trained first neural network to obtain a control strategy for the reconstructed first task, including:
[0012] Inputting the surrounding environment information, the self-state information, and the first task into the first neural network to obtain a control strategy for the reconstructed first task, where the control strategy includes the basic control instruction and the scenario control instruction.
[0013] In some embodiments, obtaining control parameters for controlling the manipulator based on the control strategy, and controlling the manipulator to execute a control action according to the control parameters, including:
[0014] Obtaining control parameters and scenario parameters for controlling the manipulator based on the control strategy, where the scenario parameters are the parameters corresponding to the scenario control instruction;
[0015] Deleting the control parameters and / or the scenario parameters to obtain simplified control parameters and / or the scenario parameters, and controlling the manipulator to execute a control action according to the simplified control parameters and / or the scenario parameters.
[0016] In some embodiments, each of the control parameters and / or the scenario parameters corresponds to a basic action, and controlling the manipulator to execute a control action according to the simplified control parameters and / or the scenario parameters, including:
[0017] Obtaining the basic actions of the manipulator and the execution order of each of the basic actions according to the simplified control parameters and / or the scenario parameters, and controlling the manipulator to sequentially execute the basic actions according to the execution order.
[0018] In some embodiments, the method further includes:
[0019] Training the first neural network based on a reinforcement learning algorithm by using the control strategy, obtaining feedback parameters, and adjusting the first neural network according to the feedback parameters to obtain an updated first neural network.
[0020] In a second aspect, the present application provides a manipulator control device, and the device includes:
[0021] An acquisition module, configured to acquire a first task of the manipulator and determine a first neural network according to the first task;
[0022] A reconstruction module, configured to reconstruct the first task by using the first neural network to obtain a control strategy for the reconstructed first task, where the control strategy includes a series of basic control instructions;
[0023] A control module, configured to obtain control parameters for controlling the manipulator based on the control strategy, and control the manipulator to perform a control action according to the control parameters, where the control parameters include parameters corresponding to the basic control instructions.
[0024] In some embodiments, the obtaining module is further configured to obtain the first task and environmental state information corresponding to the first task, where the environmental state information includes the surrounding environmental information and the self-state information of the manipulator in the simulation environment;
[0025] The reconstruction module is further configured to input the surrounding environmental information, the self-state information, and the first task into the first neural network to obtain a reconstructed control strategy for the first task, where the control strategy includes the basic control instructions and the scene control instructions.
[0026] In some embodiments, the control module is further configured to obtain control parameters and scene parameters for controlling the manipulator based on the control strategy, where the scene parameters are parameters corresponding to the scene control instructions; delete the control parameters and / or the scene parameters to obtain simplified control parameters and / or the scene parameters, and control the manipulator to perform a control action according to the simplified control parameters and / or the scene parameters.
[0027] In some embodiments, the control module is further configured to obtain the basic actions of the manipulator and the execution order of each of the basic actions according to the simplified control parameters and / or the scene parameters, and control the manipulator to sequentially perform the basic actions according to the execution order.
[0028] In some embodiments, the apparatus further includes:
[0029] A model training module, based on a reinforcement learning algorithm, trains the first neural network using the control strategy, obtains feedback parameters, and adjusts the first neural network according to the feedback parameters to obtain an updated first neural network.
[0030] In a third aspect, the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, the manipulator control method described in the first aspect above is implemented.
[0031] In a fourth aspect, the present application provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the manipulator control method described in the first aspect above is implemented.
[0032] Fifth aspect, the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is configured to run programs or instructions to implement the manipulator control method as described in the first aspect.
[0033] Sixth aspect, the present application provides a computer program product, including a computer program, which when executed by a processor implements the manipulator control method as described in the first aspect above.
[0034] One or more of the above technical solutions in the embodiments of the present application have at least the following technical effects:
[0035] The manipulator control method, device and electronic device provided by the embodiments of the present application obtain the first task of the manipulator through a server, select and determine the first neural network according to the first task, and use the first neural network to reconstruct the first task to obtain the control strategy of the reconstructed first task. The control strategy includes a series of basic control instructions. Based on the series of basic control instructions, the control parameters for controlling the manipulator are obtained. Finally, the manipulator is controlled to execute the control action according to the control parameters to complete the corresponding movement. Since a neural network model is used, self-learning of manipulator control can be realized, meeting the flexible control requirements of the manipulator, and being able to make corresponding controls accurately. And since it is completed by the server, the requirement for the computing power of the manipulator can also be reduced, and the applicable scenarios are more extensive.
[0036] The additional aspects and advantages of the present application will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] The above and / or additional aspects and advantages of the present application will become obvious and easy to understand from the description of the embodiments in conjunction with the following drawings, where:
[0038] Figure 1 is a schematic structural diagram of a manipulator control framework provided by an embodiment of the present application;
[0039] Figure 2 is a schematic flowchart of a manipulator control method provided by an embodiment of the present application;
[0040] Figure 3 is a schematic structural diagram of a manipulator control device provided by an embodiment of the present application;
[0041] Figure 4 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] Next, the technical solutions in the embodiments of the present application will be clearly described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present application.
[0043] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same type, and do not limit the number of objects. For example, the first object can be one or multiple. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally means that the related objects before and after are in an "or" relationship.
[0044] First, the overall working process of the manipulator will be described. Please refer to Figure 1 , Figure 1 which shows a schematic structural diagram of a manipulator control framework. The control of the manipulator can be based on artificial intelligence. The artificial intelligence can be installed in the cloud, such as hundreds or thousands of robots being uniformly controlled by a cloud server; the artificial intelligence can also be installed in the machine brain or robot processor owned by the robot, such as the motion control of the manipulator being controlled by a robot chip; in addition, the control of the robotic arm can also be controlled by a control chip installed in the robotic arm. Generally speaking, the control processor of the robot needs to handle the control of the entire body of the robot, and the computing power allocated to the manipulator is limited. The processing ability of the control processor of the robotic arm is weaker than that of the robot's processor.
[0045] The source of artificial intelligence is data. For example, it can be the general process of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, intelligent execution and output. In this process, data undergoes a refinement process of "data - information - knowledge - wisdom". The basic data plus the aggregation of external data, such as the aggregation of multi-modal data such as vision / audition / touch, can provide sufficient data support to help realize the control of the manipulator.
[0046] Next, in conjunction with the accompanying drawings, the manipulator control method, device and electronic device provided by the embodiments of the present application will be described in detail through specific embodiments and their application scenarios. Among them, the manipulator control method can be applied to a terminal, and can be specifically executed by hardware or software in the terminal.
[0048] Such as Figure 2As shown, this manipulator control method is applied to a server, and the method includes the following steps:
[0049] S100. Obtain the first task of the manipulator and determine a first neural network based on the first task.
[0050] It can be understood that the first task in the embodiments of this application can be obtained by the manipulator first and then sent to the server by the manipulator, or can be directly obtained by the server. Specifically, the manipulator / server receives the first task input by the operator, that is, the first task can be input through the receiving interface of the hardware device; it can also be a pre-input task, and the manipulator / server can select the first task from the pre-input tasks; it can also be a pre-set task generation rule, so that the manipulator can generate the first task by itself according to the generation rule. In some embodiments, the generation rule can be self-learned to form tasks with increasing difficulty. As an example, for example, the initial task is for the manipulator to clench its fist, then the first task with increased difficulty compared to the initial task can be for the manipulator to switch between clenching and stretching its fist, switching every 3 seconds; another example: the first task with further increased difficulty can be for the manipulator to pick up a cup from the table. In some embodiments, the first task can be a segment of control instructions, and the control instructions can be action instructions such as pick up, lift, drag, pull, etc.
[0051] It should be noted that the manipulator in the embodiments of this application can be an intelligent robot.
[0052] S200. Reconstruct the first task by using the first neural network to obtain a control strategy for the reconstructed first task, and the control strategy includes a series of basic control instructions.
[0053] In some embodiments, after the server obtains the first task, it determines the first neural network according to the task, and then uses the first neural network to reconstruct the first task to obtain the control strategy corresponding to the reconstructed first task. It can be understood that the first neural network can be understood as being used to obtain the path and result for completing the first task from a knowledge base or a skill library and disassemble them into basic control instructions one by one, so that the control strategy of the first task is composed of these basic control instructions. Each skill in the skill library or knowledge base can be specifically represented as a neural network or can be specifically represented as an operation rule. As an example, for example, the skills in the knowledge base or skill library can specifically be picking up an empty cup with one hand, or picking up a watermelon with both hands, or picking up a heavy object over 10 kg with both hands, etc.
[0054] In some embodiments, the first neural network can be pre-trained using different tasks on its skill library or knowledge base. It should be noted here that before using the first neural network, it may also include the step of obtaining the first neural network. That is, there can be multiple neural networks and skill libraries stored on the server. Then, the first neural network can be a neural network trained based on the simulation environment corresponding to the second task. That is, the first neural network can be one of the at least one pre-trained neural networks that is a mature neural network. Correspondingly, the server can determine the skill library corresponding to the first neural network as the skill library. More specifically, the operator can select the first neural network from the at least one pre-trained neural networks, and then the server obtains the first neural network selected by those skilled in the art. It can also be that the server autonomously selects the first neural network from the at least one pre-trained neural networks, where the semantic information of the first task and the semantic information of the second task are similar. Specifically, the similarity between the semantic information of the first task and the semantic information of the second task can mean using a neural network to obtain the semantic information of the first task and the second task and comparing them to determine the similarity. It can also be that the constraint conditions obtained by decomposing the first task and the second task are similar. As an example, for instance, the result of decomposing the first task is to lift something, and the constraint condition obtained by decomposing the second task is a specific thing that is particularly heavy or particularly light. Then, it can be regarded as the semantic information of the first task being similar to the semantic information of the second task. It can also be that the operating environments of the first task and the second task are similar. As an example, for instance, the first task is to lift something, and the second task is to hold something up. Then, it can be regarded as the semantic information of the first task being similar to the semantic information of the second task, etc. Of course, there can also be other ways to determine the similarity between the semantic information of the first task and the semantic information of the second task. The examples here are only for facilitating the understanding of this solution and do not exhaust all implementation methods.
[0055] In some embodiments, after determining the first task and the neural network type of the first neural network, the server can also initialize a first neural network and initially train a skill library based on the simulation environment corresponding to the first task using a reinforcement learning algorithm. In another implementation, after determining the first task and the neural network type of the first neural network, the server can also initialize a first neural network, and then those skilled in the art can configure at least one skill in the skill library according to the first task, etc. Since the skills in the skill library can be expanded in subsequent steps, the number of skills in the skill library does not need to be particularly large. The skill library consists of some basic actions of an embodied robot, such as drinking water, carrying, walking, etc.
[0056] In some embodiments, step S100 includes: obtaining the first task and the environmental status information corresponding to the first task, where the environmental status information includes the surrounding environmental information of the manipulator and its own status information in the simulation environment.
[0057] In this embodiment, the server inputs the environmental status information related to the first task into the first neural network to obtain the skills selected by the first neural network from the skill library. Specifically, the environmental status information includes the surrounding environmental information of the manipulator and the manipulator's own status information in the simulation environment corresponding to the first task. Specifically, it may include the map information around the intelligent robot, the interaction information of the intelligent robot, the movement information of adjacent intelligent robots, etc. As an example, for instance, when the embodiments of the present application are applied to the field of embodied robot intelligent robotic arms, the environmental status information may include positioning information, map information, interaction information, task information, operation information, etc.
[0058] Step S200 includes: inputting the surrounding environmental information, the own status information, and the first task into the first neural network to obtain a reconstructed control strategy for the first task, where the control strategy includes the basic control instruction and the scenario control instruction.
[0059] In this embodiment, in addition to considering the instruction information of the first task itself, the surrounding environmental information and the own status information of the manipulator are also fully considered. The control strategy reconstructed by considering these information can better fit the operation of the manipulator and achieve accurate motion control of the manipulator. For example: if the first task is to reach a certain destination, without combining the surrounding environmental information, it will move straight directly, and this moving method may cause the manipulator to encounter obstacles during operation and be unable to move. The control strategy after considering the surrounding environmental information will avoid obstacles and adopt instructions of going straight + turning to control the manipulator to complete the action; and considering the own status information of the manipulator can avoid conflicts between two tasks. For example, if the manipulator is currently executing a certain task, it needs to wait for the completion of this task before executing the first task.
[0060] In some embodiments, the environmental status information may be in the form of pictures, sequence data, or other data forms. When inputting this environmental status information into the first neural network, the neural network type of the first neural network can be determined according to the data type of the input data. For example, if the input data is picture data, the first neural network can select a convolutional neural network (CNN); if the input data is sequence data, the first neural network can select a recurrent neural network (RNN), etc. Other situations are not listed one by one here.
[0061] In some embodiments, when the first neural network outputs a control strategy, the first neural network predicts the skills to be selected based on the input first task and the environmental state information corresponding to the first task, and outputs the serial numbers of the corresponding skills, so that the server selects the corresponding skills from the skill library according to the output serial numbers. It can be understood that a unique serial number can be pre-configured for each skill in the skill library. Each skill can be understood as the most basic action control instruction when the manipulator executes an action. For example, moving forward, moving backward, turning left, turning right, etc.
[0062] It can be understood that the process of the first neural network reconstructing the first task is equivalent to disassembling and reorganizing the first task to obtain a series of basic control instructions.
[0063] It can be understood that the process of the first neural network reconstructing the first task will confirm the expected processing result. If it is found that the first task cannot be completed during the reconstruction, the output of specific execution actions will be stopped. For example, if the maximum lifting weight of the manipulator is 50 kg, and in the obtained environmental information, it is found that the weight to be lifted is 100 kg, which has exceeded the capacity of the manipulator, the first neural network needs to output information to the robot or the manipulator to stop the task and the actions of the instructions, and issue corresponding language prompts.
[0064] S300. Obtain control parameters for controlling the manipulator based on the control strategy, so as to control the manipulator to execute control actions according to the control parameters.
[0065] In some embodiments, step S300 includes: obtaining the basic actions for controlling the manipulator and the execution order of each basic action according to the control parameters, and controlling the manipulator to sequentially execute each basic action according to the execution order.
[0066] After obtaining the control strategy of the manipulator, it is equivalent to obtaining a series of basic control instructions of the manipulator. Each control instruction corresponds to a control parameter. By adjusting the control parameters of the manipulator, the manipulator sequentially executes each basic action according to the execution order, realizing the set path and / or action operation.
[0067] Furthermore, in some embodiments, obtain the basic actions of the manipulator and the execution order of each basic action according to the control parameters and / or the scenario parameters, and control the manipulator to sequentially execute the basic actions according to the execution order to complete the corresponding operation. Specifically, when controlling the manipulator to sequentially execute the basic actions according to the execution order, performance parameters such as the time, computing power, and resources used, as well as the energy, force, angle, speed, arm / finger degrees of freedom, etc. used by the basic actions, will also be estimated, and the manipulator will be controlled according to the estimated performance parameters.
[0068] In some embodiments, step S300 includes: obtaining control parameters and scenario parameters for controlling the manipulator based on the control strategy, where the scenario parameters are the parameters corresponding to the scenario control instruction. Deleting some of the control parameters to obtain simplified control parameters, and controlling the manipulator to perform a control action according to the simplified control parameters.
[0069] In some embodiments, step S300 further includes: each of the control parameters and / or the scenario parameters corresponds to a basic action. Obtaining the basic actions of the manipulator and the execution order of each basic action according to the simplified control parameters and / or the scenario parameters, and controlling the manipulator to sequentially perform the basic actions according to the execution order.
[0070] Specifically, when outputting the control strategy, the degrees of freedom of the manipulator need to be considered. That is, when determining the control strategy, judge the relationship between the degrees of freedom required to complete the first task by the control strategy and the degrees of freedom of the manipulator itself. When the degrees of freedom required to complete the first task by the control strategy are greater than the degrees of freedom of the manipulator itself, the control strategy is further simplified; when the degrees of freedom required to complete the first task by the control strategy are less than the degrees of freedom of the manipulator itself, the control strategy for performing the first task is executed. When the degrees of freedom required to complete the first task by the control strategy are less than the degrees of freedom of the manipulator itself, it means that the degrees of freedom of the control strategy for the first task are redundant, which belongs to a better control strategy and can be more learned by the neural network.
[0071] After obtaining the control strategy, the server can also delete some of the control parameters and / or scenario parameters, and control the manipulator to move with the most simplified control parameters and / or scenario parameters to avoid overly complex and unnecessary movements. The scenario parameters can be understood as the scenario elements of the first task. The fewer the scenario elements, the clearer the control strategy obtained based on the simplified scenario elements, and thus the manipulator can be better controlled. Specifically, the scenario elements for reconstructing the first task include goals, results, processes, as well as control methods, matters needing attention, etc. For example, for the drinking action, the goal is to drink water, the result is to drink the water, the process is to lift the cup to drink water, and the control method is to lift. Attention needs to be paid to whether there is water in the cup and what material the cup is made of. The drinking action is stored in the skill library as described above. If there are more scenario parameters for reconstructing the first task, the scenario parameters are simplified. The principle of simplification is to delete useless and redundant parameters, such as the color and liquid of the cup, and only retain key parameters such as the position, size, and mass of the cup. When the number of key parameters is the same as or slightly greater than the number of parameters of the basic action in the skill library, there is no need to simplify. Of course, some of the control parameters can also be deleted. If there are a lifting action and a drinking action, it can be directly simplified to a drinking action.
[0072] It should be noted that the server can be a cloud server, and the cloud server controls the manipulator through communication means.
[0073] The manipulator control method provided by the embodiments of this application obtains the first task of the manipulator through the server, selects and determines the first neural network according to the first task, and uses the first neural network to reconstruct the first task to obtain the control strategy of the reconstructed first task. The control strategy includes a series of basic control instructions. Based on a series of basic control instructions, the control parameters for controlling the manipulator are obtained. Finally, the manipulator is controlled to execute the control action according to the control parameters to complete the corresponding movement. Since this method uses a neural network model, it can achieve self-learning of manipulator control, meet the flexible control requirements of the manipulator, and can accurately make corresponding controls. Moreover, since this method is completed by the server, it can also reduce the computing power requirements for the manipulator and adapt to a wider range of scenarios.
[0074] In some embodiments, based on the reinforcement learning algorithm, the control strategy is used to train the neural network model, the feedback parameters are obtained and adjusted to obtain the trained second neural network.
[0075] Using the reinforcement learning algorithm, the control strategy is memorized and trained to perform fast control execution when encountering a similar scenario next time.
[0076] It should be noted that when the server controls the manipulator to act based on the simulation environment corresponding to the first task, at least one skill selected by the first neural network will be used to complete the first task. After the server controls the intelligent manipulator to execute the first task, it will obtain the data of the intelligent manipulator executing the first task (such as control parameters, scenario parameters), and use the reinforcement learning algorithm to train the parameters of the first neural network again to obtain the updated first neural network.
[0077] More specifically, the server inputs the environmental status information into the first neural network to obtain the skills selected by the first neural network. The environmental status information includes the environmental information around the intelligent manipulator and the self-status information of the intelligent manipulator in the simulation environment corresponding to the first task. The skills for executing the first task selected by the first neural network are used to obtain the basic control instructions. Then, the intelligent manipulator can be controlled in the simulator to execute the operations corresponding to the basic control instructions. During the execution process, the server performs an execution status acquisition operation on the skills selected by the first neural network every preset time interval until the execution status of the skills selected by the first neural network is execution end; the server obtains the data generated during the operation of the intelligent manipulator corresponding to the control instructions. The data includes any one or more of the operation path, operation speed, or operation destination of the intelligent manipulator; the server updates the parameters of the first neural network according to the data using the reinforcement learning algorithm. Among them, the concepts of intelligent manipulator, preset time interval, and execution status have been introduced in detail in the above description and will not be elaborated here. In the embodiments of the present application, the server determines whether the skills selected by the first neural network are executed end by obtaining the execution status of the skills selected by the first neural network every preset time interval. Thus, the server can iteratively update the new skill strategy and the parameters of the new skills in a timely manner according to the operation behavior information of the intelligent manipulator, which is beneficial to improving the accuracy of the training process.
[0078] In the method for controlling a manipulator provided in the embodiments of the present application, the execution subject may be a manipulator control device. In the embodiments of the present application, the manipulator control method is executed by the manipulator control device as an example to illustrate the manipulator control device provided in the embodiments of the present application.
[0079] The embodiments of the present application further provide a manipulator control device, as Figure 3 shown. The manipulator control device includes: an acquisition module 100, configured to acquire the first task of the manipulator and determine the first neural network according to the first task;
[0080] a reconstruction module 200, configured to reconstruct the first task using the first neural network to obtain a control strategy for the reconstructed first task, where the control strategy includes a series of basic control instructions;
[0081] a control module 300, configured to obtain control parameters for controlling the manipulator based on the control strategy, and control the manipulator to execute a control action according to the control parameters, where the control parameters include the parameters corresponding to the basic control instructions.
[0082] The manipulator control device provided by the embodiment of the present application obtains the first task of the manipulator through the server, selects and determines the first neural network according to the first task, and uses the first neural network to reconstruct the first task to obtain the control strategy of the reconstructed first task. The control strategy includes a series of basic control instructions. Based on the series of basic control instructions, the control parameters for controlling the manipulator are obtained. Finally, the manipulator is controlled to execute the control action according to the control parameters to complete the corresponding movement. Since this device uses a neural network model, it can achieve self-learning of manipulator control, meet the flexible control requirements of the manipulator, and be able to make corresponding controls accurately. And since this device is completed by the server, it can also reduce the computing power requirements for the manipulator and adapt to a wider range of scenarios.
[0083] In some embodiments, the obtaining module is further configured to obtain the first task and the environmental state information corresponding to the first task. The environmental state information includes the surrounding environmental information and the self-state information of the manipulator in the simulation environment.
[0084] The reconstruction module is further configured to input the surrounding environmental information, the self-state information, and the first task into the first neural network to obtain the control strategy of the reconstructed first task. The control strategy includes the basic control instructions and the scene control instructions.
[0085] In some embodiments, the control module is further configured to obtain the control parameters and the scene parameters for controlling the manipulator based on the control strategy. The scene parameters are the parameters corresponding to the scene control instructions; delete the control parameters and / or the scene parameters to obtain the simplified control parameters and / or the scene parameters, and control the manipulator to execute the control action according to the simplified control parameters and / or the scene parameters.
[0086] In some embodiments, each of the control parameters and / or the scene parameters corresponds to a basic action. The control module is further configured to obtain the basic actions of the manipulator and the execution order of each basic action according to the simplified control parameters and / or the scene parameters, and control the manipulator to sequentially execute the basic actions according to the execution order.
[0087] In some embodiments, the model training module trains the first neural network based on the reinforcement learning algorithm using the control strategy, obtains the feedback parameters, and adjusts the first neural network according to the feedback parameters to obtain the updated first neural network.
[0088] The manipulator control device in the embodiments of the present application may be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device may be a terminal or other devices other than terminals. Exemplarily, the electronic device may be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, a Mobile Internet Device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. It may also be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiments of the present application do not make specific limitations.
[0089] The manipulator control device in the embodiments of the present application may be a device with an operating system. The operating system may be the Microsoft (Windows) operating system, the Android operating system, the IOS operating system, or other possible operating systems. The embodiments of the present application do not make specific limitations.
[0090] The manipulator control device provided by the embodiments of the present application can implement Figure 2 each process implemented by the method embodiments shown. To avoid repetition, it will not be elaborated here.
[0091] The manipulator control method provided by the embodiments of the present application. The execution subject of this control method may be an electronic device or a functional module or functional entity in the electronic device that can implement this control method. The electronic devices mentioned in the embodiments of the present application include but are not limited to mobile phones, tablet computers, computers, cameras, and wearable devices, etc. Hereinafter, taking the electronic device as the execution subject as an example, the control method provided by the embodiments of the present application will be described.
[0092] In some embodiments, as Figure 4 shown, the embodiments of the present application further provide an electronic device 800, including a processor 801, a memory 802, and a computer program stored on the memory 802 and executable on the processor 801. When the program is executed by the processor 801, it implements each process of the above-mentioned manipulator control method embodiments and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0093] It should be noted that the electronic devices in the embodiments of the present application include the above-mentioned mobile electronic devices and non-mobile electronic devices.
[0094] The embodiments of the present application further provide a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements each process of the above-mentioned embodiment of the manipulator control method and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0095] Among them, the processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes computer-readable storage media, such as computer read-only memory ROM, random access memory RAM, magnetic disks or optical discs, etc.
[0096] The embodiments of the present application further provide a computer program product, including a computer program, which implements the above-mentioned manipulator control method when executed by a processor.
[0097] Among them, the processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes computer-readable storage media, such as computer read-only memory ROM, random access memory RAM, magnetic disks or optical discs, etc.
[0098] The embodiments of the present application further provide a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement each process of the above-mentioned embodiment of the manipulator control method and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0099] It should be understood that the chip mentioned in the embodiments of the present application may also be referred to as a system-on-chip, system chip, chip system, or system-on-chip, etc.
[0100] It should be noted that in this text, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0101] From the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the relevant technology, can be embodied in the form of a computer software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to enable a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present application.
[0102] The embodiments of the present application have been described above with reference to the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Those of ordinary skill in the art, under the inspiration of the present application and without departing from the spirit and scope protected by the claims of the present application, can still make many forms, all of which fall within the protection scope of the present application.
[0103] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "illustrative embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic descriptions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0104] Although embodiments of the present application have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present application. The scope of the present application is defined by the claims and their equivalents.
Claims
1. A robot control method, characterized in that: The method comprises: Acquire a first task of the manipulator, and determine a first neural network according to the first task; Reconstructing the first task using the first neural network to obtain a control strategy for the reconstructed first task, wherein the control strategy includes a series of basic control instructions; Based on the control strategy, control parameters for controlling the manipulator are acquired, and the manipulator is controlled to perform a control action according to the control parameters, wherein the control parameters include parameters corresponding to the basic control instructions.
2. The method according to claim 1, characterized in that The first task of obtaining the manipulator comprises: Acquire the first task and the environment status information corresponding to the first task, wherein the environment status information includes the surrounding environment information and the self status information of the manipulator in the simulation environment; The step of reconstructing the first task using a pre-trained first neural network to obtain a control strategy for the reconstructed first task includes: The surrounding environment information, the self-state information and the first task are input into the first neural network to obtain a reconstructed control strategy for the first task, wherein the control strategy includes the basic control instructions and the scene control instructions.
3. The method according to claim 2, characterized in that The acquiring control parameters for controlling the manipulator based on the control strategy, and controlling the manipulator to perform a control action according to the control parameters, comprises: Acquire control parameters and scene parameters for controlling the manipulator based on the control strategy, wherein the scene parameters are parameters corresponding to the scene control instructions; The control parameters and / or the scene parameters are deleted to obtain simplified control parameters and / or the scene parameters, and the manipulator is controlled to perform a control action according to the simplified control parameters and / or the scene parameters.
4. The method according to claim 3, characterized in that: Each of the control parameters and / or the scene parameters corresponds to a basic action, and controlling the manipulator to perform the control action according to the simplified control parameters and / or the scene parameters includes: The basic actions of the manipulator and the execution order of each basic action are obtained according to the simplified control parameters and / or the scene parameters, and the manipulator is controlled to perform the basic actions in sequence according to the execution order.
5. The method according to any one of claims 1 to 4, characterized in that: The method further comprises: Based on the reinforcement learning algorithm, the first neural network is trained using the control strategy to obtain feedback parameters and the first neural network is adjusted according to the feedback parameters to obtain an updated first neural network.
6. A robot control device, characterized in that: The device comprises: An acquisition module, used for acquiring a first task of the manipulator and determining a first neural network according to the first task; A reconstruction module, used to reconstruct the first task using the first neural network to obtain a control strategy for the reconstructed first task, wherein the control strategy includes a series of basic control instructions; A control module is used to obtain control parameters for controlling the manipulator based on the control strategy, and control the manipulator to perform control actions according to the control parameters, wherein the control parameters include parameters corresponding to the basic control instructions.
7. The device according to claim 6, characterized in that: The acquisition module is further used to acquire the first task and the environmental status information corresponding to the first task, wherein the environmental status information includes the surrounding environment information and the self status information of the manipulator in the simulation environment; The reconstruction module is further used to input the surrounding environment information, the self-state information and the first task into the first neural network to obtain a control strategy for the reconstructed first task, wherein the control strategy includes the basic control instructions and the scene control instructions.
8. The device according to claim 7, characterized in that: The control module is also used to obtain control parameters and scene parameters for controlling the manipulator based on the control strategy, where the scene parameters are parameters corresponding to the scene control instructions; to delete the control parameters and / or the scene parameters to obtain simplified control parameters and / or the scene parameters, and to control the manipulator to perform control actions according to the simplified control parameters and / or the scene parameters.
9. The device according to any one of claims 6 to 8, characterized in that: The device also includes: The model training module trains the first neural network based on the reinforcement learning algorithm and the control strategy, obtains feedback parameters and adjusts the first neural network according to the feedback parameters to obtain an updated first neural network.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the robot control method as described in any one of claims 1-6 is implemented.
Citation Information
Patent Citations
Neural network acquisition method and related equipment
CN112580795A
Mechanical arm manipulation skill learning method based on task decomposition
CN113927593A
Robot skill learning method based on knowledge data driven hierarchical reinforcement learning
CN116306896A
Robot control parameter adjustment method and device, electronic equipment and storage medium
CN118219248A
Manipulator control method and device and electronic equipment
CN119910642A
Cited By
Manipulator control method and device and electronic equipment
CN119910642A
Mechanical hand control method and device and electronic equipment
CN119910642B