Space robot operation method and device for complex tasks

The method equips space robots with a pre-trained skill library and reinforcement learning model to tackle complex tasks by dividing them into sub-skills, ensuring accurate task completion despite uncertain environments.

CN117508670BActive Publication Date: 2025-07-15BEIJING INST OF CONTROL ENG
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311533613.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-16
Publication Date
2025-07-15
Estimated Expiration
2043-11-16

AI Technical Summary

Technical Problem

In the prior art, it is difficult for space robots to complete complex tasks in complex and unstructured environments, especially in configurations and states of unknown targets.

Method used

Through the pre-trained skill library and reinforcement learning model, complex tasks are decomposed into multiple subtasks, and the corresponding subskills are determined according to the environment state, guiding the space robot to complete the operations in turn.

Benefits of technology

Accurate operation of unknown targets in complex environments is achieved, and the ability of space robots to complete complex tasks is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117508670B_ABST
    Figure CN117508670B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of artificial intelligence technology, and particularly relates to a spatial robot operation method and device for complex tasks. The method includes: obtaining a pre-trained skill library, where the skill library includes multiple sub-skills for completing a target task; inputting the skill library and the target task into a pre-trained reinforcement learning model to obtain the sub-skills to be executed at different times; wherein, at different times, the environmental states related to the target task are different, and each sub-skill is respectively used to execute a sub-task in the target task, and the reinforcement learning model is trained based on the skill library and a preset task; guiding the spatial robot to sequentially complete the spatial operations of the corresponding sub-tasks based on each sub-skill until the target task is completed. The present invention can enable the robot to accurately complete complex spatial tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence, and particularly to a method and device for operating a space robot for complex tasks. Background Art

[0002] With the development of technology, space robots are widely used in space operations of on-orbit equipment such as satellites, such as on-orbit maintenance, on-orbit assembly, fuel replenishment, etc. The development of space robots has greatly reduced the operation difficulty of on-orbit activities.

[0003] However, in related technologies, most space robots adopt teleoperation or traditional robot control technologies, which rely heavily on structured operation environments and deterministic operation processes and can only handle simple tasks. For complex tasks such as complex and unstructured space environments (such as the configuration and state of the operation target are not determined in advance), they cannot be well completed.

[0004] Therefore, there is an urgent need for a method and device for operating a space robot for complex tasks to solve the above technical problems. Summary of the Invention

[0005] Embodiments of the present invention provide a method and device for operating a space robot for complex tasks, which can enable the robot to accurately complete complex space tasks.

[0006] In a first aspect, embodiments of the present invention provide a method for operating a space robot for complex tasks, including:

[0007] Obtaining a pre-trained skill library, where the skill library includes multiple sub-skills for completing a target task;

[0008] Inputting the skill library and the target task into a pre-trained reinforcement learning model to obtain sub-skills to be executed at different times; wherein, at different times, the environmental states related to the target task are different, and each sub-skill is respectively used to execute a sub-task in the target task, and the reinforcement learning model is trained based on the skill library and a preset task;

[0009] Guiding the space robot to sequentially complete the space operations of the corresponding sub-tasks based on each sub-skill until the target task is completed.

[0010] In a second aspect, embodiments of the present invention further provide a device for operating a space robot for complex tasks, including:

[0011] An obtaining module, configured to obtain a pre-trained skill library, where the skill library includes multiple sub-skills for completing a target task;

[0012] An input module, configured to input the skill library and the target task into a pre-trained reinforcement learning model to obtain sub-skills to be executed at different times; wherein, at different times, the environmental states related to the target task are different, and each sub-skill is respectively used to execute a sub-task in the target task, and the reinforcement learning model is trained based on the skill library and a preset task;

[0013] A guidance module, configured to guide a space robot to sequentially complete spatial operations of corresponding sub-tasks based on each sub-skill until the target task is completed.

[0014] In a third aspect, an embodiment of the present invention further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the method described in any embodiment of this specification is implemented.

[0015] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed in a computer, the computer is made to execute the method described in any embodiment of this specification.

[0016] An embodiment of the present invention provides a method and device for operating a space robot for complex tasks. The method first trains a skill library so that the skill library includes all sub-skills for completing a target task for subsequent use. By using the skill library and a preset task to train a reinforcement learning model, when facing the target task, it can determine corresponding sub-skills according to different environmental states of the target task, and use the sub-skills to complete corresponding sub-tasks. That is to say, the trained reinforcement learning model can decompose the target task into sub-tasks composed of several sub-skills in the skill library combined in a certain order. Finally, based on each sub-skill, guide the space robot to sequentially complete the spatial operations of the corresponding sub-tasks until the target task is completed. It can be seen that the method of the present invention only needs to input the target task into the pre-trained reinforcement learning model to obtain an accurate combination of sub-skills and guide the robot to accurately complete complex space tasks. Description of the Drawings

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0018] Figure 1 It is a schematic structural diagram of a method for operating a space robot for complex tasks provided by an embodiment of the present invention;

[0019] Figure 2 It is a schematic diagram of a space robot operating system provided by an embodiment of the present invention;

[0020] Figure 3 It is a schematic structural diagram of a reinforcement learning model provided by an embodiment of the present invention;

[0021] Figure 4 It is a hardware architecture diagram of an electronic device provided by an embodiment of the present invention;

[0022] Figure 5 It is a structural diagram of a space robot operating device for complex tasks provided by an embodiment of the present invention. Detailed implementation manners

[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0024] Please refer to Figure 1 , an embodiment of the present invention provides a method for operating a space robot for complex tasks, and the method includes:

[0025] Step 100, obtain a pre-trained skill library, where the skill library includes multiple sub-skills for completing a target task;

[0026] Step 102, input the skill library and the target task into a pre-trained reinforcement learning model to obtain the sub-skills to be executed at different times; wherein, at different times, the environmental states related to the target task are different, and each sub-skill is respectively used to execute a sub-task in the target task, and the reinforcement learning model is trained based on the skill library and a preset task;

[0027] Step 104, based on each sub-skill, guide the space robot to sequentially complete the space operations of the corresponding sub-tasks until the target task is completed.

[0028] In the embodiments of the present invention, first, a skill library is trained so that the skill library includes all sub - skills for completing a target task for subsequent use. By using this skill library and a preset task to train a reinforcement learning model, when it faces the target task, it can determine the corresponding sub - skills according to different environmental states of the target task and use the sub - skills to complete the corresponding subtasks. That is to say, the trained reinforcement learning model can decompose the target task into subtasks composed of a number of sub - skills in the skill library combined in a certain order. Finally, based on each sub - skill, the space robot is guided to sequentially complete the space operations of the corresponding subtasks until the target task is completed. It can be seen that the method of the present invention only needs to input the target task into the pre - trained reinforcement learning model, and then an accurate combination of sub - skills can be obtained, and the robot can be guided to accurately complete complex space tasks.

[0029] The execution manners of the following described Figure 1 each step are shown.

[0030] First, for step 100, a pre - trained skill library is obtained, and the skill library includes a plurality of sub - skills for completing a target task.

[0031] In this step, the target task may be a module replacement task. Taking the replacement of a space module with an opening and closing hatch as an example, the robotic arm of the space robot needs to sequentially complete subtasks such as "tearing the coating - unscrewing the nut - opening the door - pulling out the old module - sending it to the warehouse - grasping the new module - pushing in the new module - closing the door - screwing on the nut". This kind of complex operation task including multiple steps has the characteristics of modularity on the one hand, and each step is clearly defined. On the other hand, some steps are universal and can essentially be learned as a single sub - skill and then combined to achieve more complex operations.

[0032] Based on this, when the space robot completes the above - mentioned tasks, it needs to have basic action strategies such as approaching, grasping, plugging and unplugging, and screwing. By combining each action strategy, a skill library Skill = {k0,..., kn} can be formed, where each sub - skill ki is an operation strategy and can be expressed as a function fi: s → A represented by a deep neural network. Where s is the current environmental state, A is the joint angular velocity of the robot at the next moment, and n is the number of sub - skills. Of course, the skill library can include more skills, which can be determined by the user according to needs to make it universal.

[0033] In addition, when training the skill library, the reinforcement learning method can be used to train each sub-skill separately. Taking the grasping skill as an example, different types and sizes of objects can be used to train the grasping sub-skill, so that the trained grasping sub-skill can grasp objects of any type and any size. Similarly, users can train each of the other sub-skills according to their needs to obtain a trained skill library.

[0034] Then, for step 102, the skill library and the target task are input into a pre-trained reinforcement learning model to obtain the sub-skills that need to be executed at different times; among them, at different times, the environmental states related to the target task are different, and each sub-skill is used to execute a sub-task in the target task. The reinforcement learning model is trained based on the skill library and the preset task.

[0035] In this step, as Figure 2 shown, the reinforcement learning model is applied to the host computer of the space robot operating system. The space robot operating system includes a service satellite and a receptor satellite, and the service satellite and the receptor satellite move around a preset space orbit. The space robot is carried on the service satellite to perform space operations on the objects to be operated on the receptor satellite. The space robot includes a robotic arm, a gripper, and a sensor module. The sensor module includes an RGBD camera, a force sensor, etc., which are used to collect the interaction data between the robotic arm and the environment, filter the collected signals, and send them to the host computer; the host computer runs the reinforcement learning model according to the input environmental perception information, outputs the corresponding sub-skills (i.e., control voltages) to the robotic arm, so that the robotic arm performs the corresponding actions and completes the sub-tasks of on-orbit operations.

[0036] In some embodiments, as Figure 3 shown, the reinforcement learning model includes a perception module, a decomposition module, an execution module, an evaluation module, and a control module;

[0037] The perception module is used to obtain the environmental states related to the target task at different times. The environmental states include the target pose, the joint angles and angular velocities of the robotic arm, and the force and torque information collected by the force sensor at the end of the robotic arm;

[0038] The decomposition module and the evaluation module are used to obtain the corresponding sub-skills from the skill library based on the environmental states at different times;

[0039] The execution module is used to output the joint angular velocity of the robotic arm of the space robot at the next moment based on the environmental state at the current moment obtained by the perception module and the sub-skill corresponding to this environmental state;

[0040] The control module is used to take the joint angular velocity output by the execution module as the desired target, take the joint angles and angular velocities of the space robot at the current moment, and the force and torque information collected by the force sensor at the end of the robotic arm as inputs, and output the joint torques of the robotic arm to control the space robot to complete corresponding space operations.

[0041] In this embodiment, by setting the above-mentioned modules, the reinforcement learning model can decompose complex tasks into combinations of basic skills through self-learning methods to solve the operation problems of various complex tasks in unstructured environments, and has strong versatility.

[0042] In some embodiments, the perception module uses a deep neural network to detect, identify and estimate the pose of the target of interest in the unstructured environment. For example, the perception module takes the images collected by the RGBD camera at the end of the robotic arm of the space robot as inputs, and outputs the target category and the position of the target in the image plane through a preset target detection and recognition algorithm; and outputs the six-degree-of-freedom pose of the target through the PoseCNN pose estimation algorithm. Among them, the preset target detection and recognition algorithms include at least one of the YOLO method, the Faster RCNN method, the Mask RCNN method, and the DeepLab method.

[0043] In addition, the decomposition module takes the current environmental state S as the input, and the output is the skill ki to be executed at the next moment, that is, the subtask. The network structure can adopt network models such as LSTM, RNN, and MLP. The execution module takes the target pose output by the perception module and the joint angles of the robotic arm as inputs and outputs the joint angular velocity of the robotic arm at the next moment. The network structure can adopt network models such as LSTM, RNN, and MLP.

[0044] In some embodiments, the reinforcement learning model is trained by the following method:

[0045] S1. Build a simulation model based on a preset task. The simulation model is communicatively connected to the perception module and the control module;

[0046] S2. Initialize the evaluation module, the decomposition module, and the execution module, and load the skill library;

[0047] S3. Use the perception module to obtain the current environmental state of the simulation model at the current moment; use the decomposition module and the evaluation module to select the sub-skill with the highest probability of success corresponding to the current environmental state from the skill library based on the current environmental state; based on the sub-skill and the current environmental state, use the execution module to output the joint angular velocity of the robotic arm of the space robot at the next moment; take the joint angular velocity as the desired target of the execution module, and take the current environmental state as the input, output the corresponding joint torques of the robotic arm, and control the space robot to complete the sub-space operation corresponding to the sub-skill based on the joint torques;

[0048] S4. Determine whether all spatial operations of the preset task have been completed. If so, execute S5; if not, use the environmental state after the execution of the sub-spatial operation as the current environmental state of the simulation model, and return to execute S3.

[0049] S5. Determine whether the expectation of the preset task is met after all spatial operations are completed. If so, determine that the training of the reinforcement learning model is completed; if not, restore the simulation model to the initial state, and re-execute S3 until the sub-skill sequence planned by the reinforcement learning model can meet the expectation of the preset task.

[0050] Through the above training, the trained reinforcement learning model can decompose complex tasks into combinations of basic skills to solve the operation problems of various complex tasks in an unstructured environment.

[0051] In some embodiments, the decomposition module uses the Monte Carlo tree method to find the sub-skill with the highest winning probability in different environmental states from the skill library.

[0052] In some embodiments, the specific steps of selecting the sub-skill with the highest winning probability in different environmental states from the skill library based on the Monte Carlo tree method are as follows:

[0053] Select the sub-skill ki with the highest winning probability from the current node according to the current environmental state. The formula for calculating the winning probability is:

[0054]

[0055] In the formula, Q(a) is the output of the evaluation module, N(a) is the number of times the sub-skill ki is visited, π(a|s; θ) is the output of the decomposition module π h 's output, and η is a hyperparameter;

[0056] After the selection, reach an unvisited node and a non-leaf node;

[0057] For each expanded leaf node, call the corresponding sub-skill ki for simulation;

[0058] Use the simulation results to backtrack and update the number of wins and the number of visits in each node that led to the simulation results.

[0059] By adopting the Monte Carlo tree method, the sequence combination of basic skills can be solved for the target task.

[0060] As Figure 4 、 Figure 5 shown, the embodiments of the present invention provide a spatial robot operation device for complex tasks. The device embodiments can be implemented by software, or by hardware or a combination of software and hardware. At the hardware level, as Figure 4As shown in the figure, it is a hardware architecture diagram of an electronic device where a space robot operation device for complex tasks provided by an embodiment of the present invention is located. In addition to Figure 4 the shown processor, memory, network interface, and non-volatile memory, the electronic device where the device is located in the embodiment usually may also include other hardware, such as a forwarding chip responsible for processing packets, etc. Taking software implementation as an example, as Figure 5 shown, as a logically meaningful device, it is formed by the CPU of its corresponding electronic device reading the corresponding computer program in the non-volatile memory into the memory and running it.

[0061] A space robot operation device for complex tasks provided by this embodiment includes:

[0062] An acquisition module 500, configured to acquire a pre-trained skill library, where the skill library includes multiple sub-skills for completing a target task;

[0063] An input module 502, configured to input the skill library and the target task into a pre-trained reinforcement learning model to obtain sub-skills that need to be executed at different times; among them, at different times, the environmental states related to the target task are different, and each sub-skill is respectively used to execute a sub-task in the target task, and the reinforcement learning model is trained based on the skill library and a preset task;

[0064] A guidance module 504, configured to guide the space robot to sequentially complete the space operations of the corresponding sub-tasks based on each sub-skill until the target task is completed.

[0065] In the embodiment of the present invention, the acquisition module 500 can be used to execute step 102 in the above method embodiment, the input module 502 can be used to execute step 102 in the above method embodiment, and the guidance module 504 can be used to execute step 104 in the above method embodiment.

[0066] In some embodiments, the reinforcement learning model includes a perception module, a decomposition module, an execution module, an evaluation module, and a control module;

[0067] The perception module is configured to acquire the environmental states related to the target task at different times, and the environmental states include the target pose, the joint angles and angular velocities of the robotic arm, and the force and torque information collected by the force sensors at the end of the robotic arm;

[0068] The decomposition module and the evaluation module are configured to acquire corresponding sub-skills from the skill library based on the environmental states at different times;

[0069] The execution module is configured to output the joint angular velocity of the robotic arm of the space robot at the next moment based on the environmental state at the current moment acquired by the perception module and the sub-skill corresponding to this environmental state;

[0070] The control module is used to take the joint angular velocity output by the execution module as the desired target, take the joint angles and angular velocities of the space robot at the current moment, and the force and torque information collected by the force sensor at the end of the manipulator as inputs, and output the joint torque of the manipulator to control the space robot to complete the corresponding space operations.

[0071] In some embodiments, the reinforcement learning model is trained by the following method:

[0072] S1. Build a simulation model based on a preset task. The simulation model is communicatively connected to the perception module and the control module;

[0073] S2. Initialize the evaluation module, the decomposition module and the execution module, and load the skill library;

[0074] S3. Use the perception module to obtain the current environmental state of the simulation model at the current moment; use the decomposition module and the evaluation module to select the sub-skill with the highest probability of success corresponding to the current environmental state from the skill library based on the current environmental state; based on the sub-skill and the current environmental state, use the execution module to output the joint angular velocity of the manipulator of the space robot at the next moment; take the joint angular velocity as the desired target of the execution module, and take the current environmental state as the input, and output the corresponding joint torque of the manipulator, and control the space robot to complete the sub-space operation corresponding to the sub-skill based on the joint torque;

[0075] S4. Judge whether all the space operations of the preset task are completed. If so, execute S5; if not, take the environmental state after the sub-space operation is completed as the current environmental state of the simulation model, and return to execute S3;

[0076] S5. Judge whether the expectation of the preset task is met after all the space operations are completed. If so, determine that the training of the reinforcement learning model is completed; if not, restore the simulation model to the initial state, and re-execute S3 until the sub-skill sequence planned by the reinforcement learning model can meet the expectation of the preset task.

[0077] In some embodiments, the decomposition module finds the sub-skill with the highest probability of success in different environmental states from the skill library based on the Monte Carlo tree method.

[0078] In some embodiments, the specific steps of selecting the sub-skill with the highest probability of success in different environmental states from the skill library based on the Monte Carlo tree method are as follows:

[0079] Select the sub-skill ki with the highest probability of success from the current node according to the current environmental state. The formula for calculating the probability of success is:

[0080]

[0081] Wherein, Q(a) is the output of the evaluation module, N(a) is the number of times the sub-skill ki is accessed, π(a|s; θ) is the output of the policy module πh, and η is a hyperparameter;

[0082] After the selection ends, an unvisited node that is not a leaf node is reached;

[0083] For each expanded leaf node, the corresponding sub-skill ki is called for simulation;

[0084] The simulation results are used to backtrack and update the number of wins and the number of visits in each node that led to the simulation results.

[0085] In some embodiments, the perception module takes the images collected by the RGBD camera at the end of the robotic arm of the space robot as input, and outputs the target category and the position of the target in the image plane through a preset target detection and recognition algorithm; and outputs the six-degree-of-freedom pose of the target through the PoseCNN pose estimation algorithm.

[0086] In some embodiments, the preset target detection and recognition algorithm includes at least one of the YOLO method, the Faster RCNN method, the MaskRCNN method, and the DeepLab method.

[0087] It can be understood that the structure schematically shown in the embodiments of the present invention does not constitute a specific limitation on a space robot operating device for complex tasks. In other embodiments of the present invention, a space robot operating device for complex tasks may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure can be implemented in hardware, software, or a combination of software and hardware.

[0088] Regarding the information interaction, execution process, etc. between the various modules in the above device, since they are based on the same concept as the method embodiments of the present invention, the specific content can be referred to the description in the method embodiments of the present invention, and will not be elaborated here.

[0089] The embodiments of the present invention also provide an electronic device, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, it implements a space robot operation method for complex tasks in any embodiment of the present invention.

[0090] The embodiments of the present invention also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the processor is enabled to execute a space robot operation method for complex tasks in any embodiment of the present invention.

[0091] Specifically, a system or device equipped with a storage medium can be provided, on which software program code for implementing the functions of any one of the above embodiments is stored, and the computer (or CPU or MPU) of the system or device is caused to read and execute the program code stored in the storage medium.

[0092] In this case, the program code read from the storage medium itself can implement the functions of any one of the above embodiments, so the program code and the storage medium storing the program code constitute a part of the present invention.

[0093] Examples of the storage medium for providing the program code include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Alternatively, the program code can be downloaded from a server computer via a communication network.

[0094] In addition, it should be clear that not only can the functions of any one of the above embodiments be realized by executing the program code read by the computer, but also by causing an operating system or the like operating on the computer based on the instructions of the program code to complete part or all of the actual operations.

[0095] In addition, it can be understood that the program code read from the storage medium is written into the memory provided in the expansion board inserted into the computer or into the memory provided in the expansion module connected to the computer, and then based on the instructions of the program code, the CPU or the like installed on the expansion board or expansion module is caused to execute part or all of the actual operations, so as to realize the functions of any one of the above embodiments.

[0096] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0097] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A spatial robot operation method for complex tasks, characterized in that Including: Obtain a pre-trained skill library, where the skill library includes multiple sub-skills for completing a target task; Input the skill library and the target task into a pre-trained reinforcement learning model to obtain the sub-skills to be executed at different times; wherein, at different times, the environmental states related to the target task are different, and each sub-skill is respectively used to execute a sub-task in the target task, and the reinforcement learning model is trained based on the skill library and a preset task; Based on each sub-skill, guide the space robot to sequentially complete the space operations of the corresponding sub-tasks until the target task is completed; The reinforcement learning model includes a perception module, a decomposition module, an execution module, an evaluation module, and a control module; The perception module is used to obtain the environmental states related to the target task at different times, and the environmental states include the target pose, the joint angles and angular velocities of the robotic arm, and the force and torque information collected by the force sensors at the end of the robotic arm; The decomposition module and the evaluation module are used to obtain the corresponding sub-skills from the skill library based on the environmental states at different times; The execution module is used to output the joint angular velocity of the robotic arm of the space robot at the next moment based on the environmental state at the current moment obtained by the perception module and the sub-skill corresponding to this environmental state; The control module is used to take the joint angular velocity output by the execution module as the desired target, take the joint angles and angular velocities of the space robot at the current moment and the force and torque information collected by the force sensors at the end of the robotic arm as inputs, and output the joint torques of the robotic arm to control the space robot to complete the corresponding space operations; The reinforcement learning model is trained by the following method: S1. Construct a simulation model based on a preset task, and the simulation model is communicatively connected to the perception module and the control module; S2. Initialize the evaluation module, the decomposition module, and the execution module, and load the skill library; S3. Use the perception module to obtain the current environmental state of the simulation model at the current moment; use the decomposition module and the evaluation module to select the sub-skill with the highest success rate corresponding to this current environmental state from the skill library based on this current environmental state; based on this sub-skill and this current environmental state, use the execution module to output the joint angular velocity of the robotic arm of the space robot at the next moment; take this joint angular velocity as the desired target of the execution module, and take this current environmental state as the input, and output the corresponding joint torques of the robotic arm, and control the space robot to complete the sub-space operation corresponding to this sub-skill based on this joint torque; S4. Judge whether all the space operations of the preset task are completed. If so, execute S5; if not, use the environmental state after completing this sub-space operation as the current environmental state of the simulation model, and return to execute S3; S5. Judge whether the expectation of the preset task is met after all the space operations are completed. If so, determine that the training of the reinforcement learning model is completed; if not, restore the simulation model to the initial state, and re-execute S3 until the sub-skill sequence planned by the reinforcement learning model can meet the expectation of the preset task.

2. The method according to claim 1, wherein The decomposition module uses the Monte Carlo tree method to find the sub - skills with the highest odds of success in different environmental states from the skill library.

3. The method according to claim 2, wherein The specific steps for selecting the sub - skills with the highest odds of success in different environmental states from the skill library based on the Monte Carlo tree method are as follows: Select the sub - skill \(k_i\) with the highest odds of success from the current node according to the current environmental state. The formula for calculating the odds of success is: In the formula, \(Q(a)\) is the output of the evaluation module, \(N(a)\) is the number of times the sub - skill \(k_i\) is visited, \(\pi(a|s;\) \(\theta)\) is the output of the policy module \(\pi_h\), and \(\eta\) is a hyper - parameter; After the selection, reach an unvisited node that is not a leaf node; For each expanded leaf node, call the corresponding sub - skill \(k_i\) for simulation; Use the simulation results to back - trace and update the number of wins and the number of visits in each node that led to the simulation results.

4. The method according to claim 1, characterized in that The perception module takes the images collected by the RGBD camera at the end of the manipulator of the space robot as input, and outputs the target category and the position of the target in the image plane through a preset target detection and recognition algorithm; and outputs the six - degree - of - freedom pose of the target through the PoseCNN pose estimation algorithm.

5. The method according to claim 4, wherein The preset target detection and recognition algorithm includes at least one of the YOLO method, the Faster RCNN method, the Mask RCNN method, and the DeepLab method.

6. A space robot operating device for complex tasks, characterized in that, It includes: An acquisition module for acquiring a pre - trained skill library, where the skill library includes multiple sub - skills for completing the target task; An input module for inputting the skill library and the target task into a pre - trained reinforcement learning model to obtain the sub - skills to be executed at different times; where, at different times, the environmental states related to the target task are different, and each sub - skill is respectively used to execute a sub - task in the target task, and the reinforcement learning model is trained based on the skill library and a preset task; A guidance module for guiding the space robot to sequentially complete the spatial operations of the corresponding sub - tasks based on each sub - skill until the target task is completed; The reinforcement learning model includes a perception module, a decomposition module, an execution module, an evaluation module, and a control module; The perception module is used to obtain the environmental states related to the target task at different times, and the environmental states include the target pose, the joint angles and angular velocities of the manipulator, and the force and torque information collected by the force sensors at the end of the manipulator; The decomposition module and the evaluation module are used to obtain the corresponding sub - skills from the skill library based on the environmental states at different times; The execution module is used to output the joint angular velocity of the manipulator of the space robot at the next moment based on the environmental state at the current moment obtained by the perception module and the sub - skill corresponding to this environmental state; The control module is used to take the joint angular velocity output by the execution module as the desired target, and take the joint angles and angular velocities of the space robot at the current moment and the force and torque information collected by the force sensors at the end of the manipulator as input, and output the joint torques of the manipulator to control the space robot to complete the corresponding spatial operations; The reinforcement learning model is trained by the following method: S1. Build a simulation model based on a preset task, where the simulation model is communicatively connected to the perception module and the control module; S2. Initialize the evaluation module, the decomposition module, and the execution module, and load the skill library; S3. Use the perception module to obtain the current environmental state of the simulation model at the current moment; use the decomposition module and the evaluation module to select the sub-skill with the highest odds corresponding to the current environmental state from the skill library based on the current environmental state; based on the sub-skill and the current environmental state, use the execution module to output the joint angular velocity of the manipulator of the space robot at the next moment; use the joint angular velocity as the desired target of the execution module, and take the current environmental state as the input to output the corresponding joint torque of the manipulator, and control the space robot to complete the sub-space operation corresponding to the sub-skill based on the joint torque; S4. Determine whether all space operations of the preset task have been completed. If so, execute S5; if not, use the environmental state after the sub-space operation is completed as the current environmental state of the simulation model, and return to execute S3; S5. Determine whether the expectation of the preset task is met after all space operations are completed. If so, determine that the reinforcement learning model training is completed; if not, restore the simulation model to the initial state, and re-execute S3 until the sub-skill sequence planned by the reinforcement learning model can meet the expectation of the preset task.

7. A computing device, comprising a memory and a processor, where a computer program is stored in the memory, and when the processor executes the computer program, the method according to any one of claims 1-5 is implemented.

8. A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed on a computer, the computer is made to execute the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Off-line reinforcement learning method and system for spatial fine operation

    CN114819179A

  • Framework for focused training of language models and techniques for end-to-end hypertuning of the framework

    US20230098783A1