Human-machine collaborative task planning method based on hierarchical reinforcement learning and large model prompt engineering

By combining hierarchical reinforcement learning and large model prompting engineering in a human-machine collaborative task planning method, the problem of low planning efficiency in existing technologies is solved, and the efficient completion of on-orbit refueling operations by space robots is achieved.

CN120031319BActive Publication Date: 2025-11-21NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510124724.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-11-21
Estimated Expiration
2045-01-26

AI Technical Summary

Technical Problem

Existing human-machine collaborative planning methods are inefficient in terms of task planning, and traditional methods limit the efficiency and flexibility of space robots in performing tasks in complex environments.

Method used

We employ a human-computer collaborative task planning method based on hierarchical reinforcement learning and large model prompting engineering. By designing state matrices and If-then rules, combined with an option-commentator architecture and few-shot thought chain prompts, we leverage large language models to accelerate the learning process and achieve efficient conversion of human knowledge into a rule base.

Benefits of technology

It improves the efficiency of task planning algorithms, reduces the workload of manually designing rules, and enables the robot to efficiently complete on-orbit refueling operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120031319B_ABST
    Figure CN120031319B_ABST
Patent Text Reader

Abstract

The application discloses a human-robot collaborative task planning method based on hierarchical reinforcement learning and large model prompt engineering. The method comprises the following steps: when a robot is in a scene of performing an on-orbit refueling operation task in space, an experimental process of the on-orbit refueling operation task is designed, a state matrix and an If-then rule are designed based on the experimental process, executable basic actions of the robot for the on-orbit refueling operation task are acquired, a state transition model is obtained based on the state matrix and the executable basic actions, a strategy function in an option-critic architecture is acquired and a plurality of rounds are set, the strategy function is initialized to obtain an initialized strategy function, in the first round, the state matrix is initialized, and based on an initial state in the initialized state matrix, the initialized strategy function and the state transition model, a track data is obtained. The application solves the technical problem of low planning efficiency in the prior art human-robot collaborative planning task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of space operation task planning technology, and more specifically, to a human-computer collaborative task planning method based on hierarchical reinforcement learning and large model prompting engineering. Background Technology

[0002] Due to the complexity of the space environment, various spacecraft experience wear and tear and malfunctions during their on-orbit operation, requiring on-orbit maintenance. Employing space robots to perform on-orbit maintenance tasks instead of astronauts can not only reduce the frequency of extravehicular activities (EVAs) for astronauts but also improve the efficiency of on-orbit maintenance operations. Traditional space robot operations mainly employ pre-programming and remote operation modes, which limit the types of tasks that can be performed and result in low efficiency. Intelligent decision-making and planning methods enable space robots to perform tasks more autonomously and flexibly in complex environments and adapt to different task requirements. In intelligent decision-making and planning methods, robot task planning, as a high-level function of the robot system, receives the task objective, generates a sequence of actions from the current state to the target state based on the state transition model, and inputs this sequence into the lower-level path planning and motion planning algorithms to complete the task.

[0003] In task planning algorithms, traditional symbolic language planning methods, while possessing relatively complete modeling capabilities and strong interpretability, suffer from complex modeling and low solution efficiency. Learning-based methods, such as those based on hierarchical reinforcement learning, can adapt to complex task planning using their hierarchical structure, but their network structures are complex, requiring the design of appropriate reward functions. In recent years, methods based on large language models have leveraged their powerful reasoning capabilities to handle diverse task requirements, but the illusion of large models can affect the reliability of planning results. Currently, with the rapid development of artificial intelligence technology, the era of human-machine intelligent collaboration has arrived. Therefore, this invention combines existing task planning methods with artificial intelligence technology to design a human-machine collaborative task planning method for spatial operation tasks. This method utilizes human knowledge to accelerate the learning process of hierarchical reinforcement learning algorithms, thereby improving the planning efficiency of task planning algorithms. Summary of the Invention

[0004] This invention provides a human-computer collaborative task planning method based on hierarchical reinforcement learning and large model prompting engineering, which at least solves the technical problem of low planning efficiency in human-computer collaborative task planning in the prior art.

[0005] According to one aspect of the present invention, a human-machine collaborative task planning method based on hierarchical reinforcement learning and large model prompting engineering is provided. The method may include: designing an experimental procedure for an on-orbit refueling operation task when a robot is performing such an operation in space; designing a state matrix and If-then rules based on the experimental procedure; obtaining the basic executable actions of the robot for the on-orbit refueling operation task; obtaining a state transition model based on the state matrix and the basic executable actions; obtaining a policy function in an option-commentator architecture and setting several rounds; initializing the policy function to obtain an initialized policy function; in the first round, initializing the state matrix; obtaining a trajectory data based on the initial state in the initialized state matrix, the initialized policy function, and the state transition model, wherein a trajectory data includes: several states, actions for each state, and a post-formation parameter for each state. The reward value for each state is determined based on the base reward value and the additional reward value. The additional reward value is determined by a large language model based on a trajectory data, If-then rules, and few-sample thought chain prompts. Based on a trajectory data from the first round, the policy function initialized in the first round is updated using the policy gradient theorem within the options. In the second round, the state matrix is ​​initialized. Based on the initial state in the initialized state matrix, the updated policy function from the first round, and the state transition model, a trajectory data is obtained. This process is repeated iteratively until the last round, at which point the updated policy function for the last round is obtained. Based on the updated policy function for the last round, the robot is controlled to complete the on-orbit refueling operation in space.

[0006] Optionally, the state matrix includes: the robot's position information in space, the toolbox's position information, the satellite's position information for performing the on-orbit refueling operation, the robot's gripping information of the four tools in the toolbox, and the current refueling task status information of the on-orbit refueling operation.

[0007] Optionally, the process of determining the If-then form rule is as follows: based on user experience knowledge, design rules suitable for reinforcement learning methods for the robot's on-orbit refueling operation task in space, and obtain the If-then form rule.

[0008] Optionally, the executable basic actions include: the basic actions of each task in the experimental process of the robot performing the on-orbit refueling operation are: the robot moves, grasps and releases the tool and performs the specific operation task; the executable basic actions are obtained based on the basic actions of each task in the experimental process of the robot performing the on-orbit refueling operation.

[0009] Optionally, in the first round, the state matrix is ​​initialized, and a trajectory data is obtained based on the initial state in the initialized state matrix, the initialized policy function, and the state transition model. This includes: inputting the initial state in the initialized state matrix into the initialized policy function to obtain the initial action corresponding to the initial state; inputting the initial state and the initial action into the state transition model to obtain the first state and the post-shaping reward value of the initial state; inputting the first state into the initialized policy function to obtain the first action corresponding to the first state; inputting the first state and the first action into the state transition model to obtain the second state and the post-shaping reward value of the first state, and repeating the iteration to obtain a trajectory data.

[0010] Optionally, the post-shaping reward value for each state is determined based on the base reward value and the additional reward value, including: the sum of the base reward value and the additional reward value for each state is used to determine the post-shaping reward value for each state.

[0011] Optionally, the additional reward value is determined based on a large language model determined by a trajectory data, If-then rules, and few-shot thought chain cues, including: constructing an additional reward function; wherein the additional reward function contains a target rule base, wherein the target rule base is determined based on a large language model determined by If-then rules and few-shot thought chain cues corresponding to a trajectory data; and inputting each state-action pair into the additional reward function to obtain the additional reward value corresponding to each state-action pair.

[0012] Optionally, the target rule base is determined based on a large language model determined by If-then rules and a few-shot thought chain prompt corresponding to a trajectory data, including: when a trajectory data is obtained, inputting the user knowledge corresponding to this trajectory data into the large language model determined by the few-shot thought chain prompt in the form of human knowledge in natural language, outputting rules about this trajectory data, and adding them to the rule base of If-then rules.

[0013] The beneficial effects of this invention are:

[0014] (1) Design a hierarchical reinforcement learning method that integrates rule bases to accelerate network training speed;

[0015] (2) Based on the few-sample thinking chain prompting project, the reasoning ability of the large language model is used to realize the efficient conversion of human knowledge into rules in the rule base, reducing the workload of manually designing rules;

[0016] (3) Utilize human knowledge to accelerate the learning process of the option-critic method, realize human-machine collaborative planning, and improve task planning efficiency. Attached Figure Description

[0017] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:

[0018] Figure 1 This is a flowchart of a human-computer collaborative task planning method based on hierarchical reinforcement learning and large model prompting engineering according to an embodiment of the present invention;

[0019] Figure 2 This is a schematic diagram of the rule-based option-commentator method model structure according to an embodiment of the present invention;

[0020] Figure 3 This is a schematic diagram of the structure of the efficient human knowledge extraction method according to an embodiment of the present invention;

[0021] Figure 4 This is a schematic diagram of a human-machine collaborative task planning method according to an embodiment of the present invention;

[0022] Figure 5 This is a schematic diagram showing the comparison results of different task planning methods according to embodiments of the present invention. Detailed Implementation

[0023] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0024] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0025] Example 1

[0026] According to embodiments of the present invention, a human-computer collaborative task planning method based on hierarchical reinforcement learning and large model prompting engineering is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system containing at least one set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0027] Figure 1 This is a flowchart of a human-computer collaborative task planning method based on hierarchical reinforcement learning and large model prompting engineering according to an embodiment of the present invention, such as... Figure 1 As shown, the method may include the following steps:

[0028] Step S101: When the robot is performing an on-orbit refueling operation in space, design an experimental procedure for the on-orbit refueling operation. Based on the experimental procedure, design a state matrix and If-then rules.

[0029] In the technical solution provided by step S101 of the present invention, when the robot performs an on-orbit refueling operation in space, the experimental process of the on-orbit refueling operation is designed as follows: cutting and moving the protective layer, unscrewing the fuel tank cap, inserting the fuel pipeline into the fuel tank valve, and delivering methane. The expression of the state matrix is ​​designed as follows:

[0030] (1)

[0031] in, , , This indicates the robot's location information. , , This indicates the location information of the toolbox. , , The first row represents the location information of the service satellite performing the on-orbit refueling operation; the second row of the state matrix represents the clamping information of the four tools, with 0 indicating no clamping and 1 indicating clamping; refuel_state represents the current on-orbit refueling operation status, with 0 indicating the initial state, 1 indicating completion of the cutting and moving of the protective layer, 2 indicating completion of the unscrewing of the fuel tank cap, 3 indicating completion of the insertion of the fuel pipeline into the fuel tank valve, and 4 indicating completion of the methane delivery task. The gripping and holding tool corresponding to the robot's position information.

[0032] Step S102: Obtain the basic executable actions of the robot for the on-orbit refueling operation task, and obtain the state transition model based on the state matrix and the basic executable actions.

[0033] In the technical solution provided by step S102 of the present invention, the robot can perform 14 basic actions, including robot movement, grasping and releasing tools, and performing specific operation tasks. Based on the state matrix and the basic actions that can be performed, a state transition model is obtained.

[0034] Step S103: Obtain the strategy function in the options-critic architecture and set the number of rounds.

[0035] In the technical solution provided by step S103 of the present invention, there are three main value functions in the option-commentator architecture, namely:

[0036] (1) Option-value function, which represents the state. Choose the option below The total revenue that can be generated is expressed as follows:

[0037] (2)

[0038] in, Indicates the robot's current state. This indicates the option selected by the robot. This indicates the action the robot is performing in its current state. This represents the current policy (i.e., the policy function in the options-commentator architecture). Let be the action value function, as shown in equation (3).

[0039] (2) Action-value function, which represents the robot's current state. And selected the option Under the premise of taking action The total revenue that can be generated is expressed as follows:

[0040] (3)

[0041] in, Indicates the robot's current state Next action The resulting reward or return, As a discount factor, This indicates the robot's current state. Next action Transition to the next state from the current state The state transition probability, The option value function represents the next state after reaching the current state, as shown in equation (4).

[0042] (3) The option-value function upon arrival in the state, which represents the value of the option when using the state. When, the state is reached The total revenue generated thereafter is expressed as follows:

[0043] (4)

[0044] in, Indicates options The next state after the current state The termination probability (termination probability function) is given below. This represents the option value function. This represents the state value function.

[0045] To enable learning within the option-commentator framework, we need to introduce an option-based policy gradient theorem, which consists of two main parts: the policy gradient theorem within the option and the policy gradient theorem for the termination function.

[0046] (1) Policy gradient theorem within options: Given a set of options for a Markov system, the policy gradient within the options is given by the gradient of the parameters. Intercontinental, expected discount return relative to and initial conditions The gradient is:

[0047] (5)

[0048] in, From The initial state options along the trajectory are discounted and weighted. , yes The parameters.

[0049] (2) Policy gradient theorem for termination function: Given a set of Markov options, the random termination function has a gradient gradient in its parameters. Intercontinental, expected discount return relative to and initial conditions The gradient is:

[0050] (6)

[0051] in, The dominant function among the options, i.e. , yes The parameters.

[0052] Step S104: Initialize the policy function to obtain the initialized policy function.

[0053] In the technical solution provided by step S104 of the present invention, the strategy function is initialized to obtain the initialized strategy function.

[0054] Step S105: In the first round, initialize the state matrix. Based on the initial state in the initialized state matrix, the initialized policy function, and the state transition model, obtain a trajectory data. The trajectory data includes: several states, the action of each state, and the post-shaping reward value of each state. The post-shaping reward value of each state is determined based on the basic reward value and the additional reward value. The additional reward value is determined based on the trajectory data, the If-then form rule, and the large language model determined by the few-sample thought chain prompt.

[0055] In the technical solution provided by step S105 of the present invention, in the first round, the state matrix is ​​initialized. Based on the initial state in the initialized state matrix, the initialized policy function, and the state transition model, a trajectory data is obtained. The trajectory data includes: several states, the action of each state, and the post-shaping reward value of each state. The post-shaping reward value of each state is determined based on the basic reward value and the additional reward value. The additional reward value is determined based on a trajectory data, If-then rules, and a large language model determined by a few-sample thought chain prompt. The termination probability function is initialized at the same time as the policy function is initialized.

[0056] Step S106: Based on a trajectory data point from the first round, update the policy function initialized in the first round using the policy gradient theorem within the options.

[0057] In the technical solution provided by step S106 of the present invention, based on a trajectory data in the first round, the policy function initialized in the first round is updated using the policy gradient theorem within the options. At the same time as updating the policy function initialized in the first round using the policy gradient theorem within the options, the termination probability function in the first round is also updated using the policy gradient theorem of the termination function. The purpose of updating the termination probability function in the first round is to obtain a better policy function.

[0058] Step S107: In the second round, initialize the state matrix. Based on the initial state in the initialized state matrix, the updated policy function in the first round, and the state transition model, obtain a trajectory data. Repeat the iteration until the last round, and obtain the updated policy function for the last round.

[0059] In the technical solution provided by step S107 of the present invention, in the second round, the state matrix is ​​initialized, and a trajectory data is obtained based on the initial state in the initialized state matrix, the updated policy function in the first round, and the state transition model. The process is repeated until the last round is reached, at which point the updated policy function for the last round is obtained.

[0060] Step S108: Based on the updated strategy function from the last round, control the robot to complete the on-orbit refueling operation in space.

[0061] In the technical solution provided by step S108 of the present invention, the robot is controlled to complete the on-orbit refueling operation in space according to the strategy function updated in the last round.

[0062] The method described in this embodiment will be further described below.

[0063] As an optional embodiment, in step S101, the state matrix includes: the robot's position information in space, the toolbox's position information, the satellite's position information for performing the on-orbit refueling operation, the robot's gripping information of the four tools in the toolbox, and the current refueling task status information of the on-orbit refueling operation.

[0064] In this embodiment, as shown in formula (1), the state matrix includes: the robot's position information in space, the toolbox's position information, the satellite's position information for performing the on-orbit refueling operation, the robot's gripping information of the four tools in the toolbox, and the current refueling task status information of the on-orbit refueling operation.

[0065] As an optional embodiment, step S101, the process of determining the If-then form rule is as follows: based on user experience knowledge, design rules suitable for reinforcement learning methods for the robot's on-orbit refueling operation task in space, and obtain the If-then form rule.

[0066] In this embodiment, rules suitable for reinforcement learning methods are designed for the robot's on-orbit refueling operation task in space based on user experience and knowledge, resulting in If-then form rules.

[0067] As an optional embodiment, step S102, the executable basic actions include: the basic actions of each task in the experimental process of the robot performing the on-orbit refueling operation are: the robot moves, grasps and releases the tool and performs the specific operation task; based on the basic actions of each task in the experimental process of the robot performing the on-orbit refueling operation, the executable basic actions are obtained.

[0068] In this embodiment, all the basic actions of the experimental procedure for the robot to perform the on-orbit refueling operation are as follows: move to the toolbox, move to the servicing satellite, clamp the cutting and moving protective layer tool, clamp the unscrewing fuel tank cap tool, clamp the insert fuel tank valve tool, clamp the refueling tool, release the cutting and moving protective layer tool, release the unscrewing fuel tank cap tool, release the insert fuel tank valve tool, release the refueling tool, perform the cutting and moving protective layer operation, perform the unscrewing fuel tank cap operation, perform the inserting fuel pipeline into the fuel tank valve operation, and perform the refueling operation.

[0069] As an optional embodiment, step S105, in the first round, initializes the state matrix, and obtains a trajectory data based on the initial state in the initialized state matrix, the initialized policy function, and the state transition model, including: inputting the initial state in the initialized state matrix into the initialized policy function to obtain the initial action corresponding to the initial state; inputting the initial state and the initial action into the state transition model to obtain the first state and the post-shaping reward value of the initial state; inputting the first state into the initialized policy function to obtain the first action corresponding to the first state; inputting the first state and the first action into the state transition model to obtain the second state and the post-shaping reward value of the first state, and repeating the iteration to obtain a trajectory data.

[0070] In this embodiment, the initial state in the initialized state matrix is ​​input into the initialized policy function to obtain the initial action corresponding to the initial state; the initial state and the initial action are input into the state transition model to obtain the first state and the post-shaping reward value of the initial state; the first state is input into the initialized policy function to obtain the first action corresponding to the first state; the first state and the first action are input into the state transition model to obtain the second state and the post-shaping reward value of the first state, and the process is repeated iteratively to obtain a trajectory data.

[0071] As an optional embodiment, step S105, the post-shaping reward value of each state is determined based on the basic reward value and the additional reward value, including: the sum of the basic reward value and the additional reward value of each state is used to determine the post-shaping reward value of each state.

[0072] In this embodiment, as shown in formula (7), the sum of the base reward value and the additional reward value of each state is determined as the post-shaping reward value of each state.

[0073] As an optional embodiment, step S105, where the additional reward value is determined based on a trajectory data, If-then rules, and a few-sample thought chain cues, includes: constructing an additional reward function; wherein the additional reward function contains a target rule base, wherein the target rule base is determined based on If-then rules and a few-sample thought chain cues corresponding to a trajectory data; and inputting each state-action pair into the additional reward function to obtain the additional reward value corresponding to each state-action pair.

[0074] In this embodiment, Figure 2 This is a schematic diagram of the rule-based option-commentator method model structure according to an embodiment of the present invention. The designed rule-based option-commentator method model structure is as follows: Figure 2 As shown. Furthermore, the rules are integrated into the reinforcement learning algorithm using reward shaping. The basic idea of ​​reward shaping is shown in formula (7), where, Based on the base reward value, To add bonus value, Additional reward value designed to represent the reward value after shaping. The specific form is shown in formula (8).

[0075] (7)

[0076] (8)

[0077] in, For the target rule base, This is an additional reward function.

[0078] With the rise of large language models such as ChatGPT and LLaMA, prompt engineering, as a mechanism for fine-tuning model output through carefully designed instructions, has achieved good results in tasks such as text summarization, information extraction, arithmetic question answering, and code generation. Inspired by this, to reduce the workload of manually designing rules, a few-shot chain-of-thought prompt is introduced. This involves designing a human-computer interaction interface based on natural language input to extract and transform human knowledge into reinforcement learning rules. The few-shot chain-of-thought prompt combines few-shot prompts and chain-of-thought prompts. By providing a small number of examples to the large language model and explaining and demonstrating the reasoning process within those examples, the large language model will also perform corresponding reasoning processes when answering. This form of prompt often leads to more accurate results. The few-shot chain-of-thought prompt enhances the reasoning ability of the large language model by adding intermediate reasoning steps to the prompt and has good generalization and interpretability.

[0079] Figure 3 This is a schematic diagram of the structure of the efficient human knowledge extraction method according to an embodiment of the present invention; as shown below. Figure 3 As shown, firstly, the task-related background information, task requirements, and example inputs and outputs are input into the large language model as basic information for prompts. Next, humans can express the suggested knowledge for the task in natural language, which is concise and clear, without needing to consider the specific form of the rules. Finally, the large language model understands and reasons about the context information of the task, converts the knowledge represented by natural language into reinforcement learning rules, and can convert the rules into executable Python code.

[0080] As an optional implementation method, the target rule base is determined based on a large language model determined by If-then rules and a few-shot thought chain prompt corresponding to a trajectory data. This includes: when a trajectory data is obtained, inputting the user knowledge corresponding to this trajectory data into the large language model determined by the few-shot thought chain prompt in the form of human knowledge in natural language, outputting rules about this trajectory data, and adding them to the rule base of If-then rules.

[0081] In this embodiment, human-machine collaboration involves forming a team of humans and machines, integrating human intelligence and artificial intelligence to solve complex problems. By combining a rule-based option-commentator method with an efficient human knowledge extraction method based on large model hinting engineering, both human intelligence and artificial intelligence can be utilized simultaneously to achieve human-machine collaborative planning, which can further improve the efficiency of task planning algorithms.

[0082] Based on the above ideas, a human-machine collaborative task planning method is proposed. Figure 4 This is a schematic diagram of a human-machine collaborative task planning method according to an embodiment of the present invention; as shown. Figure 4 As shown, specifically, when planning any task, the rule-based option-critic method, after a certain number of training rounds, outputs an imperfect intermediate planning result. Humans can then provide suggestions in natural language based on this result and relevant task information, such as suggesting an action (state-action pair) for the agent in any given state. Subsequently, human knowledge is efficiently extracted and integrated into the reinforcement learning algorithm, accelerating the learning process, improving planning efficiency, and achieving human-machine collaborative planning. Furthermore, for deterministic environment planning problems, corresponding replanning methods can be designed to further refine the planning results, obtaining a sequence of actions that the robot can execute, and adding it to the rule base of If-then rules.

[0083] In this invention, the selected task scenario is an on-orbit refueling operation of a robot. After training, the resulting reward curve is as follows: Figure 5 As shown, Figure 5 This is a schematic diagram comparing the results of different task planning methods according to embodiments of the present invention, such as... Figure 5 As shown, the reward value of the human-machine collaborative task planning method is much better than that of the original option-critic method. The planning results obtained for the on-orbit refueling operation task are shown in Table 1.

[0084] Table 1. Results of On-orbit Refueling Operation Planning

[0085]

[0086] In this embodiment of the invention, an experimental procedure for an on-orbit refueling operation is designed based on a scenario where a robot performs an on-orbit refueling operation in space. Based on the experimental procedure, a state matrix and If-then rules are designed. The basic executable actions of the robot for the on-orbit refueling operation are obtained. Based on the state matrix and the basic executable actions, a state transition model is obtained. The policy function in the option-commentator architecture is obtained, and several rounds are set. The policy function is initialized, resulting in an initialized policy function. In the first round, the state matrix is ​​initialized. Based on the initial state in the initialized state matrix, the initialized policy function, and the state transition model, a trajectory data is obtained. A trajectory data includes several states, the action of each state, and the post-shaping reward value for each state. The post-shaping reward value for each state is determined based on a base reward value and an additional reward value. The additional reward value is based on a trajectory data... The large language model is determined by if-then rules and few-sample thought chain prompts. Based on a trajectory data from the first round, the policy function initialized in the first round is updated using the policy gradient theorem within the options. In the second round, the state matrix is ​​initialized. Based on the initial state in the initialized state matrix, the updated policy function from the first round, and the state transition model, a trajectory data is obtained. This process is repeated iteratively until the last round, at which point the updated policy function for the last round is obtained. Based on the updated policy function for the last round, the robot is controlled to complete the on-orbit refueling operation in space. This solves the technical problem of low planning efficiency in existing human-machine collaborative planning tasks, achieving the technical effect of designing a human-machine collaborative task planning method for space operation tasks. It utilizes human knowledge to accelerate the learning process of hierarchical reinforcement learning algorithms, thereby improving the planning efficiency of task planning algorithms.

[0087] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0088] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0089] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some interfaces; indirect couplings or communication connections between units or modules may be electrical or other forms.

[0090] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0091] Furthermore, the functional units in the various embodiments of the present invention can be integrated into a first processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0092] The above are merely preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A human-computer collaborative task planning method based on hierarchical reinforcement learning and large model prompting engineering, characterized in that, include: In the scenario where a robot performs on-orbit refueling operations in space, design an experimental procedure for the on-orbit refueling operation. Based on the experimental procedure, design a state matrix and If-then rules. The basic executable actions of the robot for the on-orbit refueling operation are obtained, and a state transition model is obtained based on the state matrix and the basic executable actions. The state matrix includes: the robot's position information in space, the toolbox's position information, the satellite's position information for performing the on-orbit refueling operation, the robot's gripping information for the four tools in the toolbox, and the current refueling task status information for the on-orbit refueling operation. The process of determining the If-then form of the rule is as follows: Based on user experience knowledge, design rules suitable for reinforcement learning methods for the robot's on-orbit refueling operation task in space, and obtain the If-then form of the rule; Get the options - the strategy function in the critic architecture and set a number of rounds; Initialize the policy function to obtain the initialized policy function; In the first round, the state matrix is ​​initialized. Based on the initial state in the initialized state matrix, the initialized policy function, and the state transition model, a trajectory data is obtained. The trajectory data includes: several states, the action of each state, and the post-shaping reward value of each state. The post-shaping reward value of each state is determined based on the basic reward value and the additional reward value. The additional reward value is determined based on a trajectory data, If-then rules, and a large language model determined by few-sample thought chain prompts. The process of obtaining a trajectory data includes: Input the initial state from the initialized state matrix into the initialized policy function to obtain the initial action corresponding to the initial state; Input the initial state and initial action into the state transition model to obtain the reward value after shaping the first state and the initial state; Input the first state into the initialized policy function to obtain the first action corresponding to the first state; The first state and the first action are input into the state transition model to obtain the second state and the reward value after shaping the first state. This process is repeated iteratively to obtain a trajectory data. Based on a trajectory data from the first round, the policy function initialized in the first round is updated using the policy gradient theorem within the options; In the second round, the state matrix is ​​initialized. Based on the initial state in the initialized state matrix, the updated policy function in the first round, and the state transition model, a trajectory data is obtained. This process is repeated until the last round, at which point the updated policy function for the last round is obtained. Based on the updated strategy function from the last round, the robot is controlled to complete the on-orbit refueling operation in space.

2. The method according to claim 1, characterized in that, The executable basic actions include: In the experimental procedure for the robot to perform on-orbit refueling operations, the basic actions of each task are: the robot moves, grasps and releases the tool, and performs the specific operation task; Based on the experimental process of the robot performing on-orbit refueling operations, the basic actions of each task are obtained, resulting in executable basic actions.

3. The method according to claim 1, characterized in that, The post-shaping reward value for each state is determined based on a base reward value and an additional reward value, including: The sum of the base reward value and the additional reward value for each state is used to determine the post-shaping reward value for each state.

4. The method according to claim 3, characterized in that, The additional reward value is determined based on a large language model using trajectory data, If-then rules, and few-sample thought chain cues, including: Construct an additional reward function; wherein the additional reward function contains a target rule base, wherein the target rule base is determined based on a large language model determined by if-then rules and a few-sample thought chain prompt corresponding to a trajectory data; Each state-action pair is input into an additional reward function to obtain the additional reward value corresponding to each state-action pair.

5. The method according to claim 4, characterized in that, The target rule base is determined based on a large language model using If-then rules and few-sample thought chain hints corresponding to a trajectory data point, including: When a trajectory data is obtained, the user knowledge corresponding to this trajectory data is input into the large language model based on few-sample thought chain prompts in the form of human knowledge in natural language. The model outputs rules about this trajectory data and adds them to the rule base of If-then rules.

Citation Information

Patent Citations

  • Rule data dual-drive robot complex operation process man-machine hybrid decision-making method

    CN114662404A

  • Satellite group on-orbit refueling task planning method based on Lanbert transfer

    CN118426450A