A patrol mission planning method based on HTN
Through the HTN-based patrol task planning method, a domain model is constructed and a planning scheme is generated using human-computer interaction iteration, which solves the problem of inefficient complex task planning in the existing technology, and achieves more efficient task planning and better constraint processing.
Patent Information
- Application Number
- CN202210629468.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-01
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2042-06-01
AI Technical Summary
When the prior art deals with complex patrol task planning, it is difficult to effectively identify task requirements and quickly deal with a large number of complex constraints, resulting in inefficient planning.
The HTN-based patrol task planning method is adopted to analyze scientific detection tasks, build domain models, and use human-computer interaction iteration to generate planning solutions to solve the limitations of complex planning needs.
More efficient task planning is achieved, which can quickly identify task requirements, handle complex constraints, and improve the efficiency and accuracy of inspector task planning.
Smart Images

Figure CN114970360B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence automatic planning, and more specifically, to a mission planning method for scientific exploration of the surface of extraterrestrial bodies by rovers in the fields of national defense and space technology. Background Art
[0002] Mission intelligent planning technology originated from the research of state space search, theorem proving and control theory, as well as the actual needs of robotics, scheduling and other fields. Relevant planning technologies such as full order planning, partial order planning and planning graphs have been continuously developed in the process of solving problems. Patrol mission planning technology is a typical application of mission intelligent planning technology in the field of national defense and space technology. The mission planning problem of patrollers is to plan the behavior sequence of patrollers on a given driving path (such as charging, taking pictures, etc.) so that the patrollers can reach the target state as required, and the whole process meets various operational constraints and specified constraints. The aerospace measurement and control mode is complex and the emergency control requirements are diverse. Experts need to participate to solve the constraint conflicts encountered in the planning process. The system needs to add a new type of conditional constraints to represent the human collaborative process, which makes the mission planning problem solving and system design more complicated.
[0003] For the patroller task planning problem, the applicability of the HTN (Hierarchical Task Network) planning method is mainly reflected in the following aspects: First, HTN planning is based on domain knowledge reasoning, and the complex tasks are decomposed layer by layer, which is suitable for the hierarchical characteristics of high-level task objectives in patroller task planning; Second, the method of dealing with complex problems based on the idea of hierarchical decomposition enables HTN planning to solve large-scale problems, and has great application potential in the field of patroller task planning, which is growing in scale and number; Third, the HTN planning method has a strong ability to express domain knowledge, and can fully express the conditions in the task decomposition process and the changes in the system state caused by the execution of actions; Fourth, the clear logical reasoning process of HTN planning facilitates the recording of key nodes of task decomposition, which is conducive to solution repair in a dynamic execution environment, and is also conducive to collaborative planning in a distributed environment.
[0004] The field of mission planning has complex characteristics in terms of temporality, resources, etc. Different mission objectives have different task attributes, and the constraints that need to be considered when planning missions are also different, including task priority constraints, resource constraints, time window constraints, etc. General HTN planning usually emphasizes the rapid generation of feasible solutions, but patrol mission planning often has many tasks and few resources. On the basis of feasible solutions, it is also required to make full use of payloads and other related subsystems to maximize mission returns. Therefore, how to identify mission requirements, quickly and effectively handle a large number of complex constraints during the reasoning process, and improve planning efficiency are the biggest challenges that patrol mission planning needs to face.
[0005] In practical problems, a method of full manual control can be adopted. Ground personnel use the payload camera to obtain small-frame images to identify obstacles, determine obstacle distances, determine route traversability, and control the movement of the rover. This type of method is close to the way humans work. The disadvantage is that the entire planning process is carried out at a single granularity, the planning time is long, and the planning system has low intelligence and poor fault tolerance. At the research level, domestic and foreign scholars have proposed a method for autonomous control of the rover. The domain model and instance problems are constructed from the specific problem domain. The planner completes logical reasoning and conflict resolution and autonomously completes the patrol navigation. This method is mostly aimed at a local problem, and cannot reflect all the characteristics of the rover's ontology engineering, cannot effectively and completely describe the ontology knowledge of aerospace engineering, and cannot reflect the hierarchical characteristics of mission planning. Summary of the invention
[0006] In view of the above defects or improvement needs of the prior art, the present invention provides a patroller task planning method based on HTN (Hierarchical Task Network) planning, which determines the initial state of the task and the starting point of the task temporal network by analyzing the scientific detection mission of the patroller, constructs a domain model, and generates a planning scheme through a human-computer interaction iterative method to solve the technical problems faced by the prior art in dealing with complex planning requirements.
[0007] The technical solution of the present invention is as follows:
[0008] A patrol mission planning method based on HTN includes the following steps:
[0009] 1. Analysis of the rover's scientific exploration mission P= <D,I 0 ,TCN>, set domain knowledge and build domain model D=<O,M,δ> ;
[0010] Among them, D represents the domain model, I 0 represents the initial state and TCN represents the task temporal network; O represents the operator set, M represents the method set, and δ represents the state transfer function; the domain knowledge includes the detection point accessibility evaluation method, the measurement and control tracking condition calculation method, the solar altitude angle / azimuth prediction method, the energy consumption estimation strategy under different road conditions, and the scientific detection needs;
[0011] Second, according to the initial state of the patroller and external constraints, the overall planning of the long-term mission is decomposed into periodic operations between multiple target detection points;
[0012] 3. According to the domain model and the initial state of the job, the job is mapped to the operator set O, the method set M and the state transition function δ, the instantiation of the mathematical model is completed, and the behavior sequence of the patroller between the start and end detection points is determined;
[0013] 4. Expand the behavior sequence into atomic actions and construct the temporal constraint network TCN<TN,C> , submitted to the planner for instruction expansion and logic verification; TN represents a node, which is composed of atomic actions that constitute the working mode, and its attributes include the task execution time T and the current position coordinates N of the patroller; C represents an edge in the network graph, that is, a set of external constraints;
[0014] 5. Verify the correctness of the instruction logic through a human-machine iterative process. If errors are found in the logic verification, return to step 3 after manual confirmation, re-plan, and adjust the behavior sequence to eliminate defect factors; if the logic verification is correct, obtain the patroller task planning results.
[0015] Furthermore, the step 1 further comprises:
[0016] (1) According to the task configuration file, define the "working mode" to form an operator set O; each working mode contains a default action sequence, and also defines some attributes or constraints related to planning calculations; define the operator set O = {! perceive,! move,! detect,! charge,! sleep}, the operator! perceive instantiation represents the perception mode, that is, the rover obtains navigation information data and transmits the navigation information data to the ground control center; the operator! move instantiation represents the movement mode, that is, the rover receives the command from the ground control center and reaches the target position; the operator! detect instantiation represents the detection mode, that is, the payload equipment carried by the rover is powered on, obtains scientific detection data, and transmits the data to the ground control center within the communication window; the operator! charge instantiation represents the charging mode, that is, the rover adjusts the solar wing, achieves solar orientation according to regulations, and remains stationary, and the battery pack starts to charge; the operator! sleep instantiation represents the sleep mode, and other equipment of the rover is completely powered off and does not work;
[0017] (2) Using domain knowledge, construct a method set M = {m1, m2, m3, m4, m5, m6, m7, m8, m9}, where m1 indicates sleep wake-up; m2 indicates charging completion; m3 indicates power verification; m4 indicates load opening; m5 indicates load closing; m6 indicates movement start; m7 indicates movement end; m8 indicates energy consumption estimation; and m9 indicates body coordinate system conversion.
[0018] (3) Analyze the transition constraints between the working modes in the operator set O, introduce the method set M, and construct the state transition function δ: M×O→S, where δ(s, o) represents the state transition function in a certain state s. i Apply an operator o iThe successor state of , i∈[0, 1, 2, ...n].
[0019] Furthermore, in step 1, the mutual transfer relationship between the working modes is as follows:
[0020] 1) The patroller wakes up from sleep mode, executes method m1, and enters charging mode.
[0021] 2) In charging mode, method m2 is executed, indicating that charging is complete and the power supply on the patroller meets the energy conditions for starting and moving the payload. According to the remote control instructions from the ground control center and constraints such as communication windows and lighting, the patroller switches to detection mode, perception mode, and movement mode to perform detection, perception, and movement operations.
[0022] If the patroller switches to detection mode, the load will execute actions according to the program control instructions after it is turned on.
[0023] a. Execute method m5, turn off the payload and camera, and switch to perception mode and movement mode;
[0024] b. Switch to mobile mode, execute method m8, estimate the working energy consumption according to the battery discharge degree on the patroller and the time meter software theory, and then execute method m3 to determine whether the power is lower than the threshold and decide whether to execute the charging mode; switch to perception mode, execute method m3 to determine whether the power is lower than the threshold and decide whether to execute the charging mode;
[0025] If the rover enters the perception mode, the navigation camera is turned on to generate interstellar surface images, method m9 is executed, an external subroutine is called to perform local coordinate system conversion calculations, the mast pointing attitude adjustment is completed, and then according to the instructions of the ground control center,
[0026] a. Execute method m4, turn on the payload, and execute detection mode; or
[0027] b. Execute method m6 to plan the moving path to the next target location;
[0028] If the patrol vehicle switches to mobile mode, it will complete straight-line driving, autonomous obstacle avoidance and wheel control according to the mobile control parameters in the ground control center's instructions;
[0029] a. Execute method m4, turn on the payload, reset the camera, and switch to perception mode; or
[0030] b. Execute method m8 to estimate the working energy consumption according to the battery discharge degree on the patroller and the theoretical time meter software, and then execute method m3 to determine whether the power is lower than the threshold and decide whether to execute the charging mode.
[0031] Furthermore, the step 2 further comprises:
[0032] 2.1. According to the position information of the rover, the measurement and control tracking conditions, and the illumination information determined by the solar altitude angle, calculate the time for the rover to enter and exit the ground measurement and control interval, and determine the effective measurement and control tracking arc;
[0033] 2.2. Based on the position information of the rover, the effective tracking arc, and the original data information of the planetary surface topography, the accessibility of all detection points in the long period is evaluated, the driving path is determined, and the starting and ending positions of the periodic detection are calibrated;
[0034] 2.3. Determine the initial state of the operation based on the starting point location information of periodic detection 0 =<Q,C> , Q represents the state related to the rover body, including the body coordinates, antenna pointing, power, solar wing pointing, yaw mechanism attitude, gimbal mechanism attitude, mast mechanism attitude, and robotic arm. C represents the set of external constraints, including the solar altitude angle, communication link, and carrier switching plan.
[0035] Furthermore, the step 4 further comprises:
[0036] 1. According to the earliest time when the patroller enters the ground station in the effective tracking arc in step 2, the initial time of the operation is determined as T 0 ; Current location information, determined as the operation starting point coordinates N 0 , assigned to T in the task temporal planning network TCN 0 、N 0 , then T 0 N 0 =T 0 ×N 0 represents the starting point of the temporal task network;
[0037] 2. According to the default action sequence contained in the working mode determined in step 1, a set of atomic actions is obtained; the atomic actions constitute the nodes of the task temporal network, and the attributes include the action start time T and the position coordinate information N;
[0038] 3. Call method m8 in method set M to calculate the energy consumption of completing the action as the edge value C in the directed weighted graph;
[0039] 4. The planner performs forward search based on the graph plan and performs instruction expansion;
[0040] 5. The planner performs logical verification on the instructions.
[0041] Furthermore, the planner performs logical verification of instructions as follows:
[0042] Execute instruction overlap check: check whether the same instructions exist at the same time; check whether the number of delayed instructions at the same time is less than the cache limit;
[0043] Perform switch uplink carrier arrangement check: check whether the switch commands entering and leaving the same measuring station are paired; whether the carrier closing time of the same measuring station is later than the carrier opening time; whether there is only one uplink carrier at the same point frequency at the same time; whether the carrier opening time of the next station at the same point frequency is greater than the carrier closing time of the previous station;
[0044] Check the coordination of execution command arrangement and switching uplink carrier: Check whether the command is within the valid tracking arc;
[0045] Verify the overlap of execution instruction issuance time: whether the instruction issuance time is greater than or equal to the end time of the previous instruction issuance.
[0046] Advantages of the present invention:
[0047] 1) Based on the HTN concept, a mathematical expression for patrol mission planning is proposed, and a mathematical model definition P =<D,I,TCN> , as well as domain model D, initial state I 0 The mathematical model above transforms the engineering problem of rover detection into a scientific problem in mission intelligent planning, defines the complete properties of the rover scientific detection mission, and is applicable to all classic planning problem modeling and planning solving processes for action sequences.
[0048] 2) Combine domain knowledge to instantiate mathematical models, propose implementation methods for decomposing long-term tasks, and implement steps for building a temporal constraint network based on atomic actions. The above method combines HTN ideas, temporal constraint networks, and engineering experience, and is an innovation in HTN problem solving methods for dealing with temporal constraints.
[0049] 3) The basis for verifying the correctness of command logic is summarized from four aspects: command overlap, switch uplink carrier arrangement, coordination between command arrangement and switch uplink carrier, and verification of command issuance time overlap. It is a summary and condensation of engineering task experience and has high practical value.
[0050] 4) A human-machine collaborative iterative solution processing method was designed and implemented. This method can well reflect the hierarchical characteristics of the task, is suitable for the scientific exploration mission planning problems of rover in the aerospace field, and has direct application value to the rover's patrol and exploration activities on the surface of extraterrestrial bodies. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 Plan detailed processes for HTN-based tasks;
[0052] Figure 2 It is the state transfer relationship between working modes. DETAILED DESCRIPTION
[0053] In order to make the purpose, technical solution and advantages of the present invention more intuitive and clear, the present invention is further described in detail below with reference to the accompanying drawings and examples. The specific implementation cases described herein are only used to explain this invention, but not to limit the present invention. The technical features designed in each implementation method can be combined with each other as long as they do not conflict with each other.
[0054] The present invention overcomes the shortcomings of low efficiency of manual full control methods and the uncertainty of autonomous control methods. Based on a deep understanding of task requirements and characteristics, a mathematical expression P= <D,I 0 ,TCN>,D represents the domain model,I 0 The initial state and TCN represent the task temporal network. This mathematical model fully describes the elements of the rover scientific exploration mission planning and is suitable for solving the classic planning scenario of action sequence. The definition of the standard paradigm of HTN is expanded, and the domain model D=<O,M,δ> , O represents the operator set, M represents the method set, δ represents the state transfer function, the operator set O decomposes the task into "working modes" with constrained behaviors, each mode contains a default action sequence, and also defines some properties or constraints related to planning calculations. The set of working modes can be expanded according to the task scenario. The method set defines the state transition methods between working modes and the prerequisites that the execution method needs to meet. The state transfer function δ is represented by a matrix: M×O→S, O represents the operator set, and M represents the method set. It is generally applicable to classical planning for action sequences, and a hierarchical design is used to implement an iterative solution process for human-computer collaboration to complete the verification of the state transfer function δ. The planning obtains the target position that the patroller needs to pass through when traveling between detection points, the path between each point, and the action sequence that needs to be completed at each target point.
[0055] Execution process see Figure 1 , including the following steps:
[0056] (I) Analysis of the rover's scientific exploration mission P = <D,I 0 ,TCN>, set domain knowledge and build domain model D=<O,M,δ> .
[0057] 1. Domain knowledge specifically includes detection point accessibility assessment methods, measurement and control tracking condition calculation methods, solar altitude / azimuth prediction methods, energy consumption estimation strategies under different road conditions, and scientific detection needs.
[0058] 2. Use domain knowledge to build domain model D =<O,M,δ> .
[0059] The specific process is as follows:
[0060] (1) According to the task configuration file, define the "working mode" to form an operator set O. Each working mode contains a default action sequence and also defines some properties or constraints related to planning calculations. For example, optional parameters of the mode: arranging experiments, low-speed downlink, etc.; planning-related properties: power consumption, heat dissipation, whether ground assistance is required, communication bandwidth, etc.
[0061] In the process of scientific exploration of the rover, the operator set O = {! perceive,! move,! detect,! charge,! sleep} is defined. The instantiation of the operator! perceive represents the perception mode, that is, the rover obtains navigation information data and transmits the navigation information data to the ground control center; the instantiation of the operator! move represents the movement mode, that is, the rover receives the command of the ground control center and reaches the target position; the instantiation of the operator! detect represents the detection mode, that is, the payload equipment carried by the rover is powered on, obtains scientific exploration data, and transmits the data to the ground control center within the communication window; the instantiation of the operator! charge represents the charging mode, that is, the rover adjusts the solar wing, achieves the sun orientation according to the regulations, and remains stationary, and the battery pack starts to charge; the instantiation of the operator! sleep represents the sleep mode, and other equipment of the rover is completely powered off and does not work. The operator set can be expanded according to the specific application scenario.
[0062] (2) Using domain knowledge, a method set M = {m1, m2, m3, m4, m5, m6, m7, m8, m9} is constructed, which is defined as follows: m1: wake up from sleep mode; m2: charging is completed; m3: power check; m4: load is turned on; m5: load is turned off; m6: movement starts; m7: movement ends; m8: energy consumption estimation; m9: body coordinate system conversion.
[0063] The conversion relationship between working modes is as follows: Figure 2 As shown, the specific process is as follows:
[0064] 1) The patroller wakes up from sleep mode, executes method m1, and enters charging mode.
[0065] 2) In charging mode, method m2 is executed, indicating that charging is complete and the power supply on the patroller meets the energy conditions for starting and moving the payload. According to the remote control instructions from the ground control center and constraints such as communication windows and lighting, the patroller switches to detection mode, perception mode, and movement mode to perform detection, perception, and movement operations.
[0066] If the patroller switches to detection mode, the load will execute actions according to the program control instructions after it is turned on.
[0067] a. Execute method m5, turn off the payload and camera, and switch to perception mode and movement mode;
[0068] b. Switch to mobile mode, execute method m8, estimate the working energy consumption according to the battery discharge degree on the patroller and the time meter software theory, and then execute method m3 to determine whether the power is lower than the threshold and decide whether to execute the charging mode; switch to perception mode, execute method m3 to determine whether the power is lower than the threshold and decide whether to execute the charging mode;
[0069] If the rover enters the perception mode, the navigation camera is turned on to generate interstellar surface images, method m9 is executed, an external subroutine is called to perform local coordinate system conversion calculations, the mast pointing attitude adjustment is completed, and then according to the instructions of the ground control center,
[0070] a. Execute method m4, turn on the payload, and execute detection mode; or
[0071] b. Execute method m6 to plan the moving path to the next target location;
[0072] If the patrol vehicle switches to mobile mode, it will complete straight-line driving, autonomous obstacle avoidance and wheel control according to the mobile control parameters in the ground control center's instructions;
[0073] a. Execute method m4, turn on the payload, reset the camera, and switch to perception mode; or
[0074] b. Execute method m8 to estimate the working energy consumption according to the battery discharge degree on the patroller and the theoretical time meter software, and then execute method m3 to determine whether the power is lower than the threshold and decide whether to execute the charging mode.
[0075] (3) Analyze the transition constraints between the working modes in the operator set O, introduce the method set M, and construct the state transition function δ: M×O→S, where δ(s, o) represents the state transition function in a certain state s. i Apply an operator o i The successor state of the state is i∈[0, 1, 2, ...n]. The state transition function δ defines the mutual transition relationship between the working modes.
[0076] (ii) Based on the initial state of the rover and external constraints, the overall planning of the long-term mission is decomposed into periodic operations between multiple target detection points.
[0077] The specific process is as follows:
[0078] 1. Based on the position information of the rover, the measurement and control tracking conditions, and the lighting information determined by the solar altitude angle, calculate the time it takes for the rover to enter and exit the ground measurement and control interval, and determine the effective measurement and control tracking arc.
[0079] 2. Based on the information of the rover's position, effective tracking arc, and original data of the planet's surface topography, the accessibility of all detection points in a long period is evaluated, the driving path is determined, and the starting and ending positions of the periodic detection are calibrated. The original data of the planet's surface topography, taking the Chang'e-3 mission as an example, includes images obtained by the Chang'e-2 lunar ring photography and images obtained by the lander during its descent.
[0080] 3. Determine the initial state of the operation based on the starting point location information of periodic detection 0 =<Q,C> Q represents the state related to the rover, including the main coordinates, antenna pointing, power, solar wing pointing, yaw mechanism attitude, gimbal mechanism attitude, mast mechanism attitude, robotic arm, etc. C represents the set of external constraints, including solar altitude angle, communication link, carrier switching plan, etc.
[0081] (III) According to the domain model and the initial state of the job, the job is mapped to the operator set, method set and state transfer function, the mathematical model is instantiated, and the behavior sequence of the patroller between the two detection points is determined. Among them, the behavior sequence is instantiated from the operator "working mode" in the domain model, including perception, movement, detection, data transmission, charging, sleep and other behaviors. For example, assuming that the current location coordinate of the patroller is A, it is used as the starting point of this cycle operation, and the planned target point location B is manually specified on the topographic map. The two are 40m apart, and the terrain meets the conditions for job execution. The 10 process detection points arranged between point A and point B are called navigation points. According to the domain knowledge requirements, each navigation point needs to arrange three behaviors of perception, data transmission and movement, and each behavior needs to meet the requirements of power, measurement and control tracking conditions, etc. After the example analysis, the behavior sequence between point A and point B is obtained, including 10 perception, data transmission and movement behaviors, 2 charging behaviors and 1 entry and exit of the measurement and control area.
[0082] (IV) Expand the behavior sequence into atomic actions and construct the temporal constraint network TCN<TN,C> , submit the planner for instruction expansion and logic verification. TN represents a node, which is composed of atomic actions that constitute the working mode. Its attributes include the task execution time T and the current position coordinates N of the patroller; C represents the edge in the network graph, which is consistent with the definition in the initial state mathematical expression, that is, the set of external constraints; the temporal constraint network is a directed weighted graph.
[0083] The method of constructing a temporal constraint network is as follows:
[0084] 1. According to the earliest time when the patroller enters the ground station in the effective tracking arc in step 2, the initial time of the operation is determined as T 0 ; Current location information, determined as the operation starting point coordinates N 0 , assigned to T in the task temporal planning network TCN0 、N 0 . T 0 represents the initial time, N 0 represents the starting point coordinates, then T 0 N 0 =T 0 ×N 0 Represents the starting point of the temporal task network.
[0085] 2. According to the default action sequence contained in the working mode determined in step 1, a set of atomic actions is obtained; the atomic actions constitute the nodes of the task temporal network, and the attributes include the action start time T and the position coordinate information N.
[0086] 3. Call method m8 in method set M to calculate the energy consumption of completing the action as the value C of the edge in the directed weighted graph.
[0087] 4. The planner performs forward search based on the graph plan and performs instruction expansion.
[0088] 5. The planner performs logical verification on the instructions.
[0089] The planner's logical verification of instructions includes instruction overlap check, switch uplink carrier arrangement check, instruction arrangement and switch uplink carrier coordination check, and instruction issuance time overlap verification.
[0090] (1) Instruction overlap check requirements: no identical instructions exist at the same time; the number of delayed instructions at the same time does not exceed the cache limit.
[0091] (2) Inspection requirements for the switch arrangement of uplink carrier: The switch commands entering and leaving the same measuring station must be paired; the carrier switching off time at the same measuring station must be later than the carrier switching on time; there is only one uplink carrier at the same point frequency at the same time; the carrier switching on time of the next station at the same point frequency must be greater than the carrier switching off time of the previous station.
[0092] (3) Requirements for checking the coordination between command arrangement and switch uplink carrier: The command arrangement must be within the valid tracking arc.
[0093] (4) Overlap check requirements for command issuance time: The command issuance time must be greater than or equal to the end time of the previous command issuance.
[0094] Therefore, the planner performs logical verification of instructions as follows:
[0095] Execute instruction overlap check: check whether the same instructions exist at the same time; check whether the number of delayed instructions at the same time is less than the cache limit;
[0096] Perform switch uplink carrier arrangement check: check whether the switch commands entering and leaving the same measuring station are paired; whether the carrier closing time of the same measuring station is later than the carrier opening time; whether there is only one uplink carrier at the same point frequency at the same time; whether the carrier opening time of the next station at the same point frequency is greater than the carrier closing time of the previous station;
[0097] Check the coordination of execution command arrangement and switching uplink carrier: Check whether the command is within the valid tracking arc;
[0098] Verify the overlap of execution instruction issuance time: Check whether the instruction issuance time is greater than or equal to the end time of the previous instruction issuance.
[0099] (V) Verify the correctness of the instruction logic through the human-machine iteration process to obtain the patroller task planning result. In step 5, if the logic verification finds an error, after manual iteration confirmation, return to step 3 for re-planning and adjust the behavior sequence to eliminate the defect factor; if the logic verification is correct, obtain the patroller task planning result.
[0100] The deployment method of the present invention is described by taking the process of the patroller from detection point A to detection point B. The patroller task planning method based on HTN planning includes the following steps:
[0101] (I) Analysis of the rover's scientific exploration mission P = <D,I 0 ,TCN>, set domain knowledge and build domain model D=<O,M,δ> .
[0102] 1. Domain knowledge specifically includes detection point accessibility assessment methods, measurement and control tracking condition calculation methods, solar altitude / azimuth prediction methods, energy consumption estimation strategies under different road conditions, and scientific detection needs.
[0103] 2. Use domain knowledge to build domain model D =<O,M,δ> .
[0104] The specific process is as follows:
[0105] (1) According to the task configuration file, define the "working mode" operator set O = {! perceive,! move,! detect,! charge,! sleep}, where the operator! perceive instantiation represents the perception mode, i.e., the rover obtains navigation information data and transmits the navigation information data to the ground control center; the operator! move instantiation represents the movement mode, i.e., the rover receives the command from the ground control center and reaches the target position; the operator! detect instantiation represents the detection mode, i.e., the payload equipment carried by the rover is powered on to obtain scientific detection data and transmit the data to the ground control center within the communication window; the operator! charge instantiation represents the charging mode, i.e., the rover adjusts the solar wing and remains stationary after achieving solar orientation according to regulations, and the battery pack starts to charge; the operator! sleep instantiation represents the sleep mode, and other equipment of the rover is completely powered off and does not work. The operator set can be expanded according to specific application scenarios.
[0106] (2) Using domain knowledge, construct a method set M = {m1, m2, m3, m4, m5, m6, m7, m8, m9}, defined as follows:
[0107] m1: wake up from sleep mode;
[0108] m2: Charging completed;
[0109] m3: power quantity review;
[0110] m4: load on;
[0111] m5: load off;
[0112] m6: move start;
[0113] m7: end of movement;
[0114] m8: energy consumption estimation;
[0115] m9: Body coordinate system transformation.
[0116] (3) Analyze the transition constraints between the working modes in the operator set O, introduce the method set M, and construct the state transition function δ: M×O→S, where δ(s, o) represents the state transition function in a certain state s. i Apply an operator o i The successor state of the state is i∈[0, 1, 2, ...n]. The state transition function δ defines the mutual transition relationship between the working modes.
[0117] (ii) According to the initial state of the patroller and external constraints, the overall planning of the long-term task is decomposed into periodic operations between multiple target detection points. The patroller's movement from point A to point B is defined as a periodic operation.
[0118] (iii) Based on the domain model and the initial state of the job, the job is mapped to an operator set, a method set, and a state transfer function, the mathematical model is instantiated, and the behavior sequence of the patroller between the two detection points is determined.
[0119] The patroller moves from point A to point B, constructing the following action sequence:
[0120] a) Enter perception mode and turn on the navigation camera to generate interstellar surface images;
[0121] b) Execute method m9, call the external subroutine to perform local coordinate system conversion calculation, and complete the mast pointing attitude adjustment;
[0122] c) Execute method m4, turn on the load, execute the detection mode, and perform actions according to the program control instructions;
[0123] d) Execute method m5, turn off the payload and navigation camera, and enter mobile mode;
[0124] e) Execute method m6 to complete straight-line driving, autonomous obstacle avoidance and wheel control;
[0125] f) Execute method m7, and the movement ends;
[0126] g) Execute method m8 to estimate the working energy consumption according to the battery discharge degree on the patroller and the time meter software theory;
[0127] h) Execute method m3 to determine whether the power level is lower than the threshold and decide whether to execute the charging mode;
[0128] i) If yes, execute method m1 and switch to charging mode;
[0129] j) Execute method m2, indicating that charging is complete and the power supply on the patroller meets the energy conditions for starting up and moving the load.
[0130] (III) Expand the behavior sequence into atomic actions and construct the temporal constraint network TCN<TN,C> , submit the planner for instruction expansion and logic verification.
[0131] (iv) Verify the correctness of the command logic through a human-machine iterative process to obtain the patrol mission planning results.
[0132] (V) The planning is completed.
Claims
1. A patrol mission planning method based on HTN, It is characterized in that The method comprises the following steps:
1. Analysis of the rover's scientific exploration mission P=<D, I 0 ,TCN>, set domain knowledge and build domain model D=<O,M,δ>; Among them, D represents the domain model, I 0 represents the initial state and TCN represents the task temporal network; O represents the operator set, M represents the method set, and δ represents the state transfer function; the domain knowledge includes the detection point accessibility evaluation method, the measurement and control tracking condition calculation method, the solar altitude angle / azimuth prediction method, the energy consumption estimation strategy under different road conditions, and the scientific detection needs; Second, according to the initial state of the patroller and external constraints, the overall planning of the long-term mission is decomposed into periodic operations between multiple target detection points; 3. According to the domain model and the initial state of the job, the job is mapped to the operator set O, the method set M and the state transition function δ, the instantiation of the mathematical model is completed, and the behavior sequence of the patroller between the start and end detection points is determined; Fourth, expand the behavior sequence into atomic actions, construct a temporal constraint network TCN = <TN, C>, and submit it to the planner for instruction expansion and logic verification; TN represents a node, which is composed of atomic actions that constitute the working mode, and its attributes include the task execution time T and the current position coordinates N of the patroller; C represents an edge in the network graph, that is, a set of external constraints; 5. Verify the correctness of the instruction logic through a human-machine iterative process; if errors are found in the logic verification, return to step 3 after manual confirmation, re-plan, and adjust the behavior sequence to eliminate defect factors; if the logic verification is correct, obtain the patroller task planning results.
2. The patrol mission planning method based on HTN as claimed in claim 1, It is characterized in that The step 1 further comprises: (1) According to the task configuration file, define "working modes" to form an operator set O; each working mode contains a default action sequence and also defines some properties or constraints related to planning calculations; define the operator set O = {! perceive,! move,! detect,! charge,! sleep}, the operator! perceive instantiation represents the perception mode, that is, the rover obtains navigation information data and transmits the navigation information data to the ground control center; the operator! move instantiation represents the movement mode, that is, the rover receives the command from the ground control center and reaches the target position; the operator! detect instantiation represents the detection mode, that is, the payload equipment carried by the rover is powered on, obtains scientific detection data, and transmits the data to the ground control center within the communication window; the operator! charge instantiation represents the charging mode, that is, the rover adjusts the solar wing, achieves solar orientation according to regulations, and remains stationary, and the battery pack starts to charge; the operator! sleep instantiation represents the sleep mode, and other equipment of the rover is completely powered off and does not work; (2) Using domain knowledge, construct a method set M = {m1, m2, m3, m4, m5, m6, m7, m8, m9}, where m1 indicates sleep wake-up; m2 indicates charging completion; m3 indicates power verification; m4 indicates load opening; m5 indicates load closing; m6 indicates movement start; m7 indicates movement end; m8 indicates energy consumption estimation; and m9 indicates body coordinate system conversion. (3) Analyze the transition constraints between the working modes in the operator set O, introduce the method set M, and construct the state transition function δ: M×O→S, where δ(s, o) represents the state transition function in a certain state s. i Apply an operator o i The successor state of , i∈[0, 1, 2, ...n].
3. The patrol mission planning method based on HTN as claimed in claim 2, It is characterized in that In step 1, the mutual transfer relationship between the working modes is as follows: 1) The patroller wakes up from sleep mode, executes method m1, and enters charging mode; 2) In charging mode, method m2 is executed, indicating that charging is completed and the power supply on the patroller meets the energy conditions for starting up and moving the payload; according to the remote control instructions from the ground control center and constraints such as communication windows and lighting, the patroller switches to detection mode, perception mode, and movement mode respectively to perform detection, perception, and movement operations; If the patroller switches to detection mode, the load will execute actions according to the program control instructions after it is turned on. a. Execute method m5, turn off the payload and camera, and switch to perception mode and movement mode; b. Switch to mobile mode and execute method m8 to estimate the working energy consumption according to the battery discharge degree on the patroller and the time meter software theory, and then execute method m3 to determine whether the power is lower than the threshold and decide whether to execute the charging mode; switch to Sensing mode, executing method m3, judging whether the power is lower than the threshold, and deciding whether to execute the charging mode; If the rover enters the perception mode, the navigation camera is turned on to generate interstellar surface images, method m9 is executed, an external subroutine is called to perform local coordinate system conversion calculations, the mast pointing attitude adjustment is completed, and then according to the instructions of the ground control center, a. Execute method m4, turn on the payload, and execute detection mode; or b. Execute method m6 to plan the moving path to the next target location; If the patrol vehicle switches to mobile mode, it will complete straight-line driving, autonomous obstacle avoidance and wheel control according to the mobile control parameters in the ground control center's instructions; a. Execute method m4, turn on the payload, reset the camera, and switch to perception mode; or b. Execute method m8 to estimate the working energy consumption according to the battery discharge degree on the patroller and the theoretical time meter software, and then execute method m3 to determine whether the power is lower than the threshold and decide whether to execute the charging mode.
4. The patrol mission planning method based on HTN as claimed in claim 3, It is characterized in that The step 2 further comprises: 2.
1. According to the position information of the rover, the measurement and control tracking conditions, and the illumination information determined by the solar altitude angle, calculate the time for the rover to enter and exit the ground measurement and control interval, and determine the effective measurement and control tracking arc; 2.
2. Based on the position information of the rover, the effective tracking arc, and the original data information of the planetary surface topography, the accessibility of all detection points in the long period is evaluated, the driving path is determined, and the starting and ending positions of the periodic detection are calibrated; 2.
3. Determine the initial state of the operation based on the starting point location information of periodic detection 0 =<Q, C>, Q represents the state related to the rover body, including the body coordinates, antenna pointing, power, solar wing pointing, yaw mechanism attitude, gimbal mechanism attitude, mast mechanism attitude, and robotic arm; C represents the set of external constraints, including the solar altitude angle, communication link, and carrier switching plan.
5. The patrol mission planning method based on HTN as claimed in claim 4, It is characterized in that The step 4 further comprises:
1. According to the earliest time when the patroller enters the ground station in the effective tracking arc in step 2, the initial time of the operation is determined as T 0 ; Current location information, determined as the operation starting point coordinates N 0 , assigned to T in the task temporal planning network TCN 0 、N 0 , then T 0 N 0 =T 0 ×N 0 represents the starting point of the temporal task network; 2. According to the default action sequence contained in the working mode determined in step 1, a set of atomic actions is obtained; the atomic actions constitute the nodes of the task temporal network, and the attributes include the action start time T and the position coordinate information N; 3. Call method m8 in method set M to calculate the energy consumption of completing the action as the edge value C in the directed weighted graph; 4. The planner performs forward search based on the graph plan and performs instruction expansion; 5. The planner performs logical verification on the instructions.
6. The patrol mission planning method based on HTN as claimed in claim 5, It is characterized in that In step 4, the planner performs logic verification on the instruction as follows: Execute instruction overlap check: check whether the same instructions exist at the same time; check whether the number of delayed instructions at the same time is less than the cache limit; Perform switch uplink carrier arrangement check: check whether the switch commands entering and leaving the same measuring station are paired; whether the carrier closing time of the same measuring station is later than the carrier opening time; whether there is only one uplink carrier at the same point frequency at the same time; whether the carrier opening time of the next station at the same point frequency is greater than the carrier closing time of the previous station; Check the coordination of execution command arrangement and switching uplink carrier: Check whether the command is within the valid tracking arc; Verify the overlap of execution instruction issuance time: Check whether the instruction issuance time is greater than or equal to the end time of the previous instruction issuance.
Citation Information
Patent Citations
Hierarchical task network and key path method-based task optimization method
CN107341596A
Deep space exploration constraint satiable task planning method based on null actions
CN107480375A