Small spacecraft trajectory planning method based on intelligent task supervision

By constructing a six-state-quantity target spacecraft tracking dynamics model and deep reinforcement learning model, the spacecraft trajectory planning strategy is dynamically adjusted, which solves the problem that traditional methods are difficult to coordinate multi-task requirements, and achieves efficient trajectory planning and obstacle avoidance capabilities.

CN120141505AInactive Publication Date: 2025-06-13FUZHOU UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510616217.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-06-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional spacecraft trajectory planning methods are difficult to coordinate multi-task requirements, resulting in a decline in the quality of planned trajectory and may even cause collision accidents. The existing technology relies on human intelligent design behavior prioritization switching rules, and computer performance and storage capabilities are required.

Method used

The small spacecraft trajectory planning method based on intelligent mission supervision is adopted. By constructing a six-state quantity target spacecraft tracking dynamics model, combining deep reinforcement learning models, dynamic learning and adjustment strategies, the task priority of zero-space behavior is selected to guide the small spacecraft for trajectory planning.

Benefits of technology

It improves the adaptability of the spacecraft in an obstacle environment, avoids collision accidents, realizes efficient trajectory planning and obstacle avoidance capabilities, and improves the flexibility and adaptability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120141505A_ABST
    Figure CN120141505A_ABST
Patent Text Reader

Abstract

The invention relates to a small spacecraft trajectory planning method based on intelligent task supervision. The method comprises the following steps: S1, constructing a six-state-quantity target spacecraft tracking dynamics model based on position and speed; s2, acquiring position information of a small spacecraft, position information of a target spacecraft and position information of an obstacle based on a spacecraft tracking kinetic model to obtain a position information set; s3, according to the position information set, basic behaviors and composite behaviors are constructed for null-space behavior control based on the state and trajectory planning requirements of the small spacecraft, and behavior priorities are set as action information; s4, setting different reward functions based on different scenes of the small spacecraft; s5, inputting the position information and the action information into a deep reinforcement learning model for training; and S6, according to the trained deep reinforcement learning model, in trajectory planning, selecting a task priority of a null space behavior according to a learned strategy, and guiding the small spacecraft to perform trajectory planning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of small spacecraft trajectory planning, and particularly to a small spacecraft trajectory planning method based on intelligent mission supervision. Background Art

[0002] In recent years, space engineering technology has made breakthrough progress. The functions of spacecraft have expanded from traditional basic applications such as earth observation and navigation to more strategically valuable fields such as deep space exploration and interplanetary relay communication. However, with the increasingly fierce competition for near-earth space resources and the rapid development of space confrontation technology, spacecraft are facing increasingly complex orbital safety threats. In particular, the sharp increase in the number of space debris has posed a severe challenge to the in-orbit operation safety of spacecraft.

[0003] Under this background, conflicts are extremely likely to occur between the trajectory planning and obstacle avoidance tasks of spacecraft. Traditional planning methods often struggle to coordinate multi-task requirements, resulting in a decline in the quality of the planned trajectory and even potentially triggering collision accidents. Although the null space behavior control method exhibits good control performance in dynamic environments due to its advantages in handling multi-task conflicts, its inherent fixed task priority mechanism limits the adaptability of the system and is difficult to cope with the complex and changing space environment. However, researchers have proposed a finite state machine mission supervisor, a fuzzy mission supervisor, and a predictive mission supervisor. These methods have improved the dynamic performance of null space behavior control to a certain extent and ensured the stability of task execution. However, there are still obvious deficiencies in the existing technologies: on the one hand, these methods need to rely on human intelligence to design behavior priority switching rules; on the other hand, they pose relatively high requirements on the performance and storage capacity of computers. Summary of the Invention

[0004] To solve the above problems, the purpose of the present invention is to provide a small spacecraft trajectory planning method based on intelligent mission supervision, which effectively improves the adaptability of spacecraft tracking and rendezvous in an obstacle environment and avoids collision accidents.

[0005] To achieve the above purpose, the present invention adopts the following technical solutions: A small spacecraft trajectory planning method based on intelligent mission supervision, comprising the following steps S1: Construct a six-state variable target spacecraft tracking dynamics model based on position and velocity; S2: Based on the spacecraft tracking dynamics model, collect the position information of the small spacecraft, the position information of the target spacecraft, and the position information of the obstacles to obtain a position information set; S3: According to the position information set, based on the state of the small spacecraft and the requirements of trajectory planning, construct basic behaviors and composite behaviors for null space behavior control, and set behavior priorities as action information; S4: Set different reward functions based on different scenarios of the small spacecraft; S5: Input the position information and action information into the deep reinforcement learning model for training; S6: According to the trained deep reinforcement learning model, in trajectory planning, select the task priority of the null space behavior according to the learned policy to guide the small spacecraft to perform trajectory planning.

[0006] Furthermore, the non - linear dynamics model of the target spacecraft is specifically: (1) (2) (3) Wherein, 、 and represent the coordinates of the celestial body under the action of the gravitational field on the x, y, and z axes in the Earth inertial coordinate system; 、 and represent the velocities on the x, y, and z axes, 、 and are respectively the first - order derivatives of 、 and ; represents the gravitational parameter of the Earth; represents the equatorial radius of the Earth; represents the distance from the spacecraft under the action of the gravitational field to the center of the Earth; represents the Earth flattening coefficient.

[0007] Furthermore, the basic behaviors include rendezvous behavior, tracking behavior, and obstacle avoidance behavior, specifically as follows: In the rendezvous behavior, define as the controllable task variable for the small spacecraft to pursue the trajectory of the target spacecraft, the position coordinates of the small spacecraft are , and the reference velocity of the rendezvous behavior is: (4) (5) Wherein, is the position coordinate of the small spacecraft; is the pseudo - inverse matrix of , , is the identity matrix, is the Jacobian matrix of the rendezvous behavior; is the trajectory function of the target spacecraft; is the first derivative of the trajectory function of the target spacecraft, representing the velocity of the target spacecraft; is the mission gain matrix, is the mission error; indicates that the mission is a three-dimensional vector and each component is a real number.

[0008] In the obstacle avoidance behavior, define as the controllable mission variable to ensure that the small spacecraft moving along the desired trajectory of tracking and rendezvous avoids obstacles near the path. The reference velocity of the obstacle avoidance behavior is: (6) (7) where is the pseudo-inverse matrix of, , is the Jacobian matrix of the obstacle avoidance behavior; is the obstacle coordinate; is the mission gain matrix; is the mission error, is the safety distance between the small spacecraft and the obstacle; In the tracking behavior, define as the controllable mission variable to ensure the tracking behavior of the small spacecraft to the target spacecraft. The reference velocity of the tracking behavior is: (8) (9) where is the pseudo-inverse matrix of, , is the Jacobian matrix of the tracking behavior; is the desired trajectory function, is the position of the target spacecraft at the current moment; is the first derivative of the desired trajectory function; is the mission gain matrix; is the mission error, indicates that the mission is a three-dimensional vector and each component is a real number.

[0009] Furthermore, the composite behavior includes the first composite behavior and the second composite behavior, which are specific as follows for the rendezvous scenario and the tracking scenario: In the rendezvous scenario, in the first composite behavior, the rendezvous behavior is the main task and the obstacle avoidance behavior is the auxiliary task. In terms of priority, the rendezvous behavior is higher than the obstacle avoidance behavior. The speed of the first composite behavior is specifically: (10) Where is the speed of the first composite behavior in the rendezvous scenario; is the null space of the rendezvous behavior in the rendezvous scenario; is the identity matrix; In the second composite behavior, the obstacle avoidance behavior is the main task and the rendezvous behavior is the auxiliary task. In terms of priority, the obstacle avoidance behavior is higher than the rendezvous behavior. The speed of the second composite behavior is specifically: (11) Where is the speed of the second composite behavior in the rendezvous scenario; is the null space of the obstacle avoidance behavior in the rendezvous scenario; is the identity matrix.

[0010] In the tracking scenario, in the first composite behavior, the tracking behavior is the main task and the obstacle avoidance behavior is the auxiliary task. In terms of priority, the tracking behavior is higher than the obstacle avoidance behavior. The speed of the first composite behavior is specifically: (12) Where is the speed of the first composite behavior in the tracking scenario; is the null space of the tracking behavior in the tracking scenario; is the identity matrix; In the second composite behavior, the obstacle avoidance behavior is the main task and the tracking behavior is the auxiliary task. In terms of priority, the obstacle avoidance behavior is higher than the tracking behavior. The speed of the second composite behavior is specifically: (13) Where is the speed of the second composite behavior in the tracking scenario; is the null space of the obstacle avoidance behavior in the tracking scenario; is the identity matrix.

[0011] Furthermore, the priority setting is as follows: In the rendezvous scenario, when the priority is set to 0, it is the first composite behavior , and when the priority is 1, it is the second composite behavior ; In the tracking scenario, when the priority is set to 0, it is the first composite behavior , and when the priority is 1, it is the second composite behavior .

[0012] Furthermore, different reward functions are set based on different scenarios of small spacecraft, as follows: The trajectory error function between the tracking spacecraft and the target spacecraft is: ; where is the position coordinate of the target spacecraft at time; is the position coordinate of the small spacecraft at time, is the distance error between the small spacecraft and the target spacecraft at time.

[0013] The reward function for the small spacecraft rendezvous scenario is set as follows: (14) (15) (16) where is the position of the target spacecraft at time; the entire reward function consists of two parts, namely the rendezvous error reward function, the rendezvous scenario collision penalty reward function. After adding the reward values of the two, the final reward value is obtained; The reward function for the small spacecraft tracking scenario is set as follows: (17) (18) (19) where is the position of the small spacecraft at time; is the position of the target point at time t; is the tracking distance; the entire reward function consists of two parts, namely the tracking error reward function, the tracking scenario collision penalty reward function. After adding the reward values of the two, the final reward value is obtained.

[0014] Furthermore, the position information and action information are input into the deep reinforcement learning model for training, as follows: Define as the environment; is the state set of the small spacecraft, is the initial state; For the set of behaviors that are composite behaviors, the small spacecraft needs to select composite behaviors based on the environment and then complete the mission; is the experience replay pool, is the capacity of the experience replay pool, is the total number of time steps in one episode, is the safe distance between the small spacecraft and the obstacle, is the total number of training episodes; During the entire mission supervision process, the small spacecraft interacts with the environment When in state , based on a greedy method with an exploration rate that decays exponentially with parameter , a composite behavior is selected . After the small spacecraft makes the composite behavior , the environment will give the small spacecraft a reward as feedback, and then the small spacecraft enters the next state . In this way, a quadruple is generated and stored in the experience replay pool of the reinforcement learning algorithm as experience for future learning. If the number of experiences is equal to the capacity of the replay pool , the experience replay pool will remove the earliest stored quadruple to store new experiences; in trajectory planning, the reinforcement learning task supervisor will select the task priorities of the null space behaviors according to the learned policy to guide the small spacecraft in trajectory planning. A small spacecraft trajectory planning system based on intelligent mission supervision includes a processor, a memory, and a computer program stored on the memory. When the processor executes the computer program, it specifically executes the steps in the above-mentioned small spacecraft trajectory planning method based on intelligent mission supervision.

[0015] A small spacecraft trajectory planning system based on intelligent mission supervision includes a processor, a memory, and a computer program stored on the memory. When the processor executes the computer program, it specifically executes the steps in the above-mentioned small spacecraft trajectory planning method based on intelligent mission supervision.

[0016] The present invention has the following beneficial effects: 1. By introducing a deep reinforcement learning model, the present invention enables the small spacecraft to dynamically learn and adjust strategies according to the actual environment. In a complex and uncertain environment, the small spacecraft can autonomously adapt to changes in obstacles, thereby achieving efficient obstacle avoidance and trajectory planning; 2. The present invention utilizes a null space behavior control framework and combines the priority management of basic behaviors and composite behaviors, which can effectively achieve the coordination between the small spacecraft in meeting the main task and the obstacle avoidance task and avoid task conflicts. At the same time, the null space behavior control supports the small spacecraft to dynamically avoid obstacles without changing the main task, thereby improving the flexibility of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 This is a schematic diagram of the method flow of the present invention. Specific implementation mode

[0018] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments: Refer to Figure 1 , in this embodiment, a small spacecraft trajectory planning method based on intelligent task supervision is provided, including the following steps S1: Construct a six-state target spacecraft tracking dynamics model based on position and velocity; S2: Based on the spacecraft tracking dynamics model, collect the position information of the small spacecraft, the position information of the target spacecraft, and the position information of the obstacles to obtain a position information set; S3: According to the position information set, based on the requirements of the state and trajectory planning of the small spacecraft, construct basic behaviors and composite behaviors for null space behavior control, and set behavior priorities as action information; S4: Set different reward functions based on different scenarios of the small spacecraft; S5: Input the position information and action information into the deep reinforcement learning model for training; S6: According to the trained deep reinforcement learning model, in the trajectory planning, select the task priority of the null space behavior according to the learned strategy to guide the small spacecraft to perform trajectory planning.

[0019] In this embodiment, the non-linear dynamics model of the target spacecraft is specifically: (1) (2) (3) Among them, , and represent the coordinates of the celestial body under the action of the gravitational field on the x, y, and z axes in the Earth inertial coordinate system; , and represent the velocities on the x, y, and z axes, , and are respectively , and 's first-order derivatives; represents the gravitational parameter of the Earth; represents the equatorial radius of the Earth; represents the distance from the spacecraft under the action of the gravitational field to the center of the Earth; represents the Earth flattening coefficient.

[0020] In this embodiment, the basic behaviors include rendezvous behavior, tracking behavior, and obstacle avoidance behavior, which are specifically as follows: In the rendezvous behavior, define as the controllable mission variable for the small spacecraft to pursue the trajectory of the target spacecraft. The position coordinates of the small spacecraft are , and the reference velocity of the rendezvous behavior is: (4) (5) where is the position coordinate of the small spacecraft; is the pseudo-inverse matrix of , is the identity matrix, is the Jacobian matrix of the rendezvous behavior; is the trajectory function of the target spacecraft; is the first derivative of the trajectory function of the target spacecraft, representing the velocity of the target spacecraft; is the mission gain matrix, is the mission error; indicates that the mission is a three-dimensional vector and each component is a real number; In the obstacle avoidance behavior, define as the controllable mission variable to ensure that the small spacecraft avoids obstacles near the path during the process of moving along the expected trajectory of tracking and rendezvous. The reference velocity of the obstacle avoidance behavior is: (6) (7) where is the pseudo-inverse matrix of , is the Jacobian matrix of the obstacle avoidance behavior; is the obstacle coordinate; is the mission gain matrix; is the mission error, is the safety distance between the small spacecraft and the obstacle; In the tracking behavior, define as the controllable mission variable to ensure the tracking behavior of the small spacecraft for the target spacecraft. The reference velocity of the tracking behavior is: (8) (9) where is The pseudo-inverse matrix of , is the Jacobian matrix of the tracking behavior; is the desired trajectory function, is the position of the target spacecraft at the current moment; is the first derivative of the desired trajectory function; is the mission gain matrix; is the mission error, indicating that the mission is a three-dimensional vector and each component is a real number.

[0021] In this embodiment, the composite behavior includes a first composite behavior and a second composite behavior. For the rendezvous scenario and the tracking scenario, specifically as follows: In the rendezvous scenario, in the first composite behavior, the rendezvous behavior is the primary task and the obstacle avoidance behavior is the secondary task. In terms of priority, the rendezvous behavior is higher than the obstacle avoidance behavior. The speed of the first composite behavior is specifically: (10) where is the speed of the first composite behavior in the rendezvous scenario; is the null space of the rendezvous behavior in the rendezvous scenario; is the identity matrix; In the second composite behavior, the obstacle avoidance behavior is the primary task and the rendezvous behavior is the secondary task. In terms of priority, the obstacle avoidance behavior is higher than the rendezvous behavior. The speed of the second composite behavior is specifically: (11) where is the speed of the second composite behavior in the rendezvous scenario; is the null space of the obstacle avoidance behavior in the rendezvous scenario; is the identity matrix.

[0022] In the tracking scenario, in the first composite behavior, the tracking behavior is the primary task and the obstacle avoidance behavior is the secondary task. In terms of priority, the tracking behavior is higher than the obstacle avoidance behavior. The speed of the first composite behavior is specifically: (12) where is the speed of the first composite behavior in the tracking scenario; is the null space of the tracking behavior in the tracking scenario; is the identity matrix; In the second composite behavior, the obstacle avoidance behavior is the primary task and the tracking behavior is the secondary task. In terms of priority, the obstacle avoidance behavior is higher than the tracking behavior. The speed of the second composite behavior is specifically: (13) where For the speed of the second composite behavior in the tracking scenario; For the null space of the obstacle avoidance behavior in the tracking scenario; Is the identity matrix.

[0023] In this embodiment, the priority is set as follows: In the rendezvous scenario, when the priority is set to 0, it is the first composite behavior , and when the priority is 1, it is the second composite behavior ; In the tracking scenario, when the priority is set to 0, it is the first composite behavior , and when the priority is 1, it is the second composite behavior .

[0024] In this embodiment, different reward functions are set based on different scenarios of small spacecraft, as follows: The trajectory error function of the tracking spacecraft and the target spacecraft is: ; Where Is the target spacecraft Time position coordinates; Is the small spacecraft Time position coordinates, Is the distance error between the small spacecraft and the target spacecraft Time.

[0025] The reward function for the small spacecraft rendezvous scenario is set as follows: (14) (15) (16) Where Is Time target spacecraft position; The entire reward function consists of Two parts, namely Rendezvous error reward function, Rendezvous scenario collision penalty reward function. After adding the two reward values, the final reward value is obtained ; The reward function for the small spacecraft tracking scenario is set as follows: (17) (18) (19) Where Is Time small spacecraft position; Is the target point position at time t; is the tracking distance; the entire reward function consists of two parts, which are respectively the tracking error reward function, the tracking scenario collision penalty reward function. After adding the reward values of the two, the final reward value is obtained .

[0026] In this embodiment, the position information and the action information are input into the deep reinforcement learning model for training, as follows: Define as the environment; is the state set of the small spacecraft, is the initial state; is the behavior set of the composite behavior. The small spacecraft needs to select the composite behavior based on the environment and then complete the task; is the experience replay pool, is the capacity of the experience replay pool, is the total number of time steps in one episode, is the safe distance between the small spacecraft and the obstacle, is the total number of training episodes; During the entire task supervision process, the small spacecraft interacts with the environment and, at state , based on a greedy method with an exploration rate that decays exponentially with parameter , selects a composite behavior . After the small spacecraft makes the composite behavior , the environment will give the small spacecraft a reward as feedback. Then the small spacecraft enters the next state , thus generating a quadruple . The quadruple is stored in the experience replay pool of the reinforcement learning algorithm as experience for future learning. If the number of experiences is equal to the replay pool capacity , the experience replay pool will remove the earliest stored quadruple to store new experiences; in trajectory planning, the reinforcement learning task supervisor will select the task priority of the null space behavior according to the learned policy to guide the small spacecraft in trajectory planning.

[0027] A small spacecraft trajectory planning system based on intelligent task supervision, including a processor, a memory, and a computer program stored on the memory. When the processor executes the computer program, it specifically executes the steps in a small spacecraft trajectory planning method based on intelligent task supervision as described above.

[0028] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.

[0029] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0030] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0031] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0032] As described above, it is only the preferred embodiments of the present invention, and the present invention is not limited to other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still fall within the protection scope of the technical solution of the present invention.

Claims

1. A small spacecraft trajectory planning method based on intelligent mission supervision, characterized in that: The following steps are included S1: Construct a six-state target spacecraft tracking dynamics model based on position and velocity; S2: Based on the spacecraft tracking dynamics model, the position information of the small spacecraft, the position information of the target spacecraft and the position information of the obstacle are collected to obtain a set of position information; S3: Based on the position information set and the requirements of the state and trajectory planning of the small spacecraft, basic behaviors and composite behaviors are constructed for zero-space behavior control, and the behavior priority is set as action information; S4: Set different reward functions based on different scenarios of small spacecraft; S5: Input the position information and action information into the deep reinforcement learning model for training; S6: Based on the trained deep reinforcement learning model, in trajectory planning, the task priority of the zero-space behavior is selected according to the learned strategy to guide the small spacecraft to perform trajectory planning.

2. The small spacecraft trajectory planning method based on intelligent mission supervision according to claim 1 is characterized in that: The nonlinear dynamic model of the target spacecraft is specifically: (1); (2); (3); in, , and It represents the coordinates of the celestial body affected by the gravitational field on the x, y and z axes in the earth's inertial coordinate system; , and represents the speed on the x, y and z axes, , and They are , and The first derivative of ; Represents the gravitational parameters of the Earth; represents the equatorial radius of the Earth; It represents the distance from the spacecraft to the center of the earth under the influence of the gravitational field; Represents the Earth's flattening coefficient.

3. The small spacecraft trajectory planning method based on intelligent mission supervision according to claim 2 is characterized in that: The basic behaviors include intersection behavior, tracking behavior and obstacle avoidance behavior, as follows: In the intersection behavior, define is the controllable mission variable for the small spacecraft to pursue the target spacecraft trajectory. The position coordinates of the small spacecraft are , the reference speed for the rendezvous behavior for: (4); (5); in, are the position coordinates of the small spacecraft; for The pseudo-inverse matrix of , is the identity matrix, is the Jacobian matrix of the intersection behavior; is the trajectory function of the target spacecraft; is the first-order derivative of the target spacecraft’s trajectory function, representing the target spacecraft’s velocity; is the task gain matrix, is the task error; Represents that the task is a three-dimensional vector, and each component is a real number; In obstacle avoidance behavior, we define The controllable mission variables for ensuring that obstacles near the path are avoided during the small spacecraft moving along the desired trajectory of tracking and rendezvous, and the reference speed of the obstacle avoidance behavior for: (6); (7); in for The pseudo-inverse matrix of , is the Jacobian matrix of the obstacle avoidance behavior; is the obstacle coordinate; is the task gain matrix; is the task error, is the safe distance between small spacecraft and obstacles; In tracking behavior, define Controllable mission variables for small spacecraft to ensure tracking behavior of target spacecraft, reference speed of tracking behavior for: (8); (9); in for The pseudo-inverse matrix of , is the identity matrix, is the Jacobian matrix of the tracking behavior; is the expected trajectory function, is the current position of the target spacecraft; is the first-order derivative of the expected trajectory function; is the task gain matrix; is the task error, Represents a task as a three-dimensional vector, and each component is a real number.

4. The small spacecraft trajectory planning method based on intelligent mission supervision according to claim 3 is characterized in that: The composite behavior includes a first composite behavior and a second composite behavior, which are aimed at the intersection scene and the tracking scene, and are as follows: In the intersection scenario, in the first composite behavior, the intersection behavior is the main task and the obstacle avoidance behavior is the auxiliary task. In terms of priority, the intersection behavior is higher than the obstacle avoidance behavior. The speed of the first composite behavior is as follows: (10); in is the first composite behavior speed in the intersection scenario; is the null space of the intersection behavior in the intersection scenario; is the identity matrix; In the second composite behavior, obstacle avoidance is the main task and intersection is the auxiliary task. In terms of priority, obstacle avoidance is higher than intersection. The speed of the second composite behavior is as follows: (11); in is the second composite behavior speed in the intersection scenario; It is the zero space for obstacle avoidance behavior in the intersection scenario; is the identity matrix; In the tracking scenario, in the first composite behavior, the tracking behavior is the main task and the obstacle avoidance behavior is the auxiliary task. In terms of priority, the tracking behavior is higher than the obstacle avoidance behavior. The speed of the first composite behavior is as follows: (12); in To track the speed of the first composite behavior in the scene; It is the null space of tracking behavior in tracking scenarios; is the identity matrix; In the second composite behavior, obstacle avoidance is the main task and tracking is the auxiliary task. In terms of priority, obstacle avoidance is higher than tracking. The speed of the second composite behavior is as follows: (13); in To track the speed of the second composite behavior in the scene; The null space for obstacle avoidance behavior in tracking scenarios; is the identity matrix.

5. The small spacecraft trajectory planning method based on intelligent mission supervision according to claim 4 is characterized in that: The priority setting is as follows: In the intersection scenario, when the priority is set to 0, it is the first composite behavior , when the priority is 1, it is the second composite behavior ; In the tracking scenario, setting the priority to 0 is the first composite behavior , when the priority is 1, it is the second composite behavior .

6. The small spacecraft trajectory planning method based on intelligent mission supervision according to claim 5 is characterized in that: Different reward functions are set based on different scenarios of small spacecraft, as follows: The trajectory error function between the tracking spacecraft and the target spacecraft is: ; in Target spacecraft Position coordinates at the moment; For small spacecraft The position coordinates at the time, For small spacecraft and target spacecraft Time distance error; The reward function for the small spacecraft rendezvous scenario is set as follows: (14); (15); (16); in for The target spacecraft position at the moment; the entire reward function It consists of two parts, namely The intersection error reward function, The collision penalty reward function for the intersection scene. After adding the two reward values, the final reward value is obtained. ; The reward function for the small spacecraft tracking scenario is set as follows: (17); (18); (19); in for The position of the small spacecraft at any moment; is the target point position at time t; is the tracking distance; the entire reward function It consists of two parts, namely Tracking error reward function, Tracking scene collision penalty reward function, add the two reward values ​​to get the final reward value .

7. The small spacecraft trajectory planning method based on intelligent mission supervision according to claim 6 is characterized in that: The position information and action information are input into the deep reinforcement learning model for training, as follows: definition for the environment; is the state set of the small spacecraft, is the initial state; It is a set of composite behaviors. Small spacecraft needs to select composite behaviors based on the environment to complete the mission. For the experience replay pool, is the capacity of the experience replay pool, is the total time step of one round, is the safe distance between small spacecraft and obstacles, is the total training rounds; Throughout the mission monitoring process, small spacecraft and environment To interact, in state When, based on an exploration rate with parameter Exponentially decaying Greedy approach to select a composite behavior , small spacecraft are performing complex behaviors After that, the environment will give the small spacecraft a The reward is used as feedback, and then the small spacecraft enters the next state , thus generating a four-tuple , store the quadruple in the experience replay pool of the reinforcement learning algorithm As experience for future learning, if the amount of experience is equal to the capacity of the replay pool , the experience replay pool will remove the earliest stored quadruple and use it to store new experience; In trajectory planning, the reinforcement learning task supervisor selects the task priorities of the null space behavior according to the learned strategy to guide the small spacecraft in trajectory planning.

8. A small spacecraft trajectory planning system based on intelligent mission supervision, characterized in that: The method comprises a processor, a memory and a computer program stored in the memory. When the processor executes the computer program, the method specifically executes the steps in a small spacecraft trajectory planning method based on intelligent mission supervision as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Multi-spacecraft formation system control integration method based on predictive behavior control

    CN116280269A

  • Small spacecraft long-distance pursuit and close-distance rendezvous integrated control method

    CN116873230A

  • Space rendezvous guidance method considering avoidance constraint

    CN117232533A

  • Track planning method for adaptive adjustment of dynamic obstacle avoidance precision of spacecraft

    CN118605563A

  • Mechanical arm dynamic obstacle avoidance method based on intelligent task supervision

    CN119458386A