A track behavior decision-making method and system based on convolutional neural network search tree

CN117910553BActive Publication Date: 2026-08-21NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410069010.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-17
Publication Date
2026-08-21
Estimated Expiration
2044-01-17

AI Technical Summary

Technical Problem

[0003]为了解决现有技术中存在的问题,本发明提供一种基于蒙特卡洛搜索树与卷积神经网络算法的航天器轨道行为决策方法,针对多回合连续追踪条件下的复杂多约束轨道追击与防护规避问题,通过树搜索得到航天器的一系列机动动作指令,利用卷积神经网络对大量的追逃模拟结果进行学习拟合,得到能够利用双方航天器相对状态信息直接输出最佳机动指令的神经网络模型,利用双方航天器的相对信息,基于所述模型得到航天器的机动指令

Benefits of technology

[0042] Compared with existing technologies, this invention has at least the following advantages: The orbital behavior decision-making method can effectively solve continuous dynamic game problems, representing a novel orbital planning method different from existing orbital planning methods based on intelligent optimization algorithms. This method combines search trees and neural networks, possessing both the powerful global optimization capabilities of search trees and the fast computational advantages of neural networks. This significantly improves the computational efficiency of decision-making using traditional Monte Carlo search trees, enabling the decision model obtained through this method to be used on spaceborne computers with lower computing power.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117910553B_ABST
    Figure CN117910553B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of track behavior decision-making method and system based on convolutional neural network search tree, belong to space technology field.The present application first constructs the track game model under the condition of multiple rounds continuous tracking, and establishes its complex multi-constraint model, and the complex task constraint is inducted into mathematical model;The mathematical model of escape spacecraft and pursuit spacecraft is established under the track coordinate system, and the "monte carlo tree search+convolutional neural network" algorithm is proposed, based on the simulation decision-making method using monte carlo tree search, according to the relative position, relative speed information of both sides spacecraft, pursuit and escape simulation is carried out through search tree, search depth is dynamically adjusted according to the relative state relationship of both sides in search process, and the maneuvering action instruction of spacecraft is obtained through tree search, improve the overall efficiency of algorithm, so that the algorithm can be used on spacecraft, can provide effective solution for multiple rounds continuous tracking / avoidance threat for spacecraft.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of aerospace technology, specifically relating to a method and system for orbital behavior decision-making based on a convolutional neural network search tree. Background Technology

[0002] On-orbit servicing and collision avoidance are crucial means of peacefully utilizing space resources. The problem of actively approaching and avoiding a target spacecraft can be categorized as the spacecraft orbital pursuit and escape problem. Existing research on orbital pursuit and escape problems often focuses only on simple pursuit-escape scenarios. However, real-world orbital pursuit and escape scenarios are extremely complex. For escaping spacecraft, in addition to meeting the basic requirements of avoidance, it is also necessary to minimize the impact of avoidance on its original mission, i.e., it cannot stray far from its initial orbital position. Current research lacks sufficient understanding of the characteristics and coupled multiple constraints of such complex scenarios, making it difficult to solve real-world orbital pursuit and escape problems. Summary of the Invention

[0003] To address the problems existing in the prior art, this invention provides a spacecraft orbital behavior decision-making method based on Monte Carlo search tree and convolutional neural network algorithm. For complex multi-constraint orbital pursuit and evasion problems under multi-round continuous tracking conditions, a series of maneuver commands for the spacecraft are obtained through tree search. A convolutional neural network is then used to learn and fit a large number of pursuit and escape simulation results to obtain a neural network model that can directly output the optimal maneuver commands using the relative state information of both spacecraft. Based on the model, the spacecraft's maneuver commands are obtained using the relative information of both spacecraft.

[0004] To achieve the above objectives, the technical solution adopted by this invention is: a trajectory behavior decision-making method based on a convolutional neural network search tree, comprising the following steps:

[0005] Acquire relative information of the two spacecraft in a high-orbit pursuit scenario. The relative information includes the relative coordinates of the two spacecraft, the distance between the two spacecraft, the relative coordinates of the escape spacecraft E and the reference point, the distance between the escape spacecraft E and the reference point, the distance between the escape spacecraft E and the boundary, the relative velocity of the two spacecraft, the relative velocity of the escape spacecraft E, and a certain maneuver.

[0006] The relative information is input into a decision neural network model, and a certain maneuver is... One of the actions; repeated cyclically The secondary neural network feedforward computation, from the output From the data, the input action corresponding to the largest output of the decision neural network is selected as the maneuver command for the spacecraft;

[0007] The input layer of the decision neural network model passes through After convolution processing by the convolution kernel, the output vector is obtained through a fully connected network with three hidden layers. During the training of the decision neural network model, the dynamic deep Monte Carlo search tree algorithm is used to simulate the spacecraft pursuit and escape process in multiple rounds, generating a dataset for model training.

[0008] Furthermore, constructing and training the decision neural network model includes the following steps:

[0009] Establish a game theory model and a multi-constraint model for the pursuit and escape of high-orbit spacecraft;

[0010] Establish a dynamic model of the relative motion orbit of the spacecraft;

[0011] The initial attack trajectory of the pursuing spacecraft P is randomly selected, and the dynamic deep Monte Carlo search tree algorithm is used to simulate the initial trajectory. The simulation process of each initial trajectory has multiple rounds. Each round stores the state variables of both spacecraft and the win rate distribution of each action, forming a dataset. The dataset is then transformed into training samples for a neural network model through calculation.

[0012] A loss function is constructed, and the neural network model is iteratively trained using the Admas solver with training samples until the network loss function converges, resulting in a decision neural network model that can directly output the best maneuver commands using the relative state information of the two spacecraft.

[0013] Furthermore, the establishment of a high-orbit pursuit and escape game and multi-constraint model includes:

[0014] Escape spacecraft E is the escape spacecraft, and pursuit spacecraft P is the pursuit spacecraft. In an orbital pursuit mission, pursuit spacecraft P aims at escape spacecraft E through several orbital maneuvers and pursues it. The objective of escape spacecraft E is to move away from pursuit spacecraft P and prevent P from entering its vicinity. Within a safe range, and the distance from the initial orbital position is always maintained. Within the scope; for orbital pursuit missions, the escape spacecraft's game process must satisfy constraints: within Within a given timeframe, it will not enter the danger zone of the pursuing spacecraft, and its maneuvering range will not exceed the radius around its initial orbital position. The region that satisfies:

[0015]

[0016] in, In order to pursue the spacecraft and escape the distance between them, To escape the distance between the spacecraft and its initial position;

[0017] The process of pursuing spacecraft in a game of strategy must satisfy constraints: Successfully entered the escape spacecraft within the specified time. Within the range, that is, satisfying:

[0018]

[0019] In the pursuit and escape game, both spacecraft employ pulse maneuvers, and the velocity of a single pulse satisfies the following constraints:

[0020]

[0021] Considering the time delay in the precise orbital calculations of each other's spacecraft, the actual game involves the spacecraft alternating pulse maneuvers. Let the time required for precise orbit determination be... During the chase and escape game between the two spacecraft, The spacecraft was constantly in pursuit, and maneuvered to escape spacecraft E. Time came after time. At any moment, the escaping spacecraft senses the unusual movement of the pursuing spacecraft and immediately maneuvers. The two sides take turns maneuvering and competing until a winner is determined.

[0022] Furthermore, establishing a spacecraft dynamics model includes:

[0023] Using the initial position of the escape spacecraft E as a reference point, establish the relative motion equations of both spacecraft with respect to the geostationary target. Due to the target orbital eccentricity... Since the distances between the two spacecraft and the reference point are much smaller than the orbital radius of the reference spacecraft, the analytical solution of the CW equation is used as the dynamic equations for both spacecraft. These dynamic equations are then modified into a state-space form. The relative coordinates and relative velocities of the two spacecraft relative to the reference spacecraft are selected to construct the state space, and equidistant points are selected within the velocity range. A velocity value, constructing the spacecraft in Discrete action space of direction The spacecraft has a total of One optional action.

[0024] Furthermore, when simulating the initial orbit using the dynamic deep Monte Carlo search tree algorithm, the game process of alternating pulse maneuvers between the two spacecraft in the spacecraft pursuit-escape game problem is a round-swapping process, where only one spacecraft maneuvers in each round, and both sides have opportunities. There are several actions to choose from. Before actually executing an action in each round, multiple simulations and iterations are performed. After sufficient iterations, the optimal action is selected and then actually executed to achieve the best game effect.

[0025] Simulating decision-making using Monte Carlo search trees involves the following steps:

[0026] Create a search tree: based on the current state of both players' spacecraft. As the root node of the search tree, i.e., the first level of the search tree, the current turn's side has a total of Several optional maneuvers, after each maneuver... Hours, the status of both spacecraft changed from Transform into According to the corresponding action As the root node of The first child node, which is the second level of the search tree, and so on, gradually create the entire search tree;

[0027] Selection: At a certain level of the search tree, select the child node with the highest score as the action to be taken by the current turn player. The system state changes from... Transform into the selected action corresponding to ;

[0028] Expansion and Simulation: Using the rules selected in the second step, continuously expand the tree structure starting from the root node until either the escaping spacecraft E or the pursuing spacecraft P wins, thus obtaining the game result between the two parties;

[0029] Backtracking: Convert the outcome of the game between the two parties into scores and feed them back to all the child nodes traversed by the game, updating the winning count of each child node. and number of visits ;

[0030] By continuously iterating through the above expansion, simulation, and backtracking steps from the root node, the statistical win rate of the different actions available at the root node is obtained. The algorithm continuously converges to the true value, and selects the action with the highest win rate as the actual action to be executed based on the statistical win rate of different actions.

[0031]

[0032] According to the action After making maneuvers, the system state transitions to the next round. The optimal action is then selected and executed using the methods described above, ultimately completing the entire game process. After each simulation, the number of visits to each node in the "action sequence" of that game is recorded. Add 1 if the escape spacecraft E wins, then add the number of wins for each node in the "action sequence" belonging to the escape spacecraft E. Add 1; if the pursuing spacecraft P wins, then add the number of wins for each node in the "action sequence" belonging to the pursuing spacecraft P. Add 1.

[0033] The scores of the child nodes are calculated using an upper confidence interval, as shown in the following formula:

[0034]

[0035] in It is this layer The first child node The number of times each node wins a round. It is the first Total number of visits to the node It is the total number of visits to all nodes in this layer. It is an exploration constant. It is the sum of the pulse velocities along the winning path. For the weight of the fuel term, the UCT formula is divided into three parts. The first part of the formula represents... The node's value score; the second part represents the node's exploration opportunities and the number of explorations. The less the value, the higher the value; the third part is the fuel item, the higher the required pulse speed, the lower the value, and the less fuel consumed, the more search opportunities will be obtained.

[0036] Furthermore, in the expansion and simulation steps, the path traversed in a complete simulation process from the root node to the end of the simulation is called the "process sequence" of a simulation. The search depth of both parties is dynamically adjusted. For the pursuing spacecraft P, in a certain search simulation, if P successfully captures E, the search ends; if P... If E is not captured within the specified time, then the search continues. As of the deadline, for the escaped spacecraft E, if it is captured by P during a certain search simulation, the search ends; if the distance between P and E is still greater than the safe distance when they reach the minimum distance, the search continues for 4 rounds after the minimum distance is reached; when the two have actually passed each other, the search continues until... The moment has ended.

[0037] Furthermore, by using the relative information of the two spacecraft as input to the neural network, a decision neural network is constructed, denoted as... The input layer is A dimensional vector, corresponding to the current state variables of the spacecraft. The network output layer is 1-dimensional, corresponding to the win rate of a certain action in the current state. , The weights are the weights between the layers of the neural network. The network structure consists of an input layer, a convolutional layer, a fully connected layer 1, a fully connected layer 2, a fully connected layer 3, and an output layer connected in sequence. After the input layer is processed by the convolutional layer, the output vector is obtained through a fully connected network with 3 hidden layers.

[0038] The loss function is: ,in, For the first i The initial orbitals correspond to the win rate distribution of each action. This represents the win rate of a certain action in the current state.

[0039] The present invention also provides a trajectory behavior decision-making system based on a convolutional neural network search tree, including an initial information acquisition module and a decision module;

[0040] The initial information acquisition module is used to acquire the relative information of the two spacecraft in the high-orbit pursuit scenario. The relative information includes the relative coordinates of the two spacecraft, the distance between the two spacecraft, the relative coordinates of spacecraft E and the reference point, the distance between spacecraft E and the reference point, the distance between spacecraft E and the boundary, the relative velocity of the two spacecraft, the relative velocity of spacecraft E, and a certain maneuver.

[0041] The decision module is used to input the relative information into the decision neural network model, and a certain maneuver is... One of the actions; repeated cyclically The secondary neural network feedforward computation, from the output From the data, the input action corresponding to the largest output of the decision neural network is selected as the maneuver command for the spacecraft; wherein, the input layer of the decision neural network model undergoes... After convolution processing with convolution kernels, the output vector is obtained through a fully connected network with 3 hidden layers.

[0042] Compared with existing technologies, this invention has at least the following advantages: The orbital behavior decision-making method can effectively solve continuous dynamic game problems, representing a novel orbital planning method different from existing orbital planning methods based on intelligent optimization algorithms. This method combines search trees and neural networks, possessing both the powerful global optimization capabilities of search trees and the fast computational advantages of neural networks. This significantly improves the computational efficiency of decision-making using traditional Monte Carlo search trees, enabling the decision model obtained through this method to be used on spaceborne computers with lower computing power. Attached Figure Description

[0043] Figure 1 This is a schematic diagram of a typical orbital pursuit and escape mission scenario.

[0044] Figure 2 A schematic diagram of the spacecraft's pursuit and escape game.

[0045] Figure 3 This is a schematic diagram of a decision neural network. Detailed Implementation

[0046] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0047] In the description of this invention, it should be understood that the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.

[0048] A spacecraft orbital behavior decision-making method based on Monte Carlo search tree and convolutional neural network algorithm includes the following steps:

[0049] Design and construct a spacecraft pursuit and escape game model and a multi-constraint model.

[0050] Establish a dynamic model of the relative motion orbit of the spacecraft;

[0051] Based on the relative position and relative velocity information of the two spacecraft, a pursuit and escape simulation is conducted through a search tree. During the search, the search depth of the search tree is dynamically adjusted according to the relative state relationship between the two parties, and the maneuvering commands of the spacecraft are obtained through the tree search.

[0052] By using a convolutional neural network to learn and fit a large number of pursuit and escape simulation results, a neural network model is obtained that can directly output the best maneuver commands using the relative state information of the two spacecraft.

[0053] Step 1: Establish a high-orbit orbit pursuit and escape game and multi-constraint model

[0054] like Figure 1 The diagram illustrates a typical orbital pursuit mission scenario. Spacecraft E is the escaping spacecraft, and spacecraft P is the pursuing spacecraft. In this mission, spacecraft P maneuvers several times to target spacecraft E and pursue it. Spacecraft E's objective is to move away from spacecraft P and prevent it from entering its vicinity. Within a safe range, and the distance from the initial orbital position is always maintained. Within the scope, minimize the impact on the original task.

[0055] For this typical orbital pursuit mission, the spacecraft's game-theoretic process must satisfy the following constraints: Within the specified time, it shall not enter the danger zone of spacecraft P, and its maneuvering range shall not exceed the radius around its initial orbital position. The region that satisfies:

[0056]

[0057] in, Let P be the distance between spacecraft P and spacecraft E. This is the distance between spacecraft E and its initial position.

[0058] The spacecraft P game process must satisfy the following constraints: Successfully entered spacecraft E within the specified time. Within the range, that is, satisfying:

[0059]

[0060] In the pursuit and escape game, both spacecraft employ pulse maneuvers, and the velocity of a single pulse satisfies the following constraints:

[0061]

[0062] Considering the time delay in the precise orbital calculations of each other's spacecraft, the actual game involves the spacecraft alternating between pulse maneuvers. Let the time required for precise orbit determination be... The game of pursuit and escape between the two spacecraft is as follows: Figure 2 The process is shown. Spacecraft P first maneuvers to pursue and escape spacecraft E. Time came after time. At that moment, the escape spacecraft E sensed the abnormal movement of spacecraft P and immediately began maneuvering. The two sides continued this alternating maneuvering game until a winner was determined.

[0063] Step 2: Establish a spacecraft dynamics model.

[0064] Using the initial position of spacecraft E as a reference point, establish the relative motion equations of both spacecraft relative to the geostationary target. Due to the target orbital eccentricity... Since the distances between both spacecraft and the reference point are much smaller than the orbital radius of the reference spacecraft, the analytical solution of the CW equations is used as the dynamic equations for both spacecraft:

[0065]

[0066] The above equations can be rearranged into state-space form, i.e.

[0067]

[0068] in,

[0069] ,

[0070]

[0071] The state space is constructed by selecting the relative coordinates and relative velocities of the two spacecraft relative to a reference spacecraft, and its expression is as follows:

[0072]

[0073] in, Spacecraft P and spacecraft E are respectively relative to the central spacecraft. Relative coordinates of direction Spacecraft P and spacecraft E are respectively relative to the central spacecraft. Relative velocity in a direction.

[0074] Selecting equal intervals within the speed range Speed ​​value Building spacecraft in Discrete action space of direction

[0075]

[0076] Total One optional action.

[0077] Step 3: Simulate decision-making using Monte Carlo search trees

[0078] In the spacecraft pursuit-escape game problem, the alternating pulse maneuvers of both sides can be viewed as a round-by-round exchange. In each round, only one spacecraft performs a maneuver, and both sides have... Each action can be chosen. Before actually executing an action in each round, multiple simulations and iterations are performed. After sufficient iterations, the optimal action is selected, and then that optimal action is executed to achieve the best game outcome. Specifically, the dynamic deep Monte Carlo search tree algorithm includes the following four steps:

[0079] Step 1: Create a search tree. This is based on the current state of both spacecraft in the current round. As the root node of the search tree, that is, the first level of the search tree. The current turn's side has a total of... There are several selectable maneuvers, and after each maneuver... Hours, the status of both spacecraft changed from Transform into According to the corresponding action As the root node of The first child node represents the second level of the search tree. This process continues until the entire search tree is created.

[0080] Step 2: Selection. At a certain level of the search tree, the child node with the highest score is selected as the action to be taken by the current player. The system state changes from... Transform into the selected action corresponding to The scores for child nodes are calculated using the upper limit confidence interval (UCT), as shown in the following formula:

[0081]

[0082] in It is this layer The first child node The number of times each node wins a round. It is the first Total number of visits to the node It is the total number of visits to all nodes in this layer. It is an exploration constant. It is the sum of the pulse velocities along the winning path. This represents the weight of the fuel term. The UCT formula consists of three parts; the first part of the formula represents... The node's value score; the second part represents the node's exploration opportunities and the number of explorations. The less the number of times a node is explored, the larger its value will be, thus nodes that are explored less will have more exploration opportunities; the third part is the fuel item, the higher the required pulse speed, the smaller the value, thus nodes that consume less fuel will have more search opportunities.

[0083] Step 3: Expansion and Simulation. Following the rules chosen in Step 2, the tree structure is continuously expanded from the root node. This process continues until spacecraft E or spacecraft P wins. The path traversed in a complete simulation from the root node to the end of the simulation is called the "process sequence" of a simulation.

[0084] In the third step, the search depth for both sides is dynamically adjusted. For spacecraft P, if P successfully captures E during a certain search simulation, the search ends; if P... If E is not captured within the specified time, then the search continues. The search ends at a certain point. For spacecraft E, if it is captured by P during a search simulation, the search ends; if the distance between P and E is still greater than the safe distance when they reach their minimum distance, the search continues for four rounds after reaching the minimum distance; once they have actually passed each other, the search continues until... The moment has ended.

[0085] Step 4: Backtracking. The game results are converted into scores and fed back to all child nodes traversed in this game, updating the winning counts of these child nodes. and number of visits Specifically, after each simulation, the number of visits to each node in the "action sequence" of that game is recorded. Add 1. If E wins, then add the number of wins for each node in the "action sequence" belonging to spacecraft E. Increment by 1; if P wins, then increment the win count for each node in the "action sequence" belonging to P. Add 1.

[0086] By continuously iterating through the extended simulation and backtracking steps from the root node, the statistical win rate of the different actions available at the root node is obtained. The algorithm continuously converges to the true value, and selects the action with the highest win rate as the actual action to be executed based on the statistical win rate of different actions. ,Right now:

[0087]

[0088] According to the action The system then performs a maneuver, transitioning to the next round. The optimal action is then selected and executed using the same method described above. This completes the entire game process.

[0089] Step 4: The convolutional neural network learns to fit the optimal maneuver commands.

[0090] The relative information of the two spacecraft is used as the input to the neural network, including the relative coordinates of the two spacecraft. Distance between the two spacecraft Relative coordinates of spacecraft E and reference point Distance between spacecraft E and the reference point Distance between spacecraft E and the boundary Relative speed of both parties and Relative velocity of spacecraft E and A certain maneuver The above 13 variables together constitute the input data for the neural network:

[0091]

[0092] Construct a decision neural network, denoted as The input layer is... Dimension, corresponding to the current state variable of the spacecraft. The network output layer is one-dimensional, corresponding to the win rate of a certain action in the current state. , These are the parameters of the neural network, i.e., the weight coefficients between each layer. The network structure is as follows: Figure 3 As shown,

[0093] Input layer after After convolution processing with convolution kernels, the output vector is obtained through a fully connected network with 3 hidden layers.

[0094] The initial approach orbit of spacecraft P is randomly selected, and the dynamic deep Monte Carlo search tree algorithm in step 3 is used to perform simulation calculations on these initial orbits. The simulation process for each initial orbit has multiple rounds, and the state variables of both spacecraft are stored in each round. Win rate distribution for each action This constitutes a data set: Data sets middle Converted after calculation , These are the training samples for the neural network.

[0095] Construct the loss function as follows .

[0096] Using neural network training samples The neural network is iteratively trained using the Admas solver until the network loss function converges. The neural network training algorithm described in this invention is shown in Table 1.

[0097] Table 1 Algorithm Flowchart

[0098]

[0099] During use, the state information of both spacecraft is converted into an input vector, and... Each action is used sequentially as the last two dimensions of the input vector, and the process is repeated cyclically. The secondary neural network performs feedforward computation. From the output... From the data, the input action corresponding to the largest output is selected as the maneuver command for the spacecraft.

[0100] The present invention also provides a trajectory behavior decision-making system based on a convolutional neural network search tree, including an initial information acquisition module and a decision module;

[0101] The initial information acquisition module is used to acquire the relative information of the two spacecraft in the high-orbit pursuit scenario. The relative information includes the relative coordinates of the two spacecraft, the distance between the two spacecraft, the relative coordinates of spacecraft E and the reference point, the distance between spacecraft E and the reference point, the distance between spacecraft E and the boundary, the relative velocity of the two spacecraft, the relative velocity of spacecraft E, and a certain maneuver.

[0102] The decision module is used to input the relative information into the decision neural network model, and a certain maneuver is... One of the actions; repeated cyclically The secondary neural network feedforward computation, from the output From the data, the input action corresponding to the largest output of the decision neural network is selected as the maneuver command for the spacecraft; wherein, the input layer of the decision neural network model undergoes... After convolution processing by the convolution kernel, the output vector is obtained through a fully connected network with three hidden layers. During the training of the decision neural network model, the dynamic deep Monte Carlo search tree algorithm is used to simulate the spacecraft pursuit and escape process in multiple rounds, generating a dataset for model training.

Claims

1. A trajectory behavior decision-making method based on a convolutional neural network search tree, characterized in that, Includes the following steps: Acquire relative information of the two spacecraft in a high-orbit pursuit scenario. The relative information includes the relative coordinates of the two spacecraft, the distance between the two spacecraft, the relative coordinates of the escape spacecraft E and the reference point, the distance between the escape spacecraft E and the reference point, the distance between the escape spacecraft E and the boundary, the relative velocity of the two spacecraft, the relative velocity of the escape spacecraft E, and a certain maneuver. The relative information is input into a decision neural network model, and a certain maneuver is... One of the actions; repeated cyclically The secondary neural network feedforward computation, from the output From the data, the input action corresponding to the largest output of the decision neural network is selected as the maneuver command for the spacecraft; The input layer in the decision neural network model passes through After convolution processing with convolutional kernels, the output vector is obtained through a fully connected network with three hidden layers. During training, the decision neural network model utilizes a dynamic deep Monte Carlo search tree algorithm to simulate the spacecraft pursuit process in multiple rounds, generating a dataset for model training. Constructing and training the decision neural network model includes the following steps: Establish a game theory model and a multi-constraint model for the pursuit and escape of high-orbit spacecraft; Establish a dynamic model of the relative motion orbit of the spacecraft; The initial attack trajectory of the pursuing spacecraft P is randomly selected, and the dynamic deep Monte Carlo search tree algorithm is used to simulate the initial trajectory. The simulation process of each initial trajectory has multiple rounds. Each round stores the state variables of both spacecraft and the win rate distribution of each action, forming a dataset. The dataset is then transformed into training samples for a neural network model through calculation. A loss function is constructed, and the neural network model is iteratively trained using the Adams solver with training samples until the network loss function converges, resulting in a decision neural network model that can directly output the optimal maneuver command using the relative state information of both spacecraft. A high-orbit pursuit-escape game and multi-constraint model is established, including: In the orbital pursuit mission, the pursuing spacecraft P aims at the escaping spacecraft E through several orbital maneuvers and pursues E. The objective of the escaping spacecraft E is to move away from the pursuing spacecraft P and prevent P from entering its vicinity. Within a safe range, and the distance from the initial orbital position is always maintained. Within the scope; for orbital pursuit missions, the escape spacecraft's game process must satisfy constraints: within Within a given timeframe, it will not enter the danger zone of the pursuing spacecraft, and its maneuvering range will not exceed the radius around its initial orbital position. The region that satisfies: in, In order to pursue the spacecraft and escape the distance between them, To escape the distance between the spacecraft and its initial position; The process of pursuing spacecraft in a game of strategy must satisfy constraints: Successfully entered the escape spacecraft within the specified time. Within the range, that is, satisfying: In the pursuit and escape game, both spacecraft employ pulse maneuvers, and the velocity of a single pulse satisfies the following constraints: Considering the time delay in the precise orbital calculations of each other's spacecraft, the actual game involves the spacecraft alternating pulse maneuvers. Let the time required for precise orbit determination be... During the chase and escape game between the two spacecraft, The spacecraft was constantly in pursuit, and maneuvered to escape spacecraft E. Time came after time. At any moment, the escaping spacecraft senses the unusual movement of the pursuing spacecraft and immediately performs maneuvers. The two sides take turns maneuvering and competing until a winner is determined. Establishing a spacecraft dynamics model includes: Using the initial position of the escape spacecraft E as a reference point, establish the relative motion equations of both spacecraft with respect to the geostationary target. Due to the target orbital eccentricity... Since the distances between the two spacecraft and the reference point are much smaller than the orbital radius of the reference spacecraft, the analytical solution of the CW equation is used as the dynamic equations for both spacecraft. These dynamic equations are then modified into a state-space form. The relative coordinates and relative velocities of the two spacecraft relative to the reference spacecraft are selected to construct the state space, and equidistant points are selected within the velocity range. A velocity value, constructing the spacecraft in Discrete action space of direction The spacecraft has a total of One optional action.

2. The trajectory behavior decision-making method based on a convolutional neural network search tree according to claim 1, characterized in that, When using the dynamic deep Monte Carlo search tree algorithm to simulate the initial orbit, for the spacecraft pursuit-escape game problem, the game process of alternating pulse maneuvers by both sides is a round-swapping process, where only one spacecraft maneuvers in each round, and both sides have opportunities. There are several actions to choose from. Before actually executing an action in each round, multiple simulations and iterations are performed. After sufficient iterations, the optimal action is selected and then actually executed to achieve the best game effect.

3. The trajectory behavior decision-making method based on a convolutional neural network search tree according to claim 1, characterized in that, Simulating decision-making using Monte Carlo search trees involves the following steps: Create a search tree: based on the current state of both players' spacecraft. As the root node of the search tree, i.e., the first level of the search tree, the current turn's side has a total of Several optional maneuvers, after each maneuver... Hours, the status of both spacecraft changed from Transform into According to the corresponding action As the root node of The first child node, which is the second level of the search tree, and so on, gradually create the entire search tree; Selection: At a certain level of the search tree, select the child node with the highest score as the action to be taken by the current turn player. The system state changes from... Transform into the selected action corresponding to ; Expansion and Simulation: Using the rules selected in the second step, continuously expand the tree structure starting from the root node until either the escaping spacecraft E or the pursuing spacecraft P wins, thus obtaining the game result between the two parties; Backtracking: Convert the outcome of the game between the two parties into scores and feed them back to all the child nodes traversed by the game, updating the winning count of each child node. and number of visits ; By continuously iterating through the above expansion, simulation, and backtracking steps from the root node, the statistical win rate of the different actions available at the root node is obtained. The algorithm continuously converges to the true value, and selects the action with the highest win rate as the actual action to be executed based on the statistical win rate of different actions. According to the action After making maneuvers, the system state transitions to the next round. The optimal action is then selected and executed using the methods described above, ultimately completing the entire game process. After each simulation, the number of visits to each node in the sequence of actions in that game is recorded. Add 1 if the escaping spacecraft E wins, then add the number of wins for each node in the action sequence belonging to the escaping spacecraft E. Add 1; if the pursuing spacecraft P wins, then add the number of wins for each node in the action sequence belonging to the pursuing spacecraft P. Add 1.

4. The trajectory behavior decision-making method based on a convolutional neural network search tree according to claim 3, characterized in that, The scores of the child nodes are calculated using an upper confidence interval, as shown in the following formula: in The current layer The first child node The number of times each node wins a round. It is the first Total number of visits to the node It is the total number of visits to all nodes in this layer. It is an exploration constant. It is the sum of the pulse velocities along the winning path. For the weight of the fuel term, the UCT formula is divided into three parts. The first part of the formula represents... The node's value score; the second part represents the node's exploration opportunities and the number of explorations. The less the value, the higher the value; the third part is the fuel item, the higher the required pulse speed, the lower the value, and the less fuel consumed, the more search opportunities will be obtained.

5. The trajectory behavior decision-making method based on a convolutional neural network search tree according to claim 2, characterized in that, In the expansion and simulation steps, the path traversed from the root node to the end of a complete simulation is called a simulation process sequence. The search depth for both parties is dynamically adjusted. For the pursuing spacecraft P, in a certain search simulation, if P successfully captures E, the search ends; if P... If E is not captured within the specified time, then the search continues. As of the deadline, for the escaped spacecraft E, if it is captured by P during a certain search simulation, the search ends; if the distance between P and E is still greater than the safe distance when they reach the minimum distance, the search continues for 4 rounds after reaching the minimum distance. Once the two sides have actually crossed paths, the search will proceed to... The moment has ended.

6. The trajectory behavior decision-making method based on a convolutional neural network search tree according to claim 1, characterized in that, Using the relative information of the two spacecraft as input to the neural network, a decision neural network is constructed, denoted as . ,in, The input layer is the state of both spacecraft. A dimensional vector, corresponding to the current state variables of the spacecraft. The network output layer is 1-dimensional, corresponding to the win rate of a certain action in the current state. , The weights are the weights between the layers of the neural network. The network structure consists of an input layer, a convolutional layer, a fully connected layer 1, a fully connected layer 2, a fully connected layer 3, and an output layer connected in sequence. After the input layer is processed by the convolutional layer, the output vector is obtained through a fully connected network with 3 hidden layers. The loss function is: ,in, For the first i The initial orbitals correspond to the win rate distribution of each action. This represents the win rate of a certain action in the current state.

7. A trajectory behavior decision-making system based on a convolutional neural network search tree, characterized in that, It includes an initial information acquisition module and a decision-making module; The initial information acquisition module is used to acquire the relative information of the two spacecraft in the high-orbit pursuit scenario. The relative information includes the relative coordinates of the two spacecraft, the distance between the two spacecraft, the relative coordinates of spacecraft E and the reference point, the distance between spacecraft E and the reference point, the distance between spacecraft E and the boundary, the relative velocity of the two spacecraft, the relative velocity of spacecraft E, and a certain maneuver. The decision module is used to input the relative information into the decision neural network model, and a certain maneuver is... One of the actions; repeated cyclically The secondary neural network feedforward computation, from the output From the data, the input action corresponding to the largest output of the decision neural network is selected as the maneuver command for the spacecraft; wherein, the input layer of the decision neural network model undergoes... After convolution processing with convolutional kernels, the output vector is obtained through a fully connected network with three hidden layers; during the training of the decision neural network model, a dynamic deep Monte Carlo search tree algorithm is used to perform multi-round simulations of the spacecraft pursuit and escape process, generating a dataset for model training; the high-orbit orbit pursuit and escape game and multi-constraint model is established including: In the orbital pursuit mission, the pursuing spacecraft P aims at the escaping spacecraft E through several orbital maneuvers and pursues E. The objective of the escaping spacecraft E is to move away from the pursuing spacecraft P and prevent P from entering its vicinity. Within a safe range, and the distance from the initial orbital position is always maintained. Within the scope; for orbital pursuit missions, the escape spacecraft's game process must satisfy constraints: within Within a given timeframe, it will not enter the danger zone of the pursuing spacecraft, and its maneuvering range will not exceed the radius around its initial orbital position. The region that satisfies: in, In order to pursue the spacecraft and escape the distance between them, To escape the distance between the spacecraft and its initial position; The process of pursuing spacecraft in a game of strategy must satisfy constraints: Successfully entered the escape spacecraft within the specified time. Within the range, that is, satisfying: In the pursuit and escape game, both spacecraft employ pulse maneuvers, and the velocity of a single pulse satisfies the following constraints: Considering the time delay in the precise orbital calculations of each other's spacecraft, the actual game involves the spacecraft alternating pulse maneuvers. Let the time required for precise orbit determination be... During the chase and escape game between the two spacecraft, The spacecraft was constantly in pursuit, and maneuvered to escape spacecraft E. Time came after time. At any moment, the escaping spacecraft senses the unusual movement of the pursuing spacecraft and immediately performs maneuvers. The two sides take turns maneuvering and competing until a winner is determined. Establishing a spacecraft dynamics model includes: Using the initial position of the escape spacecraft E as a reference point, establish the relative motion equations of both spacecraft with respect to the geostationary target. Due to the target orbital eccentricity... Since the distances between the two spacecraft and the reference point are much smaller than the orbital radius of the reference spacecraft, the analytical solution of the CW equation is used as the dynamic equations for both spacecraft. These dynamic equations are then modified into a state-space form. The relative coordinates and relative velocities of the two spacecraft relative to the reference spacecraft are selected to construct the state space, and equidistant points are selected within the velocity range. A velocity value, constructing the spacecraft in Discrete action space of direction The spacecraft has a total of One optional action.

Citation Information

Patent Citations

  • Flooding topology construction method and device for low earth orbit satellite network, and storage medium

    CN115297045A

  • FGCM-based spacecraft orbit threat avoidance decision-making method

    CN116513489A