Method and device for avoiding multistage collision of spacecraft

By predicting and filtering multiple collision objects of spacecraft and selecting appropriate optimization models, the problem that spacecraft can only avoid one collision object in the prior art is solved, and safe avoidance in multi-level collision scenarios is achieved.

CN120135480AActive Publication Date: 2025-06-13BEIHANG UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510184296.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-06-13
Estimated Expiration
2045-02-19

AI Technical Summary

Technical Problem

Existing spacecraft's collision avoidance methods can only control collision avoidance for one collision object, and there is a risk of collision with other collision objects.

Method used

By predicting the closest distance between the spacecraft and multiple potential rendezvous objects, filter out multiple collision objects closest to the spacecraft, and select appropriate collision optimization models to control the spacecraft to avoid multiple collision objects according to the control complexity of the spacecraft circumventing these collision objects.

Benefits of technology

It realizes the effect of controlling the spacecraft to avoid one collision object while avoiding multiple other collision objects at the same time, improving the safety of the spacecraft in multi-level collision scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120135480A_ABST
    Figure CN120135480A_ABST
Patent Text Reader

Abstract

The invention provides a method and device for a spacecraft to avoid multi-level collision, belongs to the technical field of universe navigation, and can control the spacecraft not to collide with other collision objects in the process of controlling the spacecraft to avoid one collision object. The method comprises the following steps: firstly, predicting the distance between a spacecraft and a plurality of potential rendezvous objects when the spacecraft is closest to the plurality of potential rendezvous objects, then screening the plurality of potential rendezvous objects to obtain a plurality of collision objects, and finally, according to the control complexity of the spacecraft for avoiding the plurality of collision objects. And determining a target collision optimization model for controlling the spacecraft to simultaneously avoid the target collision object and other collision objects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of space navigation, and particularly to a method and device for a spacecraft to avoid multi-level collisions. Background Art

[0002] During the process of a spacecraft operating along a set orbit to perform a space mission, there are collision objects such as space debris, defunct satellites, and other on-orbit targets in space. Moreover, with the development of technology, various satellite constellations (i.e., satellite clusters composed of multiple satellites, such as Starlink, etc.) are deployed in space, and the orbital space is becoming increasingly crowded. On this basis, during the process of a spacecraft performing a space mission, the possibility of colliding with multiple objects (i.e., multi-level collisions) within a short period of time increases. Therefore, it is necessary to control the spacecraft to adjust its operating orbit to ensure that the spacecraft does not collide with other objects, so that the spacecraft can successfully complete the space mission.

[0003] However, the existing methods for a spacecraft to avoid collisions can only perform collision avoidance control on the spacecraft for one collision object. During the process of controlling the spacecraft to avoid one collision object, there is a risk that the spacecraft will collide with another collision object. Summary of the Invention

[0004] The present invention provides a method and device for a spacecraft to avoid multi-level collisions, which can control the spacecraft not to collide with other collision objects while controlling the spacecraft to avoid one collision object.

[0005] To achieve the above object, the present invention adopts the following technical solutions:

[0006] In a first aspect, the present invention provides a method for a spacecraft to avoid multi-level collisions, including: predicting the distance when the spacecraft is closest to each potential rendezvous object among multiple potential rendezvous objects. Then, screening out multiple collision objects of the spacecraft from the multiple potential rendezvous objects; the multiple collision objects are potential rendezvous objects among the multiple potential rendezvous objects whose distance when closest to the spacecraft is less than a distance threshold. And determining a target collision optimization model from multiple collision optimization models according to the control complexity of the spacecraft to avoid the multiple collision objects; wherein, the target collision optimization model is used to control the spacecraft to avoid the target collision object among the multiple collision objects, and avoid other collision objects among the multiple collision objects during the process of controlling the spacecraft to avoid the target collision object.

[0007] In a method for a spacecraft to avoid multi - level collisions provided by the present invention, the distance when the spacecraft is closest to each of multiple potential rendezvous objects is predicted, and multiple collision objects of the spacecraft are selected from the multiple potential rendezvous objects based on the above - mentioned distances. Then, based on the control complexity of the spacecraft to avoid the above - mentioned multiple collision objects, a target collision optimization model for the spacecraft to avoid the target collision object among the multiple collision objects is selected, and during the process of controlling the spacecraft to avoid the target collision object, other collision objects among the multiple collision objects are also avoided. It can be seen that in the process of the present invention to avoid a collision with the target collision object, the avoidance of collisions with other collision objects among the multiple collision objects is considered. Therefore, during the process of controlling the spacecraft to avoid one collision object, the spacecraft can be simultaneously controlled not to collide with other collision objects.

[0008] In one implementation manner of the first aspect, according to the time when the spacecraft is closest to each of the multiple collision objects, the collision probability between the spacecraft and each of the multiple collision objects, and the number of the multiple collision objects, the control complexity of the spacecraft to avoid the multiple collision objects is analyzed;

[0009] The control complexity satisfies the following formula,

[0010]

[0011] where c represents the control complexity, n represents the number of the multiple collision objects of the spacecraft, represents the average value of the collision probabilities between the spacecraft and the multiple collision objects, c vt represents the degree of dispersion of the time when the spacecraft is closest to the multiple collision objects, and σ t represents the standard deviation of the time when the spacecraft is closest to the multiple collision objects, μ t represents the average value of the time when the spacecraft is closest to the multiple collision objects.

[0012] In one implementation manner of the first aspect, according to the control complexity of the spacecraft to avoid the multiple collision objects, to determine the target collision optimization model from multiple collision optimization models, it includes: judging the control complexity and the number of the multiple collision objects, and selecting the target collision optimization model. When the number of the multiple collision objects is less than the number threshold, a collision optimization model using the Monte Carlo tree search algorithm with a continuous action space is selected. When the number of the multiple collision objects is greater than or equal to the number threshold, and the control complexity is less than the complexity threshold, a collision optimization model using the cross - entropy algorithm with a continuous action space is selected. When the number of the multiple collision objects is greater than or equal to the number threshold, and the control complexity is greater than or equal to the complexity threshold, a collision optimization model using the proximal policy optimization algorithm is selected.

[0013] In an implementation of the first aspect, the collision probability p satisfies:

[0014]

[0015] where exp(*) represents the exponential function, and R 1 represents the component of the distance when the spacecraft is closest to the collision object on the R-axis in the space-based orbital coordinate system, and S 1 represents the component of the distance on the S-axis in the space-based orbital coordinate system, and W 1 represents the component of the distance on the W-axis in the space-based orbital coordinate system, and σ R represents the joint variance of the position component of the spacecraft on the R-axis in the space-based orbital coordinate system and the position component of the collision object on the R-axis in the space-based orbital coordinate system, and σ SW represents the joint variance of the position component of the spacecraft in the S-W plane in the space-based orbital coordinate system and the position component of the collision object in the S-W plane in the space-based orbital coordinate system, and r A represents the combined radius of the spacecraft and the collision object, that is, the sum of the effective radius of the spacecraft and the effective radius of the collision object, where the effective radius refers to half of the maximum size of the space object.

[0016] In an implementation of the first aspect, the method further includes: excluding space objects that meet any of the following conditions from multiple space objects of the spacecraft. Condition 1: The apogee altitude of the collision object is less than the perigee altitude of the orbit where the spacecraft is located. Condition 2: The perigee altitude of the collision object is greater than the perigee altitude of the orbit where the spacecraft is located. Then, space objects with a time difference between passing through the intersection point less than the time threshold among the multiple space objects are determined as multiple potential rendezvous objects of the spacecraft. Wherein, the time difference between passing through the intersection point is the time difference between the moment when the spacecraft is at the orbit intersection point and the moment when the collision object is at the orbit intersection point, and the orbit intersection point is the intersection point of the orbit where the spacecraft is located and the orbit where the collision object is located.

[0017] In an implementation of the first aspect, multiple collision optimization models are pre-trained based on multi-level collision scenarios; during the pre-training process, the loss function satisfies the following formula;

[0018] LOSS = k 1 ·∑ i∈H p i + k 2 ·f + k 3 ·r

[0019] where LOSS represents the loss value, and k 1 represents the negative weight coefficient of the collision probability, ∑ i∈H p i represents the sum of the collision probabilities between the spacecraft and multiple collision objects, and k 2represents the fuel negative weight coefficient, f represents the fuel consumption, and k 4 represents the acceleration negative weight coefficient, represents the total maneuvering acceleration of the spacecraft, k 3 represents the orbit negative weight coefficient, and r represents the orbit offset of the spacecraft before and after maneuvering.

[0020] In an implementation manner of the first aspect, the method further includes: determining a maneuver control amount for controlling the spacecraft according to the target collision optimization model, and controlling the spacecraft to avoid the target collision object through the maneuver control amount.

[0021] In a second aspect, the present invention provides a device for a spacecraft to avoid multi-level collisions, including a prediction module, a screening module, and a determination module. The prediction module is used to predict the distance when each potential rendezvous object among multiple potential rendezvous objects is closest to the spacecraft. The screening module is used to screen out multiple collision objects of the spacecraft from the multiple potential rendezvous objects; the multiple collision objects are potential rendezvous objects among the multiple potential rendezvous objects whose distance when closest to the spacecraft is less than the distance threshold. The determination module is used to determine a target collision optimization model from multiple collision optimization models according to the control complexity of the spacecraft to avoid multiple collision objects. Wherein, the target collision optimization model is used to control the spacecraft to avoid the target collision object and avoid other collision objects among the multiple collision objects during the process of controlling the spacecraft to avoid the target collision object.

[0022] In a third aspect, the present invention provides an electronic device, including a processor and a memory coupled to the processor; the memory is used to store computer instructions, and when the electronic device runs, the processor executes the computer instructions stored in the memory so that the electronic device executes the method described in the first aspect or any one of its implementation manners above.

[0023] In a fourth aspect, the present invention provides a computer-readable storage medium, including computer program instructions, and when the computer program instructions are executed by a computer, the computer is made to execute the method described in the first aspect or any one of its implementation manners above.

[0024] In a fifth aspect, the present invention provides a computer program product, including computer program instructions, and when the computer program instructions run on a computer, the computer is made to execute the method described in the first aspect or any one of its implementation manners above.

[0025] The technical effects corresponding to the second to fifth aspects and their possible implementation manners above can refer to the description of the technical effects of the first aspect and its possible implementation manners above, and will not be elaborated here. Description of the Drawings

[0026] Figure 1 It is a schematic diagram of six orbital elements provided by an embodiment of the present application;

[0027] Figure 2 It is one of the schematic diagrams of a method for a spacecraft to avoid multi - stage collisions provided by an embodiment of the present application;

[0028] Figure 3 It is one of the schematic diagrams of a multi - stage collision scenario provided by an embodiment of the present application;

[0029] Figure 4 It is the second of the schematic diagrams of a multi - stage collision scenario provided by an embodiment of the present application;

[0030] Figure 5 It is the third of the schematic diagrams of a multi - stage collision scenario provided by an embodiment of the present application;

[0031] Figure 6 It is the fourth of the schematic diagrams of a multi - stage collision scenario provided by an embodiment of the present application;

[0032] Figure 7 It is the second of the schematic diagrams of a method for a spacecraft to avoid multi - stage collisions provided by an embodiment of the present application;

[0033] Figure 8 It is a schematic diagram of the area where the motion orbits of the remaining space objects are located after elimination provided by an embodiment of the present application;

[0034] Figure 9 It is the second of the schematic diagrams of a method for a spacecraft to avoid multi - stage collisions provided by an embodiment of the present application;

[0035] Figure 10 It is a schematic diagram of a satellite - based orbital coordinate system provided by an embodiment of the present application;

[0036] Figure 11 It is a schematic diagram of the structure of a device for a spacecraft to avoid multi - stage collisions provided by an embodiment of the present application. Detailed implementation manners

[0037] If terms such as "first" and "second" are used in the description and claims of the present invention, they are used to distinguish different objects rather than to describe a specific order of the objects.

[0038] In the embodiments of the present application, "and / or" represents the relationship between objects. For example, A and / or B can represent the following three situations: A exists alone, B exists alone, and both A and B exist simultaneously.

[0039] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.

[0040] In the description of the present invention, unless otherwise specified, the meaning of "a plurality of" refers to two or more. For example, a plurality of potential intersection objects refers to two or more potential intersection objects.

[0041] The method and device provided in the embodiments of the present application relate to spacecraft collision avoidance control and can be used to control a spacecraft to avoid multi-stage collisions. Specifically, according to the control complexity of the spacecraft avoiding multiple collision objects, a target collision optimization model for controlling the spacecraft to avoid a target collision object is determined. Moreover, during the process of the spacecraft avoiding the target collision object, other collision objects of the spacecraft are simultaneously avoided, so as to control the spacecraft to avoid multi-stage collisions.

[0042] First, the technical terms involved in the embodiments of the present application will be introduced below.

[0043] 1. Six orbital elements

[0044] The six orbital elements can be simply understood as some parameters of the operating orbit of a spacecraft (such as a satellite) orbiting a central celestial body (such as the Earth). The orbital elements can be used to describe the motion state of the spacecraft in space. For example, for a satellite, through the six orbital elements of the satellite, information such as the position, orbital shape, and motion speed of the satellite can be accurately described.

[0045] In the embodiments of the present application, the orbital elements of a satellite may include the semi-major axis a of the orbit, the eccentricity e, the orbital inclination i, the mean anomaly M, the right ascension of the ascending node Ω, and the argument of perigee ω.

[0046] Among them, the semi-major axis a of the orbit refers to half of the major axis of the satellite orbit, and the semi-major axis is used to describe the size of the satellite orbit.

[0047] The eccentricity e refers to the eccentricity of the satellite orbit and is used to measure the degree of deviation of the orbit.

[0048] The orbital inclination i refers to the dihedral angle between the satellite orbit and the equatorial plane of the central celestial body (such as the orbital inclination i in Figure 1 ).

[0049] The mean anomaly M refers to the angle at which the satellite moves along the satellite orbit starting from the pericenter (the point on the satellite orbit closest to the central celestial body), and the mean anomaly M is used to describe the position of the satellite on the orbit.

[0050] ReferenceFigure 1 The right ascension of the ascending node Ω refers to the angle in the equatorial plane of the central celestial body, measured eastward from the vernal equinox point V to the ascending node B (the point where the satellite orbit crosses the equator from south to north).

[0051] Continue to refer to Figure 1 The argument of perigee ω refers to the angle in the orbital plane of the satellite, measured from the ascending node B to the perigee P.

[0052] The above-mentioned semi-major axis a and eccentricity e of the orbit are used to describe the size and shape of the orbit; the orbital inclination i, the right ascension of the ascending node Ω, and the argument of perigee ω are used to describe the position of the satellite orbit; the mean anomaly M is used to describe the position of the spacecraft in the orbit.

[0053] With the development of the space exploration field, there are more and more various satellites deployed in space, making it more and more likely for a spacecraft to collide with multiple objects in a short period of time (i.e., multi-stage collision). However, the existing methods for controlling spacecraft collision avoidance can only avoid collision with one collision object. In particular, during the process of controlling a spacecraft to avoid one collision object, there is a risk that the spacecraft will collide with another collision object. The embodiments of the present application provide a method and device for a spacecraft to avoid multi-stage collision. Based on the control complexity of the spacecraft avoiding the above-mentioned multiple collision objects, a target collision optimization model for selecting a target collision object among the multiple collision objects that the spacecraft avoids is selected, and during the process of controlling the spacecraft to avoid the target collision object, other collision objects among the multiple collision objects are avoided. Therefore, the embodiments of the present application can control the spacecraft not to collide with other collision objects while controlling the spacecraft to avoid one collision object.

[0054] Exemplarily, the method for a spacecraft to avoid multi-stage collision provided by the embodiments of the present invention can be executed by an electronic device with processing functions. For example, the electronic device can be a computer, a server, etc. Taking the electronic device as a computer as an example, the hardware part of the computer can include: a processor, a memory, a network interface, a user interface, a communication bus, etc.

[0055] Among them, the processor is used to control the electronic device to execute relevant processing and calculation tasks. The processor can include a central processing unit (CPU) or other processors. The processor can be single-core or multi-core. For example, the processor can include multiple CPUs.

[0056] The memory is used to store computer instructions and related data. The memory can be a random access memory (RAM), a read only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, or an optical memory, a magnetic disk storage medium, or any other magnetic storage device, or any other medium capable of storing program code or data that can be accessed by a computer. Optionally, the memory can be integrated within the processor, or the memory can be independent of the processor.

[0057] The network interface is used for the computer to communicate with other devices or communication networks. The network interface can be a transceiver with transceiver functions. Optionally, the network interface can include standard wired interfaces, wireless interfaces (such as WI-FI interfaces, Bluetooth interfaces, 5G interfaces).

[0058] The communication bus is used to enable connection and communication between various different components. For example, the above-mentioned processor, memory, network interface, and user interface can be interconnected through the communication bus.

[0059] The user interface can include a display screen, an input unit (such as a keyboard). Optionally, the user interface can also include standard wired interfaces, wireless interfaces.

[0060] Those skilled in the art can understand that the above computer can also include more or fewer components, or combine certain components, or have different component arrangements, and the embodiments of the present application do not limit this.

[0061] As Figure 2 shown, the method for a spacecraft to avoid multi-level collisions provided by the embodiments of the present application includes S101 - S103.

[0062] S101. Predict the distance when the spacecraft is closest to each potential rendezvous object among multiple potential rendezvous objects;

[0063] In the embodiments of the present application, through the six orbital elements of the spacecraft and the six orbital elements of multiple potential rendezvous objects, the motion states of the spacecraft on its own operating orbit and the motion states of multiple potential rendezvous objects on their own operating orbits within a future period of time are simulated; meanwhile, the distance when the spacecraft is closest to each potential rendezvous object among multiple potential rendezvous objects is recorded and screened out;

[0064] In an application scenario, the positions of a spacecraft and multiple potential rendezvous objects within a future period are calculated based on the six orbital elements of the spacecraft and the six orbital elements of the multiple potential rendezvous objects. Subsequently, the distances between the spacecraft and each of the multiple potential rendezvous objects are obtained by subtracting the positions of the spacecraft and the multiple potential rendezvous objects.

[0065] The calculation formulas for obtaining the positions of the spacecraft / potential rendezvous objects in the J2000 coordinate system based on the six orbital elements of the spacecraft / potential rendezvous objects are as follows;

[0066]

[0067] Among them, r x = r cos f, r y = r sin f,

[0068] S102. Select multiple collision objects of the spacecraft from the multiple potential rendezvous objects; the multiple collision objects are the potential rendezvous objects among the multiple potential rendezvous objects whose distances are less than a distance threshold when closest to the spacecraft.

[0069] Exemplarily, the above distance threshold can be 10 km or other reasonable values, which are not limited in the embodiments of the present application.

[0070] S103. Determine a target collision optimization model from multiple collision optimization models according to the control complexity of the spacecraft to avoid the multiple collision objects.

[0071] Among them, the target collision optimization model is used to control the spacecraft to avoid the target collision object among the multiple collision objects, and to avoid other collision objects among the multiple collision objects during the process of controlling the spacecraft to avoid the target collision object.

[0072] In one implementation manner, S103 includes: judging the control complexity and the number of the multiple collision objects, and selecting the target collision optimization model.

[0073] Specifically, when the number of the multiple collision objects is less than a number threshold, the collision optimization model of the Monte Carlo tree search algorithm with a continuous action space is used as the target collision optimization model.

[0074] The above-mentioned collision optimization model of the Monte Carlo tree search algorithm using a continuous action space is a decision-making planning algorithm that uses Monte Carlo simulation for searching. In the embodiments of the present application, a decision tree is constructed with each node being the state of the spacecraft (for example, node A, node B), and the connection between nodes being the maneuver control quantity (for example, the maneuver control quantity required for the spacecraft to reach the state represented by node B from the state represented by node A).

[0075] It can be understood that the above-mentioned Monte Carlo tree search algorithm with a continuous action space realizes the spacecraft's avoidance of the target collision object through multiple iterations. Exemplarily, the process of each iteration is as follows;

[0076] 1. Selection: Starting from the root node, select an action according to the probability distribution and move down the tree. The selection process can combine the Upper Confidence Bound (UCB) strategy to balance exploration and exploitation;

[0077] 2. Expansion: Under the selected action, expand a new node to represent the new state of the spacecraft;

[0078] 3. Simulation: Start Monte Carlo simulation from the new node, simulate a series of state transitions of the spacecraft after executing this action until a target state is reached (for example, the spacecraft avoids the target collision object);

[0079] 4. Backup: Propagate the simulation results (such as rewards or costs) back to each node in the tree to update the statistical information of each node (such as average reward or number of visits);

[0080] When the number of multiple collision objects is greater than or equal to the quantity threshold and the control complexity is less than the complexity threshold, select the collision optimization model using the cross-entropy algorithm with a continuous action space as the target collision optimization model;

[0081] The above-mentioned cross-entropy algorithm with a continuous action space (CEM) is a Monte Carlo method used to optimize the role and sample the maneuver control quantity according to importance. The goal of the above cross-entropy algorithm is to minimize the cross-entropy between the data distribution obtained by random sampling and the actual data distribution, that is, to minimize the relative entropy (Kullback-Leibler Divergence), and try to make the sampling distribution the same as the actual situation. In the embodiments of the present application, the above-mentioned collision optimization model using the cross-entropy algorithm with a continuous action space finds the optimal strategy by iteratively improving the probability distribution of the maneuver actions corresponding to the maneuver control quantity. Exemplarily, the process of each iteration is as follows;

[0082] 1. Initialization: Define an initial probability distribution (e.g., normal distribution), whose mean and variance are based on prior knowledge of the problem or randomly initialized;

[0083] 2. Sample collection: Draw a large number of samples from the current probability distribution, and these samples represent possible actions or strategies;

[0084] 3. Performance evaluation: Evaluate the performance of each sample through simulation or actual execution, usually measured by calculating the cumulative reward;

[0085] 4. Sample selection: Select the top small portion of samples with the best performance, and the actions or strategies corresponding to these samples are considered as approximations of the current optimal solution;

[0086] 5. Parameter update: Update the parameters of the probability distribution (e.g., the mean and variance of the Gaussian distribution) according to the selected samples, so that the distribution is more concentrated on the excellent samples;

[0087] When the number of multiple collision objects is greater than or equal to the quantity threshold and the control complexity is greater than or equal to the complexity threshold, select a collision optimization model using the Proximal Policy Optimization algorithm;

[0088] The above Proximal Policy Optimization algorithm (also known as the PPO algorithm) is a deep reinforcement learning algorithm based on the policy-value framework (Actor-Critic). In the Actor-Critic framework, two neural networks, Actor and Critic, are set up; in the embodiments of this application, the Actor network of the Actor-Critic framework is used to determine the best avoidance actions generated under a given state, and the Critic network is used to estimate the value of the current policy actions; the above PPO algorithm is based on the policy gradient method. After sampling data by interacting with the environment, it uses stochastic gradient ascent to optimize a surrogate objective function, thereby improving the policy;

[0089] Specifically, the implementation process of the above PPO algorithm is as follows;

[0090] 1. Network initialization: Define two neural networks: Actor and Critic;

[0091] Among them, the Actor network outputs the probabilities of taking various possible actions under a specific state, while the Critic network outputs the state value estimation under the current policy; the policy is improved by optimizing the Actor network, and at the same time, the Critic network is used to guide and evaluate the improvement;

[0092] 2. Interaction and collection: Interact with the environment by executing the actions generated by the Actor network, and collect a series of empirical data, including state, action, reward, and the next state;

[0093] 3. Calculate the advantage function through TD error: Use the data of the Critic to calculate the advantage function A t , to measure the additional value of taking a certain action in a specific state compared to the average policy;

[0094] Assume r t is the immediate reward obtained at time t, and V(s t ) is the estimated value function of state s t . Then the TD error δ t and the advantage function A t can be defined as (γ is the discount factor, used to measure the importance of future rewards): A t ≈δ t =r t +γV(s t+1 )-V(s t );

[0095] 4. Calculate the objective function;

[0096] By using the importance sampling ratio r t (s,a) and the Clip clipping method to assist in calculating the target policy function L clip and the target value function L v , to update the policy and value network;

[0097] Among them, the importance sampling ratio r t (s,a') is: π θ (a|s) represents the probability of executing action a' in state s under the new policy, and π θold (a'|s) represents the probability of executing action a' in state s under the old policy;

[0098] The target value function Lv is: L v =(V(s t )-A t ) 2 ; among them, V(s t ) is the estimated value function of state s t , and A t is the advantage function;

[0099] Apply the Clip clipping method to the target value function L v to obtain the clipped target policy function L clip as:

[0100] L clip =E t [min(r t (θ)A t ,clip(r t(θ), 1 - ε, 1 + ε)A t )]

[0101] The above-mentioned target policy function L clip limits this ratio within the range of 1 - ε to 1 + ε, where ε is a small positive number used to control the clipping range; the clip method is introduced to constrain the step size of policy update to ensure the stability of the update and prevent the new policy from deviating too far from the old policy; r t (θ) represents the ratio of the new and old policies; E t [·] represents the expectation over time step t;

[0102] 5. Optimize the neural network: By the gradient ascent method, calculate the above-mentioned optimized target policy function L and at each gradient (i.e., clip and the above-mentioned target value function L v , and update the parameters of the Actor and the parameters of the Critic;

[0103] In the embodiments of the present application, the above-mentioned multiple collision optimization models (including the Monte Carlo tree search algorithm for continuous action space, the cross-entropy algorithm for continuous action space, and the proximal policy optimization algorithm) are pre-trained based on multi-level collision scenarios;

[0104] The above-mentioned multi-level collision scenarios include: a simple rendezvous scenario of a single debris (as shown in Figure 3 ), a front-back multi-level collision scenario of five debris (as shown in Figure 4 ), a debris cloud collision scenario based on probability distribution (as shown in Figure 5 ), and a complex rendezvous scenario integrating reality (as shown in Figure 6 ); among them, the debris cloud collision scenario based on probability distribution is generated in the virtual scenario according to the normal distribution based on the set average position, average velocity, and variances of position and velocity of the debris cloud; the complex rendezvous scenario integrating reality is constructed by adding real space debris data in the real space on the basis of the debris cloud collision scenario based on probability distribution; the specific situations of the above-mentioned multi-level collision scenarios in the embodiments of the present application are shown in Table 1 below; it can be understood that Figure 3 , Figure 4 , Figure 5 and Figure 6 , the black solid circles represent spacecraft, and the gray solid circles represent debris;

[0105] Table 1

[0106]

[0107]

[0108] In Table 1 above, "centroid alignment" means that the centroids of the two fragments coincide or nearly coincide at the closest moment, and "grazing" means that there is a certain distance between the centroid of the spacecraft and the space debris at the closest moment, but this distance is less than the safety distance;

[0109] During the above pre-training process, the loss function satisfies the following formula;

[0110] LOSS = k 1 ·∑ i∈H p i + k 2 ·f + k 3 ·r Formula (13)

[0111] where LOSS represents the loss value, k 1 represents the negative weight coefficient of the collision probability, ∑ i∈H p i represents the sum of the collision probabilities of the spacecraft with multiple collision objects, k 2 represents the negative weight coefficient of the fuel, f represents the fuel consumption, and k 4 represents the negative weight coefficient of the acceleration, represents the total maneuvering acceleration of the spacecraft, k 3 represents the negative weight coefficient of the orbit, r represents the orbit offset of the spacecraft before and after the maneuver;

[0112] The orbit offset r of the spacecraft before and after the maneuver satisfies the following formula;

[0113]

[0114] where a t and a 0 represent the semi-major axis of the spacecraft's orbit after and before the maneuver respectively, e t and e 0 represent the eccentricity of the spacecraft's orbit after and before the maneuver respectively, Ω t and Ω 0 represent the right ascension of the ascending node of the spacecraft's orbit after and before the maneuver respectively, i t and i 0 represent the inclination of the spacecraft's orbit after and before the maneuver respectively, ω t and ω 0 represent the argument of perigee of the spacecraft's orbit after and before the maneuver respectively, M t and M 0 represent the mean anomaly of the spacecraft's orbit after and before the maneuver respectively;

[0115] Optionally, in combination with Figure 2 , as Figure 7 shown, before S101, the method further includes S104;

[0116] S104. Screen out multiple potential rendezvous objects of the spacecraft from multiple space objects of the spacecraft;

[0117] In an application scenario, S104 includes S1041 - S1042;

[0118] S1041. Eliminate the space objects that meet any of the following conditions from multiple space objects of the spacecraft;

[0119] Condition 1: The apogee altitude of the collision object is less than the perigee altitude of the orbit where the spacecraft is located;

[0120] Condition 2: The perigee altitude of the collision object is greater than the perigee altitude of the orbit where the spacecraft is located;

[0121] Specifically, the spacecraft receives the early warning information sent by the ground monitoring station through the transmission link. The early warning information includes the positions and velocities of multiple space objects of the spacecraft at the current moment. Through the position and velocity of the spacecraft at the current moment in the J2000 mean equator and equinox geocentric coordinate system and the positions and velocities of the above - mentioned multiple space objects at the current moment, the six orbital elements (semi - major axis a of the orbit, eccentricity e, right ascension of the ascending node Ω, orbital inclination i, argument of perigee ω, and mean anomaly M) of the spacecraft and the multiple space objects are calculated; furthermore, the perigee altitude and apogee altitude of the spacecraft are calculated by using the six orbital elements of the spacecraft and the multiple space objects, and the perigee altitude and apogee altitude of each space object among the multiple space objects are calculated; after the above elimination operation, the area where the orbits of the remaining multiple space objects of the spacecraft are located is as Figure 8 shown;

[0122] Since the J2000 mean equator and equinox geocentric coordinate system is a common coordinate system in this technical field, the embodiments of this application will not elaborate on it here;

[0123] The formula for calculating the semi - major axis a of the orbit through the position and velocity is as follows;

[0124]

[0125] where ε represents the specific orbital energy, and μ represents the gravitational constant, v represents the magnitude of the velocity, and v = ||v||, v represents the velocity vector, r represents the magnitude of the position, and r = ||r||, r represents the position vector;

[0126] The formula for calculating the orbital inclination i through the position and velocity is as follows;

[0127]

[0128] where h represents the magnitude of the specific angular momentum, and h = ||h||, h represents the specific angular momentum, h = r×v, hz Denotes the component of the specific angular momentum on the z-axis of the J2000 mean equatorial geocentric coordinate system;

[0129] The formula for calculating the right ascension of the ascending node Ω through the position and velocity is as follows;

[0130]

[0131] Wherein, n denotes the magnitude of the ascending node vector, and n = ||n||, n denotes the ascending node vector h = r × v, and, n x Denotes the component of the ascending node vector on the x-axis of the J2000 mean equatorial geocentric coordinate system;

[0132] The formula for calculating the eccentricity e through the position and velocity is as follows;

[0133] e = ||e|| Formula (4)

[0134] Wherein, e denotes the eccentricity vector,

[0135] The formula for calculating the eccentricity e through the position and velocity is as follows;

[0136]

[0137] The formula for calculating the mean anomaly M through the position and velocity is as follows;

[0138] M = E - e sin E Formula (6)

[0139] Wherein, E denotes the eccentric anomaly, and f 0 Denotes the true anomaly, and

[0140] It should be noted that the perigee altitude and apogee altitude of the above spacecraft and those of multiple space objects can be obtained by inputting the six orbital elements of the spacecraft and the six orbital elements of multiple space objects into simulation software (for example, satellite simulation software Satellite Tool Kit, abbreviated as STK);

[0141] S1042. Then, the space objects among multiple space objects with the time difference between the crossing times less than the time threshold are determined as multiple potential rendezvous objects of the spacecraft;

[0142] Wherein, the above time difference between the crossing times is the time difference between the moment when the spacecraft is at the orbital intersection and the moment when the collision object is at the orbital intersection, and the orbital intersection is the intersection of the orbit where the spacecraft is located and the orbit where the collision object is located;

[0143] Specifically, for each of the multiple spatial objects, the orbital intersection point is calculated by the normal vector of the spacecraft's orbital plane and the normal vector of the spatial object's orbital plane. Then, the time when the spacecraft reaches the orbital intersection point and the time when the spatial object reaches the orbital intersection point are calculated respectively, and the difference between the two is used as the time difference at the intersection point.

[0144] It can be understood that the above normal vector of the orbital plane can be obtained by inputting the six orbital elements of the spacecraft and the six orbital elements of the multiple spatial objects into simulation software (for example, satellite simulation software Satellite Tool Kit, abbreviated as STK).

[0145] The calculation formula for the above orbital intersection point is as follows;

[0146] n 1 =τ 1 ×τ 2 Formula (7)

[0147] n 2 =τ 2 ×τ 1 Formula (8)

[0148] Among them, n 1 represents the intersection point vector on the spacecraft's orbital plane, τ 1 represents the normal vector of the spacecraft's orbital plane, τ 2 represents the normal vector of the spatial object's orbital plane, n 2 represents the intersection point vector on the spatial object's orbital plane;

[0149] The calculation formulas for the time when the spacecraft reaches the orbital intersection point and the time when the spatial object reaches the orbital intersection point are as follows;

[0150] t i-1 =t 0 +(i - 1)T s Formula (9)

[0151] Among them, t i-1 represents the time when the spacecraft / spatial object reaches the orbital intersection point for the i-th time, t 0 represents the time when the spacecraft / spatial object reaches the orbital intersection point for the first time, T s represents the orbital period, Optionally, μ = 3.986005×10 14 m 3 / s 2 ;

[0152] It can be understood that the above t 0 can be obtained by simulating the six orbital elements of the spacecraft / spatial object. The embodiments of the present application do not limit t 0The determination process will not be elaborated further;

[0153] Optionally, in combination with Figure 7 , as Figure 9 shown, before S103, the above method further includes S105;

[0154] S105. Analyze the control complexity of the spacecraft to avoid multiple collision objects according to the time when the spacecraft is closest to each collision object among the multiple collision objects, the collision probability of the spacecraft with each collision object among the multiple collision objects, and the number of the multiple collision objects;

[0155] It should be understood that the time when the spacecraft is closest to each collision object among the multiple collision objects refers to the moment when the spacecraft is closest to each collision object among the multiple collision objects;

[0156] In the embodiment of the present application, while determining the distance when the spacecraft is closest to each collision object among the multiple collision objects through S101, obtain the time when the spacecraft is closest to each collision object among the multiple collision objects; and establish a satellite-based orbital coordinate system, and determine the collision probability of the spacecraft with each collision object among the multiple collision objects according to the positions of the spacecraft and the multiple collision objects in the satellite-based orbital coordinate system and the above distance; then calculate the control complexity of the spacecraft to avoid the multiple collision objects through the above collision probability and the number of the multiple collision objects;

[0157] It can be understood that the above satellite-based orbital coordinate system (also called the RSW coordinate system) is established based on the spacecraft; specifically, as Figure 10 shown, the origin of the satellite-based orbital coordinate system is set at the center of mass of the spacecraft, the positive direction of the R axis of the satellite-based orbital coordinate system is the same as the direction from the center of the earth to the spacecraft, the positive direction of the S axis is perpendicular to the R axis and points to the moving direction of the spacecraft, and the W axis forms a right-handed coordinate system with the R axis and the S axis;

[0158] Based on the above satellite-based orbital coordinate system, the above collision probability p satisfies the following formula;

[0159]

[0160] where exp(*) represents the exponential function, R 1 represents the component of the distance when the spacecraft is closest to the collision object on the R axis of the satellite-based orbital coordinate system, S 1 represents the component of the distance on the S axis of the satellite-based orbital coordinate system, W 1 represents the component of the distance on the W axis of the satellite-based orbital coordinate system, σ RDenotes the combined variance of the position component of the spacecraft on the R-axis in the space-based orbital coordinate system and the position component of the collision object on the R-axis in the space-based orbital coordinate system, that is, the sum of the variance of the position component of the spacecraft on the R-axis in the space-based orbital coordinate system and the variance of the position component of the collision object in the R-axis direction in the space-based orbital coordinate system, σ SW Denotes the combined variance of the position component of the spacecraft in the S-W plane in the space-based orbital coordinate system and the position component of the collision object in the S-W plane in the space-based orbital coordinate system, and Denotes the orbital inclination (known quantity), σ S Denotes the combined variance of the position component of the spacecraft on the S-axis in the space-based orbital coordinate system and the position component of the collision object on the S-axis in the space-based orbital coordinate system, that is, the sum of the variance of the position component of the spacecraft on the S-axis in the space-based orbital coordinate system and the variance of the position component of the collision object in the S-axis direction in the space-based orbital coordinate system, σ W Denotes the combined variance of the position component of the spacecraft on the W-axis in the space-based orbital coordinate system and the position component of the collision object on the W-axis in the space-based orbital coordinate system, that is, the sum of the variance of the position component of the spacecraft on the W-axis in the space-based orbital coordinate system and the variance of the position component of the collision object in the W-axis direction in the space-based orbital coordinate system, σ SW Coupled by the variance of the position component of the spacecraft in the S-axis direction and the variance of the position component in the W-axis direction in the space-based orbital coordinate system, and the variance of the position component of the collision object in the S-axis direction and the variance of the position component in the W-axis direction in the space-based orbital coordinate system, r A Denotes the combined radius of the spacecraft and the collision object, that is, the sum of the effective radius of the spacecraft and the effective radius of the collision object. The effective radius refers to half of the maximum size of a space object (e.g., the spacecraft and the collision object);

[0161] The above control complexity satisfies the following formula;

[0162]

[0163] Where c represents the control complexity and n represents the number of multiple collision objects of the spacecraft, Denotes the average value of the collision probabilities between the spacecraft and multiple collision objects, c vt Denotes the degree of dispersion of the time when the spacecraft is closest to multiple collision objects, and σ t Denotes the standard deviation of the time when the spacecraft is closest to multiple collision objects (used to describe the difference between most of the values and their average value), μ t Denotes the average value of the time when the spacecraft is closest to multiple collision objects; It can be understood that the above c vtThe larger it is, the more dispersed the distribution of the time when the spacecraft is closest to multiple collision objects, the more time is left for the spacecraft to perform maneuvering evasion, and the lower the control complexity for the spacecraft to avoid multiple collision objects.

[0164] Optionally, after S103, the above method further includes S106;

[0165] S106. Determine the maneuver control amount for controlling the spacecraft according to the target collision optimization model; and control the spacecraft to avoid the target collision object through the maneuver control amount.

[0166] In summary, in a method for a spacecraft to avoid multi-level collisions provided by an embodiment of the present application, the distance when the spacecraft is closest to each potential rendezvous object among multiple potential rendezvous objects is predicted, and multiple collision objects of the spacecraft are screened out from the multiple potential rendezvous objects according to the above distances. Then, based on the control complexity of the spacecraft to avoid the above multiple collision objects, a target collision optimization model for the spacecraft to avoid the target collision object among the multiple collision objects is selected, and other collision objects among the multiple collision objects are avoided during the process of controlling the spacecraft to avoid the target collision object. Therefore, it is possible to control the spacecraft not to collide with other collision objects while controlling the spacecraft to avoid one collision object.

[0167] Correspondingly, an embodiment of the present application provides a device for a spacecraft to avoid multi-level collisions, as Figure 11 shown, including a prediction module 501, a screening module 502, and a determination module 503.

[0168] Among them, the prediction module 501 is used to predict the distance when the spacecraft is closest to each potential rendezvous object among multiple potential rendezvous objects. For example, the prediction module 501 is used to implement S101 of the above method for a spacecraft to avoid multi-level collisions.

[0169] The screening module 502 is used to screen out multiple collision objects of the spacecraft from multiple potential rendezvous objects; the multiple collision objects are potential rendezvous objects among the multiple potential rendezvous objects whose distance when closest to the spacecraft is less than the distance threshold. For example, the screening module 502 is used to implement S102 of the above method for a spacecraft to avoid multi-level collisions.

[0170] The determination module 503 is used to determine a target collision optimization model from multiple collision optimization models according to the control complexity of the spacecraft to avoid multiple collision objects; wherein, the target collision optimization model is used to control the spacecraft to avoid the target collision object and avoid other collision objects among the multiple collision objects during the process of controlling the spacecraft to avoid the target collision object. For example, the determination module 503 is used to implement S103 of the above method for a spacecraft to avoid multi-level collisions.

[0171] Each module of the above device for a spacecraft to avoid multi-level collisions can also be used to execute other steps in the above method embodiments. All relevant contents involved in the above method embodiments can be cited in the function descriptions of the corresponding functional modules, and will not be elaborated here.

[0172] An embodiment of the present application further provides an electronic device, including: a processor and a memory coupled to the processor; the memory is used to store computer instructions. When the electronic device runs, the processor executes the computer instructions stored in the memory, so that the electronic device executes the method in the above embodiment. Among them, the processor can implement the above prediction module 501, screening module 502 and determination module 503; the above memory can also be used to store the distance when the spacecraft is closest to multiple potential rendezvous objects, the control complexity of the spacecraft to avoid multiple collision objects, and multiple collision optimization models, etc.

[0173] An embodiment of the present application further provides a computer-readable storage medium, which includes a computer program. When the computer program runs on a computer, it is used to execute the method in the above embodiment.

[0174] An embodiment of the present application further provides a computer program product, which includes computer program instructions. When the computer program instructions run on a computer, it is used to execute the method in the above embodiment.

[0175] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the key points of each embodiment are the differences from other embodiments.

[0176] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for avoiding multi-level collisions of a spacecraft, characterized in that: include: predicting a closest distance between the spacecraft and each of the potential rendezvous objects among a plurality of potential rendezvous objects; Screening a plurality of collision objects for the spacecraft from the plurality of potential intersection objects according to the distance; the plurality of collision objects are potential intersection objects whose closest distance to the spacecraft is less than a distance threshold among the plurality of potential intersection objects; According to the control complexity of the spacecraft in avoiding the multiple collision objects, a target collision optimization model is determined from multiple collision optimization models; wherein the target collision optimization model is used to control the spacecraft to avoid the target collision object among the multiple collision objects, and to control the spacecraft to avoid other collision objects among the multiple collision objects in the process of avoiding the target collision object.

2. The method according to claim 1, characterized in that The method further comprises: Analyzing the control complexity of the spacecraft to avoid the multiple collision objects according to the time when the spacecraft is closest to each of the multiple collision objects, the collision probability between the spacecraft and each of the multiple collision objects, and the number of the multiple collision objects; The control complexity satisfies the following formula: Wherein, c represents the control complexity, n represents the number of the multiple collision objects of the spacecraft, represents the average value of the collision probability p between the spacecraft and the multiple collision objects, c vt represents the dispersion of the times at which the spacecraft is closest to multiple collision objects, and σ t represents the standard deviation of the time when the spacecraft is closest to multiple collision objects, μ t It represents the average of the times when the spacecraft is closest to multiple collision objects.

3. The method according to claim 1, characterized in that Determining a target collision optimization model from a plurality of collision optimization models according to the control complexity of the spacecraft in avoiding the plurality of collision objects comprises: Determining the control complexity and the number of the plurality of collision objects, and selecting the target collision optimization model; When the number of the plurality of collision objects is less than a number threshold, selecting a collision optimization model using a Monte Carlo tree search algorithm in a continuous action space as the target collision optimization model; When the number of the plurality of collision objects is greater than or equal to a number threshold, and the control complexity is less than a complexity threshold, selecting a collision optimization model using a cross entropy algorithm of a continuous action space as the target collision optimization model; When the number of the multiple collision objects is greater than or equal to a number threshold, and the control complexity is greater than or equal to a complexity threshold, a collision optimization model using a proximal strategy optimization algorithm is selected as the target collision optimization model.

4. The method according to claim 1, characterized in that The collision probability p satisfies: Wherein, exp(*) represents an exponential function, R1 represents the component of the distance between the spacecraft and the collision object on the R axis of the satellite-based orbital coordinate system when the spacecraft is closest to the collision object, S1 represents the component of the distance on the S axis of the satellite-based orbital coordinate system, W1 represents the component of the distance on the W axis of the satellite-based orbital coordinate system, and σ R represents the joint variance of the position component of the spacecraft on the R axis of the satellite-based orbital coordinate system and the position component of the collision object on the R axis of the satellite-based orbital coordinate system, σ SW represents the joint variance of the position component of the spacecraft in the SW plane of the satellite-based orbital coordinate system and the position component of the collision object in the SW plane of the satellite-based orbital coordinate system, r A represents the combined radius of the spacecraft and the collision object.

5. The method according to claim 1, characterized in that The method further comprises: Eliminate the space objects satisfying any of the following conditions from the multiple space objects of the spacecraft; Condition 1: The apogee height of the collision object is less than the perigee height of the orbit of the spacecraft; Condition 2: The perigee height of the collision object is greater than the perigee height of the orbit of the spacecraft; Then, the spatial objects whose cross-intersection time difference is less than the time threshold among the multiple spatial objects are determined as multiple potential rendezvous objects of the spacecraft; the cross-intersection time difference is the time difference between the moment when the spacecraft is located at the orbit intersection and the moment when the collision object is located at the orbit intersection, and the orbit intersection is the intersection of the orbit of the spacecraft and the orbit of the collision object.

6. The method according to claim 1, characterized in that The multiple collision optimization models are obtained based on multi-level collision scenario pre-training; during the pre-training process, the loss function satisfies the following formula; LOSS=k1·∑ i∈H p i +k2·f+k3·r Among them, LOSS represents the loss value, k1 represents the negative weight coefficient of collision probability, ∑ i∈H p i The sum of the collision probabilities between the spacecraft and multiple collision objects, k2 represents the negative fuel weight coefficient, f represents the fuel consumption, and k4 represents the negative weight coefficient of acceleration, represents the total maneuvering acceleration of the spacecraft, k3 represents the orbital negative weight coefficient, and r represents the orbital deviation of the spacecraft before and after the maneuver.

7. The method according to claim 1, characterized in that The method further comprises: Determining a maneuver control amount for controlling the spacecraft according to the target collision optimization model; The maneuver control amount is used to control the spacecraft to avoid a target collision object.

8. A device for avoiding multi-level collisions of a spacecraft, characterized in that: It includes prediction module, screening module and determination module; The prediction module is used to predict the distance between the spacecraft and each potential rendezvous object among a plurality of potential rendezvous objects when the spacecraft is closest to each potential rendezvous object; The screening module is used to screen out multiple collision objects for the spacecraft from the multiple potential intersection objects; the multiple collision objects are potential intersection objects whose closest distance to the spacecraft is less than a distance threshold among the multiple potential intersection objects; The determination module is used to determine a target collision optimization model from multiple collision optimization models according to the control complexity of the spacecraft in avoiding the multiple collision objects; wherein the target collision optimization model is used to control the spacecraft to avoid the target collision object, and to control the spacecraft to avoid other collision objects among the multiple collision objects in the process of avoiding the target collision object.

9. A computer-readable storage medium, characterized in that: The method comprises computer program instructions, which, when executed by a computer, cause the computer to perform the method according to any one of claims 1 to 7.

10. A computer program product, characterized in that The method comprises computer program instructions, which, when executed on a computer, cause the computer to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Spacecraft obstacle avoidance control method based on ellipsoid description

    CN112000132A

  • Spacecraft collision early warning method and device, control equipment and storage medium

    CN114715436A

  • Space target collision early warning method based on distribution

    CN115578889A

  • Method and device for controlling spacecraft to avoid multistage collision

    CN119239996A

  • Space debris collision avoidance method and apparatus taking orbit altitude adjustment task into consideration

    WO2024045778A1