A High-Value Spacecraft Cluster Intelligent Guarding Method

By building a behavioral rule base and capability rule base of spacecraft clusters, and combining multi-agent reinforcement learning algorithms, the problem of rapid response and intelligent decision-making of spacecraft clusters in the face of multiple space threats and target uncertainties is solved, intelligent decision-making and collaborative control of spacecraft clusters are realized, and the response speed and decision-making efficiency of multi-agent systems are improved.

CN118597444BActive Publication Date: 2025-06-20HARBIN INST OF TECH +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410759991.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-13
Publication Date
2025-06-20
Estimated Expiration
2044-06-13

AI Technical Summary

Technical Problem

The prior art is difficult to achieve rapid response and intelligent decision-making of spacecraft clusters in the face of multiple space threats and objective uncertainties, especially in multi-agent systems, which are difficult to make effective decisions and reactions in a short time.

Method used

By building a behavioral rule library and capability rule library for spacecraft clusters, combining multi-agent reinforcement learning algorithms and reward functions, a reinforcement learning model is established to realize intelligent decision-making and collaborative control of spacecraft clusters. The method includes obtaining initial data, environmental deduction, iterative training of reinforcement learning models until convergence, and applying the final decision result in a real scenario.

Benefits of technology

It realizes the ability of spacecraft clusters to make decisions and respond in a short period of time, avoids the time cost brought by the large circuit of the world, can cope with the impact of uncertain target behavior in space game tasks, and is suitable for a variety of threat scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118597444B_ABST
    Figure CN118597444B_ABST
Patent Text Reader

Abstract

The present application discloses a method for intelligent escort of a high-value spacecraft cluster, belonging to the technical field of spacecraft, including: obtaining the initial data of the spacecraft cluster, where the spacecraft cluster includes multiple spacecraft with autonomous decision-making; according to the initial data, combining the spacecraft orbit model, satellite on-board sensor model, and illumination model to conduct environmental deduction on the spacecraft cluster and obtain the current simulated environment; constructing a behavior rule library and a capability rule library for the spacecraft cluster; setting a reward function, establishing a reinforcement learning model according to the multi-agent reinforcement learning algorithm and the reward function, and iteratively training the reinforcement learning model according to the current simulated environment, behavior rule library, and capability rule library until the reinforcement learning model converges; simulating the tasks in the real scenario according to the converged reinforcement learning model to obtain the final multi-agent decision result and performing visualization processing. Making decisions for the problem of spacecraft cluster approaching can avoid the radiation or interference of other spacecraft.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to a method for intelligent escort of a high-value spacecraft cluster, belonging to the technical field of spacecrafts. Background Art

[0002] Due to the limited on-board load capacity, a single satellite does not have the ability to complete the entire process of perception, action sequence planning and maneuvering. Therefore, the entire process from threat perception to maneuvering response is usually completed through a large space-ground loop with a ground command unit. However, the battlefield situation changes rapidly. How to improve the rapid response ability to radiation or interference from other space units is an important research direction. Future spacecrafts should possess the abilities of autonomous perception, autonomous judgment, autonomous decision-making and autonomous control.

[0003] Existing research on decision-making methods mainly includes traditional methods, optimization methods, artificial intelligence methods and the comprehensive application of these methods. The traditional method is mainly the Lambert orbit change method used in the classic spacecraft rendezvous problem. The main content is to find the initial and final velocities given the initial and final positions and transfer time to obtain the spacecraft rendezvous decision. The essence of the optimization method is to transform the decision-making problem into an optimization problem based on the knowledge of other related disciplines for solution. For example, in 2008, Chen Changqing et al. used the shooting method to solve the multi-pulse rendezvous problem in "Research Progress on Optimal Impulse Rendezvous". The robustness of the direct method is better than that of the indirect method, mainly because the indirect method lacks a systematic method to determine the initial values ​​of the initial adjoint variables of the problem. Although the direct method bypasses the necessary conditions in the indirect method, the direct method also relies on initial value guessing. It is worth noting that it is not easy to generate a suitable guess value. In fact, the optimal solution found by the direct method can only be located in the neighborhood of the guess value. Compared with the indirect method and the direct method, the evolutionary algorithm explores the global optimal solution through frequent iterations without relying on the initial guess. The evolutionary algorithm mainly includes genetic algorithm, particle swarm algorithm and differential evolution algorithm. Evolutionary algorithms are simple and robust, and can be easily transplanted into more complete and accurate gravity models. For example, in 2015, Ouyang Gaoxiang et al. used GA to search for the global optimal orbit change in "Multi-pulse rendezvous guidance based on implicit gene hybrid genetic algorithm". In 2017, Liu Yuan et al. used the PSO algorithm to find the best multi-pulse rendezvous trajectory for coplanar orbits in "Optimization of multi-pulse non-planar rendezvous and docking transfer orbit". In 2022, Li Junlong et al. used the PSO algorithm to optimize the three-pulse rendezvous problem in "A three-pulse rendezvous optimization design method under multiple constraints", but used the relative motion equation, which is only applicable to coplanar circular orbits. It can be seen that the optimization of multi-pulse maneuvers using evolutionary algorithms is still the basis of this type of problem. This method has a clear mathematical form and good logic, but it is complex to solve, has a large amount of calculations, and has poor adaptability, making it difficult to make decisions in unknown environments. The rapid development of machine learning in the field of artificial intelligence has provided a new way to solve autonomous decision-making problems. For example, some progress has been made in the field of spacecraft rendezvous guidance using reinforcement learning. Wang et al. proposed a low-thrust autonomous rendezvous guidance method based on a deep deterministic policy gradient algorithm in "Autonomous Rendezvous Guidance via Deep Reinforcement Learning". HOVELL et al. proposed a deep guidance method for spacecraft guidance in "Deep reinforcement learning for spacecraft proximity operations guidance" and adopted a distributed deep deterministic policy gradient algorithm. FEDERICI et al.

[0004] The problems of behavior cloning and reinforcement learning in linear multi-impulse rendezvous missions are studied in "Machine Learning Techniques for Autonomous Spacecraft Guidance during Proximity Operations". Although the above research has solved the problem of maneuver decision-making to a certain extent, with the development of space technology and the increasing demand for increasingly complex space missions, spacecraft clusters have become an important trend for future development. By coordinating multiple satellites with each other, the R & D difficulty and maintenance cost of the space system can be reduced, and special space missions that cannot be performed by traditional single satellites can be completed. Most of the above game countermeasure models are relatively single, unable to establish corresponding countermeasure models according to specific scenarios and requirements, and it is also difficult to cope with the impact of target uncertain behaviors in space game missions on spacecraft game decision-making and control.

[0005] Based on the research on game decision-making for spacecraft to avoid target threats based on machine learning, it is also necessary to further explore the intelligent decision-making method for spacecraft to autonomously avoid space threats and approach targets. The key to solving this problem lies in designing a decision-making method that is effective for various threat methods and enables the aircraft to operate independently. At the same time, the designed decision-making method should not only consider the spatial relationship between agents and between agents and the target aircraft, but also further study the operation efficiency of the method, so that the multi-agent system can make decisions and responses in a short time. However, the current related research is all based on offline planning algorithms. Only scholars have studied collision avoidance between spacecraft formations, and no one has proposed an online interception and radiation decision-making method for facing multiple interceptors. In summary, it is of great significance to construct an aircraft intelligent agent unit and study the collaborative threat avoidance and target approach of space vehicle clusters. Summary of the Invention

[0006] The purpose of this application is to provide a high-value spacecraft cluster intelligent escort method. Based on the background of artificial intelligence application in spacecraft clusters, aiming at the need for spacecraft to avoid threats and approach space targets, an intelligent decision-making mechanism and machine learning method are applied to study the collaborative control technology for spacecraft cluster proximity rendezvous based on machine learning.

[0007] To achieve the above purpose, the first aspect of this application provides a high-value spacecraft cluster intelligent escort method, including:

[0008] Obtain the initial data of the spacecraft cluster, where the spacecraft cluster includes multiple spacecraft with autonomous decision-making capabilities;

[0009] According to the initial data, combined with the spacecraft orbit model, satellite on-board sensor model, and illumination model, conduct environmental deduction on the spacecraft cluster to obtain the current simulated environment;

[0010] Build a behavior rule library and a capability rule library for the spacecraft cluster;

[0011] Set a reward function, and establish a reinforcement learning model according to the multi-agent reinforcement learning algorithm and the reward function, so as to obtain a multi-agent decision result for the problem of spacecraft cluster approaching;

[0012] Iteratively train the reinforcement learning model according to the current simulation environment, the behavior rule library and the capability rule library until the reinforcement learning model converges;

[0013] Simulate the tasks in the real scenario according to the converged reinforcement learning model to obtain the final multi-agent decision result, and perform visualization processing on the multi-agent decision result.

[0014] In one implementation, the environmental deduction of the spacecraft cluster based on the initial data in combination with the spacecraft orbit model, the satellite on-board sensor model, and the illumination model to obtain the current simulation environment includes:

[0015] Predict the flight orbit of the spacecraft according to the initial data and the spacecraft orbit model;

[0016] Determine the payload beam pointing of the spacecraft according to the initial data and the satellite on-board sensor model;

[0017] Determine the observation impact of the illumination condition on the spacecraft according to the initial data and the illumination model.

[0018] Perform environmental deduction according to the flight orbit of the spacecraft, the payload beam pointing, and the observation impact to obtain the current simulation environment.

[0019] In one implementation, the spacecraft orbit model includes an ideal orbit model and a spacecraft orbit maneuver guidance model;

[0020] Among them, the ideal orbit model is used to obtain the position vector and velocity vector of the spacecraft at any time in the absence of the influence of other external disturbing forces, and the spacecraft orbit maneuver guidance model is used to obtain the orbit change strategy of the spacecraft considering the influence of external disturbing forces.

[0021] In one implementation, the satellite on-board sensor model is used to determine the payload beam pointing according to the size of the rotation angle of the two-axis positioning mechanism in the payload system of the spacecraft.

[0022] In one implementation, the illumination model specifically calculates the corresponding illumination condition critical angle according to the illumination condition at the current moment to determine the observation impact of the illumination condition at the current moment on the spacecraft, where the illumination condition includes: the earth occlusion condition and the earth light condition, the earth shadow condition, the sunlight and moonlight condition, and the earth atmosphere light condition.

[0023] In one embodiment, the construction of the behavior rule library and the ability rule library for the spacecraft cluster includes:

[0024] Construct the behavior rule library and the ability rule library for the spacecraft cluster, and associate and interact information among the spacecrafts in the spacecraft cluster according to the behavior rule library and the ability rule library to obtain the optimal strategy for spacecraft rendezvous. Among them, the behavior rule library includes:

[0025] The movement rule is used to describe that in the absence of enemy threats, the spacecraft moves towards the target according to the established recursive direction, where the enemy threats include radiation satellites and interference satellites;

[0026] The threat avoidance rule is used to describe that in the presence of enemy threats, the spacecraft changes its moving direction to avoid threats during the process of searching for the target;

[0027] The information transfer rule is used to describe that when the number of spacecrafts successfully approaching the target is greater than or equal to 1, the number of subsequent spacecrafts approaching the target simultaneously increases;

[0028] The path selection rule for increasing sensing information is used to describe that information is transferred between spacecrafts through electromagnetic signals, and the moving points of the spacecrafts are selected according to the safety values and offensive values of other different spacecrafts within the sensing range;

[0029] The reward and punishment rule for pheromone update is used to describe that during the process of the spacecraft approaching the target, when the distance between the spacecraft and the target is less than the preset value, a linearly increasing reward is given to the spacecraft, and when the spacecraft successfully approaches the target, a reward for approaching the target is given to the spacecraft to enhance the reference effect of the spacecraft on other spacecrafts.

[0030] In one embodiment, the ability rule library includes:

[0031] The orbit change ability rule is used to describe the thrust and torque required for the spacecraft to obtain attitude and orbit control through thrusters;

[0032] The radiation ability rule is used to describe the initial range of the radiation satellite with radiation effect. When the spacecraft is irradiated by the radiation satellite, the spacecraft is tracked by the radiation satellite;

[0033] The interference ability rule is used to describe that when the spacecraft is irradiated by the radiation satellite, the probability of the spacecraft being subjected to electromagnetic wave interference by the interference satellite increases. When the spacecraft is interfered by the interference satellite, the spacecraft is tracked by the interference satellite.

[0034] In one embodiment, the establishment of the reinforcement learning model according to the multi-agent reinforcement learning algorithm and the reward function includes:

[0035] Obtain a reward function according to the said reward and punishment rules;

[0036] Based on the reward function, establish a control model for a single spacecraft based on the proximal policy optimization algorithm;

[0037] Based on the said control model for a single spacecraft, establish control models for multiple spacecraft based on the multi-agent reinforcement learning algorithm, and obtain the reinforcement learning model according to the said control model for a single spacecraft and the said control models for multiple spacecraft.

[0038] The second aspect of this application provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps in the above first aspect or any implementation manner of the above first aspect are implemented.

[0039] The third aspect of this application provides a computer-readable storage medium. The above computer-readable storage medium stores a computer program. When the above computer program is executed by a processor, the steps in the above first aspect or any implementation manner of the above first aspect are implemented.

[0040] As can be seen from the above, this application provides a method for intelligent escort of high-value spacecraft clusters, constructs a simulation environment, a behavior rule library and a capability rule library for spacecraft clusters, and establishes a reinforcement learning model according to the multi-agent reinforcement learning algorithm and the reward function, realizing multi-star multi-task integrated reinforcement learning simulation and three-dimensional visualization simulation, and being able to complete the entire process of perception, action sequence planning and maneuvering; being able to realize multi-star training, establish corresponding models according to specific scenarios and requirements and obtain decision results, and being able to cope with the impact of target uncertain behaviors on spacecraft game decision-making and control in space game tasks; in addition, the method of this application is effective for various threat methods, and the aircraft can operate independently, not only considering the spatial relationship between agents and between agents and the target aircraft, but also further deeply studying the operation efficiency of the method, enabling the multi-agent system to make decisions and responses using the model trained by the reinforcement learning algorithm within a short time (such as within 10s), avoiding the time cost brought by the large ground-space loop; the reinforcement learning model is transferable and finally applicable to avoidance and approaching scenarios. The method provided by this application is based on the background of artificial intelligence application of spacecraft clusters, studies space cluster collaborative approaching technology and cluster aircraft decision control technology to support space cluster approaching rendezvous collaborative control technology, avoid the radiation or interference of other spacecraft, and by avoiding these spacecraft, prevent the enemy from causing radiation or interference to our high-value satellites. Description of the Drawings

[0041] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0042] Figure 1 Schematic flow chart of a high-value spacecraft cluster intelligent escort method provided by an embodiment of the present application;

[0043] Figure 2 Schematic diagram of the working principle of a payload beam positioning provided by an embodiment of the present application;

[0044] Figure 3 Schematic diagram of a J2000 inertial coordinate system provided by an embodiment of the present application;

[0045] Figure 4 Schematic diagram of the definition of the azimuth angle below the satellite body coordinate system provided by an embodiment of the present application;

[0046] Figure 5 Schematic diagram of the earth occlusion condition and the earth light condition provided by an embodiment of the present application;

[0047] Figure 6 Schematic diagram of the earth shadow condition provided by an embodiment of the present application;

[0048] Figure 7 Schematic diagram of the sunlight cone and the moonlight cone provided by an embodiment of the present application;

[0049] Figure 8 Schematic diagram of the earth-air light provided by an embodiment of the present application;

[0050] Figure 9 Schematic diagram of the satellite speed pulse provided by an embodiment of the present application;

[0051] Figure 10 Example diagram of the azimuth pitch angle below the sensor body coordinate system provided by an embodiment of the present application;

[0052] Figure 11 Schematic diagram of the PPO algorithm structure provided by an embodiment of the present application;

[0053] Figure 12 Schematic flow chart of the MAPPO training process provided by an embodiment of the present application;

[0054] Figure 13 Schematic diagram of a scenario provided by an embodiment of the present application;

[0055] Figure 14A satellite training curve graph provided by an embodiment of the present application;

[0056] Figure 15 A schematic diagram of the shortest distance between an aircraft and a target and a satellite reward curve provided by an embodiment of the present application;

[0057] Figure 16 A schematic diagram of the minimum distance between a satellite and a target, the satellite reward value, and the total reward value provided by an embodiment of the present application;

[0058] Figure 17 A schematic diagram of the trajectory curves of satellites 1-4 in the J2000 coordinate system provided by an embodiment of the present application;

[0059] Figure 18 A schematic diagram of the trajectory curves of satellites 5-8 in the J2000 coordinate system provided by an embodiment of the present application;

[0060] Figure 19 An example diagram of a satellite approaching a target in a SpaceSim simulation environment provided by an embodiment of the present application;

[0061] Figure 20 An example diagram of multiple satellites approaching a target under an optimal strategy provided by an embodiment of the present application;

[0062] Figure 21 A schematic diagram of the visualization of game information executed according to an optimal strategy provided by an embodiment of the present application. Detailed implementation manners

[0063] Next, in combination with the accompanying drawings in the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0064] Embodiment 1

[0065] The embodiment of the present application provides a method for intelligent escort of a high-value spacecraft cluster, as Figure 1 shown, the method includes:

[0066] S100 Obtain the initial data of the spacecraft cluster, where the spacecraft cluster includes multiple spacecraft with autonomous decision-making;

[0067] Optionally, the initial data includes the corresponding data required for calculating the spacecraft orbit model, the satellite on-board sensor model, and the illumination model in the subsequent steps, as well as the corresponding data required for establishing the reinforcement learning model, which will not be elaborated here.

[0068] In one embodiment, a spacecraft cluster is a cluster composed of multiple spacecrafts, that is, a multi-agent system. Each spacecraft in the cluster can be regarded as an agent unit, called a sub-spacecraft. The spacecraft cluster carries payloads and approaches the target relying on the cluster. Each agent unit has the ability to sense space objects within a certain distance or a strong maneuvering ability. Through inter-satellite information interaction and intelligent strategies within the cluster, radiation or interference from other spacecrafts (such as radiation satellites and interference satellites) can be avoided. By avoiding these spacecrafts, the radiation or interference to our high-value satellites by the enemy can be avoided.

[0069] S200 performs environmental deduction on the spacecraft cluster according to the initial data, in combination with the spacecraft orbit model, the satellite on-board sensor model, and the illumination model, and obtains the current simulated environment;

[0070] Optionally, an important task in the design and implementation of space missions is to predict the flight orbit of spacecrafts. Orbit extrapolation is a mathematical algorithm for predicting the flight orbit of spacecrafts under given initial conditions. According to different application scenarios, different extrapolation techniques are required. In the embodiments of the present application, predictions can be made through various extrapolation algorithms such as the ideal orbit model and the spacecraft orbit maneuver guidance model. Among them, the ideal orbit model is used to obtain the position vector and velocity vector of the spacecraft at any time in the absence of the influence of other external disturbing forces, and the spacecraft orbit maneuver guidance model is used to obtain the orbit change strategy of the spacecraft considering the influence of external disturbing forces.

[0071] In one embodiment, the ideal orbit model means that there is no influence of other external disturbing forces. Given the position vector and velocity vector of a satellite (i.e., a spacecraft) at time t0, the position vector and velocity vector of the satellite at any time t≥t0 can be solved. Specifically, given the six orbital elements A, e, i, Ω, ω, M and time t0, in the orbital coordinate system, the position vector R of the satellite es is

[0072]

[0073] Let R es be transformed into the J2000 coordinate system, and after arrangement, the position vector and velocity vector are respectively

[0074]

[0075] In the formula:

[0076]

[0077] where a is the semi-major axis of the orbit, f is the true anomaly, E is the mean anomaly, P is the first trigonometric conversion vector, Q is the second trigonometric conversion vector, and μ is the gravitational coefficient.

[0078] In one embodiment, the determination of the transfer orbit needs to consider factors such as the launch vehicle capability, satellite measurement and control time requirements, the launch site geographical location and launch direction restrictions. Generally speaking, the design result of the transfer orbit depends more on the size of the carrying capacity and the latitude of the launch site. The perigee height mainly depends on the capability and measurement and control requirements of the launch vehicle. The selection of the perigee height should not be too low to avoid excessive influence of atmospheric drag on the satellite orbit and attitude. It is generally not less than 200km. The size of the orbital inclination is related to the latitude of the launch site. For launch vehicles that do not change the orbital plane, the minimum inclination of the transfer orbit is the latitude of the launch site. However, for launch vehicles that can change the orbital plane, the inclination can be reduced as much as possible by using the rocket's margin. The apogee height should at least reach the geosynchronous altitude. If the launch vehicle has a margin of capacity, the apogee height can be further raised to save the consumption of propellant for satellite orbit change, but the height cannot be increased indefinitely. The limitations of on-board measurement instruments and measurement and control links need to be comprehensively considered. The design of the orbit change strategy provided in the embodiment of the present application is mainly formulated based on the theoretical transfer orbit, the parameters of the standard engine and the required fixed-point longitude, and the design results such as the number of orbit changes, the start time of each ignition, the length of ignition time, the ignition direction, and the amount of propellant consumed are obtained. This result is the basis for propellant budgeting, formulating flight procedures and launch windows, etc. Before each ignition, it is necessary to adjust the orbit change strategy based on the measurement results of the actual orbit and the calibration results of the orbit change accuracy. The following modeling and description are carried out for the satellite maneuver orbit change strategy and the control method for executing the orbit change:

[0079] Any orbital maneuver must take into account the fuel consumption Δm during the orbit change process and the engine operation time Δt when the expected effect is achieved. The consumed fuel is obtained by the Tsiolkovsky formula:

[0080]

[0081] If we assume that the engine consumption per second is is a constant, then the time required for the engine to work to achieve the speed increment is:

[0082]

[0083] Where m(t0) is the mass of the spacecraft before the orbit change, Δv is the velocity increment required for the orbit change, and I sg is the specific impulse of the engine.

[0084] Optionally, the satellite-borne sensor model is used to determine the payload beam pointing direction according to the size of the rotation angle of the dual-axis positioning mechanism in the payload system of the spacecraft.

[0085] In one embodiment, in the payload system, the payload feed is kept fixed, and a two-axis positioning mechanism is applied to enable the reflector to rotate within an angular range of ±4° in the azimuth and elevation directions, so as to realize the movement of the beam within the ±8° region. The working principle of the movable spot beam offset dual-reflector payload beam positioning is as follows Figure 2 as shown (in the figure, it is the initial pointing of the payload reflection line, and the feed and sub-reflector payload are not drawn). By changing the magnitude of the rotation angle of the two-axis positioning mechanism, the pointing of the payload beam can be changed, so as to realize earth orientation or target tracking. In the figure, O e is the origin of the geocentric coordinate system, O s is the satellite center, point A is the focus position of the sub-reflector payload, which coincides with the focus of the reflecting paraboloid initially, point B is the rotation center of the two-axis positioning mechanism, point C is the reflection center of the paraboloid payload, CD is the beam pointing, and its pointing angle is described by the azimuth angle and elevation angle in the satellite body coordinate system. In this system, the feed system remains stationary, and the change of the reflection beam pointing is realized by the deflection of the parabolic reflector. Three basic transformations are required to complete the transformation from the payload pointing to the geographical relative longitude and latitude to the rotation angle of the payload positioning mechanism: namely, the direction angle of a certain point on the ground in the satellite coordinate system, the satellite attitude change compensation, and the calculation of the rotation angle of the satellite direction angle to the rotation mechanism; all the above three basic transformation steps involve forward and reverse calculations in the calculation process, as Figure 2 shown. In the transformation calculation shown in the figure: it is relatively easy to obtain the spatial pointing of the reflection line given the rotation angle of the mechanism, while calculating the rotation angle of the mechanism according to the spatial pointing of the reflection line, due to the non-coincidence of the reflection center of the paraboloid payload and the rotation center of the positioning mechanism, and the rotation center is located on the extension line of the normal line passing through the reflection center of the paraboloid payload, the transformation between the spatial pointing angle of the payload beam and the rotation angle of the two-axis positioning mechanism shows a complex non-linear relationship, which is a non-linear equation solving problem and no analytical solution of the equation can be obtained, and only numerical approximation can be obtained through numerical methods. Therefore, it is relatively easy to calculate the direction angle of the payload pointing given the geographical relative longitude and latitude, while calculating the geographical longitude and latitude according to the spatial pointing of the reflection line requires solving a quadratic algebraic equation containing trigonometric functions, or calculating using other algorithms suitable for computer processing.

[0086] Assuming that the satellite is in an ideal state, the basic parameters related to the transformation include: r e - the radius of the earth, r s - the satellite orbit radius (to the center of the sphere), k e - the longitude of a certain point on the ground, l e - the latitude of a certain point on the ground, k s - the fixed-point longitude of the satellite, l s - the fixed-point latitude of the satellite, k p - the relative longitude with the fixed-point position of the satellite as the origin, l p - the relative latitude with the fixed-point position of the satellite as the origin, where kp = k e -k s , l p = l e -l s . First, define the coordinate system: If the longitude of point D is greater than the longitude of the satellite l p is positive, and if the latitude of point D is greater than the latitude of the satellite k p is positive. Spherical coordinate system: The center of the sphere is the origin, Y e points to the due north, Z e points to the satellite, and X e satisfies the right-hand rule. Ideal satellite coordinate system: The center of the satellite is the coordinate origin, Y s points south, Z s points to the center of the earth, and X s points east and satisfies the right-hand rule. Load rotation mechanism coordinate system: The load rotation coordinate system is represented by the two rotation axes of the load (X apm , Y apm ). X apm is on the symmetry axis of the reflector surface, one end is fixed to the reflector surface, and Y apm is on the non-symmetry plane of the reflector surface, and one end is connected to the load support. In the orbit, X apm is positive towards the east, Y apm is positive towards the south, and Z apm satisfies the right-hand rule, and it is the normal direction of the load reflector surface at its reflection center. The rotation angle of the drive mechanism is represented by (α, β), where α is the pitch angle of rotation around the X apm axis, and β is the azimuth angle of rotation around the Y apm axis (it is stipulated that α and β are positive in the counterclockwise direction). This coordinate system is also the following B-x b y b z b coordinate system. The load surface focus coordinate system and the reflection center fixed coordinate system are as Figure 2 shown, and the J2000 inertial coordinate system is as Figure 3 shown. In the load system, take A-xyz as the focus coordinate system, B-x b y b z b as the mechanism rotation coordinate system with the origin at the rotation center of the positioning mechanism, C-x c y c z c as the paraboloid reflection center fixed coordinate system. In the figure, h is the height of the initial load reflection center from the yz plane in the focus coordinate system A-xyz, β c is the angle between the incident ray AC and the yz plane, and f c is the focal length of the reflecting paraboloid. Then, in the A-xyz coordinate system, the equation of the reflecting paraboloid is

[0087] x 2+y 2 = 4f(z + f c )

[0088] The coordinates of point B, the rotation center of the driving mechanism, are

[0089] [(Lsin(β c / 2) + h), 0, -Lcos(β c / 2) - (f - h∧2 / 4f c )] T

[0090] Calculating the load direction from the fixed-axis mechanism rotation angle: If the rotation angle of the dual-axis positioning mechanism is the rotation angle α around the y b axis and the rotation angle β around the x b axis, the transformation of any point in space in the coordinate system C- x c y c z c and A-xyz can be described by the direction cosine matrix and the translation vector.

[0091] Calculating the fixed-axis mechanism rotation angle from the beam direction: According to the geometric optics principle, as shown in Figure 2 , the straight lines BC, CD, BA, and CA are coplanar. Let the reverse extension line of the reflected ray CD intersect BA at point E, and l bc be the length of the BC connection. The required rotation angle of the positioning mechanism is as follows:

[0092]

[0093] The sensor azimuth angle is measured in the right-hand direction around the +X axis, and the azimuth angle on the +Z axis is zero. The elevation angle is measured from the XY plane towards the +Z axis. As shown in Figure 4 , where, in this system, the relationship between the sensor pointing vector (x body , y body , z body ) and the sensor azimuth angle and elevation angle is:

[0094]

[0095] Converting the description (x body , y body , z body ) in this system to the description l ant = (x, y, z) in the inertial system, calculating the connection vector l sat = (x l , y l , z l ) between the radiating and the radiated satellites, and the included angle <l ant , l sat>Less than the beam half-angle, and if the radiated satellite is within the maximum range of the radiation effect, it is determined to be radiated to the target.

[0096] Optionally, the illumination model specifically calculates the corresponding critical angle of the illumination condition according to the illumination condition at the current moment to determine the observation impact of the illumination condition at the current moment on the spacecraft. Among them, the illumination conditions include: the earth occlusion condition and the earthshine condition, the earth shadow condition, the sunlight and moonlight condition, and the earth-airglow condition.

[0097] In one implementation, when there is no occlusion between the space target and the space-based observation platform, the space target can be geometrically visible to the space-based observation platform. For space targets in the Earth's orbit, the main occlusion comes from the Earth's occlusion. In addition to the occlusion caused by the Earth's body, due to the reflection of sunlight by the Earth and the atmospheric refraction phenomenon, the sensors of the space-based observation platform cannot image space targets that enter the Earth's apparent disk and its nearby areas. This situation is called the earthshine effect. The Earth's occlusion effect can be included in the earthshine effect. The observation geometry model is as Figure 5 shown. In the figure, H0 is the height of the critical line of sight from the ground under the earthshine condition, R e is the radius of the Earth, r O is the geocentric vector of the space-based observation platform in the ECI coordinate system, r T is the geocentric vector of the space target in the ECI coordinate system, r OT is the position vector of the space target relative to the space-based observation platform. When the angle θ OT between r O and r ET is greater than the critical angle θ ETO of the earthshine condition, the space target can be observed by the sensors of the space-based observation platform. Among them, the critical angle θ ETO of the earthshine condition can be expressed as

[0098]

[0099] After the space target enters the earth shadow region, since it cannot be irradiated by sunlight, the sensors of the space-based observation platform cannot image it. The observation geometry model under the earth shadow occlusion condition is as Figure 6 shown. In the figure, r S is the geocentric vector of the sun in the ECI coordinate system. During the calculation of the earth shadow, the sun can be equivalent to a light source at infinity, and the sunlight can be regarded as parallel light. Therefore, the earth shadow region can be simplified into a cylinder with the direction of the sun vector as the center line. Let the angle θ T between the geocentric vector r S of the space target and the geocentric vector r ST of the sun, and the critical angle θ ES0 corresponding when the space target enters the earth shadow region can be expressed as

[0100]

[0101] When θ ST <π - θ ES0 , the space target is not within the earth's shadow area and can be irradiated by the sensor of the space-based observation platform.

[0102] The sunlight condition means that due to the large light intensity, the sensor of the space-based observation platform cannot image the space target entering the solar disk or its nearby area. The moonlight condition means that due to the large reflected light intensity on the side of the moon facing the sun, the sensor of the space-based observation platform cannot image the space target entering the lunar disk or its nearby area. Figure 7 is the visible range under sunlight and moonlight conditions. In the figure, r M is the geocentric vector of the moon in the ECI coordinate system. The cone angle θ ST0 of the sunlight cone and the cone angle θ MT0 of the moonlight cone can be expressed as

[0103]

[0104] In the formula, θ ST1 is the solar apparent radius, θ ST2 is the solar light scattering angle, θ MT1 is the lunar apparent radius, θ MT2 is the lunar light scattering angle. When the space target is outside the sunlight cone and the moonlight cone, that is, θ ST >θ ST0 , θ MT >θ MT0 , the space target can be irradiated by the sensor of the space-based observation platform. The critical angle values for the sunlight and moonlight conditions refer to the literature, and the cone angles of the sunlight cone and the moonlight cone are respectively set to θ MT0 = 10°, θ ST0 = 10°.

[0105] When analyzing the earth's atmospheric glow, the earth cannot be simplified as a point light source or the earth's albedo light regarded as parallel light. Instead, it should be regarded as an extended curved surface light source. Therefore, it is necessary to divide the earth's surface into grids. The irradiance transfer model of the earth's atmospheric glow is as Figure 8As shown in the figure. The Earth's surface is divided into grids according to longitude and latitude. j and w represent the longitude angle and latitude angle of the ground surface element respectively; θ represents the angle between the line connecting the ground surface element and the sun (i.e., the opposite direction of the incident sunlight) and the normal of the ground surface element, which is called the illumination angle; ε represents the angle between the line connecting the ground surface element and the space-based observation and the normal of the ground surface element, which is called the emission angle; σ represents the angle between the line connecting the space-based observation and the ground surface element and the optical axis of the telescope, which is called the off-axis angle. It can be seen from the figure that the condition for the earth-atmosphere light to generate stray light on the baffle is 0 < θ, ε, σ < π / 2. When θ < π / 2, it indicates that the ground surface element is in the sunlit area; when ε < π / 2, it indicates that the reflected light of the ground surface element can be incident on the space-based observation body; when σ < π / 2, it indicates that the ground surface reflected light can enter the space-based observation entrance pupil.

[0106] S300 constructs the behavior rule library and the ability rule library of the spacecraft cluster;

[0107] Optionally, construct the behavior rule library and the ability rule library of the spacecraft cluster, so as to load the behavior rules and the ability rule library in the subsequent task scheduling, and customize various parameters in the task process according to the settings of the rule library, such as the designated two-party programs, the satellites used, the number of innings, the scene selection, etc., and at the same time output the status data to the training forwarding environment.

[0108] Optionally, regard a single spacecraft in the spacecraft cluster as an autonomous decision-making intelligent agent. They are the basic and core elements of the model. Each agent will perceive the surrounding environment and make decisions and take actions according to the rules. In the cooperative control model of the spacecraft cluster, each agent spacecraft needs to follow the following most basic rules to associate and interact information among the spacecrafts in the spacecraft cluster to obtain the optimal strategy for spacecraft rendezvous:

[0109] The movement rule is used to describe the rule for the spacecraft to recursively orbit and move towards the target without enemy threats. The movement rule first needs to clarify the internal variables of the spacecraft, that is, the status of the number of successfully approaching satellites. If the number of successfully approaching satellites = 0, it means that all spacecrafts are approaching the target. Then, first look for whether the target has been approached within the sensing radius. If so, move over, the number of successfully approaching satellites is converted to 1, and the information transfer rule is executed; if not, look for threats and sense the levels of enemy reconnaissance and interference, and then move towards the position with a lower threat level. If there is no threat within the sensing range, the spacecraft will move in the established recursive direction.

[0110] The threat avoidance rule is used to describe that in the case of enemy threats, the spacecraft changes its moving direction to avoid threats during the process of searching for the target; when there is likely to be a threat between the current point and the target point determined by the movement rule, the spacecraft will change the main direction, obtain a new target point according to the new main direction, and then judge again whether there is a threat between the current point and the target point.

[0111] An information transfer rule, which is used to describe that when the number of spacecraft successfully approaching the target is greater than or equal to 1, the number of subsequent spacecraft approaching the target simultaneously increases, so as to achieve the rendezvous of the spacecraft cluster and the target;

[0112] A path selection rule for increasing sensing information. There is a certain amount of information transfer between spacecraft, and a certain amount of information can be transmitted through electromagnetic signals. When a spacecraft goes in many wrong directions on its flight path, the remaining safety value of the spacecraft will be less than that of the spacecraft that has not taken a detour. That is, the higher the safety value, the more reference value the path taken by the spacecraft has. Therefore, a spacecraft can see the amount of threat carried by different types of spacecraft within its sensing range and the offensive value to the target. The spacecraft will make a choice between the point where the spacecraft with the highest safety value is located and the maximum offensive value according to the training results to select the moving point of the spacecraft.

[0113] A reward and punishment rule for pheromone update. Due to the influence of threats, the movement of spacecraft has greater randomness. When it moves to the target point, it may deviate from the optimal path. Then its training experience will mislead other spacecraft to also deviate from the optimal path, thus affecting the search for the target point. To prevent this kind of misguidance, the embodiment of the present application designs a reward and punishment rule. By changing the magnitude of the path offensive value, when the spacecraft is approaching the target, when the spacecraft is at a position relatively close to the target point, a gradually linearly increasing reward is given to the spacecraft; if it successfully approaches the target, the path taken by the spacecraft has a certain reference value. Therefore, a relatively large reward is given to the spacecraft to enhance the reference effect of the spacecraft on other spacecraft.

[0114] Optionally, the ability rule library includes: an orbit transfer ability rule, which is used to describe the thrust and torque required for a spacecraft to obtain attitude and orbit control through thrusters; a radiation ability rule, which is used to describe the initial range of a radiation satellite with a radiation effect. When a spacecraft is irradiated by a radiation satellite, the spacecraft is tracked by the radiation satellite; an interference ability rule, which is used to describe that when a spacecraft is irradiated by a radiation satellite, the probability of the spacecraft being subjected to electromagnetic wave interference by an interference satellite increases. When the spacecraft is interfered by an interference satellite, the spacecraft is tracked by the interference satellite.

[0115] In one implementation, a spacecraft can obtain the thrust and torque required for attitude and orbit control by installing thrusters. Assume that n thrusters are installed on the spacecraft, and assume that the directions of 3 inertial principal axes coincide with the body-fixed coordinate system Ox b y b z b The position vector d of thruster i relative to the center of mass of the spacecraft i =(x i ,y i ,zi ) e = d i T e, x i y i z i respectively represent the components of the thrust of the thruster on the three coordinate axes. e = [e x e y e z T are the three basis vectors of the body coordinate system. The unit thrust vector matrix generated by the i-th thruster is e i = (cosα i cosβ i , cosα i sinβ i , sinα i ), as shown in Figure 9 . Thus, the velocity pulse for spacecraft orbit transfer can be obtained.

[0116] In one implementation, the spacecraft radiation rule, i.e., the initial range of the satellite with radiation effect, is approximately known with an error of ±30 km, and the algorithm is shown in Table 1.

[0117] Table 1 Algorithm for Using Spacecraft Radiation Rule

[0118]

[0119] Among them, the input: the initial two target positions (the first 1 - 19 and the second 20 - 40 values). Since they are fuzzy values, the function of whether it is irradiated is false. By training the sensor to rotate so that the target is irradiated for three consecutive steps, the agent is considered to be irradiated. Then, the radiation sensor tracking instruction is executed, and the radiation sensor can always follow the target. The rotation angle of the radiation sensor during search is in gears (angle size). Observed in the xoz orbital plane and placed in the body coordinate system of the radiation sensor, it is equivalent to the azimuth angle being always 0, and the pitch angle θ takes values of ±90 degrees. The definition of the azimuth and pitch angles in the body coordinate system of the satellite is as shown in Figure 4 . The example of the azimuth and pitch angles in the body coordinate system of the sensor is as shown in Figure 10 .

[0120] In one implementation, the spacecraft interference ability rule means that after the satellite agent is irradiated, it will be more vulnerable to electromagnetic interference. At this time, the interfering satellite may be within the beam range of the high-value payload sensor of the satellite agent, and the interference opportunity of the interfering satellite will cause interference to the satellite agent. The specific algorithm used in the implementation is shown in Table 2.

[0121] Table 2 Spacecraft Interference Tools

[0122]

[0123] ​

[0124] Among them, the inputs are: the position of the agent, the velocity of the agent, the position of the interfering satellite, the maneuvering velocity and direction of the interfering satellite, such that the target is interfered with for three consecutive steps, even if the agent is interfered with. Then the sensor tracking instruction is executed, and the interference sensor can always follow the target, that is, the agent fails.

[0125] The S400 sets a reward function and establishes a reinforcement learning model according to the multi-agent reinforcement learning algorithm and the reward function to obtain a multi-agent decision result for the problem of spacecraft cluster approaching.

[0126] In one implementation, to reasonably train the problem of spacecraft cluster approaching, achieve a good learning effect, and obtain a convergent and reasonable training result, it is necessary not only to reasonably design the spacecraft behavior rules and ability rules, but also to establish a correct and effective reinforcement learning model, set a correct and reasonable reward function, and an algorithm training process, so as to accurately simulate the tasks in the real scenario and achieve a good learning effect. When making decisions for multi-agent training, the control decisions exist in the form of separate sub-processes and are isolated from each other, mainly including two modules: the reinforcement learning algorithm and the reward rule construction. Among them, the reinforcement learning algorithm (Reinforcement Learning, RL), that is, the agent decision-making algorithm, is used to receive observation data and output action control data. Reinforcement learning is one of the methods of machine learning. Compared with other machine learning methods, it focuses more on goal-oriented learning from interactions. The trainer is not told what actions to take or has reference actions for learning, but instead tries many times to find actions that can generate higher rewards and uses them as its own strategy. In the most challenging situations, such as the spacecraft game problem, actions may not only affect the direct reward, but also affect the next situation and thus all subsequent rewards. The training and solution of reinforcement learning do not require relying on an accurate mathematical model, the calculation is simple but requires many training attempts. A multi-agent system refers to a system composed of multiple agents, which is considered to be the closest natural way to view the system and provides a distributed perspective for viewing the entire system. The process by which each agent in a multi-agent environment obtains a reward by interacting with the environment to improve its own strategy and thus obtains the optimal strategy in this environment is multi-agent reinforcement learning.

[0127] Optionally, establishing a reinforcement learning model includes: obtaining a reward function according to the reward and punishment rules; establishing a single spacecraft control model based on a proximal policy optimization algorithm (PPO) according to the reward function; establishing multiple spacecraft control models based on a multi-agent reinforcement learning algorithm (Multi-agent PPO, MAPPO) according to the single spacecraft control model, and obtaining the reinforcement learning model according to the single spacecraft control model and the multiple spacecraft control models.

[0128] In one embodiment, in order to solve the problem of how to implement the spacecraft pursuit game, it is proposed to select a reinforcement learning algorithm suitable for the problem based on the characteristics of the spacecraft pursuit game problem, so as to achieve a better simulation effect on the spacecraft pursuit game. The spacecraft pursuit game problem has the following characteristics: the spacecraft kinematic equation is relatively complex, the relative motion relationship is difficult to learn, and the orbital dynamics are significantly affected by long-term games; the thrust of the spacecraft chemical propulsion is large, and the overall game time is long. It is necessary to constrain the behavior of the intelligent agent and appropriately reduce the exploration space. Otherwise, although the intelligent agent will learn the action strategy, because its initial neural network is randomly generated and there is no prior knowledge, if the action is not restricted, it will cause the behavior to be erratic and even self-contradictory, and it is almost impossible to guarantee the target. The entire game cycle is long, and the relative motion characteristics of the spacecraft orbit are more obvious. It is necessary to design the reward function for the state space and behavior space in the spacecraft reinforcement learning game, otherwise the training result is less than ideal. In order to solve the problem of continuous action space, the embodiment of this application uses PPO and MAPPO reinforcement learning algorithms.

[0129] S500 iteratively trains the reinforcement learning model according to the current simulation environment, the behavior rule library, and the capability rule library until the reinforcement learning model converges;

[0130] In one implementation, PPO is a single-agent reinforcement learning algorithm that adopts the classic actor-critic architecture. Among them, the actor network, also known as the policy network, receives local observations (obs) and outputs actions; the critic network, also known as the value network, receives states (state) and outputs action values (value) to evaluate the quality of the actions output by the actor network. It can be intuitively understood as a judge (critic) scoring (value) an actor's (actor) performance (action). This reinforcement learning algorithm proposes a new objective function, which can achieve mini-batch updates in multiple training steps, and solves the problem that the step size is difficult to determine in the traditional Policy Gradient algorithm. This algorithm defines the state space and the action space, and designs the reward function, so as to realize the control algorithm of the attacker under the one-on-one pursuit-evasion problem, making it catch up with the defender as quickly as possible, as Figure 11 shown. For the attacker spacecraft agent, at time t, its own local state value is S_t, and the input of the Actor network is its own local observation value. To reduce the size of the action space, referring to the SAC architecture of Berkeley, different from the assumption in the continuous action space that the Q function takes the state and action as input and outputs the Q value, now a state will be obtained and mapped to a vector containing the Q values of each action. In the continuous setting, π is represented by parameters θ, and these parameters are used to calculate the mean and variance of each feature in the action space. So parameterizing π maps the state to a real vector with 2|A| elements, where |A| is the dimension of the action space. In the discrete setting, a simple probability vector for each action is sufficient. Therefore, π maps the state to a probability vector with |A| elements. The propulsion direction angle is set to discrete values, ranging from -90°, 90°, with an interval of every 3 degrees, and there are a total of A = 60 optional propulsion direction angles. When marked as training, the sample function is used to sample from the normal distribution to obtain the action a i , when marked as testing, the argmax function is used to directly obtain the maximum value to obtain the action a i , removing randomness. The spacecraft executes the action a i and stores the process information (s i , a i , r i , s' i ) in the experience replay pool.

[0131] After the experience replay pool stores a certain number of experiences, it starts to randomly sample from the experience replay pool for network training. During the training process, the input of the critic network (Critic) is the state s of the agent t , and the output is the state value V corresponding to the agent RED , and its optimized loss function is:

[0132]

[0133] Wherein, V RED represents the output value of the critic network of the attacking spacecraft agent; is the target value calculated through the Bellman equation.

[0134] In one implementation, MAPPO is a variant of the PPO algorithm applied to multi-agent tasks. It also adopts the actor-critic architecture. The difference is that at this time, the critic learns a centralized value function, which can observe global information, including the information of other agents and the information of the environment. Its basic idea is: centralized learning and decentralized execution (CTDE). There are four networks in the MAPPO algorithm: Actor, Target, Actor_Critic, and Target_Critic. The Target network is periodically copied from its corresponding network. The Actor network, also called the policy network, is responsible for making decisions for the agent. Its input is the state State of the agent. The Critic network is generally called the critic and is responsible for evaluating the value of the decision made by the Actor. To meet the training of multi-agents, MAPPO improves the PPO algorithm. The Critic is extended to be able to learn using the policies of other agents, that is, each agent performs a function approximation on the policies of other agents, as shown in Figure 12 shown. In MAPPO, θ = [θ1,..., θ n represents the parameters of the decisions of n agents, and π = [π1,..., π n represents the policies of n agents. Then the cumulative expected reward of the i-th agent is:

[0135]

[0136] For the stochastic policy, the policy gradient is:

[0137]

[0138] Where, o i represents the observation of the i-th agent, and x = [o1,..., o n represents the observation vector, that is, the state. represents the centralized state-action function of the i-th agent. Since each agent independently learns its own The function, so each agent can have a different reward function, and thus can complete cooperative or competitive tasks.

[0139] For the deterministic policy, the gradient formula is:

[0140]

[0141] where D is the experience replay buffer, and the element composition is (x, x', a1,..., a n , r1,..., r n ), is the Actor_Critic network.

[0142] An inherent problem in multi-agent reinforcement learning is that due to the update and iteration of each agent's policy, the environment is dynamically unstable for a specific agent, resulting in poor simple experience replay. This situation is especially serious in competitive tasks, and often an agent overfits a strong policy against its competitor. However, this strong policy is very fragile and not what we want, because as the competitor's policy updates and changes, this strong policy is difficult to adapt to the opponent's new policy.

[0143] To better handle the above situation, MAPPO proposes an improvement to the actor-critic method, which considers the action strategies of other agents and can successfully learn strategies that require complex multi-agent coordination. Specifically, it is an idea of a set of strategies. The policy μ of the i-th agent i consists of a set with K sub-policies, and only one sub-policy is used in each training episode (abbreviated as ). For each agent, maximize the overall reward of its set of strategies:

[0144]

[0145] Construct a memory repository for each sub-policy K Overall optimize the entire set of strategies, so the update gradient of each sub-policy is:

[0146]

[0147] It can be seen that Critic borrows global information for learning, while Actor only uses local observation information. MAPPO solves the problem of unstable multi-agent training environment. If the actions of all agents are known, the environment is stable, and even if the policy is continuously updated, the environment remains constant. At the same time, since each agent independently learns its own policy, MAPPO can also be applied to cooperative or competitive tasks.

[0148] In one implementation, corresponding observation and reward information is constructed for each intelligent algorithm, and the observation and reward data organization forms that meet the requirements of the reinforcement learning algorithm are generated from the original state data. The constructed reward rules can be adjusted according to the actual situation.

[0149] S600 simulates the tasks in the real scenario according to the converged reinforcement learning model to obtain the final multi-agent decision result, and visualizes the multi-agent decision result.

[0150] Optionally, the visualization interface uses SpaceSim. By inputting satellite data and decision results into this software, the text and graphical display of the simulation results, the output of important parameters, and simulation control and task restoration can be performed through the interface.

[0151] In an application scenario, the embodiments of the present application are implemented in the form of computer simulation. The spacecraft simulation part is carried out by designing a SpaceSim aerospace simulation software, such as the environmental deduction part. The part of the reinforcement learning training calculation is written in a script using the Python language, such as the multi-agent training decision part. Among them, the SpaceSim spacecraft system simulation software can support the whole-process simulation and analysis of space missions, covering various stages such as spacecraft design, testing, launch, operation, and mission application. The architecture of the independently designed SpaceSim software can not only flexibly realize the customized development of the software and the reuse of modules, but also provide a reference for other domestic industrial software. The SpaceSim software relies on the collaborative development and application of domestic universities and aerospace research institutes, and has been successfully applied to domestic spacecraft system simulation, pre-research, and model tasks, etc. The spacecraft payload model included supports the development of satellite attitude and orbit control simulation and control mode simulation. The satellite platform model supports real-time intervention in the simulation process in the form of dynamic instructions. The simulation support tool can dynamically display the state changes such as the satellite platform orbit attitude. After comparing and verifying the accuracy and efficiency with STK, taking the satellite platform model as an example, the accuracy of its satellite orbit attitude recursion result compared with STK is better than 99%. At the same time, this software has more advantages in GPU high-performance parallel computing, space game, space optical observation, space early warning, network routing planning, task autonomous intelligent planning, and antenna gain interference calculation. It has two ways to connect with Python scripts: SpaceSimGYM or UDP network. SpaceSimGYM packages the functions in SpaceSim into PYD files for users to call, and users can directly use the functions in them to achieve the required functions; UDP is a commonly used network transmission method, and SpaceSim supports exchanging data with Python using the network. Both methods can realize the combination of space missions and reinforcement learning.

[0152] As can be seen from the above, the embodiments of the present application provide a high-value spacecraft cluster intelligent escort method, establish a reinforcement learning model to realize the collaborative approach control of the spacecraft cluster and construct a simulation environment, integrate the simulation information of the satellite payload, and realize the integrated simulation of multi-star multi-task comprehensive reinforcement learning and three-dimensional visualization simulation. It can complete the entire process of perception, action sequence planning, and maneuvering; it can realize multi-star training, establish corresponding countermeasure models according to specific scenarios and requirements, and can cope with the impact of target uncertain behaviors in space game tasks on spacecraft game decision-making and control; the multi-agent system can make decisions and responses using the model trained by reinforcement learning within 10s, avoiding the time cost brought by the large ground-space loop; the reinforcement learning model is transferable and finally applicable to the avoidance and approach scenarios.

[0153] Embodiment 2

[0154] In the embodiments of the present application, the method described in Embodiment 1 is verified through simulation experiments, and the steps are as follows: First, set the initial conditions, then train the satellite agent, and apply it to the orbital approach mission. The spacecraft observes the scene state, obtains the control quantity according to the current control strategy, and then adjusts the control strategy using the feedback of the scene to form a closed-loop training process.

[0155] Specifically, set the environmental information: In the embodiments of the present application, it is assumed that both the spacecraft and the target are near the low Earth orbit, the process time is short, and the instantaneous state information is completely known. The maximum number of rounds is 50 rounds, the duration of each round is 60S, the mass of the radiation / interference spacecraft is 8000 kg, and it is equipped with a chemical thruster with a constant thrust of 220N, a specific impulse of 300S, and supports sensitive propulsion. Therefore, the velocity increment is 12m / s every 60s. The scene step size is 60(s), and the maximum scene step size () is 50 steps, that is, the approach mission needs to be completed within 3000s.

[0156] Assume that there are a total of n red target agents, and a total of n blue satellite agents. Then the own state S t of the satellite includes its own physical state at the current moment (including the remaining fuel value Fuel ATK at the current moment), the required target information, and the information of its own satellite. Specifically, it is shown in Table 3 as follows:

[0157] Table 3 Satellite simulation scene information in the scene

[0158]

[0159] Set the scene and mechanism: To support the orbital approach process, the embodiments of the present application construct a typical scene. There are a total of three types of aircraft in the scene, namely the target satellite, the radiation and interference aircraft, and the approaching satellite. All the aircraft are distributed in the low orbit. The target satellite has 1 core aircraft RED_1, 12 radiation or interference aircraft (respectively: 6 radiation satellites and 6 electromagnetic interference satellites), and there are 8 approaching satellites, as Figure 13 shown.

[0160] The six orbital elements of the RED_1 target aircraft are shown in Table 4. The orbital altitude of the RED_1 target aircraft is about 407KM, and the radiation and interference aircraft are located near RED_1 at 100km in the outer circle and 50km in the inner circle.

[0161] Table 4 Six orbital elements of the target aircraft

[0162] Name a ecc inc raan argp nu RED_1 6778.096494 0 42 0 140.034 180

[0163] Approaching aircraft deployment: The approaching satellite has an initial distance of 200km. It approaches the target from two directions, intending to avoid radiation and interference, and approach close enough to rendezvous with RED_1. The approaching satellite initially adopts a follow-up formation. The BLUE_1, BLUE_2, BLUE_3, and BLUE_4 satellites are set to follow the target RED_1 at a fixed point distance of about 200km in one direction, and the BLUE_5, BLUE_6, BLUE_7, and BLUE_8 satellites are set to follow the target RED_1 at a fixed point distance of about 200km in another direction. The initial position speed and initial distance of the 8 approaching satellites are shown in Table 5.

[0164] Table 5 Approaching satellite initial position speed table

[0165]

[0166]

[0167] The process can be subdivided according to the time sequence:

[0168] 1) Approaching the target: Continuously advance and approach the target within 5 km. During the approach, the vehicle cannot be irradiated or interfered with.

[0169] 2) Rendezvous with the target: After reaching within 5 km, it is necessary to rendezvous with the target without being interfered by electromagnetic interference.

[0170] In this scenario, the satellite begins to approach at 2025-05-2011:20:00.000 (UTC). Eight satellites BLUE_1, BLUE_2, BLUE_3, BLUE_4, BLUE_5, BLUE_6, BLUE_7, and BLUE_8 are called, and the orbit change strategy is obtained by using the reinforcement learning algorithm. Multiple propulsion methods are used to approach the target RED_1.

[0171] The radiating and interfering spacecraft deployment is shown in Table 6:

[0172] Table 6 Radiation and interference satellite initial position velocity table (J2000 coordinate system)

[0173]

[0174] In summary, the scenario start time is: 2025-05-2011:20:00.000 (UTC), the simulation step is 60 (s), and the scenario end time is: 2025-05-2014:16:30.000 (UTC).

[0175] Next, train the satellite responsible for approaching. When training the satellite agent, select discrete action PPO, and set the Actor network architecture and Critic network architecture as shown in Table 7, and set the hyperparameters as shown in Table 8.

[0176] Table 7 Scenario Simulation Network Architecture

[0177] Actor Network Architecture Critic Network Architecture Fc1 = nn.Linear(state dimension, 64) Fc1 = nn.Linear(state dimension, 64) Tanh() Tanh() Fc2 = nn.Linear(64, 64) Fc2 = nn.Linear(64, 64) Tanh() Tanh() Fc3 = nn.Linear(64, sum of the number of actions) Fc4 = nn.Linear(64, 1)

[0178] Table 8 Scenario Simulation Hyperparameter Settings

[0179]

[0180]

[0181] The reward settings are shown in Table 9.

[0182] Table 9 Blue Side Simulation Reward Settings for the 13v8 Scenario

[0183]

[0184] The training curve is as shown in Figure 14 As shown, the shortest distance between the aircraft and the target is as shown in Figure 15 (a). It can be seen that as the training progresses to the later stage, the minimum distance continues to shrink until it converges. The satellite reward curve is as shown in Figure 15 (b). It can be seen that as the training progresses to the later stage, the reward continues to rise until it converges to the maximum value, indicating the correctness of the environment and the algorithm. Import the trained reinforcement learning model into the algorithm to obtain the final approaching orbit transfer strategy. The minimum distance between the satellite and the target under a series of strategies is as shown in Figure 16 (a). It can be seen that when the satellite does not adopt any strategy, the minimum distance between the satellite and the target is as shown by the eight straight lines in the upper half of Figure 16 (a). From the uppermost orange straight line to the pink straight line in the center of the picture, the minimum distances between Satellite_4 and the target are 85.6 km, between Satellite_3 and the target are 80.25 km, between Satellite_1 and the target are 75.13 km, between Satellite_7 and the target are 74.78 km, between Satellite_2 and the target are 70.27 km, between Satellite_6 and the target are 68.56 km, between Satellite_5 and the target are 67.08 km, and between Satellite_8 and the target are 52.67 km. Without importing the strategy, the satellite cannot approach the target in the end. After the satellite imports the strategy, the minimum distance between the satellite and the target is as shown in Figure 16(a) As shown by the eight lines in the lower part, the results under the finally obtained optimal strategy are as follows: the minimum distance of satellite_2 from the target is 4.935 km, the minimum distance of satellite_4 from the target is 4.431 km, the minimum distance of satellite_6 from the target is 4.417 km, the minimum distance of satellite_5 from the target is 4.224 km, the minimum distance of satellite_7 from the target is 4.221 km, the minimum distance of satellite_3 from the target is 4.071 km, the minimum distance of satellite_8 from the target is 3.569 km, and the minimum distance of satellite_1 from the target is 2.224 km. After importing the optimal strategy, the satellite can conduct a close approach to the target at the end.

[0185] After importing the trained reinforcement learning model, the reward values of a series of strategies obtained are as Figure 16 (b) shown. It can be seen that the reward value of the final strategy is the largest, and the maximum reward value is 57980. Therefore, the final strategy is the optimal strategy. The trajectories of the satellites and the target in the game executed according to the optimal strategy are plotted as trajectory curves in the J2000 coordinate system, as Figure 17 and 19 shown. It can be seen that all satellites can approach the target well, and the disturb step is 0 for all, indicating that the satellites are not interfered by other aircraft. In the simulation, each step is 60 seconds, and the satellites can fly 1 to 3 steps, that is, from 60 seconds to 180 seconds. Among them, the vehicle numbered satellite_6 flies to approach the corresponding target earliest, and then other satellites approach the target. It meets the design of the spacecraft cluster behavior rule: the "one-first-and-many-later" principle. Multiple satellites work together. After one satellite successfully approaches the target, the number of satellites approaching simultaneously subsequently increases greatly and lasts for a period of time.

[0186] The training results are shown in the SpaceSim simulation environment as Figure 19 shown, which is an example of satellites approaching the target in the SpaceSim simulation environment. The smaller satellites are the approaching satellites, and the conical green area represents the sensor length of the satellite. The cone touching the target represents the success of the rendezvous. Figure 20 That is, under the optimal strategy, multiple satellites conduct close approach flights to the target and ensure a certain flight time. Visualize the game information executed according to the optimal strategy. Taking satellite_1 as an example, the key information such as the propulsion speed of the satellite is visualized as Figure 21 shown.

[0187] In summary, the embodiments of the present application establish a reinforcement learning model to achieve cooperative approaching control of a spacecraft cluster, construct a simulation environment, integrate the simulation information of satellite payloads, and realize integrated simulation of multi-satellite multi-task comprehensive reinforcement learning and three-dimensional visual simulation. It can complete the entire process of perception, action sequence planning, and maneuvering; it can realize multi-satellite training, establish corresponding countermeasure models according to specific scenarios and requirements, and can cope with the impact of target uncertain behaviors in space game tasks on spacecraft game decision-making and control; the multi-agent system can make decisions and responses using the model trained by reinforcement learning within 10 s, avoiding the time cost brought by the space-ground large loop; the reinforcement learning model is transferable and is ultimately applicable to avoidance and approaching scenarios.

[0188] Embodiment 3

[0189] The embodiments of the present application provide an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. The memory is used to store software programs and modules. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory. The memory and the processor are connected by a bus. Specifically, when the processor runs the above computer program stored in the memory, any step in Embodiment 1 above is implemented.

[0190] It should be understood that in the embodiments of the present application, the so-called processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0191] The memory may include a read-only memory, a flash memory, and a random access memory, and provide instructions and data to the processor. A part or all of the memory may also include a non-volatile random access memory.

[0192] It should be understood that if the above integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of this application, it can also be completed by a computer program instructing relevant hardware. The above computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above various method embodiments can be implemented. Among them, the above computer program includes computer program code, and the above computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The above computer-readable medium can include: any entity or device capable of carrying the above computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the above computer-readable storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction.

[0193] The above embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A high-value spacecraft cluster intelligent guarding method, characterized in that: include: Acquiring initial data of a spacecraft cluster, wherein the spacecraft cluster includes a plurality of autonomous decision-making spacecraft; Based on the initial data, combining the spacecraft orbit model, the satellite onboard sensor model, and the illumination model, the environment of the spacecraft cluster is simulated to obtain the current simulated environment; Build a behavioral rule base and capability rule base for spacecraft clusters; Setting a reward function, and establishing a reinforcement learning model according to a multi-agent reinforcement learning algorithm and the reward function, so as to obtain a multi-agent decision result for the spacecraft cluster approach problem; Iteratively training the reinforcement learning model according to the current simulation environment, the behavior rule library, and the capability rule library until the reinforcement learning model converges; Simulating tasks in real scenarios according to the converged reinforcement learning model to obtain final multi-agent decision results, and visualizing the multi-agent decision results; The behavior rule base and capability rule base for constructing the spacecraft cluster include: A behavior rule library and a capability rule library of a spacecraft cluster are constructed, and each spacecraft in the spacecraft cluster is associated and information is exchanged according to the behavior rule library and the capability rule library to obtain an optimal strategy for spacecraft rendezvous. The behavior rule library includes: Movement rules are used to describe the movement of a spacecraft toward a target according to a predetermined recursive direction in the absence of enemy threats, wherein the enemy threats include radiation satellites and interference satellites; Threat avoidance rules are used to describe how a spacecraft changes its direction of movement to avoid the threat when searching for a target in the presence of an enemy threat. The information transmission rule is used to describe that when the number of spacecraft that successfully approach the target is greater than or equal to 1, the number of spacecraft that subsequently approach the target at the same time increases; Added path selection rules for perception information, which are used to describe the information transmission between spacecraft through electromagnetic signals, and select the moving point of the spacecraft according to the safety value and offensive value of other spacecraft within the perception range; The reward and punishment rules for pheromone updates are used to describe the process of a spacecraft approaching a target. When the distance between the spacecraft and the target is less than a preset value, the spacecraft is rewarded with a linearly increasing reward. When the spacecraft successfully approaches the target, the spacecraft is rewarded for approaching the target.

2. The high-value spacecraft cluster intelligent guarding method as claimed in claim 1, characterized in that: The process of performing environmental simulation on the spacecraft cluster based on the initial data and combining the spacecraft orbit model, the satellite onboard sensor model, and the illumination model to obtain the current simulated environment includes: predicting the flight orbit of the spacecraft based on the initial data and the spacecraft orbit model; Determining the payload beam pointing direction of the spacecraft according to the initial data and the satellite onboard sensor model; determining the observed effects of lighting conditions on the spacecraft based on the initial data and the lighting model; Environmental deduction is performed according to the flight orbit of the spacecraft, the payload beam pointing, and the observation influence to obtain the current simulated environment.

3. The high-value spacecraft cluster intelligent guarding method as claimed in claim 2, characterized in that: The spacecraft orbit model includes an ideal orbit model and a spacecraft orbit maneuver guidance model; Among them, the ideal orbit model is used to obtain the position vector and velocity vector of the spacecraft at any time in the absence of other external interference forces, and the spacecraft orbit maneuvering guidance model is used to obtain the spacecraft's orbit change strategy while considering the influence of external interference forces.

4. The high-value spacecraft cluster intelligent guarding method as claimed in claim 2, characterized in that: The satellite-borne sensor model is used to determine the load beam direction according to the size of the rotation angle of the dual-axis positioning mechanism in the load system of the spacecraft.

5. The high-value spacecraft cluster intelligent guarding method as claimed in claim 2, characterized in that: The illumination model is used to calculate the corresponding critical angle of illumination conditions according to the illumination conditions at the current moment, so as to determine the observation impact of the illumination conditions at the current moment on the spacecraft, wherein the illumination conditions include: earth occlusion conditions and earth light conditions, earth shadow conditions, sunlight and moonlight conditions, and earth air light conditions.

6. The high-value spacecraft cluster intelligent guarding method as claimed in claim 1, characterized in that: The capability rule base includes: The orbit change capability rule is used to describe the thrust and torque required for the spacecraft to obtain attitude and orbit control through thrusters; The radiation capability rule is used to describe the initial range of the radiation satellite with radiation effect. When the spacecraft is radiated by the radiation satellite, the spacecraft is tracked by the radiation satellite. The interference capability rule is used to describe the increased probability of a spacecraft being subjected to electromagnetic wave interference by an interfering satellite after the spacecraft is irradiated by the radiating satellite. When the spacecraft is interfered by the interfering satellite, the spacecraft is tracked by the interfering satellite.

7. The high-value spacecraft cluster intelligent guarding method as claimed in claim 1, characterized in that: The establishing of a reinforcement learning model according to the multi-agent reinforcement learning algorithm and the reward function comprises: Obtaining a reward function according to the reward and punishment rules; According to the reward function, a single spacecraft control model is established based on the proximal policy optimization algorithm; According to the single spacecraft control model, multiple spacecraft control models are established based on a multi-agent reinforcement learning algorithm, and the reinforcement learning model is obtained according to the single spacecraft control model and the multiple spacecraft control models.

8. An electronic device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method according to any one of claims 1 to 7 when executing the computer program.

9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Autonomous orbital transfer decision-making method and device for energy limited satellite based on deep reinforcement learning

    CN115828741A

  • Method and device for predicting satellite attitude in orbit, medium and electronic equipment

    CN116767516A

  • Spacecraft cluster game hunting motion planning method based on multi-agent reinforcement learning

    CN118056756A