Space target approaching intelligent trajectory optimization method and equipment
By introducing reference orbits and proportional differential control coefficients into the optimization of space target approach trajectories and combining them with deep reinforcement learning, the trajectory is optimized to achieve high-precision and low-energy approach to space targets, which solves the accuracy and usability problems of existing methods and improves the efficiency of trajectory optimization.
Patent Information
- Application Number
- CN202410431357.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-11
- Publication Date
- 2025-10-21
AI Technical Summary
Existing space target approach trajectory optimization methods have problems with accuracy, usability and efficiency. Traditional methods are inefficient and difficult to apply online. Methods based on deep reinforcement learning are difficult to converge in training, have a large action space and sparse rewards.
A reference orbit is introduced, and the target approach reference orbit is constructed through Lambert orbit change. The orbit following framework is built by combining the proportional differential control coefficient and the satellite dynamics model. The control coefficient is adjusted using deep reinforcement learning, and the trajectory is optimized to achieve high-precision and low-energy approach.
It achieves high-precision and low-energy approach to spatial targets, reduces the training difficulty of deep reinforcement learning, solves the problems of large action space and difficult training convergence, and improves the accuracy and usability of trajectory optimization.
Smart Images

Figure CN120822280A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of satellite trajectory optimization, and in particular to a method and device for optimizing the approaching intelligent trajectory of a space target. Background Art
[0002] In recent years, space activities have become increasingly frequent, and the number of space target approach missions such as space debris removal, rendezvous and docking, and in-orbit refueling has gradually increased. These missions require precise approach and trajectory optimization goals such as saving fuel. Therefore, it is necessary to develop a space target approach trajectory optimization method with high accuracy, strong availability, and high efficiency.
[0003] Among the existing methods for optimizing the approach trajectory of space targets, traditional methods represented by the pseudo-spectral method face problems such as general efficiency, difficulty in online application, and the possibility of falling into local optimality. In addition, they lack generalization capabilities and need to be re-solved for each task. Although trajectory optimization methods based on deep reinforcement learning have generalization capabilities, problems such as the large action space and sparse rewards make training difficult to converge, and further in-depth research is still needed.
[0004] Therefore, in view of the accuracy, usability and efficiency problems of existing space target approach trajectory optimization methods, it is necessary to develop an accurate, efficient and highly usable space target approach trajectory optimization method. Summary of the Invention
[0005] The present invention provides a method and device for optimizing the intelligent trajectory of approaching space targets. By introducing a reference trajectory, the training difficulty of deep reinforcement learning is reduced. This can solve the technical problems of existing trajectory optimization technologies based on deep reinforcement learning, such as large action space, sparse rewards, and difficult training convergence.
[0006] According to one aspect of the present invention, a method for optimizing an approaching intelligent trajectory of a space target is provided, the method comprising:
[0007] Construct the target approach reference orbit based on Lambert orbit change;
[0008] The orbit following framework of the reference orbit is constructed based on the proportional-differential control coefficients and the satellite dynamics model;
[0009] The proportional-differential control coefficients are adjusted using deep reinforcement learning methods, and the adjusted proportional-differential control coefficients are substituted into the orbit following frame of the reference orbit to obtain the control that needs to be applied. The corresponding control is applied to the satellite and the orbit is extrapolated until it approaches the space target, thereby obtaining the optimal trajectory for the space target to approach.
[0010] Preferably, using a deep reinforcement learning method to adjust the proportional differential control coefficient, substituting the adjusted proportional differential control coefficient into the orbit following frame of the reference orbit to obtain the control to be applied, applying the corresponding control to the satellite and performing orbit extrapolation until it approaches the space target, thereby obtaining the optimal trajectory for the space target to approach, including:
[0011] S31. Setting the intelligent agent, environmental information, action space, state space, and reward function in the deep reinforcement learning framework, wherein the intelligent agent is set to be a satellite, the environmental information is set to include the target to be approached and the reference orbit, the action space is set to include the proportional control coefficient and the differential control coefficient, the state space is set to include the position information and velocity information of the satellite, the target to be approached, and the reference orbit, and the reward function is set to include the target approach accuracy and the fuel consumption throughout the process;
[0012] S32, initialize the policy network, value network and experience pool;
[0013] S33. In the current training round, setting the current moment as the initial moment of the target approach process;
[0014] S34. The agent takes action according to the current policy network, substitutes the current action into the track following framework of the reference track, obtains the control that needs to be applied, and applies the corresponding control to the agent. The agent state is updated according to the dynamic equation. After the state is updated, the agent interacts with the environment to obtain the current reward and stores the "state-action-reward" experience in the experience pool.
[0015] S35. Determine whether the current capacity of the experience pool is greater than the preset capacity. If so, randomly select a number of experience pairs from the experience pool to update the value network and go to S36. Otherwise, return to S34.
[0016] S36, updating the policy network based on the updated value network;
[0017] S37. Determine whether the next moment reaches the target moment. If so, obtain the cumulative reward of the current training round and go to S38. Otherwise, return to S34.
[0018] S38. If the current training round satisfies the following conditions simultaneously: the cumulative reward continues to increase, the convergence value is greater than the preset convergence value, the target approach accuracy is greater than the preset accuracy value, and the fuel consumption during the entire process is less than the preset consumption value, then the training ends and the trained optimal policy network is obtained; otherwise, the training parameters are adjusted and the process returns to S32.
[0019] S39. Based on the trained optimal strategy network, the optimal action sequence of the intelligent agent is obtained. The optimal action sequence of the intelligent agent is substituted into the orbit following framework of the reference orbit to obtain the control sequence that needs to be applied. The corresponding control sequence is applied to the satellite and the orbit is extrapolated until it approaches the space target, thereby obtaining the optimal trajectory for the space target to approach.
[0020] Preferably, the training parameters include the number of policy network layers, the number of value network layers, the number of neurons, the type of activation function, the learning rate, the discount factor and the batch size.
[0021] Preferably, the kinetic equation is as follows:
[0022]
[0023]
[0024] a=a control +a0
[0025] Where r, v, and a are the actual position, velocity, and acceleration information of the satellite, respectively. a0 is the acceleration information generated by the resultant external force without control force on the satellite. control Acceleration information generated by the applied control force.
[0026] Preferably, constructing a target approaching reference orbit based on Lambert orbit change includes:
[0027] Determine the time and location of approaching the target based on mission requirements;
[0028] The current time is taken as the initial time of Lambert orbit maneuver, the current position of the satellite is taken as the initial position of Lambert orbit maneuver, the time of approaching the target is taken as the end time of Lambert orbit maneuver, and the position of approaching the target is taken as the end position of Lambert orbit maneuver;
[0029] The initial velocity of the Lambert maneuver is obtained based on the initial time, initial position, end time and end position of the Lambert maneuver;
[0030] Through orbit extrapolation, the reference orbit of the Lambert maneuver is obtained according to the initial time, initial position and initial velocity of the Lambert maneuver, thereby obtaining the position information and velocity information of the reference orbit at each moment.
[0031] Preferably, the orbit following framework for building a reference orbit based on the proportional differential control coefficient and the satellite dynamics model includes:
[0032] Based on the satellite dynamics model, the position and velocity information of the preset orbit at each moment when no control is applied are predicted;
[0033] Obtaining a position error based on the position information of the reference track at each moment and the position information of the preset track at each moment when no control is applied, and obtaining a speed error based on the speed information of the reference track at each moment and the speed information of the preset track at each moment when no control is applied;
[0034] The track following framework of the reference track is constructed based on the position error, velocity error and proportional differential control coefficient.
[0035] Preferably, the position information and speed information of the preset track at each moment when no control is applied are predicted by the following formula:
[0036]
[0037]
[0038]
[0039] The position error and velocity error are obtained by the following formula:
[0040]
[0041]
[0042] The track following framework of the reference track is constructed using the following formula:
[0043]
[0044] Where, e r is the position error, e v is the speed error, The reference orbit is at t i+1 Location information at all times, The preset orbits at t are respectively i , t i+1 Location information at all times, The reference orbit is at t i+1 Speed information at the moment, The preset orbits at t are respectively i , t i+1 Speed information at the moment, a control is the acceleration information generated by the applied control force, K P is the proportional control coefficient, K D is the differential control coefficient, For satellites at t i The acceleration information at the moment, μ is the gravitational constant.
[0045] According to another aspect of the present invention, a computer device is provided, comprising a memory, a processor, and a space target approach intelligent trajectory optimization program stored in the memory and executable on the processor, wherein the processor implements any of the above-mentioned methods when executing the space target approach intelligent trajectory optimization program.
[0046] Applying the technical solution of this invention, deep reinforcement learning is used to train control coefficients, enabling trajectory following a reference track, thereby achieving high-precision, low-energy approach to space targets. By introducing a reference track, this method reduces the difficulty of deep reinforcement learning training, addressing technical issues such as large action space, sparse rewards, and difficulty in training convergence in existing trajectory optimization techniques based on deep reinforcement learning. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] The accompanying drawings are included to provide a further understanding of the embodiments of the present invention, constitute a part of the specification, illustrate the embodiments of the present invention, and together with the description, explain the principles of the present invention. Obviously, the drawings described below are only some embodiments of the present invention, and those skilled in the art can derive other drawings based on these drawings without inventive effort.
[0048] Figure 1 A flowchart of a method for optimizing an intelligent trajectory for approaching a space target provided according to an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0049] It should be noted that, in the absence of conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. The following description of at least one exemplary embodiment is actually only illustrative and is in no way intended to limit the present invention and its application or use. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0050] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.
[0051] Unless otherwise specifically stated, the relative arrangement of the parts and steps, the numerical expressions and the numerical values set forth in these embodiments do not limit the scope of the present invention. At the same time, it should be understood that, for ease of description, the sizes of the various parts shown in the drawings are not drawn according to the actual proportional relationship. The techniques, methods and equipment known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the techniques, methods and equipment should be considered as part of the authorization specification. In all examples shown and discussed here, any specific values should be interpreted as being merely exemplary and not as limiting. Therefore, other examples of the exemplary embodiments may have different values. It should be noted that similar numbers and letters represent similar items in the following figures, and therefore, once an item is defined in one figure, it does not need to be further discussed in subsequent figures.
[0052] like Figure 1 As shown, the present invention provides a method for optimizing an approaching intelligent trajectory of a space target, the method comprising:
[0053] S10, constructing a target approach reference orbit based on Lambert orbit change;
[0054] S20, constructing an orbit following framework of the reference orbit based on the proportional differential control coefficient and the satellite dynamics model;
[0055] S30. Use deep reinforcement learning methods to adjust the proportional differential control coefficients, substitute the adjusted proportional differential control coefficients into the orbit following frame of the reference orbit, obtain the control that needs to be applied, apply corresponding control to the satellite and perform orbit extrapolation until it approaches the space target, thereby obtaining the optimal trajectory for approaching the space target.
[0056] This method uses deep reinforcement learning to train control coefficients, achieving trajectory following a reference trajectory, thereby enabling high-precision, low-energy approach to space targets. By introducing a reference trajectory, this method reduces the difficulty of deep reinforcement learning training and addresses technical issues such as large action space, sparse rewards, and difficulty in training convergence in existing trajectory optimization techniques based on deep reinforcement learning.
[0057] According to an embodiment of the present invention, in S10 of the present invention, constructing a target approaching reference orbit based on Lambert orbit change includes:
[0058] S11. Determine the time and location of approaching the target according to mission requirements;
[0059] That is, according to the mission requirements, select the specific position r_Lambert close to the target in the target's future orbit end and the approach time t_Lambert end ;
[0060] S12, take the current time as the initial time of Lambert maneuver (denoted as t_Lambert start ), the current position of the satellite is used as the initial position of the Lambert orbit change (denoted as r_Lambert start ), the time of approaching the target is taken as the end time of Lambert maneuver (denoted as t_Lambert end ), the position close to the target is taken as the end position of Lambert maneuver (denoted as r_Lambert end );
[0061] S13, based on the initial time, initial position, end time and end position of the Lambert track change, obtain the initial speed of the Lambert track change (denoted as v_Lambert start );
[0062] Among them, the Lambert orbit change problem can be solved by using Gauss method, Battin method and other methods;
[0063] S14. Obtaining a reference orbit of the Lambert orbit change according to the initial time, initial position, and initial velocity of the Lambert orbit change by orbit extrapolation, thereby obtaining position information and velocity information of the reference orbit at each moment;
[0064] Among them, the orbital extrapolation adopts the two-body model, and the initial position is r_Lambert start , the initial velocity is v_Lambert start The initial time of orbital extrapolation is t_Lambert start The orbit extrapolation ends at t_Lambert end The formula for the two-body model is as follows:
[0065]
[0066]
[0067]
[0068]
[0069]
[0070]
[0071] Where x, y, and z are the three-dimensional positions of the satellite in the Earth-inertial system under the two-body model assumption, and v is x 、v y 、v z is the three-dimensional velocity in the Earth's inertial system under the two-body model assumption, and μ is the Earth's gravitational constant;
[0072] The reference orbit of Lambert orbit change is shown as follows:
[0073]
[0074] Where, is the reference orbit of Lambert maneuver, t i is the current moment, N is the total number of moments, The reference orbit is at t i The position information at the moment is the three-dimensional position vector in the Earth's inertial system. The reference orbit is at t i The velocity information at the moment is the three-dimensional velocity vector in the Earth's inertial system.
[0075] According to one embodiment of the present invention, in S20 of the present invention, building an orbit following framework of a reference orbit based on a proportional differential control coefficient and a satellite dynamics model includes:
[0076] S21, based on the satellite dynamics model, predict the position information and velocity information of the preset orbit at each moment when no control is applied; for example, at t i At time t, the satellite dynamics model predicts that the preset orbit is at t when no control is applied. i+1 Location and speed
[0077] S22. Obtain a position error based on the position information of the reference track at each moment and the position information of the preset track at each moment when no control is applied, and obtain a speed error based on the speed information of the reference track at each moment and the speed information of the preset track at each moment when no control is applied;
[0078] S23. Complete the construction of the track following framework of the reference track based on the position error, velocity error and proportional differential control coefficient.
[0079] Specifically, in S21 of the present invention, the position information and speed information of the preset track at each moment when no control is applied are predicted by the following formula:
[0080]
[0081]
[0082]
[0083] In S22 of the present invention, the position error and the speed error are obtained by the following formula:
[0084]
[0085]
[0086] In S23 of the present invention, the track following framework of the reference track is constructed by the following formula:
[0087]
[0088] Where, e r is the position error, e v is the speed error, The reference orbit is at t i+1 The position information at the moment is the three-dimensional position vector in the Earth's inertial system. The preset orbits at t are respectively i , t i+1 The position information at the moment is the three-dimensional position vector in the Earth's inertial system. The reference orbit is at t i+1 The velocity information at the moment is the three-dimensional velocity vector in the Earth's inertial system. The preset orbits at t are respectively i , t i+1 The velocity information at the moment is the three-dimensional velocity vector in the earth's inertial system, a control is the acceleration information generated by the applied control force, K P is the proportional control coefficient, K D is the differential control coefficient, For satellites at t i The acceleration information at the moment is the three-dimensional acceleration vector in the Earth's inertial system, and μ is the Earth's gravitational constant.
[0089] In this embodiment, if a control If the actual performance index of the satellite is exceeded, it can be reduced proportionally to within the performance index range.
[0090] According to one embodiment of the present invention, in S30 of the present invention, a proportional differential control coefficient is adjusted using a deep reinforcement learning method, the adjusted proportional differential control coefficient is substituted into the orbit following frame of the reference orbit to obtain the control to be applied, the corresponding control is applied to the satellite and the orbit is extrapolated until it approaches the space target, thereby obtaining the optimal trajectory for the space target approaching, including:
[0091] S31. Setting the intelligent agent, environmental information, action space, state space, and reward function in the deep reinforcement learning framework, wherein the intelligent agent is set to be a satellite, the environmental information is set to include the target to be approached and the reference orbit, the action space is set to include the proportional control coefficient and the differential control coefficient, the state space is set to include the position information and velocity information of the satellite, the target to be approached, and the reference orbit, and the reward function is set to include the target approach accuracy and the fuel consumption throughout the process;
[0092] S32, initialize the policy network, value network and experience pool;
[0093] S33. In the current training round, setting the current moment as the initial moment of the target approach process;
[0094] S34. The agent (i.e., the satellite) takes an action (i.e., the proportional control coefficient and the differential control coefficient) according to the current policy network, substitutes the current action into the orbit following framework of the reference orbit, obtains the control that needs to be applied, and applies the corresponding control to the agent. The agent state is updated according to the dynamic equation. After the state is updated, the agent interacts with the environment (i.e., the target to be approached, the reference orbit), obtains the current reward (i.e., the target approach accuracy and fuel consumption), and stores the "state-action-reward" experience in the experience pool;
[0095] S35. Determine whether the current capacity of the experience pool is greater than the preset capacity. If so, randomly select a number of experience pairs from the experience pool to update the value network and go to S36. Otherwise, return to S34.
[0096] S36, updating the policy network based on the updated value network;
[0097] S37. Determine whether the next moment reaches the target moment. If so, obtain the cumulative reward of the current training round and go to S38. Otherwise, return to S34.
[0098] S38. If the current training round satisfies the following conditions simultaneously: the cumulative reward continues to increase, the convergence value is greater than the preset convergence value, the target approach accuracy is greater than the preset accuracy value, and the fuel consumption during the entire process is less than the preset consumption value, then the training ends and the trained optimal policy network is obtained; otherwise, the training parameters are adjusted and the process returns to S32.
[0099] S39, based on the trained optimal strategy network, obtain the optimal action sequence of the agent (i.e., K P and K D The optimal action sequence of the intelligent agent is substituted into the orbit following frame of the reference orbit to obtain the control sequence that needs to be applied. The corresponding control sequence is applied to the satellite and the orbit is extrapolated until it approaches the space target, thereby obtaining the optimal trajectory of the space target.
[0100] Specifically, in S34 of the present invention, the kinetic equation is shown below:
[0101]
[0102]
[0103] a=a control+a0
[0104] Where r, v, and a are the actual position, velocity, and acceleration information of the satellite, respectively. The position, velocity, and acceleration information are the three-dimensional position, velocity, and acceleration vectors of the Earth-inertial system. a0 is the acceleration information generated by the resultant external force without control force on the satellite. control Acceleration information generated by the applied control force.
[0105] Specifically, in S38 of the present invention, the training parameters include the number of policy network layers, the number of value network layers, the number of neurons, the type of activation function, the learning rate, the discount factor and the batch size.
[0106] The present invention also provides a computer device comprising a memory, a processor, and a space target approach intelligent trajectory optimization program stored in the memory and runnable on the processor, wherein the processor implements any of the above-mentioned methods when executing the space target approach intelligent trajectory optimization program.
[0107] In summary, the present invention provides a method and device for intelligent trajectory optimization for approaching space targets. By training control coefficients through deep reinforcement learning, the method achieves trajectory tracking relative to a reference trajectory, thereby enabling high-precision, low-energy approach to space targets. By introducing a reference trajectory, the present method reduces the difficulty of deep reinforcement learning training, addressing technical issues inherent in existing deep reinforcement learning-based trajectory optimization techniques, such as large action space, sparse rewards, and difficulty in training convergence.
[0108] Parts of the present invention that are not described in detail are well known to those skilled in the art.
[0109] In the description of the present invention, it should be understood that the directions or positional relationships indicated by directional words such as "front, back, up, down, left, right", "horizontal, vertical, perpendicular, horizontal" and "top, bottom" are usually based on the directions or positional relationships shown in the accompanying drawings. They are only for the convenience of describing the present invention and simplifying the description. Unless otherwise specified, these directional words do not indicate or imply that the device or element referred to must have a specific direction or be constructed and operated in a specific direction. Therefore, they cannot be understood as limiting the scope of protection of the present invention; the directional words "inside and outside" refer to the inside and outside relative to the outline of each component itself.
[0110] For ease of description, spatially relative terms such as "above", "above", "on the upper surface of", "above", etc. may be used herein to describe the spatial positional relationship of a device or feature to other devices or features as shown in the figures. It should be understood that spatially relative terms are intended to include different orientations of the device in use or operation in addition to the orientation described in the figures. For example, if the device in the drawings is inverted, the device described as "above other devices or structures" or "above other devices or structures" will be positioned as "below other devices or structures" or "below other devices or structures". Thus, the exemplary term "above" can include both "above" and "below". The device can also be positioned in other different ways (rotated 90 degrees or in other orientations), and the spatially relative descriptions used here are interpreted accordingly.
[0111] In addition, it should be noted that the use of terms such as "first" and "second" to limit components is only for the convenience of distinguishing the corresponding components. Unless otherwise stated, the above terms have no special meaning and therefore cannot be understood as limiting the scope of protection of the present invention.
[0112] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A space target approach intelligent trajectory optimization method, characterized in that: The method comprises: Construct the target approach reference orbit based on Lambert orbit change; The orbit following framework of the reference orbit is constructed based on the proportional-differential control coefficients and the satellite dynamics model; The proportional-differential control coefficients are adjusted using deep reinforcement learning methods, and the adjusted proportional-differential control coefficients are substituted into the orbit following frame of the reference orbit to obtain the control that needs to be applied. The corresponding control is applied to the satellite and the orbit is extrapolated until it approaches the space target, thereby obtaining the optimal trajectory for the space target to approach.
2. The method according to claim 1, characterized in that The proportional-differential control coefficients are adjusted using deep reinforcement learning methods. The adjusted proportional-differential control coefficients are substituted into the orbit-following frame of the reference orbit to obtain the required control. The corresponding control is applied to the satellite and the orbit is extrapolated until it approaches the space target. The optimal trajectory for approaching the space target is obtained, including: S31. Setting the intelligent agent, environmental information, action space, state space, and reward function in the deep reinforcement learning framework, wherein the intelligent agent is set to be a satellite, the environmental information is set to include the target to be approached and the reference orbit, the action space is set to include the proportional control coefficient and the differential control coefficient, the state space is set to include the position information and velocity information of the satellite, the target to be approached, and the reference orbit, and the reward function is set to include the target approach accuracy and the fuel consumption throughout the process; S32, initialize the policy network, value network and experience pool; S33. In the current training round, setting the current moment as the initial moment of the target approach process; S34: The agent takes action based on the current policy network and substitutes the current action into the track-following framework of the reference track to obtain the control required for the current action. This control is then applied to the agent, and the agent's state is updated according to the dynamic equation. After the state is updated, the agent interacts with the environment to obtain the current reward and stores the "state-action-reward" experience in the experience pool. S35. Determine whether the current capacity of the experience pool is greater than the preset capacity. If so, randomly select a number of experience pairs from the experience pool to update the value network and go to S36. Otherwise, return to S34. S36, updating the policy network based on the updated value network; S37. Determine whether the next moment reaches the target moment. If so, obtain the cumulative reward of the current training round and go to S38. Otherwise, return to S34. S38. If the current training round satisfies the following conditions simultaneously: the cumulative reward continues to increase, the convergence value is greater than the preset convergence value, the target approach accuracy is greater than the preset accuracy value, and the fuel consumption during the entire process is less than the preset consumption value, then the training ends and the trained optimal policy network is obtained; otherwise, the training parameters are adjusted and the process returns to S32. S39. Based on the trained optimal strategy network, the optimal action sequence of the intelligent agent is obtained. The optimal action sequence of the intelligent agent is substituted into the orbit following framework of the reference orbit to obtain the control sequence that needs to be applied. The corresponding control sequence is applied to the satellite and the orbit is extrapolated until it approaches the space target, thereby obtaining the optimal trajectory for the space target to approach.
3. The method according to claim 2, characterized in that The training parameters include the number of policy network layers, the number of value network layers, the number of neurons, the type of activation function, the learning rate, the discount factor and the batch size.
4. The method according to any one of claims 1 to 3, characterized in that The kinetic equation is shown below: a=a control +a0 Where r, v, and a are the actual position, velocity, and acceleration information of the satellite, respectively. a0 is the acceleration information generated by the resultant external force without control force on the satellite. control Acceleration information generated by the applied control force.
5. The method according to claim 1, wherein The target approach reference orbit based on Lambert orbit change includes: Determine the time and location of approaching the target based on mission requirements; The current time is taken as the initial time of Lambert orbit maneuver, the current position of the satellite is taken as the initial position of Lambert orbit maneuver, the time of approaching the target is taken as the end time of Lambert orbit maneuver, and the position of approaching the target is taken as the end position of Lambert orbit maneuver; The initial velocity of the Lambert maneuver is obtained based on the initial time, initial position, end time and end position of the Lambert maneuver; Through orbit extrapolation, the reference orbit of the Lambert maneuver is obtained according to the initial time, initial position and initial velocity of the Lambert maneuver, thereby obtaining the position information and velocity information of the reference orbit at each moment.
6. The method according to claim 1, characterized in that The orbit following framework for building a reference orbit based on the proportional-differential control coefficients and the satellite dynamics model includes: Based on the satellite dynamics model, the position and velocity information of the preset orbit at each moment when no control is applied are predicted; Obtaining a position error based on the position information of the reference track at each moment and the position information of the preset track at each moment when no control is applied, and obtaining a speed error based on the speed information of the reference track at each moment and the speed information of the preset track at each moment when no control is applied; The track following framework of the reference track is constructed based on the position error, velocity error and proportional differential control coefficient.
7. The method according to claim 6, characterized in that The position and speed information of the preset track at each moment when no control is applied are predicted by the following formula: The position error and velocity error are obtained by the following formula: The track following framework of the reference track is constructed using the following formula: Where, e r is the position error, e v is the speed error, The reference orbit is at t i+1 Location information at all times, The preset orbits at t are respectively i , t i+1 Location information at all times, The reference orbit is at t i+1 Speed information at the moment, The preset orbits at t are respectively i , t i+1 Speed information at the moment, a control is the acceleration information generated by the applied control force, K P is the proportional control coefficient, K D is the differential control coefficient, For satellites at t i The acceleration information at the moment, μ is the gravitational constant.
8. A computer device, characterized in that: The method comprises a memory, a processor, and a space target approach intelligent trajectory optimization program stored in the memory and executable on the processor, wherein the processor implements any one of the methods of claims 1 to 7 when executing the space target approach intelligent trajectory optimization program.