Satellite orbit forecasting method based on deep learning physical constraint loss
By combining CNN and BiLSTM for satellite orbit prediction, the problems of insufficient accuracy and efficiency of traditional methods are solved, and high-precision and fast orbit prediction is achieved, improving the physical consistency and computational efficiency of the model.
Patent Information
- Application Number
- CN202511094643.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-11-21
AI Technical Summary
Traditional satellite orbit prediction methods suffer from insufficient accuracy, computational complexity, long processing time, and high dependence on ground-based systems. Furthermore, existing deep learning methods have limitations in terms of model structure specificity and feature extraction effectiveness.
We employ a lightweight spatiotemporal feature extraction method based on convolutional neural networks (CNN), combine it with a bidirectional long short-term memory network (BiLSTM) for orbital temporal feature learning, and design a loss function that incorporates physical constraints to improve the physical consistency and interpretability of the model.
It achieves high-precision and rapid satellite orbit prediction, reduces dependence on ground-based systems, and improves the physical consistency and computational efficiency of the model.
Smart Images

Figure CN120995107A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the field of deep learning and orbit dynamics, and particularly relates to a satellite orbit prediction method based on deep learning physical constraint loss. BACKGROUND
[0002] For global navigation satellite system (GNSS), fast and high-precision orbit prediction is a core technology for ensuring service quality and reliability. On the one hand, satellites need to perform short-term orbit prediction for the generation of navigation ephemeris, which is then broadcast to users for time service and positioning. On the other hand, for the autonomous operation system of satellites, long-term orbit prediction is needed to provide autonomous navigation services in special situations without ground station support. At the same time, with the rapid increase in the number of space satellites, in the field of space situation awareness, fast and high-precision orbit prediction of satellites is also needed to cope with collision warning, space debris monitoring and avoidance, etc. These all pose higher demands and challenges to the precision, timeliness and robustness of orbit prediction technology.
[0003] Traditional prediction methods mainly rely on complex orbit dynamics mathematical models and force models, which may face difficulties in accurate modeling, complex calculation, time-consuming process, etc. Moreover, existing implementation schemes have a high dependence on the ground, which may occupy the satellite-ground communication bandwidth and greatly restrict the real-time performance, anti-interference performance and invulnerability of satellite autonomous operation. In recent years, artificial intelligence technology represented by deep learning has become an effective way to solve these problems. Deep learning methods have strong non-linear fitting capability and unique advantages in automatically extracting complex features from data, which can achieve higher precision and real-time orbit prediction and greatly reduce the dependence on the ground. The present application aims to address the challenges in the current satellite orbit prediction field, especially the deficiencies of traditional physical models in precision and efficiency, and the limitations of existing deep learning methods in model structure pertinence, feature extraction effectiveness, computational overhead and physical consistency, etc. Based on the "convolution + time series" hybrid neural network structure, a high-precision satellite orbit prediction method is proposed, which fuses physical constraint loss. SUMMARY
[0004] The present application aims to provide a satellite orbit prediction method based on deep learning physical constraint loss. First, a lightweight CNN module is constructed to extract spatio-temporal features of satellite orbit data, then a BiLSTM bidirectional time series neural network is used to learn the time series features of the orbit, and finally a loss function that fuses physical constraints is designed to incorporate the physical characteristics of error propagation and the basic constraints of orbit dynamics into the model design, which can effectively improve the physical consistency and interpretability of the model.
[0005] The satellite orbit prediction method based on deep learning physical constraint loss provided by the application has the implementation steps as follows:
[0006] Step one: normalizing preprocessing of input data, forming a training data set and a test data set, and constructing batch training data;
[0007] Step two: dimension expansion of sample data points in each window of the batch data formed in step one, construction of a multi-dimensional feature space of sample points, and formation of a batch input data format that can be passed into the model;
[0008] Step three: forward inference of the batch data formed in step two using the model, respectively through a CNN lightweight space-time feature extraction module and a BiLSTM bidirectional time series neural network module, to obtain the batch orbit prediction value of the next time point output by the model;
[0009] Step four: taking the orbit prediction value obtained in step three and the true value label in the training set obtained in step one as input, calculating the loss value of the current training iteration batch through a multi-task learning loss module that fuses physical constraints, and performing backward update of the model parameters to complete the model training;
[0010] Step five: through steps two to four, the test set formed in step one is used to infer and verify the model, and compared with the true value in the test set, the model test result can be obtained;
[0011] In step one, the "normalizing preprocessing of input data, forming a training data set and a test data set, and constructing batch training data" is as follows:
[0012] S11, prepare a satellite orbit sequence data with a time length of not less than 12 days, set the format to give a satellite three-axis position and velocity vector in the geocentric inertial coordinate system every 300s, and divide the total time length by 7:3 to form a training data set and a test data set respectively;
[0013] S12, for the training data set, form batch input data in the form of sliding window, set the window length to 144 sample points, and the sliding step to 1, the sample points in each sliding window as training input, predict the satellite position vector of the next time through the model;
[0014] S13, for the test data set, take the first 144 points of the sequence data as the initial test input, recursively predict the satellite position vector at the 145th time point and all subsequent time points in the test set in the form of a sliding window, the sliding window step is 1, the window size remains 144 in length, the predicted value at the previous time is taken as the last value of the next time sliding window input sequence each time the window slides, while the first value in the sliding window is removed, keeping the total length of 144 unchanged;
[0015] S14, set the number of data batch processing during training to 64 window sequence data per batch, and set the test to not use batch processing, that is, only the initial window of 144 sample data is taken as input to perform recursive prediction.
[0016] In step two, the "dimensional expansion of the sample data points in each window in the batch data formed in step one, the construction of the multi-dimensional feature space of the sample points, and the formation of the batch input data format that can be transmitted into the model" is as follows:
[0017] S21, for each window sequence data, calculate the J2 perturbation force feature at the time for each sample point;
[0018] S22, for each window sequence data, calculate the sin function encoding period feature at the time for each sample point;
[0019] S23, for each window sequence data, calculate the cos function encoding period feature at the time for each sample point;
[0020] S24, for each window sequence data, expand the dimension of each sample point, and add the feature values calculated in steps S21, S22 and S23 to the original satellite position and velocity vector features as new feature dimensions, to form a batch input data format that can be transmitted into the model Each batch training input dimension is 64x144x9.
[0021] In step three, the "forward inference of the batch data formed in step two using the model, respectively through a CNN lightweight space-time feature extraction module and a BiLSTM bidirectional time series neural network module, to obtain the next time batch orbit prediction value output by the model" is as follows:
[0022] S31, input the batch training data obtained in step two into the CNN lightweight space-time feature extraction module to obtain the output feature, the dimension is 64x9x144;
[0023] S32, input the result data obtained in step S31 into a BiLSTM bidirectional time sequence network module to obtain a satellite orbit prediction value at a next moment, and the dimension is 64*6, which respectively includes a position and a velocity vector;
[0024] In step four, the orbit prediction value obtained in step three and the true value label in the training set obtained in step one are input as inputs, a multi-task learning loss module with physical constraints is used to calculate the loss value of the current training iteration batch, and the model parameters are updated in reverse to complete the model training.
[0025] S41, the satellite orbit prediction output value obtained in step three and the true value label at the prediction moment in the training set obtained in step one are input as inputs, and are input into a physical constraint loss module to calculate the loss value of the current training batch.
[0026] S42, the model is updated in reverse, and the next batch of iteration training is performed until the loss converges, and the model training is completed.
[0027] In step five, the model is verified by using the test set formed in step one through steps two to four, and the true value in the test set is compared, and the test result of the model can be obtained.
[0028] S51, using the test data set obtained in step one, the first 144 points of the sequence data are used as initial test inputs, and the model trained in steps two to four is used to predict the satellite orbit position at all moments in the test set in the form of a sliding window as described in S13.
[0029] S52, the satellite orbit prediction value obtained in S51 is compared with the true value in the test set to obtain the test result of the model, and the index form is the three-axis position error of the satellite in the geocentric inertial system.
[0030] Through the above steps, a satellite orbit prediction method based on deep learning physical constraint loss can complete the training of the model and realize high-precision and fast prediction of the satellite orbit position through a CNN lightweight space-time feature extraction module, a BiLSTM time sequence network module and a loss module with physical constraints.
[0031] According to the design of the present application, the present application realizes a satellite orbit prediction method based on deep learning physical constraint loss, and compared with the traditional prediction method based on complex orbit dynamics integral, the present application can realize higher prediction accuracy and faster calculation speed. BRIEF DESCRIPTION OF DRAWINGS
[0032] Figure 1 is a general step flowchart.
[0033] Figure 2 is a network structure diagram of a CNN lightweight spatio-temporal feature extraction module.
[0034] Figure 3 is a BiLSTM time sequence network structure diagram.
[0035] Figure 4 is a simulation test result of the present application. DETAILED DESCRIPTION
[0036] In order to have a further understanding and knowledge of the features, objects and functions of the present application, the present application will be described in more detail with specific implementation examples and accompanying drawings and tables.
[0037] As shown in Figure 1 is a general step flowchart of the present application, first, the satellite orbit time sequence data is batch data processed, and a feature space is constructed as an input tensor of the model; through the CNN lightweight feature extraction module, the input satellite orbit data sequence is preliminarily modeled from the time dimension and the space dimension respectively; based on the output features of the CNN module, the BiLSTM time sequence neural network is used for learning the time sequence characteristics of the orbit data; through the loss module of the physical constraint, the loss of the satellite orbit position and velocity is calculated at the same time, and the physical constraint penalty of the orbit consistency is based on Kepler's law, the end-to-end multi-task training and parameter updating of the network model are realized, and the training is completed until the loss converges. The present application provides a satellite orbit prediction method based on deep learning physical constraint loss, and the specific implementation steps are as follows:
[0038] Step 1: Normalize the input data for pretreatment, form the training data set and the test data set, and construct the batch processing training data.
[0039] Firstly, the satellite orbit sequence data is preprocessed, the total time length is not less than 12 days, and the data can be obtained from STK simulation software or other publicly released satellite orbit determination data. The sequence data format is to give a satellite three-axis position and velocity vector in the geocentric inertial coordinate system every 300s, and the unit is set to meters. The total time length is divided into 7:3 proportions, respectively as the training data set and the test data set.
[0040] Then, the training data set is formed into batch processing input data in the form of sliding window, the window length is set to 144 sample points, and the sliding step is set to 1. Assuming that the total sample points of the training set are k, a total of (k-144+1) window sequences are formed, each sequence contains 144 sample points. The sample points in each sliding window are used as training input, and the satellite position vector at the next time is predicted by the model.
[0041] Then, the first 144 points of the sequence data in the test data set are taken as the initial test input, and the satellite position vector at the 145th time point and all subsequent time points in the test set are recursively predicted in a sliding window manner. The sliding window step size is 1, and the window size remains 144. The predicted value at the previous time is used as the last value of the next time sliding window input sequence, while the first value in the sliding window is removed, keeping the total length of 144 unchanged.
[0042] Finally, the number of data batch processing (Batchsize) during training is set to 64 window sequence data per batch, and the batch processing method is not used during testing. That is, only the initial window of 144 sample data is used as input for recursive prediction.
[0043] Second step: Dimension expansion is performed on the sample data points in each window of the batch processing data formed in step one to construct a multi-dimensional feature space for the sample points and form a batch input data format that can be passed into the model.
[0044] First, for each window sequence data, the J2 perturbation force feature at the current time is calculated for each sample point in the window The calculation formula is as follows:
[0045]
[0046] In the formula, R e is the Earth's radius; J2 is the Earth's flattening coefficient; (x, y, z) is the satellite's current position in the Earth-centered inertial system; is the satellite's radial unit vector pointing to the center of the Earth.
[0047] At the same time, considering the periodicity of satellite two-body orbit motion, periodic feature encoding modeling is used in the orbit time dimension using sin function and cos function as two separate feature dimensions, and the calculation is as follows:
[0048]
[0049] In the formula, Δt k represents the time difference of t k in the satellite orbit period relative to the starting time. T orbit represents the satellite orbit period, with a unit of seconds.
[0050] For each window sequence data, dimension expansion is performed for each sample point, and the feature values calculated above are added to the original satellite position and velocity vector features as new feature dimensions to form a batch input data format that can be passed into the model. For t k , the satellite orbit state feature space is constructed as follows:
[0051]
[0052] wherein, denote the position and velocity of the satellite at time t k denote the J2 perturbation characteristics and the time period characteristics of the satellite at time t k , and the input dimension of each batch is 64x144x9. It can be seen that the last three feature dimensions are related to the position and velocity of the satellite at time t. The feature space construction method adopted by the present application can fully mine the space-time characteristics of the satellite's perturbed motion from the data during model training, thereby improving the model accuracy and stability.
[0053] Step 3: Forward inference is performed on the batch data formed in step 2 using the model to obtain the next time batch orbit prediction value output by the model through a CNN lightweight space-time feature extraction module and a BiLSTM bidirectional time series neural network module.
[0054] Figure 2 The structural diagram of the CNN lightweight space-time feature extraction module is shown. Based on the model input data obtained in step 2, each batch (Batch) is a two-dimensional tensor, which contains 144 sequences, each sequence contains 9 channel feature values, and each channel represents a different feature dimension. The specific structure is introduced as follows:
[0055] First, two layers of one-dimensional convolution are used to model the time dimension, wherein the first layer of convolution (including the BN layer and the ReLU activation layer) learns the features of the time dimension, expands the channel dimension of each sequence from 9 to 64, and the second layer of convolution maps the learned features and generates a time attention weight vector with a dimension of 9x1 through the Softmax normalization function. Multiply the corresponding elements of each row sequence of the input data to obtain the time attention weighted feature data. On this basis, the module further extracts spatial features through three layers of 1D convolution, wherein the first layer of convolution expands the features from 9 dimensions to 64 dimensions, and the second and third layers of convolution perform deep feature mining based on the residual connection structure. Finally, each sequence maintains 64-dimensional features as the input of the time series network.
[0056] In the above residual connection design, the input features are divided into two paths: the main path and the residual path. The main path is nonlinearly fitted through the convolution layer, the BN layer and the ReLU activation layer, and the residual path directly transmits the input features. The two features are added at the end of the block to form the final output, which is mathematically expressed as follows:
[0057] H(x) = F(x) + x
[0058] Where x is the input, F(x) is the convolution process through the main path. The residual structure can effectively avoid the problem of backward gradient disappearance or gradient explosion of the model during training, and can directly pass the original feature information to the subsequent layer, avoid the loss of key information, and enable the network to more easily train and learn the identity mapping. In addition, the module adopts a depth separable convolution instead of a traditional convolution layer in the design. The depth separable convolution is a lightweight convolution operation that decomposes the traditional convolution operation into a depth convolution and a point-by-point convolution two processes, and is currently widely used in mobile terminals and real-time neural networks.
[0059] After the model passes through the CNN space-time feature extraction module, it enters the BiLSTM time sequence network module.
[0060] First, the data processing process of the LSTM basic unit is as follows:
[0061]
[0062] In the formula, W and b represent the network parameters and bias parameters corresponding to the gate structure respectively; X t is the input at the current time; * represents the multiplication of corresponding elements of the matrix. C t is the memory unit, F t is the forget gate; I t is the input gate, O t is the output gate, and the output result H t is the hidden state input of the next unit.
[0063] Based on the above basic unit structure, the BiLSTM bidirectional time sequence network structure used by the application is as shown in Figure 3 The input sequence feature data is propagated through the forward LSTM and the reverse LSTM at the same time, and then the output features of each are spliced in the channel to be input into the full connection layer as the final output result for track prediction at the next time.
[0064] Step 4: The track prediction value obtained in step 3 and the true value label in the training set obtained in step 1 are input together, and a multi-task learning loss module with physical constraints is used to calculate the loss value of the current training iteration batch, and the model parameters are updated in reverse to complete the model training.
[0065] After calculating the track prediction value of the current training batch (Epoch), the satellite orbit prediction output value and the true value label at the prediction time in the training set obtained in step 1 are input into the physical constraint loss module to calculate the loss value of the current training batch. The calculation process of the module is as follows:
[0066] First, the mean square error is used to calculate the position loss L pos and the velocity loss Lvel The formula is as follows:
[0067]
[0068] In the formula, n represents the number of samples, and Y represents the true label. This represents the predicted value.
[0069] Then calculate the physical constraint loss L. kepler This loss is calculated based on Kepler's third law, the core of which is that the square of the orbital period is proportional to the cube of the semi-major axis. The calculation process is as follows:
[0070] First, calculate the satellite's mechanical energy. Considering that the satellite's orbital characteristics (such as orbital shape and period) are only related to the unit mass values of energy and angular momentum, the calculation expression is as follows:
[0071]
[0072] Based on Kepler's equations and the law of conservation of energy, the semi-major axis a can be obtained through energy calculation. E for:
[0073]
[0074] Calculate the numerical value of angular momentum h:
[0075] h = ||h|| = ||r×v||
[0076] In the formula, h represents the angular momentum vector; r represents the satellite position vector; and v represents the satellite velocity vector.
[0077] Calculate the eccentricity value e:
[0078]
[0079] Calculate the semi-major axis a from angular momentum and eccentricity. p :
[0080]
[0081] Calculate Kepler constraint loss L kepler :
[0082]
[0083] In the formula, τ is the scaling factor, which ensures that the magnitude of the loss is not too small.
[0084] In summary, the total loss is calculated as follows:
[0085] L total =λ·L pos +(1-λ)·Lvel + β · L kepler
[0086] In the formula, λ and β represent weight factors of loss, which need to be adjusted according to different source data sets through experiments.
[0087] Through the above steps, the loss value of the current training batch can be obtained, then the model can be updated in reverse, and the next batch of iterative training can be performed until the loss converges, at which time the model training is completed.
[0088] Fifth step: through steps two to four, the inference verification of the model is performed using the test set formed in step one, and the true value in the test set is compared, and the model test result can be obtained.
[0089] Using the test data set obtained in step one, the first 144 points of the sequence data are taken as the test input, and the input data format with the dimension of 1x144x9 is formed according to step two. The model trained in steps two to four is used to predict the satellite orbit position at all times in the test set in the form of sliding window described in step one. The obtained prediction value is compared with the true value of the orbit in the test set, and the prediction accuracy result of all times in the test set is obtained, which is in the form of three-axis position error of the satellite in the earth-centered inertial system, with the unit of meter. The present application uses the 60-day orbit data simulated by STK as the training and test data, and gives a set of model test results as an example, as shown in Table 1. Figure 4
Claims
1. A satellite orbit prediction method based on deep learning physical constraint loss, characterized in that: The steps of this method are as follows: Step 1: Normalize the input data to form training and test datasets, and build batch training data; Step 2: Expand the dimensions of the sample data points in each window of the batch data formed in Step 1 to construct a multi-dimensional feature space for the sample points, forming a batch input data format that can be passed into the model; Step 3: Use the model to perform forward inference on the batch data formed in Step 2. Use a lightweight CNN spatiotemporal feature extraction module and a BiLSTM bidirectional temporal neural network module to obtain the batch trajectory prediction value for the next time step output by the model. Step 4: Take the orbital prediction value obtained in Step 3 and the ground truth labels in the training set obtained in Step 1 as input, and calculate the loss value of the current training iteration batch through a multi-person learning loss module that integrates physical constraints. Then, perform reverse update of the model parameters to complete the model training. Step 5: Using the test set formed in Step 1, the model is inferred and verified through Steps 2 to 4, and compared with the true values in the test set to obtain the model test results.
2. The satellite orbit prediction method based on deep learning physical constraint loss according to claim 1, characterized in that: The specific process of step one is as follows: S11. Prepare satellite orbit sequence data for a period of time not less than 12 days. The format is set to give the three-axis position and velocity vector of a satellite in the geocentric inertial coordinate system every 300 seconds. The total time length is divided into a 7:3 ratio and used as training dataset and test dataset respectively. S12. For the training dataset, batch input data is formed in the form of a sliding window. The window length is set to 144 sample points and the sliding step size is set to 1. The sample points in each sliding window are used as training input. The model predicts the satellite position vector at the next moment of the sliding window. S13. For the test dataset, take the first 144 points of the sequence data as the initial test input, and predict the satellite position vector at the 145th time point and all subsequent time points in the test set in the form of a sliding window. The sliding window step size is 1, and the window size remains unchanged at 144. Each time the window slides, the predicted value of the previous time point is used as the last value of the sliding window input sequence at the next time point, and the first and first values in the sliding window are removed, keeping the total length of 144 unchanged. S14. Set the number of data batches during training to 64 window sequence data per batch. Set the test to not use batch processing, that is, only use the initial window of 144 sample data as input for recursive prediction.
3. The satellite orbit prediction method based on deep learning physical constraint loss according to claim 1, characterized in that: The specific process of step two is as follows: S21. For each window sequence of data, calculate the J2 perturbation characteristics at that moment for each sample point. S22. For each window sequence of data, calculate the periodicity feature of the sin function encoding at that moment for each sample point. S23. For each window sequence of data, calculate the periodicity feature of the cosine function encoding at that moment for each sample point. S24. For each window sequence of data, perform dimensional expansion on each sample point, and add the feature values calculated in steps S21, S22, and S23 to the original satellite position and velocity vector features as new feature dimensions, forming a batch input data format that can be passed into the model. The input dimension for each batch of training is 64×144×9.
4. The satellite orbit prediction method based on deep learning physical constraint loss according to claim 1, characterized in that: The specific process of step three is as follows: S31. Input the batch training data obtained in step two into the CNN lightweight spatiotemporal feature extraction module to obtain the output features with dimensions of 64×9×144. S32. Input the result data obtained in step S31 into the BiLSTM bidirectional time series network module to obtain the satellite orbit prediction value at the next moment. The dimension is 64×6, which includes position and velocity vectors respectively.
5. The satellite orbit prediction method based on deep learning physical constraint loss according to claim 1, characterized in that: The specific process of step four is as follows: S41. Take the satellite orbit prediction output value obtained in step 3 and the true value label of the prediction time in the training set obtained in step 1 as input, and pass them into the physical constraint loss module to calculate the loss value of the current training batch. S42. Perform reverse update of the model and iterate the next batch of training until the loss converges, at which point the model training is complete.
6. The satellite orbit prediction method based on deep learning physical constraint loss according to claim 1, characterized in that: The specific process of step five is as follows: S51. Using the test dataset obtained in step one, take the first 144 points of the sequence data as the initial test input, and use the model trained in steps two to four to predict the satellite orbit position at all times in the test set in the form of the sliding window described in S13. S52. Compare the difference between the predicted satellite orbit value obtained in S51 and the true value in the test set to obtain the test result of the model. The index is the three-axis position error of the satellite in the geocentric inertial frame.
Citation Information
Cited By
Space debris feature embedded characterization method based on physically guided learning
CN121190781A
Satellite orbit single-point state forecasting method based on deep learning
CN122065285A
Space-time prediction method for cloud track of satellite cloud picture
CN122115507A