A multi-source data driven method for calculating extremely low orbit atmospheric parameters
Patent Information
- Application Number
- CN202610950327.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-29
- Publication Date
- 2026-09-29
AI Technical Summary
然而,极低轨道大气受到太阳活动、地磁扰动等空间环境的强烈影响,其时空变化复杂且存在高度非线性特征,传统经验模型虽然能够提供长期稳定估计,但难以捕捉突发事件引起的快速变化;物理模型虽然能够反映事件驱动的动力学过程,但计算资源消耗大且实时性不足
(1)本发明通过多源数据融合显著提升了极低轨大气参数预测的精度。利用空间环境驱动数据、经验模型输出与事件驱动型物理模型输出,结合机器学习动态生成的融合权重,能够自适应调整不同模型的贡献,充分利用各类数据的优势,克服单一模型易受突发事件影响或长期偏差的问题,实现对大气密度、温度、主要成分数密度及水平风场的高精度预测。
Smart Images

Figure CN122839253A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of ELEAD atmospheric parameter calculation technology, specifically relating to a multi-source data-driven method for calculating ELEAD atmospheric parameters. Background Technology
[0002] In the operation of low Earth orbit (LEO) spacecraft, accurately acquiring atmospheric environmental parameters is crucial for ensuring orbit prediction, attitude control, and mission safety. However, the LEO atmosphere is strongly influenced by the space environment, including solar activity and geomagnetic disturbances, resulting in complex spatiotemporal variations with highly nonlinear characteristics. While traditional empirical models can provide long-term stable estimates, they struggle to capture rapid changes caused by sudden events. Physical models, although capable of reflecting event-driven dynamic processes, are computationally expensive and lack real-time performance. Existing methods often rely on single models or simple weighted fusion, making it difficult to balance high accuracy and timeliness. Furthermore, the lack of quantification of prediction uncertainties introduces risks into spacecraft operations.
[0003] Therefore, there is an urgent need for a method that can achieve adaptive and high-precision prediction of atmospheric parameters in very low Earth orbit, while also providing uncertainty assessment, in order to meet the needs of refined management and safety assurance in complex orbital environments. Summary of the Invention
[0004] One objective of this invention is to provide a method for calculating extremely low Earth orbit (ELEO) atmospheric parameters based on multi-source data. By fusing space environment-driven data, empirical model outputs, and event-driven physical model outputs, this method achieves high-precision, dynamically adaptive prediction of ELEO atmospheric environmental parameters. Through spatiotemporal unification and feature construction, combined with machine learning models to dynamically generate fusion weights, this method overcomes the shortcomings of traditional methods in terms of accuracy and timeliness. It can effectively respond to sudden space weather events while quantifying prediction uncertainties, thereby improving the reliability and accuracy of atmospheric parameter prediction.
[0005] To achieve the above objectives, the first aspect of the present invention provides a multi-source data-driven method for calculating extremely low Earth orbit atmospheric parameters, comprising the following steps: Acquire multiple data sources, including at least spatial environment-driven data, empirical model output data, and event-driven physical model output data; Spatiotemporal unification and feature construction are performed on multiple data sources to generate the optimal feature subset; Based on machine learning models and optimal feature subsets, fusion weights are dynamically generated. Based on the fusion weights, the output data of the empirical model and the output data of the physical model are adaptively fused to output the final atmospheric environmental parameters.
[0006] According to a specific embodiment of the present invention, multiple data sources are acquired, including: Acquire space environment-driven data, which include at least the solar radiation flux F10.7 and its 81-day moving average, and the geomagnetic activity indices Kp and Ap; Acquire the output data of the empirical model, which is generated by the empirical atmospheric model and the empirical wind field model running throughout the entire time period; Obtain the output data of the event-driven physics model, which is generated by the physics model when the preset space weather event trigger conditions are met.
[0007] According to a specific embodiment of the present invention, before acquiring the output data of the event-driven physics model, a preset space weather event trigger judgment is also included; Preset space weather event trigger judgment, including: Real-time monitoring of solar radiation flux F10.7 and geomagnetic activity index Kp in space environment-driven data; When F10.7 is detected to be greater than or equal to the first threshold, or F10.7 rises to exceed the second threshold within a preset number of days, or Kp is greater than or equal to the third threshold, an event is triggered and the physical model is started to simulate the event. After the event ends, the physical model will continue running for an additional preset duration before stopping.
[0008] According to a specific embodiment of the present invention, the first threshold is 150 sfu, the preset number of days is 3 days, the second threshold is 30 sfu, the third threshold is 5, and the preset duration is 24h.
[0009] According to a specific embodiment of the present invention, spatiotemporal unification and feature construction are performed on multiple data sources to generate an optimal feature subset, including: All data sources are unified to the same time base system, and the time resolution is standardized to 1 hour through interpolation. Establish a multidimensional spatiotemporal grid covering the very low Earth orbit region. The dimensions of the multidimensional spatiotemporal grid should include at least time, latitude, longitude, and altitude. Based on a multidimensional spatiotemporal grid, an input feature set and an output label set are constructed respectively. The output label set is used to train a supervised machine learning model. The input feature set includes spatial environment-driven data and spatiotemporal location parameters, including year-day (DoY), hour (hod), geographic latitude (lat), longitude (lon), and altitude (alt). The output label set includes empirical model output data and event-driven physical model output data. Based on the input feature set and the output label set, a coarse screening is first performed using mutual information, followed by a fine screening using recursive feature elimination to obtain the optimal feature subset. Mutual information is used to evaluate the nonlinear correlation between each input feature in the input feature set and each output label in the output label set. Recursive feature elimination uses random forest or linear regression as the base model to iteratively eliminate low-importance features.
[0010] According to a specific embodiment of the present invention, before constructing the input feature set and the output label set respectively, a derivative transformation of the input feature set is also included; Derivative transformations include: Perform sine / cosine encoding on the year-to-day (DoY) and hour-to-hour (hod) data. Construct the product interaction features between different parameters in spatial environment-driven data and spatiotemporal location parameters; Discretize and hierarchically encode continuous height variables.
[0011] According to a specific embodiment of the present invention, before dynamically generating fusion weights based on a machine learning model and an optimal feature subset, the method further includes constructing a machine learning model, which is a multi-task deep neural network model. Building machine learning models includes: Construct a shared underlying network to extract high-dimensional common features from the input feature set; On top of the shared underlying network, empirical model output header, physical model output header, and fusion coefficient output header are constructed respectively. A multi-task loss function is used to jointly supervise the training of the empirical model output head, the physical model output head, and the fusion coefficient output head.
[0012] According to a specific embodiment of the present invention, based on a machine learning model and an optimal feature subset, fusion weights are dynamically generated, including: After standardizing all continuous features in the optimal feature subset, high-dimensional shared features are extracted by stacking residual blocks or fully connected layers. The high-dimensional shared features are input into the output heads of the empirical model, the physical model, and the fusion coefficients, respectively, to obtain the output labels of the empirical model. Physical model output labels and fusion weight ;in ∈[0,1], fusion weights This indicates the degree of confidence in the physical model's correction terms; Based on empirical model output labels, physical model output labels, and fusion weights The system outputs the final atmospheric environmental parameters, including density, temperature, number density of major components, and horizontal wind field. The final atmospheric environmental parameters are calculated using the following formula: in, The empirical model output header outputs the empirical model output labels; Empirical model output labels ; The physical model output labels for the physical model output header. ; Physical model output labels ; The scalar fusion weights are output by the fusion coefficient output header.
[0013] According to a specific embodiment of the present invention, after adaptively fusing the empirical model output data and the physical model output data according to the fusion weights, the method further includes quantifying prediction uncertainty. The quantification of prediction uncertainty includes: A deep ensemble strategy is employed to train multiple machine learning models with different initializations; Calculate the mean and standard deviation of the prediction results from multiple machine learning models; The standard deviation is used as a measure of prediction uncertainty and is output along with the final atmospheric environmental parameters.
[0014] Compared with the prior art, the above-described solutions of the present invention have at least the following beneficial effects: (1) This invention significantly improves the accuracy of atmospheric parameter prediction in the ultra-low orbit by fusing multi-source data. By utilizing space environment-driven data, empirical model output, and event-driven physical model output, combined with dynamically generated fusion weights from machine learning, the contribution of different models can be adaptively adjusted. This fully utilizes the advantages of various types of data, overcomes the problem that single models are easily affected by sudden events or have long-term biases, and achieves high-precision prediction of atmospheric density, temperature, main component density, and horizontal wind field.
[0015] (2) This invention can provide the quantification of uncertainty in atmospheric parameter prediction, providing a reliable basis for spacecraft operation decisions. Multiple machine learning models are trained using a deep integration strategy, and the predicted mean and standard deviation are output, which not only provide the final atmospheric environmental parameters, but also quantify the prediction confidence. Attached Figure Description
[0016] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention. It is obvious that the drawings described below are merely some embodiments of the invention, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort. In the drawings: Figure 1 This is a flowchart of a multi-source data-driven method for calculating very low Earth orbit atmospheric parameters according to the present invention. Figure 2 This is a flowchart of another multi-source data-driven method for calculating ultra-low orbit atmospheric parameters according to the present invention. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0018] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a product or device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a product or device. Without further limitation, an element defined by the phrase "comprising one" does not exclude the presence of other identical elements in the product or device that includes that element.
[0019] The following is in conjunction with the appendix Figure 1-2 Detailed description of optional embodiments of the present invention.
[0020] like Figure 1 As shown in the figure, according to a specific embodiment of the present invention, the present invention provides a multi-source data-driven method for calculating extremely low Earth orbit atmospheric parameters, comprising the following steps: S100. Acquire multiple data sources, including at least spatial environment-driven data, empirical model output data, and event-driven physical model output data. S200: Perform spatiotemporal unification and feature construction on multiple data sources to generate the optimal feature subset; S300: Dynamically generates fusion weights based on machine learning models and optimal feature subsets; S400. Based on the fusion weights, the output data of the empirical model and the output data of the physical model are adaptively fused to output the final atmospheric environmental parameters.
[0021] This invention provides a multi-source data-driven method for calculating atmospheric parameters in Extremely Low Earth Orbit (ELEO). In this embodiment, ELEO refers to the Earth's orbital region with an altitude range of approximately 100km-300km.
[0022] The atmospheric environment in this region has a direct impact on the on-orbit safety, lifespan, and mission effectiveness of spacecraft: atmospheric temperature determines the design requirements of the thermal control system; atmospheric density affects orbital maintenance fuel consumption; atomic oxygen is the main corrosive component for the aging of surface materials of low-Earth orbit spacecraft; molecular nitrogen and atomic oxygen work together to accelerate material aging; and atmospheric wind speed affects attitude control and orbital dynamics modeling.
[0023] However, due to the extreme scarcity of direct observation data in the very low Earth orbit (LEO) region, traditional methods rely on two main independent technical approaches: one is empirical models based on statistical fitting of large amounts of historical observation data, such as NRLMSISE-00, NRLMSIS2.0, JB2008, and the DTM series. These models are computationally fast but struggle to accurately depict the violent fluctuations in transient events such as geomagnetic storms; the other is physical models based on the first principles of fluid dynamics, such as the Global Ionosphere-Thermosphere Model (GITM) and the Thermosphere-Ionosphere-Mesosphere Energy and Dynamics Model (TIMEGCM). These models can more realistically reflect transient characteristics but require enormous computational resources, with a single complete simulation potentially requiring hundreds to thousands of CPU hours. Furthermore, traditional empirical models generally do not provide wind speed information, resulting in a lack of functional dimensions.
[0024] Therefore, this invention constructs a multi-task deep neural network model that organically integrates space environment-driven data, the stable baseline prediction capability of empirical models, and the transient event characterization capability of event-driven physical models. This achieves high-precision and high-efficiency collaborative calculation of multiple parameters such as atmospheric temperature, density, atomic oxygen number density, molecular nitrogen number density, and horizontal wind field. Furthermore, it achieves smooth adaptive fusion of empirical and physical models through dynamically learned fusion weights. The technical solution of this invention will be described in detail below.
[0025] As an optional embodiment, this embodiment of the invention first acquires three main types of data sources: space environment-driven data, empirical model output data, and event-driven physical model output data. These three types of data sources play different roles: space environment-driven data provides external driving conditions for the entire system; empirical model output data provides stable baseline predictions with full-time coverage; and event-driven physical model output data provides high-fidelity transient response information during periods of active space weather.
[0026] As an optional embodiment, the aforementioned space environment driving data is a quantitative characterization of the external energy sources driving changes in atmospheric state, including at least the solar radiation flux F10.7 and its 81-day moving average, and the geomagnetic activity indices Kp and Ap. In this embodiment of the invention, these space weather driving data are specifically acquired in real-time or near real-time through the NASA / GSFC OMNI database. The OMNI database integrates observational data from multiple solar observation satellites and ground magnetometers, providing continuous, quality-verified time series of space environment parameters, typically with a time resolution of 1 hour.
[0027] The F10.7 index is a classic proxy for the intensity of solar extreme ultraviolet radiation. The 81-day moving average of F10.7 reflects the long-term background level of solar activity, while the difference between its daily value and the 81-day average reflects short-term activity fluctuations. The geomagnetic activity index Kp (range 0-9, step size 1 / 3) is a comprehensive measure of the intensity of global geomagnetic disturbances, released every 3 hours; the Ap index is a linearized geomagnetic index corresponding to Kp, facilitating numerical calculations. In this embodiment of the invention, after acquiring these data, data quality control is also required: identifying and removing outliers (such as invalid data marked 999) and outliers that significantly deviate from the statistical distribution; then, linear interpolation is used to fill in the data gaps after removal, ensuring the continuity and integrity of the data sequence.
[0028] The aforementioned empirical model output data is generated continuously by the empirical atmospheric model and the empirical wind field model throughout the entire time period, including both calm and active periods of space weather. In this embodiment of the invention, the empirical atmospheric model adopted is the NRLMSIS2.0 model.
[0029] NRLMSIS2.0 is an upgraded version of the classic MSIS series models developed by the U.S. Naval Research Laboratory. It integrates long-term observational data from satellite accelerometers (such as GRACE and GOCE) and can output atmospheric temperature, total mass density, and number density of major atmospheric components based on input time (year-days, hours), geographical latitude and longitude, altitude, and space weather parameters such as F10.7 and Ap. The empirical wind field model uses HWM14, which is based on statistical regression of a large amount of satellite and ground wind field observation data and can provide estimates of horizontal wind fields at specific time points.
[0030] The advantage of empirical models lies in their extremely fast computation speed, ease of use, and ability to provide stable baseline predictions based on long-term statistical regularities.
[0031] To construct the sample dataset, this embodiment of the invention sequentially calls NRLMSIS2.0 and HWM14 for computation at each spatiotemporal point of the training sample, and merges the outputs of the two models to form the empirical model output dataset. This dataset covers all time periods and is continuous throughout the entire time period.
[0032] The event-driven physical model output data is the most unique among the three types of data sources. It features on-demand triggering, meaning the physical model only generates data when preset space weather event triggering conditions are met, rather than running continuously throughout the day. The rationale behind this design is that while physical models can explicitly solve for the number density, temperature, and velocity of various neutral and ionic components based on complete fluid dynamics equations and describe energy injection processes such as Joule heating and auroral deposition, thus theoretically reflecting the transient characteristics of the atmosphere during space weather events more realistically, their computational cost is extremely high. During calm periods, empirical models can already provide relatively reliable predictions, eliminating the need to invoke the physical model. Therefore, activating the physical model only during the most critical time periods is a key strategy for achieving a balance between high accuracy and low computational cost.
[0033] As an optional embodiment, in this embodiment of the invention, the specific mechanism for triggering the judgment of a preset space weather event is as follows: Real-time monitoring of solar radiation flux F10.7 and geomagnetic activity index Kp in space environment-driven data.
[0034] For solar activity events, the triggering conditions are: F10.7 is greater than or equal to the first threshold of 150 sfu, or F10.7 rises above the second threshold of 30 sfu within a preset number of days (3 days). For geomagnetic activity events, the triggering condition is: Kp is greater than or equal to the third threshold of 5. As long as either of the above conditions is met, the event is determined to be triggered, and the physical model is started for simulation.
[0035] As an optional implementation, when the physical model is started, the output of the empirical model before the event occurs is used as the initial field, that is, the output of NRLMSIS2.0 and HWM14 at the moment before the event provides the initial temperature, density, component density and wind field distribution for GITM; at the same time, real-time space weather observation data is used as the driving input, including the time series of F10.7 and Kp / Ap exponents.
[0036] As an optional embodiment, regarding the termination of the event and the stopping of the physical model, this embodiment of the invention designs the following mechanism: For solar activity events, the event is determined to end when F10.7 falls below 150 sfu or remains stable or decreases for 3 days; for geomagnetic activity events, the event is determined to end when Kp is less than 4 for 12 consecutive hours. After the event is determined to end, the physical model does not stop immediately, but continues to run for an additional preset duration, which can be set to 24 hours, to simulate the recovery process of atmospheric conditions after a space weather event.
[0037] Simulations of the recovery period are crucial for accurately depicting the dynamic process by which atmospheric density and composition gradually return to calm levels after the event. After an extension of 24 hours, the physical model ceases operation, and subsequent time periods continue to rely solely on empirical models.
[0038] Through the aforementioned event-driven mechanism, the computational resources of the physical model are concentrated in approximately 10%-30% of the event time period, which can save 70%-90% of the computational cost compared to continuously running the physical model, while ensuring the accuracy of atmospheric parameter simulation during the event.
[0039] As an optional implementation, after acquiring multi-source data, it is necessary to integrate these data from different sources, with varying formats and inconsistent spatiotemporal resolutions, into a standardized dataset that can be directly used for training machine learning models. This embodiment of the invention achieves this goal through the following four sub-steps.
[0040] First, we need to unify the time base and standardize the resolution.
[0041] To unify all data sources to the same time base system, this embodiment of the invention uses Coordinated Universal Time (UTC) as the unified time base. Since different data sources have varying time resolutions, this embodiment of the invention standardizes the time resolution to 1 hour using interpolation.
[0042] Specifically, for indices such as F10.7 and Kp / Ap, linear interpolation is used for time-dimension resampling. For both empirical and physical model outputs, cubic spline interpolation is used to obtain a smoother time profile, as atmospheric parameters typically exhibit smooth changes over time. Minor discrepancies between the original timestamps and the target time grid points are also handled using the aforementioned interpolation method to ensure accurate time-dimension alignment of all data.
[0043] Then, a multidimensional spatiotemporal grid was established.
[0044] Based on time unification, a multidimensional spatiotemporal grid covering the very low Earth orbit region is established. This grid has at least four dimensions: time, latitude, longitude, and altitude.
[0045] Specifically: the time dimension covers the complete available data period of OMNI data, with a time step of 1 hour after standardization; the latitude dimension covers -90° to 90°, with a resolution selectable from 1° to 5° depending on application requirements; the longitude dimension covers -180° to 180°, with the resolution consistent with the latitude; and the altitude dimension covers the very low Earth orbit altitude range of 100km to 300km, with a resolution selectable from 5km to 10km. This four-dimensional spatiotemporal grid provides a unified spatial reference framework for all subsequent calculations, ensuring that solar geomagnetic data, empirical model outputs, and physical model outputs can be accurately matched and compared at the same spatiotemporal point.
[0046] Then, the input feature set and the output label set are constructed.
[0047] Based on the aforementioned multidimensional spatiotemporal grid, input feature sets and output label sets are constructed respectively.
[0048] The input feature set includes space environment-driven data and spatiotemporal location parameters. Specifically, the space environment-driven data includes: the daily value and 81-day moving average of solar radiation flux F10.7, and the geomagnetic activity indices Kp and Ap.
[0049] The spatiotemporal location parameters specifically include: day of year (DoY, ranging from 1 to 365 / 366), hour (hod, ranging from 0 to 23), latitude (lat, ranging from -90° to 90°), longitude (lon, ranging from -180° to 180°), and altitude (alt, ranging from 100 to 300 km).
[0050] It is important to note that the annual day of day (DoY) and the hourly day of day (hod) are key time parameters for characterizing the seasonal and diurnal variations of atmospheric parameters. To enhance the model's ability to perceive this periodic information, this embodiment of the invention performs sine / cosine encoding on the annual day of day (DoY) and the hourly day of day (hod) before constructing the input feature set. This smooths the periodic boundaries and avoids the discontinuous jumps that occur between January 1st and December 31st or between 11 PM and midnight when using the original values directly.
[0051] As an optional embodiment, in order to enhance the expressive power of features and capture the coupling effect between different physical quantities, the present invention performs the following derivative transformation on the input feature set: First, construct the product interaction features between different parameters in the space environment driving data and spatiotemporal location parameters. For example, the product of geomagnetic activity index Kp and solar radiation flux F10.7 reflects the coupling effect of solar activity and geomagnetic activity, and the product of geographical latitude and geomagnetic activity index Kp reflects the difference in the degree of response of different latitudes to geomagnetic activity. Such interaction features can explicitly characterize the nonlinear synergistic effect between multiple physical fields.
[0052] Second, the continuous altitude variable `alt` is discretized and hierarchically encoded. For example, the altitude range of 100-300 km is divided into several altitude layers, and each altitude layer is assigned a one-hot encoding or embedding vector to help the model learn the differentiated features of different altitude layers. Through these derivative transformations, the input feature set is expanded from the original few basic variables to a rich feature space containing dozens of features.
[0053] As an optional embodiment, the output label set includes all atmospheric parameters in the empirical model output data and the event-driven physics model output data.
[0054] Specifically: the output labels of the empirical model include the atmospheric temperature, total mass density, atomic oxygen number density, molecular nitrogen number density, and horizontal wind field output by the empirical atmospheric model; the output labels of the physical model include the atmospheric temperature, total mass density, atomic oxygen number density, molecular nitrogen number density, and horizontal wind field output by the physical model.
[0055] It is important to note that the output labels of the physical model only have values during the active phase and are null during the quiet phase. The output label set is used to provide training targets for subsequent supervised learning models.
[0056] Finally, the optimal feature subset is selected.
[0057] As an optional embodiment, after constructing a complete input feature set and output label set, this embodiment of the invention uses a two-stage feature selection method to select the optimal feature subset in order to eliminate redundant features, reduce model complexity, and prevent overfitting.
[0058] Phase 1: Coarse Screening with Mutual Information. Mutual information is a non-linear correlation measure based on information theory, capable of capturing any form of statistical dependency between input features and output labels, not limited to linear correlation. This embodiment of the invention calculates the mutual information value between each input feature and each output label, sets a mutual information threshold, and eliminates features whose mutual information with all output labels is below the threshold, achieving preliminary coarse-grained screening.
[0059] The second stage: Recursive feature elimination for fine-tuning. Building upon the initial coarse screening using mutual information, recursive feature elimination is employed for refined feature selection. Using random forest or linear regression as the base model, the base model is first trained with all remaining features, calculating the importance score for each feature. Then, iteratively, the feature with the lowest current importance is removed, the base model is retrained, and the feature importance is re-evaluated. This process is repeated until the number of remaining features reaches the preset optimal number, or the model performance begins to decline significantly. The final retained features are the optimal feature subset. Through this two-stage selection process, the optimal feature subset retains the most informative features for atmospheric parameter prediction while eliminating redundant and noisy features, laying an efficient and robust foundation for subsequent machine learning fusion modeling.
[0060] As an optional implementation, after generating the optimal feature subset and before generating the dynamic fusion weights, it is necessary to build a machine learning model for achieving intelligent fusion.
[0061] As an optional embodiment, this embodiment of the invention uses a multi-task deep neural network model as the core fusion engine. The construction of this model includes the construction of a shared underlying network, the construction of three independent output heads, and joint supervised training of multi-task loss functions.
[0062] As an alternative implementation, a shared underlying network forms the basis of a multi-task learning architecture, which extracts high-dimensional common feature representations that are applicable to all output tasks from the optimal feature subset.
[0063] In this embodiment of the invention, the shared bottom-layer network adopts a structure of alternating stacked residual blocks and fully connected layers. The core idea of residual blocks is to introduce skip connections, so that the output of the network layer not only includes the transformation result of the input of that layer, but also retains the information of the original input, thereby effectively alleviating the gradient vanishing problem in deep network training and making it easier for the model to learn meaningful feature representations.
[0064] Specifically, each residual block contains two fully connected layers, each followed by a ReLU activation function and a Dropout regularization layer. The Dropout ratio is set between 0.1 and 0.2 to prevent overfitting. The input of the residual block is added to the output of the second fully connected layer through a skip connection, and then passed through the activation function to form the final output of the residual block.
[0065] The shared underlying network typically stacks 2 to 4 such residual blocks, ultimately outputting a high-dimensional shared feature vector with a fixed dimension.
[0066] As an optional embodiment, on top of the shared underlying network, the embodiments of the present invention construct three functionally independent but shared underlying features output heads: an empirical model output head, a physical model output head, and a fusion coefficient output head.
[0067] The aforementioned empirical model output head consists of 1-2 fully connected layers. Its input is the high-dimensional shared features extracted from the underlying network, and its output is the predicted value of the empirical model's output label. During the training phase, this output head uses the empirical model's output label constructed in step S2 as the supervision target. The role of the empirical model output head is to learn the mapping from input features to the empirical model's output, enabling the model to reproduce its output even when the empirical model is not actually run during the inference phase.
[0068] The output head of the physics model also consists of 1-2 fully connected layers, and its output is the predicted value of the physics model's output label. During the training phase, this output head uses the output label of the event-driven physics model constructed above as the supervision target.
[0069] It's important to note that since the physical model output labels only have values during the active phase, the loss function of this output header is only calculated during the active phase. The role of the physical model output header is to learn the mapping from input features to the physical model output.
[0070] The fusion coefficient output head is the most innovative design of this invention. It consists of 1-2 layers of fully connected networks, directly learning to generate the optimal fusion weights from shared features and outputting a scalar. The output header uses the sigmoid activation function to constrain the output within the interval [0, 1].
[0071] Fusion coefficient The physical meaning of is the degree of confidence in the correction terms of the physical model: when When the value approaches 0, it indicates that the system tends to completely trust the empirical model; when... When the value approaches 1, it indicates that the system tends to fully adopt the corrections to the physical model. Instead of being a pre-set fixed value, it is automatically learned and generated by the neural network based on the input features, thus realizing the adaptability and dynamism of the fusion strategy.
[0072] After the above process is completed, the embodiments of the present invention dynamically generate fusion weights based on machine learning models and optimal feature subsets and output the final atmospheric environment parameters.
[0073] Specifically, all continuous features in the optimal feature subset are standardized. This embodiment of the invention employs the Z-score standardization method, which involves subtracting the mean of each continuous feature from the training set and then dividing by the standard deviation, resulting in standardized features with zero mean and unit variance. Standardization eliminates differences in units and numerical ranges between different features, making neural network training more stable and efficient.
[0074] Next, the standardized features are input into the already trained multi-task deep neural network model. The features first undergo forward propagation through the shared underlying network, being extracted and transformed layer by layer, ultimately yielding a high-dimensional shared feature vector. This high-dimensional shared feature is then input into three output heads: the empirical model output head outputs the empirical model labels; the physical model output head outputs the physical model labels; and the fusion coefficient output head outputs the scalar fusion weights. , ∈[0,1].
[0075] Based on the above three outputs, the final atmospheric environmental parameters are calculated using the following dynamic fusion formula: in, The final output is a multidimensional atmospheric environment parameter vector, which specifically includes atmospheric density ρ, temperature T, number density of major components, and horizontal wind field components. The empirical model output label vector, which is output by the empirical model output head, represents the baseline estimation of atmospheric parameters based on long-term statistical regularities. The physical model output label vector, output by the physical model output header, characterizes the transient changes in atmospheric parameters during space weather events. Scalar fusion weights generated for the fusion coefficient output header.
[0076] It is important to note that the output of this invention not only covers the temperature and density parameters provided by traditional empirical models, but also includes atomic oxygen number density, molecular nitrogen number density, and horizontal wind field information that are not available in traditional empirical models. These parameters are crucial for spacecraft material degradation assessment and orbital dynamics modeling. It is through this multi-task learning framework that the wind speed and composition information from the physical model and the wind field empirical model are integrated into a unified output system, enabling the final atmospheric environmental parameters to achieve complete coverage in the functional dimension.
[0077] As an optional embodiment, such as Figure 2As shown, after adaptively fusing the output data of the empirical model and the output data of the physical model according to the fusion weight, the method of this embodiment further performs prediction uncertainty quantification in step S500, and outputs the final atmospheric environmental parameters and uncertainty measure, providing key information about the confidence level of the prediction results for engineering applications.
[0078] The embodiments of the present invention employ a deep integration strategy to achieve uncertainty quantification.
[0079] Specifically, the S500 performs forecast uncertainty quantification and outputs the final atmospheric environmental parameters and uncertainty measures, including: During the model training phase, multiple multi-task deep neural network models with the same network structure but different parameters are independently trained using the same training data but different random initialization seeds.
[0080] During the inference phase, for the same input sample, these multiple models are independently subjected to forward inference, resulting in multiple predictions. The mean of these multiple predictions is calculated as a more robust estimate of the final output value, and its standard deviation is also calculated, which is the measure of the uncertainty of the prediction.
[0081] The physical meaning of standard deviation is: the degree of dispersion of prediction results due to the uncertainty of model parameters. The larger the standard deviation, the greater the difference in predictions between different models, the higher the uncertainty of predictions, and the more cautious one should be in engineering applications; the smaller the standard deviation, the more consistent the predictions of each model, and the higher the reliability of the prediction results.
[0082] In engineering applications, this uncertainty metric has significant value: for example, when the prediction uncertainty is low, the spacecraft can operate according to the normal orbit maintenance strategy; when the prediction uncertainty is high, ground operators can refer to the uncertainty information and make more conservative decisions, such as adjusting the orbit maintenance plan in advance and strengthening the protective measures of surface materials, thereby achieving intelligent decision-making based on risk perception.
[0083] Through the above embodiments, this step not only achieves optimal weighted fusion of the outputs of the empirical model and the physical model, but also fully utilizes the dynamic fusion weights generated by machine learning, enabling the final atmospheric parameter output to achieve a balance between long-term stability and event responsiveness. This method can significantly improve the prediction accuracy of atmospheric parameters for very low Earth orbit (LEO) satellites, ensuring reliable prediction results even under space weather events. Simultaneously, it enhances the interpretability and operability of the model through recording and uncertainty quantification. The technical effects of this embodiment are significant, and the adaptability of the fusion process gives the method broad application potential, supporting various scenarios such as precise tracking and control of LEO satellite orbits, collision avoidance, and space weather risk assessment.
[0084] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various modifications and variations. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention. Finally, it should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to mutually. For the systems or apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and relevant parts can be referred to the method section.
[0085] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multi-source data-driven method for calculating very low Earth orbit atmospheric parameters, characterized in that, Includes the following steps: Acquire multiple data sources, including at least space environment-driven data, empirical model output data, and event-driven physical model output data; Spatiotemporal unification and feature construction are performed on the multiple data sources to generate the optimal feature subset; Based on the machine learning model and the optimal feature subset, the fusion weights are dynamically generated. Based on the fusion weights, the output data of the empirical model and the output data of the physical model are adaptively fused to output the final atmospheric environmental parameters.
2. The method according to claim 1, characterized in that, The acquisition of multiple data sources includes: The space environment driving data is acquired, which includes at least the solar radiation flux F10.7 and its 81-day moving average, and the geomagnetic activity index Kp and Ap. The empirical model output data is obtained, which is generated by the empirical atmospheric model and the empirical wind field model running throughout the entire time period. Obtain the output data of the event-driven physics model, which is generated by the physics model when the preset space weather event triggering conditions are met.
3. The method according to claim 2, characterized in that, Before acquiring the output data of the event-driven physics model, a preset space weather event trigger judgment is also included; The preset space weather event trigger judgment includes: Real-time monitoring of the solar radiation flux F10.7 and geomagnetic activity index Kp in the space environment driving data; When F10.7 is detected to be greater than or equal to the first threshold, or F10.7 rises to exceed the second threshold within a preset number of days, or Kp is greater than or equal to the third threshold, an event is triggered and the physical model is started to simulate. After the event ends, the physical model will continue to run for an additional preset duration before stopping.
4. The method according to claim 3, characterized in that, The first threshold is 150 sfu, the preset number of days is 3 days, the second threshold is 30 sfu, the third threshold is 5, and the preset duration is 24 hours.
5. The method according to claim 1, characterized in that, Spatiotemporal unification and feature construction are performed on the multiple data sources to generate an optimal feature subset, including: All data sources are unified to the same time base system, and the time resolution is standardized to 1 hour through interpolation. Establish a multidimensional spatiotemporal grid covering the very low Earth orbit region, wherein the dimensions of the multidimensional spatiotemporal grid include at least time, latitude, longitude, and altitude; Based on the multidimensional spatiotemporal grid, an input feature set and an output label set are constructed respectively. The output label set is used to train a supervised machine learning model. The input feature set includes spatial environment-driven data and spatiotemporal location parameters. The spatiotemporal location parameters include year-day (DoY), hour (hod), geographic latitude (lat), longitude (lon), and altitude (alt). The output label set includes empirical model output data and event-driven physical model output data. Based on the input feature set and the output label set, a coarse screening is first performed using mutual information, followed by a fine screening using recursive feature elimination to obtain the optimal feature subset; wherein the mutual information is used to evaluate the nonlinear correlation between each input feature in the input feature set and each output label in the output label set; the recursive feature elimination uses random forest or linear regression as the base model to iteratively eliminate low-importance features.
6. The method according to claim 5, characterized in that, Before constructing the input feature set and output label set respectively, the method further includes performing a derivative transformation on the input feature set; The derived transformation includes: Perform sine / cosine encoding on the year-day (DoY) and hour (hod) data. Construct the product interaction features between different parameters in spatial environment-driven data and spatiotemporal location parameters; Discretize and hierarchically encode continuous height variables.
7. The method according to claim 1, characterized in that, Before dynamically generating fusion weights based on the machine learning model and the optimal feature subset, the process further includes constructing the machine learning model, which is a multi-task deep neural network model. The construction of the machine learning model includes: Construct a shared underlying network to extract high-dimensional common features from the input feature set; On top of the shared underlying network, an empirical model output head, a physical model output head, and a fusion coefficient output head are constructed respectively. A multi-task loss function is used to jointly supervise the training of the empirical model output head, the physical model output head, and the fusion coefficient output head.
8. The method according to claim 7, characterized in that, The dynamic generation of fusion weights based on the machine learning model and the optimal feature subset includes: After standardizing all continuous features in the optimal feature subset, high-dimensional shared features are extracted by stacking residual blocks or fully connected layers. The high-dimensional shared features are respectively input into the output heads of the empirical model, the physical model, and the fusion coefficients to obtain the output labels of the empirical model. Physical model output labels and fusion weight ;in The fusion weights are ∈[0,1]. This indicates the degree of confidence in the physical model's correction terms; Based on the empirical model output labels, physical model output labels, and fusion weights The system outputs the final atmospheric environmental parameters, including density, temperature, main component number density, and horizontal wind field. The final atmospheric environmental parameters are calculated using the following formula: in, The empirical model output label is the empirical model output header output. The empirical model outputs labels. ; The physical model output label for the physical model output header ; The physical model outputs labels ; The scalar fusion weight is output by the fusion coefficient output head.
9. The method according to claim 1, characterized in that, After adaptively fusing the empirical model output data and the physical model output data according to the fusion weights, the method further includes quantifying prediction uncertainty. The quantification of prediction uncertainty includes: A deep ensemble strategy is employed to train multiple machine learning models with different initializations. Calculate the mean and standard deviation of the prediction results from the multiple machine learning models; The standard deviation is used as a measure of prediction uncertainty and is output together with the final atmospheric environmental parameters.