Method and system for complementing soft X-ray photon counting rate through deep learning based on physical driving
By employing a physics-driven deep learning approach and utilizing a photon count rate interpolation model based on solar wind and geomagnetic characteristic variables, the problem of missing segments in soft X-ray photon count rate data was solved. This approach achieved accurate data completion and consistency with physical laws, thereby improving the reliability of scientific research.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NAT SPACE SCI CENT CAS
- Filing Date
- 2025-12-22
- Publication Date
- 2026-05-08
AI Technical Summary
Current technologies cannot effectively process missing segments in soft X-ray photon count rate data, resulting in incomplete observation data and affecting the accuracy and reliability of scientific research.
A physics-driven deep learning approach is adopted. A photon count rate interpolation model is designed, which includes first and second auxiliary sub-networks and a main network. The photon count rate is completed by using characteristic variables such as proton number density, solar wind velocity, ion abundance ratio and geomagnetic index in solar wind, combined with causal convolutional layers and Transformer encoder blocks.
It enables accurate completion of soft X-ray photon count rate data, restores temporal continuity, enhances data integrity, ensures that the completion results conform to physical laws, and improves the reliability of scientific research.
Smart Images

Figure CN121997706A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the fields of soft X-ray, solar wind charge exchange, data completion, time series, and deep learning, and specifically relates to a physics-driven deep learning method and system for completing soft X-ray photon count rates. Background Technology
[0002] Photon count rate is widely recognized for its high signal-to-noise ratio and efficient resource utilization, playing a crucial role in space physics. In soft X-ray research, photon count rate, as a fundamental observational indicator, enables the precise quantification of soft X-ray radiation intensity and temporal variations, thus supporting a wide range of studies, including solar flare dynamics, stellar canopy plasma evolution, and high-resolution imaging of soft X-ray spatial distribution. This makes accurate photon counting indispensable for advancing physical interpretation and observational precision.
[0003] Even so, the photon count rate of soft X-rays is frequently contaminated during space exploration. For example, during observations conducted by EPIC aboard the XMM-Newton satellite, a magnetosphere reconnection event accelerated low-energy solar protons (energy less than 300 keV), scattering them along with soft X-ray photons to the focal plane. This resulted in soft proton contamination, polluting the observational data and preventing it from accurately reflecting the true soft X-ray photon count rate. These soft proton flares are sudden and have a significant impact. In XMM-Newton observations, the signal from soft proton contamination can exceed 1000% of the stationary background signal and can last from approximately 100 seconds to several hours, affecting up to 40% of the observation time. The distribution of these solar protons is similar to that of true soft X-ray photons, making them difficult to distinguish directly by spatial or morphological characteristics, thus hindering their separation from pure soft X-rays. Consequently, this contaminates a large amount of space science observational data.
[0004] During periods of soft proton contamination, the absolute count rate and its temporal distribution are altered, obscuring physical characteristics. Using contaminated data can distort the interpretation of soft X-rays, introduce systematic errors, and reduce the reliability of scientific conclusions. Therefore, contaminated soft X-ray photon count rate data cannot be used directly. In studies, contaminated data segments are typically labeled and excluded, creating gaps in the photon count rate time series. Therefore, completing this missing data is crucial. Completion not only restores the signal but also restores temporal continuity, enhances data integrity, and ensures the reliability of downstream analyses relying on complete, physically consistent photon time series.
[0005] Researchers have developed several toolkits to support photon count rate analysis. Among them, Stingray is a widely used Python library that provides basic functionality for spectral time series analysis. While it focuses on photon count rate time series, it is not well-suited for handling incomplete data. It lacks built-in functionality for gap-filling or missing data modeling, which limits its application in practical observation scenarios. This highlights the need for complementary methods that reconstruct continuous sequences before analysis, especially when signal integrity is critical for downstream tasks.
[0006] Most existing methods for processing soft X-ray photon count rate data focus primarily on signal denoising and spectral analysis, without directly addressing the interpolation of missing segments. Therefore, these techniques cannot recover continuous photon count rate sequences from incomplete and irregularly sampled observation data. Summary of the Invention
[0007] The purpose of this application is to overcome the shortcomings of existing technologies in accurately and physically complete incomplete photon count rate observation data.
[0008] To achieve the above objectives, this application proposes a physics-driven deep learning-based method for completing soft X-ray photon count rates, comprising: The preprocessed measurement data is input into the trained photon count rate interpolation model, and the completed photon count rate is output. The photon count rate interpolation model includes a first auxiliary sub-network, a second auxiliary sub-network, and a main network; The first auxiliary sub-network is used to convert the input ion abundance ratio into the SWCX scaling factor; The second auxiliary subnetwork is used to convert the input geomagnetic index into neutral hydrogen density; Both the first and second auxiliary sub-networks include a Transformer encoder block and a fully connected layer; The main network consists of three one-dimensional convolutional layers, two Transformer encoder blocks, and two fully connected layers connected in sequence.
[0009] As an improvement to the above method, the three one-dimensional convolutional layers are all causal convolutional layers with dilation rates of 1, 2 and 4, respectively, and the kernel size of each layer is 3.
[0010] As an improvement to the above method, the data processing procedure for the two fully connected layers includes: The first fully connected layer uses the ReLU activation function to compress the embedding at each time step into a low-dimensional latent space; The second fully connected layer performs a linear projection at each time step, generating the initial output sequence of the model.
[0011] As an improvement to the above method, the measurement data includes: proton number density in the solar wind, solar wind velocity, ion abundance ratio, geomagnetic index, and photon count rate in the 2.5-5.0 keV energy band.
[0012] As an improvement to the above method, the calculation process of the photon count rate interpolation model includes: The ion abundance ratio is input into the first auxiliary sub-network, and the SWCX scaling factor is output. Input the geomagnetic index into the second auxiliary sub-network to output the neutral hydrogen density; The number density of protons in the solar wind, the velocity of the solar wind, the ion abundance ratio, the geomagnetic index, the photon count rate of the 2.5-5.0 keV energy band, the SWCX scaling factor, and the neutral hydrogen density are concatenated along the feature dimension, and then broadcast together with the subset identifier, the binary mask used to distinguish missing regions, and the fixed sinusoidal position code to a common embedding space and combined to form the final multivariate input tensor. The multivariate input tensor is input into the main network, and the initial completion result, i.e., the complete photon count rate time series, is output. From the initial result sequence, the central part is extracted according to the set number of steps to obtain the final output result, which is the photon count rate time sequence after focusing on the key region.
[0013] As an improvement to the above method, the photon count rate interpolation model also includes an auxiliary physical loss module for calculating the physical theoretical value of the photon count rate; The processing procedure of the auxiliary physical loss module includes: Each subset identifier is mapped to an embedding vector, and the subset correlation coefficients are obtained through a fully connected layer. and ; Photon counting rate from 2.5–5.0 keV energy band using one-dimensional convolutional layers. Extract the invariant components from the background signal in the 2.5-5.0 keV energy range. ; use minus Time-varying soft proton signals were obtained from the background signal in the 2.5-5.0 keV energy range. ; The background signal in the target energy range of 0.5-0.7 keV is calculated using the following formula. : ; The soft X-ray photon count rate in the 0.5–0.7 keV energy range is calculated using the following formula. : ; in, SWCX scaling factor; It has a neutral hydrogen density; The density of protons in the solar wind; The speed of the solar wind; For thermal velocity; This is the actual line of sight of the satellite.
[0014] As an improvement to the above method, the training process of the photon count rate interpolation model includes: Data sets containing target sequences of proton number density, solar wind velocity, ion abundance ratio, geomagnetic index, photon count rate in the 2.5–5.0 keV energy band, and photon count rate in the 0.5–0.7 keV energy band are preprocessed; The preprocessed dataset is divided into training and test sets according to a set ratio; The ion abundance ratio in the training set is input into the first auxiliary sub-network, and the SWCX scaling factor is output. The geomagnetic index in the training set is input into the second auxiliary sub-network, which outputs the neutral hydrogen density. The number density of protons in the solar wind, the velocity of the solar wind, the ion abundance ratio, the geomagnetic index, the photon count rate of the 2.5-5.0 keV energy band, the SWCX scaling factor, and the neutral hydrogen density are concatenated along the feature dimension, and then broadcast together with the subset identifier, the binary mask used to distinguish missing regions, and the fixed sinusoidal position code to a common embedding space and combined to form the final multivariate input tensor. The multivariable input tensor is input into the main network, and the potential photon counting sequence of the first set number of steps is output. The potential photon counting sequence of the first set number of steps is input into the auxiliary physical loss module, and the completed photon counting rate and auxiliary physical loss are output. The numerical main loss of the main network and the auxiliary physical loss of the auxiliary physical loss module are weighted and combined to form the loss function of the photon count rate interpolation model. By using the backpropagation algorithm, the weights and biases of each layer are adjusted iteratively in reverse along the network layers based on the error between the imputed value calculated by the loss function and the true value, thereby minimizing the loss function and improving the imputed accuracy of the photon count rate interpolation model.
[0015] As an improvement to the above method, the preprocessing includes: Time alignment and synchronization of all data sources; Remove invalid values from the data and perform normalization. The data processed as described above is used to construct a time series dataset with a fixed window length by sliding a window with a set step size.
[0016] This application also provides a physics-driven deep learning-completed soft X-ray photon counting rate system, implemented based on the above method, the system comprising: The preprocessing module is used to preprocess the measurement data; The photon count rate completion module is used to input the data output from the preprocessing module into the trained photon count rate interpolation model and output the completed photon count rate.
[0017] Compared with existing technologies, the advantages of this application are: 1. Unlike traditional causal models that rely solely on historical information, the network in this application simultaneously considers both past and future observational data. This capability allows it to utilize richer contextual cues on both sides of the missing fragment, resulting in a more comprehensive and accurate estimation. Ablation experiments highlight the importance of this design choice: performance significantly degrades when bidirectional attention is replaced with causal attention.
[0018] 2. This application introduces physical loss to integrate domain prior knowledge into the training process. Through constraints, the completion results are made as consistent as possible with the theoretical values calculated by physical formulas, guiding the network towards solutions that satisfy physical meaning. This constraint acts as a regularization mechanism, avoiding unrealistic fluctuations that are numerically consistent but violate physical principles. Removing this constraint in ablation experiments also led to a significant decrease in completion accuracy, further emphasizing its contribution. Attached Figure Description
[0019] Figure 1 The diagram shows the data preprocessing flowchart before photon count rate completion; Figure 2 The following is used for estimation and The architecture consists of two sub-networks. Each sub-network takes 100 time-step feature variables as input, passes through a TEB, and is output via a fully connected layer. The output is a 100-step sequence of estimated latent physical quantities, which is then used to support physical information modeling in the main network. Figure 3 The diagram shows the architecture of the main network. It receives multiple feature tensors and the output tensors from the subnetworks, then passes them through stacked dilated causal convolutional layers and two TEBs, followed by two fully connected layers, to output the completed photon count rate. Figure 4 The diagram shows the overall framework of the model. The two auxiliary subnetworks first infer... and Then it is concatenated with other features and represented along with the subset of data. The mask and position code are input into the main network. The output of the main network enters the physical loss layer, where the background signal conversion between different energy bands is completed, and then the physical loss is calculated, thereby ensuring that the reconstruction result is numerically accurate and conforms to physical laws. Figure 5 The diagram shows the structure of an expanded causal convolutional block used for local feature extraction. Three causal one-dimensional convolutional layers are applied sequentially, with dilation strides of 1, 2, and 4, and each layer has a kernel size of 3. Figure 6 The diagram shows the architecture of the background transformation sub-network. This module will... Estimated values converted to the 0.5–0.7 keV energy range This supports the calculation of physical information loss downstream; Figure 7(a) shows a representative example of the model's completion performance when one time step is missing; Figure 7(b) shows a representative example of the model's completion performance when two time steps are missing; Figure 7(c) shows a representative example of the model's completion performance when 5 time steps are missing; Figure 7(d) shows a representative example of the model's completion performance when 10 time steps are missing. Detailed Implementation
[0020] The technical solution of this application will be described in detail below with reference to the accompanying drawings.
[0021] Example 1 This application presents a physics-driven deep learning-based method for completing soft X-ray photon count rate data, addressing the task of completing irregularly missing time series data in soft X-ray photon count rate datasets. Photon count rate sequences originate from satellite observations, but due to contamination or instrument limitations, these sequences often contain missing intervals. Therefore, the goal of this application's method is to reconstruct the missing values in the observation sequence. In this application, the middle data is missing, but the preceding and following observation data are available. This allows for the utilization of bidirectional temporal context to improve the accuracy of the reconstruction. The method aims to reconstruct missing photon count rate segments while utilizing surrounding contextual information and relevant physical variables.
[0022] Soft X-rays consist primarily of two parts: the solar wind charge exchange (SWCX) induced signal and the background signal. SWCX occurs when highly charged solar wind ions exchange charge with neutral atoms, most commonly in Earth's outer atmosphere or the extended atmospheres of comets and planetary bodies. This interaction results in the trapping of electrons, which then emit soft X-ray photons via radiative transitions. The intensity of these photons fluctuates with changes in solar wind conditions and geomagnetic activity. In addition, the background signal, generated by instrument artifacts, high-energy particles, and diffuse cosmic sources, also contributes to soft X-rays.
[0023] To isolate the influence of solar wind charge-exchange emission on soft X-ray observations, this application employs a spectral band segmentation method. The 0.5–0.7 keV energy band was identified as the region primarily exhibiting SWCX enhancement. Conversely, the 2.5–5.0 keV band was selected to represent a continuous background because this region has essentially no spectral line emission and is assumed to be unaffected by the SWCX process.
[0024] Based on this spectral segmentation method, this application aims to fill the gap in photon count rate within the 0.5–0.7 keV band, hereinafter referred to as... This band contains signals from solar wind charge exchange as well as background radiation, and can therefore be used as a contaminated observation. On the other hand, the 2.5–5.0 keV band is considered to contain only background components. By utilizing this assumption, we aim to reconstruct and correct soft X-ray data for the SWCX sensitive band, thereby more clearly identifying the emission structures of the heliosphere and corona.
[0025] The method in this application is based on a physically driven model that describes the process of generating soft X-ray photons through a solar wind charge exchange mechanism. Cravens proposed the X-ray emissivity generated in this process. This power density represents the X-ray power produced per unit volume: (1) in, It is the density of neutral hydrogen ( ), Solar wind proton density ( ), It is the speed of the solar wind ( ), It is an efficiency factor that encompasses all detailed atomic physics processes. ). It is power density ( This corresponds to the generation of soft X-ray energy per unit volume caused by charge exchange in the solar wind.
[0026] To estimate the intensity of soft X-rays induced by SWCX observed on Earth. Cravens employs power density per unit volume along the line of sight. The method of integration in the heliocentric coordinate system extends outward from the Sun: (2) However, the physical quantity recorded by satellite observations is the photon count rate, not the generation rate. To convert radiation energy into the number of photons, the integrand needs to be divided by the average energy of soft X-ray photons, denoted as... This yields the photon generation rate per unit volume. : (3) In this application, this framework is extended by performing integration along the actual satellite observation direction, starting from Earth's orbit and intersecting regions such as the magnetic sleeve and geocorona. This path-dependent formula better captures the spatial variations in neutral hydrogen and ion density that influence SWCX emission near Earth. Integration along the actual satellite line of sight The soft X-ray photon count rate induced by SWCX was obtained: (4) constant It is redefined as the SWCX scaling factor. To account for changes in ion energy, the rate term is further extended to include thermal velocity, resulting in... Adding the SWCX-induced portion to the background signal portion yields the soft X-ray photon count rate in the 0.5–0.7 keV energy range: (5) To simulate the background signal in the 0.5-0.7 keV energy range Introducing a background photon count rate in the 2.5–5.0 keV energy range The background signal differs across energy bands, but it always consists of two parts: an invariant component and a time-varying soft proton signal. The invariant component and the time-varying soft proton signal exhibit a linear relationship between the background signals of different energy bands. Therefore, the following formula is used to calculate the background photon count rate for the 2.5–5.0 keV energy band. This is used to represent the background signal in the 0.5-0.7 keV energy range. .
[0027] (6) in, and These represent the time-varying and invariant components of the background signal in the 0.5-0.7 keV energy range, respectively. and These represent the time-varying and invariant components of the background signal in the 2.5–5.0 keV energy range, respectively. (Coefficients) , This demonstrates the multiple relationship between the time-varying and constant components of the background signal across different energy bands, and that the values differ across different observation periods.
[0028] The final form of the model adopted in this application is Equations (5) and (6), which provide a physical basis approximation of the soft X-ray photon count rate observed in the 0.5–0.7 keV band, integrating the components caused by SWCX as well as the background components.
[0029] In practical applications, several terms in formula (5), such as empirical coefficients and neutral hydrogen density, are difficult to measure directly. Therefore, ion abundance ratio and geomagnetic perturbation index can be used to approximate these components. To facilitate data-driven modeling, this application makes two simplifying assumptions. First, the thermal rate term is ignored. Because under typical conditions, it is much smaller than the overall speed of the solar wind. Therefore, its contribution to relative velocity is negligible. Secondly, the effects of the line-of-sight integration path and satellite observation geometry are omitted, assuming that the observation geometry remains relatively stable over short time intervals.
[0030] Based on this theoretical framework, this application designs an interpolation model as a multivariate time series regression framework for interpolating contaminated data. To capture the main drivers of soft X-ray variations in the case of SWCX, the model introduces five input features: , Ion abundance ratio ( Defined as ), geomagnetic index ( )as well as These features are used as input variables in the data-based photon count rate interpolation model of this application. The selection of these features is based on their physical relevance to SWCX generation and propagation.
[0031] The selection of the above features is based on physical theory. Specifically, This represents the number density of protons in the solar wind. This parameter controls the flux of charged particles interacting with neutral atoms in Earth's exosphere, directly affecting the occurrence rate of charge exchange events. A higher proton density increases the probability of collisions, thereby increasing the production of soft X-ray photons via the solar wind (SWCX).
[0032] The speed of the solar wind controls the temporal dynamics of the interaction region. Faster solar wind flows lead to higher-energy collisions and alter the spatial extent and intensity of the X-ray radiation region.
[0033] Ion abundance ratio It is a diagnostic tool for solar coronal ionization conditions and reflects the relative abundance of ions capable of effective charge exchange. Since SWCX emission efficiency depends largely on the availability of highly charged ions, fluctuations in this ratio directly modulate the expected X-ray output.
[0034] The index captures the global response of Earth's magnetosphere to solar wind disturbances and reflects the compression or expansion of the magnetosphere region. This, in turn, affects the density and distribution of neutral particles in Earth's corona. Therefore, It can serve as an indirect indicator of the evolution of the SWCX interaction space configuration.
[0035] at last, That is, the photon counting rate in the 2.5-5.0 keV energy band provides a reference for estimating the background X-ray intensity in the 0.5-0.7 keV energy band not caused by SWCX.
[0036] Overall, the variables mentioned above collectively encode the main physical mechanisms controlling the generation of soft X-ray photons, enabling predictive models to learn the complex dependencies between solar-Earth conditions and soft X-ray radiation patterns.
[0037] To mathematically express the completion task of this application, the target sequence of photon counting rates in the 0.5–0.7 keV energy range is represented as follows: ,in This is the time step number. The region to be padded is represented by a binary mask. Specify, where This indicates a missing value; elsewhere it is... The input is modeled as a set of parallel univariate sequences: (7) in, It is the number of features, each Each input represents a distinct physical or structural characteristic variable. These inputs are processed through separate encoded streams and collectively participate in the interpolation process.
[0038] The purpose of this application is to learn mappings parameterized by deep neural networks. This makes the interpolation value For each satisfied of accomplish .
[0039] The model is designed as a multi-input, multi-output architecture, where the main output is... This represents the completed photon count rate sequence. During training, in addition to the main output, the physical loss estimated based on theoretical constraints is also output as an auxiliary supervision signal.
[0040] To guide the learning process, this application employs a composite objective function that combines three losses based on mean squared error (MSE): one to compute completion accuracy over missing intervals, one defined at boundaries to facilitate a smooth transition, and another to ensure physical consistency with the theoretical SWCX model, as shown in Equation (5). These components collectively ensure the local accuracy and physical plausibility of the completed photon count rate.
[0041] The raw data used in this application, including soft X-ray photon count rates, solar wind-related parameters, and geomagnetic indices, were initially recorded at a fixed temporal resolution. However, contamination and instrumentation errors led to the removal of corrupted data, creating discontinuities in what was originally a periodically sampled sequence. These interruptions rendered the data unsuitable for direct input into neural network models, which require temporally aligned, numerically stable input sequences where observations are valid and semantically coherent.
[0042] like Figure 1 As shown, in order to address these challenges and ensure compatibility with time series learning architectures, this application implements a structured preprocessing flow consisting of four stages: (1) time alignment and synchronization of all data sources; (2) feature scaling and normalization to ensure numerical stability during training; (3) constructing a sliding window to convert continuous time series into supervised learning samples; and (4) dividing the dataset into training and test sets for model evaluation.
[0043] Figure 1 This paper outlines a four-step preprocessing procedure to transform raw multivariate observation data into structured input-output samples required for training. The original input consists of six different sequences: , , , , and Each sequence has its own original time length. This reflects different data sources and resolutions. The first step generates a shape as... The unified dataset. The second step is to remove invalid values and perform normalization to obtain a dataset with shape [shape missing]. The data is cleaned. Next, the third step involves segmenting the sequence into overlapping sequences of 100 time steps each, generating... The final step involves dividing the dataset into training and testing sets in an 80:20 ratio, resulting in feature input-output pairs with shapes of [shapes to be filled in]. and The input and output pairs of the label have shapes respectively. and The preprocessed samples form the basis for model training and evaluation.
[0044] The first challenge in preprocessing stemmed from inconsistencies in the time representations of the input variables. Soft X-ray photon counts in the 0.5–0.7 keV and 2.5–5.0 keV bands were initially timed using the XMM-Newton method, which defines time as the number of seconds elapsed since 00:00:00 on January 1, 1998. Conversely, all solar wind parameters and geomagnetic indices were time-stamped in Coordinated Universal Time (UTC), resulting in an inconsistent time base and making direct alignment of the multivariate series impossible. To bridge this gap, the timestamps of all photon counts were converted to the UTC standard. This conversion ensured that the photon count data were temporally synchronized with the external variables and aligned on a common timeline.
[0045] Once the time base is harmonized, differences in time resolution should be resolved. Among the input variables, and Provided in one-minute increments, this matches the resolution of soft X-ray data, thus requiring no further modification. In contrast, Provided at hourly intervals, each value represents the starting condition for that hour. To convert this into a minute-by-minute sequence suitable for time modeling, this application employs a forward-padding method, applying the same... The value is assigned 60 minutes after the timestamp. This exhibits related but different resolution issues. Each The value represents the midpoint of the corresponding hour, a format that cannot be directly mapped to minute values. To address this issue, a center-aligned expansion strategy is employed: each hour... The value is extended to all minutes of that hour, effectively resampling the data while preserving the implicit time anchor.
[0046] Furthermore, some input variables exhibit a systematic time lag between the occurrence of the physical phenomenon and the availability of observational data. For example, those used as input features in the model... These values are acquired by remote sensing instruments at upstream observation points. They experience a known time delay before reaching Earth. To ensure proper temporal alignment of all input features, these variables need to be adjusted accordingly based on the known delay characteristics.
[0047] To correct for this delay, this application employs a standard physical assumption that the average propagation time from L1 to Earth is typically approximately 60 minutes. For each observation period in the dataset, the window duration remains constant, but its start and end times are shifted forward by a full hour. Within this shifted window, the average... Then use this value to estimate the propagation time. (In minutes), the formula is as follows: (8) in, It is the distance from the Sun to the Earth (1.5 million kilometers). It is the average solar wind speed (in kilometers per second) calculated over the shift window. The resulting delay is used for forward shifting. Time series allows upstream satellite measurements to be synchronized with their expected performance on Earth.
[0048] In the original data, the one marked as 99999.9 Special data points, such as values, represent invalid measurements and cannot be directly used for calculations. To address this issue, an expectation-maximization (EM) imputation method is applied within a moving time window to replace missing or corrupted entries before averaging. The EM algorithm optimizes the expected likelihood of the observed data, assuming that the observed data follows a multivariate Gaussian distribution. In each iteration, the algorithm alternates between estimating missing values based on the current parameters and updating the parameters based on the complete data.
[0049] These operations ensure that all time-series inputs are not only represented on a shared time axis but also sampled at consistent one-minute intervals, which is crucial for the integrity of sequential inputs in the model. Propagation delay correction ensures that the ratio is temporally aligned with its expected effect on soft X-rays, thereby improving the physical realism of the time series used for model training.
[0050] The second issue addressed was the presence of outliers in the input variables. Outliers in satellite-based datasets can originate from instrument-related interference and environmental disturbances during satellite operation.
[0051] Instrument limitations encompass a variety of technical and systemic issues. Satellite instruments are subject to aging, which can lead to a gradual decline in sensor accuracy and responsiveness. Software glitches can also introduce errors during data acquisition, processing, or transmission. Furthermore, inaccurate analog-to-digital conversion, noise from high-frequency heater duty cycles, and interference from onboard subsystems are all considered potential sources of measurement distortion. Thermal instabilities, especially under rapid orbital temperature variations, can further impact instrument performance. In some cases, system model errors and limited spatial or temporal resolution in sensor design can amplify these effects, particularly when resolving transient or small-scale geophysical phenomena.
[0052] Meanwhile, disturbances originating from the near-Earth space environment are a significant source of observational anomalies. Enhanced solar and geomagnetic activity—such as coronal mass ejections, solar high-energy particle events, and geomagnetic storms—can increase atmospheric drag and alter plasma conditions in the ionosphere and magnetosphere. These phenomena can compromise the consistency and continuity of time-series data.
[0053] If not properly identified and removed, these anomalous entries can introduce misleading patterns into data-driven models. Their amplitudes often far exceed the typical range of physical signals. Therefore, their inclusion can distort model optimization, especially in gradient-based learning algorithms. To prevent the network from overfitting to these non-representative inputs, it is necessary to accurately identify and remove affected time points before training.
[0054] In this application, a rule-based approach is used to detect anomalies, which relies on prior knowledge specific to the dataset. For example, It contains an invalid entry with a value of 999.99, and As mentioned before, it occasionally displays 99999.9. Similarly, It was discovered in The value indicates that there is missing or corrupt measurement data.
[0055] Following anomaly detection, all invalid entries were converted to standard missing value labels. Subsequently, time steps containing any missing value among the six variables—input features and target photon count rate—were removed from the dataset. This rigorous selection criterion ensures that each training sample input to the model is a complete observation and physically meaningful across all dimensions. To promote numerical stability during model training, all continuous variables were normalized using a min-max scaling method. The range is defined. Since this application addresses the completion task, the minimum and maximum values used for normalization are calculated from the entire dataset, including both the training and test sets. This transformation is performed independently for each variable in the dataset. Normalization ensures that all features contribute similar magnitudes to the loss function and prevents scale-related bias during gradient-based optimization.
[0056] To transform the cleaned dataset into a format suitable for time series learning, this application employs a sliding window strategy to generate fixed-length training samples. Each window contains 100 consecutive time steps, and the window slides one time step at a time, thus generating densely overlapping segments. The target variable used for interpolation is... The input features include five physical variables: , , , and Although the model generates predictions for all 100 time steps, the loss is calculated using only the values within the artificially masked regions.
[0057] Since the dataset contains several non-overlapping observation intervals, referred to as subsets in this application, a sliding window is created independently within each subset. This ensures temporal continuity within each sample and avoids unintended connections across physically independent intervals. For each sample, a corresponding subset identifier is also recorded. This is to track its observation context in the raw data.
[0058] All five input features are processed through a synchronous window into a feature-partitioned sequence. This allows for flexible handling of heterogeneous inputs while preserving the temporal structure of each physical parameter.
[0059] After constructing the complete sliding window sample set, the dataset is split into training and test sets in an 80 / 20 ratio. The split is performed at the sample level using the `train_test_split` tool from the Scikit-learn library, and a fixed random seed is used to ensure reproducibility between different run-times. Importantly, this split is applied evenly to all input features and their corresponding label sequences, using a shared sample index. This ensures perfect alignment between variables and avoids potential mismatches between input and target pairs.
[0060] In addition to the multivariate feature sequences and label sequences, the associated subset identifiers are also split using the same index. These identifiers are crucial for tracking the temporal origin of each sample and are retained in subsequent analyses. Consistent splitting of input data, label data, and subset metadata ensures the preservation of temporal context and data source integrity during training and evaluation.
[0061] All generated NumPy arrays are explicitly converted to a normalized data type: numerical inputs and targets are converted to float32, and classification identifiers are converted to int32. This type normalization contributes to computational consistency and compatibility with deep learning frameworks.
[0062] Through a preprocessing process, the raw data is systematically transformed into high-quality, model-usable input. First, time alignment and synchronization ensure the temporal consistency of all physical variables. Second, feature cleaning and normalization remove contaminated values and standardize variable ranges, improving numerical stability during training. Third, a fixed-length sliding window is constructed to convert continuous multivariate time series into structured, context-aware training samples. Finally, the dataset is divided into well-aligned training and test sets, and synchronized feature-label mapping is achieved. These steps collectively enable the model to receive consistent and physically meaningful input sequences, thereby facilitating robust learning of the temporal dependencies of soft X-ray photon emissions and supporting accurate interpolation under partial observation conditions.
[0063] The photon count rate interpolation model used in this application has a modular overall network architecture to reflect physical dependencies. It consists of two auxiliary subnetworks and one main network. The design is motivated by two key theoretical variables in equation (5). and It cannot be directly observed. Therefore, two available surrogate variables are used for approximation: respectively and Both of these relationships are non-linear and involve time delays, making them suitable as targets for specialized modeling components.
[0064] To this end, this application introduces two subnetworks to estimate from their respective proxy inputs. and The first subnetwork is... As input, the output prediction Sequence. The second subnetwork uses... The value is used as input, and the predicted value is output. The sequence. As shown in Figure 2, these two sub-networks have similar structures, including a Transformer encoder block (TEB) and a fully connected layer. The resulting... and The sequence was then integrated into the main network.
[0065] Figure 2 The following is used for estimation and The architecture consists of two sub-networks. Each sub-network takes 100 time-step feature variables as input, passes through a TEB, and is output via a fully connected layer. The output is a 100-step sequence of estimated latent physical quantities, which is then used to support physical information modeling in the main network.
[0066] The main network receives the complete set of input variables, including five physical characteristics. , , , and and estimated and Sequence. These input variables are concatenated along the feature dimension to form a physical input tensor. This tensor is further enriched by three auxiliary components: subset identifiers. A binary mask is used to distinguish missing regions, and a fixed sine position code is used. These four components are broadcast to a common embedding space and combined to form the final multivariate input tensor, which is fed into the main network for representation learning and imputation.
[0067] Figure 3 illustrates the structure of the main network, which employs a hierarchical structure combining dilated causal convolutional layers and TEBs. The convolutional layers capture local temporal patterns while preserving causal orientation, while the multi-head self-attention modules within each TEB facilitate long-range dependency modeling. Subsequently, contextualized sequence representations are fed into two fully connected layers to generate the model's output.
[0068] Figure 3 The diagram shows the architecture of the main network, which receives multiple feature tensors and the output tensors of the sub-networks. Then, it passes through stacked dilated causal convolutional layers and two TEBs, followed by two fully connected layers, and outputs the completed photon count rate.
[0069] The network's final output is a 20-step estimated photon count rate sequence and a physically guided loss value. The 20-step output is obtained by pruning the central portion from a 100-step latent sequence generated by the main network. This pruning method is consistent with the training and evaluation settings, allowing subsequent studies to focus more on simulating the missing central region. Although the model is primarily guided by this segment of information, the loss function does not strictly restrict learning to only the occluded steps. The output region includes not only the missing interval but also its adjacent context, both of which are flexibly supervised. Specifically, the numerically anchored main loss consists of two parts: a reconstruction term for the occluded segment itself, and a boundary consistency term constraining the left and right neighboring contexts.
[0070] Meanwhile, the physics-driven loss component closely approximates the physical meaning of the output of the theoretically constrained model. This loss component is derived from subnetwork inference. , And observed , , , and The theoretical value of the photon count rate is calculated according to formulas (5) and (6), and approximated using the trapezoidal integral method. The model output is then compared with the theoretical value. The comparison is performed within a fixed-length time window centered on the missing region, not just limited to the masked steps. This auxiliary loss acts as a global regularizer, encouraging data-driven reconstruction results to align with the physical laws of the specific domain. The overall training objective is a weighted combination of the numerical main loss and the auxiliary physical loss.
[0071] Figure 4 This presents an overall view of the model architecture. The two auxiliary subnetworks first infer... and Then it is concatenated with other features and represented along with the subset of data. The mask and position code are input into the main network. During training, the output of the main network enters the physical loss layer, where the background signal conversion between different energy bands is completed, and then the physical loss is calculated, thereby ensuring that the reconstruction result is numerically accurate and conforms to physical laws.
[0072] The entire architecture is encapsulated as a multi-input, multi-output structure. Each input branch corresponds to a specific source of physical or structural information, while the output includes a completed sequence and an auxiliary physical information loss. The encoding method and function of each input component will be explained in detail below.
[0073] The model receives a structured set of nine inputs. Five of these correspond to physical parameters obtained from observations of the solar wind and magnetosphere: , , , and To enhance structural modeling, three additional components are included: a scalar quantum dataset identifier. A binary mask vector is used to distinguish observation segments; a fixed sine code is used to mark artificially masked regions for supervision; and a fixed sine code is used to assign a unique position vector to each time step to capture relative and absolute temporal order. All input branches are broadcast or projected into a shared embedding dimension for consistent combination and integration.
[0074] Five physical variables and two sub-network outputs and The features are concatenated along the feature dimension and projected onto a shared embedding space. The resulting physical feature embeddings, along with the subset embeddings, mask embeddings, and positional encodings, are element-wise summed to form a composite input representation. This tensor is then used as the input to the main network.
[0075] After the input encoding process, the model begins to generate its output. This output consists of two branches: the first is an inferred sequence of photon count rates over a 100-step window, denoted as... The second is a scalar auxiliary loss used to enforce physical consistency between the model output and the theoretical expectations based on domain knowledge.
[0076] The main network first generates a complete sequence containing 100 estimates for each time step within the input window. However, this full-length output is not retained as the final result. Since all downstream training objectives, evaluation metrics, and visualization processes are limited to the middle region of the sequence, only the middle 20 time steps are ultimately extracted from the complete output. These 20 time steps, including the masked gaps and their immediate context, constitute the final output sequence used for supervision and evaluation.
[0077] The model's second output is a physics-based auxiliary loss derived from the theoretical photon emission model. This component acts as a supplementary constraint, guiding the model toward a physically plausible solution.
[0078] The first component of the main network is a set of dilated causal convolutional layers designed to process the input representation. Extracting local temporal patterns. Unlike standard convolutional layers that simultaneously include past and future context within a fixed kernel window, causal convolution strictly enforces temporal causality, ensuring that time steps... The output depends only on the time step The input value. This characteristic is particularly important in temporal modeling tasks that consider the reconstruction of missing fragments in context.
[0079] To enhance the receptive field without increasing the number of parameters or reducing resolution through pooling, this application employs dilated causal convolution, where the dilation rate... The input spacing is expanded exponentially in each layer. As shown in Figure 5, this application stacks three causal convolutional layers with dilation rates of [missing information]. =1, 2, 4. This design allows the network to aggregate multi-scale temporal information within a wider context window while retaining the fine-grained resolution required for accurate interpolation.
[0080] Figure 5 The structure of the dilated causal convolutional block for local feature extraction is shown. Three causal one-dimensional convolutional layers are applied sequentially with dilation strides of 1, 2, and 4, and the kernel size of each layer is 3.
[0081] Causal convolutions utilize only forward information at each time step, thus preventing the accidental leakage of target information from future locations. This directional constraint is particularly important when filling in missing data, as masked regions must be inferred without relying on their own true values. Furthermore, causal convolutions provide a clean and temporally coherent feature base for subsequent global modeling layers. Their strict forward property imposes a well-defined structure on the input representation, which facilitates effective bidirectional learning in subsequent Transformer-based attention modules.
[0082] To capture short-term dependencies, a kernel of size 3 is used in each layer of the causal convolution, and the ReLU activation function is applied to introduce non-linearity and improve training stability. Causal padding ensures strict temporal alignment, maintaining consistent input and output lengths. The output dimension remains fixed across layers to preserve feature consistency.
[0083] These three dilated causal convolutional layers serve as an efficient and temporally coherent mechanism for extracting local structures, smoothing fluctuations, and initializing sequence representations before entering the global attention layer.
[0084] After extracting local patterns through causal convolutions, the network captures long-range dependencies and bidirectional context through stacked TEBs. Each TEB is designed strictly according to the original architecture introduced in "Attention is All You Need," so its internal structure is not illustrated here. These modules employ a self-attention mechanism, enabling each position in the sequence to pay attention to all other positions—both preceding and following steps. This bidirectional receptive field allows the model to integrate contextual information from the entire input window, which is crucial for accurately estimating values in missing segments surrounded by observable data.
[0085] At the core of each Transformer block is a multi-head self-attention module. It executes multiple attention operations in parallel, with each attention head learning a different projection onto the input sequence. The number of attention heads is denoted as . This determines how many independent attention patterns the model can capture simultaneously. This enhances the network's ability to represent multiple temporal interactions at different time scales. Following the attention mechanism, each block also includes a feed-forward network (FFN), whose dimensions are determined by... Hyperparameter control. Larger... and Values typically increase the model's representational power, but also introduce a higher risk of overfitting and computational cost.
[0086] Each TEB further introduces residual connections, layer normalization, and dropout regularization. These design choices help stabilize training dynamics, facilitate smooth gradient flow, and reduce the likelihood of overfitting, especially in deep architectures.
[0087] By stacking Transformer layers together, the model is able to construct a rich and globally aware representation of the input sequence. This representation incorporates both forward and backward context, which helps to accurately and context-awarely reconstruct occluded regions.
[0088] To enhance physical consistency during interpolation, this application introduces an auxiliary physical loss. This physical loss encourages the completed sequence to follow theoretically expected behavior derived from space plasma physics.
[0089] The loss is calculated by comparing the completed result with the theoretical physical value.
[0090] The completion result includes the photon count rate interpolated by the model within a fixed 10-step window in the masked region, as well as the completed values at two time steps before and after the masked region. Notably, this result extends beyond the masked region's time steps to include surrounding predicted values. This allows the loss to indirectly influence the contextual completion effect, facilitating a smooth transition with the missing segment and maintaining consistency with physical constraints. To ensure comparability, the output sequence of the completion result is averaged over time steps before being compared with the theoretical physical values.
[0091] The calculation of the physical theoretical values is based on formulas (5) and (6). First, it needs to pass through the background transformation sub-network, the structure of which is as follows: Figure 6 As shown.
[0092] Figure 6 The architecture of the background transformation sub-network is shown. This module will... Estimated values converted to the 0.5–0.7 keV energy range This supports the calculation of physical information loss downstream.
[0093] Given that the observed data represents the background signal in the 2.5-5.0 keV energy range, the background signal in the target energy range of 0.5-0.7 keV needs to be calculated using formula (6). Each subset of the dataset is identified by its ID. Mapped to an embedding vector, and then passed through a fully connected layer to learn the correlation coefficients of the subsets. and Using large convolution kernels A one-dimensional convolution with a value of 31 is applied, approximating a low-pass filter, to extract the gradually varying trend, thereby extracting the invariant component in the background signal of the 2.5-5.0 keV energy range. It can then be used. Expression of time-varying soft proton signals in the 2.5-5.0 keV energy range background signal Combining this with formula (6), the background signal in the target energy range of 0.5-0.7 keV can be obtained. Then, for a 10-step time window that considers the same range as the completion result, the input values are used. , and the result after conversion and subnetwork prediction and The theoretical value is calculated according to formula (5). This involves integrating these inputs using a trapezoidal approximation integral to obtain a theoretical estimate of the total photon emission over a specified time interval. The resulting value reflects the expected cumulative behavior governed by domain knowledge.
[0094] The physics-driven loss is then defined as the mean squared error between these two scalars. This design does not require the predicted value at each individual time step to perfectly match its theoretical counterpart; instead, it forces global consistency between the model completion and the physics-based emission model throughout the time window by comparing the average value. Therefore, it acts as a soft constraint, connecting the main network and auxiliary sub-networks, encouraging physically plausible behavior in the gaps and their adjacent regions.
[0095] This physical loss offers several advantages. When the missing region is very short, traditional loss functions lack sufficient supervision to adequately constrain the reconstruction. By expanding the evaluation window, the physical information loss anchors the completed segment to a physically reasonable benchmark, preventing unrealistic or abrupt biases and improving the model's generalization ability. This loss is implemented as a differentiable custom layer and is optimized together with the main loss throughout the training process.
[0096] After the main network has two stacked Transformer layers, the sequence representation will pass through a pair of fully connected layers as the final transformation stage before generating the interpolated output.
[0097] The first fully connected layer uses the ReLU activation function to compress the embeddings at each time step into a lower-dimensional latent space, thereby achieving non-linear refinement of the globally contextualized features. This layer also applies L2 regularization to prevent overfitting by penalizing excessively large weights and promoting a smoother parameter distribution.
[0098] The second fully connected layer performs a linear projection at each time step, generating the initial output sequence of the model. The resulting output tensor has the following shape: The temporal resolution is consistent with the input. However, it's important to note that only the central region of the output, corresponding to segments artificially masked as missing, is subject to supervision and evaluation. Therefore, to improve efficiency, the final output is truncated to a central window. Specifically, the network extracts the central 20 time steps to form the final output sequence. This design ensures that all subsequent evaluations and optimizations are focused on the region of interest.
[0099] The soft X-ray photon count rate data used in this embodiment were collected by the European Space Agency's XMM-Newton satellite. The European Photon Imaging Camera System, mounted on the telescope's focal plane, consists of three imaging detectors. The MOS1 instrument, equipped with a metal-oxide-semiconductor array, was selected as the primary source of observational data. The MOS1 instrument records incident soft X-ray photons as electrical signals, which are then converted into digital count rates. These values reflect the temporal evolution of soft X-ray radiation and serve as core variables in the interpolation mission. To supplement the photon count rate data, this application also incorporates space environment parameters known to affect soft X-ray signals. These solar wind and geomagnetic index data are obtained from the Coordinated DataAnalysis Web maintained by NASA.
[0100] Specifically, the data includes five input variables: solar wind proton density. Solar wind speed Ion abundance ratio Geomagnetic index and a background photon count rate of 2.5-5.0 keV. .in and The values are taken from the OMNI_HRO_1MIN dataset. Taken from the ACE-AC_H2_SWI dataset, It is obtained from the OMNI2_H0_MRG1HR dataset. and a background photon count rate of 0.5-0.7 keV Then the MOS1 observation data from the corresponding time period is extracted.
[0101] Table 1. Sources of Feature and Label Data and Their Temporal Resolution
[0102] Among the input variables mentioned above, Its observational source is unique. It was measured by the ACE satellite, which resides near the Lagrange point L1—approximately 1.5 million kilometers upstream of Earth. Due to the limited speed of the solar wind, from... There is a measurable time lag between the recording of L1 and its associated effects in near-Earth observations. This lag must be corrected for to ensure temporal consistency with other physical variables.
[0103] The dataset used in the experiment spans from 08:21:00 on March 22, 2000 to 07:04:00 on September 4, 2008, containing a total of 61,366 raw samples at 1-minute resolution. These samples were selected from an observation list compiled by Carter based on the XMM-Newton scientific archive, comprising 103 different observation intervals, as shown in Table S1 in the appendix. Each interval represents a different subset of the dataset in the experimental setup. These observation intervals were chosen to ensure diversity in mission phase, season, and exposure time.
[0104] Following the preprocessing procedure of this application, after time alignment and outlier removal, 38,664 valid samples were ultimately retained for model training and evaluation. Each sample contains five feature values and a photon count rate label. These processed data laid the foundation for subsequent interpolation modeling experiments.
[0105] To simulate a real-world contamination scenario, the lengths and distributions of missing segments within the cleaned dataset were first analyzed. Each subset of the dataset was processed independently to identify consecutively missing sequences. The statistical results are summarized in Table 2.
[0106] Table 2. Statistics on the length of missing segments in the cleaned dataset
[0107] As shown in the table above, short-time missing gaps dominate the distribution, with over 75% of missing segments consisting of only a single time step. Based on these observations, this embodiment defines four representative imputation scenarios with missing lengths of 1, 2, 5, and 10. Single-step missing gaps, i.e., missing length 1, are the most common and therefore the primary configuration for model development and evaluation. All model training and evaluation provided in this embodiment are performed under this configuration, while the other three scenarios are used for performance comparisons under different missing lengths. In these cases, the loss weights are moderately adjusted to better address the increased imputation complexity caused by longer missing gaps.
[0108] To represent missing intervals in the sequence, this embodiment employs a mask-based strategy: in each 100-step input sequence, a set of consecutive positions is explicitly marked as unavailable. These time steps are not provided to the model and are considered as targets to be filled. As shown in Table 3, for each scenario, the masked region is fixed near the middle of the sequence.
[0109] Table 3. Mask configuration and loss weights for scenarios with different missing lengths
[0110] Each missing region uses a shape of... The binary mask is used for encoding, where a value of 1 indicates a time step with observed valid data, and a value of 0 indicates artificially created missing data, used to simulate data that is contaminated by soft protons and cannot be used directly. This binary mask serves as part of the model input, acting as a structural indicator to show which parts of the sequence are unavailable and must be padded.
[0111] All models in this embodiment were implemented and trained using TensorFlow 2.4 and Python 3.8 in a Jupyter Notebook environment. All experiments were performed on an NVIDIA GeForce RTX 3090 GPU.
[0112] The model in this embodiment uses the Adam optimizer, with an initial learning rate set to... Training was performed for 100 epochs with a batch size of 64. To ensure stable convergence of the loss during training, a cosine annealing adjustment strategy with preheating was used to adaptively decay the learning rate over time. A gradient pruning threshold of 1.0 was applied to prevent instability. Furthermore, the learning rate adjustment optimizer ReduceLROnPlateau was enabled, which halved the learning rate when the validation loss stagnated within 10 epochs, with a lower bound of [missing value]. .
[0113] Table 4 lists the hyperparameters used in this embodiment and their values. These values were optimized after experimental performance verification and reflect the most efficient configuration for reconstructing the missing soft X-ray photon count rate under different missing interval scenarios.
[0114] Table 4. Model Architecture and Key Hyperparameter Settings for Training
[0115]
[0116] The model is trained using a composite loss function consisting of two parts: a main loss and an auxiliary physical loss. These two parts work together: the main loss ensures the model accurately reconstructs missing values, while the physical loss promotes consistency with theoretical formulas. Combined, this allows the model to achieve not only local accuracy but also to naturally and reasonably integrate the completed values into the temporal and physical context of the surrounding sequences.
[0117] Although the model is guided to the gap region via a binary mask input, the loss design goes beyond simply supervising the masked segment. Both the main loss and the physical loss provide additional learned signals beyond the gap itself, enabling lightweight supervision of the gap's temporal context. Specifically, the contextual term in the main loss fixes the predicted values at the gap boundary within a reasonable range by matching the model's completed output with known values at adjacent time steps. Simultaneously, the physical information loss associates the gap segment with its context through a physically consistent integral, imposing a global constraint within the completed window.
[0118] The main loss consists of two sub-components: a numerical loss that directly supervises the missing region, and a context loss that forces a smooth transition at the gap boundaries. The numerical loss is defined as the mean squared error between the imputed value and the true value within the missing interval, while the context loss is the mean squared error between the predicted value and the true value at each of the two time steps before and after the gap. These two sub-losses are respectively composed of scalar coefficients. and Weighted. In the most common case where the missing length is 1, set... =0.1, = 0.9, emphasizing context alignment while preserving numerical supervision within the gap. This strategy is particularly important for short gaps because when internal supervision information is scarce, fully optimizing the numerical values can lead to instability. In this case, the context loss acts as a boundary anchor, stabilizing the predicted values and guiding them towards a physically reasonable range. For longer missing segments, it can be increased... To balance the two objectives, as shown in Table 3. This flexibility allows the model to adapt to different completion challenges based on the missing length.
[0119] The physical loss is evaluated over a fixed 10-step window at the center of the sequence, independent of the actual length of the missing data. For longer missing intervals, this window may completely cover the gap; for shorter gaps, its supervision extends into the context. It measures the deviation between the average photon count rate of the model completion and the theoretical value obtained through trapezoidal integration of solar wind and ionospheric parameters. In this way, the physical loss acts as an extended constraint, enforcing physical consistency not only within the gap but also in its neighboring region. By connecting missing values with adjacent observations into a unified, theory-based structure, this loss term guides the completion results toward global rationalization. Even unmasked time steps are guided by the physical loss to generate outputs that are more physically consistent.
[0120] During training, the two loss terms are combined with fixed weights: the main loss has a weight of 1.0, and the auxiliary loss has a weight of 0.3, thus ensuring that physical constraints play a supporting role in guiding the model's behavior. This design ensures that the model's completion results are both numerically accurate and physically consistent, thereby improving the overall completion quality outside the masked region.
[0121] After inverse normalization, two evaluation metrics are used to assess the completion performance of the model proposed in this embodiment: MedAE and Huber Loss. These metrics collectively reflect the completion accuracy and robustness within the masked region.
[0122] MedAE is defined as the median of the absolute differences between the completed value and the corresponding true value. Compared to mean absolute error, median absolute error exhibits greater robustness to extreme outliers or transient peaks, which are common in soft X-ray observations and are typically caused by instrument malfunctions or sudden solar events. By relying on a central tendency, MedAE provides a more stable and reliable indicator of model performance, even under noisy conditions.
[0123] (9) Furthermore, this embodiment also employs Huber Loss as an auxiliary evaluation metric. Huber Loss, through the parameter δ, applies a quadratic penalty to small errors and a linear penalty to large errors, thus combining the characteristics of mean squared error and mean absolute error. This characteristic allows it to maintain strong tolerance even when large deviations occur occasionally, while remaining sensitive to overall accuracy. In the experiment, [the following is a partial translation of the remaining text, which is incomplete and requires further context]. δ Setting it to 1.0 follows the standard conventions widely adopted in regression tasks, ensuring a balance between robustness to outliers and sensitivity to subtle changes in imputed values.
[0124] (10) While Huber Loss is not a traditional evaluation metric in time series completion tasks, its sensitivity to a mixture of small and large errors is highly consistent with the physical characteristics of our task. In this embodiment, it is used as an auxiliary reference to supplement the insights provided by MedAE.
[0125] These two metrics, when combined, provide complementary perspectives on completion quality: MedAE reflects typical biases, while Huber Loss provides additional visibility into error distribution and robustness.
[0126] To verify the effectiveness of the proposed imputation model, this embodiment compares it with two baseline methods that represent common strategies for imputing missing values in time series data: linear interpolation and a neural network method based solely on LSTM. These baseline methods provide contrasting paradigms: linear interpolation is memory-based and belongs to the traditional statistical method; while LSTM is recurrent and data-driven.
[0127] To ensure a fair comparison, all methods were trained and evaluated under the same experimental conditions, including data preprocessing steps, loss functions (excluding physical loss), and training procedures. The only difference lay in the network architecture.
[0128] Through careful design, the parameter scale of the LSTM baseline model was matched to that of the main network of the model proposed in this embodiment. It consists of three unidirectional LSTM layers with 128, 128 and 64 hidden neurons, respectively, followed by a fully connected output layer.
[0129] The completion performance of the three methods described above under four missing length scenarios is recorded in Table 5. The results show that the proposed method consistently outperforms the two baseline methods in all cases, especially when the missing length is short. Missing segments of only one or two time steps account for the vast majority of the experimental data. In these two scenarios, the proposed method achieves significantly lower MedAE and Huber Loss. As the missing length increases, the completion difficulty increases, and the advantage of the proposed method over LSTM methods decreases somewhat. However, the method still maintains superior or comparable performance and continues to significantly outperform linear interpolation methods.
[0130] Table 5. Comparison of completion performance of different methods under different missing length scenarios.
[0131] The experimental results above demonstrate that the proposed method is particularly suitable for situations where short-term missing data are predominant in space weather data. Furthermore, the robustness of the proposed method is also verified in scenarios with longer missing lengths, showcasing its generalization ability in non-trivial environments.
[0132] To investigate the contributions of each component in the proposed model, ablation experiments were conducted in a scenario with a missing length of 1. The effects of two key designs on model completion performance were verified: the bidirectional modeling capability implemented via TEB, and the introduction of physical information loss.
[0133] In the first ablation experiment, the original bidirectional Transformer encoder block was replaced with a causal Transformer block, whose attention mechanism was masked to ensure that each time step only focuses on the past. This configuration eliminates the model's ability to utilize future context. In the second ablation experiment, the physical loss was removed, thereby revoking the physical consistency constraint while keeping the main loss unchanged.
[0134] The results of the ablation study are shown in Table 6. Variants of both models lead to significant performance degradation, confirming the value of the TEB and physics loss components. The model using unidirectional attention exhibits significantly worse completion performance, especially on MedAE, highlighting the importance of leveraging future information for completion. Similarly, removing the physics loss increases MedAE and Huber Loss, indicating that domain-prior-based auxiliary physics loss not only improves physical plausibility but also enhances numerical accuracy of the completion.
[0135] Table 6. Ablation experimental results in the scenario of missing length 1
[0136] In summary, ablation experiments demonstrate that the model proposed in this application not only benefits from a flexible and expressive architecture, but also from the explicit introduction of domain-specific constraints and bidirectional temporal context.
[0137] The experimental results above demonstrate the superiority of the proposed model across various missing length scenarios. The proposed method consistently outperforms linear interpolation and baseline LSTM-based neural network methods on MedAE and Huber Loss metrics, especially when the missing interval is short. This advantage can be attributed to two key architectural innovations, which have been validated through ablation experiments.
[0138] First, this model introduces bidirectional contextual modeling through a Transformer encoder block. Unlike traditional causal models that rely solely on historical information, the network in this application simultaneously considers both past and future observations. This capability allows it to utilize richer contextual cues on both sides of the missing fragment, resulting in more comprehensive and accurate estimations. Ablation experiments highlight the importance of this design choice: performance significantly degrades when bidirectional attention is replaced with causal attention.
[0139] Secondly, a physical loss is introduced to incorporate prior domain knowledge into the training process. Through constraints, the completion results are made as consistent as possible with the theoretical values calculated from physical formulas, guiding the network towards solutions that satisfy physical meaning. This constraint acts as a regularization mechanism, avoiding unrealistic fluctuations that conform numerically but violate physical principles. Removing this constraint in ablation experiments also led to a significant decrease in completion accuracy, further emphasizing its contribution.
[0140] The model proposed in this application is primarily designed and tuned for scenarios with short missing intervals, particularly for missing lengths of 1, which is dominant in the distribution of missing lengths in the experimental dataset. This design choice helps the model perform well in the most common real-world scenarios. Although performance degrades for longer missing segments, the results show that the proposed method remains robust across a wide range of completion challenges and consistently conforms to physical laws.
[0141] The model proposed in this application has several key advantages that enable it to play a role in the task of completing soft X-ray photon count rate sequences.
[0142] First, it performs exceptionally well in the most common short-missing scenarios. For missing segments of length 1, the model significantly outperforms the LSTM baseline and linear interpolation across all evaluation metrics. As shown in Figure 7(a), the model generates highly accurate and context-sensitive completion results that closely match the true values. The model also performs well when handling missing segments of length 2.
[0143] Figures 7(a)–7(d) illustrate representative examples of the model's completion performance at different missing lengths—1, 2, 5, and 10 time steps. Each subplot shows the middle 20 time steps of the input sequence. The yellow shaded area represents the artificially masked segment requiring interpolation, the red curve represents the true photon count rate label value, the blue curve represents the completed output value of our method, the green curve represents the baseline result obtained through linear interpolation, and the green curve represents the baseline completion effect achieved through a neural network relying solely on LSTM. These visualizations highlight the model's ability to closely align with the true values in favorable contexts.
[0144] Secondly, the model demonstrates competitiveness even with longer missing intervals. Although the availability of contextual information decreases in these cases, the proposed method still generates smooth and physically consistent outputs. This robustness stems from its bidirectional sequence modeling capability and the introduced physical loss, which guides the completion results toward theoretically consistent solutions. The results shown in Figures 7(c) and 7(d) highlight the model's ability to maintain reasonable reconstruction when missing intervals are five and ten steps.
[0145] Furthermore, the model completion follows domain physics principles. Unlike purely data-driven approaches, this application's framework embeds an auxiliary physics loss derived from simplified soft X-ray emission equations. This constraint ensures that the output not only matches the numerical values but also conforms to the expected physical behavior. This is a crucial characteristic in space physics applications, where ensuring physical plausibility is just as important as completion accuracy.
[0146] This application proposes a deep learning-based framework for completing the soft X-ray photon count rate sequence missing due to soft proton contamination within a specific energy band. The method integrates multiple physical and contextual cues to achieve accurate and physically consistent completion. The model is designed as a multi-input, multi-output neural network architecture, consisting of several key modules: dilated causal convolutional layers for local pattern extraction, a Transformer encoder block for global context modeling, and an auxiliary physical loss that forces the model output to align with domain theory.
[0147] Through comprehensive experiments across scenarios with varying missing lengths, the proposed method consistently outperforms baseline methods using only LSTM with linear interpolation and parameter matching. The method performs particularly well in the most common short-interval missing scenarios. Ablation experiments further demonstrate the importance of bidirectional context modeling and physical loss in achieving robust completion. Overall, this application provides a practical and theoretically sound solution for recovering missing space physics data, with significant implications for downstream analyses such as solar flare detection, energy flux estimation, and Earth-space response modeling.
[0148] Example 2 This application also provides a physics-driven deep learning-completed soft X-ray photon counting rate system, implemented based on the above method, the system comprising: The preprocessing module is used to preprocess the measurement data; The photon count rate completion module is used to input the data output from the preprocessing module into the trained photon count rate interpolation model and output the completed photon count rate.
[0149] This application may also provide a computer device, including: at least one processor, memory, at least one network interface, and a user interface. The various components in this device are coupled together via a bus system. It is understood that the bus system is used to implement communication between these components. In addition to a data bus, the bus system also includes a power bus, a control bus, and a status signal bus.
[0150] The user interface can include a display, keyboard, or clicking device. Examples include a mouse, trackball, touchpad, or touchscreen.
[0151] It is understood that the memory in the embodiments disclosed in this application may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory may be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDRSDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchlink DRAM (SLDRAM), and Direct Rambus RAM (DRRAM). The memories described herein are intended to include, but are not limited to, these and any other suitable types of memory.
[0152] In some implementations, the memory stores elements such as executable modules or data structures, or subsets thereof, or extended sets thereof: operating systems and applications.
[0153] The operating system includes various system programs, such as the framework layer, core library layer, and driver layer, used to implement various basic business functions and handle hardware-based tasks. The application programs include various applications, such as media players and browsers, used to implement various application functions. Programs implementing the methods of the embodiments of this disclosure can be included in the application programs.
[0154] In the above embodiments, the processor can also invoke programs or instructions stored in memory, specifically programs or instructions stored in an application program, for the following purposes: Follow the steps described above.
[0155] The above methods can be applied to or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above methods can be completed by integrated logic circuits in the processor's hardware or by software instructions. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic diagrams disclosed above. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the disclosed methods can be directly implemented by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory; the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above methods.
[0156] It is understood that the embodiments described in this application can be implemented using hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described in this application, or combinations thereof.
[0157] For software implementation, the technology of this application can be implemented by executing the functional modules (e.g., procedures, functions, etc.) of this application. The software code can be stored in memory and executed by a processor. The memory can be implemented in the processor or outside the processor.
[0158] This application may also provide a non-volatile storage medium for storing a computer program. When the computer program is executed by a processor, it can implement the steps in the above method embodiments.
[0159] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application and are not intended to limit it. Although this application has been described in detail with reference to the embodiments, those skilled in the art should understand that modifications or equivalent substitutions to the technical solutions of this application do not depart from the spirit and scope of the technical solutions of this application, and should all be covered within the scope of the claims of this application.
Claims
1. A physics-driven deep learning-based method for completing soft X-ray photon count rates, comprising: The preprocessed measurement data is input into the trained photon count rate interpolation model, and the completed photon count rate is output. The photon count rate interpolation model includes a first auxiliary sub-network, a second auxiliary sub-network, and a main network; The first auxiliary sub-network is used to convert the input ion abundance ratio into the SWCX scaling factor; The second auxiliary subnetwork is used to convert the input geomagnetic index into neutral hydrogen density; Both the first and second auxiliary sub-networks include a Transformer encoder block and a fully connected layer; The main network consists of three one-dimensional convolutional layers, two Transformer encoder blocks, and two fully connected layers connected in sequence.
2. The method for completing soft X-ray photon count rate based on physics-driven deep learning according to claim 1, characterized in that, The three one-dimensional convolutional layers are all causal convolutional layers with dilation rates of 1, 2 and 4, respectively, and the kernel size of each layer is 3.
3. The method for completing soft X-ray photon count rate based on physics-driven deep learning according to claim 1, characterized in that, The data processing procedures for the two fully connected layers include: The first fully connected layer uses the ReLU activation function to compress the embedding at each time step into a low-dimensional latent space; The second fully connected layer performs a linear projection at each time step, generating the initial output sequence of the model.
4. The physics-driven deep learning-based method for completing soft X-ray photon count rates according to claim 1, characterized in that, The measurement data include: proton number density in the solar wind, solar wind velocity, ion abundance ratio, geomagnetic index, and photon count rate in the 2.5-5.0 keV energy band.
5. The physical-driven deep learning-based method for completing soft X-ray photon count rates according to claim 4, characterized in that, The calculation process of the photon count rate interpolation model includes: The ion abundance ratio is input into the first auxiliary sub-network, and the SWCX scaling factor is output. Input the geomagnetic index into the second auxiliary sub-network to output the neutral hydrogen density; The number density of protons in the solar wind, the velocity of the solar wind, the ion abundance ratio, the geomagnetic index, the photon count rate of the 2.5-5.0 keV energy band, the SWCX scaling factor, and the neutral hydrogen density are concatenated along the feature dimension, and then broadcast together with the subset identifier, the binary mask used to distinguish missing regions, and the fixed sinusoidal position code to a common embedding space and combined to form the final multivariate input tensor. The multivariate input tensor is input into the main network, and the initial completion result, i.e., the complete photon count rate time series, is output. From the initial result sequence, the central part is extracted according to the set number of steps to obtain the final output result, which is the photon count rate time sequence after focusing on the key region.
6. The physics-driven deep learning-based method for completing soft X-ray photon count rates according to claim 1, characterized in that, The photon count rate interpolation model also includes an auxiliary physical loss module, used to calculate the physical theoretical value of the photon count rate; The processing procedure of the auxiliary physical loss module includes: Each subset identifier is mapped to an embedding vector, and the subset correlation coefficients are obtained through a fully connected layer. and ; Photon counting rate from 2.5–5.0 keV energy band using one-dimensional convolutional layers. Extract the invariant components from the background signal in the 2.5-5.0 keV energy range. ; use minus Time-varying soft proton signals were obtained from the background signal in the 2.5-5.0 keV energy range. ; The background signal in the target energy range of 0.5-0.7 keV is calculated using the following formula. : ; The soft X-ray photon count rate in the 0.5–0.7 keV energy range is calculated using the following formula. : ; in, SWCX scaling factor; It has a neutral hydrogen density; The density of protons in the solar wind; The speed of the solar wind; For thermal velocity; This is the actual line of sight of the satellite.
7. The physics-driven deep learning-based method for completing soft X-ray photon count rates according to claim 6, characterized in that, The training process of the photon count rate interpolation model includes: Data sets containing target sequences of proton number density, solar wind velocity, ion abundance ratio, geomagnetic index, photon count rate in the 2.5–5.0 keV energy band, and photon count rate in the 0.5–0.7 keV energy band are preprocessed; The preprocessed dataset is divided into training and test sets according to a set ratio; The ion abundance ratio in the training set is input into the first auxiliary sub-network, and the SWCX scaling factor is output. The geomagnetic index in the training set is input into the second auxiliary sub-network, which outputs the neutral hydrogen density. The number density of protons in the solar wind, the velocity of the solar wind, the ion abundance ratio, the geomagnetic index, the photon count rate of the 2.5-5.0 keV energy band, the SWCX scaling factor, and the neutral hydrogen density are concatenated along the feature dimension, and then broadcast together with the subset identifier, the binary mask used to distinguish missing regions, and the fixed sinusoidal position code to a common embedding space and combined to form the final multivariate input tensor. The multivariable input tensor is input into the main network, and the potential photon counting sequence of the first set number of steps is output. The potential photon counting sequence of the first set number of steps is input into the auxiliary physical loss module, and the completed photon counting rate and auxiliary physical loss are output. The numerical main loss of the main network and the auxiliary physical loss of the auxiliary physical loss module are weighted and combined to form the loss function of the photon count rate interpolation model. By using the backpropagation algorithm, the weights and biases of each layer are adjusted iteratively in reverse along the network layers based on the error between the imputed value calculated by the loss function and the true value, thereby minimizing the loss function and improving the imputed accuracy of the photon count rate interpolation model.
8. The physical-driven deep learning-based method for completing soft X-ray photon count rates according to claim 1, characterized in that, The preprocessing includes: Time alignment and synchronization of all data sources; Remove invalid values from the data and perform normalization. The data processed as described above is used to construct a time series dataset with a fixed window length by sliding a window with a set step size.
9. A physics-driven deep learning-complete soft X-ray photon counting rate system, implemented based on the method described in any one of claims 1-8, characterized in that, The system includes: The preprocessing module is used to preprocess the measurement data; and The photon count rate completion module is used to input the data output from the preprocessing module into the trained photon count rate interpolation model and output the completed photon count rate.