A robust spline decomposition and extreme value extraction method for solar wind speed time series
By employing a robust regression model based on cubic B-spline basis functions and Huber loss function using quantile nodes, combined with an extreme value feedback iteration mechanism, the problem of separating trends from extreme events in solar wind speed time series was solved. This enabled stable estimation of solar wind speed trends and accurate identification of extreme events, thereby improving the accuracy and automation of space weather analysis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TIANJIN UNIV
- Filing Date
- 2025-05-14
- Publication Date
- 2026-07-24
AI Technical Summary
Traditional time series decomposition methods are unstable in trend extraction under solar wind speed anomalies, making it difficult to effectively identify extreme events, which affects the accuracy of long-term solar wind trend analysis and the accuracy of space weather event prediction.
A robust regression model based on cubic B-spline basis function matrix and Huber loss function with quantile nodes is adopted. Combined with an extreme value feedback iteration mechanism, the abnormal perturbation points in the solar wind speed time series are dynamically weighted, and the extreme fluctuation term is explicitly extracted to achieve a clear distinction between trends and extreme events.
It significantly improves the accuracy of solar wind speed trend estimation and the precision of space weather analysis, automatically identifies extreme anomalies and disturbances, and expands the applicability and convenience of space weather data analysis.
Smart Images

Figure CN120493205B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of space weather, specifically relating to a method for trend analysis and extreme event identification of solar wind speed time series based on robust spline decomposition mechanism, which is particularly suitable for the separation and identification of anomalous disturbances and extreme fluctuation events in solar wind observation data. Background Technology
[0002] Solar wind speed is a crucial observational indicator in space weather research, and its variations directly impact the Earth's magnetosphere, ionosphere, and the safety of space missions. Solar wind speed time series typically exhibit significant non-stationary characteristics, including periodic variations, trend fluctuations, and dramatic transient disturbances and extreme events. These disturbances are often closely related to extreme space weather events such as solar flares, coronal mass ejections (CMEs), and interplanetary shocks.
[0003] Traditional time series decomposition methods (such as moving average, exponential smoothing, STL decomposition, etc.) are mostly suitable for extracting trends and periods under stable or weakly volatile conditions. However, in scenarios like solar wind speed, which are characterized by high noise and strong disturbances, extreme events or short-term abnormal fluctuations often lead to serious distortion in trend estimation, thereby reducing the accuracy of analysis on the long-term trend of solar wind changes.
[0004] Furthermore, existing methods generally lack explicit identification and modeling mechanisms for solar wind extreme events, making it difficult to effectively distinguish between background trends and short-term, severe disturbances in solar wind speed time series. Especially in subsequent space weather event prediction or catastrophic event analysis, the confusion between trends and anomalous disturbances can severely impact the accuracy of forecast models and the effectiveness of space disaster risk assessment.
[0005] Therefore, how to accurately separate trend information from extreme fluctuations in solar wind speed sequences and accurately identify and capture extreme space weather events has become an important technical problem that urgently needs to be solved in the field of space weather.
[0006] [References]
[0007] [1] Castillo, Y., Pais, MA, Fernandes, J., Ribeiro, P., Morozova, AL, & Pinheiro, FJ (2021). Relating 27-day averages of solar, interplanetary medium parameters, and geomagnetic activity proxies in solar cycle 24. Solar Physics, 296(7),115.
[0008] [2]Temmer, M., Reiss, MA, Nikolic, L., Hofmeister, SJ, & Veronig, AM (2017). Preconditioning of interplanetary space due to transient CMEdisturbances. The Astrophysical Journal, 835(2),141.
[0009] [3]Hudson,MK,Elkington,SR,Li,Z.,Patel,M.,Pham,K.,Sorathia,K.,...&Leali,A.(2021).MHD-test particle simulations of moderate CME and CIR-driven geomagnetic storms at solar minimum.Space Weather, 19(12),e2021SW002882. Summary of the Invention
[0010] To address the problems of unstable trend extraction and difficulty in effectively identifying extreme events under the influence of solar wind speed anomalies in traditional time series decomposition methods, this invention proposes a robust spline decomposition and extreme value extraction method for solar wind speed time series. First, an initial trend fitting framework is constructed using a cubic B-spline basis function matrix based on quantile nodes. A robust regression model is built by introducing the Huber loss function, and dynamic weight reduction of anomaly points is achieved through multiple rounds of iterative optimization, significantly improving the stability of the solar wind speed trend term estimation. Second, in each iteration, an extreme event identification threshold is dynamically set based on the median absolute difference (MAD) of the residuals, accurately capturing and extracting the severe perturbation information corresponding to space extreme weather events in the solar wind speed time series. The final decomposition result can effectively distinguish between the long-term trend and short-term extreme fluctuations of the solar wind speed series, significantly improving the accuracy of space weather analysis and extreme event prediction. This invention does not rely on manual annotation and has advantages such as automation, strong robustness, and wide applicability, making it particularly suitable for critical mission scenarios such as space weather monitoring and forecasting, and space disaster risk assessment.
[0011] To address the aforementioned technical problems, this invention proposes a robust spline decomposition and extreme value extraction method for solar wind speed time series, comprising the following steps:
[0012] Step 1) Solar wind speed time series input: Obtain the solar wind speed time series y, which is a sequence of observations at equal time intervals; generate the corresponding time index t = [0, 1, 2, ..., N-1] for the input data length. T ;
[0013] Step 2) Constructing spline basis functions: Based on the predetermined number of spline nodes n, n node positions are selected from the time index t generated in step 1) using the quantile position method to construct cubic B-spline basis functions, thus forming spline basis functions;
[0014] Step 3) Initialize the weight vector: Construct a weight vector w of the same length as the solar wind speed time series y, and initialize all points in the weight vector w to 1;
[0015] Step 4) Robust spline fitting and extreme value feedback iteration: Set the number of iterations to N, and iterate through the sub-processes 4-1) to 4-3) below, including:
[0016] 4-1) Fitting the smooth trend term: Based on the current weight vector w, a robust regression model with the Huber loss function is used to perform a weighted fitting of the spline basis function matrix formed in step 2) with the solar wind speed time series y to obtain the smooth trend term ^.
[0017] 4-2) Calculate the extreme fluctuation term: Calculate the extreme fluctuation term based on the solar wind speed time series y and the smoothing trend term, and find the absolute median difference of the extreme fluctuation term; set the extreme value discrimination threshold T based on the absolute median difference to measure the degree of abnormality of the fluctuation.
[0018] 4-3) Update the weight vector w. Dynamically adjust the value of each point in the weight vector w according to the extreme fluctuation term and the threshold T, and calculate the updated weight vector. This completes one update of the weight vector.
[0019] Step 5) After completing N iterations, the final output includes a smoothed trend term, an extreme fluctuation term, and a weight vector w; where:
[0020] 5-1) The smoothed trend term in the output represents a combination of long-term structure and periodic patterns after removing extreme fluctuations from the solar wind speed time series, which helps to identify the basic patterns of solar activity.
[0021] 5-2) The extreme fluctuation term in the output retains the extreme value component centered at zero in the solar wind speed time series, reflecting the anomalous event characteristics of severe disturbances in solar wind speed.
[0022] 5-3) The output weight vector w represents a quantitative index of the degree of anomaly in the solar wind speed time series at each time point. It is used to determine the degree of extreme anomaly of the data points and to assist in subsequent space weather event classification, anomaly location and further anomaly diagnosis and analysis.
[0023] Furthermore, in the robust spline decomposition and extreme value extraction method for solar wind speed time series described in this invention, wherein:
[0024] The specific content of step 2) in constructing the spline basis functions includes:
[0025] 2-1) The principle for setting the number of spline nodes n is: the number of spline nodes n is set based on the complexity of the solar wind speed time series, and industry professionals know how to grasp this principle.
[0026] 2-2) After sorting the data values of the solar wind speed time series y using the quantile position method, the data is divided into n+1 intervals with equal probability according to the quantile points, and nodes are selected at the boundaries of these intervals;
[0027] 2-3) Construct cubic B-spline basis functions, where the set of cubic B-spline basis functions is represented as:
[0028] X=[Φ1(t),Φ2(t),…,Φ m (t)] (1)
[0029] In equation (1), Φ m (t) represents the m-th cubic B-spline basis function, where m represents the dimension of the spline basis function family;
[0030] 2-4) Apply the constructed cubic B-spline basis functions to each time point in the time index t to obtain the corresponding spline basis function matrix X. spline Its dimension is N×m, and each row corresponds to a spline basis function vector at a specific time point.
[0031] Step 3) initializes the weight vector, including: for the time series data y = [y1, y2, ..., y] obtained in step 1), n ] T Construct a weight vector w = [w1, w2, ..., w] with the same length as the time series data y. n ] T In the initial state, the weights of all time points are uniformly set to 1, i.e., w i =1.
[0032] Step 4-1) fitting the smoothing trend term includes: based on the spline basis function matrix X constructed in step 2-4). splineUsing the time series data y from step 1) and the current weight vector w defined in step 3), a robust regression model based on the Huber loss function is constructed; the estimated regression coefficients in the robust regression model are:
[0033]
[0034] In equation (2), w i X represents the weights at the corresponding time points in the weight vector. spline (i,:) represents matrix X spline The i-th row vector, where β represents the regression coefficient; L δ (·) represents the Huber loss function used in the robust regression model, which has the following mathematical expression:
[0035]
[0036] In equation (3), the residual term δ represents the threshold parameter of the Huber function, which is used to adjust the sensitivity of the model to extreme values;
[0037] Based on regression coefficient estimates Calculate the smoothing trend term
[0038]
[0039] Step 4-2) calculates the extreme fluctuation term, including:
[0040] 4-2-1) Subtract the smoothing trend term from the time series data y in step 1) Obtain the extreme fluctuation term The specific expression is:
[0041]
[0042] 4-2-2) Calculate the extreme fluctuation term obtained in step 4-2-1). Absolute median difference:
[0043]
[0044] In equation (6), median(·) represents the median;
[0045] 4-2-3) Based on the absolute median difference obtained in step 4-2-2), an extreme value identification threshold T is set to identify data points with abnormal fluctuations. The expression is as follows:
[0046] T = 1.5 × MAD (7).
[0047] Step 4-3) updates the weight vector w, including: based on the extreme fluctuation term obtained in step 4-2). The updated weight vector is calculated using the extreme value discrimination threshold T:
[0048]
[0049] In equation (8), This represents the updated weight value of the i-th data point. w represents the weight value of the i-th data point before the update. min The preset lower limit constant for weights is 0.2, and T is the extreme value discrimination threshold obtained in step 4-2). Let ε be the absolute value of the extreme fluctuation term corresponding to the i-th data point, and let ε be a positive constant with a value of 1 × 10⁻⁶. -8 This is used to ensure the stability of numerical calculations.
[0050] Compared with the prior art, the beneficial effects of the present invention are:
[0051] First, by introducing the Huber loss function and the extreme value feedback iteration mechanism, this invention achieves dynamic weight reduction of anomalous disturbance points in the solar wind speed time series, which significantly enhances the robustness of the smoothed trend term to anomalous events in the solar wind speed series, overcomes the problem of severe distortion of the trend term under anomalous disturbances in traditional methods, and improves the accuracy and reliability of trend estimation.
[0052] Secondly, by explicitly constructing extreme fluctuation residual terms, this invention can not only accurately extract the long-term smooth trend in the solar wind speed time series, but also retain the short-term violent disturbance information in the solar wind speed, which facilitates independent analysis and modeling of space extreme weather events in the solar wind speed, effectively improving the interpretability of the solar wind speed time series and the accuracy of space weather forecasts.
[0053] Finally, this invention employs an automated extreme value feedback iterative weight update strategy, which eliminates the need for manual pre-labeling of anomalous events and the need for pre-defined fixed anomaly discrimination rules. It can adaptively identify extreme anomalous disturbances in solar wind speed during the data decomposition process, realizing the intelligence and automation of the method and greatly expanding the applicability and convenience of space weather data analysis. Attached Figure Description
[0054] Figure 1 This is a flowchart of the method described in this invention;
[0055] Figure 2 This is the decomposition result of the solar wind speed sequence in the material studied in this invention. Detailed Implementation
[0056] This invention proposes a robust spline decomposition and extreme value extraction method for solar wind speed time series, aiming to achieve accurate separation of trend structure and extreme value perturbations in solar wind speed time series, thereby improving the accuracy of space weather prediction and the ability to identify extreme events. Specifically, this method constructs a cubic B-spline basis function matrix based on quantile nodes, uses the Huber loss function for robust regression modeling, and combines an extreme value feedback iterative update mechanism to effectively extract smooth trend structures and clearly identified extreme fluctuation terms from the original solar wind speed series, significantly enhancing the method's robustness and adaptability to solar wind anomalies.
[0057] For ease of explanation, the representation of solar wind speed time series data in this invention is defined as follows:
[0058] Let the observed solar wind speed time series be y = [y1, y2, ..., y]. n ] T Here, n represents the length of the time series, and y represents the observation at the i-th time step. The corresponding time index is t = [0, 1, 2, ..., N-1]. T The goal of the decomposition is to decompose the original sequence y into a smooth trend term ^ and an extreme volatility term ^, i.e.
[0059] The present invention will be further described below with reference to the accompanying drawings and specific embodiments, but the embodiments described do not limit the scope of application of the present invention in any way.
[0060] Taking the actual observed solar wind speed time series as an example, the specific steps of the solar wind speed trend analysis and extreme event identification method proposed in this invention are explained, such as... Figure 1 As shown, the method includes the following steps:
[0061] Step 1) Input solar wind speed time series. This step aims to complete the preparation and basic processing of the raw solar wind speed series data.
[0062] Step 1) Solar wind speed time series input: This step aims to prepare and perform basic processing of the raw solar wind speed time series data. First, obtain the solar wind speed time series observation data y, and represent it as a one-dimensional vector y = [y1, y2, ..., y]. n ] T Where n is the total number of observed data points, and the time intervals must be uniform to ensure the accuracy of subsequent fitting. A corresponding time index t = [0, 1, 2, ..., N-1] is generated based on the length of the input data. T The time index will be used as the input of the independent variable for constructing the spline basis function, ensuring that the fitting process strictly follows the time order.
[0063] After completing the above steps, the solar wind speed sequence data and its time index are ready, providing the basic input for subsequent spline basis function construction and robust regression fitting.
[0064] Step 2) Constructing spline basis functions: After completing the input of solar wind speed time series data, this step aims to construct spline basis functions for robust fitting by setting the number of nodes.
[0065] First, the number of nodes required in spline decomposition is set to n. This parameter determines the local fitting ability and smoothness of the B-spline basis function. The more nodes there are, the higher the fitting flexibility, but the smoothness decreases. Conversely, the fitting smoothness is improved, but it may lead to underfitting. Therefore, the principle for setting the number of spline nodes n is based on a comprehensive consideration of the complexity of time series data, the fitting accuracy requirements, and computational efficiency. Industry professionals know how to master this principle.
[0066] Then, based on the set number of spline nodes n, and the constructed time index sequence t = [0, 1, 2, ..., N-1], T A quantile partitioning strategy is used to determine node positions. Specifically, the time series data index is uniformly divided into n+1 intervals based on the quantile positions of the data length, and nodes are selected at the boundaries of these intervals. Compared to the traditional equidistant node partitioning strategy, this method can more effectively adapt to the non-uniform density distribution of the time series, significantly improving the adaptability and expressiveness of spline fitting. Unlike traditional equidistant partitioning, the quantile selection method can better adapt to the non-uniform distribution of observation points in the time series, thereby improving the ability of the subsequent fitted curve to capture local features.
[0067] Based on the selected set of nodes, construct spline basis functions using cubic B-splines as basis functions, where the set of cubic B-spline basis functions is represented as follows:
[0068] X=[Φ1(t),Φ2(t),…,Φ m (t)]
[0069] In the formula, Φ m (t) represents the m-th cubic B-spline basis function, where m represents the dimension of the spline basis function family;
[0070] Then, the constructed cubic B-spline basis functions are applied to each time point in time index t, that is, each time point in time index t is substituted into all B-spline basis functions to calculate the corresponding response value, and finally the spline basis function X is formed. spline Its dimension is N×m, and each row corresponds to a spline basis function vector at a specific time point. Spline basis function X spline This will serve as the input feature matrix for the subsequent robust regression model, supporting the fitting process of the trend term.
[0071] Step 3) Initialize the weight vector. To effectively handle outliers during robust regression, this step initializes the weight vector for each data point. Specifically, for time series data y = [y1, y2, ..., y...], ... n ] T Construct a weight vector w = [w1, w2, ..., w] with the same length as the solar wind speed time series data y. n ] T This is used to identify the degree of influence of each observation point on the fitting process. Initially, all weight elements are uniformly assigned a value of 1, i.e., w. i =1 indicates that in the initial stage, all data points are considered equally important and participate equally in the fitting calculation of the trend term, which is used for weight reduction of outliers in the subsequent robust fitting process. As the iteration process progresses, the weight vector w will be dynamically updated according to the magnitude of the residual, thereby gradually reducing the interference of outlier fluctuations on the trend fitting results and improving the robustness of the model to outliers.
[0072] Step 4) Robust Spline Fitting and Extreme Value Feedback Iteration: After completing the spline basis function construction and weight initialization, this step iteratively optimizes the robust fitting and extreme value feedback process by setting the number of iterations N, thereby improving its adaptability to abnormal fluctuations. Specifically, each iteration includes the following sub-steps:
[0073] 4-3) Update the weight vector w. Dynamically adjust the value of each point in the weight vector w according to the extreme fluctuation term and the threshold T, and calculate the updated weight vector. This completes one update of the weight vector.
[0074] First, based on the current weight vector w, a robust regression model is constructed using the Huber loss function, and the spline basis function X is... spline We perform a weighted fit with the solar wind speed time series y, solve for the optimal regression coefficients, and obtain the fitted smooth trend term. The Huber loss function combines the accuracy advantage of squared error with small residuals and the robustness of absolute error with large residuals, thus effectively reducing the impact of extreme values on the overall trend estimation.
[0075] The estimated regression coefficients in the robust regression model are:
[0076]
[0077] In the formula, w i X represents the weights at the corresponding time points in the weight vector. spline (i,:) represents matrix X spline The i-th row vector, where β represents the regression coefficient; L δ(·) represents the Huber loss function used in the robust regression model, which has the following mathematical expression:
[0078]
[0079] In the formula, the residual term δ represents the threshold parameter of the Huber function, which is used to adjust the sensitivity of the model to extreme values;
[0080] Based on regression coefficient estimates Calculate the smoothing trend term
[0081]
[0082] Then, the extreme fluctuation term is calculated based on the solar wind speed time series y and the smoothing trend term, and the absolute median difference of the extreme fluctuation term is obtained. An extreme value discrimination threshold T is set based on the absolute median difference to measure the degree of abnormality of the fluctuation, which is to subtract the smoothing trend term obtained from the current round of fitting from the original series y. Obtain the extreme fluctuation term
[0083] Based on this extreme volatility term, calculate the median and absolute median of the residuals.
[0084]
[0085] In the formula, median(·) represents the median; and the extreme value discrimination threshold T = 1.5 × MAD is set according to the following formula to identify data points with abnormal fluctuations. MAD is used to measure the robustness of the residual distribution and can effectively avoid the amplification effect of mean square error on extreme outliers. By using the extreme value discrimination threshold T, normal fluctuations and extreme fluctuations can be distinguished, and the degree of abnormality at each time point can be determined.
[0086] Finally, the weights of each sample point are dynamically updated based on the residual magnitude and the extreme value discrimination threshold T. Specifically, the update rule compares the current weight with the ratio of the residual normalization threshold, and uses a pruning operation to limit the weight update range.
[0087] Based on the obtained extreme fluctuation term The updated weight vector is calculated using the extreme value discrimination threshold T:
[0088]
[0089] In the formula, This represents the updated weight value of the i-th data point. w represents the weight value of the i-th data point before the update. minThe preset lower limit constant for the weights is set to 0.2 to avoid the weights being too small or ineffective. T is the obtained extreme value discrimination threshold. Let ε be the absolute value of the extreme fluctuation term corresponding to the i-th data point, and let ε be a positive constant with a value of 1 × 10⁻⁶. -8 This mechanism is used to ensure the stability of numerical calculations. Through the aforementioned extreme value feedback mechanism, the impact of extreme fluctuations can be gradually reduced in multiple iterations, making the fitting results more robust.
[0090] The above steps are repeated within the set number of iterations N until the iteration termination condition is met.
[0091] Step 5) Output the results. After completing the set number of iterations N, this step outputs the final results of the solar wind speed time series decomposition process.
[0092] First, output the smoothing trend term. This indicates the long-term trend structure and periodic change pattern of the solar wind speed time series after multiple rounds of robust iterative fitting and removal of extreme fluctuation interference. This helps to identify the basic pattern of solar activity and can effectively reveal the background evolution trend of solar wind speed on the solar activity cycle scale, providing a reliable basis for space weather background state analysis. Specifically, it reveals the long-term variation law of solar wind speed on the solar activity cycle scale, such as the steady evolution trend between slow wind and fast wind. This trend term is applicable to long-term trend analysis of solar wind speed, space meteorological background modeling, and space weather forecasting and space mission orbit dynamics simulation research. Studies have shown that parameters such as solar wind speed, density and magnetic field strength exhibit a 27-day periodic change throughout the solar cycle, especially during the maximum and decline phases of solar activity [1]. Therefore, the periodicity makes solar wind speed one of the most predictable solar wind plasma characteristics.
[0093] Secondly, output the extreme fluctuation term. Defined as the difference between the solar wind speed time series and the smoothed trend term. This extreme fluctuation term clearly retains the abnormal disturbance information in the solar wind speed series, clearly showing the abnormal event characteristics of the severe disturbance of solar wind speed in the short term, which usually corresponds to space weather extreme events such as solar flare, coronal mass ejection (CME) shock wave and interplanetary shock wave. This extreme fluctuation term is suitable for automatic identification of solar wind extreme events, rapid response analysis of space weather disaster events, research on the statistical law of extreme events, and risk assessment of space weather disasters; such as the signal characteristics of space extreme events such as solar flare, coronal mass ejection (CME) shock wave and interplanetary shock wave, which can be directly applied to the rapid identification, early warning and disaster risk analysis of space extreme weather. For example: the output extreme fluctuation term retains the extreme value component centered at zero in the solar wind speed time series, reflecting the abnormal event characteristics of the severe disturbance of solar wind speed. Studies have shown that CME events cause the solar wind speed to increase by 18% to 32% during the event, and still maintain an increase of 9% to 24% in the two days after the event, indicating that the disturbance of interplanetary space by CME may last for 3 to 6 days [2]. In contrast, CIR / HSS events typically originate from coronal holes near the solar equator, and their effects can last from several days to more than a week [3]. Therefore, by analyzing anomalous perturbations in solar wind speed, CME and CIR / HSS events can be effectively distinguished.
[0094] Finally, the weight vector w obtained from the final iteration is output, reflecting the degree of anomaly exhibited by the solar wind speed time series at each time point during the trend term fitting process. The smaller the weight value, the stronger the anomalous fluctuation at the corresponding time point. The weight vector can serve as a quantitative indicator of the degree of anomaly, assisting in the subsequent location analysis of space weather anomalies and the subsequent anomaly detection and diagnostic analysis.
[0095] Research Materials
[0096] This study focuses on solar wind speed time series and proposes a robust spline decomposition and extremum extraction method for this data. Addressing the issues of severe perturbations, non-stationary fluctuations, and local extrema in solar wind observation data, this method combines the Huber loss function with a spline fitting model to construct a robust time series decomposition framework. An iterative feedback mechanism dynamically adjusts the sample point weights, effectively suppressing the interference of extreme outliers on trend term extraction.
[0097] Figure 2The decomposition results of this method on a solar wind speed dataset collected by NASA are shown. Subplot (a) shows the solar wind speed time series and smoothing trend term; the red curve represents the solar wind speed time series in this dataset, and the black curve represents the extracted trend component, clearly removing spike interference caused by burst disturbances. Subplot (b) shows the extreme fluctuation component; the black curve represents the extracted extreme fluctuation component, and red dots mark the times identified as strong fluctuation points. These positions often correspond to abrupt solar wind events or instrument errors. Subplot (c) shows the weight vector obtained from the final iteration; the lower the weight, the more likely the time point is to be an anomalous fluctuation point. The overall results reflect the advantages of the model in terms of interpretability and stability.
[0098] This method can not only extract the main trend changes in solar wind speed sequences for prediction and long-term evolution analysis, but also explicitly identify and quantify extreme points, which has practical technical value in application scenarios such as anomaly detection, data cleaning, and space weather early warning.
[0099] Although the present invention has been described above in conjunction with the accompanying drawings, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many improvements and changes under the guidance of the present invention without departing from the spirit of the present invention, and these improvements and changes are all within the protection scope of the present invention.
Claims
1. A robust spline decomposition and extreme value extraction method for solar wind speed time series, characterized in that, The method includes the following steps: Step 1) Solar wind speed time series input: Obtain the solar wind speed time series y, which is a sequence of observations at equal time intervals; generate a corresponding time index based on the length of the input data. ; Step 2) Constructing the spline basis function: Based on the predetermined number of spline nodes n, select n node positions from the time index t generated in Step 1) using the quantile position method, and construct a cubic B-spline basis function to form the spline basis function. ; Step 3) Initialize the weight vector: Construct a weight vector w of the same length as the solar wind speed time series y, and initialize all points in the weight vector w to 1; Step 4) Robust spline fitting and extreme value feedback iteration: Set the number of iterations to N, and iterate through the sub-processes 4-1) to 4-3) below, including: Step 4-1) Fitting the smooth trend term: Based on the current weight vector w, use a robust regression model with the Huber loss function to fit the spline basis function matrix formed in Step 2). A weighted fit is performed with the solar wind speed time series y to obtain a smoothed trend term, including: Based on the spline basis function matrix constructed in step 2 The time series data y in step 1) and the current weight vector defined in step 3). A robust regression model based on the Huber loss function is constructed; the estimated regression coefficients in the robust regression model are: In equation (2), These are the weights in the weight vector corresponding to the time points. Representation matrix The row vectors Represents the regression coefficient; The Huber loss function, used in robust regression models, has the following mathematical expression: In equation (3), the residual term , This represents the threshold parameter of the Huber function, used to adjust the model's sensitivity to extreme values; Based on regression coefficient estimates Calculate the smoothing trend term : Step 4-2) Calculate the extreme fluctuation term: Calculate the extreme fluctuation term based on the solar wind speed time series y and the smoothing trend term, and find the absolute median difference of the extreme fluctuation term; set the extreme value discrimination threshold T based on the absolute median difference to measure the degree of abnormality of the fluctuation. Step 4-3) Update the weight vector w. Dynamically adjust the value of each point in the weight vector w according to the extreme fluctuation term and the threshold T, and calculate the updated weight vector. This completes one update of the weight vector. Step 5) After completing N iterations, the final output includes a smoothed trend term, an extreme fluctuation term, and a weight vector w; where: The smoothed trend term in the output represents a combination of long-term structure and periodic patterns after removing extreme fluctuations from the solar wind speed time series, which helps to identify the underlying patterns of solar activity. The extreme fluctuation term in the output retains the extreme value component centered at zero in the solar wind speed time series, reflecting the anomalous event characteristics of severe disturbances in solar wind speed; The output weight vector w represents a quantitative index of the degree of anomaly in the solar wind speed time series at each time point. It is used to determine the degree of extreme anomalies of the data points and to assist in subsequent space weather event classification, anomaly location, and further anomaly diagnosis and analysis.
2. The robust spline decomposition and extreme value extraction method for solar wind speed time series according to claim 1, characterized in that, Step 2) includes: The principle for setting the number of spline nodes n in step 2-1) is: to set the number of spline nodes n based on the complexity of the solar wind speed time series; Step 2-2) After sorting the data values of the solar wind speed time series y using the quantile position method, divide it into n+1 intervals with equal probability according to the quantile points, and select nodes at the boundaries of these intervals; Steps 2-3) Construct cubic B-spline basis functions, where the set of cubic B-spline basis functions is represented as: In equation (1), Let m represent the m-th cubic B-spline basis function, where m represents the dimension of the spline basis function family; Steps 2-4) Apply the constructed cubic B-spline basis functions to each time point in time index t to obtain the corresponding spline basis function matrix. Its dimension is N×m, and each row corresponds to a spline basis function vector at a specific time point.
3. The robust spline decomposition and extreme value extraction method for solar wind speed time series according to claim 1, characterized in that, Step 3) includes: Regarding the time series data obtained in step 1), Construct the time series data Weight vectors of consistent length In the initial state, the weight of all time points is uniformly set to 1, that is... .
4. The robust spline decomposition and extreme value extraction method for solar wind speed time series according to claim 1, characterized in that, Step 4-2) includes: Step 4-2-1) Take the time series data from step 1) Subtract the smoothing trend term from step 4-1). We obtain the extreme fluctuation term. The specific expression is: Step 4-2-2) Calculate the extreme fluctuation term obtained in step 4-2-1). Absolute median difference: In equation (6), This represents the median; Step 4-2-3) Based on the absolute median difference obtained in step 4-2-2), set the extreme value identification threshold. This is used to identify data points with abnormal fluctuations, and its expression is as follows:
5. The robust spline decomposition and extreme value extraction method for solar wind speed time series according to claim 1, characterized in that, Step 4-3) includes: Based on the extreme fluctuation term obtained in step 4-2) and extreme value discrimination threshold Calculate the updated weight vector: In equation (8), This represents the updated weight value of the i-th data point. This represents the weight value of the i-th data point before the update. This is a preset lower limit constant for the weight, with a value of 0.
2. The extreme value discrimination threshold obtained in step 4-2) is... Let be the absolute value of the extreme fluctuation term corresponding to the i-th data point. It is a positive constant, and its value is... This is used to ensure the stability of numerical calculations.