A blood donation amount prediction method based on a hierarchical prediction framework and neural network nonlinear learning correction
Patent Information
- Application Number
- CN202610693091.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-20
- Publication Date
- 2026-08-21
AI Technical Summary
[0004]这种独立预测模式存在一个未被重视的固有缺陷:预测结果在业务逻辑上不自洽
[0029] 1. This invention constructs a hierarchical time series structure to model blood donation data in layers according to the overall level, gender level, and blood type combination level. This enables the model to simultaneously depict macro-trend changes and behavioral differences between different blood donation groups, thereby improving the interpretability and stability of the prediction model.
Smart Images

Figure CN122619293A_ABST
Abstract
Description
Technical Field
[0001] This invention proposes a fusion prediction method for blood donation volume prediction, combining a hierarchical prediction framework with a nonlinear correction mechanism, based on time series modeling and deep learning techniques. The method first preprocesses historical blood donation data and constructs a hierarchical time series structure according to donor attributes and business management needs, thereby characterizing the changing patterns of different blood donation groups over time. Based on this, the SARIMAX model is used in the hierarchical prediction framework to statistically model each level of the time series, and external environmental variables are introduced to enhance the model's ability to characterize seasonal fluctuations and environmental changes, thus obtaining a basic prediction result for blood donation volume. Furthermore, based on the statistical prediction results, a nonlinear correction mechanism, namely a backpropagation neural network (BPNN), is introduced to establish a nonlinear mapping relationship between predicted and actual values, performing a learning-based correction on the basic prediction results to compensate for complex nonlinear variation characteristics that traditional time series models struggle to characterize. By combining statistical modeling with the nonlinear correction mechanism, this invention can simultaneously leverage the advantages of time series models in trend and cycle analysis, as well as the ability of neural networks to approximate nonlinear functions, thereby effectively improving the accuracy and stability of blood donation volume prediction. This invention belongs to the field of time series analysis and medical blood supply management technology, specifically involving hierarchical time series modeling, deep learning, and intelligent predictive decision-making. Specifically addressing the technical problem of "inconsistent business logic in prediction results" caused by ignoring the inherent hierarchical summation constraints in existing blood donation volume predictions, this invention proposes a hierarchical time series structure based on forced summation constraints, fundamentally ensuring the consistency between overall and component prediction results. Background Technology
[0002] With the rapid development of smart healthcare and information technology, the methods of collecting and managing medical data are undergoing profound changes. Particularly in the field of blood management and blood donation prediction, the construction of blood collection and supply information systems and the acceleration of data digitization have accumulated massive amounts of multi-dimensional blood donation data. However, the prediction of blood donation volume and the allocation of blood products often rely on the human experience of medical workers, making accurate and scientific predictions difficult. These data exhibit significant time-series characteristics, seasonal fluctuations, and complex external influencing factors, making accurate and efficient blood donation volume prediction a key issue in medical blood supply management. Accurate blood donation volume prediction is not only related to blood inventory allocation and supply-demand balance, but also directly affects emergency medical security and public health safety, thus possessing significant research significance and application value.
[0003] Traditional methods for predicting blood donation volume, whether statistical models such as ARIMA and SARIMA or various deep learning models that have emerged in recent years (such as LSTM and GRU), mostly treat blood donation volume as a single time series, or model and predict the series of men, women, and different blood types independently.
[0004] This independent forecasting model has an inherent, often overlooked, flaw: the forecast results are logically inconsistent with business logic. For example, in actual operations, blood management centers need to know both the "total blood donation volume next month" and the refined inventory allocation by gender and blood type. However, when forecasting "male blood donation volume," "female blood donation volume," and "total blood donation volume" separately, due to the independence of model construction and errors, a logical contradiction often arises: "male predicted value + female predicted value ≠ total blood donation volume predicted value." This inconsistency renders the forecast results unusable for refined management, severely weakening their decision support value. Fundamentally, this is because existing methods ignore the inherent hierarchical summation structure of blood donation data (total volume = male volume + female volume; male volume = male type A + male type B + ...).
[0005] To address the aforementioned problems, this invention proposes a blood donation volume prediction method combining a hierarchical prediction framework and a nonlinear correction mechanism. First, historical blood donation data is cleaned and processed to construct a monthly blood donation volume time series. Based on this, external variables (such as weather and holidays) are introduced, and the correlation between the monthly blood donation volume time series and these external variables is analyzed to identify key external variables. Subsequently, the blood donation volume data is hierarchically divided according to donor attributes, constructing a multi-level time series structure to reflect the periodic variation characteristics of different groups (such as gender and blood type). In the hierarchical prediction framework, the SARIMA model is used to statistically model each level of the time series, and a SARIMAX model is constructed by introducing exogenous variables to enhance the model's responsiveness to changes in the external environment. Based on the obtained statistical prediction results, a nonlinear learning module, namely a backpropagation neural network (BPNN), is further introduced to nonlinearly learn and correct the SARIMAX prediction results, thereby improving the model's ability to characterize complex time series features and its prediction accuracy.
[0006] This method combines the advantages of hierarchical time series modeling and deep learning nonlinear correction, effectively handling nonlinear features and external disturbances while capturing periodic patterns, thus achieving accurate prediction of blood donation volume. Compared with traditional models, this invention demonstrates significant advantages in prediction accuracy, stability, and robustness to sudden factors, showing promising industrialization prospects and social application value. Summary of the Invention
[0007] Compared with existing blood donation volume prediction methods, this invention proposes a blood donation volume prediction method based on hierarchical time series and neural network nonlinear correction. This method combines the advantages of time series models in characterizing trends and periodic changes with the ability of neural networks in modeling nonlinear relationships, thereby achieving high-precision prediction of blood donation volume and improving the ability of blood management departments to regulate future blood collection and supply balance.
[0008] Existing blood donation volume prediction technologies suffer from a neglected flaw: the prediction results are logically inconsistent. For example, when predicting male, female, and total blood donation volumes separately, a contradiction often arises: "Predicted male value + predicted female value ≠ predicted total blood donation volume." This inconsistency prevents blood management departments from making refined inventory allocations based on gender and blood type, severely weakening the decision-making value of the prediction information.
[0009] To address the specific technical problem of prediction consistency mentioned above, this invention proposes a blood donation volume prediction method based on a hierarchical time series structure with forced additive constraints. This method does not simply stratify the data; instead, it constructs a three-layer structure (population layer—gender layer—blood type layer) that strictly satisfies the additive identity and employs a bottom-up aggregation prediction strategy, fundamentally eliminating prediction inconsistencies mathematically. This is the essential difference between this invention and all existing single-sequence or independent multi-sequence prediction methods.
[0010] like Figure 1 As shown, the prediction method of this invention mainly includes four parts: First, data preparation and variable screening are carried out, historical blood donation data are cleaned and organized, and key external variables are screened through correlation analysis; Second, a hierarchical time series structure is constructed, and the hierarchical SARIMAX model is used to model and predict the trends of each level of the series; Then, a backpropagation neural network (BPNN) is introduced for nonlinear learning to correct the statistical prediction results; Finally, the prediction results are verified and compared through model performance evaluation to determine the prediction effect of the model.
[0011] This invention addresses the problem of insufficient prediction accuracy of existing blood donation volume prediction methods in complex environments by proposing a blood donation volume prediction method based on hierarchical time series modeling and nonlinear learning correction. This method constructs a multi-level blood donation volume time series structure to perform hierarchical modeling of the dynamic changes of different blood donation groups. Furthermore, it introduces a learning-based correction mechanism on the basis of statistical time series prediction, achieving a joint characterization of the trend and nonlinear fluctuation characteristics of blood donation volume changes. This improves the accuracy and stability of blood donation volume prediction, providing reliable data support for blood management departments to formulate blood collection and supply scheduling plans.
[0012] Unlike traditional single time series forecasting methods, this invention starts from the group differences and environmental influences of blood donation behavior, organically integrating time series modeling with nonlinear learning mechanisms. It constructs a unified blood donation volume forecasting framework through key steps such as hierarchical time series modeling, external variable-driven forecasting, and adaptive error correction. Figure 1 As shown, this invention mainly includes four parts: data preparation and variable analysis, hierarchical statistical prediction modeling, nonlinear learning correction, and model performance evaluation.
[0013] (1) Data preparation and variable analysis
[0014] First, the acquired raw blood collection records undergo data cleaning, which includes business rule filtering, time field standardization, outlier handling, and field format unification to ensure consistency in the logical structure and chronological order of the blood collection data. After data cleaning, the blood collection records are summarized along the time dimension based on the blood collection date to construct a monthly statistical time series of blood donation volume.
[0015] Based on this, multi-dimensional statistics of blood donation data were performed according to the donors' gender and blood type information. The overall blood donation volume sequence was further divided into multiple subsequences to construct a hierarchical time series system. This hierarchical structure includes an overall layer, a gender layer, and a gender and blood type combination layer. Through hierarchical modeling, it is possible to simultaneously characterize the overall blood donation trend and the behavioral differences between different blood donation groups.
[0016] Furthermore, considering that blood donation behavior is not only influenced by historical trends but also closely related to factors such as meteorological conditions, social environment, and blood collection resource allocation, this invention further conducts a systematic analysis of external environmental variables. Candidate variables include weather conditions, air quality, holidays, and the number of open blood collection points. The linear correlation between external variables and blood donation volume is assessed using the Pearson correlation coefficient, and the nonlinear dependency between external variables and blood donation volume is quantified using mutual information methods. By combining the strength of linear correlation and the degree of nonlinear dependency, key variables with significant influence are selected from the candidate variables and used as exogenous input features for subsequent predictive models, thereby enhancing the model's responsiveness to changes in the external environment.
[0017] (2) Hierarchical statistical forecasting modeling
[0018] After constructing the hierarchical time series dataset, this embodiment conducts statistical prediction modeling on the overall monthly blood donation volume series to characterize its combined influence of trends, seasonality, and external driving factors, and to provide basic input for subsequent nonlinear learning correction.
[0019] First, stationarity analysis is performed on the overall time series. Trend and seasonal components are eliminated through differencing, and the enhanced Dickey-Fuller (ADF) test is used to statistically verify stationarity. Once the series meets the stationarity condition, the lag dependency structure is further analyzed using the autocorrelation function (ACF) and partial autocorrelation function (PACF) to determine the candidate range of SARIMA model parameters. Based on this, different SARIMA(p,d,q)×(P,D,Q)_s model combinations are traversed within a predefined parameter search space. The Akaike Information Criterion (AIC) and Bayesian Information Criterion (BIC) are used as model evaluation indicators, and the parameter combination that achieves the optimal balance between fitting accuracy and model complexity is selected as the final model structure.
[0020] Based on the established SARIMA model structure, key external variables identified earlier are introduced to extend and construct the SARIMAX model, enabling joint modeling of the time series trend, seasonality, and external driving factors of blood donation volume. During model training, each exogenous variable is jointly estimated with the model parameters using maximum likelihood estimation.
[0021] In terms of hierarchical structure, each bottom-level sequence establishes a SARIMAX model and generates corresponding prediction results. Then, a bottom-up aggregation mechanism is used to aggregate the bottom-level prediction results level by level to obtain the intermediate and overall layer prediction sequences, thereby achieving unified modeling and integrated prediction of blood donation volume at different levels. Simultaneously, within the model training time window, the overall layer prediction sequence on historical data (in-sample prediction) is obtained to characterize the model's fitting effect on known data and serves as the input basis for the subsequent nonlinear error learning module, thus providing data support for subsequent prediction error correction.
[0022] (3) Nonlinear learning correction
[0023] After obtaining the basic prediction results from the statistical prediction model, the prediction results are further nonlinearly corrected. Because blood donation behavior is affected by various factors such as changes in group behavior, social activities, and complex weather conditions, the time series of blood donation volume often exhibits a certain degree of nonlinear fluctuation characteristics, which traditional statistical models cannot fully characterize.
[0024] To address this, the present invention introduces a backpropagation neural network model (BPNN) as a prediction correction module to adaptively optimize the statistical prediction results. Specifically, the overall layer fitting sequence obtained during the training phase of the statistical prediction model is used as the input feature, and the actual blood donation volume sequence is used as the learning target. A nonlinear mapping relationship between the statistical prediction results and the actual blood donation volume is established through a multi-layer feedforward neural network structure.
[0025] By continuously adjusting the network weight parameters through the backpropagation algorithm, the neural network can learn the systematic biases and uncaptured nonlinear patterns in the statistical model's prediction results. During the prediction phase, the overall layer prediction sequence generated by the statistical prediction model within the prediction window is input into the trained neural network model, thus obtaining the final blood donation volume prediction result after nonlinear correction. This correction mechanism can improve the model's ability to characterize complex fluctuations while maintaining the statistical model's trend interpretability.
[0026] (4) Model performance evaluation
[0027] To verify the effectiveness of the prediction method proposed in this invention, the performance of the model prediction results was evaluated. The mean absolute percentage error (MAPE) was used as the evaluation metric to measure the relative error level between the predicted and actual values. By comparing the error metrics of different prediction methods, the predictive performance and stability of the method of this invention in the blood donation volume prediction task can be objectively evaluated.
[0028] Compared with the prior art, the present invention has the following obvious advantages and beneficial effects:
[0029] 1. This invention constructs a hierarchical time series structure to model blood donation data in layers according to the overall level, gender level, and blood type combination level. This enables the model to simultaneously depict macro-trend changes and behavioral differences between different blood donation groups, thereby improving the interpretability and stability of the prediction model.
[0030] 2. This invention introduces external environmental variables into the statistical time series prediction model. By comprehensively analyzing factors such as weather conditions and blood collection resource allocation, the prediction model can reflect the response characteristics of blood donation behavior to changes in the external environment, thereby improving the reliability of the prediction results.
[0031] 3. This invention introduces a neural network learning mechanism to perform nonlinear correction on the output of the statistical prediction model, enabling the model to automatically identify and compensate for complex nonlinear fluctuation patterns that are difficult for statistical models to characterize, thereby further improving prediction accuracy.
[0032] 4. Experimental results show that the prediction method proposed in this invention has high prediction accuracy in blood donation volume prediction tasks.
[0033] 5. This invention mathematically guarantees the consistency of prediction results by constructing a hierarchical structure that strictly satisfies the addition constraint and adopting a bottom-up aggregation strategy.
[0034] Compared with the prior art, the innovation of this invention is mainly reflected in the following aspects:
[0035] (1) Construct a hierarchical time series structure that conforms to the characteristics of blood collection and supply business to achieve consistent prediction of multi-level data;
[0036] (2) Propose an external variable screening method based on the combination of statistical analysis and business mechanism to improve the rationality of model input features;
[0037] (3) Design a nonlinear correction mechanism that combines statistical models with neural networks to effectively improve prediction accuracy and model stability.
[0038] (4) Construct an overall prediction framework for blood collection and supply business scenarios to achieve a refined depiction of the trend of blood collection volume changes and early warning, and provide more timely decision support for blood inventory scheduling and emergency support. Attached Figure Description
[0039] Figure 1 Flowchart of a blood donation volume prediction method based on hierarchical time series and nonlinear learning correction;
[0040] Figure 2 Hierarchical structure diagram;
[0041] Figure 3 Pearson coefficient heatmap;
[0042] Figure 4 Mutual information score graph;
[0043] Figure 5 Time series differencing diagram;
[0044] Figure 6 ACF / PACF plot of the differencing sequence;
[0045] Figure 7 Schematic diagram of a nonlinear learning model;
[0046] Figure 8 Overall model architecture diagram; Detailed Implementation
[0047] like Figure 1As shown, the blood donation volume prediction method of this invention includes the following steps: First, acquire raw blood donation data and preprocess it, including data cleaning, outlier removal, and summarizing by time dimension to form a monthly blood donation volume time series. Then, based on business attributes and management needs, divide the series into multiple hierarchical structures such as the overall layer, intermediate layer, and bottom layer for subsequent hierarchical modeling and prediction. On this basis, screen the original external variables, analyze their correlation with blood donation volume, select external variables with significant relationships as auxiliary prediction factors, and construct a set of exogenous variables. Next, perform stationarity analysis and model parameter determination on each level of time series. Through differencing and combined with autocorrelation and partial autocorrelation function analysis, determine the parameter search range of the SARIMAX model and obtain a candidate model parameter set. Then, based on this parameter range, construct and train the SARIMAX model on the bottom-level time series, using the selected exogenous variables as external inputs to complete the prediction modeling of the bottom-level blood donation volume. After completing the bottom-level modeling, the prediction results of each bottom-level model are integrated step by step through a bottom-up hierarchical aggregation method to form the prediction sequences of the intermediate and overall layers, obtaining the preliminary prediction results of the overall layer. Subsequently, the fitted and predicted sequences of the overall layer are input into the BPNN layer, and the SARIMAX prediction results are corrected by nonlinear learning using a backpropagation neural network to compensate for the prediction bias caused by complex behavioral changes and nonlinear factors in the linear model. Finally, the monthly blood donation volume prediction result of the overall layer after nonlinear correction is output as the final prediction value, completing the entire prediction process.
[0048] like Figure 2 As shown, the hierarchical structure of blood donation volume is divided into three layers: the top layer is the overall layer, representing all blood donations; the middle layer is the gender layer, which divides the overall blood donation volume into male and female donations; and the bottom layer is the blood type layer, which further subdivides the blood donation volume into type A, type B, type AB, and type O based on the gender division. This structure, through a bottom-up approach, enables the hierarchical aggregation of blood donation volume from each blood type to the gender-based blood donation volume, and finally to the overall blood donation volume, thus providing a clear structural foundation for subsequent hierarchical modeling and prediction.
[0049] like Figure 3 As shown, the Pearson coefficient analysis was performed on the initial external variables, and a heat map was drawn. From the graph, it can be seen that the top three external variables with the highest correlation are the highest temperature, the lowest temperature, and the number of open blood collection points.
[0050] like Figure 4 As shown, the mutual information scores of the initial external variables were analyzed and a ranking plot was drawn. From the plot, it can be seen that the top three external variables with the highest correlation are the highest temperature, the lowest temperature, and the precipitation.
[0051] like Figure 5As shown, the original time series was differentially processed, and trend graphs of the original series, the first-order difference series, and the seasonal difference series were plotted to observe the fluctuation characteristics of the series after removing trend and seasonal components, thus providing a stationarity basis for subsequent modeling.
[0052] like Figure 6 As shown, autocorrelation function (ACF) and partial autocorrelation function (PACF) plots were drawn for the differencing time series to analyze the autocorrelation structure and lag relationship of the series, providing a basis for the construction of subsequent time series models.
[0053] like Figure 7 The diagram shows a backpropagation neural network model used for nonlinear correction. The neural network includes an input layer, a hidden layer, and an output layer. The input layer receives the predicted blood donation volume from the hierarchical SARIMAX model, and the output layer outputs the prediction result after nonlinear correction, thereby compensating for the prediction error of the linear model.
[0054] like Figure 8 The diagram shows the model architecture, illustrating the nonlinear correction process: First, the SARIMAX layer outputs preliminary predicted values and fitted values. Second, the fitted values and actual values are input into the BPNN layer to train the BPNN layer to learn the difference between the prediction and the actual values. Finally, the preliminary predicted values are input into the trained BPNN layer for nonlinear correction, and the final predicted value is output.
[0055] Based on the above description, the following is a specific implementation process, but the scope of protection of this patent is not limited to this implementation process.
[0056] Step 1: Data Preparation and Variable Selection
[0057] Currently, in the field of blood donation prediction, traditional time series models often rely solely on historical blood donation data for univariate modeling, neglecting the dynamic impact of external factors such as climate, holidays, and blood collection facilities. This makes it difficult to accurately reflect the periodicity and external driving patterns of blood donation behavior. While existing deep learning methods possess strong nonlinear modeling capabilities, they generally require large datasets and feature sizes. Directly applying them to monthly-level blood donation data can lead to overfitting and insufficient interpretability.
[0058] Step 1.1: Data Preparation
[0059] In this embodiment, the original blood collection records from 2015 to 2019 are first processed for data preparation. The data preparation includes, but is not limited to, the following operations: performing business rule verification on the blood collection records, uniformly formatting time fields such as blood collection date, unifying the format of donor identification and related coding fields, and removing or correcting abnormal or missing records to ensure the consistency of the original data in logic and format.
[0060] After data cleaning, valid blood collection records were screened based on operational criteria such as blood donation type, organization method, adequacy of donation, and initial and re-examination results. Subsequently, based on these valid blood collection records, monthly blood donation volumes were grouped and statistically analyzed according to gender and blood type to construct a hierarchical time series structure.
[0061] The hierarchical time series structure includes the following three levels:
[0062] (1) Level 0:
[0063] Monthly time series representing total blood donation volume This is used to reflect the changing trend of overall blood donation volume;
[0064] (2) Intermediate layer (Level 1):
[0065] Monthly blood donation volume time series categorized by gender, including male blood donation volume series. And the sequence of female blood donations This is used to characterize the impact of gender differences on blood donation behavior;
[0066] (3) Level 2:
[0067] Based on gender, further subdivisions are made according to blood type, with male blood donation volume subdivided by blood type. (Male, Type A) (Male, Type B) (Male AB type) and (Male blood type O); Female blood donation volume is further subdivided by blood type. (Female, Type A) (Female, Type B) (Female AB type) and (Women, blood type O) is used to characterize the changing characteristics of blood donation behavior in different groups.
[0068] The above hierarchical structure adopts a bottom-up additive constraint relationship, that is... and The summation relationship of the time series of the top layer, intermediate layer, and bottom layer is as follows:
[0069] That is, the bottom-level time series are aggregated by gender to form the intermediate-level time series, and the intermediate-level time series are further aggregated to form the top-level time series. The hierarchical relationship is as follows: Figure 2 As shown.
[0070] By employing the aforementioned hierarchical construction method, the underlying model can fully capture the dynamic characteristics of fine-grained populations while ensuring data consistency across all levels. Simultaneously, it provides interpretable, structured support for subsequent overall-level predictions. After hierarchical construction, the time series data at each level are saved as structured data files for subsequent time series modeling and prediction.
[0071] It should be noted that the above hierarchical structure is not arbitrarily divided, but rather stems from the natural organizational form of blood donation data in actual operations. In blood collection and supply management, blood donation data is typically expressed in both overall statistics and cluster statistics (such as by gender or blood type), with strict additive constraints between different levels. Modeling only the overall time series, while depicting the overall trend, fails to reflect the structural differences between different blood donation groups; conversely, modeling only fine-grained data easily overlooks the overall patterns of change.
[0072] This invention constructs the aforementioned hierarchical structure to address the technical challenge of "prediction consistency" in multi-granularity time series forecasting. Prediction consistency refers to the requirement that the sum of the prediction results of the lower-level fine-grained sequences must be strictly equal to the prediction results of the upper-level coarse-grained sequences. In traditional independent forecasting methods, when predicting "male blood donation volume," "female blood donation volume," and "total blood donation volume" separately, a logical contradiction often arises: "male + female ≠ total," preventing blood management departments from accurately allocating inventory based on the prediction results.
[0073] This invention constructs a hierarchical time series that strictly satisfies additive constraints and employs a bottom-up aggregation forecasting strategy, thereby mathematically guaranteeing the consistency of the forecast results. This design ensures that the forecast results of this invention are not only statistically optimal but also logically consistent, directly serving the practical business scenario of refined blood inventory management by gender and blood type. This is fundamentally different from the existing approach of treating each series independently and ignoring their inherent correlation, and is one of the key innovations of this invention.
[0074] Therefore, this invention constructs a hierarchical time series structure, unifying the overall layer, intermediate layers, and bottom layer into a single model. A bottom-up aggregation approach ensures consistency in prediction results across all levels, thus capturing the overall trend while also considering group differences. This hierarchical modeling method effectively solves the problem of inconsistent prediction results in traditional single-sequence prediction methods, improving the interpretability and stability of the model in practical business applications.
[0075] Step 1.2: Variable Selection
[0076] To improve the ability of time series models to depict changes in the external environment, this invention screens and models candidate external variables. After completing the hierarchical time series construction, the monthly blood donation volume time series of the total stratum is used as the model. As the subject of analysis, external factors that may affect changes in blood donation volume are screened.
[0077] Meanwhile, considering that blood donation behavior is not only influenced by historical trends but also closely related to factors such as the behavioral intentions of potential blood donors, external environmental constraints, and the capacity of blood collection and supply services, this invention first systematically constructs a set of candidate external variables from the following four dimensions based on the theory of blood donation behavior and common knowledge in the field of medical resource allocation:
[0078] (1) Meteorological environment dimension: Temperature and precipitation directly affect residents' willingness to go out and their physical comfort. Existing medical statistics show that extreme high or low temperatures and rainfall significantly reduce the amount of blood collected at street blood donation points. Therefore, this dimension selects the highest temperature, lowest temperature, precipitation, and weather conditions as candidate variables.
[0079] (2) Time-based social dimension: Holidays (such as statutory long holidays) can change people's travel and daily routines, usually leading to a surge in blood donations before the holiday and a sharp drop in blood donations after the holiday. Therefore, whether it is a holiday or not is selected as a candidate variable for this dimension.
[0080] (3) Service supply dimension: The number of open blood collection points is a key factor in determining the accessibility of blood donation services and directly constrains the maximum possible blood collection volume. Therefore, the number of open blood collection points is selected as a candidate variable in this dimension.
[0081] (4) Environmental quality dimension: Air quality (such as AQI index) may affect the outdoor activity decisions of sensitive groups, so air quality level was selected as a candidate variable.
[0082] Through the four dimensions mentioned above, a candidate set containing seven initial variables was constructed. This candidate set was not constructed arbitrarily, but rather covers the entire chain of influencing factors from "willingness to donate blood" to "ability to donate blood," demonstrating clear technical motivation and business interpretability. In the variable selection process, this invention did not directly use all candidate variables, but instead used quantitative analysis methods to screen and optimize them. On the one hand, the Pearson correlation coefficient was used to measure the linear correlation between external variables and blood donation volume; on the other hand, mutual information methods were introduced to characterize the non-linear dependencies between variables. By simultaneously considering both linear and non-linear correlations, the bias caused by relying on only a single indicator can be avoided, thus allowing for a more comprehensive assessment of the impact of variables on changes in blood donation volume.
[0083] In this embodiment, the initial candidate external variables include: the number of open blood collection points, air quality level, precipitation, whether it is a holiday, maximum temperature, minimum temperature, and weather conditions. Numerical variables (including the number of open blood collection points, maximum temperature, minimum temperature, and precipitation) are directly involved in subsequent correlation analysis; non-numerical or categorical variables (including air quality level, whether it is a holiday, and weather conditions) are converted to numerical form using label coding before analysis to ensure that different types of variables can be compared within a unified framework.
[0084] To quantitatively assess the correlation between external variables and blood donation volume, this embodiment employs both linear correlation analysis and nonlinear correlation analysis:
[0085] (1) The Pearson correlation coefficient is used to measure the degree of linear correlation between two continuous variables, and its value range is: When the correlation coefficient is close to When the correlation coefficient is 1, it indicates a strong positive linear correlation between the two variables; when it is close to 1, it indicates a strong positive linear correlation between the two variables. When the value is close to 0, it indicates a strong negative linear correlation between the two variables; when it is close to 0, it indicates a strong negative linear correlation between the two variables. When the expression is true, it indicates that there is essentially no significant linear correlation between the two variables. The formula for calculation is:
[0086]
[0087] in, Represents random variables In the Observations at each time point Indicates the first Blood donation volume at different time points The observed values, Representing variables The sample mean, Representing variables The sample mean, Indicates the sample length. Representing variables With variables The Pearson correlation coefficient is used to measure the degree of linear correlation between the two.
[0088] In this invention, by calculating the Pearson correlation coefficient between each candidate external variable and the blood donation volume time series, the degree of linear correlation between external environmental factors and blood donation volume can be assessed, thereby screening out key exogenous variables that have a significant impact on changes in blood donation volume and using them as external input features in subsequent time series prediction models.
[0089] (2) Mutual information measures the degree of statistical dependence between two variables and can characterize the nonlinear correlation between variables. When two variables are independent, their mutual information is zero; when there is a dependency between the variables, the mutual information value is greater than zero, and the larger the value, the stronger the dependency between the variables. Its calculation formula is:
[0090]
[0091] in, Represents random variables The value of , Represents random variables The value of , Represents random variables and The joint probability distribution, and Representing random variables respectively and Marginal probability distribution, Represents random variables and The mutual information value between them is used to measure the degree of non-linear dependence between them.
[0092] In this invention, by calculating the mutual information value between external variables and the blood donation volume time series, nonlinear correlation factors that are difficult to find by linear correlation analysis alone can be identified, thereby further screening key external variables that have an important impact on changes in blood donation volume and using them as exogenous input features of the time series prediction model.
[0093] This invention constructs a set of candidate external variables.
[0094]
[0095] This is used to reflect factors that may affect changes in blood donation volume, among which The highest temperature, The lowest temperature, To determine the number of open blood collection points, For precipitation, Air quality level, Whether it is a holiday or not, This refers to the weather conditions.
[0096] Before performing correlation calculations, candidate external variables undergo uniform data preprocessing. For numerical variables, such as maximum temperature, minimum temperature, precipitation, and number of blood collection points, a normalization method is used to standardize them to eliminate the influence of dimensional differences on the correlation measurement results. For non-numerical variables, such as air quality level, whether it is a holiday, and weather conditions, they are converted into discrete numerical form through numerical encoding to participate in the subsequent calculation of Pearson correlation coefficient and mutual information.
[0097] Then, for each variable The Pearson correlation coefficient between the blood donation volume and the time series of blood donations was calculated. Mutual information score Based on this, all candidate variables were ranked according to their relevance. The ranking results are shown as follows: Figure 3 and Figure 4 The variables with the highest Pearson correlation coefficients are, in order: (0.43) (0.35) and (0.13); the variables with the highest mutual information scores are, in order: (0.53) (0.42) and (0.15).
[0098] Based on the combined performance of Pearson correlation coefficient and mutual information score, this invention screens out key external factors, forming the final set of external variables:
[0099]
[0100] These variables include the highest and lowest temperatures, the number of open blood collection sites, and precipitation. These variables reflect both the direct impact of weather conditions on blood donation behavior and the combined effect of blood supply capacity and environmental factors, providing the most representative data support for the exogenous input of subsequent hierarchical models.
[0101] From a mechanistic perspective, the highest and lowest temperatures (both indicators) jointly characterize the asymmetric impact of temperature on blood donation willingness (both excessive heat and excessive cold have an inhibitory effect); precipitation serves as a direct measure of travel barriers; and the number of open blood collection points directly reflects service supply capacity. These four variables rank highly in both Pearson and mutual information analyses, indicating a significant linear trend correlation with blood donation volume, as well as a complex nonlinear threshold effect (e.g., the negative impact only becomes apparent when the temperature exceeds 30°C). This precisely confirms the necessity of employing a dual-indicator method of "linear screening + nonlinear screening" in this invention.
[0102] Furthermore, although the mutual information score of air quality levels is acceptable, its impact is often coupled with meteorological factors, and in actual forecasting, the accuracy of medium- and long-term air quality forecasts is low. Introducing this variable would reduce the engineering stability of the forecasting system, so it was discarded. Whether or not the monthly fluctuations of holidays are absorbed by seasonal components, and its impact pattern (pre-holiday peak, post-holiday trough) is relatively fixed and has already been characterized by the seasonal term in the SARIMAX model, therefore it was removed as a redundant variable. The information contained in weather conditions (such as sunny, cloudy, rainy, etc., categorical variables) has been partially replaced by precipitation and temperature, and its discrete encoding method is not conducive to model training, therefore it was not included in the final exogenous variable set.
[0103] It should be noted that this invention unifies the external variable screening process within the overall blood donation volume time series. The process proceeds by constructing a global perspective on the relationship between external variables and the target variable, thereby selecting a representative set of key external variables. Within the hierarchical time series framework, each bottom-level sequence (such as blood donation quantum sequences corresponding to different genders and blood types) exhibits statistical consistency and a subordinate relationship with the overall layer sequence. Therefore, this invention passes the overall layer selection results downwards, sharing the same set of external variables across all bottom-level prediction models. This avoids model inconsistencies caused by repeated variable selection and improves model stability and computational efficiency.
[0104] Step 2: Hierarchical Statistical Forecasting Modeling
[0105] After completing the hierarchical time series construction and external variable screening, this step, based on the time series modeling concept, performs trend modeling and prediction of blood donation volume series at each level, providing a basic input for subsequent nonlinear learning correction. By introducing external variables into each bottom-level series and constructing a SARIMAX model, the joint modeling of the trend term, seasonal fluctuations, and external influencing factors of blood donation volume is achieved.
[0106] Step 2.1: Determining the range of model order
[0107] In time series modeling, differencing is typically required to eliminate trend and seasonal fluctuations in the original series. Let the original time series be... Then the first-order difference can be expressed as the following mathematical expression:
[0108] It is used to eliminate long-term trends if the sequence has a period of The seasonal pattern can be observed through seasonal differences, expressed mathematically as follows:
[0109]
[0110] It is used to eliminate periodic effects. Among them, This represents the backshift operator, which shifts the sequence one time step back. Its mathematical expression is as follows:
[0111]
[0112] Furthermore, there is a mathematical expression:
[0113]
[0114] in, Denotes the lag order, taken in the first-order difference. Taking from seasonal differences When considering both trend and seasonal terms, a combined difference form can be constructed, with the following mathematical expression:
[0115]
[0116] in Indicates the first-order difference degree. Indicates the degree of seasonal difference. This represents the length of the seasonal cycle. This combined differencing operation effectively reduces the non-stationary characteristics of the sequence, making it meet the basic requirements for subsequent modeling.
[0117] Based on the above, the Seasonal Autoregressive Moving Average (SARIMA) model can be uniformly expressed mathematically as follows:
[0118]
[0119] The model introduces a hysteresis operator. Representing the dynamic relationship of time series as about The polynomial form of the non-seasonal autoregressive polynomial is as follows:
[0120]
[0121] This mathematical expression describes the relationship between the current value and the observations from the past p time points, where , ,…, These are autoregressive coefficients, used to represent the weight of the influence of historical observations on the current value under different lag orders, for example... This indicates the degree of influence of the previous value on the current value. Meanwhile, the mathematical expression for the non-seasonal moving average polynomial is as follows:
[0122]
[0123] This mathematical expression describes the relationship between the current value and past error terms, where , , …, These are the moving average coefficients, used to represent the weight of the historical error term on the current value under different lag orders, for example... This indicates the degree to which the error from the previous time step affects the current value. Meanwhile, the seasonality component is expressed by the following two mathematical expressions, where... Indicates a lagged operation spanning a seasonal cycle:
[0124]
[0125]
[0126] The above expression describes the influence of historical observations across seasonal cycles on the current value and the influence of the error term across seasonal cycles on the current value, where , , …, The seasonal autoregressive coefficient represents the dependency weight between the current observation and the observations over the past P seasonal cycles (each cycle being s in length), for example... This indicates the influence of the value from the previous seasonal cycle on the current value. , , …, The seasonal moving average coefficient represents the dependency weight between the current observation and the random error term over the past Q seasonal periods, for example... This indicates the impact of the error in the previous seasonal cycle on the current value.
[0127] This represents the autoregressive (AR) component, where Used to characterize the linear dependence between the current value and the observations at the past p time points under non-seasonal conditions. This section describes the impact of historical observations across seasonal cycles on the current value. It is denoted by φ (phi) and reflects the influence of historical values on the current value.
[0128] This represents the moving average (MA) component, where... Used to describe the relationship between current values and past error terms under non-seasonal conditions. This section describes the impact of the error term across seasonal cycles on the current value. It is denoted by θ (theta) and reflects the influence of historical errors on the current value.
[0129] in, and In this context, lowercase φ and θ represent the non-seasonal autoregressive term and the non-seasonal moving average term, respectively, used to characterize the short-term dependency between adjacent time points. and In this context, the uppercase Φ and Θ represent the seasonal autoregressive term and the seasonal moving average term, respectively, used to characterize long-term dependencies across seasonal cycles.
[0130] The aforementioned polynomials transform the dynamic dependence of the sequence into a finite-order lag structure, providing a basis for setting the parameters p, q, P, and Q. The SARIMA model thus uniformly handles trend, seasonal, and stochastic fluctuations, laying the theoretical foundation for predictive modeling.
[0131] In this embodiment, as Figure 5 As shown, comparative analyses were performed on the time series of monthly blood donations at the overall level using the original sequence, the first-order difference sequence, and the first-order difference combined with the seasonal difference sequence to determine the difference order parameter.
[0132] First, observing the original time series reveals a clear trend and periodic fluctuations, indicating that it is a non-stationary time series. Further first-order differencing shows that the differencing sequence fluctuates around the zero mean, effectively weakening the original trend and stabilizing the variance. This demonstrates that first-order differencing effectively eliminates the influence of long-term trends.
[0133] However, analysis of the sequence after first-order differencing still reveals a clear periodic fluctuation structure, characterized by recurring fluctuation patterns at fixed time intervals (period of 12), indicating that the seasonal component in the sequence has not been completely eliminated. Further seasonal differencing with a period of 12 results in a more random fluctuation, no longer exhibiting a clear periodic structure, and displaying approximately stationary random fluctuation characteristics overall.
[0134] Based on the above analysis, it can be determined that the time series basically meets the stationarity requirement after first-order differencing and one seasonal differencing (period of 12). Therefore, in the subsequent model construction, the differencing order will be set to: ordinary differencing order. Seasonal difference order Seasonal cycle .
[0135] Based on this, ACF and PACF plots are drawn for the stationary series after "first-order + seasonal difference" to identify the potential value ranges of non-seasonal and seasonal parameters, respectively.
[0136] In time series analysis, the autocorrelation function (ACF) and partial autocorrelation function (PACF) are important tools for analyzing the correlation structure of sequences, and are often used to determine the autoregressive order and moving average order of ARIMA or SARIMA models. Let the time series be... The sample length is The lag order is denoted as (Integer, range of values) ).
[0137] The autocorrelation function (ACF) describes the overall correlation of a series at different lags, that is, the strength of the linear relationship between the series and its lag values. The mathematical expression for the sample autocorrelation coefficient is:
[0138]
[0139] in The mean of the sequence. This refers to the lag order. By calculating different... corresponding An ACF plot can be drawn to observe how the overall correlation of the sequence decreases with lag order.
[0140] The partial autocorrelation function (PACF) represents the direct correlation between a series and a given lag value after removing the effects of intermediate lags. Partial autocorrelation coefficient. It can be obtained through an autoregressive model:
[0141]
[0142] in For the error term, That is, lag The partial autocorrelation coefficient. Plotting different... corresponding The PACF diagram can be obtained.
[0143] In summary, the ACF plot is used to observe how the overall correlation of a series decays with the lag order, thus helping to determine the order q of the non-seasonal moving average (MA) and the order Q of the seasonal MA; the PACF plot is used to observe the direct correlation after removing the effects of intermediate lags, thus helping to determine the order p of the non-seasonal autoregression (AR) and the order P of the seasonal AR.
[0144] In this embodiment, as Figure 6As shown in the ACF plot, a significant negative correlation peak exists at lag 1, suggesting that the non-seasonal moving average term q can be taken as 1. While there is no strong peak near the seasonal lag 12, there are slight oscillations, so the seasonal moving average term Q can be taken as 0–1. In the PACF plot, lag 1 shows a clear peak with a slight tail, suggesting that the non-seasonal autoregressive term p can be taken as 0–1. There is also weak significance at the seasonal lag 12, so the seasonal autoregressive term P can be taken as 0–1.
[0145] Considering that ACF / PACF is generally less sensitive to AR / MA structures of order higher than one, to avoid overlooking potential higher-order short-term and seasonal dependencies, we appropriately expand the parameter space during the actual model search. This avoids the limitations of relying solely on graphical judgments for parameter selection, improves the comprehensiveness and robustness of the model search, and sets the candidate order of SARIMAX as follows:
[0146]
[0147] This range covers both common low-order models and allows for the capture of higher-order short-term and seasonal dependencies, thus providing a more comprehensive set of candidates for subsequent optimal model selection based on AIC / BIC.
[0148] Step 2.2: Introduction of external variables
[0149] Based on the external variable screening results in step 1.2, the same exogenous input variables are introduced into each underlying SARIMAX model without further screening. These include maximum temperature, minimum temperature, number of open blood collection points, and precipitation. The SARIMAX model introduces exogenous variables based on the SARIMAX model, and its expression can be represented as:
[0150]
[0151] in, Indicates time Blood donation volume observations at specific times; Indicates time The vector of exogenous variables at time t; This is the vector of regression coefficients corresponding to the exogenous variables; This represents the random error term. Based on the autocorrelation analysis results, the model order parameter... and The range of values for is predetermined and used to constrain the subsequent model search space. The exogenous variable coefficient β and the random error term... As parameters to be estimated, they are jointly solved using the maximum likelihood estimation method during the model training phase, and their final values are determined during the parameter optimization process.
[0152] In this embodiment, the exogenous variable vector is represented as:
[0153]
[0154] in, Indicates the highest temperature. Indicates the lowest temperature. Indicates the number of open blood collection points. This indicates the amount of precipitation.
[0155] Considering that some external factors may not be available in real time during the prediction phase, this embodiment constructs one-phase lag features for the highest temperature, lowest temperature, and precipitation before using them as model inputs. The number of open blood collection points, however, is not subject to lag processing because it can be planned and obtained in advance. All the aforementioned exogenous variables are standardized before being input into the model and are introduced as exogenous input terms into each underlying SARIMAX model to enhance the model's responsiveness to changes in the external environment and improve the interpretability and accuracy of the prediction results.
[0156] Step 2.3: Low-level sequence SARIMAX modeling and parameter optimization
[0157] Based on the preliminary determination of the range of possible model orders in step 2.1 and the introduction of exogenous variables in step 2.2, SARIMAX modeling and parameter optimization are carried out on the underlying time series.
[0158] In this embodiment, the input to each underlying SARIMAX model consists of both endogenous time series and exogenous variables. The endogenous variables are the monthly blood donation volume sequences corresponding to different blood type and gender combinations, denoted as follows: The set of exogenous variables is denoted as... This model is used to characterize external factors influencing blood donation behavior, including weather conditions, the availability of blood collection sites, and social activities. The model is developed by simultaneously inputting... and This enables joint modeling of the target sequence.
[0159] During model training, the original monthly sequences are divided into a training set (60%), a validation set (20%), and a test set (20%). Model parameters are estimated only on the training set, and the performance of different order combinations is compared and selected on the validation set. The test set is used only for the final model generalization performance evaluation. To fully explore the possible values of the model order and reduce the human judgment error based on ACF and PACF, this embodiment adopts a traversal search method to train and evaluate candidate parameter combinations for each bottom-level sequence one by one.
[0160] The model selection criteria adopt the Akaike Information Criterion (AIC) and the Bayesian Information Criterion (BIC), which are defined as follows:
[0161]
[0162]
[0163] in, The number of model parameters, The likelihood function value, The sample length is denoted by . The smaller the AIC and BIC, the better the balance between fitting accuracy and complexity of the model. Therefore, by comprehensively comparing the performance of AIC and BIC on the validation set, the parameter combination with better performance in both metrics is selected as the optimal model structure for each underlying sequence.
[0164] In the actual model selection process, the values of AIC and BIC can be comprehensively evaluated. If a candidate model... By calculating the corresponding AIC and BIC, a comprehensive evaluation index can be constructed:
[0165]
[0166] in, Indicates the first The comprehensive evaluation value of each candidate model. When there is... When there are n candidate models, the optimal model structure can be expressed as:
[0167]
[0168] in, Indicates the candidate model number. This represents the total number of candidate models. This represents the index of the optimal model that minimizes the weighted sum of AIC and BIC among all candidate models.
[0169] In other words, among all candidate models, the model with the smallest sum of AIC and BIC is selected as the final optimal model structure. This method can simultaneously consider model fitting ability and model complexity, so that the selected model achieves a good balance between predictive performance and structural simplicity. Therefore, it is often used in the process of screening parameter combinations for time series models.
[0170] In addition, each underlying SARIMAX model needs to undergo a white noise test on the residual sequence using the Ljung-Box Q test. If the residual sequence does not show significant autocorrelation, it is considered that the model has fully extracted the effective structural information in the time series, and the remaining error term can be regarded as random perturbation, thereby verifying the effectiveness of the model fitting and the reliability of the prediction.
[0171] In this embodiment, independent SARIMAX models are constructed for the underlying time series of different blood type and gender combinations, and their input sequences are denoted as follows: , , , as well as , , , ,in, , , , The data correspond to the time series of blood donation volumes for male blood types A, B, AB, and O, respectively. , , , The data correspond to the time series of blood donation volumes for women with blood types A, B, AB, and O, respectively.
[0172] During the model structure determination process, the optimal model parameter combinations for each underlying sequence were obtained by combining ACF and PACF analysis with the AIC / BIC optimal selection criteria. Among them, the seasonal cycle parameter is uniformly set to... This is used to characterize the annual seasonal variation features of monthly data. Finally, the optimal model structure corresponding to each underlying sequence is shown in the table below.
[0173] Table 1. Optimal Model Structure of Each Bottom Layer
[0174] Step 2.4: Hierarchical Prediction Integration and Nonlinear Correction Input Construction
[0175] After training and predicting each of the lower-level SARIMAX models, a bottom-up hierarchical aggregation method is used to construct the intermediate and final layer time series results. Let the first... The underlying sequence in time The predicted value is Then the intermediate layer The predicted sequence is obtained by summing the underlying sequence set it contains:
[0176]
[0177] in, Indicates the intermediate layer The underlying sequence index set contained therein This represents the total number of intermediate layers; in this embodiment, M=2. Furthermore, all intermediate layers are aggregated to obtain the total layer prediction sequence:
[0178]
[0179] During the model training phase, each underlying SARIMAX model is fitted within the time window of the training and validation sets to obtain in-sample fitting values. This fitting process is performed solely based on the training and validation sets, used to characterize the model's fitting ability within a range of historically known data, and to generate a new training set for the non-linear learning correction part. The overall layer fitting sequence is then constructed using the same hierarchical aggregation method.
[0180]
[0181] in, Indicates the first The fitted values of each underlying sequence within the training and validation set windows. This represents the total number of the underlying sequences.
[0182] Therefore, the total layer fitting sequence With predicted sequence These are used to characterize the model's fitting results on historical data (training set + validation set) and its prediction results within the target prediction window, respectively, thus providing a basis for the analysis and scheduling of overall monthly blood donation volume. To further characterize the nonlinear error features that the SARIMAX model failed to fully capture, the above-mentioned fitted and predicted sequences are used together as input data for subsequent nonlinear learning correction models to construct a neural network-based error compensation and optimization model.
[0183] Step 3: Nonlinear learning correction
[0184] After completing the trend prediction of the hierarchical SARIMAX model, errors caused by nonlinear factors may still exist in the prediction results, such as changes in group blood donation behavior, the impact of extreme weather, and changes in the blood supply environment. To further improve the overall layer prediction accuracy, this embodiment introduces a nonlinear learning module, namely a backpropagation neural network (BPNN), to perform nonlinear correction on the SARIMAX prediction results.
[0185] BPNN learns the systematic deviation between SARIMAX predictions and actual blood donations through a multi-layer feedforward structure, and outputs the final prediction result after nonlinear correction.
[0186] It should be noted that this invention combines statistical models with neural networks, rather than simply superimposing them. Instead, it leverages the complementary modeling capabilities of the two types of models. The SARIMAX model effectively characterizes the trend and seasonal structure in time series and linearly models the influence of external variables, but its ability to characterize complex nonlinear fluctuation patterns is limited. In contrast, BPNN possesses excellent nonlinear function approximation capabilities and can learn residual structures that statistical models fail to capture. Therefore, by constructing a fusion framework of "statistical prediction + nonlinear correction," prediction accuracy can be improved while maintaining model interpretability.
[0187] Step 3.1: BPNN Input / Output Construction and Network Structure Design
[0188] To enable the BPNN to learn the systematic biases of the SARIMAX model over historical periods, this embodiment constructs supervised learning samples based on the historical fitting results of the training and validation sets. During the training phase, only the fitting sequences of the SARIMAX model on historical samples are used. , as the input features of the BPNN, and the corresponding real observation values at the time step. As a supervised learning objective, the nonlinear residual mapping relationship of the SARIMAX model in historical stages is learned. Therefore, at time... The input-output relationship of a BPNN is defined as follows:
[0189]
[0190]
[0191] in, Indicates input features, This represents the target value for supervised learning. This represents the actual blood donation data obtained from actual observations. This represents the historical fitting results based on the training and validation sets.
[0192] This embodiment adopts a "one-dimensional input-single output" supervised learning structure, that is, at each time point... This forms a one-dimensional input sample, used to characterize the nonlinear mapping relationship between SARIMAX predicted values and true values.
[0193] It should be noted that this invention does not perform independent modeling for a single time point, but rather constructs a training dataset from multiple time point samples and uses batch gradient descent for overall training, thereby learning a globally consistent nonlinear residual correction function.
[0194] During the prediction phase, the input to the BPNN is the prediction sequence generated by the SARIMAX model within the future prediction window. It outputs the corresponding nonlinear correction results, thereby achieving error compensation for future predictions.
[0195] Before inputting the BPNN, to improve training stability, all sequences are normalized using min–max normalization and scaled to [the appropriate scale]. Interval. Let the original sequence be... The normalized sequence is Then the normalization formula is:
[0196]
[0197] in, and These are the minimum and maximum values of the sequence within the training window, respectively.
[0198] It should be noted that the SARIMAX model, as a statistical modeling method, is not sensitive to the scale of the input data, so no normalization is required; while BPNN, as a gradient-based nonlinear model, is more sensitive to the scale of the input data.
[0199] The dataset of this invention features a small monthly data volume and relatively small fluctuations in monthly blood donation volume. Furthermore, SARIMAX has absorbed trend and seasonal components, resulting in low-frequency, continuous changes in its prediction residuals over time, meaning the error changes are relatively smooth. This step employs a lightweight three-layer feedforward neural network structure, including: Input layer: 1 node; Hidden layer 1: 16 neurons, ReLU activation function; Hidden layer 2: 8 neurons, ReLU activation function; Hidden layer 3: 4 neurons, ReLU activation function; Output layer: 1 node, linear activation. This structure ensures sufficient parameters to learn non-linear relationships while avoiding the overfitting problem of deep networks on small-sample monthly data.
[0200] Step 3.2: Network training and optimization strategies.
[0201] During network training, this example uses mean squared error (MSE) as the loss function to measure the difference between the network output and the actual blood donation volume. The formula is as follows:
[0202]
[0203] in The number of training samples, This is the output of the BPNN. During training, all 48 samples are used collectively for iterative optimization of the network parameters. The globally optimal parameter values of the network are determined through a batch training process spanning multiple epochs. Training is not updated monthly, but rather based on a unified optimization of the global loss function to ensure that the model can capture the overall nonlinear bias patterns.
[0204] In terms of parameter optimization, this example uses a gradient-based adaptive optimization method (Adam) to update the network weights. This method can adaptively adjust the learning rate based on the first and second-order statistical information of the gradient, thereby improving the stability and convergence speed of the training process.
[0205] To prevent overfitting on small sample data, an early stopping mechanism is introduced during training, terminating training prematurely when the validation set error does not decrease significantly over several consecutive rounds. Simultaneously, L2 regularization constraints are applied to the network weights to limit model complexity and improve generalization ability.
[0206] After training, this invention yields a three-layer feedforward neural network model with fixed parameters. Structurally, this model consists of an input layer, multiple hidden layers, and an output layer. Computationally, it is implemented through a cascade of linear transformations and nonlinear activation functions, and its overall structure can be represented as a nonlinear mapping function:
[0207]
[0208] This nonlinear mapping function is used to learn the residual compensation relationship between the SARIMAX model's predicted values and the actual observed values. After training, its network parameters remain fixed, and only forward propagation calculations are performed during the prediction phase, without the need for retraining or parameter updates.
[0209] Furthermore, this nonlinear mapping function can be formally represented as a residual correction operator:
[0210]
[0211] in, This represents the predicted value from the SARIMAX model. This represents the nonlinear correction amount learned by BPNN. The final prediction result is obtained by superimposing the SARIMAX prediction value and the correction amount.
[0212] Step 3.3: Nonlinear correction output
[0213] After completing the BPNN network training, the overall layer prediction sequence obtained by the SARIMAX model during the prediction phase is used. The data is input into the trained BPNN network to obtain the final blood donation volume prediction result after nonlinear correction. This correction process can compensate for the systematic bias of the SARIMAX model under extreme conditions or complex behavioral changes, making the prediction result closer to the actual blood donation volume trend.
[0214] After training, the SARIMAX prediction sequence for the test set is input into the BPNN, and the final prediction value after nonlinear correction is obtained. It can be represented as:
[0215]
[0216] in, This is a linear prediction sequence of the overall layer of SARIMAX. The nonlinear mapping function learned by the trained BPNN network is used to compensate for the systematic bias of SARIMAX under extreme weather, holiday behavior fluctuations, and differences in group blood donation, while This is the final predicted value after nonlinear correction, which is closer to the actual trend of blood donation volume.
[0217] Step 4: Model Performance Evaluation
[0218] To verify the effectiveness and stability of the proposed hierarchical SARIMAX and BPNN fusion model in blood donation volume prediction, a multi-dimensional performance evaluation of the model's prediction results was conducted. The mean absolute percentage error (MAPE) was used as the evaluation metric to measure the relative error between the model's predicted values and the actual values. Its calculation formula is as follows:
[0219]
[0220] in, This indicates the actual amount of blood donated. This indicates the predicted blood donation volume. The sample length is denoted by . The smaller the MAPE value, the closer the prediction result is to the true value, and the higher the accuracy of the model.
[0221] In the experiment, the MAPE values of the traditional SARIMA model, the SARIMAX model with external variables, the hierarchical SARIMAX model, and the hierarchical SARIMAX + BPNN fusion model proposed in this invention were first calculated. The results are shown in Table 2 below, which shows that the MAPE of the traditional SARIMA model is 20.09%, the SARIMAX model decreases to 15.57% after introducing external variables, the hierarchical SARIMAX model further decreases to 10.53%, and the hierarchical SARIMAX + BPNN fusion model proposed in this invention finally reduces the MAPE to 9.97%.
[0222]
[0223] Table 2. Model Performance Comparison
Claims
1. A blood donation volume prediction method based on a hierarchical prediction framework and neural network nonlinear learning correction, characterized in that, include: First, data preparation and variable screening are carried out. Historical blood donation data are cleaned and organized, and key external variables are screened through correlation analysis. Secondly, a hierarchical time series structure is constructed, and the hierarchical SARIMAX model is used to model and predict the trends of each level of the series. Then, a backpropagation neural network is introduced for nonlinear learning to correct the statistical prediction results.
2. The method according to claim 1, characterized in that, (1) Data preparation and variable analysis: First, the acquired raw blood collection record data is cleaned, including business rule filtering, time field standardization, outlier handling, and field formatting. After the data cleaning is completed, the blood collection records are summarized by time dimension according to the blood collection date to construct a monthly statistical time series of blood donation volume. Based on this, the blood donation data is statistically analyzed in multiple dimensions according to the donors' gender and blood type information. The overall blood donation volume sequence is further divided into multiple subsequences to construct a time series system with a hierarchical structure. This hierarchical structure includes an overall layer, a gender layer, and a gender and blood type combination layer. Through hierarchical modeling, the overall blood donation trend and the behavioral differences between different blood donation groups can be characterized simultaneously.
3. The method according to claim 1, characterized in that, (2) Hierarchical statistical forecasting modeling: First, stationarity analysis is performed on the overall time series. Trend and seasonal components are eliminated through differencing, and the enhanced Dickey-Fuller (ADF) test is used to statistically verify stationarity. Once the series meets the stationarity condition, the lag dependency structure is further analyzed using the autocorrelation function (ACF) and partial autocorrelation function (PACF) to determine the candidate range of SARIMA model parameters. Based on this, different SARIMA(p,d,q)×(P,D,Q)_s model combinations are traversed within the preset parameter search space. The Akaike Information Criterion (AIC) and Bayesian Information Criterion (BIC) are used as model evaluation indicators, and the parameter combination that achieves the optimal balance between fitting accuracy and model complexity is selected as the final model structure. Based on the established SARIMA model structure, the key external variables obtained from the aforementioned screening are introduced to extend and construct the SARIMAX model, thereby achieving joint modeling of the time series trend term, seasonal term, and external driving factors of blood donation volume. Among them, each exogenous variable is jointly estimated with the model parameters through the maximum likelihood estimation method during the model training phase. In terms of hierarchical structure, each bottom-level sequence establishes a SARIMAX model and generates corresponding prediction results. Then, a bottom-up aggregation mechanism is used to aggregate the bottom-level prediction results step by step to obtain the intermediate layer and overall layer prediction sequences, thereby realizing unified modeling and integrated prediction of blood donation volume at different levels. At the same time, within the model training time window, the overall layer prediction sequence on historical data is obtained to characterize the model's fitting effect on known data and serve as the input basis for the subsequent nonlinear error learning module.
4. The method according to claim 1, characterized in that, (3) Nonlinear learning correction: A backpropagation neural network model (BPNN) based on the backpropagation algorithm is introduced as a prediction correction module to adaptively optimize the statistical prediction results. Specifically, the overall layer fitting sequence obtained by the statistical prediction model during the training phase is used as the input feature, and the actual blood donation volume sequence is used as the learning target. A nonlinear mapping relationship between the statistical prediction results and the actual blood donation volume is established through a multi-layer feedforward neural network structure. By continuously adjusting the network weight parameters through the error backpropagation algorithm, the neural network can learn the systematic biases and uncaptured nonlinear patterns in the statistical model's prediction results. In the prediction phase, the overall layer prediction sequence generated by the statistical prediction model within the prediction window is input into the trained neural network model to obtain the final blood donation volume prediction result after nonlinear correction.
5. The method according to claim 1, characterized in that, Step 1: Data preparation and variable selection specifically includes: Step 1.1: Data Preparation First, the original blood collection records are prepared and processed. The data preparation includes, but is not limited to, the following operations: verifying the blood collection records according to business rules, uniformly formatting the time fields such as the blood collection date, unifying the format of the blood donor identification and related coding fields, and removing or correcting abnormal or missing records to ensure the consistency of the original data in logic and format. After data cleaning, valid blood collection records were screened based on blood donation type, organization method, whether the amount was sufficient, and the results of initial and re-examinations. Subsequently, based on the valid blood collection records, the monthly blood donation volume was grouped and statistically analyzed according to two dimensions: gender and blood type, and a hierarchical time series structure was constructed. The hierarchical time series structure includes the following three levels: (1) Level 0: Monthly time series representing total blood donation volume This is used to reflect the changing trend of overall blood donation volume; (2) Intermediate layer (Level 1): Monthly blood donation volume time series categorized by gender, including male blood donation volume series. And the sequence of female blood donations This is used to characterize the impact of gender differences on blood donation behavior; (3) Level 2: Based on gender, further subdivisions are made according to blood type, with male blood donation volume subdivided by blood type. (Male, Type A) (Male, Type B) (Male AB type) and (Male blood type O); Female blood donation volume is further subdivided by blood type. (Female, Type A) (Female, Type B) (Female AB type) and (Women, blood type O), used to characterize the changing characteristics of blood donation behavior in different groups; The above hierarchical structure adopts a bottom-up additive constraint relationship, that is... and The summation relationship of the time series of the top layer, intermediate layer, and bottom layer is as follows: That is, the bottom-level time series are aggregated by gender to form the intermediate-level time series, and the intermediate-level time series are further aggregated to form the top-level time series. Step 1.2: Variable Selection The candidate external variable set is systematically constructed from the following four dimensions: (1) Meteorological environment dimension: The highest temperature, lowest temperature, precipitation, and weather conditions were selected as candidate variables; (2) Time-related social dimension: Whether it is a holiday or not was selected as a candidate variable; (3) Service supply dimension: The number of open blood collection points was selected as a candidate variable; (4) Environmental quality dimension: Air quality level was selected as a candidate variable; Construct a set of candidate external variables ; This is used to reflect factors that may affect changes in blood donation volume, among which The highest temperature, The lowest temperature, To determine the number of open blood collection points, For precipitation, Air quality level, Whether it is a holiday or not, Weather conditions; Before performing correlation calculations, candidate external variables undergo unified data preprocessing. For numerical variables, such as maximum temperature, minimum temperature, precipitation, and number of blood collection points, a normalization method is used to standardize them to eliminate the influence of dimensional differences on the correlation measurement results. For non-numerical variables, such as air quality level, whether it is a holiday, and weather conditions, they are converted into discrete numerical form through numerical coding.
6. The method according to claim 1, characterized in that, Step 2: Hierarchical statistical forecasting modeling specifically involves: Step 2.1: Determining the range of model order In the process of time series modeling, let the original time series be... The first-order difference can be expressed as the following mathematical expression: ; It is used to eliminate long-term trends if the sequence has a period of The seasonal pattern can be observed through seasonal differences, expressed mathematically as follows: ; in, This represents the lag operator, which shifts the sequence one time step back. Its mathematical expression is as follows: ; Furthermore, there is a mathematical expression: ; in, Denotes the lag order, taken in the first-order difference. Taking from seasonal differences When considering both trend and seasonal terms, a combined difference form is constructed, with the following mathematical expression: ; in Indicates the first-order difference degree. Indicates the degree of seasonal difference. Indicates the length of the seasonal cycle; The seasonal autoregressive moving average (SARIMA) model is uniformly expressed mathematically as follows: ; The model introduces a hysteresis operator. Representing the dynamic relationship of time series as about The polynomial form of the non-seasonal autoregressive polynomial is as follows: ; This mathematical expression describes the relationship between the current value and the observations from the past p time points, where , , …, These are autoregressive coefficients, used to represent the weight of the influence of historical observations on the current value under different lag orders. This indicates the degree of influence of the previous value on the current value; meanwhile, the mathematical expression for the non-seasonal moving average polynomial is as follows: ; This mathematical expression describes the relationship between the current value and past error terms, where , , …, These are the moving average coefficients, used to represent the weight of the historical error term on the current value under different lag orders. This indicates the degree of influence of the error at the previous time step on the current value; the seasonality part is expressed by the following two mathematical expressions, where... Indicates a lagged operation spanning a seasonal cycle: ; ; The above expression describes the influence of historical observations across seasonal cycles on the current value and the influence of the error term across seasonal cycles on the current value, where , , …, The seasonal autoregressive coefficient represents the dependency weight between the current observation and the observations over the past P seasonal cycles, each with a cycle length of s. This indicates the influence of the value from the previous seasonal cycle on the current value; , , …, This is the seasonal moving average coefficient, representing the dependency weight between the current observation and the random error term over the past Q seasonal periods. This indicates the impact of the error from the previous seasonal cycle on the current value; Represents the autoregressive component, where Used to characterize the linear dependence between the current value and the observations at the past p times under non-seasonal conditions. The influence of historical observations across seasonal cycles on current values is represented by φ (phi), which reflects the impact of historical values on current values. Represents the moving average component, where Used to describe the relationship between current values and past error terms under non-seasonal conditions. This section describes the impact of error terms across seasonal cycles on the current value; it is denoted by θ (theta) and reflects the impact of historical errors on the current value. in, and In this context, lowercase φ and θ represent the non-seasonal autoregressive term and the non-seasonal moving average term, respectively, used to characterize the short-term dependency between adjacent time points. and In this context, the uppercase Φ and Θ represent the seasonal autoregressive term and the seasonal moving average term, respectively, which are used to characterize long-term dependencies across seasonal cycles. Comparative analysis was conducted on the time series of monthly blood donation volume at the overall level using the original sequence, the first-order difference sequence, and the first-order difference combined with seasonal difference sequence to determine the difference order parameter. Set the difference order to: ordinary difference order Seasonal difference order Seasonal cycle ; Based on this, ACF and PACF plots were drawn for the stationary series after "first-order + seasonal difference" to identify the potential value ranges of non-seasonal and seasonal parameters, respectively. In time series analysis, the autocorrelation function (ACF) and partial autocorrelation function (PACF) are important tools for analyzing the correlation structure of sequences, and are often used to determine the autoregressive order and moving average order of ARIMA or SARIMA models; let the time series be... The sample length is The lag order is denoted as (Integer, range of values) ); The autocorrelation function (ACF) describes the overall correlation of a series at different lags, that is, the strength of the linear relationship between the series and its lag values; the mathematical expression for the sample autocorrelation coefficient is: ; in The mean of the sequence. The lag order is calculated by... corresponding It can draw an ACF plot to observe how the overall correlation of the sequence decreases with lag order; The partial autocorrelation function (PACF) represents the direct correlation between a series and a given lag value after removing the effects of intermediate lags; the partial autocorrelation coefficient... The following was obtained using an autoregressive model: ; in For the error term, That is, lag The partial autocorrelation coefficient; plotting different corresponding Obtain the PACF plot; In summary, the ACF plot is used to observe how the overall correlation of a series decays with lag order, thus helping to determine the order q of the non-seasonal moving average (MA) and the order Q of the seasonal MA; the PACF plot is used to observe the direct correlation after removing the effects of intermediate lags, thus helping to determine the order p of the non-seasonal autoregression (AR) and the order P of the seasonal AR. Set the candidate order of SARIMAX as follows: ; Step 2.2: Introduction of external variables Based on the external variable screening results in step 1.2, the same exogenous input variables are introduced into each underlying SARIMAX model without further screening. These include the highest temperature, lowest temperature, number of open blood collection points, and precipitation. The SARIMAX model introduces exogenous variables based on the SARIMA model, and its expression can be represented as: ; in, Indicates time Blood donation volume observations at specific times; Indicates time The vector of exogenous variables at time t; This is the vector of regression coefficients corresponding to the exogenous variables; This is the random error term; Based on the autocorrelation analysis results, the model order parameter and The range of values for is predetermined and used to constrain the subsequent model search space; the exogenous variable coefficient β and the random error term As parameters to be estimated, they are jointly solved using the maximum likelihood estimation method during the model training phase; The vector representation of exogenous variables is: ; in, Indicates the highest temperature. Indicates the lowest temperature. Indicates the number of open blood collection points. Indicates precipitation; The maximum temperature, minimum temperature and precipitation are constructed with one-phase lag features and used as model inputs, while the number of open blood collection points is not lag-processed because it can be planned and obtained in advance; Step 2.3: Low-level sequence SARIMAX modeling and parameter optimization Based on the preliminary determination of the range of possible model orders in step 2.1 and the introduction of exogenous variables in step 2.2, SARIMAX modeling and parameter optimization are carried out on the underlying time series. The input to each underlying SARIMAX model consists of endogenous time series and exogenous variables; among them, the endogenous variables are the monthly blood donation volume series corresponding to different blood type and gender combinations, denoted as follows: The set of exogenous variables is denoted as This is used to characterize external factors influencing blood donation behavior, including weather conditions, the availability of blood collection sites, and social activities; by simultaneously inputting... and This enables joint modeling of the target sequence; During model training, the original monthly sequence is divided into training, validation, and test sets according to time. Model parameters are estimated only on the training set, and the performance of different order combinations is compared and selected on the validation set. The test set is only used for the final model generalization performance evaluation. In order to fully explore the possible values of the model order and reduce the human discrimination error based on ACF and PACF, this embodiment adopts a traversal search method to train and evaluate the candidate parameter combinations of each bottom sequence one by one. The model selection criteria adopt the Akaike Information Criterion (AIC) and the Bayesian Information Criterion (BIC), which are defined as follows: ; ; in, The number of model parameters, The likelihood function value, The sample length is denoted by . The smaller the AIC and BIC are, the better the balance between fitting accuracy and complexity of the model. Therefore, by comprehensively comparing the performance of AIC and BIC on the validation set, the parameter combination with better performance in both metrics is selected as the optimal model structure for each underlying sequence. In the actual model selection process, the values of AIC and BIC can be comprehensively evaluated; if a candidate model... By calculating the corresponding AIC and BIC, a comprehensive evaluation index can be constructed: ; in, Indicates the first The comprehensive evaluation value of each candidate model; when there exists When there are 10 candidate models, the optimal model structure is represented as follows: ; in, Indicates the candidate model number. This represents the total number of candidate models. This represents the index of the optimal model that minimizes the weighted sum of AIC and BIC among all candidate models; That is, among all candidate models, the model with the smallest sum of AIC and BIC is selected as the final optimal model structure. Independent SARIMAX models were constructed for the underlying time series data of different blood types and gender combinations, with their input sequences denoted as follows: , , , as well as , , , ,in, , , , The data correspond to the time series of blood donation volumes for male blood types A, B, AB, and O, respectively. , , , The time series correspond to the blood donation volumes of women with blood types A, B, AB, and O, respectively. During the model structure determination process, the optimal model parameter combinations for each underlying sequence were obtained by combining ACF and PACF analysis with the AIC / BIC optimal selection criteria. Among them, the seasonal cycle parameter is uniformly set as follows: It is used to characterize the annual seasonal variation features of monthly data; Step 2.4: Hierarchical Prediction Integration and Nonlinear Correction Input Construction After training and predicting each of the underlying SARIMAX models, a bottom-up hierarchical aggregation method is used to construct the time series results of the intermediate and overall layers; let the first layer be... The underlying sequence in time The predicted value is Then the intermediate layer The predicted sequence is obtained by summing the underlying sequence set it contains: ; in, Indicates the intermediate layer The underlying sequence index set contained therein This represents the total number of intermediate layers, M=2; further, by summing all intermediate layers, we obtain the total layer prediction sequence: ; During the model training phase, each underlying SARIMAX model is fitted within the time window of the training and validation sets to obtain in-sample fitting values. This fitting process is performed solely based on the training and validation sets, used to characterize the model's fitting ability within the range of historically known data, and to generate a new training set for the non-linear learning correction part; and a total layer fitting sequence is constructed using the same hierarchical aggregation method: ; in, Indicates the first The fitted values of each underlying sequence within the training and validation set windows. This represents the total number of the underlying sequences. Therefore, the total layer fitting sequence With predicted sequence These are used to characterize the model's fitting results on historical data and its prediction results within the target prediction window, respectively, thus providing a basis for the analysis and scheduling of overall monthly blood donations. The fitted sequence and the prediction sequence are used together as input data for the subsequent nonlinear learning correction model to construct an error compensation and optimization model based on a neural network.
7. The method according to claim 1, characterized in that, Step 3: Nonlinear learning correction specifically involves; A nonlinear learning module, namely the backpropagation neural network (BPNN), is introduced to perform nonlinear correction on the SARIMAX prediction results. The BPNN learns the systematic deviation between the SARIMAX prediction value and the actual blood donation volume through a multi-layer feedforward structure, thereby outputting the final prediction result after nonlinear correction. Step 3.1: BPNN Input / Output Construction and Network Structure Design In order to enable BPNN to learn the systematic biases of the SARIMAX model in historical stages, supervised learning samples are constructed based on the historical fitting results of the training set and validation set; During the training phase, only the fitted sequences of the SARIMAX model on historical samples are used. , as the input features of the BPNN, and the corresponding real observation values at the time step. As a supervised learning objective, it learns the nonlinear residual mapping relationship of the SARIMAX model in historical stages; therefore, at time... The input-output relationship of a BPNN is defined as follows: ; ; in, Indicates input features, Indicates the value of the supervised learning objective; This represents the actual blood donation data obtained from actual observations. This represents the historical fitting results based on the training and validation sets; A supervised learning structure of "one-dimensional input - single output" is adopted, that is, at each time point... This forms a one-dimensional input sample, used to characterize the nonlinear mapping relationship between the SARIMAX predicted value and the true value; During the prediction phase, the input to the BPNN is the prediction sequence generated by the SARIMAX model within the future prediction window. It outputs the corresponding nonlinear correction results, thereby achieving error compensation for future predictions; Before inputting into the BPNN, all sequences are normalized using a min–max method to scale them to the maximum value. Interval; let the original sequence be The normalized sequence is Then the normalization formula is: ; in, and These are the minimum and maximum values of the sequence within the training window, respectively. A lightweight three-layer feedforward neural network structure is adopted, including: input layer: 1 node, hidden layer 1: 16 neurons, ReLU activation function, hidden layer 2: 8 neurons, ReLU activation function, hidden layer 3: 4 neurons, ReLU activation function, and output layer: 1 node, linear activation. This structure ensures that the number of parameters is sufficient to learn non-linear relationships while avoiding the overfitting problem of deep networks on small sample monthly data. Step 3.2: Network Training and Optimization Strategies During network training, mean squared error (MSE) is used as the loss function to measure the difference between the network output and the actual blood donation volume. The formula is as follows: ; in The number of training samples, This is the output of the BPNN; during training, all 48 samples are used as a whole for iterative optimization of network parameters, and the global optimal parameter values of the network are jointly determined through a batch training process of multiple epochs; the training is not updated monthly, but is uniformly optimized based on the global loss function to ensure that the model can capture the overall nonlinear deviation pattern; In terms of parameter optimization, the network weights are updated using a gradient-based adaptive optimization method (Adam). To prevent overfitting of the model on small sample data, an early stopping mechanism is introduced during training. Training is terminated early when the validation set error does not decrease significantly within several consecutive rounds. At the same time, L2 regularization constraints are applied to the network weights to limit model complexity and improve generalization ability. After training, a three-layer feedforward neural network model with fixed parameters is obtained. Structurally, this model consists of an input layer, multiple hidden layers, and an output layer. Computationally, it is implemented by a cascade of linear transformations and nonlinear activation functions, and its overall representation is a nonlinear mapping function. ; This nonlinear mapping function can be formally represented as a residual correction operator: ; in, This represents the predicted value from the SARIMAX model. This represents the nonlinear correction amount learned by BPNN. The final prediction result is obtained by superimposing the SARIMAX prediction value and the correction amount. Step 3.3: Nonlinear correction output After completing the BPNN network training, the overall layer prediction sequence obtained by the SARIMAX model during the prediction phase is used. The data is input into the trained BPNN network to obtain the final blood donation volume prediction result after nonlinear correction. After training, the SARIMAX prediction sequence for the test set is input into the BPNN, and the final prediction value after nonlinear correction is obtained. It can be represented as: ; in, This is a linear prediction sequence of the overall layer of SARIMAX. The nonlinear mapping function learned by the trained BPNN network, and This is the final predicted value after nonlinear correction.