A platelet demand prediction system based on time series

By adopting time series decomposition and multi-model combination methods in the platelet demand prediction system, the problems of large prediction error and low efficiency in the prior art are solved, and more accurate and efficient platelet demand prediction is achieved, and assisted by intelligent decision-making.

CN119028452BActive Publication Date: 2025-06-06ZHEJIANG PROVINCIAL BLOOD CENT
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411520252.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-29
Publication Date
2025-06-06
Estimated Expiration
2044-10-29

AI Technical Summary

Technical Problem

The prior art has problems of large prediction errors and low efficiency in platelet demand prediction, and the model lacks generalization capabilities, making it difficult to achieve automated and accurate demand prediction.

Method used

A time series-based platelet demand prediction system is adopted, which includes data acquisition and alignment module, time series adjustment module, time series decomposition module, sub-sequence modeling and integration module, model evaluation module and model deployment module. The time series was decomposed by the X-13ARIMA-SEATS method, subsequence prediction was performed by combining ARIMA, Prophet, TimeGPT and other models, and multiple rounds of modeling were realized through rolling windows, and model evaluation was performed using weighted accuracy indicators.

Benefits of technology

It improves the accuracy and efficiency of platelet demand prediction, reduces prediction errors, enhances the generalization ability of the model, and assists decision-making through intelligent and graphical prediction systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119028452B_ABST
    Figure CN119028452B_ABST
Patent Text Reader

Abstract

The present invention provides a platelet demand prediction system based on time series, including a data acquisition and regularization module, a time series adjustment module, a time series decomposition module, a subsequence modeling and integration module, a model evaluation module and a model deployment module. The platelet demand prediction system based on time series of the present invention uses the X‑13ARIMA‑SEATS method to decompose the platelet clinical supply sequence into its trend, seasonality and residual subsequences, constructs TimeGPT and ARIMA models (or uses the SNAIVE method) for trend and seasonality respectively to perform subsequence prediction, and recombines the prediction results into a prediction of the entire sequence based on multiplicative decomposition, thereby constructing a trend×seasonal decomposition-combination model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of platelet demand prediction, and in particular relates to a platelet demand prediction system based on time series. Background Art

[0002] Platelet transfusion plays a vital role in modern medicine. It is not only used to preventively reduce the risk of bleeding, but also to manage active bleeding therapeutically. Especially in patients with blood diseases and tumors, platelet transfusion is one of the common treatment methods. However, the clinical supply of platelets faces many challenges. In recent years, the clinical supply of platelets in Hangzhou has grown rapidly. This growth is affected by many factors, such as the population size and age structure of the region, the configuration of medical beds, seasonal factors, patients with indications, the knowledge of clinical doctors on blood transfusion, medical reform and social and economic development. Moreover, the various factors are intricately related and difficult to accurately measure. In addition, the shelf life of platelets is short. When stored in a platelet-specific blood bag, the shelf life is 5 days, and when stored in an ordinary blood bag, the shelf life is only 24 hours. On the one hand, it is necessary to ensure the clinical supply of platelets, and on the other hand, it is necessary to prevent platelets from being expired and scrapped. These have brought great challenges to the platelet collection and supply management of blood stations. Using statistical methods, we analyzed the time series data of clinical platelet supply in Zhejiang Provincial Blood Center over the past 20 years and established a prediction model. By predicting clinical platelet demand, we can optimize platelet collection plans, improve platelet supply capacity and reduce blood waste caused by platelet expiration.

[0003] Existing studies on component blood supply prediction in China, such as Liu Yanyan et al., 2021; Peng Rongrong et al., 2020; Xie Shuhong et al., 2021, adopt the classic Box-Jenkins research paradigm, namely: (1) Stationary analysis and processing of time series: Use ADF-Augmented Dickey-Fuller, PP-Phillips-Perron test or KPSS-Kwiatkowski-Phillips-Schmidt-Shin test to determine whether the sequence is stationary. If it is stationary, proceed to the next step; if it is not stationary, try to perform differences and transformations to achieve stationarity. (2) Model identification and parameter estimation: Draw ACF Autocorrelation Function, autocorrelation function and PACF-Partial Autocorrelation Function, partial autocorrelation function graphs, propose alternative ARIMA models based on their tailing and truncation conditions, fit the model to obtain parameters and perform significance tests on them, and determine the best model based on information criteria such as AIC-AkaikeInformation Criterion and BIC-Bayesian Information Criterion. (3) Model test: Determine whether the model residual sequence is white noise by observing the ACF graph of the residual sequence and performing the Ljung-Box test. (4) Model prediction: Use the optimal model.

[0004] This type of research has the following defects: model identification is based on empirical judgments such as ACF and PACF, and there is a possibility that the globally optimal model parameters cannot be selected. It relies on the personal experience of researchers and cannot be automated. The model prediction range is generally the last six months or one year of the data set. Error indicators such as RMSE and MAPE are calculated through such a single narrow range prediction. The error indicators are easily affected by randomness and external events, so there are large fluctuations in prediction errors. The model lacks generalization ability and may lead to incomplete and unreliable model evaluation.

[0005] Recent studies in the field of platelet prediction abroad, such as MOTAMEDI et al., 2024, predicted platelet demand using a large clinical data set provided by four hospitals in Canada. (1) This study used the STL method to decompose the time series into trend, seasonality, and residual subsequences, thereby analyzing the impact of sequence patterns such as weekdays, weekends, and holidays. (2) The univariate ARIMA and Prophet models and the multivariate Lasso regression, random forest, and LSTM models were used to predict platelet demand. (3) The rolling window analysis method was used to fit the model with different window moving steps on a 2-year or 8-year data window, and the mean error indicator was reported. This study differs from the present invention in the following ways: This study used the STL method to decompose the time series, but only used it for sequence pattern analysis, and still fitted the entire sequence when fitting the model; This study conducted a rolling window analysis, but used the arithmetic mean in the error indicator reporting, giving equal importance to the prediction errors of the models fitted in each time window; From the perspective of the MAPE indicator, this study did not perform well in predicting the daily blood supply of four hospitals in Canada. Even with a retraining cycle of only 1 day and a data window size of 8 years, the average predicted MAPE was still close to 20%.

[0006] Therefore, how to solve the problems of large prediction errors and low efficiency of complex time series and provide a time series-based platelet demand prediction system with a trend × seasonality decomposition-combination model is a technical problem that needs to be solved urgently by those skilled in the art. Summary of the invention

[0007] The purpose of the present invention is to provide a platelet demand prediction system based on time series to address the problems in the prior art.

[0008] To this end, the above-mentioned purpose of the present invention is achieved through the following technical solutions:

[0009] A platelet demand prediction system based on time series, characterized by comprising a data acquisition and regularization module, a time series adjustment module, a time series decomposition module, a subsequence modeling and integration module, a model evaluation module and a model deployment module, wherein:

[0010] In the S1, data acquisition and regularization module, based on the clinical blood supply data collected by Zhejiang Provincial Blood Center, the data reading, format regularization, and descriptive statistics functions are integrated, and the regularized data are output to the time series adjustment module;

[0011] In S2, the time series processing module, outlier processing is performed on the received time series data, the required difference order is identified to meet the stationarity, and the Box-Cox transformation is performed to meet the normality; the adjusted time series is decomposed using the X-13ARIMA-SEATS method to obtain the trend subsequence, seasonal subsequence and residual subsequence. The decomposed time series enters the rolling window and is rolled out to the subsequence modeling and integration module batch by batch and year by year;

[0012] In S3, the subsequence modeling and integration module, the ARIMA, Prophet, TimeGPT, and SNAIVE models are fitted to the subsequence fitting models in the received window for subsequence prediction, and the one-year forecast value of the entire sequence is obtained by combining them; this module is executed through a rolling window loop to obtain all the forecast values ​​of the specified modeling method combination over the specified time span, and output to the model evaluation module;

[0013] In S4, the model evaluation module calculates the MAPE and RMSE accuracy indicators for the received prediction values, and improves the accuracy of recent predictions by exponential weighting, thereby calculating the weighted MAPE and weighted RMSE indicators to evaluate the overall performance of the prediction system and determine whether it can be deployed;

[0014] S5. Model deployment module: Use Shiny to encapsulate the model that has passed the model evaluation into an interactive web application, so as to realize the intelligent and graphical blood supply prediction work and assist in decision making.

[0015] While adopting the above technical solutions, the present invention may also adopt or combine the following technical solutions:

[0016] As the preferred technical solution of the present invention:

[0017] The data acquisition and regularization module in S1 includes the following steps:

[0018] S11. Obtain monthly data on clinical platelet supply from the Zhejiang Provincial Blood Center platform;

[0019] S12. Check the format of the input data to see whether it contains a time column and a value column. If the input data does not meet the conditions, an exception is thrown to indicate that the input data is incorrect. If the data meets the conditions, the time column is extracted as a time series index and the data is encapsulated as a tisbble time series data frame.

[0020] S13. Perform descriptive statistics on the sequence.

[0021] As the preferred technical solution of the present invention:

[0022] The time series processing module in S2 includes the following steps:

[0023] S21. Draw time series graphs and seasonal graphs, and analyze time series trends, seasonality and other patterns

[0024] S22. Detect outliers in time series: temporarily use the STL method to remove the trend and seasonality of the series, use the median and inner quartile range to detect anomalies in the residuals, and mark the data points outside the inner quartile range of the median plus or minus a specified multiple as outliers;

[0025] S23, outlier processing: remove outliers and use Kalman smoothing algorithm to perform smooth interpolation processing on outliers;

[0026] S24. Combined with the KPSS stationarity test, the minimum difference and the minimum seasonal difference that make the time series stable are identified and applied;

[0027] S25. Use Guerrero's (1993) algorithm to select the optimal Box-Cox transformation Lambda value and apply the transformation;

[0028] S26. Perform X-13ARIMA-SEATS decomposition (introducing Chinese holidays as exogenous regressor variables) to decompose the time series into trend, seasonality, and residual subsequences.

[0029] As the preferred technical solution of the present invention:

[0030] The S3 neutron sequence modeling and integration module includes the following steps:

[0031] S31. Use the TimeGPT API or the R language fable package to fit the TimeGPT model or ARIMA model to the trend subsequence and generate a trend subsequence forecast;

[0032] S32. Fit the ARIMA model to the seasonal subseries or use the seasonal naive method to generate seasonal subseries forecasts;

[0033] S33, using a multiplication model to synthesize the overall prediction value of the sequence;

[0034] S34, perform Ljung-Box test on model residuals;

[0035] S35. Use the rolling window method to cyclically run S31, S32, S33, and S34 to achieve overall prediction of the sequence within the specified time range and model residual test, and pass the prediction results and test results to the model evaluation module.

[0036] As the preferred technical solution of the present invention:

[0037] The model evaluation module in S4 includes the following steps:

[0038] S41. Combined with the true value of the sequence, calculate the forecast error indicators: RMSE, MAPE;

[0039] S42. Assign weights to the forecast errors of each year by means of index weighting;

[0040] S43, calculate the weighted MAPE and weighted RMSE of the entire model;

[0041] S44: Determine the availability of the model based on the model's weighted error index, forecast performance in each year, and residual analysis. If available, proceed to S5.

[0042] As the preferred technical solution of the present invention:

[0043] The model deployment module in S5 includes the following steps:

[0044] S51. Report the parameters of the tested model, draw a line chart based on the predicted results and true values ​​of the tested model, and perform a static display of the model;

[0045] S52. Use Shiny to build a web application, including data uploading, sequence visualization and adjustment, model selection and adjustment, and prediction result downloading functions;

[0046] S53. Connect the network application to the blood center server and the smart blood platform, use the network application to debug the model, generate and download prediction results, and then provide auxiliary decision-making.

[0047] Compared with the prior art, a platelet demand prediction system based on time series of the present invention has the following beneficial effects: in the present invention, the Hyndman-Khandakar stepwise algorithm is adopted in ARIMA model identification to automatically search for ARIMA model parameters and find the optimal parameters with a small amount of calculation; the present invention introduces the TimeGPT-1 model based on the Transformer architecture to improve the prediction performance of complex time series; the present invention not only utilizes time series decomposition to analyze sequence patterns, but also selects models for construction and subsequence prediction according to the characteristics of subsequences, thereby achieving lower model prediction errors; the present invention selects exponential weighting of the prediction error indicators of the models fitted to each time window according to the temporal importance of the model prediction results, thereby obtaining the weighted average of error indicators such as MAPE and RMSE, and taking into account both overall considerations and emphasis on recent prediction effects in model evaluation.

[0048] The present invention discloses a time series-based platelet demand prediction system which uses the X-13ARIMA-SEATS method to decompose the platelet clinical supply sequence into its trend, seasonality and residual subsequences, constructs TimeGPT and ARIMA models for the trend and seasonality respectively / or uses the SNAIVE method to perform subsequence prediction, recombines the prediction results into a prediction of the sequence as a whole based on multiplicative decomposition, thereby constructing a trend × seasonality decomposition-combination model, realizes multiple rounds of modeling through a rolling window method, uses a weighted precision index to achieve a more comprehensive and reliable model evaluation, and constructs a more complete platelet demand prediction system by introducing a new model and improving the evaluation method. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 It is a structural schematic diagram of a platelet demand prediction system based on time series of the present invention. DETAILED DESCRIPTION

[0050] The present invention will be described in further detail with reference to the accompanying drawings and specific embodiments.

[0051] Example 1

[0052] like Figure 1 As shown, a mixing prediction method and system based on time series data of the present invention are further described in detail.

[0053] S1. As shown in the attached figure, in the data acquisition and regularization module, based on the clinical platelet supply data collected by Zhejiang Provincial Blood Center, the data reading, format regularization, and descriptive statistics functions are integrated, and the regularized data are output to the time series adjustment module;

[0054] S2. As shown in the attached figure, in the time series processing module, outlier processing is performed on the received time series data, the required difference order is identified to meet the stationarity, and the Box-Cox transformation is performed to meet the normality; the adjusted time series is decomposed using the X-13ARIMA-SEATS method to obtain the trend subsequence, seasonal subsequence and residual subsequence. The decomposed time series enters the rolling window and is rolled out to the subsequence modeling and integration module batch by batch and year by year;

[0055] S3. As shown in the attached figure, in the subsequence modeling and integration module, the ARIMA, Prophet, TimeGPT, and SNAIVE models are fitted to the subsequence fitting models in the received window for subsequence prediction, and the one-year forecast value of the entire sequence is obtained by combining them; this module is executed cyclically through a rolling window to obtain all the forecast values ​​of the specified modeling method combination over the specified time span, and output to the model evaluation module;

[0056] S4. As shown in the attached figure, in the model evaluation module, the MAPE and RMSE accuracy indicators are calculated for the received prediction values, and the recent prediction accuracy is improved by exponential weighting, so as to calculate the weighted MAPE and weighted RMSE indicators to evaluate the overall performance of the prediction system and determine whether it can be deployed;

[0057] S5. As shown in the attached figure, the model deployment module: Shiny is used to encapsulate the model that has passed the model evaluation into an interactive network application, so as to realize the intelligent and graphical blood supply prediction work and provide auxiliary decision-making.

[0058] In the present invention, the data acquisition and regularization module in S1 includes the following steps:

[0059] S11. Obtain monthly data on clinical platelet supply from the Zhejiang Provincial Blood Center platform;

[0060] S12. Check the format of the input data to see whether it contains a time column and a value column. If the input data does not meet the conditions, an exception is thrown to indicate that the input data is incorrect. If the data meets the conditions, the time column is extracted as a time series index and the data is encapsulated as a tisbble time series data frame.

[0061] S13. Perform descriptive statistics on the sequence.

[0062] Furthermore, the time series processing module in S2 includes the following steps:

[0063] S21. Draw time series graphs and seasonal graphs, and analyze time series trends, seasonality and other patterns;

[0064] S22. Detect outliers in time series: temporarily use the STL method to remove the trend and seasonality of the series, use the median and inner quartile range to detect anomalies in the residuals, and mark the data points outside the inner quartile range of the median plus or minus a specified multiple as outliers;

[0065] S23, outlier processing: remove outliers and use Kalman smoothing algorithm to perform smooth interpolation processing on outliers;

[0066] S24. Combined with the KPSS stationarity test, the minimum difference and the minimum seasonal difference that make the time series stable are identified and applied;

[0067] S25. Use Guerrero's (1993) algorithm to select the optimal Box-Cox transformation Lambda value and apply the transformation;

[0068] S26. Perform X-13ARIMA-SEATS decomposition (introducing Chinese holidays as exogenous regressors) to decompose the time series into trend, seasonality, and residual subsequences.

[0069] Series t =Trend t ×Seasonality t ×Irregular t .

[0070] Furthermore, the subsequence modeling and integration module in S3 includes the following steps:

[0071] S31. Use the TimeGPT API or the R language fable package to fit the TimeGPT model or ARIMA model to the trend subsequence and generate a trend subsequence forecast;

[0072] S32. Fit the ARIMA model to the seasonal subseries or use the seasonal naive method to generate seasonal subseries forecasts;

[0073] S33, using the multiplication model to synthesize the overall prediction value of the sequence

[0074]

[0075] S34, perform Ljung-Box test on model residuals;

[0076] S35. Use the rolling window method to cyclically run S31, S32, S33, and S34 to achieve overall prediction of the sequence within the specified time range and model residual test, and pass the prediction results and test results to the model evaluation module.

[0077] Furthermore, the model evaluation module in S4 includes the following steps:

[0078] S41. Combined with the true value of the sequence, calculate the forecast error indicators: RMSE, MAPE;

[0079] S42. Assign weights to the forecast errors of each year by means of index weighting; S43. Calculate the weighted MAPE and weighted RMSE of the overall model

[0080]

[0081] S44: Determine the availability of the model based on the model's weighted error index, forecast performance in each year, and residual analysis. If available, proceed to S5.

[0082] Furthermore, the model deployment module in S5 includes the following steps:

[0083] S51. Report the parameters of the tested model, draw a line chart based on the predicted results and true values ​​of the tested model, and perform a static display of the model;

[0084] S52. Use Shiny to build a web application, including data uploading, sequence visualization and adjustment, model selection and adjustment, and prediction result downloading functions;

[0085] S53. Connect the network application to the blood center server and the smart blood platform, use the network application to debug the model, generate and download prediction results, and then assist in decision-making.

[0086] In this application, the Box-Cox transformation refers to a power transformation method proposed by George Box and David Cox in 1964 for making data closer to a normal distribution.

[0087] The X-13ARIMA-SEATS method refers to a complex algorithm for seasonal adjustment and forecasting of time series. The method was developed by the U.S. Census Bureau to improve the accuracy of time series analysis by combining the ARIMA model and signal extraction techniques.

[0088] The SEATS method is a signal extraction technology used for time series analysis, and its full name is Seasonal Extraction in ARIMA Time Series.

[0089] ARIMA model refers to the Autoregressive Integrated Moving Average model, which is a statistical method used for time series analysis and prediction.

[0090] The Prophet model refers to an open source time series forecasting tool developed by Facebook, which is mainly used to process time series data with strong seasonality.

[0091] The TimeGPT model refers to a time series prediction method based on a generative pre-training model (GPT), which generates future time points by learning patterns in historical data.

[0092] The SNAIVE model refers to a time series prediction model based on the Naive Bayes algorithm, which assumes that future values ​​only depend on the most recent past values.

[0093] The MAPE accuracy index refers to the mean absolute percentage error, which is a commonly used indicator to evaluate the accuracy of time series forecasts.

[0094] The RMSE accuracy index refers to the root mean square error, which is another commonly used indicator to evaluate the accuracy of time series forecasting.

[0095] Shiny is an R package for building interactive web-based applications.

[0096] Tsibble is a data frame format for processing time series data in R language. It provides a convenient way to store and manipulate time series data.

[0097] The platelet demand prediction system based on time series of the present invention comprises a data acquisition and regularization module, a time series adjustment module, a time series decomposition module, a subsequence modeling and integration module, a model evaluation module and a model deployment module; wherein the data acquisition and regularization module collects data based on the Zhejiang Provincial Blood Center and regularizes it into a time series format; the time series processing module analyzes the time series morphology, performs stabilization, Box-Cox transformation and other processing on it, and after adjustment, uses the X-13ARIMA-SEATS method to decompose the sequence multiplication into trend, seasonality and residual subsequences; the subsequence modeling and integration module uses ARIMA, Prophet, TimeGPT etc. models to model and predict the subsequences, and integrate to obtain the overall prediction results of the sequence; the model evaluation module models and integrates the subsequences through a rolling window method, and calculates the weighted MAPE index to evaluate the model performance; the model deployment module uses Shiny to build an interactive Internet application to realize the intelligent and graphical work of blood supply prediction, and then provide auxiliary decision-making, introduce more advanced prediction models and more comprehensive evaluation methods to improve the accuracy of prediction and the reliability of evaluation; the present invention provides a platelet demand prediction system based on time series, which improves the prediction performance through the use of a decomposition-combination model, realizes multiple rounds of modeling through a rolling window method, and uses weighted accuracy indicators to achieve more comprehensive and reliable model evaluation.

[0098] The above-mentioned specific implementation methods are used to explain the present invention and are only preferred embodiments of the present invention, rather than limiting the present invention. Any modifications, equivalent substitutions, improvements, etc. made to the present invention within the spirit of the present invention and the protection scope of the claims shall fall within the protection scope of the present invention.

Claims

1. A platelet demand prediction system based on time series, characterized in that: It includes data acquisition and regularization module, time series adjustment module, time series decomposition module, subsequence modeling and integration module, model evaluation module and model deployment module, among which: S1, data acquisition and regularization module: integrates clinical platelet supply data, including data reading, format regularization, and descriptive statistics functions, and outputs the integrated data to the time series adjustment module to obtain time series data; S2, time series processing module: outlier processing is performed on the received time series data, the required difference order is identified to meet the stationarity, and the Box-Cox transformation is performed to meet the normality; the adjusted time series is decomposed using the X-13ARIMA-SEATS method to obtain trend subsequences, seasonal subsequences and residual subsequences. The decomposed time series enters the rolling window and is rolled out to the subsequence modeling and integration module batch by batch and year by year to obtain the subsequences within the window; S3, subsequence modeling and integration module: fits ARIMA, Prophet, TimeGPT, and SNAIVE models to the subsequence fitting models received in the window for subsequence prediction, and recombines the prediction results based on multiplicative decomposition to obtain the one-year prediction value for the entire sequence; executes this module in a rolling window loop to obtain all the prediction values ​​of the specified modeling method combination over the specified time span, and outputs them to the model evaluation module; S4, model evaluation module: calculates the MAPE and RMSE accuracy indicators for the received prediction values, improves the accuracy of recent predictions by exponential weighting, and thus calculates the weighted MAPE and weighted RMSE indicators to evaluate the overall performance of the prediction system and determine whether it can be deployed; S5, model deployment module: use Shiny to encapsulate the model that has passed the model evaluation into an interactive network application, so as to realize the intelligent and graphical blood supply prediction work and provide auxiliary decision-making.

2. The platelet demand prediction system based on time series according to claim 1, characterized in that: Collect target data by digital means; pre-design data structure and data regularization scripts; store data in the tsibble time series data frame format of R language.

3. The platelet demand prediction system based on time series according to claim 1, characterized in that: The KPSS stationarity test is used to determine the minimum number of differences required for series stationarization. The Box-Cox lambda value is automatically selected through an iterative algorithm that automatically selects the Box-Cox transformation parameters, thereby achieving full automation of the time series adjustment process. The SEATS method is applied in X-13ARIMA-SEATS for seasonal adjustment, and Chinese holidays are introduced as exogenous regressor variables.

Citation Information

Patent Citations

  • Parameter-adaptive electricity consumption prediction method and system

    CN113298308A