Medical data analysis and management method and system
By judging stationarity in medical data analysis, determining the difference order, autoregression and sliding average order, and building an ARIMA model, the problem of insufficient accuracy of drug sales data analysis model in traditional methods is solved, and high-precision drug sales forecasting and inventory management are achieved.
Patent Information
- Application Number
- CN202510359417.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-06-17
AI Technical Summary
Traditional medical data analysis methods ignore the problem of data stationarity when processing drug sales data, resulting in poor accuracy of the analysis model, unable to effectively capture the real trends and rules behind the data, and it is difficult to meet the high-precision requirements of drug sales forecasting and inventory management.
By judging the stationaryness of the original drug sales data, determine the target differential order, and stabilize the non-stationary data; determine the target autoregression and sliding average order based on the stationary target time series data, characterize the dynamic characteristics of the data; perform parameter estimation to obtain the autoregression and sliding average coefficients, and construct an ARIMA model to analyze and manage medical data.
It has achieved effective and stable drug sales data and dynamic characterization, improved the model's ability to capture data laws, improved the accuracy of drug sales forecasts, assisted pharmaceutical companies to reasonably arrange production and inventory, reduce costs, and helped medical institutions to make preparations for drug procurement and reserves in advance to ensure patients' drug use needs.
Smart Images

Figure CN120164594A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure belongs to the technical field of data analysis and management, and more specifically, relates to a method and system for pharmaceutical data analysis and management. Background Art
[0002] In the pharmaceutical industry, accurate data analysis and management are crucial for the operation decisions of enterprises and medical institutions. Currently, with the rapid development of the pharmaceutical market, the scale of drug sales data is becoming increasingly large and complex. Traditional pharmaceutical data analysis methods often ignore the stationarity of data when processing these drug sales data, resulting in poor accuracy of the analysis models constructed based on non-stationary data, being unable to effectively capture the true trends and laws behind the data, and being difficult to meet the high-precision requirements for drug sales prediction, inventory management, etc. in practical applications. Summary of the Invention
[0003] The purpose of the present disclosure is to provide a method and system for pharmaceutical data analysis and management to improve the accuracy of drug sales prediction and inventory management.
[0004] In the first aspect of the embodiments of the present disclosure, a method for pharmaceutical data analysis and management is provided, including: Determining the target difference order of the target model based on the stationarity of the original time series data, where the original time series data is drug sales data arranged in chronological order within a historical time span, and the stationarity includes stationary and non-stationary; Determining the target autoregressive order and the target moving average order of the target model based on the target correlation function of the target time series data, where the target time series data is time series data with stationary stationarity; Performing parameter estimation based on the target difference order, the target autoregressive order, and the target moving average order to obtain the target parameters of the target model, where the parameters of the target model include autoregressive coefficients and moving average coefficients; Constructing the target model based on the target parameters and analyzing and managing pharmaceutical data based on the target model.
[0005] In the second aspect of the embodiments of the present disclosure, a system for pharmaceutical data analysis and management is provided, including: A first order determination module for determining the target difference order of the target model based on the stationarity of the original time series data, where the original time series data is drug sales data arranged in chronological order within a historical time span, and the stationarity includes stationary and non-stationary; A second order determination module for determining the target autoregressive order and the target moving average order of the target model based on the target correlation function of the target time series data, where the target time series data is time series data with stationary stationarity; A model parameter determination module for performing parameter estimation based on the target difference order, the target autoregressive order, and the target moving average order to obtain the target parameters of the target model; An analysis and management module for constructing a target model based on the target parameters and analyzing and managing medical data based on the target model.
[0006] In a third aspect of the embodiments of the present disclosure, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of the above-mentioned medical data analysis and management method are implemented.
[0007] In a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned medical data analysis and management method are implemented.
[0008] The beneficial effects of the medical data analysis and management method and system provided by the embodiments of the present disclosure are as follows: By determining the target difference order by judging the stationarity of the original drug sales data, the non-stationary data can be made stationary, ensuring that the model can effectively process various actual sales data. Secondly, based on the stationary target time series data, the target autoregressive and moving average orders are determined to characterize the dynamic characteristics of the data and improve the model's ability to capture data patterns. Then, parameter estimation is performed to obtain the autoregressive and moving average coefficients, further improving the model to make it fit the actual sales situation. Finally, the target model constructed based on this predicts the drug sales trend, assisting pharmaceutical enterprises to reasonably arrange production and inventory, reducing costs; it can also help medical institutions to make advance preparations for drug procurement and reserves to ensure the drug demand of patients, thereby realizing the efficient analysis and management of medical data. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0010] Figure 1 It is a flowchart of the medical data analysis and management method provided by an embodiment of the present disclosure; Figure 2 It is a structural block diagram of the medical data analysis and management system provided by an embodiment of the present disclosure; Figure 3 It is a schematic block diagram of the electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0011] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present disclosure. However, those skilled in the art should clearly understand that the present disclosure can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present disclosure.
[0012] To make the objectives, technical solutions, and advantages of the present disclosure clearer, the following will be described through specific embodiments with reference to the accompanying drawings.
[0013] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a medical data analysis and management method provided for an embodiment of the present disclosure. The method includes: S101: Determine the target difference order of the target model based on the stationarity of the original time series data. The original time series data is drug sales data arranged in chronological order within a historical time span. The stationarity includes stationary and non-stationary.
[0014] In the medical field, the original time series data is drug sales data arranged in chronological order within a historical time span. The original time series data may be affected by various factors, such as historical data, seasonal changes (for example, some diseases are highly prevalent in specific seasons, resulting in fluctuations in the sales volume of related drugs), etc., making the data have non-stationary characteristics.
[0015] In this embodiment, the stationary time series data does not change over time. Unit root tests, KPSS tests, etc. can be used to judge the stationarity of the original time series data. If the original time series data does not meet the stationarity condition, then the data needs to be differenced.
[0016] For example, after performing a first-order difference on the drug sales data of a certain drug, the ADF test still shows non-stationarity. After performing a second-order difference, the test passes, so the target difference order is 2. If the original data is itself stationary, then the target difference order is 0.
[0017] S102: Determine the target autoregressive order and the target moving average order of the target model based on the target correlation function of the target time series data. The target time series data is time series data with stationary stationarity.
[0018] In this embodiment, the target model can be an Autoregressive Integrated Moving Average (ARMA) model. The target time series data is time series data that has been determined to be stationary. When constructing an ARIMA model, stationarity is a basic requirement, and the ARIMA model can work effectively on stationary data.
[0019] The target correlation functions can be the Autocorrelation Function (ACF) and the Partial Autocorrelation Function (PACF). The ACF represents the correlation between observations at different lag orders in a time series, and the PACF represents the direct correlation between two observations after controlling for the effects of intermediate lag orders. By plotting the ACF and PACF diagrams of the target time series data, their truncation and trailing characteristics can be observed.
[0020] The target autoregressive order and the target moving average order can be determined based on the characteristics of the ACF and PACF diagrams.
[0021] S103: Perform parameter estimation based on the target differencing order, the target autoregressive order, and the target moving average order to obtain the target parameters of the target model. The parameters of the target model include autoregressive coefficients and moving average coefficients.
[0022] In this embodiment, let the target differencing order be denoted as d, the target autoregressive order be denoted as p, and the target moving average order be denoted as q.
[0023] After determining the target differencing order d, the target autoregressive order p, and the target moving average order q, an Autoregressive Integrated Moving Average ARIMA model can be constructed. The mathematical expression of the ARIMA model is:
[0024] Where, ( represents the autoregressive coefficient), ( represents the moving average coefficient), represents the differencing operator, represents, represents the original time series data, represents the white noise error term.
[0025] In this embodiment, the maximum likelihood estimation method can be used for parameter estimation.
[0026] For example, given an ARIMA model structure and observed data (target time series data), find a set of autoregressive coefficients and moving average coefficients such that the probability of the observed data occurring is maximized. Solve for the parameter values that maximize the likelihood function through the Newton-Raphson algorithm, and the maximized parameter values are the target parameters of the target model.
[0027] S104: Construct a target model based on the target parameters, and analyze and manage medical data based on the target model.
[0028] In this embodiment, based on the autoregressive coefficients, moving average coefficients, and the determined target differencing order, target autoregressive order, and target moving average order, a complete ARIMA model (target model) is constructed. The target model can capture the internal laws and changing trends of drug sales data.
[0029] In the analysis and management of medical data, the target model is used for prediction and decision support. For example, by inputting historical drug sales data into the model, future drug sales volumes can be predicted. For pharmaceutical companies, this can reasonably arrange production plans, avoid inventory backlogs or shortages, and reduce costs; for medical institutions, it can prepare drug purchases and reserves in advance to ensure patients' medication needs. In addition, the target model can also be used to analyze the impact of different factors on drug sales, helping enterprises and institutions formulate more effective marketing strategies and management decisions, such as adjusting prices and promoting key products.
[0030] As can be seen from the above, in this embodiment, the target differencing order is determined by judging the stationarity of the original drug sales data, which can make non-stationary data stationary and ensure that the model can effectively process various actual sales data. Secondly, based on the stationary target time series data, the target autoregressive and moving average orders are determined to characterize the dynamic characteristics of the data and improve the model's ability to capture data patterns. Then, parameter estimation is performed to obtain autoregressive and moving average coefficients, further improving the model to make it fit the actual sales situation. Finally, the target model constructed based on this predicts the drug sales trend, assisting pharmaceutical companies in reasonably arranging production and inventory, reducing costs; and also helping medical institutions prepare drug purchases and reserves in advance to ensure patients' medication needs, thereby achieving the efficient analysis and management of medical data.
[0031] In an embodiment of the present disclosure, it further includes:[[]] Perform windowing processing on the first time series data to obtain multiple target subsequences, where the first time series data is the time series data obtained by normalizing the original time series data; Extract features from the multiple target subsequences to obtain multiple feature vectors corresponding to the multiple target subsequences; Classification analysis is performed on multiple eigenvectors to obtain the stationarity of the original time series data.
[0032] In this embodiment, the original time series data is the drug sales data arranged in chronological order within the historical time span, and these data have different dimensions and value ranges. Normalization is to map the data to the interval [0,1], so that each data point has the same influence in the model.
[0033] The normalized first time series data is subjected to window processing, that is, the entire time series is divided into multiple target subsequences of fixed length. The purpose of window processing is to convert the continuous time series into multiple independent samples. Each target subsequence can be regarded as a local time series fragment, which contains the information of the original data in that time period.
[0034] Perform feature extraction on each target subsequence, extract key information that can represent its characteristics and regularities from each subsequence, and convert this information into a feature vector. A feature vector is a multidimensional vector, in which each dimension represents a specific feature of the subsequence.
[0035] The autocorrelation coefficient can be used to extract features from multiple target subsequences to obtain the distribution characteristics, fluctuations, and correlation information between data points of the subsequences. By extracting features from each target subsequence, the corresponding feature vector can be obtained.
[0036] The purpose of classifying and analyzing multiple eigenvectors is to determine the stationarity of the original time series data based on the characteristics and rules of these eigenvectors.
[0037] From the above, it can be concluded that this embodiment can refine complex data by windowing the normalized original drug sales time series. Feature extraction can mine the key features of each subsequence and convert them into feature vectors. Classification analysis can judge the stability based on this and improve the accuracy of judgment.
[0038] In one embodiment of the present disclosure, determining the target difference order of the target model based on the stationarity of the original time series data includes: In response to the stationarity of the original time series data being stationary, the target difference order of the target model is zero; In response to the stationarity of the original time series data being non-stationary, the original time series data is differentiated until a first condition is satisfied, and a target differential order of the target model is determined; The first condition is that the variance of the target time series data is less than the preset variance, and the target time series data is the time series data obtained by performing any differential operation on the original time series data.
[0039] In this embodiment, the Augmented Dickey-Fuller (ADF) test can be used to determine the stationarity of the original time series data. Stationarity means that the statistical properties (such as mean, variance, and autocovariance) of the time series do not change over time. If the statistical properties of the original time series data meet the stationarity requirements, it is determined to be stationary.
[0040] In this embodiment, when the stationarity of the original time series data is stationary, it indicates that the original time series data already meets the basic requirements of stationarity for modeling, and no differencing operation is required to make it stationary. Therefore, the target differencing order of the target model is directly determined to be zero. The purpose of the differencing operation is to transform a non-stationary sequence into a stationary sequence. If the original time series data is already stationary, there is no need to perform differencing again.
[0041] In this embodiment, when the stationarity of the original time series data is non-stationary, in order to make the data meet the modeling requirements, a differencing operation needs to be performed on it. The differencing operation eliminates non-stationary factors such as trends and seasonality in the data by calculating the differences between adjacent data points.
[0042] After performing a first-order differencing on the original time series data, the first time series data is obtained.
[0043] Assume the original time series data is , and perform a first-order differencing. The calculation method of the first-order differencing is:
[0044] where represents the first time series data.
[0045] Then calculate the variance of the first time series data and compare it with a preset variance. The preset variance is a threshold set according to the actual situation and experience, and is used to measure the degree of fluctuation of the time series data.
[0046] If the variance of the first time series data is greater than the preset variance, it means that the first time series data still has large fluctuations and the non-stationarity has not been completely eliminated. It is necessary to perform differencing on the target time series data again to obtain the second time series data, and repeat the above steps of calculating the variance and comparing.
[0047] The second-order differencing is to perform a differencing operation on the sequence after the first-order differencing again, that is:
[0048] where represents the second time series data.
[0049] Repeat this process until the variance of the second time series data is less than the preset variance, i.e., the first condition is satisfied. The second time series data that satisfies the first condition is used as the target time series data. At this time, the number of differencing times performed is the target differencing order of the target model. Because when the variance is less than the preset variance, it indicates that the fluctuation degree of the time series data has been reduced to an acceptable range, and the stationarity of the data has been improved. It can be considered that the data at this time meets the modeling requirements.
[0050] From the above, it can be concluded that in this embodiment, the target differencing order of the target model is reasonably determined according to the stationarity of the original time series data, which provides a basis for constructing the target model and ensures that the target model can effectively process and analyze drug sales data.
[0051] In an embodiment of the present disclosure, determining the target autoregressive order and the target moving average order of the target model based on the target correlation function of the target time series data includes: Determining the target autoregressive order and the target moving average order of the target model based on the truncation and tailing of the target correlation function.
[0052] In this embodiment, the target correlation function includes ACF and PACF.
[0053] ACF is the correlation between observations at different lag orders in a time series. For example, for an original time series data , ACF(k) represents and The correlation between them, where k is the lag order.
[0054] After PACF controls the influence of intermediate lag orders, it measures the direct correlation between two observations. That is, PACF(k) represents after excluding The influence of these intermediate terms, and The correlation between them.
[0055] If the PACF of the target time series data truncates after the p-th order (i.e., the partial autocorrelation coefficient rapidly approaches 0 after the p-th order), and the ACF tails (the autocorrelation coefficient gradually decays), then this time series conforms to the Autoregressive (AR) model, and the target autoregressive order is p. This is because in the AR model, Only has a direct linear relationship with its first p-period values After exceeding the p-th order, the partial autocorrelation coefficient should theoretically be 0, so it shows truncation.
[0056] If the ACF of the target time series data is truncated after the q-th order and the PACF has a trailing tail, then the time series may conform to the Moving Average (MA) model, and the target moving average order is q. In the MA model, is a linear combination of the current and the previous q-period white noise error terms. After exceeding the q-th order, the autocorrelation coefficient should theoretically be 0, thus showing a truncated characteristic.
[0057] In practical applications, by observing the truncated and trailing characteristics of the target correlation function and combining information criteria, the target autoregressive order p and the target moving average order q can be more accurately determined, and then a suitable ARIMA model can be constructed to analyze and manage medical data.
[0058] It can be concluded from the above that this embodiment provides an effective method for accurately determining the autoregressive order and the moving average order of the target model by utilizing the truncated and trailing characteristics of the target correlation function, which helps to construct a time series model that better fits the actual data.
[0059] In an embodiment of the present disclosure, determining the target autoregressive order and the target moving average order of the target model based on the truncated and trailing of the target correlation function includes: Determining the initial autoregressive order and the initial moving average order of the target model based on the truncated and trailing of the target correlation function; Constructing an order matrix based on the initial autoregressive order and the initial moving average order; Constructing multiple initial models based on the order matrix; Based on that the fitting of each initial model satisfies the second condition, taking the initial autoregressive order and the initial moving average order corresponding to the initial model that satisfies the second condition as the target autoregressive order and the target moving average order of the target model, and the second condition is the initial model corresponding to the fitting greater than the preset fitting.
[0060] In this embodiment, for the stationary target time series data, the truncated and trailing characteristics of its ACF and PACF are used to initially determine the initial autoregressive order and the initial moving average order of the target model.
[0061] If the PACF is truncated after orders and the ACF has a trailing tail, it can be initially judged that the initial autoregressive order is .
[0062] If the ACF is truncated after orders and the PACF has a trailing tail, then the initial moving average order is initially determined to be .
[0063] Based on the obtained initial autoregressive order and the initial moving average order Construct an order matrix. The order matrix can consider taking different combinations of p and q values within a certain range centered around the initial order.
[0064] For example, it can be considered that p takes values within , and q takes values within (a and b are integers determined according to experience or actual situation), so as to obtain a series of different (p, q) combinations and form an order matrix.
[0065] Construct multiple initial models according to each group of (p, q) combinations in the order matrix. Each initial model is an ARIMA(p, d, q) model, where d is the target difference order previously determined according to the stationarity of the original time series data. Constructing multiple initial models through different (p, q) combinations is to search for the model order that best fits the data within a wider range.
[0066] For each initial model, evaluate its goodness of fit. The goodness of fit can be measured by the mean squared error, Akaike information criterion, etc. The second condition is that the initial model whose goodness of fit is greater than the preset goodness of fit, that is, select those models that perform better than the preset threshold in statistical indicators.
[0067] The preset goodness of fit is a standard set according to actual needs and experience, used to judge whether the fitting degree of the model to the data is good enough. When the goodness-of-fit index of a certain initial model meets the second condition, the corresponding initial autoregressive order and initial moving average order of the model are used as the target autoregressive order and target moving average order of the target model.
[0068] It can be concluded from the above that in this embodiment, the order range is initially determined by using the target correlation function, and then by constructing multiple models and screening, the autoregressive order and moving average order suitable for the target time series data can be determined more accurately, so as to construct a more accurate medical data analysis model and provide more reliable support for the analysis and management of medical data.
[0069] In an embodiment of the present disclosure, it further includes: Calculate the goodness of fit of each initial model based on the first formula; The first formula is:
[0070] Wherein, represents the goodness of fit of any initial model, represents the likelihood function value of any initial model, represents the number of initial models, represents the types of different drugs in the medical data, represents the The standard deviation corresponding to each type of drug, represents the quantity of the original time series data.
[0071] In this embodiment, is used to measure the balance between the goodness of fit of the model to the data and the complexity. In model selection, it is desired to find a model with the smallest value, which means that while fitting the data, the complexity of the model is relatively low, avoiding overfitting.
[0072] represents the likelihood function value of any initial model, which is the likelihood function value of any initial model. The likelihood function measures the probability of the observed data occurring given the model parameters. The larger the likelihood function value, the better the model fits the data.
[0073] represents the number of initial models. During the calculation process, reflects the complexity of the model. The more parameters the model contains, the higher the complexity, for the greater the contribution, reflecting the penalty mechanism for the model complexity.
[0074] refers to the types of different drugs in the medical data. The medical data involves the sales situations of various different drugs, and the data characteristics of each drug may vary. By considering the types of different drugs, the model can more comprehensively consider the diversity of the data when evaluating the fitness.
[0075] Standard deviation reflects the degree of fluctuation of the data of this type of drug. Data with larger fluctuations may be more difficult to fit. By incorporating the standard deviation into the formula, the model can take into account the stability of the data when evaluating the fitness.
[0076] is the penalty term for the model complexity. As the model complexity k increases, this part of the value will increase; at the same time, the sample size n will also affect the degree of penalty; in addition, the average value of the sum of the standard deviations of the data of different drugs will increase the penalty term, meaning that the greater the data fluctuation, the heavier the penalty for the model complexity.
[0077] It can be concluded from the above that in this embodiment, by calculating the value of each initial model, we can compare the fitness of all initial models and select the autoregressive order and the moving average order corresponding to the model with the smallest value as the order of the target model, thereby constructing a more suitable medical data analysis model.
[0078] In one embodiment of the present disclosure, parameter estimation is performed based on the target difference order, the target autoregressive order, and the target moving average order to obtain the target parameters of the target model, including: Calculating the weight of each data point in the original time series data based on the time distance and the degree of data fluctuation of the original time series data to obtain time series weight data; Constructing a weighted likelihood function based on the time series weight data; Obtaining the target parameters of the target model based on the weighted likelihood function.
[0079] In this embodiment, the original time series data is drug sales data arranged in chronological order within a historical time span. The time distance is based on a reasonable assumption that recent data often has a greater influence and reference value on the current situation. For example, when predicting future drug sales trends, the sales data of the recent few months may be more reflective of the current market demand and changes than data from several years ago. Therefore, a time reference point (such as the current time) is set, and by calculating the time distance between each data point and the reference point, a suitable function (such as an exponential decay function) is used to determine the time distance weight, so that data points closer to the reference point obtain higher weights.
[0080] The degree of data fluctuation reflects the stability and reliability of the data. Data points with larger fluctuations contain more noise or uncertainties and have a relatively smaller impact on parameter estimation; while data points with smaller fluctuations can better reflect the internal laws of the data and should be given higher weights. The degree of fluctuation of each data point is measured by calculating the moving standard deviation of the data, and then a corresponding function (such as an inverse proportional function) is used to determine the degree of fluctuation weight, that is, the greater the degree of fluctuation, the lower the weight.
[0081] The time distance weight and the degree of fluctuation weight are combined (such as multiplied) to obtain time series weight data.
[0082] In this embodiment, the likelihood function measures the probability of the occurrence of observed data under the condition of known model structure (ARIMA model structure determined by the target difference order, the target autoregressive order, and the target moving average order) and parameters. For time series data, the traditional likelihood function assumes that the contribution of each data point to the model is the same.
[0083] In actual situations, since the importance of different data points is different (reflected by the weights calculated previously), it is necessary to construct a weighted likelihood function. Specifically, the conditional probability density function of each data point is adjusted according to its weight to construct a weighted likelihood function
[0084] Wherein, is the conditional probability density function of the t-th data point, is the autoregressive coefficient vector, is the moving average coefficient vector, is the weight of the t-th data point. In this way, the weighted likelihood function can more accurately reflect the actual situation of the data, highlight the role of important data points, and reduce the influence of noise data points.
[0085] The maximum likelihood estimation method can be used for parameter estimation. The goal is to find a set of parameters (autoregressive coefficients and moving average coefficients) that maximize the value of the likelihood function, that is, to maximize the probability of the observed data occurring. For the weighted likelihood function, the Newton-Raphson algorithm is also used to find the parameter values that maximize the weighted likelihood function.
[0086] During the optimization process, the values of the autoregressive coefficients and moving average coefficients are continuously adjusted, and the value of the weighted likelihood function is calculated until a parameter combination that maximizes the weighted likelihood function is found. This set of parameters is the target parameters of the target model, which can make the constructed ARIMA model better fit the original time series data, thereby providing more accurate model support for subsequent medical data analysis and management.
[0087] It can be concluded from the above that in this embodiment, the weights are determined according to the time distance and fluctuation degree of the data points, the weighted likelihood function is constructed and parameter estimation is carried out, which can make more effective use of the information of the time series data and improve the accuracy and reliability of the target model parameter estimation.
[0088] Corresponding to the medical data analysis and management method in the above embodiment, Figure 2 is the structural block diagram of the medical data analysis and management system provided by an embodiment of the present disclosure. For the sake of convenience of description, only the parts related to the embodiments of the present disclosure are shown. Refer to Figure 2 The medical data analysis and management system 20 includes: a first order determination module 21, a second order determination module 22, a model parameter determination module 23, and an analysis and management module 24. Among them, the first order determination module 21 is used to determine the target difference order of the target model based on the stationarity of the original time series data. The original time series data is the drug sales data arranged in chronological order within the historical time span, and the stationarity includes stationary and non-stationary; The second order determination module 22 is used to determine the target autoregressive order and the target moving average order of the target model based on the target correlation function of the target time series data. The target time series data is the time series data with stationary stationarity; The model parameter determination module 23 is used to perform parameter estimation based on the target difference order, the target autoregressive order, and the target moving average order to obtain the target parameters of the target model; The analysis and management module 24 is used to construct a target model based on target parameters and analyze and manage medical data based on the target model.
[0089] In an embodiment of the present disclosure, the medical data analysis and management system 20 further includes: a stationarity analysis module, specifically used for: Performing windowing processing on the first time series data to obtain a plurality of target subsequences, where the first time series data is the time series data obtained by normalizing the original time series data; Performing feature extraction on the plurality of target subsequences to obtain a plurality of feature vectors corresponding to the plurality of target subsequences; Performing classification analysis on the plurality of feature vectors to obtain the stationarity of the original time series data.
[0090] In an embodiment of the present disclosure, the first order determination module 21 is specifically used for: In response to the stationarity of the original time series data being stationary, the target difference order of the target model is zero; In response to the stationarity of the original time series data being non-stationary, performing differencing on the original time series data until a first condition is satisfied, and determining the target difference order of the target model; The first condition is that the variance of the target time series data is less than a preset variance, and the target time series data is the time series data obtained by performing any differencing on the original time series data.
[0091] In an embodiment of the present disclosure, the second order determination module 22 is specifically used for: Determining the target autoregressive order and the target moving average order of the target model based on the truncation and trailing of the target correlation function.
[0092] In an embodiment of the present disclosure, the second order determination module 22 is specifically further used for: Determining the initial autoregressive order and the initial moving average order of the target model based on the truncation and trailing of the target correlation function; Constructing an order matrix based on the initial autoregressive order and the initial moving average order; Constructing a plurality of initial models based on the order matrix; Based on the fitting of each initial model satisfying a second condition, taking the initial autoregressive order and the initial moving average order corresponding to the initial model that satisfies the second condition as the target autoregressive order and the target moving average order of the target model, and the second condition is the initial model corresponding to the fitting being greater than the preset fitting.
[0093] In an embodiment of the present disclosure, the second order determination module 22 is specifically further used for: Calculating the fitting of each initial model based on the first formula; The first formula is as follows:
[0094] Wherein, represents the fitness of any initial model, represents the likelihood function value of any initial model, represents the number of initial models, represents the types of different drugs in the medical data, represents the th standard deviation corresponding to the drugs of the th type, and
[0095] In an embodiment of the present disclosure, the model parameter determination module 23 is specifically configured to: calculate the weight of each data point in the original time series data based on the time distance and the degree of data fluctuation of the original time series data to obtain time series weight data; construct a weighted likelihood function based on the time series weight data; and obtain the target parameters of the target model based on the weighted likelihood function. Based on the time series weight data, construct a weighted likelihood function; Based on the weighted likelihood function, obtain the target parameters of the target model.
[0096] Refer to Figure 3 Figure 3 which is a schematic block diagram of an electronic device provided in an embodiment of the present disclosure. As shown in Figure 3 the electronic device 300 in this embodiment may include: one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The above-mentioned processors 301, input devices 302, output devices 303, and memories 304 communicate with each other through a communication bus 305. The memory 304 is used to store a computer program, and the computer program includes program instructions. The processor 301 is used to execute the program instructions stored in the memory 304. Among them, the processor 301 is configured to call the program instructions to execute the functions of each module / unit in the above-mentioned system embodiments, such as Figure 2 the functions of the modules 21 to 24 shown in
[0097] It should be understood that in the embodiments of the present disclosure, the so-called processor 301 may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0098] The input device 302 may include a touchpad, a fingerprint acquisition sensor (for acquiring the fingerprint information and the direction information of the fingerprint of the user), a microphone, etc., and the output device 303 may include a display (such as an LCD), a speaker, etc.
[0099] The memory 304 may include a read-only memory and a random access memory, and provide instructions and data to the processor 301. A part of the memory 304 may also include a non-volatile random access memory. For example, the memory 304 may also store information about the device type.
[0100] In specific implementation, the processor 301, the input device 302, and the output device 303 described in the embodiments of the present disclosure may implement the implementation manners described in the first and second embodiments of the medical data analysis and management method provided by the embodiments of the present disclosure, and may also implement the implementation manner of the electronic device described in the embodiments of the present disclosure, which will not be elaborated herein.
[0101] In another embodiment of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, all or part of the processes in the method of the above embodiment are implemented. It can also be completed by instructing relevant hardware through the computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of the above various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or system capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0102] The computer-readable storage medium can be the internal storage unit of the electronic device in any of the foregoing embodiments, such as the hard disk or memory of the electronic device. The computer-readable storage medium can also be an external storage device of the electronic device, such as a plug-in hard disk equipped on the electronic device, a smart media card (Smart Media Card, SMC), a secure digital (SecureDigital, SD) card, a flash card (Flash Card), etc. Further, the computer-readable storage medium can also include both the internal storage unit and the external storage device of the electronic device. The computer-readable storage medium is used to store the computer program and other programs and data required by the electronic device. The computer-readable storage medium can also be used to temporarily store the data that has been output or will be output.
[0103] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present disclosure.
[0104] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described electronic devices and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0105] In several embodiments provided by the present application, it should be understood that the disclosed electronic devices and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed coupling or direct coupling or communication connection between each other can be an indirect coupling or communication connection through some interfaces or units, and can also be an electrical, mechanical or other form of connection.
[0106] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of the present disclosure.
[0107] In addition, each functional unit in various embodiments of the present disclosure can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0108] The above is only the specific implementation manner of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present disclosure can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.
Claims
1. A medical data analysis and management method, characterized in that: include: Determining a target difference order of a target model based on the stationarity of original time series data, wherein the original time series data is drug sales data arranged in chronological order within a historical time span, and the stationarity includes stationarity and non-stationarity; Determining a target autoregressive order and a target sliding average order of a target model based on a target correlation function of target time series data, wherein the target time series data is time series data with stationarity; Perform parameter estimation based on the target difference order, the target autoregressive order and the target sliding average order to obtain target parameters of the target model, where the parameters of the target model include autoregressive coefficients and sliding average coefficients; A target model is constructed based on the target parameters, and medical data is analyzed and managed based on the target model.
2. The medical data analysis and management method according to claim 1, characterized in that: Also includes: Performing window processing based on first time series data to obtain multiple target subsequences, wherein the first time series data is time series data after normalization processing of original time series data; Performing feature extraction on the multiple target subsequences to obtain multiple feature vectors corresponding to the multiple target subsequences; The multiple feature vectors are classified and analyzed to obtain the stationarity of the original time series data.
3. The medical data analysis and management method according to claim 1, wherein: The step of determining the target difference order of the target model based on the stationarity of the original time series data includes: In response to the stationarity of the original time series data being stationary, the target difference order of the target model is zero; In response to the stationarity of the original time series data being non-stationary, the original time series data is differentiated until a first condition is satisfied, and a target differential order of a target model is determined; The first condition is that the variance of the target time series data is less than a preset variance, and the target time series data is time series data obtained by performing any differential operation on the original time series data.
4. The medical data analysis and management method according to claim 1, wherein: The target correlation function based on the target time series data determines the target autoregressive order and the target sliding average order of the target model, including: The target autoregressive order and target moving average order of the target model are determined based on the truncation and tailing of the target correlation function.
5. The medical data analysis and management method according to claim 4, characterized in that: The method of determining the target autoregressive order and the target sliding average order of the target model based on the truncation and tailing of the target correlation function includes: Determine the initial autoregressive order and initial sliding average order of the target model based on the truncation and tailing of the target correlation function; Constructing an order matrix based on the initial autoregressive order and the initial sliding average order; constructing a plurality of initial models based on the order matrix; Based on the fact that the fit of each initial model satisfies the second condition, the initial autoregressive order and initial sliding average order corresponding to the initial model that meets the second condition are used as the target autoregressive order and target sliding average order of the target model, and the second condition is that the initial model corresponding to the preset fit has a fit greater than that of the initial model.
6. The medical data analysis and management method according to claim 5, characterized in that: Also includes: Calculate the goodness of fit of each initial model based on the first formula; The first formula is: in, represents the fit of any initial model, represents the likelihood function value of any initial model, represents the number of initial models, Represents the types of different drugs in medical data, Indicates The standard deviation corresponding to each type of drug, Represents the number of original time series data.
7. The medical data analysis and management method according to claim 1, wherein: The performing parameter estimation based on the target difference order, the target autoregressive order and the target sliding average order to obtain the target parameters of the target model includes: The weight of each data point in the original time series data is calculated based on the time distance of the original time series data and the degree of fluctuation of the data to obtain the time series weight data; Constructing a weighted likelihood function based on the time series weight data; The target parameters of the target model are obtained based on the weighted likelihood function.
8. A medical data analysis and management system, characterized in that: include: A first order determination module is used to determine a target differential order of a target model based on the stationarity of original time series data, wherein the original time series data is drug sales data arranged in chronological order within a historical time span, and the stationarity includes stationarity and non-stationarity; A second order determination module is used to determine a target autoregressive order and a target sliding average order of a target model based on a target correlation function of target time series data, wherein the target time series data is time series data with stationarity being stationary; A model parameter determination module, used to perform parameter estimation based on the target difference order, the target autoregressive order and the target sliding average order to obtain target parameters of the target model; The analysis and management module is used to construct a target model based on the target parameters and analyze and manage the medical data based on the target model.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.