Intermittent time series prediction method and system oriented to price-quantity relationship and application of intermittent time series prediction method and system
By introducing the binary mixed distribution model MOD-DeepVAR in time series prediction, the problem of difficulty in capturing multimodal characteristics and nonlinear relationships in the prior art is solved, and higher prediction accuracy and stability are achieved.
Patent Information
- Application Number
- CN202510273146.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-27
AI Technical Summary
When processing complex multivariate time series data, it is difficult to capture the multimodal characteristics and nonlinear relationships of the data, resulting in limited accuracy and reliability of prediction results.
A binary mixed distribution model MOD-DeepVAR based on quantity-price relationship is proposed. A potential subdistribution is generated through multiple RNN networks, and a mixed Gaussian distribution is generated through a fully connected layer. Combined with negative log likelihood loss function and L2 regularization, the model parameters are optimized to improve prediction accuracy.
Effectively capturing the multimodal characteristics and nonlinear relationships of data improves the accuracy and stability of intermittent time series prediction, especially when the system changes significantly or the external environment fluctuates greatly, it can better adapt to market changes.
Smart Images

Figure BDA0005303322300000031 
Figure BDA0005303322300000041 
Figure BDA0005303322300000043
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of time series prediction, and relates to an intermittent time series prediction method, system and application for vector-price relationship. Background Art
[0002] In practical applications such as financial markets and sales forecasting, there are usually complex interactions between prices and sales volumes. Traditional intermittent time series prediction methods often only focus on single-variable modeling, ignoring the mutual influence between variables, and unable to fully capture the non-linear relationships and potential patterns in the dynamic changes of the market, thus limiting the accuracy and reliability of prediction results. By jointly modeling sales volume and price, the internal connection between price changes and sales volume can be revealed, the dynamic changes of the market can be captured more accurately, and thus the accuracy of prediction can be improved.
[0003] The mixture distribution model is a probability model that describes the situation where data is generated by multiple different probability distributions. For data with multi-modal characteristics, a single distribution often has difficulty accurately fitting the overall data structure. In this case, the data can be regarded as composed of multiple sub-distributions, each sub-distribution having a different weight. The mixture distribution model can effectively capture the multi-modal characteristics of the data and provide more flexible fitting ability by combining multiple sub-distributions. In the bivariate time series data of the price-volume relationship, the mixture distribution model can simultaneously fit the multi-modal characteristics of sales volume and price, and identify potential states or different market mechanisms in the data.
[0004] For traditional multi-variable time series prediction tasks, the DeepVAR model performs well. This model is a deep learning-based intermittent time series prediction method that uses an autoregressive method to recursively generate predicted values for future time points. However, in the original DeepVAR model, the data is assumed to follow a single distribution (such as a Gaussian distribution), which may be insufficiently flexible when dealing with actual multi-objective variable data. Real-world time series data often exhibits complex multi-modal characteristics, especially when the system dynamics change significantly or the external environment fluctuates greatly, and obvious time series characteristics are also shown. For example, intermittent time series often show that the values in some time periods are zero or extremely low, while peaks appear at other time points. The multi-modal characteristics of this time series data make the single distribution assumption unable to accurately fit its characteristics, because at different time points, the data may follow different distributions. The DeepVAR model has limitations in dealing with data sets with multi-modal forms. Summary of the Invention
[0005] To address the deficiencies of the existing technologies, the objective of the present invention is to provide an intermittent time series prediction method oriented to volume-price relationships. The prediction method is based on a binary mixture distribution model of volume-price relationships and can effectively handle the intermittent demand prediction tasks of commodities. By jointly modeling prices and sales volumes and using a mixture distribution model to capture the multi-modal characteristics of the data. The core idea of the implementation process of the prediction method of the present invention is as follows: Input the binary time series data into multiple RNN networks, and each RNN network is responsible for generating a latent distribution. Then, map the hidden states of each RNN to the parameters of the sub-distributions through a fully connected layer, including the mean and covariance matrix of each sub-distribution. After obtaining these sub-distribution parameters, use the concatenation of the hidden vectors of the RNN network to generate the weights of the mixture distribution through a fully connected layer, thereby forming the final mixture Gaussian distribution. The specific implementation process for achieving the objective of the present invention includes the following steps:
[0006] Step 1: Collection and preprocessing of time series data of commodity sales volumes and prices
[0007] The dataset contains time series data of multiple commodities, which are numbered according to commodity numbers such as commodity grades, department affiliations, product categories, and geographical regions. Although the sales data contains rich information, since sales on the platform do not occur daily, the data has a high degree of sparsity, and in the case of no sales volume, the sales price data is usually missing. To solve this problem, the present invention cleans the data, eliminates duplicate records, and fills in the missing price data with the transaction price of the most recent date, and finally constructs a complete time series dataset of commodity sales volumes and sales prices.
[0008] Step 2: Construct a MOD-DeepVAR binary mixture distribution model and perform training optimization; existing DeepVAR usually adopts a single network architecture and models time series based on a single distribution assumption.
[0009] Based on paired commodity volume-price time series data, the present invention proposes a MOD-DeepVAR model based on a binary mixture distribution to fit its joint distribution. The modeling process is as follows: Use the training data of the commodity sales volume and price dataset as the input of the MOD-DeepVAR model in the form of time series, and input them into multiple parallel sub-networks of the RNN network for training respectively; On the basis of the outputs of multiple networks, construct the overall mixed joint probability distribution by weighted mixing of the Gaussian sub-distributions generated by each sub-network; The model output is the parameters of the trained binary Gaussian joint probability distribution.
[0010] For the MOD-DeepVAR hybrid distribution model established in Step 2, a loss function based on negative log-likelihood is adopted, combined with L2 regularization of the Gaussian sub-distribution, to enhance the fitting ability for complex data. Negative log-likelihood is used to measure the difference between the predicted output and the true value, while L2 regularization helps prevent the model from overfitting and ensures the generalization ability of the model; the stochastic gradient descent method (SGD) is used to optimize the model parameters, and the best parameters of the model are learned by maximizing the log-likelihood; after each parameter update, the loss function values of various commodities on the training set are calculated, and the training is repeated until the loss function value converges to the optimum. After the training is completed, the model outputs the optimized hybrid joint probability distribution parameters. These parameters are used for the subsequent joint prediction of sales volume and price to improve the accuracy of the intermittent time series prediction of bulk commodities.
[0011] Step 3: Through the model trained and optimized in Step 2, perform joint prediction of commodity sales volume and price.
[0012] Specific measures taken for the above steps also include:
[0013] In Step 1 above, the present invention aggregates all data by day and commodity type, and respectively counts the sales volume data and price data of various commodities within that day to construct time series data. The basic form is as follows:
[0014] (SKU,z 1,1 ,z 1,2 ,…,z 1,T )
[0015] (SKU,z 2,1 ,z 2,2 ,…,z 2,T )
[0016] Among them, SKU is the commodity number, which is composed of commodity grade, department attribution, product category, geographical area, etc. This feature is the unique identifier of the commodity; z 1,t is the daily sales volume data of a certain SKU on the t-th day, and z 2,t is the daily price data of a certain SKU on the t-th day; T is the total number of days of historical daily sales volume data.
[0017] In Step 2 above, the MOD-DeepVAR binary hybrid distribution model is a multivariate time series prediction model based on deep autoregression, which is specifically used to process complex time series data with mutual correlations between multiple variables. Let z t =(z 1,t ,z 2,t ) represent the sales volume and price of the commodity at the t-th time point. Using the historical quantity-price time series data of the previous T-τ days to predict the quantity-price time series of the next τ days. The constructed model is:
[0018] p(z T-τ+1 ,…,z T ) = MOD - DeepVAR(z1, z2, …, z T-τ ),
[0019] The framework diagram of the MOD - DeepVAR model is shown in Figure 1 , including modules such as the input layer, RNN network layer, fully - connected layer, and output layer. Data enters the RNN network through the input layer. The parameters generated by the RNN network are used to generate sub - distribution parameters through the fully - connected layer, which are used to construct a binary mixture distribution. The model training and optimization are based on the mixture distribution and actual data, and finally the prediction results are output, forming a complete and efficient prediction model framework.
[0020] The construction and training of the MOD - DeepVAR model specifically include the following three steps:
[0021] Step 2.1. Build multiple RNN networks and output hidden vectors
[0022] The MOD - DeepVAR model uses the RNN architecture, which can specifically be an LSTM network or a GRU network. Define information such as the number of cell layers and hidden layers of the neural network according to the input parameters. Define [1, T - τ] as the training range, where T represents the total length of the time series, and [T + τ - 1, T] as the prediction range. What needs to be predicted is the next τ steps. The MOD - DeepVAR assumes that at each time point t, the intermittent time - series data follows a Gaussian mixture distribution, which is composed of K sub - distributions generated by K time - series networks. For the target variable z t , the probability density function of the model is:
[0023]
[0024] where, θ t = {μ 1,t ,..., μ K,t , ∑ 1,t ,..., ∑ K,t , α 1,t ..., α K,t} represents the set of mixture - distribution parameters at the t - th time point; represents the probability density function of the k - th binary Gaussian sub - distribution; α k,t represents the weight of the k - th sub - distribution at the t - th time point; μ k,t represents the mean vector of the k - th sub - distribution at the t - th time point; ∑ k,t represents the covariance matrix of the k - th sub - distribution at the t - th time point.
[0025] At each time step, use the target observation value z of the previous moment t-1and the K RNN network hidden states h at the previous moment k,t-1 (k = 1, ..., K) as inputs to obtain the K RNN network hidden states h k,t (k = 1, ..., K) at the current moment, which is specifically expressed as:
[0026] h k,t = RNN k (h k,t-1 , z t-1 , Θ k ),
[0027] where Θ k represents the internal parameters of the k-th RNN network layer.
[0028] Step 2.2. Calculate the sub-distribution parameters
[0029] For the mixture sub-distribution parameters, the hidden state h k,t (k = 1, ..., K) at time t is used as the input vector of the fully connected layer of the mixture Gaussian distribution. After being transformed by a linear function and an exponential function, the mean parameter and covariance matrix of the sub-distribution in the mixture Gaussian distribution are obtained. The mean parameter is denoted as μ k,t , and the specific calculation method is as follows:
[0030]
[0031] where respectively represent the transpose of the weight parameters of the fully connected layer of the mean of the k-th Gaussian distribution, and b μ,k respectively represent the bias parameters of the fully connected layer of the mean of the k-th Gaussian distribution.
[0032] The covariance matrix is denoted as ∑ k,t , and the specific calculation method is as follows:
[0033]
[0034] where the covariance matrix ∑ k,t is composed of the diagonal matrix D k,t and the low-rank matrix V k,t ; the diagonal matrix D k,t represents the variances of each variable, and is obtained by mapping the fully connected layer weights and biases b d,k (k = 1, ..., K) to positive values through the exponential function exp(·). Here, diag(·) means taking the diagonal elements and performing a transformation through the exponential function to ensure that the diagonal elements are positive, thereby ensuring the positive definiteness of the covariance matrix; the low-rank matrix V k,t is obtained through the fully connected layer weights and biases b d,k(k = 1, ..., K) is calculated.
[0035] Step 2.3. Calculate the sub - distribution weights
[0036] For the mixture - distribution weights, the K RNN network hidden states h at the current time k,t are concatenated and denoted as H t . The concatenated H t is input into a multinomial - distribution fully - connected layer. After transformation by a linear function and a Softmax activation function, the weights α of the sub - distributions in the mixture Gaussian distribution are obtained t , and the calculation method is as follows:
[0037] H t = Concat(h 1,t , h 2,t ,..., h K,t )
[0038]
[0039] where represents the transpose of the weight parameters of the multinomial - distribution fully - connected layer, and b α represents the bias parameters of the multinomial - distribution fully - connected layer.
[0040] By using the Softmax activation function, the weights are normalized so that their sum is 1, assigning weights to each sub - distribution and determining its contribution to the final mixture distribution;
[0041] Combining the parameters and weights of the sub - distributions, a complete mixture Gaussian distribution can be constructed.
[0042] During the process of optimizing the model, the negative log - likelihood is used as the main loss function, and a regularization term is added. The loss function is expressed as follows:
[0043]
[0044] where q k,t represents the quantile of the true non - zero values of all samples in the same training batch at time point t, and λ k is the regularization coefficient of the mean of the k - th sub - distribution, used to control the weight of the regularization term and achieve a balance with the negative log - likelihood loss term.
[0045] The constructed loss function consists of two parts: the first part is the negative log - likelihood loss term, which is used to measure the difference between the mixture Gaussian distribution generated by the model and the real data; the second part is the L2 regularization term of the mean of the Gaussian sub - distributions, aiming to make the mean μ of the k - th Gaussian sub - distribution at each time point t k,t as close as possible to the quantile qk,t Avoid the influence of extreme values on the mean. The model assumes that K Gaussian sub - distributions are modeled as Gaussian distributions with means close to each quantile, thereby enhancing the model's fitting ability for intermittent time - series distributions.
[0046] Finally, the model minimizes the loss function through the mini - batch gradient descent algorithm and updates the network parameters using backpropagation to further improve the prediction accuracy. Through this optimization method, the MOD - DeepVAR model not only captures the complex dynamic characteristics of the data but also can effectively handle the means and weights of different sub - distributions.
[0047] The present invention also provides a prediction system for implementing the above - mentioned sequence prediction method. The prediction system includes: a data acquisition module, a model construction module, an optimization module, and a prediction module;
[0048] The data acquisition module is used to collect the time - series data of commodity sales volume and price, and perform data pre - processing to generate a time - series data set containing commodity sales volume and price;
[0049] The model construction module is used to design and construct a MOD - DeepVAR model based on a binary mixture distribution, generate potential sub - distributions using multiple parallel RNN sub - networks, and generate an overall mixture Gaussian distribution through weighted mixing;
[0050] The loss function calculation module is used to calculate the loss function combining negative log - likelihood loss and L2 regularization to optimize the model parameters;
[0051] The prediction module is used to perform joint prediction of commodity sales volume and price based on the optimized mixture Gaussian distribution parameters, and output the prediction results of sales volume and price in the future time period.
[0052] The present invention also provides the above - mentioned prediction method, or the application of the above - mentioned prediction system in supermarket retail prediction, parts sales, etc.
[0053] The beneficial effects of the present invention include: By constructing a binary mixture distribution model MOD - DeepVAR, effectively combining the bivariate relationship between commodity sales volume and price, it can fully capture the multi - peak characteristics existing in market dynamics, and avoid the limitations of a single - distribution model when dealing with complex data. At the same time, it enhances the adaptability of the model to market changes, effectively dealing with the complexity and uncertainty in the actual market, especially when the system dynamic fluctuations are obvious or the external environment changes violently. Compared with traditional single - distribution models, the present invention can more accurately describe the impact of price fluctuations on sales volume, thereby improving the generalization ability of the model in different market environments, and reducing more than five percentage points in indicators such as RMSE and MASE.
[0054] The present invention more effectively models the relationship between price and sales volume through a mixture distribution model, significantly improving the accuracy and stability of intermittent time series prediction. This method provides an innovative and efficient solution for predicting tasks dealing with the relationship between price and sales volume, offers important support for relevant market decisions and resource optimization, and has broad application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for description in the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0056] Figure 1 It is a framework diagram of the MOD-DeepVAR model of the present invention.
[0057] Figure 2 It is a flowchart of the prediction method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0058] Combined with the following specific embodiments and drawings, the present invention will be further described in detail. The processes, conditions, experimental methods, etc. for implementing the present invention, except for the specifically mentioned content below, are all common knowledge and well-known common sense in the art, and the present invention has no special restrictive content.
[0059] Aiming at the problem that the multi-modal characteristics of data and the influence of external factors such as price are not fully considered in intermittent time series prediction, the present invention proposes a binary mixture distribution model (MOD-DeepVAR) for the relationship between quantity and price, and realizes a prediction method for intermittent time series through the model, as Figure 2 shown. This method improves the prediction accuracy of intermittent time series data by modeling the complex interaction relationship between commodity price and sales volume. The method includes: Step 1, collect relevant data including commodity sales volume and price, and construct a time series data set through preprocessing; Step 2, design a MOD-DeepVAR model with a multi-network structure, generate different latent sub-distributions through multiple RNN networks, and capture the multi-modal characteristics of data in the form of a mixture distribution; design a loss function module, combine the negative log-likelihood loss with the L2 regularization of the mean of the Gaussian sub-distribution to enhance the fitting effect of the model on complex data; Step 3, through the model trained and optimized in Step 2, perform joint prediction of commodity sales volume and price. The binary mixture distribution model constructed in the present invention for intermittent time series prediction can effectively help commodity production enterprises accurately grasp market dynamics, reduce production costs, increase production efficiency, and enhance competitiveness.
[0060] The present invention will be further described in detail through the following specific embodiments.
[0061] Embodiment 1
[0062] Obtain the dataset from the M5 competition on Kaggle. This dataset is the daily sales volume dataset of commodities from Walmart supermarkets in California, Texas, and Wisconsin in the United States. Specifically, the feature extraction includes numbering the time series data of various commodities according to commodity levels, departments, product categories, and geographical regions to form features, and this feature uniquely identifies each commodity; the preprocessing includes filling in missing data, aggregating daily sales volumes, and matching price and sales data at the same time, which altogether involves 30,490 SKU sales volume time series; the data form is (SKU, z 1,1 , z 1,2 , …, z 1,1913 ), (SKU, z 2,1 , z 2,2 , …, z 2,1913 ), z 1,t is the daily sales volume data of a certain SKU on the t-th day, and z 2,t is the daily price data of a certain SKU on the t-th day; the time series is statistically counted at a daily frequency, and the time range of the data is from 2011 to 2016, with a total of 1,913 time points. The training set altogether covers 1,898 days, and the test set is the last 15 days in the dataset.
[0063] Next, construct the MOD-DeepVAR multi-mixture distribution model. Set the number of sub-networks to 2, and use the training data of the 30,490 SKU sales volume and price datasets in the form of time series as the input of the MOD-DeepVAR model and input it into each sub-network for training. The hidden vectors obtained from training generate a mixture distribution through a fully connected layer. Adopt a loss function based on negative log-likelihood, combined with L2 regularization of the Gaussian sub-distribution, to enhance the fitting ability for complex data; use the stochastic gradient descent method (SGD) to optimize the model parameters and learn the optimal parameters of the model by maximizing the log-likelihood; after each parameter update, calculate the loss function values of various commodities on the training set, and repeat the training until the loss function value converges to the optimal. After the training is completed, the model outputs the optimized mixture joint probability distribution parameters and saves them.
[0064] Finally, construct a prediction scheme. Through the commodity daily sales volume data prediction model trained by MOD-DeepVAR, predict the sales volume in the final fifteen days and accumulate to obtain the half-month sales volume.
[0065] The MOD-DeepVAR model implemented in the present invention is compared with the original DeepVAR model and the other two multivariate time series prediction models. Using the four evaluation metrics of MAE, RMSE, MASE, and RMSSE as a measure, the evaluation of the instance prediction results on the test set is reported. Among them, MAE is used to measure the average absolute difference between the predicted value and the actual value; RMSE measures the square root of the mean of the squared prediction errors; MASE is a relative error metric that enables model comparison between different time series; RMSSE is a scaled derivative of RMSE and has unique advantages when evaluating prediction models for intermittent time series. The results are shown in Table 1.
[0066] Table 1 Comparison of prediction results under different models
[0067]
[0068] The smaller the value of each evaluation metric, the better the prediction result of the model. The prediction results of the MOD-DeepVAR model implemented in the present invention are better than those of other baseline models in all metrics. The RMSE of MOD-DeepVAR is reduced by 5.96% compared with DeepVAR, and the error is lower than that of other models. Its MAE has decreased by 6.57%, indicating that the model can more accurately approximate the actual sales volume when predicting the total sales volume in the next half month. In addition, RMSSE and MASE have decreased by 6.44% and 5.42% respectively, indicating that the model not only performs well in absolute error but also shows good stability in relative error.
[0069] Example 2
[0070] Obtain the dataset from the Predict-future-sales competition on Kaggle. This dataset is a daily sales dataset of goods from 1C Company, one of the largest software companies in Russia. Specifically, the feature extraction includes numbering the time series data of multiple goods according to the product grade, department, and product category to form features, which uniquely identify each kind of good; the preprocessing includes filling in missing data, aggregating daily sales volume, and matching price with sales data, involving a total of 3296 SKU sales time series; the data form is (SKU, z 1,1 , z 1,2 , …, z 1,1034 ), (SKU, z 2,1 , z 2,2 , …, z 2,1034 ), where z 1,t is the daily sales volume data of a certain SKU on the t-th day, and z 2,tis the daily price data of a certain SKU on the t-th day; the time series is statistically counted at a daily frequency, and the time range of the data is from 2013 to 2015, with a total of 1034 time points. The training set covers a total of 1019 days, and the test set is the last 15 days in the dataset.
[0071] Next, construct the MOD-DeepVAR multivariate mixture distribution model. Set the number of subnets to 2, and use the training data of 3296 SKU sales volume and price datasets in the form of time series as the input of the MOD-DeepVAR model, and input it into each subnet for training. The obtained hidden vectors are used to generate a mixture distribution through a fully connected layer. Adopt a loss function based on negative log-likelihood, combined with L2 regularization of Gaussian sub-distributions, to enhance the fitting ability for complex data; use the Stochastic Gradient Descent (SGD) method to optimize the model parameters, and learn the optimal parameters of the model by maximizing the log-likelihood; after each parameter update, calculate the loss function value of various commodities on the training set, and repeat the training until the loss function value converges to the optimum. After the training is completed, the model outputs the optimized mixture joint probability distribution parameters and saves them.
[0072] Finally, construct a prediction scheme. Through the daily sales volume data prediction model of commodities trained by MOD-DeepVAR, predict the sales volume in the final fifteen days, and accumulate to obtain the half-month sales volume.
[0073] The MOD-DeepVAR model implemented in the present invention is compared with the original DeepVAR model and the other two multivariate time series prediction models, with MAE, RMSE, MASE, and RMSSE as evaluation indicators. The results are shown in Table 2.
[0074] Table 2 Comparison of prediction results under different models
[0075]
[0076] The smaller the value of each evaluation indicator, the better the prediction result of the model. The prediction results of the MOD-DeepVAR model implemented in the present invention are better than those of other baseline models in all indicators. The RMSE of MOD-DeepVAR is reduced by 5.8% compared with DeepVAR, and its MAE drops by 16.5%, showing a great improvement compared with other models. In addition, RMSSE and MASE are reduced by 5.3% and 5.6% respectively, indicating that the model not only performs well in absolute error but also shows good stability in relative error.
[0077] References
[0078] [1] Salinas D, Bohlke-Schneider M, Callot L, et al. High-Dimensional Multivariate Forecasting with Low-Rank Gaussian Copula Processes[C] / / Proceedings of the 33rd International Conference on Neural Information Processing Systems. 2019: 6827–6837.
[0079] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code. The solutions in the embodiments of the present application can be implemented using the object-oriented programming language Python.
[0080] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or multiple flows and / or blocks
[0081] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in Figure 1 one or more of the flows Figure 1 or multiple flows and / or blocks
[0082] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable apparatus to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one process or a plurality of processes and / or blocks Figure 1 one process or a plurality of processes and / or blocks Figure 1 steps for implementing the functions specified in one block or a plurality of blocks.
[0083] Although the preferred embodiments of the present application have been described, additional changes and modifications can be made by those skilled in the art once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present application.
[0084] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.
[0085] The protection scope of the present invention is not limited to the above embodiments. Without departing from the spirit and scope of the inventive concept of the present invention, changes and advantages that can be conceived by those skilled in the art are included in the present invention, and the appended claims are taken as the protection scope.
Claims
1. A method for predicting intermittent time series based on vector-price relationship, characterized in that: The sequence prediction method comprises: Step 1: Collect time series data including product sales volume and price, number them according to product numbers, clean the data and fill in missing data to generate a complete time series data set; Step 2: Build the MOD-DeepVAR model, use multiple RNN sub-networks to generate different potential sub-distributions, and weighted mix the parameters of each sub-distribution to generate an overall mixed Gaussian distribution, and perform training optimization; Step 3: Use the model trained and optimized in step 2 to jointly predict product sales volume and price.
2. The prediction method according to claim 1, characterized in that: In step 1, the product number includes: product grade, department affiliation, product category, and geographical area; and / or, Eliminate duplicate data, aggregate product sales volume and price data by day, and fill in missing price data points with the transaction price of the most recent date; and / or, The sales volume data and price data of the goods are constructed as time series data. The time series data format is as follows: (SKU,z 1,1 ,With 1,2 ,…,With 1,T ), (SKU,z 2,1 ,With 2,2 ,…,With 2,T ), SKU is the product number, including product grade, department affiliation, product category, and geographic area information, and is the unique identifier of the product. 1,t is the daily sales data of a SKU on day t, z 2,t is the daily price data of a certain SKU on the tth day; T is the total number of days of historical daily sales data.
3. The prediction method according to claim 1, characterized in that: In step 2, the joint distribution of the quantity-price time series data of the paired commodities is fitted by the MOD-DeepVAR model based on binary mixed distribution; the model uses the historical quantity-price time series data of the past t-τ days to predict the quantity-price time series of the future τ days, which is expressed as follows: p(z T-τ+1 ,…,with T )=MOD-DeepVAR(z1,z2,…,z T-τ ), Among them, z t =(z 1,t ,z 2,t ) represents the sales volume and price of the product at time point t; and / or, The MOD-DeepVAR model is optimized using stochastic gradient descent, combining negative log-likelihood loss with an L2 regularization term on the mean of the Gaussian sub-distribution.
4. The prediction method according to claim 3, characterized in that: During the training process of the MOD-DeepVAR model, the training data of the commodity sales volume and price data set are used as the input of the MOD-DeepVAR model in the form of time series, and are respectively input into multiple parallel sub-networks of the RNN network for training; based on the output of multiple networks, the Gaussian sub-distributions generated by each sub-network are weighted mixed to construct an overall mixed joint probability distribution; the model output is the binary Gaussian joint probability distribution parameters obtained through training.
5. The prediction method according to claim 1, characterized in that: In step 2, the construction and training of the MOD-DeepVAR model includes the following steps: Step 2.
1. Use the RNN architecture to build the MOD-DeepVAR model, define the training range and prediction range; input historical data step by time and calculate the corresponding hidden state vector; Step 2.
2. Calculate the mean and covariance matrix of each potential sub-distribution through the hidden state vector output by each RNN sub-network for subsequent generation of mixed Gaussian distribution; Step 2.
3. Concatenate the hidden state vectors of the RNN sub-network at the current moment, calculate the weights through the fully connected layer of the multinomial distribution, and use the Softmax activation function to normalize the weights so that their sum is 1. Assign weights to each sub-distribution to determine its contribution to the final mixed distribution.
6. The prediction method according to claim 5, characterized in that: In step 2.1, the training range is [1, T-τ], the prediction range is [T+τ-1, T], T represents the total length of the time series; at each time step, the target observation value z at the previous moment is used t-1 and the K RNN hidden states h at the previous moment k,t-1 As input, we get the K RNN hidden states h at the current moment k,t ; and / or, In step 2.2, for the mixed sub-distribution parameters, the hidden state h at time t k,t As the input vector of the mixed Gaussian distribution fully connected layer, after the transformation of linear function and exponential function, the mean parameter and covariance matrix of the sub-distribution of the mixed Gaussian distribution are obtained; and / or, In step 2.3, for the mixed distribution weight, the K RNN network hidden states h at the current moment are k,t Splice, denoted as H t , the spliced H t Input to a multinomial distribution fully connected layer, after the transformation of linear function and Softmax activation function, the weight α of the neutron distribution of the mixed Gaussian distribution is obtained. t .
7. The prediction method according to claim 3, characterized in that: In the process of optimizing the model, the negative log-likelihood is used as the loss function, and a regularization term is added; the loss function is expressed as follows: Among them, q k,t represents the quantile of the true non-zero value of all samples in the same training batch at time point t, λ k is the regularization coefficient of the mean of the kth subdistribution, μ k,t represents the mean parameter of the neutron distribution of the mixed Gaussian distribution, p(z t θ t ) represents the probability density of the model, z t represents the target variable, θ t Represents the set of mixture distribution parameters at the tth time point.
8. The prediction method according to claim 7, characterized in that: The loss function is minimized through the mini-batch gradient descent algorithm, and the network parameters are updated using back propagation to improve the prediction accuracy.
9. A prediction system for implementing the prediction method according to any one of claims 1 to 8, characterized in that: The prediction system includes: a data acquisition module, a model building module, an optimization module, and a prediction module; The data collection module is used to collect time series data of commodity sales volume and price, and perform data preprocessing to generate a time series data set containing commodity sales volume and price; The model building module is used to design and build a MOD-DeepVAR model based on a binary mixed distribution, using multiple parallel RNN sub-networks to generate potential sub-distributions, and generating an overall mixed Gaussian distribution through weighted mixing; The loss function calculation module is used to calculate the loss function combining negative log-likelihood loss and L2 regularization to optimize model parameters; The prediction module is used to perform a joint prediction of commodity sales volume and price based on the optimized mixed Gaussian distribution parameters, and output the sales volume and price prediction results in a future time period.
10. Application of the forecasting method according to any one of claims 1 to 8, or the forecasting system according to claim 9 in supermarket retail forecasting and parts sales.