Data prediction method and device, electronic equipment and storage medium
By extracting feature and calculating the historical timing data of the target item, the similarity probability distribution data is generated, and the problem of low accuracy of intermittent data prediction in the prior art is solved, achieving a more efficient and stable prediction effect.
Patent Information
- Application Number
- CN202311514966.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-14
- Publication Date
- 2025-05-16
AI Technical Summary
When predicting intermittent data, the prior art cannot keenly capture the change pattern of long-tail data, resulting in low prediction accuracy and unstable.
By extracting historical feature data and predictive feature data based on the historical time series data of the target item, calculating feature weight data, generating similarity probability distribution data, and finally obtaining the prediction result of the target item. The method includes steps such as feature generation, weight determination, probability distribution calculation and prediction generation.
It improves the accuracy and stability of predicting intermittent data, can capture the data change patterns more sensitively, and provide more reliable prediction results.
Smart Images

Figure CN120011776A_ABST
Abstract
Description
Background Art
[0002] When predicting intermittent data, relevant technologies can make predictions through time series models: by separately modeling historical data for sales and non-sales periods, and then outputting a forecast of the previous day's sales; or by creating some strong intermittent features through machine learning models to predict intermittent data.
[0003] Since intermittent data itself is relatively sparse, the relevant technology does not take into account the distribution of the data and cannot keenly capture the changing patterns of long-tail data. The prediction results of the relevant technology for this data are not ideal, and the accuracy of predicting intermittent data is low and unstable.
[0004] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute the prior art known to ordinary technicians in the field. Summary of the invention
[0005] The present disclosure provides a data prediction method, device, electronic device and computer-readable storage medium, which at least to some extent overcome the problem of low accuracy in predicting intermittent data in the related art.
[0006] Other features and advantages of the present disclosure will become apparent from the following detailed description, or may be learned in part by the practice of the present disclosure.
[0007] According to one aspect of the present disclosure, a data prediction method is provided, comprising: obtaining historical feature data and predicted feature data based on historical time series data of a target item; obtaining feature weight data based on the historical time series data, wherein the feature weight data includes a weight coefficient corresponding to the historical feature data; obtaining similarity probability distribution data based on the historical feature data, the predicted feature data and the feature weight data; and obtaining a prediction result corresponding to the target item based on the similarity probability distribution data.
[0008] In one embodiment of the present disclosure, the similarity probability distribution data obtained based on the historical feature data, the predicted feature data and the feature weight data includes: obtaining a similarity matrix based on the feature weight data, the historical feature data and the predicted feature data; converting the similarity matrix into discrete probability density data based on a normalized exponential function; and obtaining the similarity probability distribution data based on the historical time series data and the discrete probability density data.
[0009] In one embodiment of the present disclosure, the historical time series data includes at least one of the following: historical sales data, target item identification, and sales date; wherein, obtaining the similarity probability distribution data based on the historical time series data and the discrete probability density data includes: obtaining the similarity probability distribution data based on the historical sales data and the discrete probability density data.
[0010] In one embodiment of the present disclosure, obtaining a similarity matrix based on the feature weight data, the historical feature data, and the predicted feature data includes: performing regularization processing on the historical feature data to obtain a historical regularization result; performing regularization processing on the predicted feature data to obtain a predicted regularization result; performing dot product processing on the historical regularization result, the predicted regularization result, and the feature weight data to obtain the similarity matrix.
[0011] In one embodiment of the present disclosure, obtaining the prediction result corresponding to the target item according to the similarity probability distribution data includes: determining the prediction result corresponding to the target item according to a preset position of the similarity probability distribution data.
[0012] In one embodiment of the present disclosure, obtaining feature weight data according to the historical time series data includes: performing fitting processing on the historical time series data to obtain the feature weight data.
[0013] In one embodiment of the present disclosure, the historical feature data and the predicted feature data include at least one of the following: median, mean, cumulative number of zero sales, non-zero sales interval, proportion of zero values, coefficient of variation, number of peak points, spectrum analysis, trend strength, seasonal strength, k-order autocorrelation, seasonal cycle, data length, and k-order partial autocorrelation test.
[0014] According to another aspect of the present disclosure, there is also provided a data prediction device, comprising:
[0015] The feature generation module obtains historical feature data and predicted feature data based on the historical time series data of the target item;
[0016] A weight determination module, which obtains feature weight data according to the historical time series data, wherein the feature weight data includes a weight coefficient corresponding to the historical feature data;
[0017] A probability distribution module, which obtains similarity probability distribution data according to the historical feature data, the predicted feature data and the feature weight data;
[0018] The prediction generation module obtains the prediction result corresponding to the target item according to the similarity probability distribution data.
[0019] According to another aspect of the present disclosure, an electronic device is also provided, including: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute any one of the above-mentioned data prediction methods by executing the executable instructions.
[0020] According to another aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the data prediction method described in any one of the above is implemented.
[0021] The data prediction method, device, electronic device and computer-readable storage medium provided by the embodiments of the present disclosure obtain historical feature data and predicted feature data based on the historical time series data of the target item; fit the historical feature data to obtain feature weight data of each feature corresponding to the target item; obtain a similarity matrix based on the feature weight data, the historical feature data and the predicted feature data; convert the similarity matrix into discrete probability density data based on a normalized exponential function; obtain similarity probability distribution data based on the historical time series data and the discrete probability density data; determine the prediction result corresponding to the target item based on the median of the similarity probability distribution data, which can improve the accuracy and stability of predicting intermittent data.
[0022] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification are used to explain the principles of the present disclosure. Obviously, the accompanying drawings described below are only some embodiments of the present disclosure, and for ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without creative work.
[0024] Figure 1 A flow chart of a data prediction method in an embodiment of the present disclosure is shown;
[0025] Figure 2 A schematic diagram of data prediction in an embodiment of the present disclosure is shown;
[0026] Figure 3 A flow chart of another data prediction method in an embodiment of the present disclosure is shown;
[0027] Figure 4 A schematic diagram of historical time series data characteristics in an embodiment of the present disclosure is shown;
[0028] Figure 5 A basic model fitting schematic diagram in an embodiment of the present disclosure is shown;
[0029] Figure 6 A schematic diagram of sample regularization formed by splicing historical feature data and predicted feature data in an embodiment of the present disclosure is shown;
[0030] Figure 7 A schematic diagram of generating a similarity matrix in an embodiment of the present disclosure is shown;
[0031] Figure 8 A schematic diagram of determining a prediction result in an embodiment of the present disclosure is shown;
[0032] Fig. 9 A schematic diagram of a data prediction device in an embodiment of the present disclosure is shown; and
[0033] Fig.10 A structural block diagram of an electronic device in an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0034] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that the disclosure will be more comprehensive and complete and to fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0035] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the figures represent the same or similar parts, and their repeated description will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0036] The present exemplary implementation is described in detail below with reference to the accompanying drawings and embodiments.
[0037] First, a data prediction method is provided in an embodiment of the present disclosure, and the method can be executed by any electronic device with computing and processing capabilities.
[0038] Figure 1 A flow chart of a data prediction method in an embodiment of the present disclosure is shown as follows: Figure 1 As shown, the data prediction method provided in the embodiment of the present disclosure includes the following steps:
[0039] S102, obtaining historical feature data and predicted feature data according to the historical time series data of the target item.
[0040] In one embodiment, the historical time series data includes at least one of the following: historical sales data, target item identification, and sales date; taking the historical time series data including historical sales data, target item identification, and sales date as an example, as shown in Table 1, a data specification definition table:
[0041] Table 1 Data specification definition table
[0042] obj_no ds y s0001 2020-01-01 20
[0043] Among them, obj_no is the merchant's item ID, ds is the sales date, and y is the historical sales data, that is, the true value of historical sales.
[0044] In one embodiment, different types of features corresponding to the historical time series data of the target item are extracted, and these features include but are not limited to at least one of the following: correlation, data distribution, entropy, stability, trend, discontinuity, number of peaks, maximum value, median, mean, number of spikes, cumulative number of zero sales, non-zero sales interval, proportion of zero values, coefficient of variation, spectrum analysis, trend strength, seasonal strength, k-order autocorrelation, seasonal cycle, data length, k-order partial autocorrelation test, etc.
[0045] In one embodiment, historical feature data is extracted and predicted feature data is predicted based on historical time series data, that is, historical time series data, historical feature data, and / or predicted feature data include but are not limited to at least one of the following: median, mean, cumulative number of zero sales, non-zero sales interval, proportion of zero values, coefficient of variation, number of peak points, spectrum analysis, trend strength, seasonal strength, k-order autocorrelation, seasonal cycle, data length, k-order partial autocorrelation test, etc.
[0046] In one embodiment, historical time series data can be analyzed and predicted using one or more methods including tsfeatures based on R package, tsfresh based on Python, hctsa based on Matlab, tsfresh based on Python, or catch22 based on Python to obtain historical feature data and predicted feature data.
[0047] S104, obtaining feature weight data according to the historical time series data, wherein the feature weight data includes a weight coefficient corresponding to the historical feature data.
[0048] In one embodiment, historical time series data is fitted to obtain feature weight data; the feature weight data includes but is not limited to historical feature data, and weight coefficients corresponding to the historical feature data, etc.; through the method of fitting historical time series data, the contribution weight of each feature is calculated in real time, which can prevent overfitting on the one hand, and take into account the contribution degree of each feature on the other hand.
[0049] In one embodiment, historical feature data is extracted based on historical time series data; and the historical feature data is fitted to obtain feature weight data.
[0050] S106, obtaining similarity probability distribution data according to the historical feature data, the predicted feature data and the feature weight data.
[0051] In one embodiment, a similarity matrix is obtained based on historical feature data, predicted feature data and feature weight data, and the similarity matrix is converted into discrete probability density data based on a normalized exponential function; similarity probability distribution data is obtained based on historical time series data and discrete probability density data.
[0052] In one embodiment, regularization is performed on historical feature data to obtain historical regularization results; regularization is performed on predicted feature data to obtain predicted regularization results; dot product processing is performed on the historical regularization results, the predicted regularization results, and the feature weight data to obtain a similarity matrix; based on a normalized exponential function, the similarity matrix is converted into discrete probability density data; similarity probability distribution data is obtained based on historical sales data and discrete probability density data in the historical time series data.
[0053] S108, obtaining a prediction result corresponding to the target item according to the similarity probability distribution data.
[0054] In one embodiment, the predicted values in the similarity probability distribution data are sorted from small to large or from large to small, and the predicted result corresponding to the target item is determined according to the preset position of the similarity probability distribution data; it should be noted that the preset position can be set according to user needs or historical data, and the preset position can be the median, that is, the predicted result corresponding to the target item is determined according to the preset position of the similarity probability distribution data.
[0055] In one embodiment, the prediction result corresponding to the target item is determined by the median of the similarity probability distribution data, the number with the largest probability in the similarity probability distribution data, etc.
[0056] In one embodiment, historical time series data of one or more items within a period of time is obtained, a prediction model is established and trained based on S102-S108, and the historical time series data corresponding to the target item is input into the prediction model to obtain a prediction result.
[0057] In the above embodiment, the current historical time series data is analyzed to extract the time series features during the history and prediction period, namely, the historical feature data and the prediction feature data; the weight coefficient W of each feature of the item is obtained by model fitting; the historical feature data, the prediction feature data, and the feature weight coefficient are combined to calculate the similarity; the similarity is sorted to determine the similarity probability distribution data; according to the determined probability, the position with the largest expectation in the similarity probability distribution data is found as the prediction result and the prediction result is output; the similarity is calculated through the characteristics of the time series itself, and then the prediction probability distribution is obtained through the similarity between each step of the prediction and the history, and then the prediction result at that time is obtained, and the accuracy and stability of predicting intermittent data are improved through deep feature processing, regularization, distribution fitting, etc.
[0058] Figure 2 A schematic diagram of data prediction in an embodiment of the present disclosure is shown. Figure 2 As shown:
[0059] Calculate similarity: Perform deep feature processing and analysis on the current historical time series data, extract the time series features of the historical and forecast periods, i.e., historical feature data and forecast feature data; obtain the weight coefficient W of each feature X of the item through historical period model fitting; combine the historical feature data H, forecast feature data P, and feature weight coefficient W to calculate similarity;
[0060] Prediction process: Sort the similarities to determine the similarity probability distribution data; according to the determined probability, find the position with the largest expectation in the similarity probability distribution data (usually 50%) as the prediction result and output the prediction result;
[0061] In the above embodiment, similarity is calculated by the characteristics of the time series itself, and then the prediction probability distribution is obtained by the similarity between each step prediction and history, and then the prediction result at that time is obtained. At the same time, the accuracy of the accuracy estimation is improved through methods such as feature deep processing, regularization, and distribution fitting.
[0062] Based on the same inventive concept, the present disclosure also provides a data prediction device in the following embodiments. Since the principle of solving the problem in the device embodiment is similar to that in the above method embodiment, the implementation of the device embodiment can refer to the implementation of the above method embodiment, and the repeated parts will not be repeated.
[0063] Figure 3 A flow chart of another data prediction method in an embodiment of the present disclosure is shown as follows: Figure 3 As shown, the data prediction method provided in the embodiment of the present disclosure includes the following steps:
[0064] S302, obtaining a similarity matrix according to feature weight data, historical feature data, and predicted feature data.
[0065] In one embodiment, the features extracted and analyzed from historical time series data include correlation, data distribution, entropy, stability, trend, discontinuity, number of peaks, etc. These features can reflect the basic form of a time series, such as trend and fluctuation characteristics; the maximum value, median, mean, number of peaks, etc. of the historical time series data can be obtained, and two strongly correlated intermittent time series features can also be obtained: the cumulative number of 0 sales, non-zero sales intervals, etc.
[0066] In one embodiment, historical time series data can be analyzed and predicted using one or more methods including tsfeatures based on R package, tsfresh based on Python, hctsa based on Matlab, tsfresh based on Python, or catch22 based on Python to obtain historical feature data and predicted feature data.
[0067] For example, using Python-based tsfresh and catch22, as shown in Table 2, the historical time series data feature table extracts and analyzes the historical time series data.
[0068] Table 2 Historical time series data characteristics table
[0069] feature introduce mean Mean percentile_25 / 50 / 75 Median, upper and lower quartiles interval_0 The percentage of 0 values cum_0 Accumulated 0 value none_0_mean_interval Non-zero sales interval cv2 Coefficient of variation y_number_peak Number of peaks fft_aggregated Spectrum Analysis trend_strength Trend Strength season_strength Seasonal intensity acf-k k-order autocorrelation pacf-k k-order partial autocorrelation test sesonal_period Seasonal Cycle data_len Data length
[0070] Figure 4 A schematic diagram of historical time series data characteristics in an embodiment of the present disclosure is shown as follows: Figure 4 As shown, the horizontal axis is time and the vertical axis is sales volume. Taking the historical time series data of three years as an example, January, March, May, July, August, September, October and December of the first year, March-July, September-December of the second year, January-March, June, July, September, November and December of the third year correspond to the cumulative number of zero values, January and August of the second year and April of the third year correspond to the peak number feature, that is, the number of peak points is 3; June of the first year corresponds to the maximum value feature. It should be noted that the weight coefficient of each feature corresponds to the historical feature data and the predicted feature data corresponding to the feature.
[0071] In one embodiment, the historical feature data is regularized to obtain a historical regularization result; the predicted feature data is regularized to obtain a predicted regularization result; the historical regularization result, the predicted regularization result and the feature weight data are dot-producted to obtain a similarity matrix.
[0072] In one embodiment, the weight of each feature needs to be considered in the similarity calculation process. Some features contribute differently to the prediction, and the importance of each feature is different. By calculating the weight of each feature, the similarity is finally obtained. The whole process is mainly divided into three parts: 1) basic model fitting; 2) regularization; 3) matrix similarity calculation.
[0073] 1) Basic model fitting
[0074] Considering the different contribution of each feature, instead of treating each feature equally, the contribution weight of each feature is calculated in real time through the method of fitting historical time series data. This can prevent overfitting on the one hand, and take into account the contribution of each feature on the other hand. Figure 5 A basic model fitting schematic diagram in an embodiment of the present disclosure is shown; Figure 5 As shown, by fitting, a weight matrix W is obtained; wherein y is the target value, tn is the feature data, wnn is the weight coefficient corresponding to the feature data, and multiple weight coefficients wnn corresponding to the feature data constitute the matrix W.
[0075] 2) Regularization
[0076] After obtaining the weight of each feature, the sample formed by splicing the historical feature data and the predicted feature data is normalized. The main idea of Normalization is to calculate the p-norm of each sample, and then divide each element in the sample by the p-norm, that is, divide the historical feature data and the predicted feature data by the p-norm respectively; the result of this processing is that the p-norm of each processed sample is equal to 1, so that the data distribution is approximately processed into a normal state.
[0077] The calculation formula of p-norm is shown in formula (1):
[0078]
[0079] Among them, x n is a feature sample; ||x|| P is the p-norm of each sample.
[0080] Figure 6 FIG. 4 is a schematic diagram showing sample regularization formed by splicing historical feature data and predicted feature data in an embodiment of the present disclosure; Figure 6 As shown, the sample formed by concatenating the historical feature data history feature and the predicted feature data further feature is regularized to obtain the regularization result Normalization.
[0081] 3) Matrix similarity calculation
[0082] Assuming that there is a certain relationship between the historical feature data and the predicted feature data, similarity is used here to describe this relationship. Taking the lag and week features needed to predict a day as an example, if a feature of the historical feature data is closer to the feature value of the feature data to be predicted, the larger the calculated inner product is, the higher the similarity is. Therefore, a diagonal matrix of weights is introduced at the end of the model to perform weighted summation of each feature.
[0083] In one embodiment, p=2, and the cosine similarity of the two vectors can be obtained by performing a dot product of the 2-norm regularization result of the historical feature data and the predicted feature data with the weight coefficient. The similarity matrix calculation formula is:
[0084]
[0085] Among them, his df is the historical regularization result, that is, the historical data matrix, with a dimension of n_his*n_features;
[0086] diag(weight) is diag(weight); the dimension is n_features*n_features;
[0087] It is the prediction regularization result; the dimension is prediction days*n_features.
[0088] Based on the above process, the construction of all similarity weights for prediction has been completed, and the results are roughly shown in Table 3 Similarity Weight Construction Table:
[0089] Table 3 Similarity weight construction table
[0090] Prediction Day Historical similar sales Weight predirt_day0 His_day1 0.15 predirt_day0 His_day2 0.2 predirt_day0 His_dayk 0.15 … 0.5 predirt_day1 His_day4 0.1 predirt_day1 His_day9 0.1 predirt_day1 His_daym 0.4 predirt_day1 … 0.4 … … …
[0091] Figure 7 A schematic diagram of generating a similarity matrix in an embodiment of the present disclosure is shown; Figure 7As shown, f_mn is the feature in the historical regularization result, the horizontal dimension is the feature number, and the vertical dimension is the historical days; f_np is the feature in the prediction regularization result, the horizontal dimension is the prediction days, and the vertical dimension is the feature number; wnn is the weight coefficient corresponding to the feature data, and the similarity matrix is obtained. In the similarity matrix, s_np is the feature similarity value in the similarity matrix, the horizontal dimension is the prediction days, and the vertical dimension is the historical days; taking the similarity of the first prediction day as an example, the similarity of the first prediction day is: s_11, s_21…s_n1; based on the normalized exponential function softmax, the similarity matrix is converted into discrete probability density data; based on the historical sales data and discrete probability density data, the similarity probability distribution data, i.e. p1, p2…pn, is obtained.
[0092] S304: Based on the normalized exponential function, the similarity matrix is converted into discrete probability density data.
[0093] S306, obtaining similarity probability distribution data according to the historical time series data and the discrete probability density data.
[0094] In one embodiment, similarity probability distribution data is obtained based on historical sales data and discrete probability density data.
[0095] S308, determining a prediction result corresponding to the target item according to a preset position of the similarity probability distribution data.
[0096] In one embodiment, the modeling is performed based on the non-uniform distribution sampling of weights and the similarity matrix is normalized in order to obtain a reasonable probability density for the historical data when calculating the probability density function (pdf). Since e0=1, e1=e, in the normalized exponential function softmax, each historical data will obtain a weight of [1, e]. This is done to eliminate the influence of unreasonable weight distribution when the index is large. Then, each column is converted into a discrete probability density function using the softmax function. The greater the similarity, the greater the probability density function. The discrete probability density function calculation formula (3) is as follows:
[0097]
[0098] Among them, p(x i ) is discrete probability density data;
[0099] similariy in is the value in the similarity matrix.
[0100] The discrete cumulative probability density function is calculated based on the historical sales data and distribution, as shown in formula (4):
[0101]
[0102] Among them, p(x i =sales i ) is the cumulative similarity probability distribution function;
[0103] Sales i For predicted data.
[0104] In this way, the similarity probability distribution function is obtained, and the median of the distribution is taken as the prediction result.
[0105] Figure 8 A schematic diagram of determining a prediction result in an embodiment of the present disclosure is shown; Figure 8 As shown, based on the normalized exponential function, the similarity matrix is converted into discrete probability density data. The discrete probability density data takes sales volume 1 corresponding to p=0.17, sales volume 0 corresponding to p=0.15, sales volume 1 corresponding to p=0.05, sales volume 0 corresponding to p=0.13, sales volume 9 corresponding to p=0.1, sales volume 0 corresponding to p=0.1, sales volume 1 corresponding to p=0.1, sales volume 2 corresponding to p=0.1, and sales volume 0 corresponding to p=0.1 as an example, and the total probability corresponding to each sales volume is calculated, that is, sales volume 0 corresponds to p=0.48, sales volume 1 corresponds to p=0.32, sales volume 2 corresponds to p=0.1, and sales volume 9 corresponds to p=0.1. The sales volumes are sorted from small to large, and the median of the distribution is sales volume 1, that is, the predicted result is sales volume 1.
[0106] In the above embodiment, similarity calculation is performed based on time series features and then through model fitting, similarity is extracted based on time series features, probability density function is obtained, and prediction is performed through the probability density function learned in real time. The accuracy and stability of predicting intermittent data can be improved through methods such as deep feature processing, regularization, and distribution fitting.
[0107] Fig. 9 A schematic diagram of a data prediction device in an embodiment of the present disclosure is shown. Fig. 9 As shown, the data prediction device 9 includes: a feature generation module 901, a weight determination module 902, a probability distribution module 903, and a prediction generation module 904;
[0108] The feature generation module 901 obtains historical feature data and predicted feature data according to the historical time series data of the target item;
[0109] In one embodiment, the feature generation module 901 is also used to extract different types of features corresponding to the historical time series data of the target object; extract historical feature data based on the historical time series data, and predict feature data.
[0110] A weight determination module 902 obtains feature weight data according to the historical time series data, wherein the feature weight data includes a weight coefficient corresponding to the historical feature data;
[0111] In one embodiment, the weight determination module 902 is also used to perform fitting processing on the historical time series data to obtain feature weight data.
[0112] In one embodiment, the weight determination module 902 is further used to extract historical feature data based on historical time series data; and perform fitting processing on the historical feature data to obtain feature weight data.
[0113] The probability distribution module 903 obtains similarity probability distribution data according to the historical feature data, the predicted feature data and the feature weight data;
[0114] In one embodiment, the probability distribution module 903 is also used to obtain a similarity matrix based on historical feature data, predicted feature data and feature weight data, and convert the similarity matrix into discrete probability density data based on a normalized exponential function; and obtain similarity probability distribution data based on historical time series data and discrete probability density data.
[0115] In one embodiment, the probability distribution module 903 is also used to perform regularization processing on historical feature data to obtain historical regularization results; perform regularization processing on predicted feature data to obtain predicted regularization results; perform dot product processing on the historical regularization results, the predicted regularization results and the feature weight data to obtain a similarity matrix; based on a normalized exponential function, convert the similarity matrix into discrete probability density data; and obtain similarity probability distribution data based on the historical time series data and the discrete probability density data.
[0116] The prediction generation module 904 obtains the prediction result corresponding to the target item according to the similarity probability distribution data.
[0117] In one embodiment, the prediction generation module 904 is also used to sort the predicted values in the similarity probability distribution data from small to large or from large to small, and determine the predicted result corresponding to the target item according to the preset position of the similarity probability distribution data; it should be noted that the preset position can be set according to user needs or historical data, and the preset position can be the median, that is, the predicted result corresponding to the target item is determined according to the preset position of the similarity probability distribution data.
[0118] In one embodiment, the prediction generation module 904 is further used to determine the prediction result corresponding to the target item by using the median of the similarity probability distribution data, the number with the largest probability share in the similarity probability distribution data, and the like.
[0119] In the above embodiment, the current historical time series data is analyzed to extract the time series features during the history and prediction period, namely, the historical feature data and the prediction feature data; the weight coefficient W of each feature of the item is obtained by model fitting; the historical feature data, the prediction feature data, and the feature weight coefficient are combined to calculate the similarity; the similarity is sorted to determine the similarity probability distribution data; according to the determined probability, the position with the largest expectation in the similarity probability distribution data is found as the prediction result and the prediction result is output; the similarity is calculated through the characteristics of the time series itself, and then the prediction probability distribution is obtained through the similarity between each step of the prediction and the history, and then the prediction result at that time is obtained, and the accuracy and stability of predicting intermittent data are improved through deep feature processing, regularization, distribution fitting, etc.
[0120] The system architecture may include terminal devices, networks, and servers.
[0121] The network is a medium used to provide a communication link between terminal devices and servers. It can be a wired network or a wireless network.
[0122] Optionally, the wireless network or wired network described above uses standard communication technology and / or protocol. The network is usually the Internet, but it can also be any network, including but not limited to a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a dedicated network or any combination of a virtual private network). In some embodiments, technologies and / or formats including Hyper Text Mark-up Language (HTML), Extensible Markup Language (XML), etc. are used to represent data exchanged through the network. In addition, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), Internet Protocol Security (IPsec) can also be used to encrypt all or some links. In other embodiments, customized and / or dedicated data communication technologies can also be used to replace or supplement the above data communication technologies.
[0123] The terminal device may be any electronic device, including but not limited to a smart phone, a tablet computer, a laptop computer, a desktop computer, etc., and may be used to display prediction results, etc.
[0124] Optionally, the client of the application installed in different terminal devices is the same, or the client of the same type of application based on different operating systems. Based on the different terminal platforms, the specific form of the client of the application can also be different, for example, the application client can be a mobile client, a PC client, etc.
[0125] The server can be a server that provides various services, such as a background management server that provides support for the device operated by the user using the terminal device. The background management server can analyze and process the received request and other data, and feed back the processing results to the terminal device; for example, it can analyze the current historical time series data, extract the time series features of the history and prediction period, that is, the historical feature data and the prediction feature data; obtain the weight coefficient W of each feature of the item through model fitting; combine the historical feature data, the prediction feature data, and the feature weight coefficient to calculate the similarity; sort the similarities to determine the similarity probability distribution data; according to the determined probability, find the position with the maximum expectation in the similarity probability distribution data as the prediction result and output the prediction result, etc.
[0126] Optionally, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The terminal can be a smart phone, tablet computer, laptop computer, desktop computer, etc., but is not limited to this. The terminal and the server can be directly or indirectly connected via wired or wireless communication, which is not limited in this application.
[0127] Those skilled in the art will appreciate that various aspects of the present disclosure may be implemented as systems, methods or program products. Therefore, various aspects of the present disclosure may be specifically implemented in the following forms, namely: complete hardware implementation, complete software implementation (including firmware, microcode, etc.), or a combination of hardware and software, which may be collectively referred to herein as "circuits", "modules" or "systems".
[0128] Refer to the following Fig.10 The electronic device 1000 according to this embodiment of the present disclosure is described. Fig.10 The electronic device 1000 shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0129] like Fig.10As shown, the electronic device 1000 is in the form of a general computing device. The components of the electronic device 1000 may include but are not limited to: the at least one processing unit 1010, the at least one storage unit 1020, and a bus 1030 connecting different system components (including the storage unit 1020 and the processing unit 1010).
[0130] The storage unit stores program codes, which can be executed by the processing unit 1010, so that the processing unit 1010 executes the steps described in the above “exemplary method” section of this specification according to various exemplary embodiments of the present disclosure.
[0131] For example, the processing unit 1010 can execute the following steps of the above-mentioned method embodiment: obtain historical feature data and predicted feature data based on the historical time series data of the target item; fit the historical feature data to obtain feature weight data of each feature corresponding to the target item; obtain a similarity matrix based on the feature weight data, historical feature data, and predicted feature data; based on a normalized exponential function, convert the similarity matrix into discrete probability density data; obtain similarity probability distribution data based on the historical time series data and the discrete probability density data; and determine the prediction result corresponding to the target item using the median of the similarity probability distribution data.
[0132] The storage unit 1020 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 10201 and / or a cache storage unit 10202 , and may further include a read-only storage unit (ROM) 10203 .
[0133] The storage unit 1020 may also include a program / utility 10204 having a set (at least one) of program modules 10205, such program modules 10205 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.
[0134] Bus 1030 may represent one or more of several types of bus structures, including a memory unit bus or memory unit controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.
[0135] The electronic device 1000 may also communicate with one or more external devices 1040 (e.g., keyboards, pointing devices, Bluetooth devices, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 1000, and / or communicate with any device that enables the electronic device 1000 to communicate with one or more other computing devices (e.g., routers, modems, etc.). Such communication may be performed via an input / output (I / O) interface 1050. Furthermore, the electronic device 1000 may also communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) via a network adapter 1060. As shown, the network adapter 1060 communicates with other modules of the electronic device 1000 via a bus 1030. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 1000, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0136] Through the description of the above implementation, it is easy for those skilled in the art to understand that the example implementation described here can be implemented by software, or by software combined with necessary hardware. Therefore, the technical solution according to the implementation of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the implementation of the present disclosure.
[0137] In an exemplary embodiment of the present disclosure, a computer-readable storage medium is also provided, which may be a readable signal medium or a readable storage medium. A program product capable of implementing the above method of the present disclosure is stored thereon. In some possible implementations, various aspects of the present disclosure may also be implemented in the form of a program product, which includes a program code, and when the program product is run on a terminal device, the program code is used to cause the terminal device to execute the steps according to various exemplary implementations of the present disclosure described in the above “Exemplary Method” section of this specification.
[0138] For example, when the program product in the embodiment of the present disclosure is executed by a processor, the following steps are implemented: historical feature data and predicted feature data are obtained based on the historical time series data of the target item; feature weight data of each feature corresponding to the target item is obtained by fitting the historical feature data; a similarity matrix is obtained based on the feature weight data, the historical feature data, and the predicted feature data; based on a normalized exponential function, the similarity matrix is converted into discrete probability density data; similarity probability distribution data is obtained based on the historical time series data and the discrete probability density data; and the prediction result corresponding to the target item is determined by the median of the similarity probability distribution data.
[0139] More specific examples of computer-readable storage media in the present disclosure may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0140] In the present disclosure, a computer readable storage medium may include a data signal propagated in baseband or as part of a carrier wave, wherein a readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A readable signal medium may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0141] Alternatively, the program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the foregoing.
[0142] In a specific implementation, the program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., using an Internet service provider to connect through the Internet).
[0143] It should be noted that, although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into multiple modules or units to be embodied.
[0144] In addition, although the steps of the method in the present disclosure are described in a specific order in the drawings, this does not require or imply that the steps must be performed in this specific order, or that all the steps shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps, etc.
[0145] Through the description of the above implementation, it is easy for those skilled in the art to understand that the example implementation described here can be implemented by software, or by software combined with necessary hardware. Therefore, the technical solution according to the implementation of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the implementation of the present disclosure.
[0146] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses or adaptations of the present disclosure, which follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The description and examples are intended to be exemplary only, and the true scope and spirit of the present disclosure are indicated by the appended claims.
Claims
1. A data prediction method, characterized in that: include: According to the historical time series data of the target item, historical feature data and predicted feature data are obtained; Obtaining feature weight data according to the historical time series data, wherein the feature weight data includes a weight coefficient corresponding to the historical feature data; Obtaining similarity probability distribution data according to the historical feature data, the predicted feature data and the feature weight data; According to the similarity probability distribution data, a prediction result corresponding to the target object is obtained.
2. The data prediction method according to claim 1, characterized in that: The similarity probability distribution data obtained according to the historical feature data, the predicted feature data and the feature weight data includes: Obtaining a similarity matrix according to the feature weight data, the historical feature data, and the predicted feature data; Based on a normalized exponential function, converting the similarity matrix into discrete probability density data; The similarity probability distribution data is obtained according to the historical time series data and the discrete probability density data.
3. The data prediction method according to claim 2, characterized in that: The historical time series data includes at least one of the following: historical sales data, target item identification, and sales date; Wherein, obtaining the similarity probability distribution data according to the historical time series data and the discrete probability density data includes: obtaining the similarity probability distribution data according to the historical sales data and the discrete probability density data.
4. The data prediction method according to claim 3, characterized in that: The obtaining of a similarity matrix according to the feature weight data, the historical feature data, and the predicted feature data comprises: Performing regularization processing on the historical feature data to obtain a historical regularization result; Performing regularization processing on the prediction feature data to obtain a prediction regularization result; The historical regularization result, the predicted regularization result and the feature weight data are subjected to dot product processing to obtain the similarity matrix.
5. The data prediction method according to claim 1 or 2, characterized in that: Obtaining the prediction result corresponding to the target item according to the similarity probability distribution data includes: The prediction result corresponding to the target object is determined according to the preset position of the similarity probability distribution data.
6. The data prediction method according to any one of claims 1 to 4, characterized in that: The obtaining of feature weight data according to the historical time series data comprises: The historical time series data is fitted to obtain the feature weight data.
7. The data prediction method according to any one of claims 1 to 4, characterized in that: The historical characteristic data and the predicted characteristic data include at least one of the following: median, mean, cumulative number of zero sales, non-zero sales interval, proportion of zero values, coefficient of variation, number of peak points, spectrum analysis, trend strength, seasonal strength, k-order autocorrelation, seasonal cycle, data length, and k-order partial autocorrelation test.
8. A data prediction device, characterized in that: include: The feature generation module obtains historical feature data and predicted feature data based on the historical time series data of the target item; A weight determination module, which obtains feature weight data according to the historical time series data, wherein the feature weight data includes a weight coefficient corresponding to the historical feature data; A probability distribution module, which obtains similarity probability distribution data according to the historical feature data, the predicted feature data and the feature weight data; The prediction generation module obtains the prediction result corresponding to the target item according to the similarity probability distribution data.
9. An electronic device, characterized in that: include: processor; as well as A memory, configured to store executable instructions of the processor; The processor is configured to execute the data prediction method according to any one of claims 1 to 7 by executing the executable instructions.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the data prediction method according to any one of claims 1 to 7 is implemented.