Online retail data prediction system and method
By using specific missing value filling methods and prediction model groups in online retail sales data, the missing value problem in online retail sales data prediction is solved, and the prediction accuracy is improved.
Patent Information
- Application Number
- PCT/CN2023/127689
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-10-30
- Publication Date
- 2025-05-08
AI Technical Summary
The existing retail sales forecasting methods are difficult to effectively deal with the missing value problem in online retail sales data, especially the lack of sales in January and February, resulting in large forecast errors.
The linear spline interpolation-dial method is used to fill the physical retail sales data, and the non-physical retail sales data is filled with the segmented linear function fit-spline interpolation method to form a complete data set, and multiplication decomposition and STL decomposition are used to establish a prediction model group for prediction.
By accurately filling in missing values and establishing matching prediction model groups, the prediction accuracy of online retail sales is improved and prediction errors are reduced.
Smart Images

Figure CN2023127689_08052025_PF_FP_ABST
Abstract
Description
Online retail data prediction system and method Technical Field
[0001] The present invention belongs to the field of big data and business intelligence, and specifically relates to an online retail data prediction system and method. Background Art
[0002] Accurate predictions of online retail sales are not only the basis for the government to formulate retail policies and development plans, but also the foundation for e-commerce and logistics companies to formulate development strategies. They have important guiding significance for the online sales industry.
[0003] Currently, retail sales forecasts are primarily categorized into two types: offline retail sales and micro-online enterprise retail sales. Offline retail data and micro-online enterprise retail sales consist solely of physical retail sales, with stable trends and simple data characteristics. Simple offline retail sales forecasting methods can achieve good forecasting results. However, macro-online retail sales data consists of both physical and non-physical retail sales (non-physical items primarily include virtual products such as e-books, audio, video, and game coins, as well as service-related products such as software memberships and paid software or website services). Physical and non-physical retail data have completely different characteristics, resulting in greater volatility and complexity in macro-online retail data. Using offline retail sales forecasting methods to predict macro-online retail sales can result in significant forecast errors.
[0004] Because data can be missing due to mechanical failure or human negligence, missing value filling methods have long been a key challenge in big data and artificial intelligence research. Currently, effective missing value filling methods have been proposed for data with diverse characteristics (such as data with categorical attributes, streaming data, and small-scale data). However, the online retail sales dataset published by the National Bureau of Statistics has unique characteristics: the sum of sales for January and February each year is known, but the specific monthly sales figures are missing. This further increases the forecast error when applying existing retail sales forecasting methods to online retail sales. Furthermore, existing technical means are unable to effectively address this missing value filling problem.
[0005] Therefore, the present invention proposes an online retail data prediction system and method.
[0006] Summary of the Invention
[0007] In order to make up for the deficiencies of the prior art and solve the technical problems existing in the background technology, the present invention proposes an online retail data prediction system and method.
[0008] The present invention is implemented through the following technical solution: an online retail data prediction method, the method comprising the following steps:
[0009] S1: Data collection: Acquire all online physical retail sales data and online non-physical retail sales data, where the online physical retail sales data is recorded as a first data set, and the online non-physical retail sales data is recorded as a second data set;
[0010] S2: Data filling: using linear spline interpolation-bisection method to fill the first data set to obtain a first complete data set, and using piecewise linear function fitting-spline interpolation method to fill the second data set to obtain a second complete data set;
[0011] S3: Data analysis: within a preset time period before the target prediction period, performing data noise reduction, data rounding, or outlier correction on the first complete data set and the second complete data set, respectively, to obtain a first partial data set and a second partial data set;
[0012] S4: Model building: Use multiplication decomposition to build the first prediction model group, and use STL decomposition to build the second prediction model group;
[0013] S5: Model prediction: importing the first part of the data set and the second part of the data set into the first prediction model group and the second prediction model group respectively to obtain a first prediction value and a second prediction value respectively;
[0014] S6: Prediction value integration: add the first prediction value and the second prediction value to obtain a target prediction value.
[0015] Preferably, the specific steps of obtaining the first complete data set in S2 are as follows:
[0016] S21: Filling the data of any year in the first data set using a linear spline interpolation method to obtain a first filling value, where the first filling value includes two filling values, each filling value corresponds to a missing value;
[0017] S22: Calculate the first filling weight corresponding to the missing value using formulas (1) and (2);
[0018] In formula (1) and formula (2), k here represents the year, W1 ε is the first filling weight for the first missing value, the first filling weight corresponding to the second missing value, is the first fill value for a missing value in the kth year, The first fill value for the second missing value in the kth year;
[0019] S23: multiplying the first filling weight by the sum of the two missing values for each year to obtain a second filling value for each year;
[0020] S24: Filling the second filling value of each year into the first data set to obtain a data set to be fitted;
[0021] S25: inputting the to-be-fitted data set into the first prediction model group for fitting, and obtaining a fitting error;
[0022] S26: Using a dichotomy method, the first filling weight is adjusted multiple times. During each adjustment, it is determined whether the difference between the fitting error before and after the adjustment is greater than a preset threshold. If not, the adjusted first filling weight is determined to be the optimal filling weight.
[0023] S27: Multiply the optimal filling weight by the sum of the two missing values in each year, and fill the result into the first data set to obtain a first complete data set.
[0024] Preferably, the specific steps of obtaining the second complete data set in S2 are as follows:
[0025] S210: Decomposing the second data set into a first item and a second item;
[0026] S211: fitting the first item using a piecewise linear function to obtain a first filling value;
[0027] S213: interpolating the second item using a cubic spline interpolation method to obtain a second filling value;
[0028] S214: Add the first filling value and the second filling value to obtain an initial filling value;
[0029] S215: Calculate the filling weight according to the initial filling value and using formula (3) and formula (4);
[0030] In formula (3) and formula (4), k represents the year, K represents the total number of years, and W1 v is the filling weight for the first missing value, The filling weight corresponding to the second missing value, is the initial filling value for the first missing value in year k, is the initial filling value for the second missing value in year k;
[0031] S216: Multiply the filling weight by the sum of the two missing values, and fill the result into the second data set to obtain a second complete data set.
[0032] Preferably, the preset duration is 12 months.
[0033] Preferably, the steps of establishing the first prediction model group in S4 are as follows:
[0034] S41: Decomposing the first portion of the data set into a first component, a second component, and a third component by multiplication decomposition, wherein the first component, the second component, and the third component correspond to the first prediction model, the second prediction model, and the third prediction model;
[0035] S42: Importing the first component, the second component, and the third component into corresponding prediction models for prediction, respectively, to obtain a first component prediction result, a second component prediction result, and a third component prediction result;
[0036] S43: Multiply the first component prediction result, the second component prediction result, and the third component prediction result to obtain a first prediction value.
[0037] Preferably, the steps of establishing the second prediction model group in S4 are as follows:
[0038] S410: Decomposing the second data set into a fourth component and a fifth component by using an STL decomposition method, wherein the fourth component and the fifth component correspond to a fourth prediction model and a fifth prediction model;
[0039] S411: Importing the fourth component and the fifth component into corresponding prediction models for prediction, respectively, to obtain a fourth component prediction result and a fifth component prediction result;
[0040] S412: Add the fourth component prediction result and the fifth component prediction result to obtain a second prediction value.
[0041] An online retail data prediction system, which utilizes the online retail data prediction method described above, and includes a retail data collection platform, a data prediction service platform, and an execution platform;
[0042] The retail data collection platform is used to collect and transmit online retail data, and the retail data collection platform includes a data integration unit, a data transmission unit and a feedback monitoring unit;
[0043] The data prediction service platform receives retail data and uses it to train a prediction model, outputs retail data prediction results, and the data prediction service platform includes a data processing unit, a data prediction unit and a data storage unit;
[0044] The execution platform receives the retail data forecast results and transmits them to the logistics, raw material supply and sales network maintenance departments, which then take corresponding countermeasures.
[0045] Preferably, the data integration unit operates based on the Internet of Things platform to collect logistics and sales information of online retail stores.
[0046] The beneficial effects of the present invention are:
[0047] The present invention addresses the missing value problem of the online retail sales dataset with unique characteristics, fills the missing values more accurately according to the data trend, and uses a prediction model group that matches the characteristics of the complete dataset after filling to perform predictions, thereby improving the prediction accuracy of online retail sales. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] FIG1 is a flow chart of the method of the present invention;
[0049] FIG2 is a flow chart of obtaining a first complete data set in the present invention;
[0050] FIG3 is a flow chart of obtaining a second complete data set in the present invention;
[0051] FIG4 is a flow chart of obtaining a first prediction value in the present invention;
[0052] FIG5 is a flow chart of obtaining a second prediction value in the present invention;
[0053] FIG6 is a system block diagram of the prediction system of the present invention; DETAILED DESCRIPTION
[0054] The present invention will be further described below in conjunction with specific examples. It should be understood that these examples are intended to illustrate the present invention only and are not intended to limit the scope of the invention. The experimental methods in the following examples, for which specific conditions are not specified, are generally performed under conventional conditions or as recommended by the manufacturer.
[0055] Unless otherwise defined, all professional and scientific terms used herein have the same meaning as those familiar to those skilled in the art. The reagents or raw materials used in the present invention can be purchased through conventional channels. Unless otherwise specified, the reagents or raw materials used in the present invention are used in a conventional manner in the art or in accordance with the product instructions. In addition, any method and material similar to or equivalent to the described content can be applied to the method of the present invention. The present invention is further described with reference to the accompanying drawings and specific embodiments. The preferred embodiments and materials described in the present invention are for demonstration purposes only.
[0056] The online retail data provided by the National Bureau of Statistics is missing monthly data for January and February, revealing only the combined total for January and February. Using existing retail sales forecasting methods for this incomplete time series results in significant forecast errors. Furthermore, macro online retail sales data is comprised of physical and non-physical retail sales, which have distinct characteristics. Therefore, macro online retail data is more volatile and complex, resulting in significant forecast errors using existing retail sales forecasting methods.
[0057] To address the aforementioned issues in the prior art, embodiments of the present invention disclose an online retail data forecasting system and method. It should be noted that the execution entity of the embodiments of the present invention can be a computer or any electronic device with data processing capabilities. For ease of description, the following detailed description will be based on the electronic device as the execution entity.
[0058] Example 1:
[0059] As shown in FIG1 , FIG1 is a flow chart of an online retail data prediction method provided by an embodiment of the present invention.
[0060] S1: Obtain all online physical retail sales data and online non-physical retail sales data, where the online physical retail sales data is recorded as the first dataset and the online non-physical retail sales data is recorded as the second dataset, wherein both the first dataset and the second dataset include data from several years, and the data for each year includes two missing values;
[0061] The present invention uses data on my country's total online retail sales and online physical goods retail sales published by the National Bureau of Statistics, covering a period of 60 months from January 2015 to December 2019. It should be noted that the present invention uses this data period as an example only and is not intended to limit the present invention to this period.
[0062] As shown in Table 1, which presents data on my country's online retail sales from 2015 to 2019, this data includes both online physical retail sales and online non-physical retail sales. As can be seen from the table, specific values for January and February of each year are missing (a total of 10 months from 2015 to 2019). For example, while monthly retail sales data from March to December 2015 are known, the specific values for January and February are missing, resulting in the only known sum of the data for January and February being 475.1 billion yuan.
[0063] Table 1: my country's total online retail sales data from 2015 to 2019 (unit: 100 million yuan)
[0064] As shown in Table 2, Table 2 is: my country's online physical retail sales data from 2015 to 2019. The online physical retail sales data is taken as the first data set. Similarly, the specific values of January and February of each year in the first data set are missing, and only the sum of the data for January and February can be obtained.
[0065] Table 2: my country's online physical retail sales data from 2015 to 2019 (unit: 100 million yuan)
[0066] Subtracting the online retail sales of physical goods from the total online retail sales in my country yields the total online retail sales of non-physical goods, as shown in Table 3. Table 3 shows the data for my country's online non-physical retail sales from 2015 to 2019. This data is used as the second dataset. Similarly, the specific values for January and February of each year are missing in the second dataset, so only the sum of these two months is available. Non-physical goods primarily include virtual products such as e-books, audio, video, and game coins, as well as service products such as software memberships and paid software or website services.
[0067] Table 3: my country’s online non-physical retail sales data from 2015 to 2019 (unit: 100 million yuan)
[0068] At this point, the electronic device obtains the first data set and the second data set. In the embodiment of the present invention, the first data set is the online physical retail sales data, and the second data set is the online non-physical retail sales data, which will not be described in detail later.
[0069] S2: Data filling: using linear spline interpolation-bisection method to fill the first data set to obtain a first complete data set, and using piecewise linear function fitting-spline interpolation method to fill the second data set to obtain a second complete data set;
[0070] Since the data characteristics of the first and second data sets are completely different, different filling methods should be used when filling missing values. This will make the filled values more consistent with the actual situation, so that when the complete data set obtained after filling is used to predict retail sales, the prediction error can be reduced.
[0071] In one embodiment, as shown in FIG2 , FIG2 is a flow chart of obtaining a first complete data set in the present invention. The steps of obtaining the first complete data set include:
[0072] S21: Filling the data of any year in the first data set using a linear spline interpolation method to obtain a first filling value, where the first filling value includes two filling values, each filling value corresponds to a missing value;
[0073] S22: Calculate the first filling weight corresponding to the missing value using formulas (1) and (2);
[0074] In formula (1) and formula (2), k here represents the year, W1 ε is the first filling weight for the first missing value, the first filling weight corresponding to the second missing value, is the first fill value for a missing value in the kth year, The first fill value for the second missing value in the kth year;
[0075] S23: multiplying the first filling weight by the sum of the two missing values for each year to obtain a second filling value for each year;
[0076] S24: Filling the second filling value of each year into the first data set to obtain a data set to be fitted;
[0077] S25: inputting the to-be-fitted data set into the first prediction model group for fitting, and obtaining a fitting error;
[0078] S26: Using a dichotomy method, the first filling weight is adjusted multiple times. During each adjustment, it is determined whether the difference between the fitting error before and after the adjustment is greater than a preset threshold. If not, the adjusted first filling weight is determined to be the optimal filling weight.
[0079] S27: Multiply the optimal filling weight by the sum of the two missing values in each year, and fill the result into the first data set to obtain a first complete data set.
[0080] Online physical retail sales data exhibits significant trends and seasonality. Furthermore, due to the Spring Festival holiday for express delivery companies, there is a significant discrepancy between actual retail sales in January and February (the Spring Festival typically falls in February). Directly using traditional spline interpolation to fill missing values can result in significant discrepancies between the results and actual data, leading to larger prediction errors. Therefore, this embodiment of the present invention employs linear spline interpolation combined with a binary adjustment method to fill missing values in the first dataset.
[0081] The following is an explanation of the process of filling missing values based on the data in Table 2.
[0082] Linear spline interpolation method is used to fill missing values in January and February 2016. The missing values in January correspond to the first missing value, and the missing values in February correspond to the second missing value. The first filling values in January and February are 3432.67 and 3310.33 respectively. According to the formula Obtain That is, the first filling weight for January is 0.51, and the first filling weight for February is 0.49. The filling weights for January and February in other years are replaced by the above first filling weights, that is, the first filling weight for January in all years is 0.51, and the first filling weight for February is 0.49. It is worth noting that since the first filling weights need to be adjusted later, even if the linear spline interpolation filling results of other years (such as 2017) or the average of the linear spline interpolation filling results of various years are used here as the benchmark for calculating the initial weights, it will not affect the final result.
[0083] Multiply the above first filling weight by the sum of the online physical retail sales data for January and February of each year to obtain the second filling value for each year. Taking 2016 as an example, the second filling value for January is 0.51 × 5053 = 2577.03, and the second filling value for February is 0.49 × 5053 = 2475.97. Calculate the second filling value for all years and apply it to the first dataset. This performs a preliminary filling on the first dataset to obtain the dataset to be fitted.
[0084] The dataset to be fitted is input into the first forecast model group for fitting, and the fitting error is obtained. The first forecast model group is trained using the first dataset and includes an ARIMA model, a seasonal last value forecast model, and a moving average model. The fitting error is measured using the Mean Absolute Percentage Error (MAPE). A smaller MAPE indicates a higher fitting accuracy. The fitting results are shown in Table 4. In Table 4, when the value is 0.51, the MAPE is 2.8502%.
[0085] Table 4: W1 ε The fitting effect when taking the endpoints of the interval [0, 0.51] and [0.51, 1]
[0086] Use the dichotomy method to adjust the first filling weight several times. The specific steps are as follows:
[0087] The first fill weight mentioned above will be W1 ε The weight interval [0, 1] is divided into two intervals, [0, 0.51] and [0.51, 1], W1 ε The fitting errors were calculated using the endpoints of the intervals, 0, 0.51, and 1, respectively, according to the above steps. Since the ARIMA model does not support input values of 0, weights of 0 and 1 were used instead of 0 and 1 in the calculations. The calculated fitting errors are shown in Table 4. To accurately determine the interval in which the optimal weights lie, the midpoints of the intervals [0, 0.51] and [0.51, 1] were taken and the corresponding fitting results were calculated. The results are shown in Table 5.
[0088] Table 5: W1 ε The fitting effect when taking the midpoint of the interval [0, 0.51] and [0.51, 1]
[0089] In Table 5, the MAPE value corresponding to 0.755 is 2.8383%, and the MAPE value corresponding to 0.255 is 2.8618%. Obviously, the fitting accuracy is higher when 0.755 is taken. Therefore, the optimal filling weight for January of each year can be determined. Similarly, the midpoints of the interval [0.51, 0.755] and the interval [0.755, 1] are taken and the corresponding fitting errors are calculated. The results are shown in Table 6.
[0090] Table 6: W1 ε The fitting effect when taking the midpoint of the interval [0.51, 0.755] and [0.755, 1]
[0091] Similarly, it can be determined Repeatedly use the bisection method to narrow the interval and determine whether the difference in fitting errors between the two endpoints of the interval is greater than 0.001%. If not, take the midpoint of the interval as the optimal filling weight. The interval obtained at this time is [0.846875, 0.8775]. The fitting effects of the corresponding endpoint values and the midpoint of the interval are shown in Table 7. The optimal filling weight is 0.8621875, that is, then
[0092] Table 7: W1 ε The fitting effect when taking the endpoints and midpoints of the interval [0.846875, 0.8775]
[0093] Multiply the optimal filling weight by the sum of the two missing values for each year and fill it into the first dataset to obtain the first complete dataset. Taking 2016 as an example, the final filling value for January is 0.8621875×5053≈4356.63, and the final filling value for February is 0.1378125×5053≈696.37. Calculate the final filling values for all years to obtain the first complete dataset.
[0094] In one embodiment, as shown in FIG3 , FIG3 is a flow chart of obtaining a second complete data set in the present invention; the steps of obtaining the second complete data set include:
[0095] S210: Decomposing the second data set into a first item and a second item;
[0096] S211: fitting the first item using a piecewise linear function to obtain a first filling value;
[0097] S213: interpolating the second item using a cubic spline interpolation method to obtain a second filling value;
[0098] S214: Add the first filling value and the second filling value to obtain an initial filling value;
[0099] S215: Calculate the filling weight according to the initial filling value and using formula (3) and formula (4);
[0100] In formula (3) and formula (4), k represents the year, K represents the total number of years, and W1 v is the filling weight for the first missing value, The filling weight corresponding to the second missing value, is the initial filling value for the first missing value in year k, is the initial filling value for the second missing value in year k;
[0101] S216: Multiply the filling weight by the sum of the two missing values, and fill the result into the second data set to obtain a second complete data set.
[0102] The online non-physical retail sales series has a weak trend in the early stages, a strong trend in the middle stages, and a weak trend in the later stages. Furthermore, its volatility increases over time. Using a single interpolation method cannot capture the segmented trend characteristics of online non-physical retail sales. To capture and simulate the data characteristics of online non-physical retail sales, this embodiment of the present invention uses a piecewise linear function fitting method combined with cubic spline interpolation to fill missing values in the second dataset. The specific steps are as follows:
[0103] The second dataset was decomposed into a trend term and a residual term. A piecewise linear function was used to fit the trend term to obtain the trend term filling result. Based on the growth characteristics of online non-physical retail sales data, which first increases gradually, then steepens, and finally increases gradually again, a three-segment linear function was constructed to fit the trend. Because linear function fitting does not support the use of time data as an independent variable, the monthly index of online non-physical retail sales was renumbered, with January 2015 being numbered 1, February 2015 being numbered 2, and so on, until December 2019 was numbered 60. A piecewise linear model was fitted, and missing values were filled in for the trend data to obtain the first filling value. The turning point was selected using the mean absolute error (MAE) as the evaluation metric. A smaller MAE indicates a higher fitting accuracy.
[0104] The residual term is interpolated and filled using cubic spline interpolation to obtain the second filling value. The non-physical retail data has two missing values located at the beginning of the sequence. The cubic spline interpolation method does not allow interpolation at the starting or ending points of the sequence. Therefore, before performing cubic spline interpolation, the value of the previous moment in the sequence is filled. The present invention uses the difference between the result of least squares regression of the overall data and the result of piecewise linear function fitting to fill the residual data for December 2014, that is, the residual data numbered 0.
[0105] Add the first fill value and the second fill value to get the initial fill value. Use the formula Calculate the filling weights for January and February. Multiply the filling weights by the sum of the two missing values for each year to obtain the final filling value for each year, and fill it into the second dataset to obtain the second complete dataset.
[0106] S3: Data analysis: Based on a preset time period before the target prediction period, the first complete data set and the second complete data set are subjected to data noise reduction, data rounding, or outlier correction, respectively, to obtain a first partial data set and a second partial data set.
[0107] In one embodiment, the preset time period is 12 months. For example, if the target forecast period is May 2020, to forecast online retail sales data for May 2020, it is necessary to first obtain online physical retail data and online non-physical retail data from May 2019 to April 2020 from the first complete dataset and the second complete dataset, i.e., the first and second datasets.
[0108] S4: Model establishment: using multiplication decomposition or STL decomposition to establish the first prediction model group and the second prediction model group respectively;
[0109] S5: Model prediction: importing the first part of the data set and the second part of the data set into the first prediction model group and the second prediction model group respectively to obtain a first prediction value and a second prediction value respectively;
[0110] The first prediction model group includes an ARIMA (Autoregressive Integrated Moving Average) model, a seasonal final value prediction model, and a moving average prediction model. All of the above models are trained using the first data set.
[0111] The second prediction model group includes a BP (Back Propagation) neural network model and a grey waveform model, both of which are trained using the second data set.
[0112] In one embodiment, as shown in FIG4 , FIG4 is a flow chart of obtaining the first prediction value in the present invention.
[0113] The steps for establishing the first prediction model group in S4 are as follows:
[0114] S41: Decomposing the first portion of the data set into a first component, a second component, and a third component by multiplication decomposition, wherein the first component, the second component, and the third component correspond to the first prediction model, the second prediction model, and the third prediction model;
[0115] Since online physical retail data has a stable trend and a clear cycle (seasonality), we use multiplicative decomposition to identify stable trends and fixed seasonal fluctuations. Using a decomposition period of 12, we multiplicatively decompose the first part of the dataset into a trend term, a seasonal term, and a residual term. Since the trend term component has linear characteristics, we use the linear ARIMA model to predict the trend term. Since the seasonal term component is the same in the same month of each year, we use the seasonal last value prediction model, which uses the seasonal term of the same month in previous periods to predict the future seasonal term. Since the residual fluctuates slightly and fluctuates around 1, we use a moving average model to predict the residual term.
[0116] S42: Importing the first component, the second component, and the third component into corresponding prediction models for prediction, respectively, to obtain a first component prediction result, a second component prediction result, and a third component prediction result;
[0117] The trend term, seasonal term and residual term are input into the corresponding prediction model respectively for prediction, and the trend term prediction results, seasonal term prediction results and residual term prediction results are obtained.
[0118] S43: Multiply the first component prediction result, the second component prediction result, and the third component prediction result to obtain a first prediction value.
[0119] After obtaining the first forecast value through the above process, the forecast value of online physical retail sales in the target forecast period can be obtained.
[0120] As shown in Table 8, the "spline interpolation-bisection adjustment" filling method in Table 8 is the filling method used by the present invention for the first data set, and the mean filling method is a commonly used filling method in the prior art. The smaller the MAPE, the smaller the prediction error.
[0121] Table 8: Comparison of prediction errors of online physical retail sales based on different filling methods
[0122] Taking the online physical retail data from October 2019 to December 2019 as an example, the predicted value of the "spline interpolation-bisection adjustment" filling method is closer to the true value, and its prediction error is smaller than the mean filling method. Therefore, it can be concluded that the filling method provided by the embodiment of the present invention is used for filling, and then the online physical retail sales are predicted based on the more complete data after filling, so that the prediction accuracy is improved.
[0123] In one embodiment, as shown in FIG5 , FIG5 is a flow chart of obtaining the second prediction value in the present invention.
[0124] The steps for establishing the second prediction model group in S4 are as follows:
[0125] S410: Decomposing the second data set into a fourth component and a fifth component using an STL (Seasonal-Trend decomposition procedure based on Loess) decomposition method, wherein the fourth component and the fifth component correspond to the fourth prediction model and the fifth prediction model;
[0126] The trend of online non-physical retail data changes over time and lacks a clear cycle. Therefore, STL decomposition is used to capture non-stationary trends and seasonal fluctuations. Unlike other decomposition methods, STL decomposition uses a robust locally weighted regression method to smooth the series, which can accurately identify nonlinear trends in the series. Therefore, STL decomposition is used to decompose the data into trend, seasonal, and residual terms. Since the seasonality of my country's online non-physical retail sales series is not obvious and is relatively chaotic, the seasonal and residual terms are added together to obtain a reconstruction term with large fluctuations. Therefore, the fourth component is the trend term, and the fifth component is the reconstruction term. BP neural networks are suitable for predicting smooth curves, so BP neural networks are used to predict the smooth trend term obtained by STL decomposition. Gray waveform prediction is an image prediction method that is extremely suitable for predicting highly volatile data, so gray waveform prediction is used to predict the reconstruction term.
[0127] S411: Importing the fourth component and the fifth component into corresponding prediction models for prediction, respectively, to obtain a fourth component prediction result and a fifth component prediction result;
[0128] The trend item and the reconstruction item are input into the corresponding prediction model respectively to obtain the trend item prediction result and the reconstruction item prediction result.
[0129] S412: Add the fourth component prediction result and the fifth component prediction result to obtain a second prediction value;
[0130] The trend item prediction result and the reconstruction item prediction result are added together to obtain a second prediction value, that is, the prediction value of online non-physical retail sales in the target prediction period.
[0131] As shown in Table 9, the "piecewise linear function fitting - spline interpolation" filling method in Table 9 is the filling method used by the present invention for the second data set. The mean filling method is a commonly used filling method in the prior art. A smaller MAPE indicates a smaller prediction error. Taking the online non-physical retail data from October 2019 to December 2019 as an example, the predicted value of the "piecewise linear function fitting - spline interpolation" filling method is closer to the true value, and its prediction error is smaller than that of the mean filling method. Therefore, it can be concluded that using the filling method provided by the embodiment of the present invention for filling, and then predicting online physical retail sales based on the more complete data after filling, can improve prediction accuracy.
[0132] Table 9: Comparison of prediction errors of online non-physical retail sales based on different filling methods
[0133] There is no limitation on the order of the first prediction model group and the second prediction model group in the above step S4. The first prediction model group can be executed first, or the second prediction model group can be executed first, or both can be executed separately and in parallel.
[0134] As shown in Table 10: Taking the online retail data from October 2019 to December 2019 as an example, the predicted value of the filling method used in the present invention is closer to the true value, and its prediction error is smaller than the mean filling method. Therefore, it can be concluded that the filling method provided by the embodiment of the present invention is used for filling, and then the more complete data after filling is used to predict the online physical retail sales, so that the prediction accuracy is improved.
[0135] Table 10: Comparison of online retail sales forecast errors based on different filling methods
[0136] In an online retail data prediction method provided by an embodiment of the present invention, missing values in an online retail sales dataset with unique characteristics are accurately filled based on data trends. Based on the characteristics of the complete dataset after filling, a prediction model group that matches the characteristics is used for prediction, thereby improving the accuracy of online retail sales prediction.
[0137] Example 2:
[0138] An online retail data prediction system, as shown in the system block diagram in FIG6 , uses the online retail data prediction method described in the first embodiment above. The prediction system includes a retail data collection platform, a data prediction service platform, and an execution platform.
[0139] The retail data collection platform is used to collect and transmit online retail data, and the retail data collection platform includes a data integration unit, a data transmission unit, and a feedback monitoring unit. The data integration unit collects sales record data from major sales stores on the retail platform and sends it to the statistics department server through the existing wireless and confidential data transmission means of the data transmission unit. During this process, the feedback monitoring unit, that is, the verification personnel on the existing sales platform and the automatic verification and screening system, monitor whether the data transmitted by the retail stores on the platform is true and consistent with the actual sales situation. It can mark suspicious data and suspend transmission from the data integration unit to the data transmission unit. At the same time, it notifies the relevant regulatory authorities to verify, thereby ensuring that the transmitted sales record data is consistent with reality and improving the final accuracy of online retail data prediction.
[0140] The data prediction service platform receives retail data and uses it to train a prediction model, outputting online retail data prediction results, and the data prediction service platform includes a data processing unit, a data prediction unit, and a data storage unit; wherein the data processing module is used to classify the data, separate online physical retail sales data and online non-physical retail sales data in the online retail data, and record the online physical retail sales data as a first data set and the online non-physical retail sales data as a second data set; then, the classified first data set and the second data set are transmitted to the data prediction unit for model training and prediction;
[0141] Specifically, the data prediction unit is equipped with relevant computing and processing software, and performs data filling in S2 and data analysis in S3 according to the online retail data prediction method in Example 1. In this way, the first prediction value and the second prediction value can be obtained by using multiplication decomposition or STL decomposition, and the target prediction value can be obtained after adding them. By repeatedly inputting data accumulated from multiple historical years into the data model for training, a practical and feature-matched prediction model group can be obtained, effectively improving the prediction accuracy of online retail sales. In this way, in the prediction of future years, the trained practical prediction model group can be used to calculate the monthly online retail sales forecast data for January and February, and while the forecast data continues to be stored in the data storage unit, the forecast data results are sent to the execution platform.
[0142] The execution platform receives online retail data forecasts and transmits them to the logistics, raw material supply, and sales network maintenance departments, who then implement corresponding countermeasures. Based on these forecasts, the platform estimates specific sales figures for each month, including January and February. The logistics department then notifies major express delivery companies to ensure capacity and personnel arrangements. Raw material supply factories stock up in advance based on projected sales, and the sales network maintenance department ensures network quality during peak online retail sales periods. This approach ensures smooth online sales and improves consumers' online retail experience.
[0143] The data integration unit collects logistics and sales information from online retail stores based on the Internet of Things platform. Specifically, through the Internet of Things platform, it contacts online retail stores and logistics companies to record the delivery status and receipt status of online physical orders. During the online physical retail process, after the estimated longest delivery period after the order is placed, the logistics company's delivery data and the online physical retail order transaction data are compared to verify the differences. The system identifies stores with large differences between their specific delivery data and online physical retail order transaction data, marks them through the feedback monitoring unit, and notifies the regulatory authorities for verification. In addition, if problems are found during verification, the feedback monitoring unit can promptly feedback to the data processing unit, delete the false problem data, and adjust and replace it with real data to ensure that the obtained first data set is more in line with reality, thereby improving the accuracy of online retail sales forecasts.
[0144] It should be noted that, in this document, relational terms such as first and second, etc. are merely used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations.
[0145] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the invention can be embodied in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims rather than the foregoing description, and it is intended that all variations that come within the meaning and range of equivalents of the claims be embraced therein.
Claims
1. A method for predicting online retail data, characterized in that: The method comprises the following steps: S1: Collecting data: Acquire all online physical retail sales data and online non-physical retail sales data, wherein the online physical retail sales data is recorded as a first data set, and the online non-physical retail sales data is recorded as a second data set; S2: data filling: using linear spline interpolation-bisection method to fill the first data set to obtain a first complete data set, and using piecewise linear function fitting-spline interpolation method to fill the second data set to obtain a second complete data set; S3: Data analysis: within a preset time period before the target prediction period, the first complete data set and the second complete data set are subjected to data noise reduction, data rounding or outlier correction respectively to obtain a first partial data set and a second partial data set; S4: Model establishment: using multiplication decomposition to establish the first prediction model group, and using STL decomposition to establish the second prediction model group; S5: Model prediction: importing the first part of the data set and the second part of the data set into the first prediction model group and the second prediction model group respectively, to obtain a first prediction value and a second prediction value respectively; S6: Prediction value integration: add the first prediction value and the second prediction value to obtain a target prediction value.
2. The online retail data prediction method according to claim 1, characterized in that: The specific steps of obtaining the first complete data set in S2 are as follows: S21: Filling any year of data in the first data set with a linear spline interpolation method to obtain a first filling value, where the first filling value includes two filling values, each of which corresponds to a missing value; S22: Calculate the first filling weight corresponding to the missing value using formulas (1) and (2); In formula (1) and formula (2), k here represents the year, is the first filling weight for the first missing value, the first filling weight corresponding to the second missing value, is the first fill value for a missing value in the kth year, is the first fill value for the second missing value in the kth year; S23: multiplying the first filling weight by the sum of the two missing values for each year to obtain a second filling value for each year; S24: Filling the second filling value of each year into the first data set to obtain a data set to be fitted; S25: inputting the to-be-fitted data set into the first prediction model group for fitting, and obtaining a fitting error; S26: Using the dichotomy method to adjust the first filling weight multiple times, and each time the fitting error is determined Whether the difference before and after the adjustment is greater than a preset threshold, if not, determining the adjusted first filling weight as the optimal filling weight; S27: Multiply the optimal filling weight by the sum of the two missing values for each year, and fill the result into the first data set to obtain a first complete data set.
3. The online retail data prediction method according to claim 1, characterized in that: The specific steps of obtaining the second complete data set in S2 are as follows: S210: decomposing the second data set into a first item and a second item; S211: fitting the first item using a piecewise linear function to obtain a first filling value; S213: interpolating the second item using a cubic spline interpolation method to obtain a second filling value; S214: Add the first filling value and the second filling value to obtain an initial filling value; S215: Calculate the filling weight according to the initial filling value and using formula (3) and formula (4); In formula (3) and formula (4), k represents the year, K represents the total number of years, is the filling weight for the first missing value, The filling weight corresponding to the second missing value, is the initial filling value for the first missing value in the kth year, is the initial filling value for the second missing value in the kth year; S216: Multiply the filling weight by the sum of the two missing values, and fill the second data set with the missing values to obtain a second complete data set.
4. The online retail data prediction method according to claim 1, characterized in that: The preset duration is 12 months.
5. The online retail data prediction method according to claim 1, characterized in that: The steps for establishing the first prediction model group in S4 are as follows: S41: decomposing the first part of the data set into a first component, a second component and a third component by multiplication decomposition, and the first component, the second component and the third component correspond to the first prediction model, the second prediction model and the third prediction model; S42: Importing the first component, the second component and the third component into corresponding prediction models for prediction respectively, and obtaining a first component prediction result, a second component prediction result and a third component prediction result; S43: Multiply the first component prediction result, the second component prediction result and the third component prediction result to obtain a first Predicted value.
6. The online retail data prediction method according to claim 1, characterized in that: The steps for establishing the second prediction model group in S4 are as follows: S410: decomposing the second part of the data set into a fourth component and a fifth component by using an STL decomposition method, wherein the fourth component and the fifth component correspond to a fourth prediction model and a fifth prediction model; S411: Importing the fourth component and the fifth component into corresponding prediction models for prediction, respectively, to obtain a fourth component prediction result and a fifth component prediction result; S412: Add the fourth component prediction result and the fifth component prediction result to obtain a second prediction value.
7. An online retail data prediction system, wherein the prediction system uses the online retail data prediction method according to any one of claims 1 to 6, characterized in that: The prediction system includes a retail data collection platform, a data prediction service platform and an execution platform; The retail data collection platform is used to collect and transmit online retail data, and the retail data collection platform includes a data integration unit, a data transmission unit and a feedback monitoring unit; The data prediction service platform receives retail data and uses it to train a prediction model, outputs retail data prediction results, and the data prediction service platform includes a data processing unit, a data prediction unit, and a data storage unit; The execution platform receives the retail data forecast results and transmits them to the logistics, raw material supply and sales network maintenance departments, which then take corresponding countermeasures.
8. An online retail data prediction system according to claim 7, characterized in that: The data integration unit operates based on the Internet of Things platform to collect logistics and sales information of online retail stores.
Citation Information
Patent Citations
Network retail prediction method, equipment and medium
CN114282951A
Sales prediction method based on deep learning algorithm
CN116385038A
Big data-based sales information prediction method and device
CN116596582A