Product order sales volume prediction method and system

CN120298031APending Publication Date: 2025-07-11GUANGDONG UNIV OF EDUCATION
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510266790.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-07-11

Smart Images

  • Figure CN120298031A_ABST
    Figure CN120298031A_ABST
Patent Text Reader

Abstract

The invention relates to a product order sales volume prediction method and system. The method comprises the following steps: S1, obtaining sales volume data of a target commodity in a historical time period; s2, selecting demand data of a target commodity in the training set in a historical time period, cleaning and sorting the data, extracting and deriving related features, and performing feature engineering and data preprocessing on the data; s3, performing preliminary training and evaluation by using a LightGBM regression model, using empirical parameters, and taking a root mean square error (RMSE) as an evaluation index; optimizing the model by using grid search automatic parameter adjustment; according to the sales volume difference characteristics among the sales areas, carrying out partition independent training, and analyzing the sales characteristics of each area; and S4, predicting data of a plurality of months later according to the sales volume data of the target commodity in the historical time period. By adopting the LightGBM regression model method, large-scale data sets and high-dimensional features can be processed, the model can be quickly established and optimized in the training process, and the prediction efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular, to a method and system for predicting the sales volume of product orders. Background Art

[0002] Sales volume prediction refers to estimating the cumulative sales quantity of a commodity within a future period of time. A typical application scenario is in an e-commerce platform, where a merchant sets a reasonable inventory arrangement based on the sales volume prediction results of each sold commodity to avoid losses caused by overstocking or insufficient quantity of goods. However, in existing actual production applications, the prediction results can only be preliminarily estimated based on the historical sales amount of each commodity and then adjusted in combination with manual experience. Facing a large number of commodities, it requires a large amount of manpower, and at the same time, the prediction accuracy cannot be guaranteed, resulting in high costs.

[0003] Currently, when using machine learning algorithms to complete the sales volume prediction task, it is mainly analyzed as a time series prediction problem, and the historical sales amount sequence of a commodity is used to predict the sales volume within a future period of time. Commonly used traditional methods include the Autoregressive Moving Average model (ARMA model for short), the Autoregressive Integrated Moving Average model (ARIMA model for short), etc. After these methods smooth the non-stationary time series data, they calculate the relevant parameters of the sequence for subsequent regression analysis. However, this type of method only predicts the next sequence based on the past time series and does not consider the influence of the commodity's attribute characteristics on consumers' purchase choices. Therefore, the prediction accuracy is poor, and at the same time, the similarity between commodities is not utilized, resulting in the trained prediction model being only applicable to one sequence and unable to be generalized.

[0004] Some other algorithms in the prior art require artificial selection of features as inputs for prediction, and the weight of each feature needs to be adjusted separately, resulting in very low training efficiency. Using neural network models such as LSTM and GRU, combined with the commodity's own attributes, the sales volume of a commodity within a future period of time can be predicted. However, when predicting the sales volume of subsequent days, the predicted values of the previous sales volume need to be used, which will lead to the repeated accumulation of errors when calculating the total sales volume prediction result as the prediction time window becomes longer.

[0005] How to comprehensively consider various information affecting the sales volume of commodities to improve the accuracy of sales volume prediction, reasonably design the network structure, and achieve end-to-end training is an urgent problem to be solved. Summary of the Invention

[0006] To solve the technical problems existing in the prior art, the present invention provides a method and system for predicting the sales volume of product orders. By adopting the LightGBM regression model method, it can handle large-scale data sets and high-dimensional features, and can quickly establish and optimize the model during the training process, improving the prediction efficiency.

[0007] The method of the present invention is implemented by the following technical solutions: A method for predicting the sales volume of product orders, comprising the following steps:

[0008] S1. Obtain the sales volume data of the target commodity in the historical time period, including order date, sales area code, product code, product category code, product sub-category code, sales channel name, product price, and order demand quantity; obtain the sales volume data of all commodities having the same product category number as the target commodity in the sales platform, and use this data as the basic data for model input;

[0009] S2. Select the demand data of the target commodity in the historical time period from the basic data, clean and sort the data, and extract and derive relevant features, including region, month, sales channel, promotion date, different time periods, and seasonal influence; and perform feature engineering and data preprocessing on the data, and divide the data into training sets and test sets according to daily, weekly, and monthly granularities;

[0010] S3. Use the LightGBM regression model for preliminary training and evaluation, use empirical parameters, and use the root mean square error RMSE as the evaluation index; at the same time, use grid search to automatically tune the parameters to further optimize the model; according to the sales volume difference characteristics between sales regions, perform partition-independent training, analyze the sales characteristics of each region, and further optimize the model;

[0011] S4. Predict the data for the next several months according to the sales volume data of the target commodity in the historical time period; first construct and predict the data for the first month, and add the data of this month to the basic data as the training set; then construct and predict the data for the second month until the prediction result of the final month of the target commodity is obtained.

[0012] The system of the present invention is implemented by the following technical solutions: A system for predicting the sales volume of product orders, comprising:

[0013] Data acquisition module: used to obtain the sales volume data of the target commodity in the historical time period, including single date, sales area code, product code, product category code, product sub-category code, sales channel name, product price, and order demand quantity; obtain the sales volume data of all commodities having the same product category number as the target commodity in the sales platform, and use this data as the basic data for model input;

[0014] Data processing module: Select the demand data of the target product in the historical time period from the basic data, clean and organize the data, and extract and derive relevant features, including region, month, sales channel, promotion date, different time periods, and seasonal influence; and perform feature engineering and data preprocessing on the data, and divide the data into training set and test set according to daily, weekly, and monthly granularities;

[0015] Training and evaluation module: Conduct preliminary training and evaluation by using the LightGBM regression model, use empirical parameters, and use the root mean square error RMSE as the evaluation index; at the same time, use grid search for automatic hyperparameter tuning to further optimize the model; according to the sales volume difference characteristics between sales regions, conduct independent training for each region to deeply understand the sales characteristics of each region and further optimize the model;

[0016] Prediction module: Predict the data for several months in the future based on the sales volume data of the target product in the historical time period; first construct and predict the data for the first month, and add the data of this month to the basic data as the training set; then construct and predict the data for the second month until the prediction result of the final month of the target product is obtained.

[0017] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0018] 1. By adopting the LightGBM regression model method, the present invention obtains the result of feature importance ranking, which helps to understand which factors have a greater impact on the prediction result in sales volume prediction; at the same time, it can better explain the prediction result of the model and perform feature selection according to the importance information.

[0019] 2. The present invention can maintain good prediction performance in the face of data quality problems such as missing values, outliers, and noise; at the same time, it has good generalization ability and can adapt to different data distributions and unseen situations.

[0020] 3. The present invention can process large-scale data sets and high-dimensional features, and can quickly establish and optimize the model during the training process to improve the prediction efficiency. Brief Description of the Drawings

[0021] Figure 1 is the flowchart of the method of the present invention;

[0022] Figure 2 is the flowchart of data cleaning and preprocessing;

[0023] Figure 3 is the flowchart of data prediction using LightGBM. Detailed Embodiments

[0024] The present invention will be further described in detail below in conjunction with embodiments and the accompanying drawings, but the embodiments of the present invention are not limited thereto.

[0025] Embodiment

[0026] As Figure 1 shown, a method for predicting the sales volume of product orders in this embodiment includes the following steps:

[0027] S1. Obtain the sales volume data of the target commodity in the historical time period, including the order date, sales area code, product code, product category code, product sub-category code, sales channel name, product price, and order demand quantity; obtain the sales volume data of all commodities having the same product category number as the target commodity in the sales platform, and use this data as the basic data for model input.

[0028] S2. Select the demand data of the target commodity in the historical time period from the basic data, clean and sort the data, and extract and derive relevant features, including region, month, sales channel, promotion date, different time periods, seasonal influence conditions, etc.; and perform feature engineering and data preprocessing on the data, and divide the data into training sets and test sets according to daily, weekly, and monthly granularities.

[0029] S3. Use the LightGBM regression model for preliminary training and evaluation, use empirical parameters, and use the root mean square error RMSE as the evaluation index; at the same time, use grid search to automatically tune the parameters to further optimize the model; according to the sales volume difference characteristics between sales regions, perform independent training for each region, deeply understand the sales characteristics of each region, and further optimize the model.

[0030] S4. Predict the data for the next three months according to the sales volume data of the target commodity in the historical time period; first construct and predict the data for the first month, and add the data of this month to the basic data as the training set; then construct and predict the data for the second month until the prediction results for the final three months of the target commodity are obtained.

[0031] Specifically, in this embodiment, the sales volume data of the target commodity in the historical time period includes the sales volume values of the commodity in the past period of time, and when there is no sales volume, a filling operation with 0 is performed. The sales volume of the commodity itself has a certain change pattern or periodic law. In addition, the sales volume of commodities with the same product category number and sub-category commodity numbers can also provide auxiliary information on the sales situation of the target commodity. Therefore, in the embodiments of the present invention, feature extraction is simultaneously performed on the historical sales volume of the target commodity and the sales volume of commodities with the same product number.

[0032] Specifically, in this embodiment, the sales channels include online and offline.

[0033] Specifically, in this embodiment, all the products in the sales platform that have the same major product number as the target product are obtained and are called similar products. In this embodiment, each product has a detailed product number, which represents the product-level number. At the same time, each product has a major category type number. The major category numbers of products in the same major category type are the same, while the detailed category numbers of products in the same major category type are different, and the product numbers of products in different major categories are different.

[0034] Specifically, in this embodiment, the characteristics of the promotion date are analyzed, considering the impact of different date natures on product sales. The date types are divided into two categories, namely: promotion day and non-promotion day, and are quantified with values 1 and 0 in sequence.

[0035] The characteristics of the sales channels are analyzed, considering the impact of different sales channels on product sales. The sales method types are divided into two categories, namely: online sales and offline sales, and are quantified with values 1 and 0 in sequence.

[0036] According to the high and low price ranges, considering the impact of different price ranges on product sales, the product price types are divided into four categories, namely low price, medium price, relatively high price, and high price. Among them, the distribution of product prices is viewed. Those less than 25% of the price are low prices, 25% - 50% of the price are medium prices, 50% - 75% of the price are relatively high prices, and 75% - 100% of the price are high prices.

[0037] According to the date characteristics, considering the impact of the beginning, middle, and end of the month on product sales according to which day of the month the date is. Among them, if the date is from 1 to 10 days, it is the beginning of the month; if the date is from 10 to 20 days, it is the middle of the month; otherwise, it is the end of the month.

[0038] Basic analysis of product prices is carried out. By calculating the maximum value, minimum value, average value, median value, and standard deviation of product prices, the general situation of product price data is initially understood. The demand quantity from large to small is low-price products, medium-price products, relatively high-price products, and high-price products in turn. From this, it can be known that the lower the price, the greater the product demand; on the contrary, the higher the price, the smaller the product demand.

[0039] This embodiment considers the problem of predicting product sales in the specific scenario of an e-commerce platform. Based on the differences in the basic attributes and historical sales situations among different products, it effectively utilizes the promotion activities on the platform and the price information of products in the future time, and is implemented based on a neural network model, which can greatly improve the accuracy of predicting the future sales volume of products.

[0040] Specifically, the shipment data of an enterprise for distributors reflects information such as the prices and demands of the enterprise's products in different sales regions, including: order_date (order date), sales_region_code (sales region code), item_code (product code), first_cate_code (product major category code), second_cate_code (product detailed category code), sales_chan_name (sales channel name), item_price (product price), and ord_qty (order demand quantity).

[0041] Specifically, each commodity corresponds to a static feature, which includes order date, sales region code, product code, product major category code, product detailed category code, sales channel name, price, and order demand quantity. In the embodiments of the present invention, the static feature includes these six aspects; among them, "order date" is the date of a certain demand quantity; one "product major category code" corresponds to multiple "product detailed category codes"; "sales channel name" is divided into online and offline. "Online" refers to e-commerce platforms such as Taobao and JD.com, and "offline" refers to offline physical distributors.

[0042] Specifically, in this embodiment, the features established by constructing the model specifically include:

[0043] Obtain the time series of commodity promotion activities of the target commodity in the historical time period, determine whether the platform commodity is a promotional commodity, and construct a promotion column in the dataset;

[0044] Obtain the order date of the target commodity in the historical time period, and construct time series of different dimensions according to the order date;

[0045] Obtain the sales price of the target commodity in the historical time period, and re-perform more detailed binning encoding on the price;

[0046] Obtain the order date and sales channel name of the target commodity in the historical time period, and use map mapping to encode "channels" and "beginning of the month, middle of the month, end of the month";

[0047] Obtain the season when the order of the target commodity is generated in the historical time period, and classify and label-encode each regional season specifically as "spring": 0, "summer": 1, "autumn": 2, "winter": 3;

[0048] Obtain the order date of the target commodity in the historical time period, and re-separate for each date which day of the week the current date is, and whether it is a working day, a rest day, etc.;

[0049] Introduce lag features, introduce lags for the target variable of sales, and the maximum lag used is 60 days; lag features are a classic method for transforming time series prediction problems into supervised learning problems;

[0050] Obtain the sales information of the target product within the historical time period, and calculate the rolling average of weekly sales, such as rolling minimum, maximum or sum, etc.;

[0051] Obtain the sales information of the target product within the historical time period, and create a demand trend feature, which is negative if the daily demand is greater than the average value of the entire duration (d_1 - d_1206).

[0052] As Figure 2 shown, in step S2, data cleaning includes handling missing values, removing duplicate values, and outlier detection; missing values are handled using filling or deletion strategies, duplicate values are removed through de-duplication operations, and outliers are detected and processed through statistical analysis;

[0053] In terms of feature engineering, the date feature is refined, features such as year, month, week, and day are added, and features such as promotion days and sales channels are numerically processed;

[0054] During the data preprocessing process, the maximum-minimum normalization method is used for data standardization processing, and the normalization formula is:

[0055]

[0056] where, x i represents the original value of the i-th data point, min(x) and max(x) represent the minimum and maximum values of the data set respectively, w i represents the weight of the i-th data point, and the weights of obtaining no less than N target data are w1, w2, w3, …, w N ;

[0057] The target data and weights are represented as vectors:

[0058] Data vector x = [x1, x2, …, x N ;

[0059] Weight vector W = [w1, w2, …, w N .

[0060] Specifically, in this embodiment, by deeply analyzing the data set, for predictions at different granularities, time series with different dimensions are obtained according to the order date, and derivative features suitable for model prediction exploration are constructed, specifically including:

[0061] For predictions at the daily granularity, a new feature 'D' is constructed, which specifically represents the number of days within the historical time period;

[0062] For the prediction with a weekly granularity, a new feature 'W' is constructed, which specifically represents the number of weeks within the historical time period;

[0063] For the prediction with a monthly granularity, a new feature 'M' is constructed, which specifically represents the number of months within the historical time period.

[0064] Obtain the dataset and perform preliminary modeling data preparation.

[0065] Obtain the order dates of the target product within the historical time period and construct time series of different dimensions, specifically including:

[0066] According to the requirements of time series analysis, sort the information of the target product within the historical time period in ascending order by the order date;

[0067] Divide the dataset into a training set and a test set. Specifically, the first 70% of the data is used as the training set for model training and parameter estimation; the last 30% of the data is used as the test set for evaluating the model's performance and generalization ability;

[0068] Obtain the order demand quantity of the target product within the historical time period as the target column; all other columns are used as feature columns;

[0069] Divide the training set and the test set into feature data and target data respectively for subsequent time series analysis and model training;

[0070] Obtain all feature columns in the dataset and assign the result to variable X to maintain the feature datasets of the training set and the test set in subsequent operations.

[0071] Perform feature engineering and feature encoding for subsequent model training analysis, specifically including:

[0072] Obtain the product code, product category code, and product sub-category code of the target product within the historical time period and perform data type conversion respectively. Specifically, in this embodiment, the feature columns are converted to string type;

[0073] Obtain three types of features of the target product within the historical time period and perform one-hot encoding to convert the categorical variables into binary features;

[0074] Through in-depth analysis of the dataset and construction of derivative features, a Light Gradient Boosting Machine (LGBM) model is used for prediction exploration.

[0075] Specifically, in this embodiment, the LightGBM model is used to predict the sales data, as Figure 3 shown, the steps of the prediction method include:

[0076] S51. Obtain the training set data;

[0077] The LightGBM (Lightweight Gradient Boosting Machine) is based on Gradient Boosting, which is an ensemble learning method that improves the performance of the model by constructing multiple trees. Specifically:

[0078] Use the Leaf-wise growth strategy. Each time, find the leaf with the largest split gain from all current leaves, and then split it. Repeat this process in a loop.

[0079] Update and obtain the dataset, define a parameter dictionary that contains a series of hyperparameter settings for the model, and pass it to the LightGBM model object.

[0080] Preliminarily train and predict the model, and calculate the Root Mean Square Error (RMSE) as the performance metric of the model. The calculation formula is as follows:

[0081]

[0082] Specifically, RMSE represents the average deviation between the predicted value and the true value. The smaller its value, the more accurate the prediction result, that is, the better the model performance.

[0083] To further improve the accuracy of the model, grid search for automatic hyperparameter tuning is introduced here. Specifically, the grid search algorithm lists the possible values of each hyperparameter, then traverses the entire parameter grid, that is, the parameter space, and evaluates the performance of each possible parameter combination in this grid. Finally, return the combination of model hyperparameters with the best performance.

[0084] S52. Calculate the expected model using the LightGBM model.

[0085] According to the target requirements, it is necessary to predict the data for the next three months. Subsequently, the data for the first month will be predicted first, and then the data for this month will be added to the original data as the training set, and then the data for the second month will be predicted, and so on, until the final prediction result is obtained.

[0086] For the prediction at the daily granularity, based on the features constructed above, construct the data to be predicted for the first month, with the daily demand being 0, and add it to the original data for subsequent model establishment.

[0087] Specifically, concatenate the values of the four columns of the sales area code, product category code, product subcategory code, and product code according to the specified format, and store the result in a column named "zuhe", that is, the subsequent combined value.

[0088] Remove duplicates from the unique values in the combined value column, and use a function to obtain the number of unique values after deduplication.

[0089] Generate the information of the storage date, combined value, and order demand for predicting the first month.

[0090] Construct and improve features suitable for subsequent model building based on the above data information;

[0091] Since the delay caused by lag features introduced many null values, the data for the first 60 days was deleted because a 60-day delay had been introduced;

[0092] Obtain a new feature dataset, train LightGBM regression models for multiple sales regions, and plot the changes in the loss functions (RMSE and MAPE) with the number of training rounds, specifically including:

[0093] During the training process, extract the loss function values of the training set and the test set, and plot two subgraphs representing the changes in RMSE and MAPE with the number of training rounds respectively;

[0094] Filter the data corresponding to the region according to the sales region code, and divide the training set, validation set and their target variables (order demand) for training and evaluating the performance of the model;

[0095] Process the predicted values, replace the values less than 0 with 0, indicating the occurrence of outliers, otherwise keep the integer part of other values;

[0096] Read the data of the first month constructed previously and replace the predicted demand values of each product in January;

[0097] Extract the data of January, filter the required columns: order date, sales region code, product code, product category code, product sub-category code and order demand; calculate the demand of the target product within the time period;

[0098] According to the average value of the predicted demand grouped by sales region code, product category code and product sub-category code of the target product within the time period, as the predicted value of the new product;

[0099] Combine the product combinations of the two datasets, compare whether there are the same products in the two datasets, if a product combination does not exist in the January dataset, it is regarded as a new product; merge the two datasets and delete the product combination column;

[0100] The predicted value of the target new product is the average value of predicting the same type of product;

[0101] When predicting the second month subsequently, merge the data predicted in the first month, construct the data for the second month on this basis, and finally predict the data for February;

[0102] When predicting the third month, the steps are the same as above;

[0103] For the prediction at the weekly granularity, based on the dataset for the daily granularity prediction, the W column was reconstructed; similarly, the W column is the number of weeks from the current date.

[0104] Specifically, for the feature engineering at the weekly granularity, since the span of a week has become larger, an additional rolling average of a 14-day sliding window statistic was constructed for more accurate prediction.

[0105] For the prediction at the monthly granularity, based on the dataset for the weekly granularity prediction, the M column was reconstructed; similarly, the M column is the number of months from the current date.

[0106] Specifically, for the feature engineering at the monthly granularity, since the span of a month has become larger, an additional rolling average of a 30-day sliding window statistic was constructed for more accurate prediction.

[0107] S53. Conduct model testing on the expected data results.

[0108] In the embodiments of the present application, the root mean square error RMSE and the mean absolute error MAE are used to test the expected sales volume data obtained by the LightGBM model, and the error between the predicted value and the actual value is calculated.

[0109] Based on the same inventive concept, the present invention provides a product order sales volume prediction system, including:

[0110] Data acquisition module: used to acquire the sales volume data of the target commodity in the historical time period, including single date, sales area code, product code, product category code, product sub-category code, sales channel name, product price, and order demand quantity; acquire the sales volume data of all commodities with the same product category number as the target commodity in the sales platform, and use these data as the basic data for model input.

[0111] Data processing module: select the demand data of the target commodity in the historical time period from the basic data, clean and sort the data, and extract and derive relevant features, including region, month, sales channel, promotion date, different time periods, seasonal influence situation, etc.; and perform feature engineering and data preprocessing on the data, and divide the data into training sets and test sets according to the daily, weekly, and monthly granularities.

[0112] Training and evaluation module: conduct preliminary training and evaluation by using the LightGBM regression model, use empirical parameters, and use the root mean square error RMSE as the evaluation index; at the same time, use grid search for automatic parameter tuning to further optimize the model; according to the sales volume difference characteristics between each sales area, conduct independent training for each area to deeply understand the sales characteristics of each area and further optimize the model.

[0113] Prediction module: Predict the data for the next three months based on the sales data of the target product in the historical time period; first construct and predict the data for the first month, and add the data of this month to the basic data as the training set; then construct and predict the data for the second month until the final three-month prediction results of the target product are obtained.

[0114] The product order sales prediction system of the present invention discovers, through data analysis, that there are seasonal patterns / holiday patterns / high correlations of promotional activities in sales, that is, different products play different roles in different time and space, indirectly affecting the sales volume of products; at the same time, considering the problem of predicting the sales volume of products in the specific scenario of an e-commerce platform, based on the differences in the basic attributes and historical sales situations of different products, effectively utilize the platform and product-related information in different regions in the future time, and is implemented based on the LightGBM model, which can improve the accuracy of predicting the future sales volume of products.

[0115] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. A method for predicting the sales volume of product orders, characterized in that, It includes the following steps: S1. Obtain the sales data of the target product within the historical time period, including order date, sales area code, product code, product category code, product sub-category code, sales channel name, product price, and order demand quantity; obtain the sales data of all products with the same product category number as the target product in the sales platform, and use this data as the basic data for model input; S2. Select the demand data of the target product in the historical time period from the basic data, clean and organize the data, and extract and derive relevant features, including region, month, sales channel, promotion date, different time periods, and seasonal influence; perform feature engineering and data preprocessing on the data, and divide the data into training sets and test sets according to daily, weekly, and monthly granularities; S3. Use the LightGBM regression model for preliminary training and evaluation, use empirical parameters, and use the root mean square error RMSE as the evaluation index; at the same time, use grid search for automatic hyperparameter tuning to further optimize the model; according to the sales volume difference characteristics between sales regions, perform independent training for each region, analyze the sales characteristics of each region, and further optimize the model; S4. Predict the data for the next few months based on the sales data of the target product within the historical time period; first construct and predict the data for the first month, and add the data for this month to the basic data as the training set; then construct and predict the data for the second month until the prediction result for the final month of the target product is obtained.

2. The method for predicting the sales volume of a product order according to claim 1, wherein The features constructed in step S1 specifically include: Obtain the time series of product promotion activities of the target product within the historical time period, determine whether the platform product is a promotional product, and construct a promotion column in the dataset; Obtain the order dates of the target product within the historical time period, and construct time series with different dimensions based on the order dates; Obtain the sales prices of the target product within the historical time period, and re-perform more detailed binning encoding on the prices; Obtain the order dates and sales channel names of the target product within the historical time period, and use map mapping to encode the "channel" and "beginning, middle, end of the month"; Obtain the seasons when the orders of the target product are generated within the historical time period, and classify and label-encode each regional season specifically as "spring": 0, "summer": 1, "autumn": 2, "winter": 3; Obtain the order dates of the target product within the historical time period, and re-separate which day of the week the current date is, and whether it is a working day or a rest day for each date; Introduce lag features, introduce lags for the target variable of sales, and set the maximum delay days to be used; lag features transform the time series prediction problem into a supervised learning problem; Obtain the sales information of the target product within the historical time period, and calculate the rolling average, rolling minimum, maximum, or sum of weekly sales; Obtain the sales information of the target product within the historical time period, create a demand trend feature, and if the daily demand is greater than the average value of the entire duration (d_1 - d_1206), it is negative.

3. A method for predicting the sales volume of product orders according to claim 1, characterized in that, In step S2, data cleaning includes handling missing values, removing duplicate values, and outlier detection; for handling missing values, strategies of filling or deleting are adopted, duplicate values are removed through de-duplication operations, and outliers are detected and processed through statistical analysis.

4. A method for predicting the sales volume of product orders according to claim 1, characterized in that, In step S2, for feature engineering, the date feature is refined, features of year, month, week, and day are added, and promotional day and sales channel features are numerically processed; During the data preprocessing process, the maximum-minimum normalization method is used for data standardization, and the normalization formula is: where x i represents the original value of the i-th data point, min(x) and max(x) respectively represent the minimum and maximum values of the data set, and w i represents the weight of the i-th data point. The weights of no less than N target data are w1, w2, w3, …, w N ; The target data and weights are represented as vectors: Data vector x = [x1, x2, …, x N ; The weight vector W = [w1, w2, …, w N .

5. A method for predicting the sales volume of a product order according to claim 1, characterized in that, By deeply analyzing the dataset, for predictions at different granularities, time series with different dimensions are obtained based on the order date, and derivative features suitable for model prediction exploration are constructed, specifically including: For predictions at the daily granularity, a new feature 'D' is constructed, which specifically represents the number of days within the historical time period; For predictions at the weekly granularity, a new feature 'W' is constructed, which specifically represents the number of weeks within the historical time period; For predictions at the monthly granularity, a new feature 'M' is constructed, which specifically represents the number of months within the historical time period.

6. A method for predicting the sales volume of a product order according to claim 5, characterized in that, Obtain the order dates of the target product within the historical time period, and construct time series with different dimensions, specifically including: According to the requirements of time series analysis, the information of the target product within the historical time period is sorted in ascending order according to the order date; The dataset is divided into a training set and a test set. The first 70% of the data is used as the training set for model training and parameter estimation; the last 30% of the data is used as the test set for evaluating the model's performance and generalization ability; Obtain the order demand quantity of the target product within the historical time period as the target column; all other columns are used as feature columns; The training set and the test set are respectively divided into feature data and target data for subsequent time series analysis and model training; Obtain all feature columns in the dataset, and assign the result to variable X to maintain the feature datasets of the training set and the test set in subsequent operations.

7. A method for predicting the sales volume of product orders according to claim 1, characterized in that, Perform feature engineering and feature encoding for subsequent model training analysis, specifically including: Obtain the product code, product category code, and product sub-category code of the target product within the historical time period, and perform data type conversion respectively, and convert the feature columns to string type; Obtain the three types of features of the target product within the historical time period, and use one-hot encoding for processing to convert categorical variables into binary features.

8. A method for predicting the sales volume of product orders according to claim 1, characterized in that In step S3, the LightGBM regression model is used to predict the sales volume data, and the specific process is as follows: S51. Obtain the training set data; Construct multiple trees: Use the Leaf-wise growth strategy. Each time, find the leaf with the largest split gain from all current leaves, and then split, and so on in a loop; Update and obtain the dataset, define a parameter dictionary, which contains a series of hyperparameter settings of the model, and pass it to the LGBM model object; Preliminarily train and predict the model, and calculate the root mean square error RMSE as the performance metric of the model. The calculation formula is as follows: Among them, RMSE represents the average deviation between the predicted value and the true value; Grid search for automatic hyperparameter tuning is introduced. The grid search algorithm lists the possible values of each hyperparameter, then traverses the entire parameter grid, i.e., the parameter space, and evaluates the performance of each possible parameter combination in this grid. Finally, it returns the combination of model hyperparameters with the best performance; S52. Use the LightGBM model to calculate the expected model; According to the target requirements, predict the data for several months in the future. First, predict the data for the first month, then add the data of this month to the original as the training set, and then predict the data for the second month until the prediction results for the final month are obtained; For the prediction with a daily granularity, based on the constructed features, construct the data to be predicted for the first month, with the daily demand being 0, and add it to the original data; Concatenate the values of the four columns of sales area code, product category code, product sub-category code, and product code in the specified format, and store the result in the column named "zuhe", that is, the subsequent combined value; Remove duplicates from the unique values in the combined value column, and use a function to obtain the number of unique values after deduplication; Generate the information of the storage date, combined value, and order demand for predicting the first month; Construct and improve the features suitable for subsequent model establishment; Obtain the new feature dataset, train the LightGBM regression models for multiple sales areas, and plot the change of the loss function with the number of training rounds. Specifically, it includes: During the training process, extract the loss function values of the training set and the test set, and plot two subgraphs representing the changes of RMSE and MAPE with the number of training rounds respectively; Filter the data corresponding to the corresponding area according to the sales area code, and divide the training set, validation set, and their target variables for training and evaluating the performance of the model; Process the predicted values, replace the values less than 0 with 0, indicating the occurrence of outliers, and otherwise retain the integer part of other numerical values; Read the data of the first month constructed before, and replace the predicted demand values of each product in January; Extract the data of the first month, and filter the required columns: order date, sales area code, product code, product category code, product sub-category code, and order demand; calculate the demand of the target commodity within the time period; According to the average value of the predicted demand grouped by sales area code, product category code, and product sub-category code of the target commodity within the time period, use it as the predicted value of the new product; Combine the products of the two datasets, compare whether there are the same products in the two datasets. If a certain product combination does not exist in the January dataset, it is regarded as a new product; the two datasets are merged, and the product combination column is deleted; The predicted value of the target new product is the average value of predicting the same type of products; When predicting the second month in the future, merge the predicted data of the first month, and on this basis, construct the data of the second month and predict the data of the second month; When predicting the nth month, the steps are the same as above; For the prediction with a weekly granularity, based on the dataset for daily granularity prediction, reconstruct the W column; similarly, the W column is the number of weeks from the current date; For the feature engineering with a weekly granularity, an additional rolling average of a 14-day sliding window statistic is constructed; For the prediction at the monthly granularity, based on the dataset predicted at the weekly granularity, column M was reconstructed; similarly, column M is the number of months from the current date. For the feature engineering at the monthly granularity, an additional rolling average of a 30-day sliding window statistic was constructed. S53. Conduct model testing on the expected data results. The root mean square error (RMSE) and the mean absolute error (MAE) were used to test the expected sales volume data obtained from the LightGBM model, and the error between the predicted value and the actual value was calculated.

9. A product order sales volume prediction system, characterized in that, Including: Data acquisition module: used to obtain the sales volume data of the target product within the historical time period, including single date, sales area code, product code, product category code, product sub-category code, sales channel name, product price, and order demand; obtain the sales volume data of all products with the same product category number as the target product in the sales platform, and use this data as the basic data for model input. Data processing module: select the demand data of the target product within the historical time period in the training set, clean and organize the data, and extract and derive relevant features, including region, month, sales channel, promotion date, different time periods, and seasonal impact; perform feature engineering and data preprocessing on the data, and divide the data into training sets and test sets according to daily, weekly, and monthly granularities. Training and evaluation module: conduct preliminary training and evaluation by using the LightGBM regression model, use empirical parameters, and use the root mean square error (RMSE) as the evaluation index; at the same time, use grid search for automatic hyperparameter tuning to further optimize the model; according to the sales volume difference characteristics between sales regions, conduct independent training for each region to deeply understand the sales characteristics of each region and further optimize the model. Prediction module: predict the data for the next few months based on the sales volume data of the target product within the historical time period; first construct and predict the data for the first month, and add the data of this month to the original as the training set; then construct and predict the data for the second month until the prediction result of the final month of the target product is obtained.

Citation Information

Cited By

  • Multi-strategy dynamic prediction method, system and device based on seasonal commodities and storage medium

    CN120707054A