Wine market demand prediction system based on big data
Through the alcohol market demand prediction system combined with big data and random forest algorithm, the problem of low prediction accuracy in the existing technology is solved, and more accurate market demand prediction and sales guidance are achieved.
Patent Information
- Application Number
- CN202510474377.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing alcohol consumer purchase prediction method is based on a simple mathematical model, with low prediction results, unable to accurately predict consumption trends, and insufficient data processing, resulting in the model being unable to provide effective marketing guidance.
The wine market demand forecasting system based on big data is adopted, and multiple aspects of data are obtained through the prediction platform module, combined with the random forest algorithm for data processing and analysis, including consumer surveys, seller data, inventory data, etc., and the random forest classification model is used for consumer classification and prediction.
It achieves a more accurate forecast of the demand for alcohol market, can determine peak seasons and off-seasons, understand consumer preferences and sales, and improves the accuracy of forecasts and the diversity of data.
Smart Images

Figure CN120338868A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of liquor market demand forecasting, and particularly to a liquor market demand forecasting system based on big data. Background Art
[0002] With the development of information technology, the prediction of consumer purchases has attracted more and more attention from various enterprises. At present, the prediction of liquor consumer purchases is mainly based on the prediction of simple mathematical models. By means of consumer consumption data, fitting consumption curves, and predicting consumer consumption trends, the preprocessing of data is simple, the degree of information mining is low, and the accuracy of prediction results is low, which cannot provide effective guidance in the liquor marketing process. At the same time, liquor consumption data has complex characteristics and redundant types. When directly using classification models for operations, the operations are complex and the amount of operations is large, and overfitting is likely to occur; the liquor sample set has uneven distribution, and the samples with high-consumption labels are usually less. Training classification models with unbalanced sample sets is likely to be biased towards the majority class, resulting in the situation where the model cannot accurately predict other classes.
[0003] A liquor consumption prediction method based on random forest with the publication number of CN118735580A includes: Step 1: Collect consumer attribute data as modeling data, Step 2: Screen consumer attribute data and retain features with higher correlation with the sample label, Step 3: Balance the sample set, Step 4: Use the data of the sample set to train a random forest classification model, and use the random forest classification model to classify consumers according to the recent data of consumers and predict the future liquor consumption ability of consumers.
[0004] In the above technical solution, the required data cannot be fully obtained, the amount of data used for prediction cannot be expanded, and the accuracy of prediction cannot be fully guaranteed, so improvements are needed. Summary of the Invention
[0005] The purpose of the present invention is to solve the disadvantages existing in the prior art, and to propose a liquor market demand forecasting system based on big data.
[0006] In order to achieve the above purpose, the present invention adopts the following technical solution:
[0007] A liquor market demand forecasting system based on big data includes a prediction platform module, and a background data module, a consumer survey module, and a sales prediction module are provided in the prediction platform module;
[0008] Both the background data module and the consumer survey module are connected to the sales prediction module;
[0009] The background data module includes a supplier data acquisition module, a logistics data acquisition module, a distributor data acquisition module, and an inventory data acquisition module;
[0010] The consumer survey module includes an online survey data acquisition module and a consumer venue questionnaire module;
[0011] The chart output module connected to the prediction platform module is connected to the sales prediction module.
[0012] Compared with the prior art, the present application can fully obtain data in multiple aspects such as sales and supply, so as to fully ensure the diversity and accuracy of the data required in the prediction system, and ensure that the obtained data are all the data required for prediction and have strong relevance, which can ensure the accuracy of prediction, and uses the random forest method to fully utilize the obtained data.
[0013] Preferably, the consumer venues in the consumer venue questionnaire module include: KTVs, hotels, and restaurants; the questionnaires in the consumer venue questionnaire module include: paper survey documents and electronic survey documents.
[0014] Furthermore, it can fully obtain the liquor preferences and liquor sales situations of consumers in these areas, and can determine the peak season and off-season, as well as the required amount of liquor during the corresponding time periods.
[0015] Preferably, the distributor data acquisition module includes: an offline distributor data module and an online distributor data module;
[0016] The data sources of the offline distributor data module include large supermarkets, shopping malls, liquor specialty stores, stores that can sell liquor, and group purchase distributors;
[0017] The data sources of the online distributor data module include: sales data of multiple e-commerce platforms and sales data of multiple live broadcast platforms.
[0018] Furthermore, it fully counts the centralized out-of-stock situations, sales situations, and return and replacement situations, and fully and comprehensively understands the sales channels.
[0019] Preferably, the inventory data acquisition module includes a manufacturer inventory data module and a supplier inventory data module;
[0020] The inventory data acquisition module includes a monthly inventory data module, a year-on-year inventory data module, a month-on-month inventory data module, and an inventory prediction module.
[0021] Furthermore, it understands the inventory data.
[0022] Preferably, the sales prediction module uses the random forest algorithm for prediction; the random forest algorithm includes the following steps: data preparation, bootstrap sampling, constructing decision trees, and integrated prediction.
[0023] Furthermore, make full use of the data to improve the efficiency and quality of the algorithm.
[0024] Preferably, the random forest algorithm includes:
[0025] Step 1: Data preparation: Collect consumer attribute data as modeling data, perform data formatting on the consumer attribute data, and normalize the feature values of the consumer attribute data to unify the order of magnitude of different attribute data. The consumer attribute data includes basic data, purchase data, and behavior data.
[0026] Step 2: Bootstrap sampling to screen consumer attribute data, including:
[0027] Step 21: Calculate the Pearson coefficient between the features of the consumer attribute data, remove highly correlated features, and obtain a preliminary screening result.
[0028] Step 22: According to the preliminary screening result, form a sample set. Denote the feature data set of the sample set as X and the number of features as m. Select to ignore the first feature M1, use the remaining m - 1 features for training and verification, calculate the prediction accuracy rate, repeat the calculation of the prediction accuracy rate for ignoring other features, obtain the influence relationship of each feature on the prediction accuracy rate of the random forest classification model, and retain the features with higher correlation with the sample label according to the prediction accuracy rate ranking.
[0029] Step 3: Construct a decision tree and balance the sample set: Calculate the sample set balancing target value for the number of samples under different sample labels in the sample set. For the number of samples exceeding the sample set balancing target value, remove samples until the sample set quantity reaches the sample set balancing target value. For the number of samples less than the sample set balancing target value, add samples until the sample set quantity reaches the sample set balancing target value.
[0030] Step 4: Integrated prediction: Use the data of the sample set to train the random forest classification model, and use the random forest classification model to classify consumers according to the recent data of consumers and predict the future consumption ability of consumers for alcoholic beverages.
[0031] Preferably, in step 21, the expression of the Pearson coefficient r is used:
[0032]
[0033] Calculate the degree of correlation between feature x and feature y. Among them, when |r| ≥ 0.8, x and y are highly correlated; when 0.5 ≤ |r| < 0.8, x and y are moderately correlated; when 0.3 ≤ |r| < 0.5, x and y are lowly correlated; when |r| < 0.3, x and y are considered uncorrelated.
[0034] Preferably, in step 22, training and validation are performed, including:
[0035] First, randomly divide the feature set of the remaining m - 1 features into k parts, select the first part of the samples as the validation set, and the remaining k - 1 parts of the samples as the training set for training the random forest classification model. Then use the random forest classification model to predict the sample categories of the validation set and calculate the prediction accuracy rate. Next, select the second part of the samples as the validation set, and the remaining k - 1 parts of the samples as the training set, and repeat the process of training the model and calculating the prediction accuracy rate until the prediction accuracy rate of the random forest classification model has been verified for each part of the samples. Finally, take the average of the prediction accuracy rates calculated each time as the model prediction accuracy rate for evaluating the ignored feature M1.
[0036] Preferably, in step 3, use the formula:
[0037]
[0038] Calculate the sample set balancing target value Ntarget, where n is the number of labels and Yi is the number of samples under the i - th label;
[0039] Construct a data set with M + 1 dimensions using the features retained in step 2 and the sample labels, and calculate the minimum spatial distance of the data set:
[0040]
[0041] x1, x2 are two samples in the data set, D(x1, x2) represents the distance between the two samples, M + 1 is the dimension of the samples, x1,j represents the j - th dimension of sample 1, and x2,j represents the j - th dimension of sample 2. Remove the data with the minimum D(x1, x2) value until the number of the sample set reaches the balancing target value Ntarget.
[0042] The beneficial effects of the present invention are:
[0043] 1. Sufficiently obtain the liquor preferences and liquor sales situations of consumers in these regions, and can determine the peak season and off - season, as well as the required amount of liquor during the corresponding time periods; fully count the centralized out - of - warehouse situations, sales situations, and return and exchange situations, and fully and comprehensively understand the sales channels; understand the inventory data; through the above - mentioned solutions, fully obtain liquor sales and inventory data;
[0044] 2. The algorithm of the random forest can ensure the full use of data to obtain more accurate prediction data. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 It is a connection block diagram of the liquor market demand prediction system based on big data proposed by the present invention;
[0046] Figure 2 The block diagram of the consumer venue questionnaire in the liquor market demand prediction system based on big data proposed by the present invention;
[0047] Figure 3 The block diagram of the distributor data acquisition in the liquor market demand prediction system based on big data proposed by the present invention;
[0048] Figure 4 The block diagram of the inventory data acquisition in the liquor market demand prediction system based on big data proposed by the present invention;
[0049] Figure 5 The sales prediction step diagram in the liquor market demand prediction system based on big data proposed by the present invention. Detailed implementation manners
[0050] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments.
[0051] Refer to Figures 1-5 , the liquor market demand prediction system based on big data includes a prediction platform module. During actual production and preparation, corresponding programs can be set in the prediction platform to fully enable the programs to process the acquired data more accurately, facilitating the prediction of the liquor market demand. The prediction platform module is provided with a background data module, a consumer survey module, and a sales prediction module; both the background data module and the consumer survey module are connected to the sales prediction module to fully import the data into the sales prediction module for the sales prediction module to perform corresponding predictions.
[0052] Refer to Figure 1 , the background data module includes a supplier data acquisition module, a logistics data acquisition module, a distributor data acquisition module, and an inventory data acquisition module; it is used to fully acquire data and ensure the quality of the data.
[0053] Refer to Figure 1 , the consumer survey module includes an online survey data acquisition module and a consumer venue questionnaire module, which are used to understand consumers' preferences and determine the proportion of various types of liquor in the corresponding liquor consumption, and obtain more real market feedback through multi-point sampling.
[0054] Refer to Figure 1 , the chart output module connected to the prediction platform module, the chart output module is connected to the sales prediction module, and after the prediction is completed, it can be directly output in the form of a chart, providing more explicit data support for the staff.
[0055] Refer to Figure 1 、2 , the consumption places in the consumption place questionnaire module include: KTVs, hotels, and restaurants; the questionnaires in the consumption place questionnaire module include: paper survey documents and electronic survey documents. When actually obtaining, a broader consideration is taken, and all areas where liquor sales can be obtained are considered.
[0056] Refer to Figure 1 , 3 , the distributor data acquisition module includes: an offline distributor data module and an online distributor data module; comprehensively understand the sales situation, and fully obtain the return and exchange situation, as well as the corresponding feedback from the purchasers.
[0057] Refer to Figure 1 , 3 , the data sources of the offline distributor data module include large supermarkets, shopping malls, liquor specialty stores, stores that can sell liquor, and group purchase distributors; the data sources of the online distributor data module include: sales data from multiple e-commerce platforms and sales data from multiple live streaming platforms; expand the data acquisition sources to truly achieve big data acquisition.
[0058] Refer to Figure 1 , 4 , the inventory data acquisition module includes a manufacturer inventory data module and a supplier inventory data module; the inventory data acquisition module includes a monthly inventory data module, a year-on-year inventory data module, a month-on-month inventory data module, and an inventory prediction module.
[0059] Refer to Figure 1 , 5 , the sales prediction module uses the random forest algorithm for prediction; the random forest algorithm includes the following steps: data preparation, bootstrap sampling, constructing decision trees, and integrated prediction. The random forest algorithm includes:
[0060] Step 1: Data preparation: Collect consumer attribute data as modeling data, perform data formatting processing on the consumer attribute data, and perform normalization processing on the feature values of the consumer attribute data to unify the order of magnitude of different attribute data to the same order of magnitude. The consumer attribute data includes basic data, purchase data, and behavior data.
[0061] Step 2: Bootstrap sampling, screening consumer attribute data, including:
[0062] Step 21: Calculate the Pearson coefficient between the features of the consumer attribute data, remove highly correlated features, and obtain a preliminary screening result.
[0063] Step 22: According to the preliminary screening results, a sample set is formed. Denote the set of feature data of the sample set as X, the number of features as m. Select to ignore the first feature M1, use the remaining m - 1 features for training and validation, calculate the prediction accuracy rate, repeat the calculation of the prediction accuracy rate for ignoring other features, obtain the influence relationship of each feature on the prediction accuracy rate of the random forest classification model, and according to the ranking of the prediction accuracy rate, retain the features with higher correlation with the sample labels;
[0064] Step 3: Construct a decision tree and balance the sample set: Calculate the sample set balancing target value for the number of samples under different sample labels in the sample set. For the number of samples exceeding the sample set balancing target value, remove the samples until the number of the sample set reaches the sample set balancing target value. For the number of samples less than the sample set balancing target value, add samples until the number of the sample set reaches the sample set balancing target value;
[0065] Step 4: Integrated prediction: Use the data of the sample set to train the random forest classification model, and use the random forest classification model to classify consumers according to the recent data of consumers, and predict the future consumption ability of consumers for alcoholic beverages; Before using the data of the sample set to train the random forest classification model in Step 4, set the number of decision trees of the random forest classification model to 300, the number of randomly selected features to m which is the number of sample features, do not limit the growth depth of the decision tree, and only stop splitting the decision tree when the Gini index of the node cannot decrease, and use the balanced sample set data to train the random forest classification model.
[0066] In the present invention, it is characterized in that the expression of the Pearson coefficient r is used in Step 21:
[0067]
[0068] Calculate the degree of correlation between feature x and feature y. Among them, when |r| ≥ 0.8, x and y are highly correlated; when 0.5 ≤ |r| < 0.8, x and y are moderately correlated; when 0.3 ≤ |r| < 0.5, x and y are lowly correlated; when |r| < 0.3, x and y are regarded as uncorrelated.
[0069] Step 21: Calculate the Pearson coefficient between the features of the consumer attribute data, remove the highly correlated features, and obtain the preliminary screening results.
[0070] In the present invention, the training and validation in Step 22 include:
[0071] First, randomly divide the feature set of the remaining m - 1 features into k parts. Select the first part of the samples as the validation set, and the remaining k - 1 parts of the samples as the training set to train the random forest classification model. Then use the random forest classification model to predict the sample categories of the validation set and calculate the prediction accuracy rate. Next, select the second part of the samples as the validation set, and the remaining k - 1 parts of the samples as the training set, and repeat the process of training the model and calculating the prediction accuracy rate until the prediction accuracy rate of the random forest classification model has been verified for each part of the samples. Finally, take the average of the prediction accuracy rates calculated each time as the prediction accuracy rate of the model for evaluating the ignored feature M1.
[0072] According to the preliminary screening results, construct a sample set. Denote the feature data set of the sample set as X, and the number of features as m. Select to ignore the first feature M1, and use the remaining m - 1 features for training and verification to calculate the prediction accuracy rate. Repeat the calculation of the prediction accuracy rate for ignoring other features to obtain the influence relationship of each feature on the prediction accuracy rate of the random forest classification model. According to the prediction accuracy rate ranking, retain the features with a higher correlation with the sample labels. The sample balancing module balances the sample set: calculate the sample set balancing target value for the number of samples under different sample labels in the sample set. For the number of samples exceeding the sample set balancing target value, remove the samples until the number of the sample set reaches the sample set balancing target value. For the number of samples less than the sample set balancing target value, add samples until the number of the sample set reaches the sample set balancing target value. The model prediction module uses the data of the sample set to train the random forest classification model, and uses the random forest classification model to classify consumers based on the recent data of consumers and predict the future consumption ability of consumers for alcoholic beverages.
[0073] In the present invention, in step 3, the formula:
[0074]
[0075] is used to calculate the sample set balancing target value Ntarget, where n is the number of labels, and Yi is the number of samples under the i-th label;
[0076] Construct a data set with M + 1 dimensions from the features and sample labels retained in step 2, and calculate the minimum spatial distance of the data set:
[0077]
[0078] x1 and x2 are two samples in the data set, D(x1, x2) represents the distance between the two samples, M + 1 is the dimension of the samples, x1,j represents the j-th dimension of sample 1, and x2,j represents the j-th dimension of sample 2. Remove the data with the minimum D(x1, x2) value until the number of the sample set reaches the balancing target value Ntarget.
[0079] The above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution of the present invention and its inventive concept, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.
Claims
1. A liquor market demand prediction system based on big data, including a prediction platform module, characterized in that: The prediction platform module is provided with a background data module, a consumer survey module and a sales prediction module; Both the background data module and the consumer survey module are connected to the sales prediction module; The background data module includes a supplier data acquisition module, a logistics data acquisition module, a distributor data acquisition module and an inventory data acquisition module; The consumer survey module includes an online survey data acquisition module and a consumer venue questionnaire module; A chart output module is connected to the prediction platform module, and the chart output module is connected to the sales prediction module.
2. The big data-based liquor market demand forecasting system according to claim 1, characterized in that: The consumer venues in the consumer venue questionnaire module include: KTVs, hotels, and restaurants; the questionnaires in the consumer venue questionnaire module include: paper survey documents and electronic survey documents.
3. The system for predicting the demand of the liquor market based on big data according to claim 1, characterized in that: The distributor data acquisition module includes: an offline distributor data module and an online distributor data module; The data sources of the offline distributor data module include large supermarkets, shopping malls, liquor specialty stores, stores that can sell liquor, and group purchase distributors; The data sources of the online distributor data module include: multi-e-commerce platform sales data, multi-live platform sales data.
4. The liquor market demand prediction system based on big data according to claim 1, wherein: The inventory data acquisition module includes a manufacturer inventory data module and a supplier inventory data module; The inventory data acquisition module includes a monthly inventory data module, an inventory data year-on-year module, an inventory data month-on-month module and an inventory prediction module.
5. The liquor market demand prediction system based on big data according to claim 1, characterized in that: The sales prediction module uses the random forest algorithm for prediction; the random forest algorithm includes the following steps: data preparation, bootstrap sampling, constructing decision trees and integrated prediction.
6. The big data-based liquor market demand forecasting system according to claim 5, wherein: The random forest algorithm includes: Step 1: Data preparation: Collect consumer attribute data as modeling data, perform data formatting processing on the consumer attribute data, and perform normalization processing on the feature values of the consumer attribute data to unify the order of magnitude of different attribute data to the same order of magnitude. The consumer attribute data includes basic data, purchase data and behavior data. Step 2: Bootstrap sampling, screening consumer attribute data, including: Step 21: Calculate the Pearson coefficient between the features of the consumer attribute data, remove highly correlated features, and obtain a preliminary screening result. Step 22: According to the preliminary screening result, form a sample set. Denote the feature data set of the sample set as X and the number of features as m. Select to ignore the first feature M1, use the remaining m - 1 features for training and verification, calculate the prediction accuracy rate, repeat the calculation of the prediction accuracy rate for ignoring other features, obtain the influence relationship of each feature on the prediction accuracy rate of the random forest classification model, and retain the features with higher correlation with the sample label according to the prediction accuracy rate ranking. Step 3: Construct a decision tree and balance the sample set: Calculate the sample set balancing target value for the number of samples under different sample labels in the sample set. For the number of samples exceeding the sample set balancing target value, remove samples until the number of the sample set reaches the sample set balancing target value. For the number of samples less than the sample set balancing target value, add samples until the number of the sample set reaches the sample set balancing target value; Step 4: Ensemble prediction: Use the data in the sample set to train a random forest classification model, and use the random forest classification model to classify consumers based on the recent data of consumers, and predict the future consumption ability of consumers for alcoholic beverages.
7. The big data-based liquor market demand prediction system according to claim 6, characterized in that: The expression of the Pearson correlation coefficient r used in Step 21: Calculate the degree of correlation between feature x and feature y. Among them, when |r| ≥ 0.8, x and y are highly correlated; when 0.5 ≤ |r| < 0.8, x and y are moderately correlated; when 0.3 ≤ |r| < 0.5, x and y are lowly correlated; when |r| < 0.3, x and y are considered uncorrelated.
8. The big data-based liquor market demand prediction system according to claim 6, characterized in that: The training and validation performed in Step 22 includes: First, randomly divide the feature set of the remaining m - 1 features into k parts, select the first part of the samples as the validation set, and the remaining k - 1 parts of the samples as the training set to train the random forest classification model, and use the random forest classification model to predict the sample categories of the validation set and calculate the prediction accuracy rate; then, select the second part of the samples as the validation set, and the remaining k - 1 parts of the samples as the training set, and the process of training the model and calculating the prediction accuracy rate until the prediction accuracy rate of the random forest classification model is verified for each part of the samples; finally, take the average of the prediction accuracy rates calculated each time as the evaluation of the prediction accuracy rate of the model that ignores feature M1.
9. The system for predicting the demand of the liquor market based on big data according to claim 6, wherein: The formula used in Step 3: Calculate the sample set balancing target value Ntarget, where n is the number of labels and Yi is the number of samples under the i-th label; Construct a data set with M + 1 dimensions from the features and sample labels retained in Step 2, and calculate the minimum spatial distance of the data set: x1 and x2 are two samples in the data set, D(x1, x2) represents the distance between the two samples, M + 1 is the dimension of the samples, x1,j represents the j-th dimension of sample 1, and x2,j represents the j-th dimension of sample 2. Remove the data with the minimum D(x1, x2) value until the number of the sample set reaches the balancing target value Ntarget.
Citation Information
Patent Citations
Beer sales monitoring system based on big data
CN113011914A
Baijiu sales management method and system based on big data
CN116308495A
Wine consumption prediction method based on random forest
CN118735580A
Retail industry inventory management and demand prediction system
CN119227889A
Alchol commodity marketing and sales method by on-line
KR1020110119314A