Supermarket fresh commodity sales volume prediction method based on big data analysis

By combining historical sales data, weather data, holiday information and real-time promotions in supermarket fresh products, using machine learning algorithms for multi-dimensional data modeling, the problems of inaccurate prediction and lack of real-time in the existing technology are solved, and higher sales forecast accuracy and real-time performance are achieved, reducing inventory loss and improving sales management efficiency.

CN120047177APending Publication Date: 2025-05-27HANGZHOU LIUXIAOXIANG TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411860173.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-17
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

In the sales volume forecast of fresh food in supermarkets, external factors such as weather, holidays and promotional activities are not effectively considered, resulting in inaccurate predictions and lack of real-timeness.

Method used

Using a method based on big data analysis, combining historical sales data, weather data, holiday information and real-time promotion activities, multi-dimensional data modeling is carried out through machine learning algorithms to improve the accuracy and real-timeness of sales forecasts.

Benefits of technology

It effectively reduces inventory loss, improves the sales management efficiency of fresh food products, and can accurately and timely generate replenishment suggestions, reduces inventory waste, and improves supply chain management efficiency.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The invention relates to a supermarket fresh commodity sales volume prediction method based on big data analysis, and belongs to the technical field of commodity sales volume prediction, and the method comprises the following operation steps: 1, carrying out the data collection and preprocessing through sensor data, historical sales data and consumer behavior analysis; 2, summarizing the data of the three angles of the person and goods yard and processing the data; missing values and abnormal values are processed, and the abnormal values and features are selected. 3, establishing a data model; feature engineering is established, data segmentation is carried out, a prediction model is selected, and then prediction and evaluation are carried out. And 4, dynamically adjusting a pricing strategy based on the predicted demand quantity and supply condition. And 5, according to the predicted sales volume data, planning sales promotion activities, order quantities and distribution arrangements in advance. And the accuracy and the real-time performance of sales volume prediction are improved. The inventory loss is effectively reduced, and the sales management efficiency of fresh commodities is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of commodity sales volume prediction, and particularly relates to a method for predicting the sales volume of fresh supermarket products based on big data analysis. Background Art

[0002] Fresh supermarket products have the characteristics of short shelf life and large demand fluctuations. Therefore, sales forecasting is crucial for reducing inventory losses and optimizing supply chain management. Traditional forecasting methods often rely on historical sales data and ignore the influence of external factors such as weather, holidays, and promotions, resulting in inaccurate forecasts.

[0003] Currently, the methods of using historical sales data and machine learning methods in the market for sales volume prediction are greatly affected by external factors. The results of sales volume prediction are often inaccurate, and it is impossible to timely respond to changes in the replenishment of supermarket products, lacking accuracy and real-time performance. Summary of the Invention

[0004] The present invention mainly solves the deficiencies existing in the prior art, and provides a method for predicting the sales volume of fresh supermarket products based on big data analysis. By combining historical sales data, weather data, holiday information, and real-time promotion activities, and using machine learning algorithms for multi-dimensional data modeling, the accuracy and real-time performance of sales volume prediction are improved. Effectively reduce inventory losses and enhance the sales management efficiency of fresh products.

[0005] The above technical problems of the present invention are mainly solved by the following technical solutions: A method for predicting the sales volume of fresh supermarket products based on big data analysis, including the following operating steps: The first step: Collect and preprocess data from three levels: sensor data, historical sales data, and consumer behavior analysis; collect data related to people, goods, and the venue.

[0006] The second step: Summarize the data from the three perspectives of people, goods, and the venue and perform data processing; process missing values and outliers, and perform outlier and feature selection.

[0007] The third step: Establish a data model; establish a feature engineering, perform data splitting, select a prediction model, and then perform prediction and evaluation.

[0008] The fourth step: Dynamically adjust the pricing strategy based on the predicted demand and supply situation.

[0009] The fifth step: According to the predicted sales volume data, plan promotion activities, order quantities, and distribution arrangements in advance.

[0010] Preferably, collect data related to people, including customers' purchase habits, consumption frequencies, and average shopping basket values; customers' geographical locations and purchase location information; consumers' preferences and purchase history data recorded through membership cards and apps.

[0011] Preferably, collect data related to goods, including obtaining the most basic data sources through the data interfaces of the store sales system, including the daily or hourly sales quantities and sales amounts; including the specific sales situations of each fresh food category; the fluctuations in commodity prices will also affect the sales volume, and these data need to be recorded. Fresh food categories such as vegetables, fruits, meats, etc.

[0012] Preferably, collect data related to the venue, record the impacts of various promotional activities, discounts, coupons, etc. on sales volume, including data on TV advertisements, online advertisements, discount promotions, and festival promotional activities, and consumers' behaviors such as comments, likes, and shares on social media, which reflect consumers' interests in fresh products. At the same time, weather conditions and seasonal changes have a greater impact on the demand for fresh products.

[0013] Preferably, associate different data sources to ensure the accuracy of each data item and prepare for subsequent data preprocessing; the data sources include sales data, inventory data, and external environment data; update the data regularly every week through scheduled tasks to ensure that the data sources always remain up-to-date for real-time demand forecasting and packaging optimization.

[0014] Preferably, for numerical features, fill in the missing values with the mean or median of the column; median filling is more robust to outliers compared to the mean; for categorical features, fill in the missing values with the mode of the column; predict the missing values through simple regression or machine learning models; when the outliers are data entry errors, directly delete them; when the outliers are likely to be data noise or extreme values, replace them with the mean, median, or boundary values of the upper and lower quartiles.

[0015] Preferably, the numerical features are converted into a standard normal distribution with a mean of 0 and a variance of 1. The formula is: $z = \frac{x - \mu}{\sigma}$, which is applicable to the case where there are large differences in the dimensions between features; the numerical features are scaled to the interval [0, 1], which is applicable to the case where the numerical distribution ranges vary greatly. The formula is: $x_{norm} = \frac{x - \min(x)}{\max(x) - \min(x)}$; a new binary feature is created for each category. When a feature has N different categories, one-hot encoding will generate N binary features, with each category corresponding to one feature; the Pearson correlation coefficient between features is calculated to measure the linear correlation between continuous features. A value close to 1 or -1 indicates a high correlation, and this part of the features needs to be de-duplicated.

[0016] Preferably, the combination of weather and promotional activity features reflects the impact of promotional activities on sales volume under specific weather conditions; the sales volume data of each commodity in the past few days is used as a feature column; the average sales volume, maximum sales volume, and minimum sales volume in the past N days are calculated to capture short-term trends; the dataset is divided into a training set, a validation set, and a test set. The commonly used ratio is 70% for the training set, 15% for the validation set, and 15% for the test set. Ensure that the test set data does not appear in the training to avoid data leakage; for time series data, the training set and the test set are usually divided in chronological order to ensure that the test data is future data.

[0017] Preferably, in the first step: the time series model ARIMA is adopted. ARIMA is a classic statistical method for time series prediction, especially suitable for time series data without seasonal changes; it combines three components: autoregression - AR, differencing - I, and moving average - MA, and is applicable to analyzing stationary time series, that is, time series data with limited data fluctuation ranges and relatively large correlations between data and time. The time series model ARIMA is applicable to this data.

[0018] Step 1: Using the ADF test, it is obtained that the p-value is less than the significance level of -0.05, indicating that the null hypothesis is rejected and the data is stationary.

[0019] Step 2: By plotting the PACF graph to observe the significance between values in the data, the optimal parameter p value of the time series model is determined; by plotting the ACF graph to observe the significance between values in the data, the optimal parameter q value of the time series model is determined.

[0020] Step 3: The obtained optimal parameters are passed into the model for training and prediction to obtain the test result result1.

[0021] Step 2: Adopt the random forest regression algorithm. The random forest regression algorithm is a regression algorithm based on the idea of ensemble learning. It makes predictions by constructing multiple decision trees, and then obtains the final prediction result by averaging the output results of all trees for regression problems or majority voting for classification problems.

[0022] Step 1: Obtain the maximum depth or minimum number of samples of the random forest through cross-validation and grid search.

[0023] Step 2: Obtain the final prediction model based on the obtained maximum depth or minimum number of samples to predict the result and get result2.

[0024] Step 3: Train through deep learning, use a neural network model to handle regression problems, and predict continuous values. Different from traditional regression algorithms, deep learning captures complex patterns in data through multiple layers of non-linear transformations, is suitable for large-scale and complex data sets, performs excellently when the features are highly non-linear, and is also a relatively good model adopted by current big data. Regression algorithms include linear regression, random forest regression, etc.

[0025] Step 1: Standardize the data using the unit variance method so that the data conforms to the input data type of the neural network model.

[0026] Step 2: Define the structure of the neural network to train the model, and optimize the weights by comparing the loss function. Defining the structure of the neural network includes the input layer, hidden layer, and output layer.

[0027] Step 3: Use the optimized model to predict and obtain the prediction result result3.

[0028] Step 4: By comparing the results of different prediction models result1, result2, and result3, and using the root mean square error, mean absolute error, and root mean square error for error comparison, evaluate the best prediction model ARIMA as the final prediction model.

[0029] The present invention can achieve the following effects: The present invention provides a method for predicting the sales volume of fresh supermarket products based on big data analysis. Compared with the prior art, by combining historical sales data, weather data, holiday information, and real-time promotion activities, and using machine learning algorithms for multi-dimensional data modeling, the accuracy and real-time performance of sales volume prediction are improved. Effectively reduce inventory losses and enhance the sales management efficiency of fresh products.

[0030] It can automatically generate replenishment suggestions accurately and in a timely manner, effectively reduce inventory waste, improve the efficiency of supply chain management, and is applicable to the demand forecasting scenarios of supermarkets and fresh food distribution enterprises. Detailed implementation manners

[0031] The technical solutions of the invention will be further specifically described below through embodiments.

[0032] Embodiment: A method for predicting the sales volume of fresh food products in a supermarket based on big data analysis, including the following operating steps: The first step: Collect and preprocess data from three aspects: sensor data, historical sales data, and consumer behavior analysis; collect data related to people, goods, and the venue.

[0033] Collect data related to people, including customers' purchase habits, consumption frequencies, and average shopping basket values; customers' geographical locations and purchase location information; consumers' preferences and purchase history data recorded through membership cards and apps.

[0034] Collect data related to goods, including obtaining the most basic data sources through the data interface of the store sales system, including the daily or hourly sales quantity and sales amount; including the specific sales situation of each fresh food category; the fluctuation of commodity prices will also affect the sales volume, and these data need to be recorded.

[0035] Collect data related to the venue, record the impacts of various promotional activities, discounts, coupons, etc. on the sales volume, including data on TV advertisements, online advertisements, discount promotions, and festival promotional activities, and consumers' comments, likes, shares, etc. on social media, which reflect consumers' interest in fresh food products. At the same time, weather conditions and seasonal changes have a greater impact on the demand for fresh food products.

[0036] Associate different data sources to ensure the accuracy of each data item and prepare for subsequent data preprocessing; the data sources include sales data, inventory data, and external environment data; update the data regularly every week through a scheduled task to ensure that the data sources always remain up-to-date for real-time demand forecasting and packaging optimization.

[0037] The second step: Summarize the data from the three perspectives of people, goods, and the venue and perform data processing; handle missing values and outliers, and perform outlier and feature selection.

[0038] For numerical features, use the mean or median of the column to fill in the missing values; median filling is more robust to outliers than mean filling; for categorical features, use the mode of the column to fill in the missing values; predict the missing values through simple regression or machine learning models; when the outlier is a data entry error, directly delete it; when the outlier is likely to be data noise or an extreme value, replace it with the boundary value of the mean, median, or upper and lower quartiles.

[0039] Convert the numerical features into a standard normal distribution with a mean of 0 and a variance of 1. The formula is: \(z=\frac{x - \mu}{\sigma}\), which is applicable to the case where there are large differences in the dimensions between features; scale the numerical features to the interval [0, 1], which is applicable to the case where the numerical distribution ranges vary greatly. The formula is: \(x_{norm}=\frac{x - \min(x)}{\max(x)-\min(x)}\); create a new binary feature for each category. When a feature has N different categories, one-hot encoding will generate N binary features, with each category corresponding to one feature; calculate the Pearson correlation coefficient between features, which is used to measure the linear correlation between continuous features. A value close to 1 or -1 indicates a high correlation, and this part of the features needs to be de-duplicated.

[0040] Step 3: Establish a data model; establish feature engineering, perform data splitting, select a prediction model, and then conduct prediction and evaluation.

[0041] Combine the weather and promotion activity features to reflect the impact of promotion activities on sales volume under specific weather conditions; use the sales volume data of each commodity in the past few days as a feature column; calculate the average sales volume, maximum sales volume, and minimum sales volume in the past N days to capture short-term trends; divide the dataset into a training set, a validation set, and a test set. The commonly used ratio is 70% for the training set, 15% for the validation set, and 15% for the test set. Ensure that the test set data does not appear in the training to avoid data leakage; for time series data, the training set and the test set are usually divided in chronological order to ensure that the test data is future data.

[0042] S1: Adopt the time series model ARIMA. ARIMA is a classic statistical method for time series prediction, especially suitable for time series data without seasonal variations; it combines three components: autoregressive - AR, differencing - I, and moving average - MA, and is applicable to analyzing stationary time series, that is, time series data with limited fluctuation ranges and relatively large correlations between data and time. The time series model ARIMA is applicable to this data.

[0043] Step 1: Use the ADF test to obtain a p-value less than the significance level of -0.05, indicating that the null hypothesis is rejected and the data is stationary.

[0044] Step 2: Observe the significance among values in the data by plotting the PACF graph to determine the optimal parameter p value of the time series model; observe the significance among values in the data by plotting the ACF graph to determine the optimal parameter q value of the time series model.

[0045] Step 3: Pass the obtained optimal parameters into the model for training and prediction to obtain the test result result1.

[0046] S2: Adopt the random forest regression algorithm. The random forest regression algorithm is a regression algorithm based on the idea of ensemble learning. It makes predictions by constructing multiple decision trees and then obtains the final prediction result by averaging the output results of all trees for regression problems or by majority voting for classification problems.

[0047] Step 1: Obtain the maximum depth or minimum number of samples of the random forest through cross-validation and grid search.

[0048] Step 2: Obtain the final prediction model based on the obtained maximum depth or minimum number of samples to predict the result and get result2.

[0049] S3: Train through deep learning, use a neural network model to handle regression problems and predict continuous values. Different from traditional regression algorithms, deep learning captures complex patterns in the data through multiple layers of nonlinear transformations, is suitable for large-scale and complex data sets, performs excellently when the features are highly nonlinear, and is also a relatively good model adopted by current big data.

[0050] Step 1: Make the data conform to the input data type of the neural network model by using the standardization method of unit variance for the data.

[0051] Step 2: Define the structure of the neural network to train the model and optimize the weights by comparing the loss function.

[0052] Step 3: Use the optimized model to predict to obtain the prediction result result3.

[0053] S4: By comparing the results of different prediction models result1, result2, and result3, and using the root mean square error, mean absolute error, and root mean square error for error comparison, evaluate the best prediction model ARIMA as the final prediction model.

[0054] Fourth step: Dynamically adjust the pricing strategy based on the predicted demand and supply situation.

[0055] Fifth step: According to the predicted sales data, plan promotional activities, order quantities, and distribution arrangements in advance.

[0056] In summary, the method for predicting the sales volume of fresh food products in a supermarket based on big data analysis combines historical sales data, weather data, holiday information, and real-time promotion activities, and uses machine learning algorithms to perform multi-dimensional data modeling, improving the accuracy and real-time performance of sales volume prediction. It effectively reduces inventory losses and enhances the sales management efficiency of fresh food products.

[0057] Implement the collection and integration of historical data, external factors such as weather, holidays, promotions, etc., and real-time data. Use a multi-layer neural network model to model time series data and capture complex seasonal and non-linear relationships. Provide dynamic predictions and automatic replenishment suggestions to support users in adjusting and optimizing.

[0058] The above are only specific embodiments of the present invention, but the structural features of the present invention are not limited thereto. Any changes or modifications made by those skilled in the art within the scope of the present invention are covered by the patent scope of the present invention.

Claims

1. A method for predicting sales volume of fresh products in supermarkets based on big data analysis, characterized in that The steps are as follows: Step 1: Collect and pre-process data through three levels: sensor data, historical sales data, and consumer behavior analysis; collect data related to people, goods, and venues; Step 2: Summarize the data from the three perspectives of people, goods and venues and process the data; handle missing values ​​and outliers, and perform outlier and feature selection; Step 3: Establish a data model; establish feature engineering, perform data segmentation, select a prediction model, and then perform prediction and evaluation; Step 4: Dynamically adjust pricing strategies based on predicted demand and supply conditions; Step 5: Plan promotional activities, order quantities and delivery arrangements in advance based on the predicted sales data.

2. The method for predicting sales volume of fresh products in supermarkets based on big data analysis according to claim 1 is characterized by: Collect data related to people, including customer purchasing habits, consumption frequency and average shopping basket value; customer geographic location and purchase location information; consumer preferences and purchase history data recorded through membership cards and apps.

3. The method for predicting sales volume of fresh products in supermarkets based on big data analysis according to claim 2 is characterized by: Collect data related to goods, including obtaining the most basic data sources through the data interface of the store sales system, including daily or hourly sales quantity and sales revenue; including the specific sales situation of each fresh product category; commodity price fluctuations will also affect sales volume, and these data need to be recorded.

4. The method for predicting sales volume of fresh products in supermarkets based on big data analysis according to claim 3 is characterized by: Collect market-related data and record the impact of various promotional activities, discounts, coupons, etc. on sales, including data on TV advertising, online advertising, discount promotions and festival promotions. Consumers' comments, likes, sharing and other behaviors on social media reflect consumers' interest in fresh products. At the same time, weather conditions and seasonal changes have a great impact on the demand for fresh goods.

5. The method for predicting sales volume of fresh products in supermarkets based on big data analysis according to claim 4 is characterized by: Associate different data sources to ensure the accuracy of each data item and prepare for subsequent data preprocessing; data sources include sales data, inventory data, and external environment data; Update data regularly every week through scheduled tasks to ensure that the data source is always up to date for real-time demand forecasting and packaging optimization.

6. The method for predicting sales volume of fresh products in supermarkets based on big data analysis according to claim 1, characterized in that: For numerical features, use the mean or median of the column to fill missing values; median filling is more robust to outliers than the mean; for categorical features, use the mode of the column to fill missing values; predict missing values ​​through simple regression or machine learning models; when outliers are data entry errors, delete them directly; when outliers are likely to be data noise or extreme values, replace them with the mean, median, or the boundary values ​​of the upper and lower quartiles.

7. The method for predicting sales volume of fresh products in supermarkets based on big data analysis according to claim 6 is characterized by: Convert the numerical features to a standard normal distribution with a mean of 0 and a variance of 1. The formula is: z=x−μσz = \frac{x - \mu}{\sigma}z=σx−μ. It is applicable to situations where the dimensions of the features differ greatly. Scale the numerical features to the interval [0, 1]. It is applicable to situations where the numerical distribution ranges differ greatly. The formula is: xnorm=x−min (x)max(x)−min(x)x_{norm} =\frac{x - \min(x)}{\max(x) - \min(x)}xnorm=max(x)−min(x)x−min(x). Create a new binary feature for each category. When a feature has N different categories, one-hot encoding will generate N binary features, one feature for each category. Calculate the Pearson correlation coefficient between features to measure the linear correlation between continuous features. Values ​​close to 1 or -1 indicate a high correlation, and these features need to be deduplicated.

8. The method for predicting sales volume of fresh products in supermarkets based on big data analysis according to claim 1, characterized in that: The weather and promotion features are combined to reflect the impact of promotions on sales under specific weather conditions. The sales data of each product in the past few days is used as a feature column. Calculate the average sales, maximum sales, and minimum sales over the past N days to capture short-term trends; Divide the data set into training set, validation set and test set. The commonly used ratio is 70% training set, 15% validation set and 15% test set. Ensure that the test set data does not appear in training to avoid data leakage; for time series data, the training set and test set are usually divided in chronological order to ensure that the test data is future data.

9. The method for predicting sales volume of fresh products in supermarkets based on big data analysis according to claim 8, characterized in that: Step 1: Use the time series model ARIMA. ARIMA is a classic statistical method for time series forecasting, especially for time series data without seasonal changes. It combines three components: autoregression (AR), difference (I), and moving average (MA). It is suitable for analyzing stationary time series, i.e., data with a limited fluctuation range. The data has a large correlation with time. The time series model ARIMA is suitable for this data. Step 1: Using the ADF test, the p-value is less than the significant level -0.05, indicating that the null hypothesis is rejected and the data is stable; Step 2: Observe the significance between the values ​​in the data by drawing the PACF graph, so as to determine the optimal parameter p value of the time series model; observe the significance between the values ​​in the data by drawing the ACF graph, so as to determine the optimal parameter q value of the time series model; Step 3: The obtained optimal parameters are passed into the model for training and prediction to obtain the test result 1; Step 2: Use the random forest regression algorithm. The random forest regression algorithm is a regression algorithm based on the idea of ​​ensemble learning. It makes predictions by constructing multiple decision trees, and then obtains the final prediction result by averaging the output results of all trees for regression problems or by majority voting for classification problems. Step 1: Obtain the maximum depth or minimum number of samples of random forest through cross validation and grid search; Step 2: Obtain the final prediction model by obtaining the maximum depth or minimum number of samples to predict the result and obtain result2; Step 3: Train through deep learning and use neural network models to handle regression problems and predict continuous values. Unlike traditional regression algorithms, deep learning uses multiple layers of nonlinear transformations to capture complex patterns in data. It is suitable for large-scale and complex data sets and performs well when there is a high degree of nonlinearity between features. It is also a good model currently used in big data. Step 1: Standardize the data to unit variance to make it conform to the input data type of the neural network model; Step 2: Define the structure of the neural network to train the model and optimize the weights by comparing the loss function; Step 3: Use the optimized model to predict and obtain the prediction result result3; Step 4: By comparing the results of result1, result2, and result3 of different prediction models, the error method uses the comparison of root mean square error, mean absolute error, and root mean square error to evaluate the best prediction model ARIMA as the final prediction model.

Citation Information

Cited By

  • Retail store goods reporting data processing method and system based on sales prediction model

    CN120807030A

  • Supply and transportation management method of fresh fish meat

    CN121032368A