Consumer shopping behavior prediction method based on machine learning
By collecting and processing massive transaction data and using decision tree machine learning models to predict consumer shopping behavior, it solves the problem of difficulty in accurately predicting consumer shopping behavior in the existing technology, and improves prediction accuracy and efficiency of marketing resources.
Patent Information
- Application Number
- CN202510249898.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-03
AI Technical Summary
The existing technology is difficult to effectively use massive transaction data to mine consumer shopping models, making it difficult for companies to accurately predict consumer shopping behavior.
By collecting data on e-commerce platform transaction records, offline shopping receipts, browsing history and social media comments, feature extraction and standardization are performed, and decision tree machine learning models are used to train and predict consumers' future shopping behavior.
It improves the accuracy of forecasting shopping preferences for different consumer groups, allowing companies to push product advertisements and promotional activities that meet their interests for different groups and reasonably allocate marketing resources.
Smart Images

Figure CN120088037A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of machine learning and data analysis, and particularly relates to a method for predicting consumer shopping behavior based on machine learning. Background Art
[0002] In today's business environment, understanding consumer shopping behavior is crucial for enterprises to formulate precise marketing strategies, optimize inventory management, and improve customer satisfaction. Traditional market research methods are often time-consuming and laborious, and it is difficult to capture the dynamically changing shopping preferences of consumers. With the booming development of e-commerce, a vast amount of transaction data has been recorded. How to utilize this data to mine potential shopping patterns of consumers has become an urgent problem to be solved.
[0003] Machine learning is committed to studying how to improve the performance of a system itself through a computer using experience (usually data). Simply put, it is to let the computer learn rules and patterns from a large amount of data, and then use these rules for prediction or decision-making. With its powerful automatic learning and pattern recognition capabilities, machine learning technology provides an effective way to solve the problem of predicting consumer shopping behavior. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to overcome the above-mentioned disadvantages of the prior art and provide a method for predicting consumer shopping behavior based on machine learning.
[0005] The technical solution adopted to solve the above technical problem is as follows: It includes the following specific steps:
[0006] Step 1: Data collection: Collect consumer data from e-commerce platform transaction records, offline physical store shopping receipts, consumer browsing histories, and social media comments, and clean the collected data.
[0007] Step 2: Feature engineering: Extract features from the cleaned data and perform standardization processing.
[0008] Step 3: Model training: Select a decision tree machine learning model and use the divided training set to train the model to optimize the model performance.
[0009] Step 4: Behavior prediction: Input the processed new consumer data into the trained model and output the prediction result of the consumer's future shopping behavior.
[0010] Step 5: Result feedback: Feed back the prediction result to the relevant departments, and optimize the model based on the comparative analysis of the actual shopping behavior data and the prediction result.
[0011] Through the above technical solutions, enterprises can push product advertisements and promotional activities that match the interests of different groups, reasonably allocate marketing resources, and invest more resources in the channels preferred by consumers.
[0012] Further, in the data collection step, establish a data interface with the e-commerce platform, obtain transaction data according to the data opening protocol, deploy scanning devices in offline stores to enter shopping receipts, and use web crawlers to capture social media comment information, and comply with the platform rules when capturing.
[0013] Through the above technical solutions, the accuracy and effectiveness of data collection are greatly improved, thereby improving the accuracy of model training.
[0014] Further, in the feature engineering step, use the Scikit-learn library of Python for feature extraction and standardization operations.
[0015] Further, in the model training step, determine the optimal model parameters through the cross-validation method. In the behavior prediction step, the model runs on a high-performance server cluster to ensure the timeliness of prediction. In the result feedback step, regularly collect actual shopping behavior data every month, analyze the error reasons using statistical methods, and retrain and adjust the model according to the analysis results.
[0016] Further, the feature engineering step uses the Z-score standardization formula to preprocess the data. Let the original data feature value be x, the standardized value be z, the mean of this feature be μ, and the standard deviation be σ. Then the formula is:
[0017]
[0018] Where:
[0019]
[0020] x i is each data point in the dataset, and n is the number of data points;
[0021]
[0022] After Z-score standardization, the mean of the data becomes 0, the standard deviation becomes 1, and the new data distribution is more conducive to the processing of subsequent machine learning models.
[0023] Further, for the machine learning model, use the accuracy formula for verification, as follows:
[0024]
[0025] Among them, A is the accuracy result. TP (True Positive) represents the true positive example, that is, the sample quantity that is actually a positive class and is predicted as a positive class. TN (True Negative) represents the true negative example, that is, the sample quantity that is actually a negative class and is predicted as a negative class. FP (False Positive) represents the false positive example, that is, the sample quantity that is actually a negative class but is predicted as a positive class. FN (False Negative) represents the false negative example, that is, the sample quantity that is actually a positive class but is predicted as a negative class.
[0026] Further, the mean square error formula is used to adjust the amount of money used by consumers for shopping, and the specific expression is as follows:
[0027]
[0028] Where n is the number of samples, y i is the true value of the i-th sample, is the predicted value of the i-th sample.
[0029] Through the above technical solution, the parameters learned by the model can make the predicted value as close as possible to the true value.
[0030] The beneficial effects of the present invention are as follows: By performing feature extraction and standardization processing on the cleaned data, selecting a decision tree machine learning model, and training the model using the divided training set, the performance of the model is optimized, greatly improving the prediction accuracy of the shopping preferences of different consumer groups. It enables enterprises to push product advertisements and promotional activities that meet the interests of different groups, reasonably allocate marketing resources, and invest more resources in the channels preferred by consumers. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 is the flow chart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0032] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0033] As Figure 1 shown, the consumer shopping behavior prediction method based on machine learning in this embodiment includes the following specific steps:
[0034] Step 1: Data collection: Collect consumer data from e-commerce platform transaction records, offline store shopping receipts, consumer browsing histories, and social media comments, and clean the collected data;
[0035] Step 2: Feature Engineering: Extract features from the cleaned data and perform standardization processing;
[0036] Step 3: Model Training: Select a decision tree machine learning model and use the divided training set to train the model to optimize its performance;
[0037] Step 4: Behavior Prediction: Input the processed new consumer data into the trained model and output the prediction results of consumers' future shopping behaviors;
[0038] Step 5: Result Feedback: Feed back the prediction results to the relevant departments and optimize the model based on the comparative analysis of the actual shopping behavior data and the prediction results.
[0039] In the data collection step, establish a data interface with the e-commerce platform, obtain transaction data according to the data opening protocol, deploy scanning devices in offline stores to enter shopping receipts, and use web crawlers to capture social media comment information, and comply with the platform rules during the capture.
[0040] In the feature engineering step, use the Scikit-learn library of Python for feature extraction and standardization operations.
[0041] In the model training step, determine the optimal model parameters through the cross-validation method. In the behavior prediction step, the model runs on a high-performance server cluster to ensure the timeliness of prediction. In the result feedback step, collect actual shopping behavior data regularly every month, use statistical methods to analyze the error reasons, and retrain and adjust the model according to the analysis results.
[0042] In the feature engineering step, the Z-score standardization formula is used to preprocess the data. Let the original data feature value be x, the standardized value be z, the mean of this feature be μ, and the standard deviation be σ. Then the formula is:
[0043]
[0044] Where:
[0045]
[0046] x i is each data point in the dataset, and n is the number of data points;
[0047]
[0048] After Z-score standardization, the mean of the data becomes 0, the standard deviation becomes 1, and the new data distribution is more conducive to the processing of subsequent machine learning models.
[0049] The accuracy formula is used to verify the machine learning model, which is as follows:
[0050]
[0051] Among them, A is the accuracy result, TP (True Positive) represents the true positive example, that is, the sample quantity that is actually a positive class and is predicted as a positive class, TN (True Negative) represents the true negative example, that is, the sample quantity that is actually a negative class and is predicted as a negative class, FP (False Positive) represents the false positive example, that is, the sample quantity that is actually a negative class but is predicted as a positive class, and FN (False Negative) represents the false negative example, that is, the sample quantity that is actually a positive class but is predicted as a negative class.
[0052] For the amount of money consumers spend on shopping, the mean squared error formula is used for adjustment, so that the parameters learned by the model can make the predicted value as close as possible to the true value. The specific expression is as follows:
[0053]
[0054] Among them, n is the number of samples, y i is the true value of the i-th sample, is the predicted value of the i-th sample.
[0055] In the data collection stage, data interfaces are established with major e-commerce platforms, and transaction data is obtained regularly according to their data opening protocols. At the same time, data collection devices are deployed in offline stores to scan and enter consumer shopping receipts. For social media data, web crawler technology is used to capture brand- and product-related comment information on the premise of complying with platform rules.
[0056] In model training, a neural network model is selected, and the network structure is built using the PyTorch framework. By experimenting with different combinations of hyperparameters multiple times, the cross-validation method is used to determine the optimal model parameters.
[0057] When predicting behaviors, the consumer data collected in real time is quickly sent into the trained model through the previous processing process. The model runs on a high-performance server cluster to ensure the timeliness of prediction.
[0058] The above is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention.
Claims
1. A method for predicting consumer shopping behavior based on machine learning, characterized in that: The specific steps include: Step 1: Data collection: Collect consumer data from e-commerce platform transaction records, offline store shopping receipts, consumer browsing history, and social media comments, and clean the collected data; Step 2: Feature Engineering: Perform feature extraction and standardization on the cleaned data; Step 3: Model training: Select a decision tree machine learning model, use the divided training set to train the model, and optimize the model performance; Step 4: Behavior prediction: The new consumer data is processed and input into the trained model to output the predicted results of the consumer’s future shopping behavior; Step 5: Result feedback: Feedback the prediction results to relevant departments, and optimize the model based on comparative analysis of actual shopping behavior data and prediction results.
2. The method for predicting consumer shopping behavior based on machine learning according to claim 1, characterized in that: In the data collection step, a data interface is established with the e-commerce platform, transaction data is obtained in accordance with the data openness agreement, scanning equipment is deployed in offline stores to enter shopping receipts, and social media comment information is captured using web crawlers, and platform rules are followed when crawling.
3. The method for predicting consumer shopping behavior based on machine learning according to claim 2, characterized in that: In the feature engineering step, Python's Scikit-learn library is used to perform feature extraction and standardization operations.
4. The method for predicting consumer shopping behavior based on machine learning according to claim 3, characterized in that: In the model training step, the optimal model parameters are determined by a cross-validation method. In the behavior prediction step, the model runs on a high-performance server cluster to ensure the timeliness of the prediction. In the result feedback step, actual shopping behavior data is collected regularly every month, and the causes of errors are analyzed using statistical methods. The model is retrained and adjusted based on the analysis results.
5. The method for predicting consumer shopping behavior based on machine learning according to claim 4, characterized in that: The feature engineering step uses the Z-score standardization formula to preprocess the data. Assume that the original data feature value is x, the standardized value is z, the mean value of the feature is μ, and the standard deviation is σ, then the formula is: in: x i is each data point in the data set, and n is the number of data points; After Z-score standardization, the mean of the data becomes 0 and the standard deviation becomes 1. The new data distribution is more conducive to subsequent machine learning model processing.
6. The method for predicting consumer shopping behavior based on machine learning according to claim 5, characterized in that: The accuracy formula for machine learning models is used for verification, as follows: Where A is the accuracy result, TP (True Positive) represents true positive examples, that is, the number of samples that are actually positive and predicted to be positive, TN (True Negative) represents true negative examples, that is, the number of samples that are actually negative and predicted to be negative, FP (False Positive) represents false positive examples, that is, the number of samples that are actually negative but predicted to be positive, and FN (False Negative) represents false negative examples, that is, the number of samples that are actually positive but predicted to be negative.
7. The method for predicting consumer shopping behavior based on machine learning according to claim 6, characterized in that: The mean square error formula is used to adjust the amount of money used for consumer shopping so that the parameters learned by the model can make the predicted value as close to the true value as possible. The specific expression is as follows: Where n is the number of samples, y i is the true value of the i-th sample, is the predicted value of the ith sample.