Cigarette aggregation payment amount prediction method and system based on linear regression

By constructing a cigarette sales data model based on linear regression, the existing prediction methods are solved, and the prediction effect with higher accuracy and adaptability is achieved, providing enterprises with scientific sales strategies and resource allocation basis.

CN120013578APending Publication Date: 2025-05-16CHENGDU BRANCH OF SICHUAN TOBACCO CO
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411910268.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing cigarette sales forecasting methods are mainly based on empirical analysis, and have failed to make full use of digital payment data, with low prediction accuracy, making it difficult to support the company's refined management needs.

Method used

Using a linear regression-based method, a linear regression model is constructed by collecting and preprocessing cigarette sales data, and the model is trained using historical data to predict the aggregate payment amount in the target time period.

Benefits of technology

It significantly improves prediction accuracy, enhances the adaptability of the model, and can support the company's sales strategy and resource allocation decisions more scientifically.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013578A_ABST
    Figure CN120013578A_ABST
Patent Text Reader

Abstract

The invention discloses a cigarette aggregate payment amount prediction method based on linear regression, and the method comprises the steps: collecting historical cigarette sales data of a target store, carrying out the data preprocessing, and generating a prediction data set; taking the historical aggregate payment amount as a dependent variable, taking other historical cigarette sales data as an independent variable, constructing a linear regression model, and completing training by using a prediction data set; collecting sales data of a target store in a target time period, preprocessing the sales data, and inputting the preprocessed sales data into the trained linear regression model to obtain an aggregate payment amount prediction value; through a scientific data preprocessing method, including coding, dimension reduction processing and standardization processing of fixed-class variables, the redundancy and dimension difference between the variables are remarkably reduced, and the quality of model input data is ensured. Meanwhile, by constructing the linear regression model, the inherent rule of the cigarette sales data is accurately captured, and the prediction precision is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing for aggregate payment, and in particular to a method and system for predicting the aggregate payment amount of cigarettes based on linear regression. Background Art

[0002] With the continuous development of digital technology and smart retail models, the cigarette industry has gradually entered a new era of digital operations. As an important part of tobacco sales, the promotion and application of digital stores have significantly improved industry efficiency. In digital stores, aggregated payment has become the main payment method for consumers. Its transaction data not only reflects sales performance, but also plays an important guiding role in the company's inventory management, marketing strategy and operational decisions.

[0003] However, the sales forecasting methods commonly used in the market are mostly based on empirical analysis, which fails to fully utilize the characteristics of digital payment data, has low forecasting accuracy, and is difficult to effectively support the refined management needs of enterprises. In particular, there are significant challenges in the following aspects:

[0004] 1. Data diversity and complexity: Cigarette sales involve multi-dimensional business data, including sales amount, member consumption behavior, store attributes, etc. The correlation and redundant information of these data bring difficulties to traditional prediction models.

[0005] 2. The problem of inconsistent dimensions: Data of different dimensions (such as amount, quantity, and rating) have different dimensions. If they are not properly standardized, it will easily lead to poor model training results.

[0006] 3. Insufficient adaptability of existing models: Existing methods often lack in-depth mining and processing of multi-dimensional data features and cannot meet the needs of rapid updates in a dynamic market environment.

[0007] Therefore, there is an urgent need for a cigarette aggregate payment amount prediction model based on scientific methods, which can fully tap the potential value of multidimensional data, improve prediction accuracy and model adaptability, and provide a scientific basis for the company's sales strategy and resource allocation. Summary of the invention

[0008] In order to solve the technical problems existing in the above-mentioned prior art, the present invention aims to provide a method and system for effectively utilizing the multi-dimensional historical sales data of digital cigarette stores, constructing a high-precision prediction model, and accurately predicting the aggregated payment amount in a target time period, thereby overcoming the problems of insufficient accuracy, insufficient data processing and poor model adaptability existing in the existing prediction methods.

[0009] In order to achieve the above-mentioned invention object, the technical solution provided by the present invention includes:

[0010] The method for predicting the aggregate payment amount of cigarettes based on linear regression includes the following steps:

[0011] S1. Collect historical cigarette sales data of target stores, perform data preprocessing, and generate a prediction data set;

[0012] S2. Use the historical aggregate payment amount as the dependent variable and other historical cigarette sales data as the independent variable to build a linear regression model and complete the training using the prediction data set;

[0013] S3. Collect sales data of target stores in the target time period, input the pre-processed data into the trained linear regression model, and obtain the predicted value of the aggregated payment amount;

[0014] The pretreatment method comprises:

[0015] S11. Convert categorical variables into quantitative variables;

[0016] S12, calculating the correlation between the variables, and performing dimensionality reduction processing on the groups of variables whose correlation is higher than the preset value;

[0017] S13. Standardize all variables.

[0018] Preferably, the cigarette sales data includes: aggregate payment amount, store star rating, total cigarette scan amount, number of members, number of consumer members, specification coverage, number of scan days, number of scan customers, business format, stall, and market type.

[0019] Preferably, the method of constructing a linear regression model and completing training using a prediction data set includes:

[0020] Construct a linear regression model: Y = β0 + β1X1 + β2X2 + … + β n X n +∈; where Y is the predicted value of the aggregate payment amount, β0 is the intercept, and β n is the regression coefficient of the independent variable, X n is the independent variable, n is the number of independent variables, ∈ is the error term;

[0021] The linear regression model is fitted according to the prediction data set, and the fitting effect is tested according to the historical aggregate payment amount to complete the training of the linear regression model.

[0022] Preferably, the method of converting categorical variables into quantitative variables includes: using sequence coding to digitize the store star rating data and gear data.

[0023] Preferably, the method for converting categorical variables into quantitative variables includes: using one-hot encoding to perform dummy variable processing on business format data and market type data.

[0024] Preferably, the method for standardizing all variables includes:

[0025] Z = (X-Mean) / Std; where Z is the standardized variable; X is the variable to be processed; Mean is the mean of the variable to be processed; Std is the standard deviation of the variable to be processed.

[0026] Preferably, the dimensionality reduction processing method includes: screening variables with higher representativeness or predictive ability in each group of variables whose correlation is higher than a preset value, and deleting redundant variables.

[0027] Preferably, the linear regression model includes: Y=0.334+0.982*total cigarette scanning amount-0.118*store star+0.075*number of consumer members-0.033*specification coverage-0.082*scanning days-0.017*number of scanning customers-0.024*gear+0.04*market type0.464*business format_traditional convenience store+0.026*business format_modern single convenience store-0.293*market type_city network.

[0028] The present invention also provides a system, wherein the system is used to implement the above method.

[0029] Beneficial Effects

[0030] The present invention uses scientific data preprocessing methods, including coding, dimensionality reduction and standardization of categorical variables, to significantly reduce the redundancy and dimensional differences between variables, thereby ensuring the quality of model input data. At the same time, by constructing a linear regression model, the inherent laws of cigarette sales data can be accurately captured, and the prediction accuracy can be greatly improved. The prediction results provide a scientific basis for enterprises to formulate sales strategies, optimize inventory management and carry out precision marketing. Enterprises can flexibly adjust resource allocation according to the prediction results, improve operational efficiency, and reduce cost waste caused by excessive or insufficient stocking. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 A schematic diagram of a flow chart of a method for predicting the aggregated payment amount of cigarettes based on linear regression in a preferred embodiment of the present invention;

[0032] Figure 2 A schematic diagram of historical cigarette sales data of a target store in a preferred embodiment of the present invention;

[0033] Figure 3 A schematic diagram of the result of converting a categorical variable into a quantitative variable in a preferred embodiment provided by the present invention;

[0034] Figure 4 A schematic diagram of the result of performing dimensionality reduction processing on each group of variables with correlations higher than a preset value in a preferred embodiment provided by the present invention;

[0035] Figure 5 A schematic diagram of the results of standardizing variables in a preferred embodiment provided by the present invention;

[0036] Figure 6 A fitting curve diagram of a preferred embodiment of the present invention for training and optimizing a linear regression model based on historical cigarette sales data of a large number of stores. DETAILED DESCRIPTION

[0037] In order to make the purpose, technical solution and advantages of the present invention clearer, the present invention is further described below in conjunction with the accompanying drawings. In the description of the present invention, it should be understood that the terms "upper", "lower", "front", "back", "left", "right", "top", "bottom", "inside", "outside" and the like indicate directions or positional relationships based on the directions or positional relationships shown in the accompanying drawings, which are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operated in a specific direction, and therefore cannot be understood as limiting the present invention.

[0038] Embodiment 1

[0039] like Figure 1 As shown, the present invention discloses a method for predicting the aggregate payment amount of cigarettes based on linear regression, comprising the steps of:

[0040] S1. Collect historical cigarette sales data of target stores, perform data preprocessing, and generate a prediction data set.

[0041] The target store refers to a specific digital cigarette store, whose sales data will be collected and analyzed to build a prediction model and generate prediction results of the aggregated payment amount.

[0042] The historical cigarette sales data refers to all business data related to cigarette sales of the target store in the past period of time, which constitutes the basis for model construction and training. Figure 2 As shown, the following key historical cigarette sales data examples are given:

[0043] 1. Core sales indicators

[0044] Aggregate payment amount: the amount of cigarette sales completed by the target store through aggregate payment methods during the historical period.

[0045] Total cigarette scan amount: the total amount of all cigarette scan payments in the target store, used to reflect the sales scale.

[0046] 2. Customer behavior data

[0047] Number of members: The total number of registered members of the target store, reflecting the potential of customer resources.

[0048] Number of spending members: The number of members who have made spending decisions within a specific historical period, indicating the size of active customers.

[0049] Number of customers who scanned the QR code: the number of unique customers who scanned the QR code to pay within a specific time period.

[0050] 3. Product and service related data

[0051] Specification coverage rate: the proportion of cigarette specifications sold in stores to the total number of specifications available for sale, reflecting the product diversity and market appeal.

[0052] Scan code payment days: The number of days in which scan code payment occurs in the store within a specific time period, which represents the activity of store operations.

[0053] 4. Store attributes

[0054] Store star rating: reflects the service quality and customer evaluation of the target store.

[0055] Position: The positioning level of the target store in the market (such as low-end market, mid-to-high-end market, etc.).

[0056] Business format: the operating model of the target store (such as convenience store, fresh food store, etc.).

[0057] Market type: the market environment where the store is located (such as city, town, etc.).

[0058] 5. Data time range

[0059] The historical period covered by the data (e.g. January 2023).

[0060] The preprocessing refers to a series of operations to clean, convert and optimize the raw data, with the purpose of improving data quality, eliminating noise and redundancy, so that it can be used more efficiently to build a linear regression model. Preprocessing is a key step in data analysis and modeling. There are many ways to implement it. In some preferred embodiments, the following preferred preprocessing methods are given, including:

[0061] S11. Convert the categorical variables into quantitative variables. Specifically, convert the non-numeric categorical data into numerical data that the model can accept. In other preferred embodiments, sequence coding is used to digitize the store star data and gear data, and the variables with sequential relationships (such as "store star") are mapped to numerical values ​​in sequence, such as "no star" is 1 and "three stars" is 4. One-hot coding is used to dummy-code the business data and market type data, such as Figure 3As shown in the figure, the unordered categorical variables (such as "business format") are converted into dummy variables, and each category is mapped to an independent numerical column. Specifically, the column name is the original value, and the value in the row is 1 (present) or 0 (not present) to indicate whether the data example belongs to the category.

[0062] S12, calculate the correlation between the variables, and perform dimensionality reduction processing on the groups of variables whose correlation is higher than the preset value. By calculating the correlation coefficient between the variables, redundant variables are found and removed, the number of input features is reduced, and the model complexity and overfitting risk are reduced. In some preferred embodiments, it is considered to complete the dimensionality reduction processing by screening the variables with higher representativeness or predictive ability in the groups of variables whose correlation is higher than the preset value, and deleting the redundant variables, such as Figure 4 As shown in the figure, through the correlation analysis of "number of members", "number of consumer members" and "number of days of scanning code", it is found that the correlation coefficient between "number of members" and "number of consumer members" is 0.6901, indicating that there is a high positive correlation between these two features and there is a lot of redundant information. Consider selecting "number of consumer members" as the feature variable between "number of members" and "number of consumer members".

[0063] S13, standardize all variables. Convert data of different dimensions to the same scale to eliminate model deviation caused by dimension differences. In other preferred embodiments, consider adjusting the distribution of variable values ​​to a standard normal distribution with a mean of 0 and a standard deviation of 1 to facilitate subsequent model processing, and the following standardization method is preferred:

[0064] Z = (X-Mean) / Std; where Z is the standardized variable; X is the variable to be processed; Mean is the mean of the variable to be processed; Std is the standard deviation of the variable to be processed. Figure 5 As shown in the figure, the standardized variable values ​​of characteristic variables such as "aggregate payment amount of cigarettes", "total scanned amount of cigarettes", "number of consumer members", "number of scanned customers", "gear position" and "store star rating" fluctuate around 0. A value greater than 0 indicates that it is higher than the average level, and a value less than 0 indicates that it is lower than the average level.

[0065] Those skilled in the art should know that preprocessing also includes processing null values ​​or abnormal values ​​in the data, and the present invention does not make further limitations.

[0066] S2. Use the historical aggregate payment amount as the dependent variable and other historical cigarette sales data as the independent variable, build a linear regression model and use the prediction data set to complete the training.

[0067] The linear regression model is a basic regression analysis method used to study the linear relationship between a dependent variable (target variable) and one or more independent variables (characteristic variables). It is widely used in machine learning and statistics to predict the relationship between continuous numerical variables or explanatory variables. In the present invention, there is a relatively obvious linear relationship between the aggregate payment amount of cigarettes (target variable) and characteristic variables (independent variables) such as store star rating, total cigarette code scanning amount, and number of consumer members. For example: the sales amount is usually positively correlated with the code scanning amount and the number of members. The impact of variables such as market type or gear position on the payment amount can be described by a linear relationship. The linear regression model can capture these linear correlations intuitively and effectively.

[0068] Specifically, in some preferred embodiments, the method of constructing a linear regression model and completing training using a prediction data set includes:

[0069] Construct a linear regression model: Y = β0 + β1X1 + β2X2 + … + β n X n +∈; where Y is the predicted value of the aggregate payment amount, which is the dependent variable; β0 is the intercept, which indicates the predicted value of the target variable when all independent variables are zero. n is the independent variable regression coefficient, which indicates the influence of each independent variable on the target variable. n is the independent variable, n is the number of independent variables, and ∈ is the error term, which is used to reflect the part that is not explained or captured in the model, reflecting the randomness or unmodeled factors, so as to ensure the integrity of the model and the rationality of the prediction.

[0070] The linear regression model is fitted according to the predicted data set, and the fitting effect is tested according to the historical aggregate payment amount to complete the training of the linear regression model. The process of training the linear regression model is to determine the regression coefficient β n The value of , so that the model's prediction error for the target variable is minimized. In some preferred embodiments, the least squares method (OLS) can be used for optimization, the goal is to minimize the sum of squared errors, and the mean square error (MSE), determination coefficient (R 2 ) and other indicators to evaluate the performance and accuracy of the model.

[0071] In some preferred embodiments, by collecting historical cigarette sales data of a large number of stores, such as Figure 6 As shown in the figure, after training and optimizing the linear regression model, the following better model results can be obtained:

[0072] Y = 0.334 + 0.982 * total amount of cigarettes scanned - 0.118 * store star + 0.075 * number of consumer members - 0.033 * specification coverage - 0.082 * number of days to scan - 0.017 * number of customers to scan - 0.024 * stall + 0.04 * market type 0.464 * format_traditional convenience store + 0.026 * format_modern single convenience store - 0.293 * market type_city network. Among them, it was found in the training that the regression coefficients of the two parameters of format_hotel_dummy variable and format_restaurant_dummy variable were zero, indicating that they had no effect on the aggregate payment amount, so it was considered to remove these two independent variables.

[0073] S3. Collect sales data of target stores in the target time period, input the pre-processed data into the trained linear regression model, and obtain the aggregate payment amount prediction value. The aggregate payment amount prediction value can be used to support enterprise decision-making, such as adjusting inventory allocation in the target time period, formulating sales strategies, or evaluating store performance.

[0074] The above shows and describes the basic principles, main features and advantages of the present invention. It should be understood by those skilled in the art that the present invention is not limited to the above embodiments, and the above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, which fall within the scope of the present invention to be protected. The scope of protection of the present invention is defined by the attached claims and their equivalents.

Claims

1. A method for predicting the aggregate payment amount of cigarettes based on linear regression, characterized in that: Includes steps: S1. Collect historical cigarette sales data of target stores, perform data preprocessing, and generate a prediction data set; S2. Use the historical aggregate payment amount as the dependent variable and other historical cigarette sales data as the independent variable to build a linear regression model and complete the training using the prediction data set; S3. Collect sales data of target stores in the target time period, input the pre-processed data into the trained linear regression model, and obtain the aggregate payment amount prediction value; The pretreatment method comprises: S11. Convert categorical variables into quantitative variables; S12, calculating the correlation between the variables, and performing dimensionality reduction processing on the groups of variables whose correlation is higher than the preset value; S13. Standardize all variables.

2. The method for predicting the aggregated payment amount of cigarettes based on linear regression according to claim 1, characterized in that: The cigarette sales data include: aggregate payment amount, store star rating, total cigarette scan amount, number of members, number of consumer members, specification coverage, number of scan days, number of scan customers, business format, stall, and market type.

3. The method for predicting the aggregated payment amount of cigarettes based on linear regression according to claim 1, characterized in that: The method of constructing a linear regression model and completing training using a prediction data set includes: Construct a linear regression model: Y = β0 + β1X1 + β2X2 + … + β n X n +∈; where Y is the predicted value of the aggregate payment amount, β0 is the intercept, and β n is the regression coefficient of the independent variable, X n is the independent variable, n is the number of independent variables, ∈ is the error term; The linear regression model is fitted according to the prediction data set, and the fitting effect is tested according to the historical aggregate payment amount to complete the training of the linear regression model.

4. The method for predicting the aggregated payment amount of cigarettes based on linear regression according to claim 2, characterized in that: The method for converting categorical variables into quantitative variables includes: using sequence coding to perform numerical processing on store star data and gear data.

5. The method for predicting the aggregated payment amount of cigarettes based on linear regression according to claim 2, characterized in that: The method for converting categorical variables into quantitative variables includes: using one-hot encoding to perform dummy variable processing on business format data and market type data.

6. The method for predicting the aggregated payment amount of cigarettes based on linear regression according to claim 2, characterized in that: The method for standardizing all variables includes: Z = (X-Mean) / Std; where Z is the standardized variable; X is the variable to be processed; Mean is the mean of the variable to be processed; Std is the standard deviation of the variable to be processed.

7. The method for predicting the aggregated payment amount of cigarettes based on linear regression according to claim 1, characterized in that: The dimension reduction processing method includes: screening variables with higher representativeness or predictive ability in each group of variables with correlation higher than a preset value, and deleting redundant variables.

8. The method for predicting the aggregated payment amount of cigarettes based on linear regression according to claim 2, characterized in that: The linear regression model includes: Y=0.334+0.982*total cigarette scanning amount-0.118*store star+0.075*number of consumer members-0.033*specification coverage-0.082*scanning days-0.017*number of scanning customers-0.024*gear+0.04*market type0.464*business format_traditional convenience store+0.026*business format_modern single convenience store-0.293*market type_city network.

9. A system, characterized in that: The system is used to implement the method according to any one of claims 1 to 8.