A method for analyzing cigarette consumers
By preprocessing and clustering cigarette consumption data, building a cigarette attribute correlation map and optimizing the association rules, the problems of cigarette consumer group division and mining attribute preference correlation patterns are solved, and accurate portraits and personalized marketing support for cigarette consumers are achieved.
Patent Information
- Application Number
- CN202411552845.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-01
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2044-11-01
AI Technical Summary
In the analysis of cigarette consumption data, it is difficult to accurately divide consumer groups and explore their attribute preference correlation patterns, especially for high nicotine and low tar preference groups.
By obtaining massive cigarette consumption data, preprocessing and clustering, building a cigarette attribute correlation map, learning low-dimensional vector representation of attributes, discovering the cigarette attribute preference correlation mode of different subdivided groups, and optimizing and combining the association rules to form a subdivided group portrait model.
It realizes an accurate portrait of cigarette consumers, provides effective support for cigarette product innovation and personalized marketing, and can predict different groups' preferences for new products and formulate corresponding product formulas and marketing strategies.
Smart Images

Figure CN119494671B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information technology, and in particular, to a method for analyzing cigarette consumers. Background Art
[0002] In the analysis of cigarette consumption data, there are challenges in accurately dividing consumer groups and mining the associated patterns of their attribute preferences. First, the processing of massive cigarette consumption data requires preprocessing and cleaning using a distributed big data framework to construct a structured data set containing attribute preference information such as tobacco leaf moisture content, nicotine content in smoke, tar content in smoke, etc. Second, based on these attributes and factors such as price sensitivity and taste preference, appropriate clustering algorithms need to be designed to divide consumer segments. However, the correlations between different attributes are complex, and traditional clustering methods may have difficulty capturing these subtle relationships. Therefore, it is considered to construct a cigarette attribute association graph and use graph embedding algorithms to learn the low-dimensional representations between attributes. Although this method can better express the attribute relationships, how to efficiently mine the attribute preference association rules for different segments, especially for the groups with high nicotine content and low tar content preferences, remains an unsolved problem. In addition, how to optimize and combine the mined association rules to balance the relationships between different attributes and take into account the preference differences within the groups, so as to construct an accurate consumer portrait model, is also a technical difficulty that needs in-depth study. Summary of the Invention
[0003] The present invention provides a method for analyzing cigarette consumers, mainly including:
[0004] Obtain massive cigarette consumption data, preprocess the data, and through data cleaning and data integration operations, obtain a structured consumption data set. The consumption data set includes cigarette price sensitivity, cigarette taste preference types, cigarette appearance preference styles, cigarette brand loyalty, and the frequency of purchasing cigarettes. Among them, the cigarette taste preference types include tobacco leaf moisture content preference values, nicotine content in smoke preference ranges, and tar content in smoke preference interval attributes;
[0005] According to the cigarette attributes in the consumption data set, cluster and group consumers, and divide consumers with similar attribute preferences into the same segment. The cigarette attributes include cigarette price sensitivity, cigarette taste preference types, cigarette appearance preference styles, cigarette brand loyalty, and the frequency of purchasing cigarettes;
[0006] For each consumer segment obtained by clustering, construct a cigarette attribute association graph for the segment, learn the low-dimensional vector representations of the cigarette attributes within each segment in the graph, and measure the correlations between different attributes through the distances between the vectors;
[0007] Based on the cigarette attribute association maps of each sub-group, discover the association patterns in cigarette attribute preferences among different sub-groups, obtain the association rules for the leaf moisture content preference, cigarette price sensitivity, and appearance preference style of the high nicotine content preference group, as well as the association rules for the range of mainstream nicotine preference, taste preference type, and brand loyalty of the low tar content preference group, and form representative consumer portraits;
[0008] According to the association rules of the high nicotine content preference group and the low tar content preference group, perform an optimized combination of the association rules. By balancing the relationships between different attributes and taking into account the optimal association combination of the preferences within the sub-group consumers, obtain the sub-group portrait model;
[0009] Apply the sub-group portrait model to actual cigarette product R & D and marketing decisions. Predict the preferences of the high nicotine content preference group and the low tar content preference group for the new product attribute combinations through the portrait model, and formulate product formula adjustment and marketing promotion strategies in combination with the demographic attribute distributions of different preference groups.
[0010] The technical solution provided by the embodiments of the present invention may include the following beneficial effects:
[0011] The present invention discloses an analysis method for cigarette consumers. The method first obtains a large amount of cigarette consumption data, and obtains a structured consumption data set through data preprocessing, including attributes such as price sensitivity and taste preference. Then, cluster and group the consumers, construct the cigarette attribute association maps of each group, and learn the low-dimensional vector representation of the attributes. For the high nicotine content and low tar content preference groups, discover the association patterns of their cigarette attribute preferences, and form representative consumer portraits. Through processing such as attribute data normalization and sample resampling, optimize the association rule combination to obtain the sub-group portrait model. Finally, apply the model to cigarette product R & D and marketing decisions, predict the preferences of different groups for new products, and formulate corresponding product formulas and marketing strategies. The present invention realizes the accurate portrait of cigarette consumers and provides effective support for cigarette product innovation and personalized marketing. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 It is a flowchart of an analysis method for cigarette consumers of the present invention.
[0013] Figure 2 It is a schematic diagram of an analysis method for cigarette consumers of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0014] The technical solution of the present invention will be clearly and completely described below in conjunction with embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0015] As Figure 1-2 , a method for analyzing cigarette consumers in this embodiment may specifically include:
[0016] Step S101, obtain a large amount of cigarette consumption data, preprocess the data, and obtain a structured consumption data set through data cleaning and data integration operations. The consumption data set includes cigarette price sensitivity, cigarette taste preference types, cigarette appearance preference styles, cigarette brand loyalty, and the frequency of purchasing cigarettes. Among them, the cigarette taste preference types include the preference value of tobacco leaf moisture content, the preference range of nicotine content in the smoke, and the attribute of the preference interval of tar content in the smoke.
[0017] Obtain the price sensitivity, brand loyalty, and purchase frequency data of consumers, use the K-means clustering algorithm to cluster consumers, and obtain the taste preference feature vectors of different consumer groups; according to the taste preference feature vectors, extract visual features such as colors and textures from the cigarette appearance preference style data, and use a support vector machine to classify the appearance styles; if the support vector machine classification is completed, use a content-based recommendation algorithm to generate a personalized recommendation list of appearance styles and taste combinations for each consumer group; obtain the geographical location distribution data and seasonal consumption trends, and use the ARIMA time series analysis method to predict the cigarette sales volume in each region; according to the sales volume prediction results and the promotion activity response data, construct a random forest model to quantify the non-linear impact of promotion activities on sales volume, and determine the optimal promotion strategy combination through feature importance; apply the optimal promotion strategy combination to the personalized recommendation list to obtain the final cigarette product sales plan.
[0018] Specifically, a consumer portrait model is constructed based on cigarette price sensitivity, brand loyalty, and purchase frequency data. The K-means clustering algorithm is used to group consumers. For each consumer group, the mean and variance of the preferred values of tobacco leaf moisture content, the preferred range of nicotine content in the smoke, and the preferred interval of tar content in the smoke are calculated to obtain the taste preference feature vectors of different consumer groups. The price sensitivity index of each consumer group is calculated using the price elasticity model, and this index is added as a feature to the taste preference feature vector. Visual features such as color and texture are extracted from the cigarette appearance preference style data, and a support vector machine is used to classify the appearance styles. Combining the consumers' purchase history records, a content-based recommendation algorithm is used to generate personalized recommendation lists of appearance styles and taste combinations for each consumer group. Based on the geographical location distribution data and seasonal consumption trends, the ARIMA time series analysis method is used to predict the cigarette sales volume in each region. Combining the promotion activity response data, a random forest model is constructed to quantify the non-linear impact of promotion activities on sales volume, and the optimal promotion strategy combination is determined through feature importance. Regarding the consumer group characteristics and cigarette taste preference types, specific attributes such as tobacco leaf moisture content, nicotine content in the smoke, and tar content in the smoke are used as feature vectors and input into the content-based recommendation system. This system inputs the consumers' demographic characteristics and historical purchase data and outputs the cigarette taste types and appearance styles that the consumers are most likely to prefer. According to the prediction results of the recommendation system, the expected demand for each type of cigarette is calculated and compared with the current inventory to generate inventory adjustment suggestions.
[0019] When building a consumer portrait model, the K-means clustering algorithm is used to cluster 100,000 consumers, with the number of clusters set to 5 and the number of iterations set to 100, resulting in 5 consumer groups with different characteristics. For each group, calculate the mean and variance of the preference values for the moisture content of tobacco leaves. For example, the mean of group A is 13% and the variance is 0.5%. Calculate the preference range for the nicotine content in the smoke. For example, the range of group B is 0.8mg - 1.2mg. Calculate the preference interval for the tar content in the smoke. For example, the interval of group C is 8mg - 10mg. Use the price elasticity model to calculate the price sensitivity index for each group. For example, the index of group D is -1.5, indicating that it is more sensitive to price. Extract color and texture features from the data on the preferred styles of cigarette appearance, and use the support vector machine for classification, with the kernel function set to RBF and the penalty parameter C set to 1.0, obtaining style categories such as "traditional classic" and "fashionable luxury". Based on the content-based recommendation algorithm combined with the consumer purchase history, generate a recommendation list for each group. For example, recommend "fashionable luxury" style cigarettes with a moisture content of 12% - 14%, a nicotine content of 0.9mg - 1.1mg, and a tar content of 9mg - 11mg for group E. Use the ARIMA model to predict the cigarette sales volume in each region, with the parameters set as p = 1, d = 1, q = 1, and predict the sales volume trend for the next 3 months. Build a random forest model to quantify the impact of promotional activities, with the number of trees set to 100 and the maximum depth set to 10. Feature importance analysis shows that price discounts have the greatest impact on sales, with a weight of 0.4. Use attributes such as the moisture content of tobacco leaves, the nicotine content in the smoke, and the tar content in the smoke as feature vectors and input them into the content-based recommendation system. This system uses cosine similarity to calculate the matching degree between consumer preferences and product features, with the threshold set to 0.8, and outputs the cigarette types with a matching degree greater than the threshold. According to the prediction results of the recommendation system, calculate the expected demand for each type of cigarette. For example, the expected demand for a certain brand is 100,000 cartons, and the current inventory is 80,000 cartons, generating an adjustment suggestion to increase the inventory by 20,000 cartons.
[0020] Step S102: According to the cigarette attributes in the consumption dataset, cluster and group consumers, and divide consumers with similar attribute preferences into the same sub-group. The cigarette attributes include cigarette price sensitivity, cigarette taste preference type, cigarette appearance preference style, cigarette brand loyalty, and frequency of purchasing cigarettes.
[0021] According to the attribute data in the consumption dataset, perform standardization processing to obtain the standardized attribute data. Use the principal component analysis method to perform dimensionality reduction processing on the standardized attribute data to obtain the dimensionality-reduced feature vectors. Use the silhouette coefficient method to analyze the dimensionality-reduced feature vectors to determine the optimal number of clusters. According to the optimal number of clusters, perform K-means clustering analysis on the dimensionality-reduced feature vectors to obtain the segmented groups. For the segmented groups, calculate the average value and standard deviation of each attribute within the group to construct a group feature description matrix. Based on the group feature description matrix, use the decision tree algorithm to construct a consumer classification model to determine the segmented group to which the new consumer belongs. If there are missing values in the new consumer data, use the multiple imputation method to process the missing values. Input the new consumer data into the consumer classification model, and determine the segmented group to which the new consumer belongs by comparing the attribute values with the thresholds of the decision nodes.
[0022] Specifically, according to the attributes such as cigarette price sensitivity, taste preference type, appearance preference style, brand loyalty, purchase frequency, tobacco leaf moisture content preference, nicotine content in mainstream smoke preference, and tar content in mainstream smoke preference in the consumption dataset, all attributes are standardized to ensure that each attribute has the same weight in subsequent analyses. The principal component analysis method is used to reduce the dimension of the standardized attributes. By calculating the eigenvalues and eigenvectors, the principal components with a cumulative variance contribution rate greater than 85% are selected as the feature vectors after dimension reduction. The local outlier factor algorithm is used to identify and remove outliers to improve data quality. The silhouette coefficient method is used to determine the optimal number of clusters, and this number is used as a parameter for the K-means clustering algorithm. Cluster analysis is performed on the feature vectors after dimension reduction. By calculating the Euclidean distance from the sample points to the cluster centers, consumers with similar attribute preferences are divided into the same segment. For each segment, the mean and standard deviation of each attribute within the segment are calculated to construct a group feature description matrix. The significant features of each group are extracted from the matrix. The significant features are defined as the features whose attribute values exceed one standard deviation of the group mean, such as high price sensitivity, strong preference for mint flavor, high loyalty to a specific brand, etc. Based on the feature description matrix of the segments, a decision tree algorithm is used to construct a consumer classification model. The multiple imputation method is used to handle the missing values in the new consumer data to ensure data integrity. The processed new consumer data is input into the decision tree, and by comparing the attribute values with the thresholds of the decision nodes, the new consumers are assigned to the most matching segments. The cross-validation method is used to calculate the classification accuracy to evaluate the model performance. The cigarette attributes in the consumption dataset are standardized, and each attribute value is converted into a standard normal distribution with a mean of 0 and a standard deviation of 1. The principal component analysis method is used for dimension reduction, calculating the eigenvalues and eigenvectors, and the first 5 principal components are selected as the feature vectors after dimension reduction, with a cumulative variance contribution rate reaching 87.3%. The local outlier factor algorithm is used, with a threshold set to 1.5, to identify and remove 2.1% of the outliers. By calculating the silhouette coefficients under different numbers of clusters, the optimal number of clusters is determined to be 6, and the silhouette coefficient is 0.68. The number of clusters 6 is input into the K-means algorithm to perform cluster analysis on the feature vectors after dimension reduction, resulting in 6 consumer segments. The mean and standard deviation of each attribute within each segment are calculated to construct a 6×8 group feature description matrix. The significant features are defined as the features whose attribute values exceed one standard deviation of the group mean. For example, the price sensitivity of group 1 is 1.8, which is 1.2 standard deviations higher than the mean, and the mint flavor preference of group 2 is 2.3, which is 1.5 standard deviations higher than the mean. Based on the feature description matrix, a decision tree algorithm is used to construct a consumer classification model, with the depth of the tree set to 4 and the minimum number of samples in a leaf node set to 50. The multiple imputation method is used to handle the missing values in the new consumer data, with the number of imputations set to 5 to generate a complete dataset.Input new consumer data into the decision tree. By comparing the attribute values with the thresholds of the decision nodes, for example, if the price sensitivity > 1.5, it goes to the left subtree, otherwise to the right subtree. Eventually, the new consumers are classified into one of the six segmentation groups. The 5-fold cross-validation method is used to calculate the classification accuracy, and the average accuracy is 85.7% with a standard deviation of 2.1.
[0023] Step S103: For each consumer segmentation group obtained by clustering, construct the cigarette attribute association graph of this group respectively, learn the low-dimensional vector representation of the cigarette attributes within each segmentation group in the graph, and measure the correlation between different attributes through the distance between vectors.
[0024] Use the maximum spanning tree algorithm to construct the cigarette attribute association graph, and determine the connection relationship between attribute nodes by calculating the mutual information value between attributes. Apply the node2vec algorithm for graph embedding according to the cigarette attribute association graph to obtain the low-dimensional vector representation of attribute nodes. Use the cosine similarity to calculate the distance between attribute vectors, normalize the distance value to between 0 and 1, and obtain the correlation score matrix between different attributes. If the correlation score matrix has been obtained, use the hierarchical clustering algorithm to group the attributes to obtain the attribute association clusters. For the attribute association clusters, use the t-SNE algorithm to project the high-dimensional attribute vectors onto a two-dimensional plane to generate the attribute relationship visualization graph, and the attribute relationship visualization graph is used to show the correlation intensity and clustering results between different cigarette attributes.
[0025] Specifically, based on the data of each consumer segment obtained by clustering, the maximum spanning tree algorithm is used to construct the cigarette attribute association graph for each group, and the connection relationship between attribute nodes is determined by calculating the mutual information value between attributes. Five-fold cross-validation is used to determine the optimal mutual information threshold, and edge connections are established between attribute nodes with mutual information values higher than this threshold. The node2vec algorithm is applied to the constructed cigarette attribute association graph for graph embedding, mapping each attribute node in the graph to a low-dimensional vector space. The best hyperparameters, including vector dimension, window size, and walk length, are determined through grid search to obtain the low-dimensional vector representation of each attribute. The cosine similarity is used to calculate the distance between attribute vectors, and the calculated distance values are normalized to between 0 and 1 to obtain the correlation score matrix between different attributes. The average value and standard deviation of the correlation scores are calculated to identify abnormally high or low correlations. Based on the correlation score matrix, the hierarchical clustering algorithm is used to group the attributes. The silhouette coefficient is used to determine the optimal number of clusters, and the termination condition of hierarchical clustering is controlled by setting the minimum cluster size to form attribute association clusters. Each association cluster represents a group of highly correlated cigarette attributes. The t-SNE algorithm is used to project the high-dimensional attribute vectors onto a two-dimensional plane to generate a visualization of the attribute relationships, intuitively showing the correlation strength and clustering results between different cigarette attributes. For the 5 consumer segments obtained by clustering, the maximum spanning tree algorithm is used to construct the cigarette attribute association graph respectively. The mutual information values between 12 cigarette attributes are calculated, and the optimal mutual information threshold is determined to be 0.3 through five-fold cross-validation. Taking group A as an example, an attribute association graph with 9 nodes and 8 edges is obtained. The node2vec algorithm is applied for graph embedding, and the best hyperparameters are determined through grid search: the vector dimension is 64, the window size is 10, and the walk length is 80. A 64-dimensional attribute vector representation is obtained, such as the price sensitivity attribute vector being [0.21, -0.15,..., 0.33]. The cosine similarity between attribute vectors is calculated to obtain a 12×12 correlation score matrix. In the matrix, the correlation score between price sensitivity and purchase frequency is 0.82, indicating a high correlation. The average value of the correlation scores is calculated to be 0.45, and the standard deviation is 0.18. Three pairs of abnormally high correlations, such as scores > 0.81, and two pairs of abnormally low correlations, such as scores < 0.09, are identified. The hierarchical clustering algorithm is used to group the attributes, and the optimal number of clusters is determined to be 4 through the silhouette coefficient, and the minimum cluster size is set to 2. Four attribute association clusters are obtained, such as {price sensitivity, purchase frequency, income level} being one cluster. The t-SNE algorithm is used to project the 64-dimensional attribute vectors onto a two-dimensional plane to generate a visualization graph, intuitively presenting the relative positions and clustering results of the attributes. For example, price sensitivity and purchase frequency are close in the graph and are both located in the upper left area.
[0026] For the construction of a cigarette attribute association graph for consumer segmentation groups, define the cigarette attributes for each segmentation group. Each attribute serves as a node in the graph, and the relationship between attributes serves as an edge. By analyzing the distance between attributes, identify the correlation of attribute combinations. If the distance between two attributes is small, the correlation between the two attributes is strong.
[0027] Obtain the characteristics of consumer segmentation groups, and define the cigarette attribute set for each group according to the characteristics; use the Neo4j graph database to store the association graph of the cigarette attribute set, take each attribute in the cigarette attribute set as a node in the graph, and the relationship between attributes as an edge; use the Node2Vec algorithm to map the nodes in the attribute association graph to a low-dimensional vector space to obtain the vector representation of each attribute node; calculate the Euclidean distance between attribute vectors according to the vector representation of the attribute nodes to construct a distance matrix; if the distance in the distance matrix is less than a preset distance threshold, it is determined as a strongly correlated attribute; use the hierarchical clustering algorithm to group the strongly correlated attributes to form attribute combinations; perform visualization processing on the attribute association graph through Gephi software to generate the visualization result of the attribute association graph, where the size of the node represents the importance of the attribute, and the thickness of the edge represents the strength of the relationship between attributes.
[0028] Specifically, according to the characteristics of consumer segments, a specific set of cigarette attributes is defined for each segment. The Neo4j graph database is used to store the attribute association graph. Each attribute is used as a node in the graph, and the relationship between attributes is used as an edge. The weight of the edge is determined by calculating the Pearson correlation coefficient of the attribute values. The Node2Vec algorithm is used to map the nodes in the attribute association graph to a low-dimensional vector space. The optimal vector dimension is determined through five-fold cross-validation, and the vector representation of each attribute node is obtained. The Euclidean distance between the attribute vectors is calculated to construct a distance matrix. The interquartile range method is used to determine the distance threshold, and the distance less than the first quartile is determined to be strongly correlated. Based on the distance matrix, the hierarchical clustering algorithm is used to group the attributes to form attribute combinations. The optimal number of clusters is determined by calculating the silhouette coefficient for different numbers of clusters, and the minimum cluster size is set to 2 to control the termination condition of hierarchical clustering, obtaining the final attribute combination result. Each combination represents a group of strongly correlated cigarette attributes. The Gephi software is used to visualize the constructed attribute association graph, generating the graph visualization result. The importance of the attribute is represented by the node size, and the strength of the relationship between attributes is represented by the thickness of the edge, intuitively showing the relationship strength and clustering result between different cigarette attributes. For a certain consumer segment, 10 specific cigarette attributes are defined, including price sensitivity, taste preference, etc. The Neo4j graph database is used to store the attribute association graph. The weight of the edge is determined by calculating the Pearson correlation coefficient of the attribute values. For example, the correlation coefficient between price sensitivity and purchase frequency is 0.75, indicating a strong correlation. The Node2Vec algorithm is applied for graph embedding, and the optimal vector dimension is determined to be 64 through five-fold cross-validation, obtaining the 64-dimensional vector representation of each attribute node. The Euclidean distance between the attribute vectors is calculated to construct a 10×10 distance matrix. The interquartile range method is used to determine the distance threshold, and the first quartile is 0.3. The distance less than 0.3 is determined to be strongly correlated. Based on the distance matrix, Ward's method is used for hierarchical clustering. By calculating the silhouette coefficient for different numbers of clusters, from 2 to 8, the optimal number of clusters is determined to be 4, and the silhouette coefficient is 0.68. The minimum cluster size is set to 2, obtaining 4 attribute combinations. For example, {price sensitivity, purchase frequency, income level} is a combination. The Gephi software is used to visualize the attribute association graph. The node size is set according to the PageRank value, ranging from 10 to 50 pixels, and the edge thickness is set according to the correlation coefficient, ranging from 1 to 5 pixels. The visualization result shows that the price sensitivity node is the largest, such as 50 pixels, and the connection edge with the purchase frequency is the thickest, such as 5 pixels, intuitively reflecting the relationship strength and clustering result between attributes.
[0029] Step S104: Based on the cigarette attribute association graphs of each sub-group, discover the association patterns in the cigarette attribute preferences of different sub-groups, obtain the association rules of the leaf moisture content preference, cigarette price sensitivity, and appearance preference style for the high nicotine content preference group, as well as the association rules of the range of flue-cured tobacco nicotine content preference, taste preference type, and brand loyalty for the low tar content preference group, and form a representative consumer portrait.
[0030] Use the Apriori algorithm to extract high-frequency attribute combinations, obtain the optimal minimum support threshold through five-fold cross-validation, and the minimum support threshold is used to screen the attribute association patterns with frequencies greater than the preset threshold; according to the high-frequency attribute combinations, use the FP-Growth algorithm to analyze the association relationships between attribute combinations, and determine the optimal minimum confidence and minimum lift thresholds through grid search, and the optimal minimum confidence and minimum lift thresholds are used to obtain strong association rules that meet the conditions; for the high nicotine content preference group and the low tar content preference group, extract the association rules of specific attribute combinations from the strong association rules, and the association rules of the specific attribute combinations are used to calculate the group association strength; convert the association rules into group feature vectors, use one-hot encoding to obtain binary features, and the binary features are used to represent the group attribute preferences; if the dimension of the binary features is too high, use principal component analysis to reduce the dimension of the binary features to obtain the reduced-dimensional feature vectors; use the K-means algorithm to cluster the reduced-dimensional feature vectors, and determine the optimal number of clusters by calculating the silhouette coefficients of different numbers of clusters, and the optimal number of clusters is used to form a representative consumer portrait.
[0031] Specifically, for the cigarette attribute association graph of each segment group, the Apriori algorithm is used to extract high-frequency attribute combinations. The optimal minimum support threshold is determined through five-fold cross-validation, and the attribute association patterns with relatively high frequencies are screened out. The FP-Growth algorithm is used to analyze the association relationships among high-frequency attribute combinations. The optimal minimum confidence and minimum lift thresholds are determined through grid search, and strong association rules that meet the conditions are obtained. According to the mined association rules, for the group with a preference for high nicotine content and the group with a preference for low tar content, the association rules of their specific attribute combinations are extracted, and the association strengths of these two groups' specific attribute combinations are calculated using conditional probability. The association rules are transformed into group feature vectors, and the rules are transformed into binary features using one-hot encoding. Based on the group feature vectors, principal component analysis is used to reduce the dimension of high-dimensional features, and the principal components with an explained variance ratio exceeding 85% are retained. The K-means algorithm is used to cluster the features after dimension reduction, and the optimal number of clusters is determined by calculating the silhouette coefficients for different numbers of clusters, forming a representative consumer portrait. The radar chart visualization tool is used to generate a portrait radar chart, and the attributes with the greatest contribution of the principal components are selected as the dimensions of the radar chart to intuitively display the attribute preference characteristics of different groups.
[0032] For the attribute association graph of a certain cigarette consumer group, the Apriori algorithm is applied to extract high-frequency attribute combinations. The optimal minimum support threshold is determined to be 0.15 through five-fold cross-validation, and 87 high-frequency attribute combinations are screened out. Subsequently, the FP-Growth algorithm is used to analyze the association relationships among these combinations. Through grid search, the optimal minimum confidence is found to be 0.6 and the minimum lift is 1.2, and 53 strong association rules are obtained. For the group with a preference for high nicotine content, the rule "tobacco leaf moisture content 12 - 14% → low cigarette price sensitivity" is extracted, and the conditional probability is 0.78. In the group with a preference for low tar content, the rule "smoke nicotine content 0.6 - 0.8 mg → preference for light taste" is found, and the association strength is 0.85. These rules are transformed into 200-dimensional binary feature vectors through one-hot encoding. Principal component analysis is used to reduce the dimension of the features, and the first 8 principal components are retained, with the cumulative explained variance ratio reaching 87%. The K-means algorithm is used to cluster the features after dimension reduction, and the silhouette coefficients for the number of clusters from 2 to 10 are calculated to determine that the optimal number of clusters is 5, and the silhouette coefficient is 0.68. Finally, the original attributes corresponding to the 6 principal components with the greatest contribution, such as price sensitivity and nicotine content preference, are selected as the dimensions of the radar chart, and a consumer portrait radar chart of 5 groups is generated to intuitively display the feature distributions of each group in these 6 dimensions.
[0033] According to the data of cigarette price, tobacco leaf moisture content, nicotine content in mainstream smoke, taste type, and appearance style attribute preferences, normalize the attribute data with different dimensions, map the attribute data to the same scale space, resample the samples of the high-nicotine-preference group and the low-tar-preference group, balance the sample sizes of different sub-groups, and adapt to the association pattern characteristics of different groups by setting different support and confidence thresholds.
[0034] For the cigarette attribute data, use the min-max normalization method for normalization to obtain the standardized attribute data set. According to the standardized attribute data set, use the SMOTE algorithm to resample the samples of the high-nicotine-preference group and the low-tar-preference group to generate a balanced sample data set. If the balanced sample data set is obtained, apply the Apriori algorithm to it for association rule mining, and set the support and confidence threshold combinations through the grid search method. From the association rule mining results, calculate the lift of each rule and determine whether it is greater than the preset threshold. If the rule lift is greater than the preset threshold, screen this rule as an effective association pattern. For the selected effective association patterns, use ten-fold cross-validation to evaluate their stability and generalization ability, and determine the rules that are stable on the validation set.
[0035] Specifically, for attribute data such as cigarette price, tobacco leaf moisture content, nicotine content in mainstream smoke, taste type, and appearance style, the min-max normalization method is used for normalization processing to map attribute data with different dimensions to a unified interval from 0 to 1, obtaining a standardized attribute data set. According to the standardized attribute data, the SMOTE algorithm is used to resample the samples of the high-nicotine-preference group and the low-tar-preference group. By generating synthetic samples, the sample numbers of the two groups are balanced. The balanced sample number is set to 1.5 times the sample number of the larger group, obtaining a balanced sample data set. The Apriori algorithm is applied to the balanced sample data set for association rule mining, and different combinations of support and confidence thresholds are set through the grid search method. The support threshold range is set from 0.1 to 0.5 with a step size of 0.1; the confidence threshold range is set from 0.5 to 0.9 with a step size of 0.1. Association rule mining is carried out separately for the high-nicotine-preference group and the low-tar-preference group, and indicators such as the number of rules and average confidence of different groups are compared to adapt to the association mode characteristics of different groups. Based on the mined association rules, the lift of each rule is calculated, and the rules with a lift greater than 1 are selected as effective association patterns. For the possible conflicting rules, the rule confidence is used as the weight for rule merging to form the characteristic association rule sets of the high-nicotine-preference group and the low-tar-preference group. The ten-fold cross-validation is used to evaluate the stability and generalization ability of the association rules, and the rules that perform stably on the validation set are retained. When processing the cigarette attribute data, first, the min-max normalization method is used to map the cigarette price from the range of 50 - 200 yuan to the interval of 0 - 1, the tobacco leaf moisture content from 10% - 18% to 0 - 1, and the nicotine content in mainstream smoke from 0.5 - 1.5 mg to 0 - 1. Subsequently, the SMOTE algorithm is used to resample the high-nicotine-preference group with an original sample size of 5000 and the low-tar-preference group with an original sample size of 3000, and the sample numbers of both groups are adjusted to 7500. The Apriori algorithm is applied to the balanced samples for association rule mining, and the support thresholds (0.1, 0.2, 0.3, 0.4, 0.5) and confidence thresholds (0.5, 0.6, 0.7, 0.8, 0.9) are set through the grid search, with a total of 25 combinations. The high-nicotine-preference group obtains the optimal rule set with 87 rules and an average confidence of 0.82 at a support of 0.2 and a confidence of 0.7; the low-tar-preference group obtains the optimal rule set with 62 rules and an average confidence of 0.89 at a support of 0.3 and a confidence of 0.8. The rule lift is calculated, and the rules with a lift greater than 1 are selected. The high-nicotine-preference group retains 73 rules, and the low-tar-preference group retains 58 rules. For conflicting rules, such as "high price → high quality" with a confidence of 0.8 and "high price → low quality" with a confidence of 0.6, the weighted combination with a confidence of 0.71 is obtained by confidence weighting to get "high price → high quality".Through ten-fold cross-validation, rules with a confidence fluctuation less than 10% on the validation set are retained. Finally, 65 rules are retained for the high nicotine content preference group, and 52 rules are retained for the low tar content preference group. These rules constitute the characteristic association rule sets of the two groups, reflecting the preference characteristics of different consumer groups.
[0036] Step S105, according to the association rules of the high nicotine content preference group and the low tar content preference group, perform an optimized combination of the association rules. By balancing the relationships between different attributes and taking into account the optimal association combination of the preferences within the segmented groups, a segmented group portrait model is obtained.
[0037] The weighted average method is used to calculate the rule weights, where the support, confidence, and lift are respectively assigned preset proportion weights to obtain the weighted association rule set. According to the weighted association rule set, an association rule network is constructed using the adjacency matrix representation method, with attributes as nodes and rules as edges, and the weight of the edge is set to the comprehensive weight of the corresponding rule. The Louvain algorithm is applied to the association rule network for community discovery. By adjusting the resolution parameter and the step size, the community division result with the largest modularity is selected to identify the closely associated attribute clusters. If attribute clusters are identified, each attribute cluster is transformed into a binary feature vector using one-hot encoding, and combined with the original demographic characteristics to construct a multi-dimensional feature vector of the segmented group. For the multi-dimensional feature vector, the sensitivity analysis method is used. By changing each attribute value one by one and observing the impact on the final portrait, the importance weights of each attribute in the portrait model are determined. According to the sensitivity analysis results, different weights are assigned to each attribute in the feature vector to form a weighted feature vector. The silhouette coefficient is used to evaluate the clustering effect of the portrait model. By comparing the silhouette coefficients under different weight combinations, the optimal weight configuration is selected to obtain the segmented group portrait model.
[0038] Specifically, for the association rules of the high nicotine preference group and the low tar preference group, the weighted average method is used to calculate the rule weights. The support, confidence, and lift are respectively assigned weights of 0.3, 0.4, and 0.3, and the weights of each rule are comprehensively calculated to obtain the weighted association rule set. The adjacency matrix representation method is used to construct an association rule network, with attributes as nodes and rules as edges. The weight of the edge is set to the comprehensive weight of the corresponding rule to generate a rule network diagram. The Louvain algorithm is applied to the constructed rule network diagram for community discovery. By adjusting the resolution parameter from 0.5 to 2.0 with a step size of 0.1, the community division result with the largest modularity is selected to identify closely related attribute clusters, and each attribute cluster represents a group of interrelated cigarette attribute preference characteristics. Based on the identified attribute clusters, one-hot encoding is used to transform each attribute cluster into a binary feature vector, and combined with the original demographic characteristics, a multi-dimensional feature vector of the segmented population is constructed. The sensitivity analysis method is used to determine the importance weights of each attribute in the portrait model by changing each attribute value one by one and observing the impact on the final portrait. According to the sensitivity analysis results, different weights are assigned to each attribute in the feature vector to form a weighted feature vector. The silhouette coefficient is used to evaluate the clustering effect of the portrait model, and by comparing the silhouette coefficients under different weight combinations, the optimal weight configuration is selected to finally form the portrait model of the segmented population.
[0039] Weighted average calculations are performed on 100 association rules for the high nicotine preference group and 80 association rules for the low tar preference group. For example, for the rule "high nicotine content → strong taste", the support is 0.3, the confidence is 0.8, and the lift is 1.5. After weighting, the comprehensive weight is 0.87. A 10×10 adjacency matrix is used to represent the rule network, where the matrix element a[i][j] represents the rule weight from attribute i to attribute j. For example, a[high nicotine content][strong taste] = 0.87. The Louvain algorithm is applied for community discovery, with the resolution parameter set from 0.5 to 2.0 and a step size of 0.1, for a total of 16 iterations. The maximum modularity of 0.68 is obtained at a resolution of 1.3, and 5 attribute clusters are identified. The attribute clusters are transformed into 25-dimensional binary feature vectors. For example, cluster 1 [1,0,1,1,0] indicates that this cluster contains the 1st, 3rd, and 4th attributes. Combining 5 demographic characteristics such as age and income, a 30-dimensional feature vector is constructed. Sensitivity analysis is carried out. For example, when the nicotine preference is increased by 10%, the portrait similarity changes by 5%, and accordingly a weight of 0.5 is assigned; when the price sensitivity is increased by 10%, the similarity changes by 3%, and a weight of 0.3 is assigned. These weights are used to weight the 30-dimensional feature vector to obtain a weighted feature vector. The silhouette coefficients under different weight combinations are calculated, and the optimal combination gives a silhouette coefficient of 0.72, based on which the final portrait model of the segmented population is determined.
[0040] 0.68, and 5 attribute clusters are identified. The attribute clusters are transformed into 25-dimensional binary feature vectors. For example, cluster 1 [1,0,1,1,0] indicates that this cluster contains the 1st, 3rd, and 4th attributes. Combining 5 demographic characteristics such as age and income, a 30-dimensional feature vector is constructed. Sensitivity analysis is carried out. For example, when the nicotine preference is increased by 10%, the portrait similarity changes by 5%, and accordingly a weight of 0.5 is assigned; when the price sensitivity is increased by 10%, the similarity changes by 3%, and a weight of 0.3 is assigned. These weights are used to weight the 30-dimensional feature vector to obtain a weighted feature vector. The silhouette coefficients under different weight combinations are calculated, and the optimal combination gives a silhouette coefficient of 0.72, based on which the final portrait model of the segmented population is determined.
[0041] Step S106: Apply the segmented population portrait model to the actual cigarette product R & D and marketing decisions. Predict the preferences of high nicotine preference groups and low tar preference groups for the new product attribute combinations through the portrait model, and formulate product formula adjustment and marketing promotion strategies in combination with the demographic attribute distributions of different preference groups.
[0042] Obtain the importance weights of each attribute of the cigarette product, where the importance weights are determined by the attribute importance in the portrait model; construct a cigarette product attribute scoring function according to the importance weights, and the scoring function comprehensively scores each attribute by using the weighted summation method; input the attribute values of the new product into the scoring function to obtain the preference score of the target group for the new product; use the preference score as the dependent variable and the product attributes as the independent variables to construct a multiple linear regression model; if the variance inflation factor is greater than the preset threshold, perform regularization processing on the corresponding variables; determine the influence degree of different attributes on the preference score according to the regression coefficients of the multiple linear regression model; optimize the product formula by using the grid search method, where the grid search method takes maximizing the preference score of the target group as the optimization goal to obtain the optimal product attribute combination; obtain the demographic attribute distribution of the target group, use the optimal product attribute combination and the target group characteristics as inputs, and use the decision tree algorithm to construct a marketing strategy decision model; determine targeted marketing promotion strategy suggestions according to the output results of the decision tree algorithm; conduct sensitivity analysis on the marketing promotion strategy suggestions to judge the influence degree of different attribute changes on the final decision.
[0043] Specifically, according to the sub-group portrait model, a scoring function for cigarette product attributes is constructed. The scores of each attribute are comprehensively calculated by weighted summation, and the weights are determined according to the importance of the attributes in the portrait model. The numerical values of the attributes of the new product are input into the scoring function to obtain the preference scores of the high-nicotine-preference group and the low-tar-preference group for the new product. Using the preference scores as the dependent variable and the product attributes as the independent variables, a multiple linear regression model is constructed. During the construction process, the variance inflation factor is used to detect multicollinearity, and regularization is performed on variables with a VIF greater than 10. The influence degree of different attributes on the preference scores is analyzed through the regression coefficients to identify the key influencing factors. Based on the results of the regression model, the grid search method is used to optimize the product formula. The value range of the attributes is set as the constraint condition, and the maximization of the preference scores of the target group is used as the optimization goal to obtain the optimal combination of product attributes. Five-fold cross-validation is used to evaluate the model stability, and the model with the smallest average prediction error is selected. Combining the demographic attribute distributions of different preference groups, the decision tree algorithm is used to construct a marketing strategy decision model. The optimal product attribute combination and the characteristics of the target group are used as inputs, and targeted marketing promotion strategy suggestions are output. Decision rules are set to transform the decision tree results into specific marketing strategies, such as "if the age group is 25-35 years old, then use social media promotion". Sensitivity analysis is carried out to evaluate the influence degree of different attribute changes on the final decision and identify the key decision factors.
[0044] When constructing the scoring function for cigarette product attributes, the weights of the four attributes of nicotine content, tar content, moisture content, and price are set to 0.3, 0.25, 0.25, and 0.2 respectively. The new product for the high-nicotine-preference group is scored, and the preference score of 0.85 is obtained. Using these scores to construct a multiple linear regression model, it is calculated that the VIF values of tar content and nicotine content are 12 and 15 respectively, exceeding the threshold of 10. Ridge regression is applied to these two variables for regularization. The regression analysis results show that the influence coefficient of nicotine content on the preference score is the largest, which is 0.6. The grid search method is used to optimize the product formula. The nicotine content range is set to 0.8 - 1.2 mg, the tar content is 6 - 10 mg, the moisture content is 12 - 16%, and the price is 50 - 100 yuan. The step sizes are 0.1 mg, 1 mg, 1%, and 10 yuan respectively, and the optimal combination is obtained: nicotine content 1.1 mg, tar content 8 mg, moisture content 14%, and price 80 yuan. Five-fold cross-validation is used to evaluate the model, and the average prediction error is 0.03. A decision tree marketing strategy model is constructed, and the optimal product attributes and characteristics of the target group such as age 25 - 35 years old and monthly income 8000 - 12000 yuan are input, and the strategy suggestion "use social media promotion and emphasize the taste characteristics of the product" is output. Sensitivity analysis shows that a 1% change in nicotine content results in a 5% probability of decision change, which is the most influential factor.
[0045] The above embodiments are only one of the preferred embodiments of the present invention and should not be used to limit the protection scope of the present invention. Any meaningless changes or polish made on the main design concept and spirit of the present invention, as long as the technical problems solved are still consistent with those of the present invention, should be included within the protection scope of the present invention.
Claims
1. A method for analyzing cigarette consumers, characterized in that: The method comprises: Obtain massive amounts of cigarette consumption data, pre-process the data, and obtain a structured consumption data set through data cleaning and data integration operations. The consumption data set includes cigarette price sensitivity, cigarette flavor preference type, cigarette appearance preference style, cigarette brand loyalty, and cigarette purchase frequency. The cigarette flavor preference type includes the preferred value of tobacco leaf moisture content, the preferred range of smoke nicotine content, and the preferred interval attributes of smoke tar content; According to the cigarette attributes in the consumption data set, consumers are clustered and grouped, and consumers with similar attribute preferences are divided into the same segment group. The cigarette attributes include cigarette price sensitivity, cigarette flavor preference type, cigarette appearance preference style, cigarette brand loyalty, and cigarette purchase frequency; For each consumer segment group obtained by clustering, a cigarette attribute association map of the group is constructed respectively, and the low-dimensional vector representation of the cigarette attributes within each segment group in the map is learned. The correlation between different attributes is measured by the distance between vectors. Based on the association maps of cigarette attributes of each segmented group, the association patterns of cigarette attribute preferences of different segmented groups were found, and the association rules of tobacco leaf moisture content preference, cigarette price sensitivity and appearance preference style of the high nicotine content preference group, and the association rules of smoke nicotine content preference range, flavor preference type and brand loyalty of the low tar content preference group were obtained, forming a representative consumer portrait; According to the association rules of the high nicotine preference group and the low tar preference group, the association rules are optimized and combined. By balancing the relationship between different attributes and taking into account the optimal association combination of consumer preferences within the segmented group, a segmented group portrait model is obtained. Apply the segmented group portrait model to actual cigarette product development and marketing decisions. Use the portrait model to predict the preferences of high-nicotine preference groups and low-tar preference groups for new product attribute combinations. Combined with the demographic attribute distribution of different preference groups, formulate product formula adjustment and marketing promotion strategies, including: Obtaining the importance weight of each attribute of the cigarette product, wherein the importance weight is determined by the attribute importance in the portrait model; Constructing a cigarette product attribute scoring function according to the importance weights, wherein the scoring function uses a weighted summation method to comprehensively calculate the scores of each attribute; Inputting the attribute values of the new product into the scoring function to obtain the preference score of the target group for the new product; Using the preference scores as dependent variables and product attributes as independent variables, a multiple linear regression model is constructed; If the variance inflation factor is greater than the preset threshold, the corresponding variable is regularized; Determining the degree of influence of different attributes on the preference score based on the regression coefficients of the multivariate linear regression model; A grid search method is used to optimize the product formula, wherein the grid search method takes maximizing the preference score of the target group as the optimization goal to obtain the optimal product attribute combination; Obtain the demographic attribute distribution of the target group, take the optimal product attribute combination and target group characteristics as input, and use the decision tree algorithm to build a marketing strategy decision model; Determine targeted marketing promotion strategy recommendations based on the output results of the decision tree algorithm; A sensitivity analysis is recommended for the marketing promotion strategy to determine the impact of changes in different attributes on the final decision.
2. The method according to claim 1, characterized in that The massive cigarette consumption data is obtained, the data is preprocessed, and a structured consumption data set is obtained through data cleaning and data integration operations. The consumption data set includes cigarette price sensitivity, cigarette flavor preference type, cigarette appearance preference style, cigarette brand loyalty, and cigarette purchase frequency, wherein the cigarette flavor preference type includes the tobacco leaf moisture content preference value, the smoke nicotine amount preference range, and the smoke tar amount preference interval attributes, including: Obtain consumers’ price sensitivity, brand loyalty, and purchase frequency data, use the K-means clustering algorithm to group consumers, and obtain the taste preference feature vectors of different consumer groups; Extracting visual features such as color and texture from the cigarette appearance preference style data according to the taste preference feature vector, and classifying the appearance style using a support vector machine; If the support vector machine classification is completed, a content-based recommendation algorithm is used to generate a personalized appearance style and taste combination recommendation list for each consumer group; Obtain geographic location distribution data and seasonal consumption trends, and use ARIMA time series analysis to predict cigarette sales in each region; Based on the sales forecast results and promotional activity responsiveness data, a random forest model is constructed to quantify the nonlinear impact of promotional activities on sales, and the optimal promotion strategy combination is determined by feature importance; The optimal promotion strategy combination is applied to the personalized recommendation list to obtain the final cigarette product sales plan.
3. The method according to claim 1, characterized in that According to the cigarette attributes in the consumption data set, consumers are clustered and grouped, and consumers with similar attribute preferences are divided into the same segmented group. The cigarette attributes include cigarette price sensitivity, cigarette flavor preference type, cigarette appearance preference style, cigarette brand loyalty, and cigarette purchase frequency, including: Perform standardization processing on each attribute data in the consumption data set to obtain standardized attribute data; The principal component analysis method is used to reduce the dimension of the standardized attribute data and obtain the feature vector after dimension reduction; Using the silhouette coefficient method to analyze the eigenvector after dimension reduction to determine the optimal number of clusters; According to the optimal number of clusters, K-means cluster analysis is performed on the feature vector after dimension reduction to obtain segmented groups; For the segmented groups, the average value and standard deviation of each attribute in the group are calculated to construct a group characteristic description matrix; Based on the group feature description matrix, a consumer classification model is constructed using a decision tree algorithm to determine the segmented group to which the newly added consumer belongs; If there are missing values in the newly added consumer data, multiple imputation methods are used to handle the missing values; The newly added consumer data is input into the consumer classification model, and the segment group to which the newly added consumer belongs is determined by comparing the attribute value with the threshold of the decision node.
4. The method according to claim 1, characterized in that: For each consumer segment group obtained by clustering, a cigarette attribute association map of the group is constructed respectively, and a low-dimensional vector representation of the cigarette attributes within each segment group in the map is learned, and the correlation between different attributes is measured by the distance between vectors, including: The maximum spanning tree algorithm is used to construct the cigarette attribute association map, and the connection relationship between attribute nodes is determined by calculating the mutual information value between attributes. Applying the node2vec algorithm to perform graph embedding according to the cigarette attribute association graph to obtain a low-dimensional vector representation of the attribute node; The distance between attribute vectors is calculated using cosine similarity, and the distance value is normalized to between 0 and 1 to obtain the correlation score matrix between different attributes; If the correlation score matrix has been obtained, a hierarchical clustering algorithm is used to group the attributes to obtain attribute association clusters; For the attribute association cluster, the t-SNE algorithm is used to project the high-dimensional attribute vector onto a two-dimensional plane to generate an attribute relationship visualization diagram, which is used to display the correlation strength and clustering results between different cigarette attributes; It also includes: constructing a cigarette attribute association map for consumer segment groups, defining the cigarette attributes for each segment group, using each attribute as a node in the map, and the relationship between attributes as an edge. By analyzing the distance between attributes, the correlation of attribute combinations is identified. If the distance between two attributes is small, the correlation between the two attributes is strong.
5. The method according to claim 4, characterized in that The cigarette attribute association map for constructing consumer segment groups is described as follows: for each segment group, its cigarette attribute is defined, each attribute is used as a node in the graph, and the relationship between the attributes is used as an edge. By analyzing the distance between the attributes, the correlation of the attribute combination is identified. If the distance between the two attributes is small, the correlation between the two attributes is strong, including: Obtain the characteristics of consumer segment groups and define the cigarette attribute set for each group based on the characteristics; A Neo4j graph database is used to store the association graph of the cigarette attribute set, each attribute in the cigarette attribute set is used as a node in the graph, and the relationship between the attributes is used as an edge; Use the Node2Vec algorithm to map the nodes in the attribute association graph to a low-dimensional vector space to obtain the vector representation of each attribute node; Calculate the Euclidean distance between attribute vectors according to the vector representation of the attribute nodes, and construct a distance matrix; If the distance in the distance matrix is less than a preset distance threshold, it is determined to be a strongly correlated attribute; Using a hierarchical clustering algorithm to group the strongly correlated attributes to form attribute combinations; The attribute association map is visualized by using Gephi software to generate a visualization result of the attribute association map, in which the size of the node represents the importance of the attribute, and the thickness of the edge represents the strength of the relationship between the attributes.
6. The method according to claim 1, characterized in that Based on the cigarette attribute association map of each segmented group, the association pattern of cigarette attribute preferences of different segmented groups is found, and the association rules of tobacco leaf moisture content preference, cigarette price sensitivity and appearance preference style of the high nicotine content preference group, and the association rules of smoke nicotine content preference range, taste preference type and brand loyalty of the low tar content preference group are obtained, forming a representative consumer portrait, including: The Apriori algorithm is used to extract high-frequency attribute combinations, and the optimal minimum support threshold is obtained through five-fold cross validation. The minimum support threshold is used to filter attribute association patterns whose occurrence frequency is greater than a preset threshold. According to the high-frequency attribute combination, the FP-Growth algorithm is used to analyze the correlation between the attribute combinations, and the optimal minimum confidence and minimum lift thresholds are determined through grid search. The optimal minimum confidence and minimum lift thresholds are used to obtain strong association rules that meet the conditions; For the high nicotine content preference group and the low tar content preference group, extracting association rules of specific attribute combinations from the strong association rules, and the association rules of the specific attribute combinations are used to calculate group association strength; The association rules are converted into group feature vectors, and one-hot encoding is used to obtain binary features, where the binary features are used to represent group attribute preferences; If the binary feature dimension is too high, principal component analysis is used to reduce the binary feature dimension to obtain the reduced feature vector; Clustering the reduced-dimensional feature vectors using a K-means algorithm, and determining an optimal number of clusters by calculating silhouette coefficients of different numbers of clusters, wherein the optimal number of clusters is used to form a representative consumer portrait; It also includes: normalizing attribute data of different dimensions according to cigarette price, tobacco moisture content, smoke nicotine content, flavor type, and appearance style attribute preference data, mapping the attribute data to the same scale space, resampling samples of high nicotine preference groups and low tar preference groups, balancing the sample size of different sub-groups, and adapting to the association pattern characteristics of different groups by setting different support and confidence thresholds.
7. The method according to claim 6, characterized in that According to the attribute preference data of cigarette price, tobacco leaf moisture content, smoke nicotine content, flavor type, and appearance style, attribute data of different dimensions are normalized, attribute data are mapped to the same scale space, samples of high nicotine content preference group and low tar content preference group are resampled, sample sizes of different subdivided groups are balanced, and different support and confidence thresholds are set to adapt to the characteristics of association patterns of different groups, including: For cigarette attribute data, the minimum-maximum normalization method is used to normalize them and obtain the standardized attribute data set; According to the standardized attribute data set, the SMOTE algorithm is used to resample the samples of the high nicotine preference group and the low tar preference group to generate a balanced sample data set; If the balanced sample data set is obtained, the Apriori algorithm is applied to it to mine association rules, and the support and confidence threshold combinations are set by the grid search method; From the association rule mining results, calculate the improvement of each rule and determine whether it is greater than the preset threshold; If the rule improvement is greater than the preset threshold, the rule is filtered as a valid association mode; For the screened effective association patterns, ten-fold cross validation is used to evaluate their stability and generalization ability, and to determine the rules that are stable on the validation set.
8. The method according to claim 1, characterized in that The association rules of the high nicotine preference group and the low tar preference group are optimized and combined, and the segmented group portrait model is obtained by balancing the relationship between different attributes and taking into account the optimal association combination of consumer preferences within the segmented group, including: The weighted average method is used to calculate the rule weights, where the support, confidence and lift are assigned preset weights respectively to obtain the weighted association rule set; According to the weighted association rule set, an association rule network is constructed using an adjacency matrix representation method, with attributes as nodes, rules as edges, and edge weights set to the comprehensive weights of corresponding rules; Applying the Louvain algorithm to the association rule network for community discovery, by adjusting the resolution parameter and the step size, selecting the community division result with the largest modularity, and identifying closely related attribute clusters; If attribute clusters are identified, one-hot encoding is used to convert each attribute cluster into a binary feature vector, and combined with the original demographic characteristics, a multidimensional feature vector of the segmented group is constructed; For the multidimensional feature vector, a sensitivity analysis method is used to determine the importance weight of each attribute in the portrait model by changing each attribute value one by one and observing the degree of influence on the final portrait; According to the sensitivity analysis results, different weights are assigned to each attribute in the feature vector to form a weighted feature vector; The silhouette coefficient is used to evaluate the clustering effect of the portrait model. By comparing the silhouette coefficients under different weight combinations, the optimal weight configuration is selected to obtain the segmented group portrait model.
Citation Information
Patent Citations
Cigarette brand recommendation algorithm based on consumer modeling
CN111275459A
Consumer evaluation method and device based on RFM model, and medium
CN115115265A