Construction method of county tobacco market order intelligent evaluation model based on multi-source data fusion and feature optimization

By integrating multi-source data and feature optimization, multi-dimensional data is combined and an intelligent evaluation model is constructed, which solves the complexity of market order assessment in tobacco monopoly management, realizes efficient and visualized market order assessment and risk warning, and improves regulatory effectiveness.

CN122134393APending Publication Date: 2026-06-02FUJIAN TOBACCO CO LONGYAN CO

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
FUJIAN TOBACCO CO LONGYAN CO
Filing Date
2026-02-11
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing technologies are insufficient for effectively assessing the order of the county-level cigarette market in tobacco monopoly management. They suffer from problems such as complex indicator systems, high data collection costs, insufficient model interpretability and generalization ability, and lack of systematic feature screening processes, resulting in insufficient robustness and reliability of the assessment models.

Method used

By employing a multi-source data fusion and feature optimization approach, multi-dimensional data is integrated, and key features are selected through Pearson correlation analysis and unsupervised learning weight analysis. A classification model is then constructed by combining support vector machine, random forest, backpropagation neural network, and Naive Bayes algorithm to conduct intelligent assessment of the order of the cigarette market in the county.

Benefits of technology

It has enabled efficient and visual assessment and risk warning of the county-level cigarette market order, improved the ability of regulatory resources to be allocated according to risk levels, provided a more scientific and objective assessment system, and enhanced the interpretability and generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122134393A_ABST
    Figure CN122134393A_ABST
Patent Text Reader

Abstract

This invention relates to a method for constructing an intelligent assessment model of county-level cigarette market order based on multi-source data fusion and feature optimization, belonging to the field of tobacco monopoly supervision and public governance technology. The method integrates heterogeneous data from multiple sources, including licensing management and market supervision, to construct a 5-dimensional, 31-item initial indicator set. Through Pearson correlation analysis, indicators significantly correlated with expert scores are selected. Combined with unsupervised learning (K-means, hierarchical clustering, GNN) weight analysis (using the F-value of variance analysis to measure feature discriminative power), six core features, including the licensed operation rate (K3) and market purification rate (K9), are ultimately selected. Based on these selected features, a classification model is constructed using supervised machine learning algorithms such as Random Forest (RF) to achieve intelligent judgment of market order as "good," "medium," or "poor." This invention overcomes the curse of dimensionality and overfitting risks, providing an automated and scientific assessment tool for tobacco supervision, and can be extended to the field of public governance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of tobacco monopoly supervision and public governance technology, specifically involving a method for constructing an intelligent assessment model of county-level cigarette market order based on multi-source data fusion and feature optimization. Background Technology

[0002] The tobacco monopoly system serves as the cornerstone of the tobacco industry's governance system. County-level markets are the final stage for the implementation of monopoly laws, regulations, and policies. Accurate identification and scientific assessment of their order status have become core decision-making bases for monopoly management departments at all levels to achieve proactive risk warning, optimal allocation of regulatory resources, and differentiated and precise intervention. However, market order is a complex system, its state being coupled with multiple subsystems such as legal compliance, economic behavior patterns, regulatory intervention intensity, and social psychological perception, exhibiting significant implicitness, multidimensionality, and dynamic evolutionary characteristics. This complexity makes it difficult to directly, robustly, and comparablely measure it using a single or simply linearly summed indicator. How to overcome this measurement bottleneck and construct a scientific, objective, and efficient evaluation system is a core challenge in promoting the modernization of the tobacco monopoly management system and capabilities.

[0003] The academic and practical communities have been continuously exploring quantitative assessments of market order. Early research relied heavily on expert experience, employing subjective weighting methods such as the Analytic Hierarchy Process (AHP) and the Delphi method to construct comprehensive evaluation indices. With the deepening of management informatization, research has begun to incorporate objective business data such as the number of cases investigated, market inspection coverage, and cigarette sales volatility, attempting to monitor anomalies and assess market conditions by setting empirical thresholds. In recent years, the rise of big data and artificial intelligence technologies has driven the construction of more complex evaluation systems. For example, some local tobacco monopoly bureaus have attempted to construct comprehensive evaluation methods encompassing multiple dimensions, including market ecology, license management, anti-counterfeiting and anti-smuggling efforts, and internal supervision; or designed a three-tiered hierarchical evaluation system of "dimension-element-indicator" to more comprehensively depict market conditions. These efforts provide a valuable framework for understanding market order from multiple perspectives.

[0004] Despite this, existing research still faces several deep-seated problems that urgently need improvement. First, in constructing indicator systems, there is a general tendency to pursue a "large and comprehensive" approach, resulting in complex indicator dimensions, high data collection costs, and often severe multicollinearity among indicators. This not only creates a "data black box," weakening the model's interpretability, but may also reduce the model's generalization ability and robustness due to noise accumulation. Second, in terms of model evaluation methods, most studies remain at the level of deterministic index synthesis or traditional linear regression analysis, failing to fully meet the essential needs of market order assessment. Determining market order as "good," "medium," or "poor" is essentially a multi-class decision problem, with fuzzy and uncertain category boundaries. Traditional continuous value fitting models struggle to provide probabilistic decision support for risk classification and cannot effectively capture the complex nonlinear relationships that may exist between indicators and the state of order. Third, there is a lack of a systematic feature selection process driven by statistical inference. Simply relying on business experience or simple correlation coefficient ranking may lead to the omission of key predictive information or the introduction of irrelevant noise variables, thus affecting the performance and reliability of the final model. Summary of the Invention

[0005] The purpose of this invention is to overcome the shortcomings of existing technologies and provide a method for constructing an intelligent assessment model of county-level cigarette market order based on multi-source data fusion and feature optimization.

[0006] To achieve the above objectives, the technical solution of the present invention is: a method for constructing an intelligent assessment model of county-level cigarette market order based on multi-source data fusion and feature optimization, comprising:

[0007] Multi-dimensional data fusion: Integrate heterogeneous business data from multiple sources, including tobacco monopoly management, market supervision, case information, and administrative services, and construct an initial indicator system with five dimensions, including licensing management, market supervision, risk structure of business entities, abnormal circulation control of commodities, and administrative services and social co-governance. The initial indicator system includes 31 indicators.

[0008] Statistical feature optimization: Pearson correlation analysis was used to screen indicators that have a significant linear correlation with market order scores, and unsupervised learning weight analysis was combined to calculate the contribution of each feature to the variance of the clustering results. Key distinguishing factors were identified from the data structure, and finally, a subset of core features was selected.

[0009] Supervised machine learning modeling: Based on a subset of core features, a classification model is constructed using at least one of the following algorithms: Support Vector Machine (SVM), Random Forest (RF), Backpropagation Neural Network (BPNN), and Naive Bayes (NB) to classify the order of the county-level cigarette market into "good," "medium," and "poor" levels.

[0010] Model performance evaluation: The model performance is evaluated by confusion matrix and classification accuracy, and the model parameters are optimized to improve generalization ability.

[0011] Furthermore, the data sources in the multi-dimensional data fusion include the integrated platform of the tobacco monopoly bureau, the DingTalk three-line collaboration system, the real cigarette abnormal flow management and reporting system, retailer satisfaction survey data and marketing report system, forming a balanced panel dataset containing 7 county-level units and 12 monthly time points, with a total of 84 valid observation samples.

[0012] Furthermore, the unsupervised learning weight analysis includes constructing three models: K-means clustering, hierarchical clustering, and graph neural network (GNN) clustering. Based on the cluster labels, the F-statistic of each feature is calculated using one-way ANOVA. The magnitude of the F-value quantifies the significance of the difference between the corresponding features in different clusters, that is, it reflects the contribution or discriminative ability of the corresponding features to the formation of the unsupervised clustering structure.

[0013] Furthermore, the objective function of the K-means clustering model is to minimize the sum of squared deviations within clusters (WCSS), and its calculation formula is as follows:

[0014]

[0015] Where k is the preset number of clusters; It is the set of all points in the i-th cluster; x is the cluster to which x belongs. A data point; μ i It is the centroid of the i-th cluster; It is the sum of the "intra-cluster sums of squares" of all clusters;

[0016] The graph neural network (GNN) clustering model achieves clustering through node embedding learning, and its message passing and update formulas are as follows:

[0017]

[0018]

[0019] in, For node v, message aggregation at layer k. For the message function of the k-th layer, Let v be the embedding representation of node v at the k-th layer. Let u be the set of neighboring nodes of node v. As edge features, This is the update function for the k-th layer.

[0020] Furthermore, the final selected core feature subset for statistical feature optimization includes six indicators: licensed operation rate K3, market purification rate K9, proportion of Class B households K16, number of genuine cigarettes seized K19, complaints about license management K29, and reports of tobacco-related cases K30.

[0021] Furthermore, in supervised machine learning modeling, Support Vector Machine (SVM) uses Radial Basis Function (RBF) as the kernel function; Random Forest (RF) model constructs multiple decision trees through bootstrapping and random feature selection, and its predicted output adopts a majority voting mechanism; Backpropagation Neural Network (BPNN) constructs a multi-layer feedforward neural network with one hidden layer, the number of neurons in the hidden layer is adaptively determined according to the dimension of the input layer, and it is trained using the Levenberg-Marquardt algorithm; Naive Bayes (NB) classifier calculates the posterior probability based on Bayes' theorem.

[0022] Furthermore, the model performance is evaluated using a confusion matrix to calculate the overall classification accuracy, the formula of which is:

[0023] Overall classification accuracy = (Number of correctly classified samples / Total number of samples) × 100%

[0024] Furthermore, the proposed method ultimately recommends an intelligent evaluation model based on the Random Forest (RF) model and a subset of core features, which can be used for automated monthly evaluation and risk warning of the county-level cigarette market order.

[0025] Furthermore, the market order score is determined by an evaluation panel organized by the Tobacco Monopoly Administration, composed of senior experts in monopoly, marketing, and internal management. Based on a set of pre-set qualitative evaluation standards, the panel uses the Delphi method to conduct back-to-back independent scoring and calculates the arithmetic mean. The market order score is divided into three levels: "good" (3 points), "medium" (2 points), and "poor" (1 point).

[0026] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method described above.

[0027] The present invention also provides a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the method described above.

[0028] Compared with the prior art, the present invention has the following beneficial effects:

[0029] The intelligent assessment framework constructed in this invention has significant practical application value. It can be directly integrated into the tobacco monopoly management information system to achieve monthly automated assessment, visualization, and risk alerts of county-level market order, helping to shift regulatory resources from "average allocation" to "risk-based hierarchical allocation" and improving regulatory efficiency. Its methodological significance is even more profound, providing a referable technical path for all public management fields involving multi-source data fusion, complex system state diagnosis, and efficiency assessment (such as food safety supervision, tax risk identification, and ecological environment monitoring). Its innovation is mainly reflected in three aspects:

[0030] 1. Methodological Integration and Innovation: It successfully integrates and connects traditional statistical inference (correlation analysis, analysis of variance) with machine learning algorithms (random forest, etc.), forming a collaborative framework that "uses statistical screening to ensure the interpretability and robustness of the model, and uses machine learning to improve classification performance and nonlinear characterization capabilities," providing a new methodological paradigm for handling similar complex system evaluation problems.

[0031] 2. Problem-Oriented Feature Engineering Practice: Addressing the common pain points of indicator redundancy and collinearity in public governance evaluation, a replicable dynamic feature optimization process with statistical significance as its core and incorporating double testing was designed and validated. This process is universally applicable and can be transferred to other fields.

[0032] 3. Deepening and expanding application scenarios: Market order assessment is explicitly defined as a probabilistic multi-classification problem, and the superiority of ensemble learning models in this scenario is verified. This promotes the transformation of assessment work from static, deterministic "comprehensive scoring" to dynamic, probabilistic "risk warning and classification," which is more in line with the decision-making needs of precise regulation. Attached Figure Description

[0033] Figure 1 A heatmap of Pearson correlation analysis of multidimensional indicators and market order.

[0034] Figure 2 The accuracy of three unsupervised machine learning models in determining market order.

[0035] Figure 3 Weight analysis for three unsupervised machine learning models (first five).

[0036] Figure 4 The accuracy of four supervised machine learning models based on all indicators in determining market order.

[0037] Figure 5 This is the cumulative feature importance curve of RF based on the full feature set.

[0038] Figure 6The accuracy of four supervised machine learning models based on statistically optimized feature index sets in determining market order is given.

[0039] Figure 7 This is a block diagram illustrating the implementation of the method of the present invention. Detailed Implementation

[0040] The technical solution of the present invention will now be described in detail with reference to the accompanying drawings.

[0041] This invention provides a method for constructing an intelligent assessment model of county-level cigarette market order based on multi-source data fusion and feature optimization, comprising:

[0042] Multi-dimensional data fusion: Integrate heterogeneous business data from multiple sources, including tobacco monopoly management, market supervision, case information, and administrative services, and construct an initial indicator system with five dimensions, including licensing management, market supervision, risk structure of business entities, abnormal circulation control of commodities, and administrative services and social co-governance. The initial indicator system includes 31 indicators.

[0043] Statistical feature optimization: Pearson correlation analysis was used to screen indicators that have a significant linear correlation with market order scores, and unsupervised learning weight analysis was combined to calculate the contribution of each feature to the variance of the clustering results. Key distinguishing factors were identified from the data structure, and finally, a subset of core features was selected.

[0044] Supervised machine learning modeling: Based on a subset of core features, a classification model is constructed using at least one of the following algorithms: Support Vector Machine (SVM), Random Forest (RF), Backpropagation Neural Network (BPNN), and Naive Bayes (NB) to classify the order of the county-level cigarette market into "good," "medium," and "poor" levels.

[0045] Model performance evaluation: The model performance is evaluated by confusion matrix and classification accuracy, and the model parameters are optimized to improve generalization ability.

[0046] The following is a detailed implementation process of the present invention.

[0047] This invention discloses a method for constructing an intelligent assessment model of county-level cigarette market order based on multi-source data fusion and feature optimization, which mainly includes the following:

[0048] 1. Data Sources and Indicator System Construction

[0049] The data for this invention comes from the F Province Tobacco Monopoly Bureau's integrated platform, DingTalk's three-line collaborative system, the real cigarette abnormality flow management and reporting system, retailer satisfaction survey data, and marketing report system. Seven counties (cities, districts) under the jurisdiction of L City were selected as research units, and monthly panel data for 12 months from January to December 2024 were collected to form a balanced panel dataset containing 7 cross-sectional individuals and 12 time points, totaling 84 valid observation samples.

[0050] Based on a systematic review of relevant literature and in-depth discussions with business experts, this invention constructs an initial indicator system for evaluating the county-level cigarette market order, comprising 5 primary dimensions and 31 secondary indicators (denoted as K1-K31), as shown in Table 1. This system aims to comprehensively depict the market ecosystem:

[0051] (1) Licensing Management Dimension: This dimension includes five indicators: license issuance rate (K1), license floor area ratio (K2), license-holding operation rate (K3), license-displaying operation rate (K4), and license conformity rate (K5). This dimension reflects the legality of market access and the standardization of basic management, and is the logical starting point and legal foundation of market supervision.

[0052] (2) Market Supervision Dimension: This dimension includes nine indicators: abnormal information feedback and handling rate (K6), inspection coverage rate (K7), overall market modeling hit rate (K8), market purification rate (K9), monthly random inspection completion rate (K10), number of cases with a value of over 50,000 yuan (K11), total number of cases (K12), number of tobacco-related criminal cases (K13), and number of shut-down operations (K14). This dimension directly measures the daily supervision intensity, law enforcement efficiency, and crackdown results of the monopoly management department, and is a core means of maintaining market order.

[0053] (3) Risk structure dimension of business entities: It includes four indicators: the proportion of Class A households (K15), the proportion of Class B households (K16), the proportion of Class C households (K17), and the proportion of Class D households (K18). This dimension is based on the historical compliance records and risk ratings of retail businesses, reflecting the risk distribution and structural characteristics within the market, and is the direct basis for implementing classified and graded, differentiated and precise supervision.

[0054] (4) Dimension of Abnormal Commodity Circulation Control: This dimension includes 10 indicators: the number of genuine cigarettes seized (K19), the number of counterfeit cigarettes seized (K20), the number of smuggled cigarettes seized (K21), cases of more than 5 cigarettes flowing out of the province but outside the city (K22), cases of more than 5 cigarettes flowing out of the province but outside the city (K23), the total number of cigarettes flowing out of the province but outside the city (K24), the total number of cigarettes flowing out of the province but outside the city (K25), the number of households with more than 5 cartons flowing out at one time (K26), the number of households with more than 30 cartons flowing out within one year (K27), and the cigarette sales-to-volume ratio (K28). This dimension assesses the standardized circulation of cigarettes within the jurisdiction, focusing on the persistent problem of "abnormal circulation of legal cigarettes" in monopoly management, and is key to maintaining the price system and business order.

[0055] (5) Administrative Services and Social Co-governance Dimension: This dimension includes three indicators: satisfaction rate with license management (K29), reports of tobacco-related cases (K30), and per capita cigarette consumption (K31). This dimension reflects the satisfaction of retailers and the general public with administrative licensing and regulatory services, as well as their enthusiasm for participating in market co-governance. It reveals the deep-seated health status of market order from the perspectives of subjective perception and social feedback.

[0056] Table 1. Evaluation Index System for County-Level Cigarette Market Order

[0057]

[0058] The dependent variable (model prediction target) of this invention is the "monthly comprehensive score of the county-level cigarette market order" (Group). This score is determined by a review panel organized by the L City Tobacco Monopoly Bureau, composed of senior experts in monopoly, marketing, and internal management. The panel uses a pre-defined set of qualitative evaluation criteria (Table 2) and employs the Delphi method to conduct independent back-to-back scoring, calculating the arithmetic mean. The score categorizes market order into three levels: "Good" (assigned a value of 3), "Medium" (assigned a value of 2), and "Poor" (assigned a value of 1). This expert score serves as a training and validation label for the supervised machine learning model, providing the model with prior experience based on rich expert judgments.

[0059] Table 2 Comprehensive Scoring Standards for County-Level Cigarette Market Order

[0060]

[0061] 2. Statistical Feature Optimization Strategy

[0062] To select core influencing factors from 31 initial indicators, avoid model overfitting, and improve interpretability, this invention employs a composite statistical strategy for feature optimization:

[0063] 2.1 Pearson Correlation Analysis

[0064] The Pearson correlation coefficients (r) and their significance levels (p-values) between all 31 independent variables and market order expert ratings (Group) were calculated. This analysis aimed to preliminarily identify indicators with significant linear statistical associations to the order ratings. A significance level of α = 0.05 was set, and indicators satisfying p < 0.05 were retained as preferred features to eliminate statistically irrelevant noise variables.

[0065] 2.2 Weights in Unsupervised Learning Classification Models

[0066] To objectively assess the discriminative importance of each feature based on its inherent data structure, this invention introduces unsupervised clustering analysis as an auxiliary verification method for feature weights. Three unsupervised clustering models were constructed to perform clustering analysis on 84 samples (the number of clusters was set to 3, referencing the number of categories in the expert scoring):

[0067] (1) K-means clustering: The samples are divided by minimizing the sum of squared deviations within the cluster (WCSS). The optimal number of clusters is determined by the silhouette coefficient, and a stable result is selected by running the algorithm multiple times.

[0068]

[0069] Where k is the preset number of clusters; It is the set of all points in the i-th cluster; x is the cluster to which x belongs. A data point (vector); μ i J is the centroid (center point) of the i-th cluster; J is the sum of the "intra-cluster sums of squares" of all clusters.

[0070] (2) Hierarchical clustering: Ward's minimum variance method is used for agglomerative hierarchical clustering. A dendrogram is constructed based on the distance matrix between samples and the categories are cut.

[0071] (3) Graph Neural Network Clustering (GNN Clustering): Each observed sample is regarded as a node in a graph, and a nearest neighbor graph is constructed based on feature similarity. The low-dimensional embedding representation of the node is learned through the graph neural network, and then K-means is applied in the embedding space for clustering.

[0072]

[0073]

[0074] After obtaining the unsupervised clustering results (i.e., each sample is assigned a cluster label), this invention treats each original feature variable (K1-K31) as the dependent variable and the cluster label as the categorical independent variable, and performs a one-way ANOVA. The F-statistic for each feature is calculated. The magnitude of the F-value quantifies the significance of the feature's differences between different clusters, reflecting the feature's contribution to the formation of the unsupervised clustering structure or its discriminative power. A larger F-value indicates that the feature is more important for distinguishing different data clusters (which may correspond to different potential patterns of market order).

[0075] 2.3 Comprehensive Selection of Key Features

[0076] By combining the results of Pearson correlation analysis (significance of indicators) and unsupervised learning weight analysis (ranking of indicators by ANOVAF value), indicators that demonstrate high importance in both aspects are selected to constitute the "statistically optimized feature indicator set" used in subsequent supervised machine learning modeling in this invention. This process ensures that the selected features are not only significantly linearly correlated with the target variable, but also possess good class discrimination ability in the data space.

[0077] 3. Construction of supervised learning classification models

[0078] To comprehensively evaluate and compare the performance of different supervised learning classification models in the classification and evaluation of the cigarette market order, this invention constructs four classic supervised machine learning classification models based on the aforementioned feature sets (including the "full feature set" and the "statistically optimized feature index set"):

[0079] (1) Support Vector Machine (SVM): Its core principle is to find an optimal hyperplane that maximizes the margin between samples of different classes in the feature space. For the nonlinear classification problem of this invention, the radial basis function (RBF) is used as the kernel function. The penalty coefficient C and the kernel function scale parameter are optimized by grid search and 5-fold cross-validation.

[0080]

[0081] In the formula, For mapping functions, Here, represents the normal vector and intercept of the hyperplane, respectively. Since mapping functions have complex forms, their inner products are difficult to compute. Therefore, the kernel method can be used, defining the inner product of mapping functions as a kernel function: This avoids the explicit calculation of the inner product. Represents the feature vector of the sample. This represents the similarity between two samples.

[0082] (2) Random Forest (RF): As an ensemble learning algorithm, it makes predictions by constructing a large number of decision trees and combining their outputs (majority votes). Each tree uses bootstrap during training and randomly selects some features to find the optimal split point when splitting nodes, thereby enhancing the generalization ability and robustness of the model. In this invention, the number of decision trees is set to 100.

[0083]

[0084] In the formula, This represents the prediction results from the random forest. (x) represents the predictions of m decision trees for sample X; M is the total number of decision trees; mode is the mode.

[0085] (3) Back Propagation Neural Network (BPNN): Construct a multi-layer feedforward neural network with one hidden layer. The number of neurons in the hidden layer is adaptively determined according to the dimension of the input layer. The Levenberg-Marquardt algorithm is used for training to accelerate convergence. The convergence error is set to 0.0001, and the maximum number of training iterations is 1000.

[0086] (4) Naive Bayes Classifier (Back Propagation Neural Network, NB): Based on Bayes' theorem, under the assumption of conditional independence of features, it calculates the posterior probability of a sample belonging to each category and assigns the sample to the category with the highest posterior probability. It is assumed that continuous features follow a Gaussian distribution.

[0087]

[0088] In the formula, This is the posterior probability, which is the probability that a sample belongs to a class similar to C after observing feature X. Likelihood is the probability that a feature will appear when a sample belongs to class C. This is the prior probability, i.e., the probability distribution of category C in the population; Evidence factor, which is the probability of feature X appearing in all categories.

[0089] All models were implemented using the Statistics and Machine Learning Toolbox and Neural Network Toolbox in MATLAB software.

[0090] 3. Model Performance Evaluation

[0091] To comprehensively evaluate the performance of the classification models, this invention uses a confusion matrix for evaluation. The classification accuracy of each model is calculated, which is the proportion of all samples that are correctly classified.

[0092] To comprehensively evaluate the performance of the classification models, this invention uses the confusion matrix and its derived overall classification accuracy as the core evaluation metrics. The confusion matrix is ​​a C×C square matrix (where C is the number of classes), where the columns represent the true classes and the rows represent the predicted classes. It visually displays the model's prediction performance on each true class, including the number of correctly classified and misclassified samples. The overall classification accuracy is calculated as the proportion of all samples correctly classified. All models are trained and evaluated on the same dataset containing 84 samples.

[0093] 4. Results and Analysis

[0094] 4.1 Feature Optimization Results Based on Correlation Analysis and Unsupervised Learning Weight Analysis

[0095] Pearson correlation analysis results show that ( Figure 1 Of the 31 initial indicators, 6 showed a statistically significant correlation with the market order expert score (Group) at the 0.05 significance level (Table 3). Specifically:

[0096] The rate of licensed operation (K3) was significantly positively correlated with market order (r=0.240, P=0.028). This indicates that the higher the coverage of legal business entities, the better the baseline level of market order.

[0097] The market purification rate (K9) showed a significant positive correlation with market order (r=0.228, P=0.037). This directly confirms the close link between effective regulatory actions and the outcome of a sound market order.

[0098] Reports of tobacco-related cases (K30) showed a significant positive correlation with market order (r=0.236, P=0.031). This result may reflect a positive correlation between the activity of social supervision and the health of market order; that is, the better the market order, the stronger the public's awareness of rights protection and supervision, or the more accessible the reporting channels.

[0099] The proportion of Category B households (K16) showed a highly significant negative correlation with market order (r=-0.450, P=0.000). As a higher-risk retail group, the increase in its proportion strongly indicates that the overall market order is facing downward pressure.

[0100] The quantity of genuine cigarettes seized (K19) showed a significant negative correlation with market order (r=-0.232, P=0.034). This indicator directly measures the scale of "abnormal genuine cigarette circulation," and the larger the value, the more prominent the problem of disorder in the internal channels of the market and the worse the order.

[0101] Complaints regarding licensing management (K29) showed a highly significant negative correlation with market order (r=-0.316, P=0.003). This suggests that disputes or dissatisfaction during the administrative licensing process are a significant negative factor affecting market stability and retailers' trust in regulatory agencies.

[0102] These six indicators passed the statistical significance test and constituted the core characteristic indicators of the initial selection. They respectively represent key dimensions of tobacco monopoly and market order, including license management, regulatory results, risk structure, circulation control and social feedback.

[0103] Table 3. Indicators significantly correlated with market order (P<0.05)

[0104]

[0105] Note: * after the numbers in the table indicates a significant correlation at the 0.05 level; ** indicates a highly significant correlation at the 0.01 level.

[0106] Unsupervised learning weight analysis provides another perspective on feature importance. One-way ANOVA was performed on the results of three models: K-means, hierarchical clustering, and GNN clustering, to calculate the F-value (contribution) of each feature to the cluster label. Figure 2 , Figure 3 It can be seen that although the clustering accuracy of the three unsupervised clustering models themselves does not exceed 70% (the main problem lies in their weak ability to distinguish samples from the "poor" market, resulting in many misclassifications to the "medium" market), they show a high degree of consistency in their ranking of feature importance. Among the top five important features, market purification rate (K9), the proportion of B-class households (K16), the number of genuine cigarettes seized (K19), and complaints about license management (K29) appear frequently. This highly overlaps with the six indicators previously selected by Pearson correlation analysis, especially K16, K19, K29, and K9, which were identified as key factors under two completely different analytical methods, verifying their status as core indicators influencing market order.

[0107] Based on the significance results of Pearson correlation analysis and the high F-score ranking of unsupervised learning weight analysis, this invention ultimately determines six indicators as the "statistically optimized feature indicator set": licensed operation rate (K3), market purification rate (K9), proportion of Class B operators (K16), number of genuine cigarettes seized (K19), complaints about license management (K29), and reports of tobacco-related cases (K30). These six indicators will be used for the subsequent construction and comparison of supervised machine learning models. This subset greatly compresses the feature dimensions while ensuring statistical significance.

[0108] 4.2 Performance Comparison of Classification Models Based on the Full Feature Set

[0109] To compare model performance, this invention uses all 31 indicators as input features to construct four classification models: SVM, RF, BPNN, and NB.

[0110] Performance comparison Figure 4 As shown, it can be seen from this:

[0111] The RF model performed optimally overall, achieving a total classification accuracy of 95.24%. Its confusion matrix shows that the model achieved completely accurate identification of samples in the "Good" (Category 1) and "Medium" (Category 2) categories (10 and 51 samples correctly, respectively), with only 4 misclassified as "Medium" out of 23 samples in the "Poor" (Category 3) category. This demonstrates that RF can extremely effectively capture the complex relationship between high-dimensional features and ordered categories.

[0112] The SVM model performed second best, with an overall accuracy of 79.76%. Its misclassifications mainly occurred in the "good" and "poor" classes. Some "good" samples were misclassified as "medium," and some "poor" samples were misclassified as "medium," indicating that the model's preference for intermediate states or its insufficient distinction between extreme states was not clear enough.

[0113] The BPNN model achieved an accuracy of 70.24%. Its misclassifications were relatively scattered, with misclassifications occurring between the three categories, possibly due to the network structure or training parameters not yet being optimal.

[0114] The NB model performed the worst, with an accuracy of only 46.43%. This model was almost unable to effectively distinguish between the "medium" and "poor" classes, misclassifying a large number of "medium" samples as "poor". This indicates that in high-dimensional scenarios where features may be correlated, its strong assumption of "conditional independence" is seriously out of touch with reality, leading to a significant drop in performance.

[0115] The near-perfect performance (95.24%) of the RF model on the full feature set may imply the risk of overfitting, meaning the model is too complex and overlearns noise and specific patterns in the training data, casting doubt on its generalization ability on unseen data. Furthermore, feature importance analysis of the RF model reveals that its cumulative importance curve shows (…). Figure 5 The cumulative contribution of the top 6 most important features exceeds 85%. This confirms the feasibility of feature dimensionality reduction from within the model itself, that is, only a few core features are sufficient to support the model in making most of the correct judgments, providing an intrinsic basis for adopting statistically optimized feature index sets.

[0116] 4.3 Performance Comparison of Classification Models Based on Statistically Optimized Feature Index Sets

[0117] To further verify the performance of the feature optimization method, this invention uses the six core indicators (K3, K9, K16, K19, K29, K30) selected as inputs to reconstruct and evaluate the same four classification models.

[0118] Performance comparison Figure 6 As shown, it can be seen from this:

[0119] The RF model remains in the lead, with an overall accuracy of 92.86%. Compared to 95.24% using the full feature set, the accuracy only decreased slightly by 2.38 percentage points. Its confusion matrix shows that the model accurately identified the "good" class, with only a very small number of misclassifications (approximately 3 cases each) for the "medium" and "poor" classes. This demonstrates that even using a carefully selected subset representing only 19.35% (6 / 31) of the original features, the RF model can still maintain extremely high classification performance.

[0120] The performance of the BPNN model was significantly improved, with accuracy increasing from 70.24% to 82.14%. Feature dimensionality reduction effectively reduced noise input, simplified the relationships that the network needs to learn, and may make the training process more stable, thereby improving model performance.

[0121] The accuracy of the SVM model decreased from 79.76% to 69.05%. This may be because SVM is more sensitive to the representation of the feature space, and a significant reduction in features may have resulted in the loss of some information useful for mapping the SVM kernel function.

[0122] The accuracy of the NB model improved slightly, from 46.43% to 53.57%, but its performance was still poor, which once again confirmed its limitations on this dataset.

[0123] In summary, the RF model based on the optimized feature index set, while reducing the feature dimension by 80.65%, achieves a significant reduction in model complexity, a substantial decrease in overfitting risk, and a greatly enhanced business interpretability with only a 2.38 percentage point loss in accuracy. From a cost-benefit perspective, data collection and processing costs are reduced by over 80%, while model prediction efficiency (measured by the reduction in feature dimension and overfitting risk) is fundamentally improved. This strongly demonstrates the effectiveness and necessity of the statistical feature optimization strategy proposed in this invention. Therefore, the RF model constructed based on the statistically optimized feature index set (K3, K9, K16, K19, K29, K30), due to its optimal balance in high accuracy (92.86%), high robustness (avoiding overfitting), strong interpretability (six indicators with clear business meaning), and high efficiency (low data requirements), is established as the optimal intelligent assessment model for the county-level cigarette market order recommended by this invention.

[0124] 5. Conclusion

[0125] This invention addresses the practical bottlenecks in assessing the order of the county-level cigarette market under the monopoly system, such as redundant indicators, strong subjectivity, and weak model interpretability. It innovatively proposes an integrated analysis method of "multi-source data fusion—statistical feature optimization—machine learning modeling." Through empirical analysis of multi-source business data from L City, F Province, the system systematically implemented the entire process from initial indicator construction and dual feature optimization based on statistical inference to the construction and performance comparison of multiple supervised learning models. The study successfully screened six core influencing factors from 31 candidate indicators and established a high-performance classification and evaluation model based on the random forest algorithm and constructed using optimized features.

[0126] 5.1 Business interpretability of core feature factors and their correlation mechanism

[0127] The six indicators ultimately selected are not statistical coincidences; they constitute a logically rigorous and hierarchically structured market order observation system that deeply reflects the core logic of monopoly management.

[0128] The licensing rate (K3) is the "legitimacy foundation" of market order. It measures the basic compliance level of market operators. A high licensing rate means clear regulatory targets and broad legal coverage, which is a prerequisite for any effective management. Its significant positive correlation aligns with the institutional economics principle that "clear property rights (operating rights) are the foundation for the effective operation of the market."

[0129] The market purification rate (K9) is a "direct result" of regulatory effectiveness. This indicator integrates compliance findings from routine inspections and is an immediate reflection of how regulatory actions translate into a compliant market state. Its positive correlation directly verifies the causal relationship that "effective regulation leads to good order."

[0130] The proportion of Category B households (K16) serves as a "structural early warning system" for market risks. As a group subject to close monitoring due to past violations, its proportion directly quantifies the concentration of "existing risks" in the market. Its extremely strong negative impact (r=-0.450) confirms the business logic of "controlling key households is controlling the entire market," and changes in this indicator are often a leading signal of a systemic deterioration in market order.

[0131] The quantity of genuine cigarettes seized (K19) serves as a "barometer" of the disorder in the distribution order. It directly addresses the core pain point of the monopoly system—"abnormal circulation of legitimate products"—reflecting deep-seated market imbalances caused by price inversions, regional price differences, and arbitrage. Its negative correlation indicates that the more rampant the abnormal circulation, the more severely the normal business order and price system are disrupted.

[0132] License management complaints (K29) serve as a "social sensor" for administrative compliance. The number of complaints not only reflects the standardization, transparency, and fairness of the administrative licensing process but also reflects retailers' trust in regulatory agencies. Excessive complaints erode regulatory authority and, from a socio-psychological perspective, undermine the foundation of market stability; this negative correlation reveals the importance of "regulation itself being regulated."

[0133] The reporting of tobacco-related cases (K30) serves as a "trust index" for social co-governance. Its positive correlation may have a dual meaning: on the one hand, in a well-ordered market, consumers and the public are likely to have a higher awareness of their rights and a greater willingness to report violations; on the other hand, it also reflects public trust in regulatory agencies' ability to effectively handle reports. It is an important window for understanding hidden illegal activities in the market and assessing the credibility of regulators.

[0134] These six indicators form a complete closed loop from basic access (K3) → risk structure (K16) → circulation status (K19) → regulatory results (K9) → administrative feedback (K29) and social feedback (K30), realizing a comprehensive and three-dimensional portrayal of market order from static structure to dynamic behavior, and from objective results to subjective perception.

[0135] 5.2 The relative advantages of supervised learning models in governance effectiveness evaluation

[0136] The results of this invention demonstrate that supervised learning models (especially ensemble learning models such as Random Forest) exhibit significant advantages in classification and evaluation tasks. Compared to unsupervised clustering methods (this invention also attempted K-means, but its accuracy was less than 70%, and the results needed to be assigned business meaning afterward), the greatest advantage of supervised learning lies in its ability to directly utilize the professional domain knowledge of "expert ratings" as prior experience for training. This allows the model's learning objectives to proactively align with the comprehensive judgment standards of human experts, and the evaluation results naturally possess business interpretability, directly serving the "good, average, poor" management classification decision. Compared to traditional linear regression or comprehensive index methods, machine learning models represented by Random Forest can automatically capture the complex nonlinear relationships and interaction effects between multiple core factors. For example, when "high proportion of Category B households" and "large quantity of genuine cigarettes seized" occur simultaneously, their negative impact on market order may not be a simple addition, but rather a multiplication. This complex pattern can be effectively learned by Random Forest through different combinations of splits of multiple decision trees.

[0137] 5.3 The Necessity and Deep Mathematical Logic of Statistically Driven Feature Optimization

[0138] Directly inputting 31 high-dimensional indicators into a machine learning model without screening can easily lead to the "curse of dimensionality" and overfitting. The model may learn a large amount of noise or spurious statistical associations specific to the samples ("machine illusion"), resulting in excellent performance on the training set (e.g., the RF model with a full feature set achieves a judgment accuracy of 95.24%), but poor generalization ability on new data. This invention uses a dual statistical checkpoint of Pearson correlation analysis (linear screening) and unsupervised learning weight analysis (structural screening) to force modeling to be based on a solid foundation of statistically significant associations and data structure discriminativeness: mathematically, Pearson correlation screening ensures that there is a significant linear co-variance trend between the selected features and the target variable, laying the statistical foundation for prediction; unsupervised weight verification starts from the essence of data distribution, using the F-value of variance analysis to identify those features that can maximize the differentiation of class clusters. This is similar to performing an "unsupervised pre-classification" before modeling, screening out the features that contribute the most to class separation.

[0139] While this feature optimization strategy may lose some weak or complex nonlinear signals (which may explain the slight decrease in the accuracy of the RF model for optimizing the feature set), it greatly enhances the model's robustness, interpretability, and generalization potential. Ultimately, each factor in the model withstands both statistical testing and rigorous business logic scrutiny, transforming it from a "black box" into a highly interpretable, data-driven "white box" system for decision analysis.

[0140] To achieve a scientific, intelligent, and efficient assessment of the tobacco monopoly market order in counties, this invention takes L City in F Province as an example, integrates multi-source business data, and successfully constructs and verifies an integrated analysis method system of "multi-dimensional feature fusion—statistical feature optimization—machine learning modeling". Figure 7 ).

[0141] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method described above.

[0142] The present invention also provides a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the method described above.

[0143] The above are preferred embodiments of the present invention. Any changes made to the technical solution of the present invention that do not exceed the scope of the technical solution of the present invention shall fall within the protection scope of the present invention.

Claims

1. A method for constructing an intelligent assessment model of county-level cigarette market order based on multi-source data fusion and feature optimization, characterized in that, include: Multi-dimensional data fusion: Integrate heterogeneous business data from multiple sources, including tobacco monopoly management, market supervision, case information, and administrative services, and construct an initial indicator system with five dimensions, including licensing management, market supervision, risk structure of business entities, abnormal circulation control of commodities, and administrative services and social co-governance. The initial indicator system includes 31 indicators. Statistical feature optimization: Pearson correlation analysis was used to screen indicators that have a significant linear correlation with market order scores, and unsupervised learning weight analysis was combined to calculate the contribution of each feature to the variance of the clustering results. Key distinguishing factors were identified from the data structure, and finally, a subset of core features was selected. Supervised machine learning modeling: Based on a subset of core features, a classification model is constructed using at least one of the following algorithms: Support Vector Machine (SVM), Random Forest (RF), Backpropagation Neural Network (BPNN), and Naive Bayes (NB) to classify the order of the county-level cigarette market into "good," "medium," and "poor" levels. Model performance evaluation: The model performance is evaluated by confusion matrix and classification accuracy, and the model parameters are optimized to improve generalization ability.

2. The method for constructing an intelligent assessment model of county-level cigarette market order based on multi-source data fusion and feature optimization as described in claim 1, characterized in that, The data sources in the multidimensional data fusion include the integrated platform of the tobacco monopoly bureau, the DingTalk three-line collaboration system, the real cigarette abnormal flow management and reporting system, retailer satisfaction survey data and marketing report system, forming a balanced panel dataset containing 7 county units and 12 monthly time points, with a total of 84 valid observation samples.

3. The method for constructing an intelligent assessment model of county-level cigarette market order based on multi-source data fusion and feature optimization as described in claim 1, characterized in that, Unsupervised learning weight analysis includes constructing three models: K-means clustering, hierarchical clustering, and graph neural network (GNN) clustering. Based on the cluster labels, the F-statistic of each feature is calculated using one-way ANOVA. The magnitude of the F-value quantifies the significance of the difference between the corresponding features in different clusters, that is, it reflects the contribution or discriminative ability of the corresponding features to the formation of the unsupervised clustering structure.

4. The method for constructing an intelligent assessment model of county-level cigarette market order based on multi-source data fusion and feature optimization as described in claim 3, characterized in that, The objective function of the K-means clustering model is to minimize the sum of squared deviations within clusters (WCSS), and its calculation formula is as follows: Where k is the preset number of clusters; It is the set of all points in the i-th cluster; x is the cluster to which x belongs. A data point; μ i It is the centroid of the i-th cluster; It is the sum of the "intra-cluster sums of squares" of all clusters; The graph neural network (GNN) clustering model achieves clustering through node embedding learning, and its message passing and update formulas are as follows: in, For node v, message aggregation at layer k. For the message function of the k-th layer, Let v be the embedding representation of node v at the k-th layer. Let u be the set of neighboring nodes of node v. As edge features, This is the update function for the k-th layer.

5. The method for constructing an intelligent assessment model of county-level cigarette market order based on multi-source data fusion and feature optimization according to claim 1, characterized in that, The final selected core feature subset includes six indicators: licensed operation rate (K3), market purification rate (K9), proportion of Class B businesses (K16), number of genuine cigarettes seized (K19), complaints about license management (K29), and reports of tobacco-related cases (K30).

6. The method for constructing an intelligent assessment model of county-level cigarette market order based on multi-source data fusion and feature optimization according to claim 1, characterized in that, In supervised machine learning modeling, Support Vector Machine (SVM) uses Radial Basis Function (RBF) as its kernel function; Random Forest (RF) constructs multiple decision trees through bootstrapping and random feature selection, and its predicted output uses a majority voting mechanism; Backpropagation Neural Network (BPNN) constructs a multi-layer feedforward neural network with one hidden layer, the number of neurons in the hidden layer is adaptively determined according to the dimension of the input layer, and it is trained using the Levenberg-Marquardt algorithm; Naive Bayes (NB) classifier calculates the posterior probability based on Bayes' theorem.

7. The method for constructing an intelligent assessment model of county-level cigarette market order based on multi-source data fusion and feature optimization according to claim 1, characterized in that, Model performance is evaluated using a confusion matrix to calculate the overall classification accuracy, the formula of which is: Overall classification accuracy = (Number of correctly classified samples / Total number of samples) × 100% Furthermore, the proposed method ultimately recommends an intelligent evaluation model based on the Random Forest (RF) model and a subset of core features, which can be used for automated monthly evaluation and risk warning of the county-level cigarette market order.

8. The method for constructing an intelligent assessment model of county-level cigarette market order based on multi-source data fusion and feature optimization according to claim 1, characterized in that, The market order score is determined by an evaluation panel organized by the Tobacco Monopoly Administration, composed of senior experts in monopoly, marketing, and internal management. Based on a set of pre-set qualitative evaluation standards, the panel uses the Delphi method to conduct back-to-back independent scoring and calculates the arithmetic mean. The market order score is divided into three levels: "good" (3 points), "medium" (2 points), and "poor" (1 point).

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method as described in any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1 to 8.