Single coal multimodal transport explainable customer portrait construction method

By screening causal characteristic variables through causal discovery and clustering algorithms, and combining the LightGBM model and SHAP analysis, an interpretable customer portrait for coal multimodal transport was constructed, which solved the problem of inaccurate customer demand identification in existing technologies and achieved a clear explanation of customer decision-making logic and scientific service strategy.

CN120561700BActive Publication Date: 2025-10-17JILIN UNIVERSITY
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511046194.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-10-17
Estimated Expiration
2045-07-29

AI Technical Summary

Technical Problem

Existing research finds it difficult to construct a highly interpretable customer profile in coal intermodal transport, resulting in inaccurate identification of customer needs and affecting the scientific nature and interpretability of service strategies.

Method used

The causal discovery method is used to screen causal characteristic variables, combined with k-prototype hybrid clustering and LightGBM model, and SHAP analysis is used to improve the model interpretability and construct an interpretable customer portrait for coal multimodal transport.

Benefits of technology

By identifying key variables through causal discovery and clustering algorithms, the interpretability and accuracy of customer portraits are improved, supporting the precise classification of coal intermodal transport customers and the optimization of service strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120561700B_ABST
    Figure CN120561700B_ABST
Patent Text Reader

Abstract

The present application belongs to the field of intelligent transportation system planning and industrial artificial intelligence, and specifically relates to a single coal multimodal transport explainable customer portrait construction method, comprising: A, obtaining customer coal multimodal transport mode selection questionnaire and historical single coal multimodal transport waybill data, and performing data processing and data fusion; B, using the FCI algorithm for causal discovery, constructing a causal diagram, screening out characteristic variables having a causal relationship with customer multimodal transport mode selection, and simultaneously eliminating pseudo-correlated features; C, using the k-prototype hybrid clustering method to obtain customer categories; D, training a LightGBM classification model to predict the category to which the customer belongs, testing the portrait accuracy, and combining the SHAP explainable analysis method to describe the customer portrait, so as to provide decision support with scientificity and business landing for coal multimodal transport customer accurate classification and single service strategy optimization.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of intelligent transportation system planning and industrial artificial intelligence, and particularly relates to a single-order coal multimodal transport explainable customer portrait construction method, which is particularly suitable for customer classification in the single-order service of the coal multimodal transport scenario. BACKGROUND

[0002] Coal transportation is transforming from a single transportation mode to multimodal transportation modes such as iron-water combined transportation and highway-railway combined transportation. The single-order transportation mode realizes one-time commissioning, one-time payment and one-order bottoming through the whole set of bills of lading. However, the customer demand is significantly different, and it has been unable to meet the requirement of accurately identifying the characteristics of coal customer demand to ensure the accuracy of customer service.

[0003] Customer portrait (user portrait) is a process of describing or classifying customer individual characteristics. The application of customer portrait involves multiple research directions, but the existing research lacks research on coal customer portrait under the single-order system. The customer demand of the single-order coal multimodal transport, as an efficient and environmentally friendly transportation mode, has unique complexity and diversity, and the existing research results are difficult to apply directly.

[0004] Existing research such as the one published by Wang Yuan et al. in Power Engineering and Technology in April 2022, “Electricity user behavior portrait based on information gain and Spearman correlation coefficient”, proposes an electricity user behavior portrait method based on information gain and Spearman correlation coefficient. This method can efficiently cluster analyze the electricity consumption data of electricity users. Considering the information gain of features and the redundancy between features, an adaptive evaluation coefficient of feature set is constructed to select the optimal feature set. Through GA, the optimal feature set is efficiently and quickly solved, and the electricity consumption characteristics of electricity users are described based on the optimal feature set. For example, Ma Xin et al. in Operations Research and Management in January 2021, “Airline customer grouping research based on improved contour coefficient method”, calculate the Pearson correlation coefficient between attributes to draw a heat map to determine the linear correlation between attributes. At the same time, with the help of random forest, the original data is analyzed, and the weight of each attribute is calculated as a supplement to the selected variables of the heat map. Feature selection is achieved through the above methods.

[0005] The above-mentioned customer portrait construction methods are mostly based on the statistical correlation between variables to construct a feature system, but the correlation only reflects the surface covariation between variables, and cannot distinguish between driving factors and confounding interference in causal mechanisms. For example, the transportation distance and the selection of multimodal transportation mode may produce a pseudo-correlation through transportation cost, rather than a direct causal effect. This limitation makes the portrait model susceptible to confounding variables, making it difficult to reveal the real causes of customer behavior decisions, and thus reducing the scientificity and explainability of strategy formulation.

[0006] For example, existing research, Chinese patent CN118212034A discloses a kind of personalized recommendation method based on e-commerce data user portrait, proposes a kind of personalized recommendation method based on e-commerce data user portrait, in this method, according to the features of the comprehensive user portrait constructed, XGBoost method is used for classification prediction, to predict the items that users may be interested in.Weihuaxia et al. published in Information Science 2024 03, a method for dynamic portrait of small and micro enterprises based on BO-XGBoost comprehensive quality, proposed a portrait label classification algorithm based on XGBoost model of Bayesian optimization, realized the automatic extraction and dynamic update of small and micro enterprise comprehensive quality portrait label, the label classification accuracy of the comprehensive quality dynamic portrait method based on BO-XGBoost reaches 95.71%, and the overall performance of the model is good.

[0007] The model used in the above research has high prediction accuracy, but ignores the key element of model interpretability.In practical application scenarios, especially for coal multimodal transport industry, a black box model with high accuracy alone cannot help enterprises understand the logic behind customer behavior.Therefore, it is crucial to introduce tools that can improve model interpretability. SUMMARY

[0008] In view of the above problems, the purpose of the present application is to provide a single coal multimodal transport interpretable customer portrait construction method, which uses causal discovery method to identify key variables affecting customer multimodal transport mode selection; further apply clustering algorithm to group customers, thereby constructing preliminary customer portrait; then, test the accuracy of the portrait by training LightGBM model to ensure the accuracy of the portrait; finally, combine SHAP theory to analyze the constructed customer portrait in depth, clearly explain the influence mechanism of each feature on customer classification, and improve the interpretability of coal customer portrait.

[0009] The single coal multimodal transport interpretable customer portrait construction method provided by the present application comprises the following steps:

[0010] A, obtain customer coal multimodal transport mode selection questionnaire and historical single coal multimodal transport waybill data, and perform data processing and data fusion;

[0011] B, use FCI algorithm for causal discovery to construct causal diagram, select feature variables with causal relationship with customer multimodal transport mode selection, and eliminate pseudo-correlated features;

[0012] C, use k-prototype hybrid clustering method to cluster customers to obtain customer categories;

[0013] D, train LightGBM classification model to predict the category to which the customer belongs, test the accuracy of the portrait, and describe the customer portrait by combining SHAP interpretable analysis method.

[0014] As preferred in the present application, the following step is further included in step A:

[0015] A1, obtain historical coal single transport waybill data, design a customer coal intermodal transport mode selection questionnaire, distribute the questionnaire to coal customers, and collect the questionnaire answers filled out by the customers;

[0016] A2, process the waybill and questionnaire data, screen and delete duplicate records, supplement missing values, correct abnormal values, and obtain a structured waybill dataset after processing; for all customer coal waybill data, for the number of coal customers, for the set of all coal waybill data attributes, for the total number of attributes of each coal waybill data, for the th coal waybill data, the th data in the waybill data is represented as , each questionnaire answer is a sample, and the results in each sample are numbered and processed separately, and all the sample data after numbering are summarized to obtain summary data;

[0017] A3, data fusion, associate the questionnaire and waybill data through unique identifiers or fuzzy matching, wherein the questionnaire filling time is consistent with the waybill data recording period.

[0018] As preferred in the present application, the following step is further included in step B:

[0019] B1, data preparation, collect data related to coal customer intermodal transport mode selection from questionnaire results, the classification variables in the data include: coal characteristics composed of types and quantities, transportation factors composed of transportation distance, transportation time and transportation cost, and customer information composed of enterprise size and transportation frequency, and the classification variables in the data are encoded for FCI algorithm calculation.

[0020] As preferred in the present application, the following step is further included in step B:

[0021] B2, FCI algorithm execution;

[0022] B2.1, set the target variable as intermodal transport mode , is the number of intermodal transport modes, and the set of all coal waybill data attributes is the candidate feature set including all waybill data attributes in the questionnaire;

[0023] B2.2. Construct a complete graph. Consider all variables in the waybill data as nodes and construct a completely undirected graph. This means that all nodes are connected by edges, but the edges have no direction and only indicate the association between the variables.

[0024] B2.3, conditional independence test, partial correlation coefficient test for continuous variables, if the given variable Under the condition of and intermodal transport direct correlation, according to 、 and Pearson correlation coefficient between two pairs 、 、 calculate and In no Partial correlation coefficient under the influence of ;

[0025] ;

[0026] ;

[0027] If the partial correlation coefficient Not significant, indicating that the test variable and intermodal transport There is no direct causal relationship between the selection and the correlation between the given variables lead to;

[0028] For discrete variables, chi-square test is performed to test the variables and intermodal transport Whether it is independent, calculate the expected frequency ,in, Indicates the The number of samples of each coal type, Indicates the The number of samples of intermodal transport mode, the number of coal customers That is, it is expressed as the total sample size;

[0029] ;

[0030] Combined with The coal type is Number of samples of intermodal transport modes , calculate the chi-square statistic ;

[0031] ;

[0032] Calculating degrees of freedom , find the probability under the chi-square test , if the value is less than the significance level, reject the null hypothesis that the test variable is independent of the given variable ;

[0033] B2.4, edge deletion and retention, according to the result of conditional independence test, delete the given variable if the test variable is independent of the given variable ;

[0034] B2.5, determine partial directed edges, determine the direction of some edges by further conditional independence test, if it is found that there is an edge between the test variable and the given variable , and under the condition of a new variable , the test variable has an impact on the given variable , but the given variable has no impact on the test variable , the edge is directed from the test variable to the given variable , and a partial directed graph is obtained;

[0035] B2.6, determine the remaining undirected edges by pre-set rules, the pre-set rules include collision structure rules and direction propagation rules,

[0036] collision structure rules, if there are three variables , , constitute a path, denoted as , the symbol represents an undirected edge, and and are not adjacent, that is, there is no direct connection edge, when controlling other variables , and are independent, the control variable , and and are related, then is a collision node, the path is a collision structure, denoted as , where the symbols and represent directed edges;

[0037] direction propagation rules, if there is a directed path , and and​ Not adjacent, then the orientation is ;

[0038] B2.7, potential variable processing, for each potential variable, update the graph structure by adding a virtual edge to represent its influence on the observed variable, and distinguish real edges and virtual edges, wherein the real edge is the observed causal relationship, and the virtual edge is the potential variable influence;

[0039] B2.8, build a causal relationship graph, repeat steps B2.3 to B2.7, use conditional independence test and preset rules to determine the direction of more edges until no more edges can be determined, and finally obtain a directed acyclic graph.

[0040] As a preferred embodiment of the present application, the following steps are further included in step B:

[0041] B3, key variable determination and analysis;

[0042] B3.1, key variable identification, in the directed acyclic graph obtained in step B2, the variables connected to the intermodal mode selection node by direct or indirect directed edges are the key variables having causal relationship with the intermodal mode selection of the coal customer; wherein, according to the causal relationship graph, the characteristic variables having direct or indirect causal relationship with the intermodal mode selection node are retained, and the variables having no causal relationship with the intermodal mode selection node are removed, that is, the pseudo-correlation features are removed;

[0043] B3.2, causal relationship analysis, according to the direction and path of the directed edge, analyze the causal relationship between the key variable and the intermodal mode selection;

[0044] B3.3, result verification and interpretation, compare the results obtained by the FCI algorithm with the artificial verification to verify the rationality and interpretability of the results.

[0045] As a preferred embodiment of the present application, the following steps are further included in step C:

[0046] C1, according to the results of key variable screening in step B, extract the required data from the waybill data , input the mixed data set after causal relationship selection, and the mixed data set includes numerical features and category features;

[0047] C2, determine the optimal cluster number K by silhouette coefficient;

[0048] C2.1, traverse different cluster numbers , calculate the distance between any two customer samples, in the k-prototype clustering method, the Euclidean distance is used when the attribute is a numerical feature , the numerical features are mapped to the interval [0, 1] using the max-min normalization method, and then the Euclidean distance of the numerical features is calculated ;

[0049] ;

[0050] wherein, is the number of clusters , the normalized value of the th attribute of the th customer, is the number of clusters , the normalized value of the th attribute of the th customer;

[0051] When the attribute is a category feature, the Hamming distance is used The original code of the category feature is retained, and if the category attributes of two customers are different, the Hamming distance is 1, and if they are the same, it is 0;

[0052] ;

[0053] wherein, is the number of clusters , the mapped value of the th attribute of the th customer, is the number of clusters , the mapped value of the th attribute of the th customer, is the serial number of the total number of attributes;

[0054] The distance between two customer samples is , is the weight factor of the category attribute;

[0055] C2.2, traverse different cluster numbers values, and according to the distances between the obtained customer samples, respectively calculate the corresponding silhouette coefficients;

[0056] ;

[0057] ;

[0058] wherein, is the cohesion degree, representing the average distance of the customer point to other customer points in the belonging cluster, is the separation degree, representing the average distance to all customer points of the other cluster closest to the belonging cluster, is the customer point ​The silhouette coefficient, is the global silhouette coefficient;

[0059] C2.3. Comparison of different numbers of clusters Global silhouette coefficient at value , select the one with the largest global silhouette coefficient The value is taken as the optimal number of clusters K;

[0060] C3. Based on the determined K clusters of customer categories, extract the typical features of each cluster to form customer grouping labels. Extract the statistical distribution of each cluster's features to define customer labels that are understandable to the business.

[0061] As the preferred embodiment of the present invention, step D also includes the following steps:

[0062] D1. Set the target variable to the customer cluster label generated by the k-prototype clustering method for The index of the cluster, the input feature is the key variable obtained in step B, all the waybill data Divide the dataset into training and test datasets in a ratio of 7:3 and train the LightGBM model.

[0063] D2. Select weighted average To evaluate the accuracy of customer portraits and verify the accuracy of portraits;

[0064] ;

[0065] Where, For the The number of customer samples in a cluster, Expressed as the number of coal customers, i.e. the total sample size, For the Client's Fraction;

[0066] D3. Use the SHAP interpretable analysis method to interpret the constructed LightGBM model, combine the feature density bee swarm diagram and feature impact waterfall diagram to explain the classification model prediction results, and analyze the characteristics of different categories of customer portraits. The SHAP value calculation method is as follows

[0067] ;

[0068] Where, is the model benchmark value, which is the average prediction value of the LightGBM model on the training set. Characterized by For samples The SHAP value contribution, feature density bee colony chart and feature influence waterfall chart are visualization results of the SHAP explainable analysis method;

[0069] The SHAP value is used as feature importance to sort the features and identify the features that have the greatest impact on customer classification.

[0070] The beneficial effects of the present application are as follows:

[0071] 1. The present application provides a new solution for customer clustering by integrating causal discovery, hybrid clustering and explainable model technology. The causal feature screening process based on the FCI algorithm breaks through the limitations of traditional correlation analysis, identifies key variables that have a causal relationship with customer selection behavior by combining questionnaire survey results and actual shipment data, and improves the business interpretation of features; combining the LightGBM classification model and SHAP explainable analysis improves the explainability of the customer portrait. It provides decision support for coal multimodal transport customer precise classification and "one order" service strategy optimization with scientificity and business landing.

[0072] 2. The present application can accurately depict the reasons for customer behavior and improve the explainability of the portrait. The FCI algorithm in causal discovery is used to mine key variables that affect coal customer selection of multimodal transport from questionnaire and multimodal transport shipment multi-source data, draw a causal diagram, and identify the core variables that truly drive customer selection. The core variables mined through the causal diagram can accurately depict the driving logic of customer decision-making, making the portrait model mechanismally explainable.

[0073] 3. The present application applies LightGBM and SHAP to coal multimodal transport customer portrait construction. LightGBM, with its efficient hybrid feature processing capability, can efficiently integrate complex and diverse numerical features in the coal multimodal transport scenario, such as transportation volume and transportation cost, as well as categorical features such as customer type and coal type, and achieve high classification accuracy through fine parameter tuning to test the accuracy of the portrait. On this basis, SHAP explainability analysis is introduced to accurately identify the core factors affecting customer classification, and SHAP is used for visualization of the prediction model. BRIEF DESCRIPTION OF DRAWINGS

[0074] Other objects and results of the present application will become more apparent and easily understood with reference to the following description in conjunction with the accompanying drawings, and as a more complete understanding of the present application is acquired. In the drawings:

[0075] Fig. 1 is a technical roadmap of the present application;

[0076] Fig. 2 is a logic roadmap of the present application. DETAILED DESCRIPTION

[0077] BRIEF DESCRIPTION OF DRAWINGS Figs. 1-2 The application will be further described in detail below in conjunction with the accompanying drawings and specific examples.

[0078] The embodiment of the application provides a single-order coal multimodal transport explainable customer portrait construction method, which comprises the following steps:

[0079] A, obtaining a customer coal multimodal transport mode selection questionnaire and historical single-order coal multimodal transport waybill data, and performing data processing and data fusion;

[0080] A1, obtaining historical single-order coal transport waybill data, designing a customer coal multimodal transport mode selection questionnaire, distributing the questionnaire to coal customers, and collecting the questionnaire reply results filled by the customers;

[0081] A2, processing the waybill and questionnaire data, screening and deleting duplicate records, supplementing missing values, correcting abnormal values, obtaining a structured waybill data set after processing, and recording as the coal transport waybill data of all customers, as the number of coal customers (i.e. the number of coal transport waybill data), as the set of all coal transport waybill data attributes, as the total number of attributes of each coal transport waybill data, as the th coal transport waybill data, the th data in the waybill data is expressed as , each questionnaire reply result is taken as a sample, the results in each sample are numbered (the parameter options are converted into numerical values) respectively, and all the sample data after numbering are summarized to obtain summary data;

[0082] A3, data fusion, associating the questionnaire and the waybill data through a unique identifier (such as an enterprise unified social credit code) or fuzzy matching (enterprise name + geographical location), wherein the questionnaire filling time is consistent with the waybill data recording period;

[0083] B, using a Fast Causal Inference (FCI) algorithm to construct a causal diagram, screening out characteristic variables having a causal relationship with customer multimodal transport mode selection, and eliminating pseudo-correlated features;

[0084] B1, data preparation, collect data related to the selection of coal customer intermodal mode from the survey results of the questionnaire, the classification variables in the data include: coal characteristics composed of categories and quantities, transportation factors composed of transportation distance, transportation time and transportation cost, customer information composed of enterprise size and transportation frequency, encode the classification variables in the data for the calculation of FCI algorithm; for example, for coal categories, customer enterprise types and other classification variables, methods such as one-hot encoding can be used to convert them into numerical variables.

[0085] B2, FCI algorithm execution;

[0086] B2.1, set the target variable as intermodal mode , the number of intermodal mode categories, and the set of all coal shipment data attributes is the candidate feature set contains all shipment data attributes in the questionnaire (such as transportation cost, distance, coal categories, etc.);

[0087] B2.2, construct a complete graph, take all shipment data variables as nodes, and construct a complete undirected graph, that is, all nodes are connected by edges, and the edges have no direction, only indicating the existence of association between variables;

[0088] B2.3, conditional independence test, for continuous variables, test the direct correlation between the test variable (such as transportation cost) and the intermodal mode under the condition of the given variable (such as transportation distance), calculate the partial correlation coefficient of , and under the influence of , , ; ;

[0089] ;

[0090] ;

[0091] If the partial correlation coefficient is not significant (such as the probability value is greater than the significance level 0.05), it indicates that the test variable has no direct causal relationship with the selection of intermodal mode , and the correlation is caused by the given variable ​​​As a result, the indirect impact of transportation distance on both is excluded (for example, the longer the transportation distance, the higher the cost may be, which in turn affects the choice of intermodal transport mode);

[0092] For discrete variables, chi-square test is performed to test the variables (For example, coal types, total Coal types) and intermodal transport methods (common Are the two methods independent? Calculate the expected frequency ,in, Indicates the The number of samples of each coal type, Indicates the The number of samples of intermodal transport mode, the number of coal customers That is, it is expressed as the total sample size;

[0093] ;

[0094] Combined with The coal type is Number of samples of intermodal transport modes , calculate the chi-square statistic ;

[0095] ;

[0096] Calculating degrees of freedom , find the probability under the chi-square test Value, if If the value is less than the significance level, the null hypothesis is rejected and the test variable is considered Not independent, with intermodal transport Related;

[0097] B2.4. Deletion and retention of edges: Delete given variables based on the results of the conditional independence test. Independent test variables and intermodal transport If, after controlling for the two variables of transportation cost and transportation volume, it is found that the conditions for transportation distance and intermodal mode selection are independent, then the edge between the two variables of transportation distance and intermodal mode is deleted. After step B2.4, a sparse undirected graph is obtained. The edges in the undirected graph indicate the existence of a direct causal relationship between the variables.

[0098] B2.5. Determine some directed edges and determine the direction of some edges through further conditional independence test. If the test variable and given variables There is an edge between them, and given a new variable Under the condition of on the given variable has an effect, but on the given variable on the test variable has no effect, remove the edge from the test variable to the given variable This step results in a partially directed graph, in which the direction of some edges is determined, while the direction of other edges is still undetermined;

[0099] B2.6, determine the direction of the remaining undirected edges by pre-set rules, which include collision structure rules and direction propagation rules,

[0100] collision structure rules, if there exist three variables , , form a path, denoted as , the symbol represents an undirected edge, and and are not adjacent, i.e., there is no direct connection edge, when the other variables are controlled, and are independent, the control variable is controlled, and are related, then is a collision node, and the path is a collision structure, denoted as , where the symbols and represent directed edges;

[0101] direction propagation rules, if there exists a directed path , and and are not adjacent, then the direction is ;

[0102] B2.7, latent variable processing, the relationship between latent variables cannot be directly determined by observed data, but it is inferred that there is a relationship between the latent variable and the target variable based on theory or experience. To consider the influence of latent variables, a latent variable existence test is required. For each latent variable, a virtual edge is added to represent its influence on the observed variable, the graph structure is updated, and the real edge and virtual edge are distinguished, where the real edge is the observed causal relationship and the virtual edge is the influence of the latent variable;

[0103] B2.8, build a causal relationship graph, repeat steps B2.3 to B2.7 to determine the direction of more edges using conditional independence tests and pre-set rules until no more edges can be determined. Finally, a directed acyclic graph is obtained, which is the causal relationship graph between variables inferred from data and FCI algorithm.

[0104] B3, Key variable determination and analysis;

[0105] B3.1, Key variable identification, in the directed acyclic graph obtained in step B2, the variables connected with the intermodal mode selection node through direct or indirect directed edges are the key variables that have causal relationship with the intermodal mode selection of coal customers; for example, if the variables such as transportation cost and transportation distance are connected with the intermodal mode selection node through directed edges, these variables are key variables. Among them, according to the causal relationship diagram, the characteristic variables with direct or indirect causal relationship with the intermodal mode selection node are retained, and the variables without causal relationship with the intermodal mode selection node are removed, that is, the pseudo-correlation characteristics are eliminated.

[0106] B3.2, Causal relationship analysis, according to the direction and path of the directed edge, the causal relationship between the key variable and the intermodal mode selection is analyzed; for example, if there is a directed path from transportation distance to transportation cost, and then from transportation cost to intermodal mode selection, it shows that transportation distance may affect the selection of intermodal mode of coal customers by affecting transportation cost.

[0107] B3.3, Result verification and explanation, the results obtained by FCI algorithm are compared and verified with artificial verification (with actual business knowledge and experience) to ensure the rationality and interpretability of the results; if some causal relationships do not conform to the actual experience, it is necessary to carefully check the data and algorithm execution process to determine whether there are data errors, improper variable selection or unreasonable algorithm parameter setting problems.

[0108] C, Clustering to obtain customer categories using k-prototype hybrid clustering method;

[0109] C1, According to the results of key variable screening in step B, extract the required data from the waybill data , input the hybrid data set after causal relationship selection, and the hybrid data set includes numerical features and category features;

[0110] C2, Determine the optimal number of clusters K by the silhouette coefficient;

[0111] C2.1, Traverse different cluster numbers =2,3,4, ), calculate the distance between any two customer samples. In the k-prototype clustering method, the Euclidean distance is used when the attribute is a numerical feature . For numerical features, normalization processing is needed first to eliminate the influence of dimension, and the maximum and minimum normalization method is used to map the numerical features to the interval [0, 1], and then the Euclidean distance of the numerical features is calculated ;

[0112] ;

[0113] Where, is the number of clusters When, Customer's The standardized value of the attribute, is the number of clusters When, Customer's The standardized value of each attribute;

[0114] Hamming distance is used when the attribute is a categorical feature , retain the original encoding of the category features (such as coal types are mapped to 0 / 1 / 2). If the category attributes of two customers are different, the Hamming distance is 1, and 0 if they are the same;

[0115] ;

[0116] Where, is the number of clusters When, Customer's The value of the attribute mapping, is the number of clusters When, Customer's The value of the attribute mapping, is the ordinal number of the total number of attributes;

[0117] The distance between two customer samples , is the weight factor of the classification attribute, which is used to balance the two types of features;

[0118] C2.2. Traversing different numbers of clusters value( =2,3,4, ), calculate the corresponding silhouette coefficients based on the distances between the obtained customer samples;

[0119] ;

[0120] ;

[0121] Where, is the cohesion, indicating customer points The mean of the distances to other customer points in the cluster, is the separation degree, which represents the mean of all customer points in other clusters closest to the cluster to which it belongs. Points for customers The silhouette coefficient, Global silhouette coefficient;

[0122] C2.3, comparing different clustering numbers Global silhouette coefficient at value , selecting the clustering number K at which the global silhouette coefficient is maximum as the optimal clustering number K;

[0123] C3, according to the determined K cluster customer categories, extracting the typical features of each cluster to form customer grouping labels, and extracting the statistical distribution of each cluster feature (such as the median of transportation cost, the proportion of coal type), defining business understandable customer labels, such as cluster 1: transportation cost > 0.7, transportation distance > 500km, coal type = power coal, which can be marked as "high cost power coal customer".

[0124] D, training LightGBM (Light Gradient Boosting Machine) classification model to predict the category of customers, and testing the accuracy of the portrait, combined with SHAP (SHapley Additive exPlanations) interpretable analysis method to describe the customer portrait;

[0125] D1, setting the target variable as the customer clustering label generated by the k-prototype clustering method is the index of the cluster, and the input features are the key variables obtained in step B (such as transportation cost, transportation distance, coal type, etc.), and all shipping order data are divided into training set and test set in the ratio of 7:3, and the LightGBM model is trained;

[0126] D2, selecting the weighted average as an index to evaluate the accuracy of the customer portrait, and verifying the accuracy of the portrait;

[0127] ;

[0128] In the formula, is the number of customer samples in the th cluster, represents the number of coal customers, i.e. the total sample size, is the score of the th customer;

[0129] For example: the of cost-sensitive long-distance customers (accounting for 30%) = 0.85, and the of other categories = 0.7, then the weighted = 0.3 x 0.85 + 0.7 x 0.7 = 0.745;

[0130] D3, the constructed LightGBM model is explained and analyzed using the SHAP explainable analysis method, the feature density bee colony diagram and the feature influence waterfall diagram are combined to explain the classification model prediction result, the features of different category customer portraits are analyzed, and the SHAP value calculation method is as follows

[0131] ;

[0132] In the formula, is a model benchmark value, which is the average prediction value of the LightGBM model on the training set, is a feature The SHAP value of the sample contribution, the feature density bee colony diagram and the feature influence waterfall diagram are the visualization results of the SHAP explainable analysis method.

[0133] The size of the SHAP value is used as the feature importance, the features are sorted, and the features that have the greatest impact on customer classification (such as transportation cost and transportation distance) are identified.

[0134] The above is only a specific embodiment of the present application, but the protection scope of the present application is not limited thereto, any person skilled in the art can easily think of changes or replacements within the technical range disclosed by the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for constructing an explainable customer profile for coal multimodal transport under a single order system, characterized by: The following steps are involved: A. Obtain customer coal intermodal transport mode selection questionnaires and historical coal multimodal transport waybill data under the single-bill system, and perform data processing and data integration; B. Use the causal discovery FCI algorithm to construct a causal graph, screen out characteristic variables that have a causal relationship with the customer's choice of intermodal mode, and eliminate spurious correlation features; FCI algorithm execution; B2.

1. Let the target variable be the intermodal mode Y = {y1,y2,…,y c }, c is the number of intermodal transport modes, and the set of all coal waybill data attributes is the candidate feature set A = {a1, a2, ..., a M }Variables containing all the waybill data attributes in the questionnaire; B2.

2. Construct a complete graph. Consider all variables in the waybill data as nodes and construct a completely undirected graph. This means that all nodes are connected by edges, but the edges have no direction and only indicate the association between the variables. B2.3, conditional independence test, partial correlation coefficient test for continuous variables, B2.4, Deletion and retention of edges, according to the results of the conditional independence test, delete the given variable a k The independent test variable a j The edges between and the intermodal mode Y result in a sparse undirected graph, where the edges in an undirected graph indicate that there is a direct causal relationship between the variables; B2.

5. Determine some directed edges and determine their directions through further conditional independence tests. B2.

6. Determine the remaining undirected edges using pre-set rules. The pre-set rules include collision structure rules and directional propagation rules. B2.

7. Latent variable processing: For each latent variable, add a virtual edge to represent its influence on the observed variable. Update the graph structure and distinguish between real edges and virtual edges. Real edges represent the causal relationship between observations, while virtual edges represent the influence of latent variables. B2.

8. Build a causal relationship graph. Repeat steps B2.3 through B2.7, using conditional independence tests and pre-defined rules to determine the directions of more edges until no more directions can be determined. This results in a directed acyclic graph. C. Use k-prototype hybrid clustering method to cluster and obtain customer categories; D. Train the LightGBM classification model to predict the customer category, test the accuracy of the portrait, and use the SHAP interpretable analysis method to describe the customer portrait.

2. The method for constructing an explainable customer profile for coal multimodal transport under a single order system according to claim 1 is characterized in that: Step A also includes the following steps: A1. Obtain historical coal single-bill shipping data, develop a customer survey questionnaire on coal intermodal transport mode selection, distribute the questionnaire to coal customers, and collect their responses. A2. Process the waybill and questionnaire data, filter and delete duplicate records, supplement missing values, and correct outliers. After processing, a structured waybill dataset is obtained. Let X = [X1, X2, ..., X n ] is the coal waybill data of all customers, n is the number of coal customers, A={a1,a2,……,a M } is the set of all coal waybill data attributes, M is the total number of attributes of each coal waybill data, m∈{1,2,3,…,M} is the mth coal waybill data, and the i-th data in the waybill data X is represented by X i =[a i1 ,a i2 ,……,a iM ], each questionnaire response result is taken as a sample, the results in each sample are numbered respectively, and all the numbered sample data are summarized to obtain the summary data; A3. Data fusion: Associating the questionnaire with the waybill data through a unique identifier or fuzzy matching. The time when the questionnaire is filled out is consistent with the recording period of the waybill data.

3. The method for constructing an explainable customer profile for coal multimodal transport under a single order system according to claim 1 is characterized in that: Step B also includes the following steps: B1. Data Preparation: Collect data related to coal customers' choice of intermodal transport methods from the questionnaire survey results. The categorical variables in the data include: coal characteristics consisting of type and quantity; transportation factors consisting of transportation distance, transportation time, and transportation cost; and customer information consisting of enterprise size and transportation frequency. Encode the categorical variables in the data for use in the FCI algorithm calculation.

4. The method for constructing an explainable customer profile for coal multimodal transport under a single order system according to claim 3 is characterized in that: Step B2 also includes the following steps: If a variable a is given k Under the condition of j Direct correlation with intermodal mode Y, according to a k 、a j Pearson correlation coefficient between each pair of Y Calculate a j and Y in the absence of a k Partial correlation coefficient under the influence of a j ,a k ,∈A={a1,a2,…,a M }; If the partial correlation coefficient Not significant, indicating that the test variable a j There is no direct causal relationship with the choice of intermodal mode Y, and its correlation is determined by the given variable a k lead to; For discrete variables, chi-square test is performed to test variable a j Is it independent of the intermodal mode Y? Calculate the expected frequency E bc , where R b represents the number of samples of the bth coal type, C c represents the number of samples of the cth intermodal transport mode, and the number of coal customers n is represented as the total sample size; The number of samples using the cth intermodal transport method in combination with the bth coal type O bc , calculate the chi-square statistic χ 2 ; Calculate the degrees of freedom df = (b-1) × (c-1), find the probability p value under the chi-square test, if the p value is less than the significance level, reject the null hypothesis and consider the test variable a j Not independent, but related to intermodal mode Y; If the test variable a is found j and given variable a k There is an edge between them, and given a new variable a z Under the condition of j For a given variable a k Has an impact, but given variable a k For the test variable a j No effect, change the edge from test variable a j Points to the given variable a k , and obtain a partial directed graph; Collision structure rule: if there are three variables a1, a2, and a3 forming a path, expressed as a1-a2-a3, the symbol - represents an undirected edge, and a1 and a3 are not adjacent, that is, there is no direct connection edge. When controlling other variables a4, a1 and a3 are independent, and when controlling variable a2, a1 and a3 are related, then a2 is a collision node, and the path is a collision structure, expressed as a1→a2←a3, where the symbols → and ← represent directed edges; Directional propagation rule: if there is a directed path a1→a2-a3, and a1 and a3 are not adjacent, then the direction is a1→a2→a3.

5. The method for constructing an explainable customer profile for coal multimodal transport under a single order system according to claim 4 is characterized in that: Step B also includes the following steps: B3. Identification and analysis of key variables; B3.

1. Identify key variables. In the directed acyclic graph obtained in step B2, variables that have direct or indirect directed edges connected to the intermodal mode selection node are key variables that have a causal relationship with the coal customer's intermodal mode selection. Based on the causal relationship graph, retain the feature variables that have a direct or indirect causal relationship with the intermodal mode selection node, and remove the variables that have no causal relationship with the intermodal mode selection node, thereby eliminating spurious correlation features. B3.

2. Causal relationship analysis: Analyze the causal relationship between key variables and intermodal transport mode selection based on the direction and path of directed edges; B3.

3. Result verification and interpretation: Compare and verify the results obtained by the FCI algorithm with those obtained by manual verification to ensure the rationality and interpretability of the results.

6. The method for constructing an explainable customer profile for coal multimodal transport under a single order system according to claim 1 is characterized in that: Step C also includes the following steps: C1. According to the results of key variable screening in step B, from the waybill data X=[X1,X2,……,X n ] extract the required data and input the mixed data set after causal relationship selection, the mixed data set includes numerical features and category features; C2, determine the optimal number of clusters K by silhouette coefficient; C2.

1. Traverse different values ​​of the number of clusters k and calculate the distance d between any two customer samples. In the k-prototype clustering method, when the attribute is a numerical feature, the Euclidean distance d1 is used. Use the maximum and minimum normalization method to map the numerical feature to the interval [0, 1], and then calculate the Euclidean distance d1 of the numerical feature; Where a pmk When the number of clusters is k, the standardized value of the mth attribute of the pth customer, a qmk When the number of clusters is k, the standardized value of the mth attribute of the qth customer; When the attribute is a categorical feature, the Hamming distance d2 is used, and the original encoding of the categorical feature is retained. If the categorical attributes of two customers are different, the Hamming distance d2 is 1, and if they are the same, it is 0; d2=∑ k δ(a prk ,a qrk ); Where a prk When the number of clusters is k, the value of the rth attribute mapping of the pth customer, a qrk When the number of clusters is k, the value of the rth attribute mapping of the qth customer, m and r are the ordinal numbers of the total number of attributes; The distance between two customer samples is d = d1 + γd2, where γ is the weight factor of the classification attribute; C2.

2. Traverse different values ​​of the number of clusters k and calculate the corresponding silhouette coefficient based on the distance between each customer sample; In the formula, a(p) is the cohesion, which represents the mean distance from customer point p to other customer points in the cluster to which it belongs; b(p) is the separation, which represents the mean distance from all customer points in the cluster closest to it; S(p) is the silhouette coefficient of customer point p, S avg is the global silhouette coefficient; C2.

3. Comparison of the global silhouette coefficient S under different clustering numbers k avg , select the k value when the global silhouette coefficient is the largest as the optimal number of clusters K; C3. Based on the determined K clusters of customer categories, extract the typical features of each cluster to form customer grouping labels. Extract the statistical distribution of each cluster's features to define customer labels that are understandable to the business.

7. The method for constructing an explainable customer profile for coal multimodal transport under a single order system according to claim 1 is characterized in that: Step D also includes the following steps: D1. Set the target variable to the customer cluster label y generated by the k-prototype clustering method u ,u∈{1,2,3,…,K} is the index of K clusters, the input features are the key variables obtained in step B, all the waybill data X are divided into training set and test set in the ratio of 7:3, and the LightGBM model is trained; D2, select weighted average F1 weighted To evaluate the accuracy of customer portraits and verify the accuracy of portraits; Where n u is the number of customer samples in the u-th cluster, n is the number of coal customers, i.e. the total sample size, F1 u is the F1 score of the customers in the u-th cluster; D3. Use the SHAP interpretable analysis method to interpret the constructed LightGBM model, combine the feature density bee swarm diagram and feature impact waterfall diagram to explain the classification model prediction results, and analyze the characteristics of different categories of customer portraits. The SHAP value calculation method is as follows Where y base is the model benchmark value, which is the average prediction value of the LightGBM model on the training set, v (a u ) is the feature v for sample a u The contribution of SHAP values, feature density bee swarm plots, and feature influence waterfall plots are visualization results of the SHAP interpretable analysis method; The SHAP value is used as the feature importance to sort the features and identify the features that have the greatest impact on customer classification.

Citation Information

Patent Citations

  • Personalized recommendation method based on e-commerce data user portraits

    CN118212034A

  • Power customer tag generation method based on big data clustering technology

    CN114444573A

  • Interpretable visualization method and system based on chronic disease dynamic prediction

    CN120108739A