Railway freight customer portrait construction method, device and computer equipment
By using the ACO-FCM clustering algorithm to perform multiple clustering and chi-square analysis on railway freight customers, the problem of lack of detailed freight service information in existing technologies is solved, the precise construction of railway freight customer portraits and the accuracy of market analysis are achieved, and more targeted marketing strategies are supported.
Patent Information
- Application Number
- CN202411448587.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-17
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2044-10-17
AI Technical Summary
Existing rail freight customer demand analysis methods rely on the RFM model and the KFAV model, which lack detailed freight service information and cannot provide specific insights into customer preferences, hindering the development of more precise marketing strategies.
The ACO-FCM clustering algorithm is used to perform multiple clustering processing on the transportation data of railway freight customers. The Calinski-Harabasz Index, Davies-Bouldin Index and Silhouette Coefficient index are used to determine the optimal total number of clusters. Combined with chi-square analysis, railway freight customer portraits are generated.
It achieves precise clustering and demand analysis of railway freight customers, builds accurate customer portraits for market analysis, and helps to formulate more targeted marketing strategies.
Smart Images

Figure CN119313378B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of railway freight customer management, and specifically relates to a railway freight customer portrait construction method and its device and computer equipment. Background Art
[0002] Meeting diverse customer needs is a key component in increasing rail freight volume. Only by truly understanding customer needs can we better meet them, thereby improving customer satisfaction and encouraging more customers to choose rail freight transportation. Existing customer demand analysis primarily relies on RFM and KFAV models. These traditional models, constructed using data from rail freight ticket systems, lack detailed information on freight services. These methods fail to provide specific insights into customer preferences for rail services, hindering the development of more precise marketing strategies. Summary of the Invention
[0003] The purpose of the present invention is to provide a method for clustering different railway customer groups using the ACO-FCM clustering algorithm, and to perform demand analysis on the clustered groups to construct a railway freight customer portrait. The method has the advantages of good clustering effect and accurate market analysis, and is helpful to formulate more targeted marketing strategies.
[0004] To achieve the above-mentioned object, the present invention provides a method for constructing a railway freight customer profile, the method comprising the following steps:
[0005] Step S10: Acquire transportation data of a plurality of railway freight customers, the transportation data including first transportation data corresponding to attribute-type demand information and second transportation data corresponding to transportation-type demand information, wherein the attribute-type demand information includes transportation category, transportation volume, cargo category, transportation region, and transportation mode, and the transportation-type demand information includes freight type requirements, transportation distance requirements, transportation speed requirements, transportation punctuality requirements, delivery and pickup requirements, normal temperature warehousing requirements, circulation processing requirements, loading and unloading requirements, business acceptance requirements, payment method requirements, information consultation requirements, and insurance consultation requirements;
[0006] Step S20: Applying the ACO-FCM algorithm to perform multiple clustering processes on the first transportation data to obtain multiple clustering information, wherein the total number of clusters K corresponding to each clustering process is different, and each total number of clusters corresponds to one clustering information, K is any integer from 2 to 20, wherein each clustering information includes a clustering result with K clusters, a Calinski-Harabasz Index, a Davies-Bouldin Index, and a Silhouette Coefficient;
[0007] Step S30: Based on the Calinski-Harabasz Index, Davies-Bouldin Index, and Silhouette Coefficient in the plurality of clustering information, determine an optimal total number of clusters M from a plurality of different total numbers of clusters using the elbow method, and use the clustering result corresponding to the optimal total number of clusters M as a target clustering result, wherein the target clustering result includes M target clusters, each of the target clusters having different label features;
[0008] Step S40: Perform chi-square analysis on the second transportation data corresponding to each target cluster to obtain demand preference data of each target cluster, combine the label features and demand preference data of each target cluster, and generate a railway freight customer profile.
[0009] The present invention provides a railway freight customer portrait construction device, the construction device comprising:
[0010] A first acquisition module is configured to acquire transportation data of a plurality of railway freight customers, the transportation data including first transportation data corresponding to attribute-type demand information and second transportation data corresponding to transport-type demand information, wherein the attribute-type demand information includes transportation category, transportation volume, cargo category, transportation region, and transportation mode, and the transportation-type demand information includes freight type requirements, transportation distance requirements, transportation speed requirements, transportation punctuality requirements, delivery and pickup requirements, normal temperature warehousing requirements, circulation processing requirements, loading and unloading requirements, business acceptance requirements, payment method requirements, information consultation requirements, and insurance consultation requirements;
[0011] a second acquisition module, configured to apply an ACO-FCM algorithm to perform multiple clustering processes on the first transportation data to obtain a plurality of clustering information, wherein the total number of clusters K corresponding to each clustering process is different, and each total number of clusters corresponds to a piece of clustering information, K is any integer from 2 to 20, wherein each piece of clustering information includes a clustering result having K clusters, a Calinski-Harabasz Index, a Davies-Bouldin Index, and a Silhouette Coefficient;
[0012] a determination module, configured to determine an optimal total number of clusters M from a plurality of different total numbers of clusters using an elbow method based on the Calinski-Harabasz Index, the Davies-Bouldin Index, and the Silhouette Coefficient in the plurality of clustering information, and use the clustering result corresponding to the optimal total number of clusters M as a target clustering result, wherein the target clustering result includes M target clusters, each of the target clusters having different label features;
[0013] A construction module is used to perform chi-square analysis on the second transportation data corresponding to each target cluster, obtain the demand preference data of each target cluster, combine the label features and demand preference data of each target cluster, and generate a railway freight customer portrait.
[0014] The present invention further provides a computer device, comprising: a processor and a memory, wherein a computer program is stored in the memory, and when the processor executes the computer program, the computer device implements the method described above.
[0015] The present invention also provides a computer-readable storage medium. When instructions in the computer-readable storage medium are executed by a processor, the processor executes the method described above.
[0016] The beneficial effects of the present invention include at least:
[0017] The present invention provides a method for constructing a railway freight customer profile. The method comprises the following steps: obtaining transport data of a plurality of railway freight customers, the transport data comprising first transport data corresponding to attribute-type demand information and second transport data corresponding to transport-type demand information; applying an ACO-FCM algorithm to perform multiple clustering processes on the first transport data to obtain a plurality of clustering information; determining an optimal total number of clusters M based on a CH Index, a DB Index and a Silhouette Coefficient in the plurality of clustering information using an elbow method, and taking a clustering result corresponding to the optimal total number of clusters M as a target clustering result; performing a chi-square analysis on the second transport data corresponding to each target cluster to obtain demand preference data of each target cluster, combining the label features and demand preference data of each target cluster to generate a railway freight customer profile; the present invention utilizes an ACO-FCM clustering algorithm to cluster different railway customer groups, performs demand analysis on the clustered groups, and constructs a railway freight customer profile. The method has the advantages of good clustering effect and accurate market analysis, and is helpful in formulating more targeted marketing strategies.
[0018] In addition to the above-described objects, features and advantages, the present invention has other objects, features and advantages. The present invention will be further described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 A schematic diagram of the steps of a method for constructing a railway freight customer profile according to an embodiment of the present invention;
[0020] Figure 2 A flowchart of the steps for obtaining each cluster information in the method for constructing a railway freight customer profile provided by one embodiment of the present invention;
[0021] Figure 3 A flowchart of step S210 in the method for constructing a railway freight customer profile according to one embodiment of the present invention;
[0022] Figure 4 A flowchart of step S220 in the method for constructing a railway freight customer profile according to one embodiment of the present invention;
[0023] Figure 5 Schematic diagram of CH Index values of ACO-FCM algorithm, FCM algorithm and SAGAFCM algorithm under different cluster numbers;
[0024] Figure 6 Schematic diagram of the Silhouette Coefficient index values of the ACO-FCM algorithm, FCM algorithm, SAGAFCM algorithm and K-mode algorithm under different cluster numbers;
[0025] Figure 7 Schematic diagram of the DBIndex index values of the ACO-FCM algorithm, FCM algorithm, SAGAFCM algorithm and K-mode algorithm under different cluster numbers;
[0026] Figure 8 The box diagram of the Silhouette Coefficient index value of the ACO-FCM algorithm, FCM algorithm, SAGAFCM algorithm and K-mode algorithm under 100 iterations;
[0027] Figure 9 Schematic diagram of the objective function value of the FCM algorithm after 100 iterations;
[0028] Figure 10 Schematic diagram of the objective function value of the ACOFCM algorithm after 100 iterations;
[0029] Figure 11 A radar diagram showing the average values of all cluster groups on multiple transport demand information;
[0030] Figure 12 Radar diagram of target cluster 1 on multiple transport demand information
[0031] Figure 13 Radar diagram of target cluster 2 on multiple transport demand information
[0032] Figure 14 Radar diagram of target cluster 3 on multiple transport demand information
[0033] Figure 15 Radar diagram of target cluster 4 on multiple transport demand information
[0034] Figure 16 Radar diagram of target cluster 5 on multiple transport demand information
[0035] Figure 17 Radar diagram of target cluster 6 on multiple transport demand information
[0036] Figure 18 Radar diagram of target cluster 7 on multiple transport demand information
[0037] Figure 19 Radar diagram of target cluster 8 on multiple transport demand information
[0038] Figure 20 Radar diagram of target cluster 9 on multiple transport demand information
[0039] Figure 21 Radar diagram of target cluster 10 on multiple transport demand information
[0040] Figure 22 Radar diagram of target cluster 11 on multiple transport demand information
[0041] Figure 23 A module diagram of a railway freight customer profile building device provided by one embodiment of the present invention. DETAILED DESCRIPTION
[0042] The embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0043] Please refer to Figures 1 to 4 The present invention provides a method for constructing a railway freight customer profile, the method comprising the following steps:
[0044] Step S10: Acquire transportation data of multiple railway freight customers, wherein the transportation data includes first transportation data corresponding to attribute demand information and second transportation data corresponding to transportation demand information, wherein the attribute demand information includes transportation category, transportation volume, cargo category, transportation area and transportation mode, and the transportation demand information includes freight type demand, transportation distance demand, transportation speed demand, transportation punctuality demand, delivery and pickup demand, normal temperature warehousing demand, circulation processing demand, loading and unloading demand, business acceptance demand, payment method demand, information consultation demand and insurance consultation demand.
[0045] Preferably, the transportation data of railway freight customers are obtained based on a questionnaire.
[0046] Preferably, the step S10 includes:
[0047] Step (a): Based on the data disclosed in the literature database, multiple demand information related to railway freight train products and multiple parameter options corresponding to each demand information are obtained, and based on the characteristics of the demand information, the multiple demand information are divided into attribute demand information related to customer operations and transportation demand information used to describe the customer's transportation needs. The attribute demand information includes transportation category, transportation volume, cargo category, transportation area and transportation mode. The transportation demand information includes freight type requirements, transportation distance requirements, transportation speed requirements, transportation punctuality requirements, delivery and pickup requirements, normal temperature warehousing requirements, circulation processing requirements, loading and unloading requirements, business acceptance requirements, payment method requirements, information consultation requirements and insurance consultation requirements.
[0048] In the present invention, the literature database includes the CNKI platform and the Web of science platform.
[0049] Specifically, a comprehensive literature review was conducted on railway freight trains, railway freight trains, and intermodal transport. A preliminary search was conducted, selecting the CNKI dataset and the Web of Science dataset as the primary data sources. The search terms included "railway freight transport," "railway freight trains," "intermodal transport," "road-rail combined transport," "sea-rail combined transport," and "air-rail combined transport." The search period was 2003 to 2023, yielding a total of 13,510 articles. A secondary search was conducted within the scope of the preliminary search, adding the search term "after customer demand," yielding a total of 110 articles. Manual screening was conducted, selecting 31 articles for further in-depth review after manually reviewing the article titles and abstracts.
[0050] The present invention determines the paper data to be studied by combining multiple searches and manual screening to ensure relevance to the needs of railway freight customers.
[0051] In this invention, expert consultation was used to identify multiple requirements and corresponding parameter options for each requirement, as well as to categorize the requirements. The expert team consisted of three academic professors, two technical personnel from planning and design institutes, two railway professionals, two port professionals, two product suppliers, and one freight forwarder. The expert list summarized demographic information, education level, career path, and interview type, revealing that 66.6% were male and 33.3% were female. Furthermore, 25% held doctoral degrees, 41.6% held master's degrees, and 83.3% had research experience in the railway industry.
[0052] In the present invention, multiple pieces of demand information and multiple parameter options corresponding to each piece of demand information are detailed in Table 1.
[0053] surface Railway Freight Customer Demand Information Form
[0054]
[0055]
[0056] Step (b): converting each demand information and a plurality of parameter options corresponding to each demand information into a single-choice question or a multiple-choice question for display, thereby generating a questionnaire template.
[0057] The questionnaire template includes single-choice questions or multiple-choice questions formed by converting all the demand information. That is, if the number of demand information is 18, then the questionnaire template includes 18 multiple-choice questions.
[0058] For ease of understanding, let's take an example to illustrate the content generation multiple-choice question corresponding to sequence number 1 in Table 1 as follows:
[0059] Which of the following transport categories does your company belong to? [Single-choice question]
[0060] 0Freight forwarding
[0061] ○Transportation companies (ports, shipping, trucking, etc.)
[0062] 0Logistics companies
[0063] 0Commercial distribution companies
[0064] Factories and mines
[0065] ○Other
[0066] The content corresponding to sequence number 4 in Table 1 generates multiple choice questions as follows:
[0067] To which regions does your company mainly transport goods? [Multiple choice]
[0068] ○Domestic
[0069] Europe
[0070] Asia
[0071] ○Other areas
[0072] Step (c): Based on the questionnaire template, obtain multiple questionnaire response results, and use each questionnaire response result as a sample.
[0073] The questionnaire template is sent to railway freight customers by means of on-site distribution, email distribution, WeChat distribution, etc., and the questionnaire response results after the railway freight customers fill out are collected.
[0074] It can be understood that a single-choice question corresponds to one selection result, and a multiple-choice question corresponds to at least one selection result.
[0075] Step (d) numbering the response results in each sample respectively, and summarizing all the sample data after numbering to obtain summary data, wherein the numbering process is to convert the parameter options into numerical values.
[0076] In the present invention, the numbering processing refers to converting the railway freight customer's single-choice question selection results or multiple-choice question selection results into numbers. When the demand information is displayed as a single-choice question, the numbering processing result is selected from any integer from 1 to A, where A is the total number of parameter options; when the demand information is displayed as a multiple-choice question, the numbering processing result is 0 or 1, where 0 indicates not selected and 1 indicates selected.
[0077] For ease of understanding, based on the example above, the transport category in the questionnaire template generates a single-choice question with a total of 6 parameter options. The first parameter option corresponds to the number 1, the second parameter option corresponds to the number 2... The sixth parameter option corresponds to the number 6. When the railway freight customer's response result is the third parameter option, the corresponding number under the transport category question is 3. If the railway freight customer is the 5th customer and the transport category question is the first demand information, the transport data corresponding to the 5th row and 1st column of the summary data is 3. The transport region in the questionnaire template generates a multiple-choice question. The column headers of this demand information include domestic, Europe, Asia and other regions. If the selection result is domestic and Europe, then after numbering, the number corresponding to the parameter option domestic is 1, the number corresponding to the parameter option Europe is 1, the number corresponding to the parameter option Asia is 0, and the number corresponding to the parameter option other regions is 0.
[0078] In the present invention, the summary data may be a summary table, wherein the data portion of the summary table is composed of multiple rows and multiple columns of data, and each row of data corresponds to a questionnaire response result of a railway freight customer.
[0079] In the present invention, the summary table includes a table header, a row header and a data part, wherein the table header includes multiple demand information, each demand information is a single option or multiple options, when the demand information is a single option, the demand information is displayed as the question corresponding to the demand information or the name of the demand information, when the demand information is multiple options, the demand information is displayed as the parameter option corresponding to the demand information; the row header is a serial number, and the data part includes multiple rows and multiple columns of transportation data.
[0080] For ease of understanding, let's assume that 200 survey responses were collected, representing feedback from 200 rail freight customers. There are 18 demand information items, and each selection is a single option. The summary table consists of 200 rows and 18 columns of transportation data. Each row of transportation data corresponds to a survey response, and each column of data corresponds to a single selection.
[0081] Assume that 200 survey responses were collected, representing feedback from 200 rail freight customers. There are 18 demand information items, 17 of which have single-choice options, and one has multiple-choice options with five parameter options. The summary table consists of 200 rows x (17 + 5) columns of transportation data. Each row of transportation data corresponds to a survey response, 17 columns correspond to the selection of a parameter option in each of the 17 demand information items, and the remaining five columns correspond to the selection of each parameter option.
[0082] Step (e): classifying the summarized data to obtain the first transportation data corresponding to the attribute-type demand information and the second transportation data corresponding to the transportation-type demand information.
[0083] In the present invention, the attribute demand information includes transport category, transport volume, cargo category, transport area and transport mode, a total of 5 demand information, and each demand information is a single option. Based on the above example, the data part of the first transport data consists of 200 rows × 5 columns of transport data.
[0084] In this embodiment, the five columns of transportation data of the first transportation data correspond to the first to fifth columns of data in the data portion of the summary table.
[0085] In the present invention, the transportation demand information includes freight type requirements, transportation distance requirements, transportation speed requirements, transportation punctuality requirements, delivery / pickup requirements, normal temperature warehousing requirements, cold chain warehousing requirements, circulation processing requirements, loading and unloading requirements, business acceptance requirements, payment method requirements, information consultation requirements and insurance consultation requirements, a total of 13 demand information. When each demand information is a single option, the data part of the second transportation data consists of 200 rows × 13 columns of transportation data.
[0086] In this embodiment, the 13 columns of transportation data of the second transportation data correspond to the 6th to 18th columns of data in the data portion of the summary table.
[0087] Step S20: Apply the ACO-FCM algorithm to perform multiple clustering processes on the first transport data to obtain multiple railway freight customer clustering information, where the total number of clusters K corresponding to each clustering process is different, and each total number of clusters K corresponds to one railway freight customer clustering information, where K is any integer from 2 to 20, wherein the clustering information includes cluster centers, clustering results, Calinski-Harabasz Index (CH Index), Davies-Bouldin Index (DB Index), and Silhouette Coefficient.
[0088] Specifically, when K=2, the ACO-FCM algorithm is applied to cluster the first transport data to obtain a clustering information; when K=3, the ACO-FCM algorithm is applied to cluster the first transport data to obtain a clustering information; when K=4, the ACO-FCM algorithm is applied to cluster the first transport data to obtain a clustering information; ... when K=20, the ACO-FCM algorithm is applied to cluster the first transport data to obtain a clustering information; that is, a total of 19 clustering information are obtained.
[0089] The steps for obtaining clustering information of each railway freight customer include:
[0090] Step S210: Based on the cluster center K, the first transportation data is used as input data and the ant colony algorithm ACO is used to determine K target cluster centers.
[0091] In the present invention, the first transportation data is input as a matrix X(i, v), where X(i, v) represents the v-th attribute class demand information of the i-th customer.
[0092] In the present invention, the attribute demand information includes transport category, transport volume, cargo category, transport area and transport mode.
[0093] The step of determining the first cluster center by using the ant colony algorithm ACO includes:
[0094] Step S211: Initialize parameters and ant positions, and set the maximum number of iterations t max , where each ant position randomly selects an initial node as the starting point and sets the current path to be empty; the parameters include heuristic information, pheromone matrix N, total number of clusters K, number of ants R, pheromone evaporation rate , and threshold q.
[0095] In this embodiment, the proposed data R is 200.
[0096] In this embodiment, K is any positive integer from 2 to 20.
[0097] In this embodiment, the maximum number of iterations t max 100~500 times.
[0098] Step S212: Check the flag bit. If the flag bit is less than the maximum number of iterations t max , execute step S213; the flag is equal to the maximum number of iterations t max , execute step S219.
[0099] Step S213: Based on the pheromone matrix of the last iteration, the path selection probability of each path is obtained, and the selection path of each ant is determined based on the generated random number r. When the random number r is less than the threshold q, the path corresponding to the highest path selection probability is used as the ant selection path; when the random number r is greater than the threshold q, a path is randomly selected as the ant selection path based on the path selection probability of each path.
[0100] It can be understood that, in the first iteration, the pheromone matrix of the previous iteration is the initialized pheromone matrix, in the second iteration, the pheromone matrix of the previous iteration is the pheromone matrix updated in the first iteration, in the third iteration, the pheromone matrix of the previous iteration is the pheromone matrix updated in the second iteration, and so on.
[0101] The path selection probability calculation formula for each path is as follows:
[0102]
[0103] in, Is the heuristic information, indicating the heuristic value from node i to node j, which is the inverse of the distance from node i to node j. is the pheromone concentration from node i to node j at the tth iteration, is the set of next nodes that the ant is allowed to choose at the current node i, and It is the parameter that regulates pheromones and heuristic information.
[0104] It should be noted that, in this step, the random number r is used to determine the randomness of the ants in selecting paths, and the random number r is between [0, 1].
[0105] It can be understood that each path corresponds to an ant, node i corresponds to each customer, and node j corresponds to each cluster group.
[0106] Step S214: Determine the clustering of each sample based on the selection path of each ant, and obtain multiple membership matrices corresponding to multiple ants. The number of the membership matrices is the same as the number of the ants.
[0107] For ease of understanding, let's assume the number of clusters, K, is 5, and the number of railway freight customers is 200. Then, the membership matrix for each ant consists of 200 rows x 5 columns. If the first railway freight customer (sample 1) belongs to the first cluster, its membership with the first cluster is 1, and its membership with the other four clusters is 0. If the second railway freight customer (sample 2) belongs to the third cluster, its membership with the third cluster is 1, and its membership with the other four clusters is 0. The same applies to the membership data for the other samples.
[0108] Assuming that the number of ants is three, the number of membership matrices is three.
[0109] Step S215: Based on the first transport data and the membership matrix of each ant, calculate the K cluster centers of each ant and the path measurement values corresponding to the K cluster centers of each ant in order from small to large according to the serial number of the ant, and output the K cluster centers of the last ant, where each cluster corresponds to one cluster center, and the cluster centers are calculated according to the first preset formula.
[0110] In this step, the first preset formula is as follows:
[0111]
[0112] Among them, C(k, v) represents the cluster center of the k-th cluster in the v-th attribute class demand information, X(i, v) represents the v-th attribute class demand information of the ith customer, and w(i, k) represents the membership of the ith customer to the k-th cluster. When it is 1, it means that the ith customer belongs to the k-th cluster.
[0113] The calculation formula of the path metric value F is:
[0114]
[0115] in, is the total number of clusters, is the total number of samples, V is the total number of attribute class demand information, For samples and cluster centers The Euclidean distance between .
[0116] In this specific application process, after the cluster center of each ant is calculated, the path metric value is calculated and then saved. The cluster center of the next ant will cover the cluster center of the previous ant, so only the cluster center of the last ant will be output in the end.
[0117] Step S216: Sort the path metric value of each ant in ascending order, determine the s ants corresponding to the first s path metric values, and use the s selected paths corresponding to the s ants as s target paths, where s is a positive integer.
[0118] In this embodiment, s is 2. In other embodiments, s may also be other positive integers such as 3, 4, or 5.
[0119] For ease of understanding, let's take an example and assume that the total number of ants is 200, and the two ants with the smallest path metric values are ant 10 and ant 20. Then the path selected by ant 10 is taken as one target path, and the path selected by ant 20 is taken as another target path, thereby determining two target paths.
[0120] Step S217: Locally optimize the s target strip paths respectively to obtain s new target strip paths, and update the pheromone matrix of the previous iteration based on the s new target paths to obtain the pheromone matrix of the current iteration.
[0121] Wherein, the calculation formula for the element pheromone concentration in the pheromone matrix is:
[0122]
[0123] is the pheromone concentration from path i to path j at the tth iteration, is the updated pheromone concentration, is the pheromone evaporation rate, which indicates the decay probability of the pheromone, t is the number of iterations, is the sum of the metrics of the first s paths.
[0124] Specifically, a random number array (1 * total number of samples) is generated, and a threshold pls is used to determine whether to adjust the cluster assignment of each sample. If the random number in the random number array is less than the threshold pls, the cluster of the sample is excluded from the current cluster, a new cluster is randomly selected, and the cluster assignment of the sample is updated to the newly selected cluster. If the random number in the random number array is equal to or greater than the threshold pls, the cluster assignment remains unchanged.
[0125] It should be noted that the sample here refers to railway freight customers.
[0126] If the adjusted target path is better than the original path, the new target path will replace the original target path. If the adjusted target path is not better than the original target path, the original target path will be used as the new target path.
[0127] In the present invention, the number of rows of the pheromone matrix is the same as the total number of samples, and the number of columns is the same as the total number of clusters.
[0128] Step S218: After executing step S217, the flag bit is increased by 1, and the process returns to execute step S212.
[0129] Step S219: Obtain the target cluster center of each cluster, where the target cluster center is the cluster center of the last ant corresponding to the maximum number of iterations.
[0130] In the present invention, the target cluster centers corresponding to the K clusters can be expressed as: , where k1 represents the center of cluster 1, k2 represents the center of cluster 2, and k n represents the center of cluster n, where n=k.
[0131] Step 220: Based on the K target cluster centers and the total number of clusters K, perform secondary clustering on the railway freight customers using the fuzzy C-means algorithm (FCM) to obtain clustering information corresponding to the total number of clusters K. The clustering information includes a clustering result with K clusters, a Calinski-Harabasz Index, a Davies-Bouldin Index, and a Silhouette Coefficient.
[0132] This step includes:
[0133] Step S221: Based on the total number of clusters K and the target cluster center of each cluster, an initialized membership matrix is obtained using a uniform distribution method.
[0134] In the present invention, the initialization of the membership matrix U0 adopts a uniform distribution in the range [0, 1], wherein each element u in the membership matrix U0 is ik Indicates the degree to which customer i belongs to the kth cluster.
[0135] For ease of understanding, let's take an example. The data of the membership matrix U0 can be shown in Table 2:
[0136] Table 2 Example of membership degree matrix U0
[0137]
[0138] As can be seen from Table 1, the total number of samples, i.e. the number of railway freight transport customers, is 5, the total number of clusters K=3, and the degree of membership of sample 1 (customer 1) to cluster 1 is u 11 =0.268346074, and the degree of membership to cluster 2 is 0.172450204…….
[0139] Step S222: Based on the first transport data and the membership matrix updated in the last iteration, a second preset formula is used to obtain the cluster center of each cluster updated in the current iteration, wherein, in the first iteration, the membership matrix updated in the last iteration is the initialized membership matrix.
[0140] In this step, the second preset formula is as follows:
[0141]
[0142] Among them, C(k, v) represents the cluster center of the k-th cluster in the v-th attribute class demand information, X(i, v) is the v-th attribute class demand information of the i-th customer, is the degree of membership of the i-th customer to the k-th cluster, m is the fuzzy coefficient, and N is the total sample size.
[0143] It can be understood that in the first iteration, the cluster centers in step 222 are the target cluster centers, that is, the K target cluster centers finally output by the ant colony algorithm. In other iterations, the cluster centers in step 222 are the cluster centers obtained by iterative calculation using the second preset formula.
[0144] Step S223: Based on the cluster center updated in the current iteration of each cluster, the new membership degree of each railway freight customer to each cluster is obtained using the membership degree calculation formula, and the membership degree matrix updated in the current iteration is obtained.
[0145] The membership degree calculation formula is:
[0146]
[0147] in, is a data point and cluster centers the distance between them; is a data point and cluster centers The distance between It represents the cluster center of the j-th cluster in the v-th attribute class demand information; m is the fuzzy coefficient, and K is the total number of clusters.
[0148] Step S224: subtract the membership matrix updated in the current iteration from the membership matrix updated in the previous iteration. When the absolute value of the maximum difference obtained by the subtraction is greater than or equal to the preset convergence accuracy, iteratively execute step S222; when the absolute value of the maximum difference obtained by the subtraction is less than the preset convergence accuracy or the number of iterations reaches the preset maximum number of iterations, execute step S225.
[0149] In the present invention, the convergence accuracy is 0.00001.
[0150] In the present invention, when the absolute value of the maximum difference is less than the preset convergence accuracy, the iteration end condition is met; correspondingly, when the absolute value of the maximum difference is greater than or equal to the preset convergence accuracy, the iteration end condition is not met and iteration needs to be continued.
[0151] For ease of understanding, let's take an example. Assume that the new membership matrix and the current membership matrix are both 200*5 matrices. Then, subtracting the two matrices also results in a 200*5 matrix. If the maximum value of each element in the matrix obtained by subtraction is 0.00001074 (absolute value), it is compared with the convergence precision of 0.00001. If it is greater than the convergence precision, the iteration condition is not met and the iteration needs to continue; if the maximum value of each element in the matrix obtained by subtraction is 0.000009 (absolute value), it is compared with the convergence precision of 0.00001. If it is less than the convergence precision, the iteration end condition is met and the next iteration is not performed.
[0152] In the present invention, when the number of iterations reaches a preset maximum number of iterations of 1000 and the convergence condition is still not satisfied, reaching the preset maximum number of iterations can be used as an iteration termination condition.
[0153] Step S225: Based on the membership matrix updated in the current iteration and the cluster center updated in the current iteration of each cluster, cluster information is obtained, where the cluster information includes a clustering result with K clusters, a Calinski-Harabasz Index (CH Index), a Davies-Bouldin Index (DB Index), and a Silhouette Coefficient.
[0154] In this step, the membership matrix updated based on the current iteration and the cluster centers updated based on the current iteration of each cluster can be understood as the membership matrix and K cluster centers obtained in the last iteration.
[0155] The clustering result is obtained as follows:
[0156] Obtain the membership degree data of each railway freight customer (each sample) to each cluster based on the membership degree matrix updated in the current iteration;
[0157] Determining the cluster of each railway freight customer based on the sizes of the membership degree data of the multiple clusters corresponding to each railway freight customer, wherein the cluster corresponding to each railway freight customer has the largest membership degree data;
[0158] The railway freight customers corresponding to each cluster are aggregated to obtain clustering results, which include the set R of each cluster. k .
[0159] For ease of understanding, let's take an example. For example, when the first railway freight customer's membership degree data for the three clusters are A1, A2, and A3, if A3 is the largest, then the first railway freight customer belongs to cluster 3. According to this method, a cluster label is assigned to each railway freight customer, and then all railway freight customers are assigned to the corresponding cluster according to the cluster label, forming a set R for each cluster. k .
[0160] In a specific embodiment, the cluster labels can be represented as cluster 1, cluster 2, etc., or as the first cluster, the second cluster, etc., or as cluster A, cluster B, etc., which only needs to be able to distinguish each cluster.
[0161] The calculation formula of the Calinski-Harabasz Index (CH Index) is:
[0162]
[0163] Where N is the total number of customers; K is the total number of clusters; is the between-class divergence, is the intra-class divergence.
[0164] The calculation formula of the Davies-Bouldin Index (DB Index) is:
[0165]
[0166] Where K is the total number of clusters; is the average distance of cluster i; is the average distance of cluster j; is the cluster center and The distance between them.
[0167] The calculation formula of the Silhouette Coefficient index is:
[0168]
[0169] Where a(i) is the average distance from the i-th customer to other points in the same cluster (intra-cluster closeness); b(i) is the average distance from the i-th customer to all points in the nearest other cluster (inter-cluster separation).
[0170] Step S30: Based on the clustering information of multiple railway freight customers, determine the optimal total cluster number M among multiple different total cluster numbers based on the elbow method, and use the clustering result corresponding to the optimal total cluster number as the railway freight customer clustering result, wherein the clustering result includes M clusters, each of which has different label features.
[0171] It can be understood that the optimal total number of clusters M is an integer between 2 and 20.
[0172] Specifically, the Calinski-Harabasz Index (CH Index), Davies-Bouldin Index (DB Index), and Silhouette Coefficient corresponding to different clustering categories are plotted. The trends of these three indicators as a function of the total number of clusters are then analyzed. When the number of clusters increases to a certain critical point, the rate of decline of the curve slows significantly, forming an "elbow." This point in the curve is considered a reasonable number of clusters, as further increasing the number of clusters will not significantly reduce the error.
[0173] Please refer to Figures 5 to 7 As can be seen from the figure, the Silhouette Coefficient and CH Index tend to be stable around the 10th total number of clusters, while the DB Index drops sharply after the total number of clusters exceeds 11. Therefore, it is ideal to choose a total number of clusters of 10 or 11. Among them, the total number of clusters of 11 shows better clustering performance and shows a lower DB Index. In this invention, the optimal total number of clusters is 11.
[0174] In the present invention, the clustering results corresponding to a total number of clusters being 11 (ie, K=11) are used as the railway freight customer clustering results.
[0175] This paper adopts two methods to evaluate clustering performance. The first method compares the objective function values between the traditional FCM method and the ACO-FCM method, with the aim of demonstrating the effectiveness of the ACO-FCM method in terms of clustering performance. Figure 9 and Figure 10 The second method compares the external indicators of the ACO-FCM method with the traditional FCM algorithm, SAGAFCM algorithm and K-modes algorithm. The above comparison is to further verify the classification tightness and stability of the ACO-FCM method, as shown in Figures 5 to 8 shown.
[0176] The ACO-FCM method has a strong global search capability through ant path selection, pheromone update iteration and local optimization, which can avoid falling into the problem of local optimal solution and help find a better initial center point of the target cluster center, thereby improving the clustering performance of the FCM algorithm; at the same time, since the selection of the initial cluster center has a great influence on the results of FCM clustering, the ACOFCM method optimizes the initial cluster center through ACO, making the clustering results more accurate and stable.
[0177] Step S40: Perform chi-square analysis on the second transportation data corresponding to each target cluster to obtain demand preference data of each target cluster, combine the label features and demand preference data of each target cluster, and generate a railway freight customer profile.
[0178] In the present invention, the second transport data is input in a statistical table format or a statistical analysis format. Specifically, the summary table corresponding to the summary data can be first used as a basic table, and then the basic table can be updated to obtain a table corresponding to the data to be analyzed, and finally the table or the statistical analysis format data corresponding to the table can be input as input data.
[0179] In the present invention, the second transportation data is input in a statistical table format.
[0180] In one embodiment, performing chi-square analysis on the second transportation data corresponding to each target cluster to obtain demand preference data of each target cluster includes:
[0181] Step 1: Based on the summary table and the railway freight customer clustering result, an updated summary table is obtained.
[0182] This step can be:
[0183] Step 1) Delete multiple columns of data corresponding to the first transport data on the summary table, and add M columns to obtain an intermediate summary table, wherein the headers of the M columns respectively correspond to a cluster serial number.
[0184] In the present invention, M=11, that is, railway freight customers are divided into 11 clusters, and 11 columns are added to the intermediate summary table. The header of each column displays the serial number of the 11 clusters, and the headers of the new 11 columns are Cluster 1, Cluster 2, Cluster 3, Cluster 4, Cluster 5, Cluster 6, Cluster 7, Cluster 8, Cluster 9, Cluster 10, and Cluster 11.
[0185] Step 2) Based on the cluster serial number of each railway freight customer, update the data portion of the intermediate summary table to obtain an updated summary table, wherein the data in the newly added M columns are represented by 0 or 1, where 0 represents no and 1 represents yes.
[0186] For ease of understanding, let's take an example and assume that the first railway freight customer is divided into Cluster 6 after the clustering process in step S20. Then, except for the column corresponding to Cluster 6, which is input as 1, all other columns are input as 0.
[0187] Assume that the summary table consists of 200 rows × 18 columns of transportation data, the first transportation data corresponds to 5 columns, and M columns are 11 columns, then the updated summary table consists of 200 rows × 24 columns of data.
[0188] This step may also include:
[0189] Step 1) Add M columns to the summary table to obtain an intermediate summary table, wherein the headers of the M columns correspond to a cluster number respectively.
[0190] Step 2) Based on the cluster serial number of each railway freight customer, update the data part of the intermediate summary table to obtain an updated summary table, wherein the data of the newly added M column is represented by 0 or 1, 0 represents no, and 1 represents yes.
[0191] For ease of understanding, let's take an example and assume that the summary table consists of 200 rows and 18 columns of transportation data. The updated summary table then consists of 200 rows and 29 columns of data.
[0192] It is understandable that when performing the chi-square analysis, inputting the first transport data simultaneously will not affect the chi-square analysis result. In practical applications, inputting the first transport data simultaneously will simplify the updating steps of the summary table.
[0193] Step 2: Using the updated summary table as input data, perform chi-square analysis using SPSS application software to obtain the demand preference data of each cluster.
[0194] In another embodiment, performing chi-square analysis on the second transportation data corresponding to each target cluster to obtain demand preference data of each target cluster includes:
[0195] Step 1: Based on the second transport data in the summary table and the clustering results, obtain M data sets corresponding one-to-one to M target clusters, wherein the target data set includes the second transport data of all railway freight customers in the cluster corresponding to the target data set, and the target data set is any data set among the M data sets.
[0196] For ease of understanding, for example, assuming that the total number of railway freight customers is 200, which are divided into 11 target clusters, and the first target cluster includes 20 railway freight customers, then the first data set corresponding to the first target cluster in the M data sets includes the second transportation data of the 20 railway freight customers in the first target cluster.
[0197] In the present invention, the data of each data set is summarized in a table format.
[0198] Based on the above example, the first data set includes 20 rows of data, and the column data includes the second transportation data corresponding to the transportation demand information.
[0199] Step 2: Input the M data sets into SPSS application software for chi-square analysis to obtain the demand preference data of each target cluster.
[0200] Please refer to Figures 8 to 19 In the present invention, the generated railway freight customer portrait is as follows:
[0201] Cluster 1 (target cluster 1) label characteristics: logistics and transportation companies that mainly transport fresh goods with a transportation volume of less than 500,000 tons.
[0202] Cluster 1 (target cluster 1) has a wide range of transportation areas and modes, with no apparent preference. In terms of transportation distance, 80% of users prefer medium-distance transportation of 300-800 km. While 51.43% of users purchase fresh produce, 68.57% do not require cold chain storage. In terms of distribution and processing, unlike other clusters, users have a greater demand for packaging and metering, with 34.29% and 28.57% of users, respectively. Regarding payment methods, 60% of users prefer regular settlement and a one-ticket system.
[0203] Cluster 2 (target cluster 2) label characteristics: commercial distribution companies that mainly transport non-fresh goods with a transportation volume of less than 500,000 tons.
[0204] Cluster 2 (target cluster 2) has demand preference data: its users transport only domestically, do not have dedicated railway lines, and mainly rely on road-rail intermodal transport, with no involvement in air-rail intermodal transport; 30% of them have a strong demand for metering; in terms of loading and unloading, a small group requires door-to-door loading and unloading; in terms of payment methods, users mainly prefer to pay freight in advance and have freight and miscellaneous fees calculated separately; in terms of information consultation, they show a strong preference for freight rate consultation, with 90% requiring freight rate consultation.
[0205] Cluster 3 (target cluster 3) label characteristics: logistics and transportation companies with a transportation volume of less than 500,000 tons, mainly transporting non-bulk goods and fresh goods.
[0206] Cluster 3 (target cluster 3) demand preference data: its users only transport domestically, mainly by rail and road-rail combined transport, and do not involve air-rail combined transport, showing a strong preference for fast transportation; none of them require room temperature storage and cold chain storage, and in terms of circulation processing, 94.74% do not need circulation processing services; in terms of payment methods, they tend to prefer the pre-deposit first-deduction one-ticket system; in terms of information consultation, unlike other cluster groups, they show a high demand for ticket system consultation and claims consultation.
[0207] Cluster 4 (target cluster 4) label characteristics: Freight forwarding companies that mainly transport bulk cargo with a transportation volume of less than 500,000 tons.
[0208] Cluster 4 (target cluster 4) demand preference data: its users transport goods worldwide, using various modes of transportation; transportation distances are primarily over 800 km; unlike other cluster groups, they prefer a one-ticket payment method with periodic settlements and a strong preference for short-term room-temperature storage.
[0209] Cluster 5 (target cluster 5) label characteristics: transportation and logistics companies with a full range of transportation volumes that mainly transport non-fresh goods.
[0210] Cluster 5 (target cluster 5) demand preference data: its transportation range covers domestic, European, and Asian markets, mainly using road-rail and sea-rail transport. It requires a punctuality error of less than 1 day and high timeliness. In terms of distribution processing, it has high demands for packaging and sorting, and freight and miscellaneous expenses need to be calculated separately. In terms of information consultation, there is a strong demand for freight rate consultation, transportation plan consultation, transportation capacity consultation, and transportation technology consultation.
[0211] Cluster 6 (target cluster 6) label characteristics: transportation and logistics companies that mainly transport non-fresh goods and have a transportation volume of less than 1 million tons.
[0212] Cluster 6 (target cluster 6) demand preference data: Transportation primarily targets domestic and European markets, primarily using rail, combined road-rail, and sea-rail transport. They show no strong preference for transport speed, requiring on-time delivery within a day. Most do not require room temperature or cold chain storage. Regarding distribution processing, they have high demands for packaging, metering, sorting, and assembly.
[0213] Cluster 7 (target cluster 7) label characteristics: commercial circulation enterprises and factories and mining enterprises that transport various goods with a volume of less than 1 million tons.
[0214] Cluster 7 (target cluster 7) demand preference data: its transportation areas mainly include domestic, Asia, and other regions except Europe. In terms of delivery and pickup needs, self-collection is the main method, with strong demand for packaging and measurement. In terms of payment methods, pre-deposit and periodic settlement are the main methods.
[0215] Cluster 8 (target cluster 8) label characteristics: factories and mining enterprises that transport bulk cargo with a volume of 500,000 to 1 million tons.
[0216] Cluster 8 (target cluster 8) demand preference data: its users have global transportation, mainly by rail, road-rail, and sea-rail. In terms of freight types, all are full truckload and container transport. The punctuality requirements are more relaxed than those of other clusters. In terms of warehousing, neither room temperature warehousing nor cold chain warehousing is required. There are higher demands for packaging, metering, sorting, and assembly. In terms of information consultation, compared with other clusters, the proportion of paid consultation is higher.
[0217] Cluster 9 (target cluster 9) label characteristics: factories, mines and other enterprises that transport various goods with a volume of less than 1 million tons.
[0218] Cluster 9 (target cluster 9) demand preference data: Users have global transportation needs, using various modes. For delivery and pickup, they primarily choose self-service delivery and pickup. They also have demand for various distribution and processing services, primarily seeking freight rate consultation.
[0219] Cluster 10 (target cluster 10) label features: logistics companies that transport bulk cargo with a volume of 500,000 to 1,000,000 tons.
[0220] Cluster 10 (target cluster 10) demand preference data: The main transportation areas are domestic and European, with a preference for fast transportation. In terms of payment methods, there is a high demand for advance payment and regular settlement. In terms of information consultation, there is a preference for transportation capacity consultation, freight rate consultation, transportation plan consultation, and claims consultation.
[0221] Cluster 11 (target cluster 11) label features: 1 Logistics companies that transport bulk cargo with a transportation volume of less than 500,000 tons.
[0222] Cluster 10 (target cluster 10) demand preference data: mainly domestic transportation, mainly by rail, road-rail transport, and sea-rail transport, requiring fast transportation with an error of no more than 1 day, with strong time constraints and high demand for measurement and assembly.
[0223] From the demand preference data, it can be seen that the cluster groups are significant for freight type, transportation distance, cold chain storage, circulation processing, information consultation and insured transportation, among which insured transportation shows a significant difference at the 0.05 level (x 2 =19.125, p=0.039<0.05), and the rest showed significance at the 0.01 level, that is, there was a significant correlation.
[0224] Overall, users prefer fast transportation, requiring punctuality within 3 days, and are relatively balanced in terms of seasonality. Dedicated railway lines are generally divided into two situations: those with no dedicated railway line and those with normal use of dedicated railway lines; both room temperature storage and cold chain storage show no need for storage; there is relatively little demand for circulation and processing services; in terms of loading and unloading, users generally carry out their own loading and unloading; in terms of business and services, they tend to accept and confirm orders online, hope for freight rate consulting services, most choose railway insured transportation, and do not need supply chain finance for the time being.
[0225] The construction method further includes analyzing the correlation between demand preferences. Preferably, the Kendall coefficient is used to analyze the correlation between the demand preferences.
[0226] Certain demand characteristics revealed by correlation analysis may appear simultaneously in the same customer group, which can help understand why certain customer groups show consistency in specific demand characteristics.
[0227] See also Figure 23 The present invention further provides a railway freight customer profile building device 100, characterized in that the building device 100 includes:
[0228] A first acquisition module 101 is configured to acquire transportation data of a plurality of railway freight customers, wherein the transportation data includes first transportation data corresponding to attribute-based demand information and second transportation data corresponding to transport-related demand information, wherein the attribute-based demand information includes transportation category, transportation volume, cargo category, transportation region, and transportation mode, and the transport-related demand information includes freight type requirements, transportation distance requirements, transportation speed requirements, transportation punctuality requirements, delivery and pickup requirements, normal temperature warehousing requirements, circulation processing requirements, loading and unloading requirements, business acceptance requirements, payment method requirements, information consultation requirements, and insurance consultation requirements;
[0229] A second acquisition module 102 is configured to apply the ACO-FCM algorithm to perform multiple clustering processes on the first transportation data to obtain a plurality of clustering information, wherein the total number of clusters K corresponding to each clustering process is different, and each total number of clusters corresponds to a piece of clustering information, where K is any integer from 2 to 20, wherein each piece of clustering information includes a clustering result having K clusters, a Calinski-Harabasz Index, a Davies-Bouldin Index, and a Silhouette Coefficient.
[0230] a determination module 103 configured to determine an optimal total number of clusters M from a plurality of different total numbers of clusters using an elbow method based on the Calinski-Harabasz Index, the Davies-Bouldin Index, and the Silhouette Coefficient in the plurality of clustering information, and to use the clustering result corresponding to the optimal total number of clusters M as a target clustering result, wherein the target clustering result includes M target clusters, each of the target clusters having different label features;
[0231] The construction module 104 is used to perform chi-square analysis on the second transportation data corresponding to each target cluster, obtain the demand preference data of each target cluster, combine the label features and demand preference data of each target cluster, and generate a railway freight customer profile.
[0232] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the working process of the construction device described above can refer to the contents of the aforementioned method embodiment and will not be repeated here.
[0233] The present invention also provides a computer device, Figure 23 The construction device shown can be deployed in the computer device. The computer device includes a memory, a processor, a communication interface, and a bus. The memory, processor, and communication interface are interconnected via the bus. Furthermore, the computer device may include multiple processors, so that different processors can implement the functions of the different modules described above.
[0234] The memory can be a read-only memory, a static storage device, a dynamic storage device, or a random access memory. The memory can store executable code. When the executable code stored in the memory is executed by the processor, the processor and the communication interface are used to execute the railway freight customer profile construction method provided in the embodiments of the present application. The memory can also include software modules and data required for other running processes, such as an operating system. The operating system can be LINUX, UNIX, WINDOWS™, etc.
[0235] The processor may be a general-purpose central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), a graphics processing unit (GPU), or one or more integrated circuits.
[0236] The processor may also be an integrated circuit chip with signal processing capabilities. During implementation, some or all of the functions of the railway freight customer profile construction method of the present application may be performed by hardware integrated logic circuits or software instructions in the processor. The aforementioned processor may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component. The various methods, steps, and logic block diagrams disclosed in the embodiments of the present application may be implemented or executed. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in conjunction with the embodiments of the present application may be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in a memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the railway freight customer profile construction method of the embodiments of the present application.
[0237] A communication interface uses a transceiver module, such as, but not limited to, a transceiver, to enable communication between a computer device and other devices or a communication network. For example, a communication interface can be any one or a combination of the following devices: a network interface (such as an Ethernet interface), a wireless network card, or other network access device.
[0238] A bus may include a pathway that transfers information between various components of a computer device (eg, memory, processor, communication interface).
[0239] A communication path is established between each of the above-mentioned computer devices via a communication network. Each computer device is used to implement part of the functions of the railway freight customer profile construction method provided in the embodiment of the present application. Any computer device can be a computer device in a cloud data center (e.g., a server) or a computer device in an edge data center.
[0240] The descriptions of the processes corresponding to the above figures have different focuses. For parts that are not described in detail in a certain process, please refer to the relevant descriptions of other processes.
[0241] In the above embodiments, all or part of the embodiments may be implemented using software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments may be implemented in the form of a computer program product. The computer program product providing the data synchronization cloud service includes one or more computer instructions. When these computer program instructions are loaded and executed on a computer device, they fully or partially implement the process or functions of the railway freight customer profile construction method provided in the embodiments of the present application.
[0242] The computer device may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. Computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium stores computer program instructions for providing data synchronization cloud services.
[0243] An embodiment of the present application also provides a storage medium, which is a non-volatile computer-readable storage medium. When the instructions in the storage medium are executed by the processor, the railway freight customer portrait construction method provided in the embodiment of the present application is implemented.
[0244] An embodiment of the present application also provides a computer program product containing instructions. When the computer program product is run on a computer, it enables the computer to execute the railway freight customer portrait construction method provided by the embodiment of the present application.
[0245] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk, etc.
[0246] The above content is a further detailed description of the present invention in conjunction with specific preferred embodiments, and the specific implementation of the present invention should not be considered to be limited to these descriptions. For those skilled in the art of the present invention, several simple deductions and substitutions can be made without departing from the concept of the present invention, and all of these should be considered to fall within the scope of protection of the present invention.
Claims
1. A method for constructing a railway freight customer profile, characterized in that: The construction method comprises the following steps: Step S10: Acquire transportation data of a plurality of railway freight customers, the transportation data including first transportation data corresponding to attribute-type demand information and second transportation data corresponding to transportation-type demand information, wherein the attribute-type demand information includes transportation category, transportation volume, cargo category, transportation region, and transportation mode, and the transportation-type demand information includes freight type requirements, transportation distance requirements, transportation speed requirements, transportation punctuality requirements, delivery and pickup requirements, normal temperature warehousing requirements, circulation processing requirements, loading and unloading requirements, business acceptance requirements, payment method requirements, information consultation requirements, and insurance consultation requirements; Step S20: Applying the ACO-FCM algorithm to perform multiple clustering processes on the first transportation data to obtain multiple clustering information, wherein the total number of clusters K corresponding to each clustering process is different, and each total number of clusters corresponds to a clustering information, K is any integer from 2 to 20, wherein each clustering information includes a clustering result with K clusters, a Calinski-Harabasz Index, a Davies-Bouldin Index, and a Silhouette Coefficient; Step S30: Based on the Calinski-Harabasz Index, Davies-Bouldin Index, and Silhouette Coefficient in the plurality of clustering information, determine an optimal total number of clusters M from a plurality of different total numbers of clusters using the elbow method, and use the clustering result corresponding to the optimal total number of clusters M as a target clustering result, wherein the target clustering result includes M target clusters, each of the target clusters having different label features; Step S40: Perform chi-square analysis on the second transportation data corresponding to each target cluster to obtain demand preference data of each target cluster, combine the label features and demand preference data of each target cluster, and generate a railway freight customer profile.
2. The method for constructing a railway freight customer profile according to claim 1, characterized in that: The step S10 includes: Based on the data disclosed in the literature database, multiple demand information related to railway freight train products and multiple parameter options corresponding to each demand information are obtained, and based on the characteristics of the demand information, the multiple demand information are divided into attribute demand information related to customer operations and transportation demand information used to describe the customer's transportation needs. The attribute demand information includes transportation category, transportation volume, cargo category, transportation area and transportation mode. The transportation demand information includes freight type requirements, transportation distance requirements, transportation speed requirements, transportation punctuality requirements, delivery and pickup requirements, normal temperature warehousing requirements, circulation processing requirements, loading and unloading requirements, business acceptance requirements, payment method requirements, information consultation requirements and insurance consultation requirements; Convert each demand information and multiple parameter options corresponding to each demand information into a single-choice question or a multiple-choice question for display, and generate a questionnaire template; Based on the questionnaire template, obtain multiple questionnaire response results and take each questionnaire response result as a sample; The response results in each sample are numbered respectively, and all the sample data after numbering are summarized to obtain summary data, wherein the numbering process is to convert the parameter options into numerical values; The summary data is classified to obtain the first transportation data corresponding to the attribute-type demand information and the second transportation data corresponding to the transportation-type demand information.
3. The method for constructing a railway freight customer profile according to claim 1, characterized in that: The steps for obtaining each cluster information include: Step S210: Based on the total number of clusters K, the first transportation data is used as input data and an ant colony algorithm ACO is used to determine K target cluster centers; Step S220: Based on the K target cluster centers and the total number of clusters K, the railway freight customers are secondary clustered using the fuzzy C-means algorithm FCM to obtain clustering information corresponding to the total number of clusters K, wherein the clustering information includes a clustering result with K clusters, a Calinski-Harabasz Index, a Davies-Bouldin Index, and a Silhouette Coefficient.
4. The method for constructing a railway freight customer profile according to claim 3, characterized in that: The step S210 includes: Step S211: Initialize parameters and ant positions, and set the maximum number of iterations t max , where each ant position randomly selects an initial node as the starting point and sets the current path to be empty; the parameters include heuristic information, pheromone matrix, total number of clusters K, number of ants R, pheromone evaporation rate , and threshold q; Step S212: Check the flag bit. If the flag bit is less than the maximum number of iterations t max , iteratively execute step S213; the flag bit is equal to the maximum number of iterations t max , execute step S219; Step S213: Based on the pheromone matrix of the previous iteration, the path selection probability of each path is obtained, and the selected path of each ant is determined based on the generated random number r. When the random number r is less than the threshold q, the path corresponding to the highest path selection probability is used as the selected path of the ant; when the random number r is greater than the threshold q, a path is randomly selected based on the path selection probability of each path as the selected path of the ant; Step S214: determining the clustering of each sample based on the selection path of each ant, and obtaining a plurality of membership matrices corresponding to the plurality of ants, wherein the number of the membership matrices is the same as the number of the ants; Step 215: Based on the first transport data and the membership matrix of each ant, calculate the K cluster centers of each ant and the path metric values corresponding to the K cluster centers of each ant in ascending order according to the ant's sequence number, and output the K cluster centers of the last ant. Each cluster corresponds to one cluster center, and the cluster centers are calculated according to a first preset formula, wherein the first preset formula is: Where C(k, v) represents the cluster center of the k-th cluster in the v-th attribute class demand information, X(i, v) represents the v-th attribute class demand information of the ith customer, and w(i, k) represents the membership degree of the ith customer to the k-th cluster. When it is 1, it means that the ith customer belongs to the k-th cluster. Step S216: Sort the path metric values of each ant in ascending order, determine the s ants corresponding to the first s path metric values, and use the s selected paths corresponding to the s ants as the s target paths, where s is a positive integer. Step S217: Locally optimize the s target strip paths to obtain s new target strip paths, and update the pheromone matrix of the previous iteration based on the s new target paths to obtain the pheromone matrix of the current iteration, wherein the calculation formula for the element pheromone concentration in the pheromone matrix is: is the pheromone concentration from path i to path j at the tth iteration, is the updated pheromone concentration, is the pheromone evaporation rate, which indicates the decay probability of the pheromone, t is the number of iterations, is the sum of the metrics of the first s paths; Step S218: After executing step S217, the flag bit is increased by 1, and the process returns to step S212; Step S219: Obtain the target cluster center of each cluster, where the target cluster center is the cluster center of the last ant corresponding to the maximum number of iterations.
5. The method for constructing a railway freight customer profile according to claim 4, characterized in that: In step S213, the path selection probability of each path is calculated as follows: in, Is the heuristic information, indicating the heuristic value from node i to node j, which is the inverse of the distance from node i to node j. is the pheromone concentration from node i to node j at the tth iteration, is the set of next nodes that the ant is allowed to choose at the current node i, and It is the parameter that regulates pheromones and heuristic information.
6. The method for constructing a railway freight customer profile according to claim 4, characterized in that: In step S216, the calculation formula of the path metric value is: in, is the total number of clusters, is the total number of samples, V is the total number of attribute class demand information, For samples and cluster centers The Euclidean distance between .
7. The method for constructing a railway freight customer profile according to any one of claims 4 to 6, characterized in that: The step S220 includes: Step S221: Based on the total number of clusters K and the target cluster center of each cluster, an initialized membership matrix is obtained using a uniform distribution method; Step S222: Based on the first transportation data and the membership matrix updated in the last iteration, a second preset formula is used to obtain the cluster center of each cluster updated in the current iteration. In the first iteration, the membership matrix updated in the last iteration is the initialized membership matrix, and the second preset formula is: Among them, C(k, v) represents the cluster center of the k-th cluster in the v-th attribute class demand information, X(i, v) is the v-th attribute class demand information of the i-th customer, is the degree of membership of the i-th customer to the k-th cluster, m is the fuzzy coefficient, and N is the total number of samples; Step S223: Based on the cluster center of each cluster updated in the current iteration, a membership degree calculation formula is used to obtain the new membership degree of each railway freight customer to each cluster, and the membership degree matrix updated in the current iteration is obtained. The membership degree calculation formula is: in, is a data point and cluster centers the distance between them; is a data point and cluster centers The distance between represents the cluster center of the jth cluster in the vth attribute class demand information; m is the fuzzy coefficient, K is the total number of clusters; Step S224: Subtract the membership matrix updated in the current iteration from the membership matrix updated in the previous iteration. When the absolute value of the maximum difference obtained by the subtraction is greater than or equal to the preset convergence precision, step S222 is iteratively executed; when the absolute value of the maximum difference obtained by the subtraction is less than the preset convergence precision or the number of iterations reaches the preset maximum number of iterations, step S225 is executed. Step S225: Based on the membership matrix updated in the current iteration and the cluster center of each cluster updated in the current iteration, cluster information is obtained, where the cluster information includes a clustering result with K clusters, a Calinski-Harabasz Index, a Davies-Bouldin Index, and a Silhouette Coefficient.
8. A railway freight customer portrait construction device, characterized in that: The construction device comprises: A first acquisition module is configured to acquire transportation data of a plurality of railway freight customers, the transportation data including first transportation data corresponding to attribute-type demand information and second transportation data corresponding to transport-type demand information, wherein the attribute-type demand information includes transportation category, transportation volume, cargo category, transportation region, and transportation mode, and the transportation-type demand information includes freight type requirements, transportation distance requirements, transportation speed requirements, transportation punctuality requirements, delivery and pickup requirements, normal temperature warehousing requirements, circulation processing requirements, loading and unloading requirements, business acceptance requirements, payment method requirements, information consultation requirements, and insurance consultation requirements; a second acquisition module, configured to apply an ACO-FCM algorithm to perform multiple clustering processes on the first transportation data to obtain a plurality of clustering information, wherein the total number of clusters K corresponding to each clustering process is different, and each total number of clusters corresponds to a piece of clustering information, K is any integer from 2 to 20, wherein each piece of clustering information includes a clustering result having K clusters, a Calinski-Harabasz Index, a Davies-Bouldin Index, and a Silhouette Coefficient; a determination module, configured to determine an optimal total number of clusters M from a plurality of different total numbers of clusters using an elbow method based on the Calinski-Harabasz Index, the Davies-Bouldin Index, and the Silhouette Coefficient in the plurality of clustering information, and use the clustering result corresponding to the optimal total number of clusters M as a target clustering result, wherein the target clustering result includes M target clusters, each of the target clusters having different label features; A construction module is used to perform chi-square analysis on the second transportation data corresponding to each target cluster, obtain the demand preference data of each target cluster, combine the label features and demand preference data of each target cluster, and generate a railway freight customer portrait.
9. A computer device, characterized in that: The computer device includes: a processor and a memory, wherein a computer program is stored in the memory. When the processor executes the computer program, the computer device implements the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that When the instructions in the computer-readable storage medium are executed by a processor, the processor performs the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
ACO-FCM clustering algorithm-based road network partition and evaluation method
CN110610186A
Product recommendation method and device, computer equipment and storage medium
CN117436968A