An enterprise customer acquisition list sorting method and system based on an association map
Through the enterprise customer list sorting method based on the correlation map, the graph convolution neural network algorithm is used to sort the enterprise customer list, which solves the problems of insufficient data mining and poor interpretability in the existing technology, and achieves more efficient marketing efficiency and precise marketing effects.
Patent Information
- Application Number
- CN202210615500.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-31
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-05-31
AI Technical Summary
In the enterprise development scenario, the graph calculation method based on statistical behavior relies on manual extraction of features and insufficient data mining; while the graph learning algorithm represented by graph convolutional neural networks has strong data mining capabilities, its learning results are poorly interpretable and rely on a large number of successful samples.
A method of sorting enterprise customer list based on correlation graphs is proposed. By obtaining customer information and relationship information, a heterogeneous graph is constructed, and a graph convolution neural network algorithm is used to diffuse and reunite the positive sample node encoding, calculate the similarity and number of important relationships between each node and the positive sample, and perform weighted scoring to sort customer information.
It improves the ability to mine graph data and interpretability of results, helps account managers to automatically analyze the close relationship between target companies and banks, improve marketing efficiency, and achieve precise marketing.
Smart Images

Figure CN114997919B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular, to a method and system for sorting enterprise customer acquisition lists based on an association graph. Background Art
[0002] Customer expansion has always been an area that the corporate business of banks attaches importance to. An important link in the current corporate marketing system of the banking industry is the acquisition and screening of customer acquisition lists. Since the number of all customers in the market is extremely large and scattered, a major pain point is that it is difficult for the customer managers of branch banks to grasp the marketing focus after receiving the "flood irrigation" marketing lists, losing many opportunities to expand the market. Using a graph can analyze the complex relationships between enterprises and banks, and then analyze the success rate of enterprise marketing customer acquisition.
[0003] The following existing technologies exist for enterprise customer acquisition:
[0004] Technical solution of the first prior art:
[0005] Adopting a graph calculation method based on statistical behavior is a currently commonly used algorithm for analyzing the success rate of enterprise marketing customer acquisition. By calculating graph metrics such as node centrality and pageRank to obtain the characteristics of each enterprise node, and then using machine learning algorithms to score the customer acquisition success index.
[0006] Disadvantages of the first prior art:
[0007] Although this technical method has a certain degree of interpretability, the relationship network of the enterprise-level graph is very complex, and some important relationships are often hidden in deep associations. The traditional graph calculation method based on statistical behavior is limited by the deviation and ability of manual feature extraction, and the in-depth mining of data is insufficient. And graph learning algorithms represented by graph convolutional neural networks automatically extract graph structure features during the machine learning process using graph convolution, and have stronger result prediction ability than the graph calculation method based on statistical behavior.
[0008] Technical solution of the second prior art:
[0009] Graph learning algorithms represented by graph convolutional neural networks are also one of the currently commonly used technical solutions. Through graph convolution operators, aggregate the feature of neighbor nodes of the target node in the graph, and then use a neural network model to learn data. In the scenario of enterprise marketing customer acquisition, this technical solution can better learn the association characteristics of enterprises that have successfully acquired customers, so as to predict the success rate of customer acquisition for other target customer enterprises.
[0010] Disadvantages of the second prior art:
[0011] Although graph learning algorithms represented by graph convolutional neural networks have strong data mining and learning capabilities, they also rely on the number of learning samples. In the enterprise marketing and customer acquisition scenario, the number of enterprises that have successfully acquired customers as positive samples is not large, and the learning ability of graph learning algorithms will also be affected to a certain extent. In addition, for problems such as predicting the success index of enterprise marketing and customer acquisition, the poor interpretability of learning results is also a shortcoming of graph learning algorithms represented by graph convolutional neural networks. Summary of the Invention
[0012] The purpose of the present invention is to overcome the problems existing in the above-mentioned prior art that traditional graph recommendation models based on statistical behavior seriously rely on manual extraction of node features and have problems such as inability to deeply mine data. Although graph deep learning represented by graph neural networks has strong graph data mining and learning capabilities, its interpretability is inferior to that of traditional graph recommendation models based on statistical behavior. Therefore, a method and system for ranking enterprise customer acquisition lists based on association graphs are provided to improve graph data mining capabilities and the interpretability of results.
[0013] The purpose of the present invention can be achieved through the following technical solutions:
[0014] A method for ranking enterprise customer acquisition lists based on association graphs includes the following steps:
[0015] Step 1: Obtain customer information and relationship information between multiple customers, construct a heterogeneous graph with customer information as nodes and relationships as edges;
[0016] Step 2: Initialize the encoding of each node and the edge weight in the heterogeneous graph;
[0017] Step 3: Use the customer information of customers who have successfully acquired customers in the customer information as positive samples;
[0018] Step 4: Use the graph convolutional neural network algorithm to minimize the L2 norm difference of the positive sample node encoding after two-degree propagation in the graph, perform diffusion and reunion of the positive sample node encoding, and then update the encoding of all nodes in the graph;
[0019] Step 5: Use the distance between the node encoding and the positive sample node encoding in the vector space as the similarity, and calculate the similarity between each node and the positive sample;
[0020] Step 6: Calculate the number of important relationships of each customer in the graph, and the important relationships are selected in advance from the relationship information;
[0021] Step 7: Perform weighted scoring according to the similarity between each node and the positive sample and the number of important relationships obtained in Step 5 and Step 6 respectively to obtain the score of each node, and rank the customer information according to this score for the possibility of successful customer acquisition.
[0022] Furthermore, in step 4, the calculation expression of the graph convolutional neural network algorithm is as follows:
[0023]
[0024] In the formula, in a graph containing N nodes, X ∈ R^(N×C) represents the matrix composed of the feature vectors of the nodes, where C represents the feature dimension of each node. represents the adjacency matrix after degree matrix normalization, and W (0) and W (1) represent the network weights of the first and second layers of the neural network respectively, and Z represents the reconstructed result of the output node encoding.
[0025] Furthermore, in step 4, the expression of the loss function in the training process of the graph convolutional neural network algorithm is as follows:
[0026]
[0027] In the formula, i represents the i-th positive sample, the number of positive samples is n, and y i is the reconstructed result of the encoding of the i-th positive sample. is the mean value of the reconstructed encodings of all positive samples in the vector space.
[0028] Furthermore, the calculation expression of the similarity is as follows:
[0029]
[0030] In the formula, represents the distance from the node encoding of the target customer to the mean value of the node encodings of the positive samples in the vector space. is a normalization function. std(x) represent the reciprocal mean and standard deviation of the distances between the node encodings of all target customers and the mean value of the positive sample encodings in the vector space respectively.
[0031] Furthermore, the customer information includes in-bank target customers, in-bank non-target customers, out-of-bank target customers, and in-bank individual customers.
[0032] Furthermore, the relationships include executive appointments, guarantees, fund flows, trade, and industrial chains.
[0033] Furthermore, the edge weights corresponding to each relationship are set according to the importance of the relationship.
[0034] Furthermore, the important relationships include the fund flow relationship and the executive appointment relationship.
[0035] Furthermore, the node encoding is a multi-dimensional vector during initialization.
[0036] The present invention also provides an enterprise customer acquisition list ranking system based on an association graph, including a memory and a processor. The memory stores a computer program, and the processor calls the computer program to execute the steps of the method described above.
[0037] Compared with the prior art, the present invention has the following advantages:
[0038] (1) The present invention proposes a scoring model based on graph computing and graph convolutional neural network technology for the corporate marketing customer acquisition scenario in the banking industry, ranks the corporate marketing list of the bank from the perspective of the association relationship between the enterprise and the bank, and preferentially finds enterprises with a higher customer acquisition success index.
[0039] (2) The method of the present invention also sorts out important association relationships, and uses graph convolutional neural network to deeply mine the association degree between the target customer and the positive sample of the successfully acquired customers. Compared with the problem of low marketing efficiency caused by the previous method of the bank using the 'large marketing list' and handing it over to the branch bank customer manager for customer acquisition, the method of the present invention uses the graph to help the customer manager automatically analyze the closeness of the relationship between the target enterprise customer and the bank, grasp the marketing focus, improve the efficiency, and achieve precise marketing.
[0040] (3) The method proposed by the present invention combines the respective advantages of graph learning and graph computing. It can not only deeply mine graph structure data and learn the similarity between the target enterprise customer node and the successfully acquired enterprise, but also combine some graph computing results to score together and predict the customer acquisition success index of the target enterprise. The method of the present invention has good interpretability, that is, it is hoped that the target customers with higher rankings show strong associations with the positive samples (customers who have successfully acquired customers) in terms of important relationships such as executive appointments and guarantees on the graph, and it is also hoped that these target customers show the characteristics of a high number of associations with the executive customers in the bank and a large amount of capital transactions with the enterprises in the bank. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 It is the first process schematic diagram of a method for ranking an enterprise customer acquisition list based on an association graph provided in an embodiment of the present invention;
[0042] Figure 2 It is the second process schematic diagram of a method for ranking an enterprise customer acquisition list based on an association graph provided in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.
[0044] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0045] It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0046] Embodiment 1
[0047] As Figure 1 shown, this embodiment provides a method for ranking enterprise customer acquisition lists based on an association graph, including the following steps:
[0048] Step 1: Obtain customer information and relationship information between multiple customers, and construct a heterogeneous graph with customer information as nodes and relationships as edges.
[0049] Step 2: Initialize the node encodings and edge weights in the heterogeneous graph.
[0050] Step 3: Use the customer information of customers who have successfully acquired customers in the customer information as positive samples.
[0051] Step 4: Use the graph convolutional neural network algorithm with a double convolutional layer to minimize the L2 norm difference of the positive sample node encodings after two-degree propagation in the graph, perform diffusion and reunion of the positive sample node encodings, and then update all the node encodings in the graph.
[0052] Step 5: Use the distance between the node encoding and the positive sample node encoding in the vector space as the similarity, and calculate the similarity between each node and the positive sample.
[0053] Step 6: Calculate the number of important relationships of each customer in the graph. The important relationships are selected from the relationship information in advance.
[0054] Step 7: Based on the similarity between each node obtained in Steps 5 and 6 and the positive samples, and the number of important relationships, perform weighted scoring to obtain the score of each node. According to this score, rank the likelihood of successful customer acquisition for customer information.
[0055] In this embodiment, the customer node categories include in-bank target customers, in-bank non-target customers, out-of-bank target customers, and in-bank individual customers. The relationships include executive positions, guarantees, fund flows, trade, and industrial chains. The important relationships include fund flow relationships and executive position relationships. Edge weights corresponding to each relationship are set according to the importance of the relationship.
[0056] In this embodiment, weighted scoring is performed based on three indicators: the similarity between the out-of-bank target customer node and the positive sample, the number of fund relationships with in-bank corporate customers, and the number of executive relationships with in-bank individual customers. Based on the comparison of the scored values, the likelihood of successful customer acquisition for out-of-bank target customers is ranked, and corporate target customers with higher scores are pushed to bank account managers.
[0057] As Figure 2 shown, the specific implementation process of this embodiment is described in detail below:
[0058] 1. Starting from 37,692 real enterprise target customers, through relationships such as letters of guarantee, guarantees, investments, executives, partners, groups, industrial chains, factoring, bank drafts, commercial drafts, fund inflows, and outflows that can be obtained within the bank, in-bank and out-of-bank customers are associated to construct a heterogeneous graph spectrum with a total of 227,738 enterprise and individual nodes and 1,574,752 relationships.
[0059] 2. Using the graph spectrum association, calculate the number of relationships between each node and in-bank executive customers, and the number of fund associations with in-bank enterprises as two indicators for each node.
[0060] 3. Match the node list with the in-bank information customer table to obtain 816 enterprises that have opened accounts as positive sample nodes.
[0061] 4. Set the initial encodings for the four types of nodes:
[0062] Out-of-bank target customers: [0, 0, 0, 0, 0];
[0063] In-bank target customers (positive samples): [1, 1, 1, 1, 1];
[0064] In-bank non-target customers: [1, 0, 0, 0, 1];
[0065] In-bank individual customers: [1, 0, 1, 0, 0];
[0066] The constructed node encodings are all a set of five-dimensional vectors in the vector space, serving as the node feature input for the graph convolutional neural network model.
[0067] 5. Set the edge weights of the graph convolution:
[0068] Letter of guarantee, guarantee, investment relationship - 0.75 (weight);
[0069] Executive, partner, group relationship - 0.95 (weight);
[0070] Industrial chain, factoring, bank note, commercial paper relationship - 0.2 (weight);
[0071] Fund inflow and outflow relationship - 0.08 (weight);
[0072] Self - loop relationship - 1.5 (weight);
[0073] For relatively important relationships, such as executive appointment and guarantee, the features of positive samples are more likely to be transmitted to associated nodes during graph convolution. Conversely, there will be a significant loss when the feature values of positive samples are propagated through relationships such as fund flow.
[0074] 6. The forward propagation expression of the graph convolution neural network algorithm with a double - convolution layer is:
[0075]
[0076] In the formula, in a graph containing N nodes, \(X\in R^{N\times C}\) represents the matrix composed of the feature vectors of the nodes, where C represents the feature dimension of each node. represents the adjacency matrix after degree - matrix normalization. The purpose of normalization is to suppress the possible gradient explosion / vanishing problems in the network. \(W\) (0) and \(W\) (1) represent the network weights of the first and second - layer neural networks respectively. \(Z\) represents the reconstructed result of the output node encoding.
[0077] Back - propagation to solve for \(W\) (0) and \(W\) (1) The loss function is expressed as follows:
[0078]
[0079] In the formula, \(i\) represents the \(i\) - th positive sample, the number of positive samples is \(n\), \(y\) i is the reconstructed encoding result of the \(i\) - th positive sample, is the mean value of the reconstructed encodings of all positive samples in the vector space.
[0080] 7. Obtain the similarity between each node and the positive sample set. The similarity is calculated as follows:
[0081]
[0082] In the formula, It represents the distance from the node encoding of the target customer to the mean of the positive sample node encodings in the vector space. is a normalization function. std(x) respectively represent the reciprocal mean and standard deviation of the distances between the node encodings of all target customers and the mean of the positive sample encodings in the vector space. The closer the vector space distance between the node and the mean of the positive sample set, the higher the similarity.
[0083] 8. For each node, the weights of the weighted sum of the similarity with the positive sample set, the number of relationships with in-house executive customers, and the number of in-house enterprise fund associations are set to 0.5, 0.3, and 0.2 respectively. Based on this weight, a score is given to evaluate the success index of the target enterprise's customer acquisition.
[0084] This scoring system first hopes that the target customers with higher rankings show strong associations with important relationships such as executive appointments and guarantees with the positive samples (customers with successful customer acquisition) on the graph, and also hopes that these target customers show the characteristics of a high number of associations with in-house executive customers and a large number of financial transactions with in-house enterprises.
[0085] This embodiment also provides an enterprise customer acquisition list sorting system based on an association graph, including a memory and a processor. The memory stores a computer program, and the processor calls the computer program to execute the steps of the above-mentioned enterprise customer acquisition list sorting method based on the association graph.
[0086] The above has described in detail the preferred specific embodiments of the present invention. It should be understood that those of ordinary skill in the art can make many modifications and variations according to the concept of the present invention without creative labor. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field based on the concept of the present invention through logical analysis, reasoning, or limited experiments on the basis of the prior art should be within the protection scope determined by the claims.
Claims
1. An enterprise customer acquisition list sorting method based on an association graph, characterized in that, it includes the following steps: Step 1: Obtain customer information and the relationship information between multiple customers. Using the customer information as nodes and the relationships as edges, construct a heterogeneous graph; Step 2: Initialize the node encodings and edge weights in the heterogeneous graph; Step 3: Use the customer information of the customers who have successfully acquired customers in the customer information as positive samples; Step 4: Use the graph convolutional neural network algorithm to minimize the L2 norm difference of the positive sample node encodings after two-degree propagation in the graph, perform the diffusion and reunion of the positive sample node encodings, and then update all the node encodings in the graph; Step 5: Use the distance between the node encoding and the positive sample node encoding in the vector space as the similarity, and calculate the similarity between each node and the positive sample; Step 6: Calculate the number of important relationships of each customer in the graph, and the important relationships are pre-selected from the relationship information; Step 7: Perform weighted scoring according to the similarity and the number of important relationships obtained in Step 5 and Step 6 respectively for each node and the positive sample, obtain the score of each node, and sort the customer information according to the likelihood of successful customer acquisition based on this score; The customer information includes in-bank target customers, in-bank non-target customers, out-of-bank target customers, and in-bank individual customers; The relationships include executive appointments, guarantees, fund flows, trade, and industrial chains; Set the edge weights corresponding to each relationship according to the importance of the relationship; The important relationships include fund flow relationships and executive appointment relationships; When initializing, the node encoding is a multi-dimensional vector; In Step 4, the calculation expression of the graph convolutional neural network algorithm is: Wherein, in a graph containing N nodes, X ∈ R^(N×C) represents a matrix composed of the feature vectors of the nodes, where C represents the feature dimension of each node. represents the adjacency matrix after degree matrix normalization, and W (0) , W (1) respectively represent the network weights of the first and second layer neural networks, and Z represents the reconstructed result of the output node encoding. In Step 4, the expression of the loss function in the training process of the graph convolutional neural network algorithm is: where \(i\) represents the \(i\)-th positive sample, the number of positive samples is \(n\), and \(y\) i is the coding reconstruction result of the \(i\)-th positive sample, is the mean value of the reconstructed codes of all positive samples in the vector space; The calculation expression of the similarity is: In the formula, represents the distance from the node encoding of the target customer to the mean of the positive sample node encodings in the vector space, is a normalization function, std(x) respectively represent the reciprocal mean and standard deviation of the distances between the node encodings of all target customers and the mean of the positive sample encodings in the vector space.
2. An enterprise customer acquisition list sorting system based on an association graph, characterized in that, it includes a memory and a processor. The memory stores a computer program, and the processor calls the computer program to execute the steps of the method according to Claim 1.
Citation Information
Patent Citations
Hybrid recommendation method based on graph convolutional neural network
CN110674407A
Target identification method and device based on artificial intelligence, electronic equipment and medium
CN113901236A