Knowledge graph-based group customer identification management and control method and system
By building an enterprise relationship map based on knowledge graph and machine learning methods, automatically generate optimal query rules, identify and layer-based management of enterprise relationships, solving the problem of inefficiency in traditional methods when identifying company relationships, and achieving more efficient risk assessment and compliance management.
Patent Information
- Application Number
- CN202510381504.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-11
AI Technical Summary
Traditional methods are inefficient and intuitive in identifying relationships between companies, especially when dealing with large corporate groups and cross-industry businesses, which are difficult to meet the bank's risk management and compliance supervision needs.
The relationship map between enterprises is constructed based on the knowledge graph, and the optimal graph database query statement is automatically generated using machine learning technology. The relationship relationship between enterprises is identified through the comprehensive correlation strength scoring formula, and corresponding control rules are configured.
It improves the efficiency and accuracy of identifying corporate relationships, can better grasp the relationship status between customers, reduce the cost of manual intervention and investigation, and improve the efficiency of risk assessment and compliance management.
Smart Images

Figure CN120298100A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of relationship recognition and control, and specifically, to a method and system for identifying and controlling group customers based on a knowledge graph. Background Art
[0002] In the current corporate credit business of banks, identifying the associated relationships between companies is crucial for risk assessment and compliance management. Through effective risk assessment, banks can promptly detect loan fraud behaviors by companies manipulated by loan intermediaries and prevent the occurrence of pre-loan risks; from the perspective of compliance management, it can help banks identify group customer relationships that should be established but have not yet been established, effectively prevent the problem of unauthorized approvals by enterprises, and ensure compliance.
[0003] Traditional bank investigation methods mainly rely on manual investigation or querying multiple relationship tables through SQL to identify associated relationships. These methods are not only time-consuming, laborious, and resource-consuming, but also often difficult to discover potential associations when facing complex multi-level relationship networks. Therefore, traditional methods have certain limitations, especially in dealing with complex scenarios such as large enterprise groups and cross-industry businesses, where the identification efficiency and accuracy are greatly reduced, and they cannot meet the needs of banks in modern risk management and compliance supervision.
[0004] To solve these problems, the group customer identification method based on a knowledge graph has significant advantages. A knowledge graph can organize various relationships between enterprises in the form of a graph structure, thereby revealing potential multi-level and complex associated networks between enterprises. This method can not only more accurately identify the actual relationships between companies, but also help banks improve efficiency in risk assessment and compliance inspections, reduce manual intervention and investigation costs. By constructing a relationship graph between enterprises, banks can better master the associated status of customers and provide a solid foundation for effective credit decision-making and compliance supervision. Therefore, the present invention provides a method and system for identifying and controlling group customers based on a knowledge graph. Summary of the Invention
[0005] In view of the problems in the related art, the present invention proposes a method and system for identifying and controlling group customers based on a knowledge graph to overcome the above-mentioned technical problems existing in the existing related technologies.
[0006] To this end, the specific technical solutions adopted by the present invention are as follows:
[0007] According to one aspect of the present invention, there is provided a method for identifying and controlling group customers based on a knowledge graph, including the following steps:
[0008] S1. Based on the basic data, MAC address data, and relationship data of enterprises, construct a relationship graph between enterprises, where the relationship graph includes an equity relationship graph, a fiduciary payment relationship graph, a login device relationship graph, a business cooperation relationship graph, and a legal litigation relationship graph;
[0009] S2. Use machine learning technology to automatically generate an optimal graph database query statement, and convert the optimal graph database query statement into a query rule for the optimal graph database; identify the association relationships between enterprises based on the query rules of the optimal graph database;
[0010] S3. Calculate the comprehensive association strength between enterprises according to the comprehensive association strength scoring formula, and determine the association level of the association strength between enterprises based on the comparison result of the comprehensive association strength and the preset association strength threshold; configure corresponding control rules for control based on the association level of the association strength between enterprises.
[0011] Further, the construction of the relationship graph between enterprises based on the basic data, MAC address data, and relationship data of enterprises includes the following steps:
[0012] S11. Obtain the basic data and MAC address data of enterprises, where the basic data includes enterprise name, unified social credit code, customer identity information, registered capital, industry, customer number, and affiliated branch;
[0013] S12. Obtain relationship data, including the relationships between enterprises and the relationships between enterprises and MAC addresses, and organize the relationship data into the data format input to the graph database, where the data format input to the graph database is node-relationship-node;
[0014] S13. Based on graph database technology, combine the company node data, MAC address node data, and relationship data to construct a relationship graph between enterprises, where the relationship graph includes an equity relationship graph, a fiduciary payment relationship graph, a login device relationship graph, a business cooperation relationship graph, and a legal litigation relationship graph.
[0015] Further, the nodes in the data format input to the graph database represent different companies or MAC addresses, the relationships represent the equity relationship, fiduciary payment relationship, login device relationship, business cooperation relationship, and legal litigation relationship between enterprises, and the relationship layer number represents the penetration layer number of the relationship between two nodes.
[0016] Further, when constructing the relationship graph between enterprises, it also includes defining relationship labels for each relationship, where the definition of relationship labels for each relationship includes:
[0017] Define the label of equity relationship as GQ, the label of entrusted payment relationship as PAY, the login device relationship as LOGIN, the business cooperation relationship as COOP, and the legal litigation relationship as LIT.
[0018] Further, in the equity relationship graph, the node attribute is the enterprise name, and the relationship attribute is the shareholding ratio;
[0019] In the entrusted payment relationship graph, the node attribute is the enterprise name, and the relationship attribute is the entrusted payment amount;
[0020] In the login device relationship graph, the node attribute is the enterprise name and the MAC address information, and the relationship attribute is the login time;
[0021] In the business cooperation relationship graph, the node attribute is the enterprise name, and the relationship attributes are the cooperation amount and the cooperation time;
[0022] In the legal litigation relationship graph, the node attribute is the enterprise name, and the relationship attributes are the case amount and the case time.
[0023] Further, the method of using machine learning technology to automatically generate the optimal graph database query statement and convert the optimal graph database query statement into the query rules of the optimal graph database; identifying the association relationship between enterprises based on the query rules of the optimal graph database includes the following steps:
[0024] S21. Obtain the multi-layer paths between enterprises from the database, count the feature data of each path, and label the associated paths in the historical risk events as high-risk paths;
[0025] S22. Use the machine learning algorithm based on decision tree to model and train the path features, and based on the trained decision tree model, automatically generate the optimal graph database query statement in combination with the input risk scenario;
[0026] S23. Convert the optimal graph database query statement into the query rules of the optimal graph database, and identify the association relationship between enterprises according to the query rules of the optimal graph database.
[0027] Further, the multi-layer paths between enterprises include the combined paths of equity relationship, payment relationship, device relationship, cooperation relationship and litigation relationship, and the feature data includes enterprise nodes, relationship attributes and path length.
[0028] Further, the comprehensive association strength scoring formula is:
[0029] C 综合 = α1·C 股权 + α2·C 受托支付 + α3·C 设备关系 + α4·C 商业合作+α5·C 法律诉讼
[0030] In the formula, C 综合 represents the comprehensive correlation strength, C 股权 represents the equity strength, C 受托支付 represents the entrusted payment strength, C 设备关系 represents the equipment relationship strength, C 商业合作 represents the business cooperation strength, C 法律诉讼 represents the legal litigation strength, and α1, α2, α3, α4, α5 respectively represent the weights of the equity strength, entrusted payment strength, equipment relationship strength, business cooperation strength, and legal litigation strength.
[0031] Further, the association levels for determining the association strength between enterprises according to the comparison result between the comprehensive correlation strength and the preset correlation strength threshold include:
[0032] When the comprehensive correlation strength is greater than the first correlation strength threshold, it is determined as the first association level;
[0033] When the comprehensive correlation strength is greater than or equal to the second correlation strength threshold and less than or equal to the first correlation strength threshold, it is determined as the second association level;
[0034] When the comprehensive correlation strength is less than the second correlation strength threshold, it is determined as the third association level;
[0035] Among them, the first correlation strength threshold is greater than the second correlation strength threshold, and the association strengths of the first association level, the second association level, and the third association level decrease in turn.
[0036] According to another aspect of the present invention, a group customer identification and control system based on a knowledge graph is provided. The group customer identification and control system based on the knowledge graph includes a relationship graph construction module, an association relationship identification module, and an enterprise control module;
[0037] Among them, the relationship graph construction module is used to construct a relationship graph between enterprises based on the basic data, MAC address data, and relationship data of the enterprises. Among them, the relationship graph includes an equity relationship graph, an entrusted payment relationship graph, a logged-in equipment relationship graph, a business cooperation relationship graph, and a legal litigation relationship graph;
[0038] The association relationship identification module is used to automatically generate an optimal graph database query statement by using machine learning technology, and convert the optimal graph database query statement into a query rule of the optimal graph database; identify the association relationship between enterprises based on the query rule of the optimal graph database;
[0039] The enterprise control module is used to calculate the comprehensive association strength between enterprises according to the comprehensive association strength scoring formula, and determine the association level of the association strength between enterprises according to the comparison result between the comprehensive association strength and the preset association strength threshold; based on the association level of the association strength between enterprises, configure corresponding control rules for control.
[0040] The beneficial effects of the present invention are as follows:
[0041] 1) The present invention can automatically generate the optimal graph database query statement by using machine learning technology, effectively mine and analyze multi-layer equity relationships, entrusted payment relationships, login device relationships, business cooperation relationships, and legal litigation relationships, thereby revealing potential business groups or risk groups, and effectively solving the problem that the traditional data analysis method relying on tables is often inefficient and not intuitive enough when dealing with complex and multi-layer equity relationships.
[0042] 2) The present invention constructs a graph through five relationships, identifies the association relationships between enterprises from multiple aspects, and proposes a method for determining the strength of the association relationship, stratifying the association relationships between enterprises, and configuring different control rules for associated enterprises with different strengths; in addition, the present invention can use graph computing to identify the association relationships between enterprises, and can be effectively applied at both the risk identification and compliance management levels. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0044] Figure 1 is a flowchart of a method for identifying and controlling group customers based on a knowledge graph according to an embodiment of the present invention;
[0045] Figure 2 is an equity relationship graph in a method for identifying and controlling group customers based on a knowledge graph according to an embodiment of the present invention;
[0046] Figure 3 is a entrusted payment relationship graph in a method for identifying and controlling group customers based on a knowledge graph according to an embodiment of the present invention;
[0047] Figure 4 is a login device relationship graph in a method for identifying and controlling group customers based on a knowledge graph according to an embodiment of the present invention;
[0048] Figure 5It is a business cooperation relationship diagram in a method for identifying and controlling group customers based on a knowledge graph according to an embodiment of the present invention;
[0049] Figure 6 It is a legal litigation relationship diagram in a method for identifying and controlling group customers based on a knowledge graph according to an embodiment of the present invention. Specific embodiments
[0050] To further illustrate the embodiments, the present invention provides accompanying drawings, which are part of the disclosure of the present invention. They are mainly used to illustrate the embodiments and can be combined with the relevant descriptions in the specification to explain the operation principle of the embodiments. With reference to these contents, those of ordinary skill in the art should be able to understand other possible implementation manners and the advantages of the present invention. The components in the drawings are not drawn to scale, and similar component symbols are usually used to represent similar components.
[0051] According to an embodiment of the present invention, a method and system for identifying and controlling group customers based on a knowledge graph are provided.
[0052] Now, the present invention will be further described in conjunction with the accompanying drawings and specific embodiments. As Figures 1-6 shown, according to an embodiment of the present invention, a method for identifying and controlling group customers based on a knowledge graph is provided, including the following steps:
[0053] S1. Based on the basic data, MAC address data, and relationship data of enterprises, construct a relationship graph between enterprises, where the relationship graph includes an equity relationship graph, a fiduciary payment relationship graph, a logged-in device relationship graph, a business cooperation relationship graph, and a legal litigation relationship graph;
[0054] Specifically, the constructing of the relationship graph between enterprises based on the basic data, MAC address data, and relationship data of enterprises includes the following steps:
[0055] S11. Obtain the basic data and MAC address data of enterprises, where the basic data includes enterprise name, unified social credit code, customer identity information (i.e., whether it is a customer of our bank), registered capital, industry, customer number, and affiliated branch, etc.;
[0056] S12. Obtain relationship data, including the relationships between enterprises and the relationships between enterprises and MAC addresses, and organize these relationship data into the data format for input to the graph database, where the data format for input to the graph database is node-relationship-node;
[0057] Specifically, a node can be the company itself or a certain MAC address, and different nodes represent different companies or MAC addresses.
[0058] The relationship can be the relationship between companies, namely equity relationship, entrusted payment relationship, business cooperation relationship, legal litigation relationship, or the relationship between a company and a MAC address, namely the logged-in device relationship. As long as there is any one of the above five relationships between any two nodes, a connection is made between these two nodes. One relationship corresponds to one connection, and each connection will be marked with the relationship type to distinguish what kind of relationship this connection is. The relationship is directional. For example, for the relationship from A to B, the direction arrow is from A to B.
[0059] The relationship level represents the penetration level of the relationship between two nodes. For example, if there is a relationship between A and B, then the relationship between A and B is at the 1st level. If there is another relationship between B and C, then the relationship between A and C is at the 2nd level.
[0060] Among them, the business cooperation relationship and the legal litigation relationship need to be extracted from the publicly available unstructured text data through NLP. Specifically, for the business cooperation relationship, it is necessary to extract the relationship between enterprises from news reports, enterprise announcements, investment agreements, and bid winning and tendering information. For the legal litigation relationship, it is necessary to extract the relationship between enterprises from court announcements and judgment documents. Use the BERT relationship extraction algorithm to identify statements such as "Enterprise A and Enterprise B reached a cooperation agreement" and "Enterprise A sued Enterprise B", and take the cooperation amount, cooperation time, or case amount and case time as relationship attributes.
[0061] In addition, although the legal litigation relationship appears as an adversarial relationship on the surface, it essentially reflects the cooperation intersection between enterprises, especially when the cooperation relationship ends due to factors such as interest conflicts or contract disputes. By analyzing legal litigation data, it is possible to effectively reveal hidden group associations or risk exposure points.
[0062] S13. Based on the graph database technology, combine the company node data, MAC address node data, and relationship data to construct a relationship map between enterprises. Among them, the relationship map includes an equity relationship map, an entrusted payment relationship map, a logged-in device relationship map, a business cooperation relationship map, and a legal litigation relationship map.
[0063] Specifically, deploy a graph database (such as the neo4j graph database) on the server, and then import the sorted company node data, MAC address node data, equity relationship data, entrusted payment relationship data, device relationship data, business cooperation relationship data, and legal litigation relationship data into the neo4j graph database through the neo4j_admin tool.
[0064] When importing, it is necessary to define the relationship tags for each type of relationship. In this embodiment, the tag for the equity relationship can be defined as GQ, the tag for the entrusted payment relationship can be defined as PAY, the tag for the login device relationship can be defined as LOGIN, the tag for the business cooperation relationship can be defined as COOP, and the tag for the legal litigation relationship can be defined as LIT.
[0065] S2. Use machine learning technology to automatically generate the optimal graph database query statement, and convert the optimal graph database query statement into the query rules of the optimal graph database; identify the association relationships between enterprises based on the query rules of the optimal graph database;
[0066] Cypher is the query statement of neo4j. By writing Cypher, the desired potential companies can be queried. Different risk scenarios have different rules. For example, corresponding Cypher can be written according to the penetration level of the equity relationship and the shareholding ratio limit of each relationship, so as to query the companies that meet the conditions.
[0067] For the five relationships, the match function is needed to match the relationships between any two companies (a and b) where the number of layers of the relationship (r) satisfies between n and m layers (at least n layers and at most m layers). Then, for the equity, entrusted payment, and legal litigation relationships, the direction of the relationship also needs to be set. If a holds shares in b, a makes an entrusted payment to b, or a sues b, the direction is set from a to b, and vice versa. If it is the login device or business cooperation relationship, the direction of the relationship does not need to be set, and the direction between a and b is set as undirected, that is, without an arrow. Finally, by setting the relationship tags, the corresponding relationships can be queried. For example, by setting the relationship tag as GQ, the equity relationship can be queried. By setting the relationship tag as PAY, the entrusted payment relationship can be queried. By setting the relationship tag as LOGIN, the login device relationship can be queried. By setting the relationship tag as COOP, the business cooperation relationship can be queried. By setting the relationship tag as LIT, the legal litigation relationship can be queried.
[0068] The query will traverse all company nodes and return all nodes and paths that meet the above relationships. Specifically, as long as any company node a satisfies that starting from a, it penetrates n to m layers and then reaches b, and the relationship is GQ or PAY or LOGIN or COOP or LIT, the graph database will return all nodes (including the basic data of each node, such as enterprise name, unified social credit code or MAC address, etc.) and relationships (including the attributes of each relationship, such as shareholding ratio or entrusted payment amount or login time or cooperation amount or case amount) on this path. Then, all paths that meet this condition will also be returned. For example, if the two paths a→c→d→e→b and a→c→f→q→b both meet these conditions, then both of these paths will be returned (including all node data and relationship data on the paths).
[0069] To achieve the above content, the following query basic formulas can be used (for equity or entrusted payment relationships or legal litigation relationships):
[0070] match p=(a)-[r:LABEL*n..m]->(b) return p
[0071] Or (for login device relationships or business cooperation relationships)
[0072] match p=(a)-[r:LABEL*n..m]-(b) return p
[0073] Among them, the match clause is used to find and match the nodes and relationships that meet the requirements in the graph database.
[0074] p: This is a variable name used to store the matched path (path), which contains nodes and relationships.
[0075] =: It means assigning the matched path to the variable p.
[0076] (a): This is a node pattern used to match a node and assign the node to the variable a. In this scenario, a refers to a company, which is the starting company of the equity or entrusted payment or login device relationship.
[0077] -[r:LABEL*n..m]-> or -[r:LABEL*n..m]-
[0078] -[r]: This is a relationship pattern used to match the relationship between nodes and assign the relationship to the variable r.
[0079] :LABEL: Here, LABEL is the label of the relationship, indicating to match a specific type of relationship, connected to the previous r with a colon. If the label of the equity relationship is GQ, the label of the entrusted payment relationship is PAY, the label of the login device relationship is LOGIN. When querying the equity relationship, LABEL can be written as GQ. When querying the entrusted payment relationship, LABEL can be written as PAY. When querying the login relationship, LABEL can be written as LOGIN. When querying the business cooperation relationship, LABEL can be written as COOP. When querying the legal litigation relationship, LABEL can be written as LIT.
[0080] *n..m: This part represents the length range of the matched relationships:
[0081] *n means that there must be at least n layers of such relationships.
[0082] m means that there can be at most m layers of such relationships.
[0083] *n..m together represents that the length of the relationship can be any value between n and m. For example, *2..5 means that at least 2 layers of relationships are matched and at most 5 layers of relationships are matched. For example, for the equity relationship, it represents matching 2 to 5 layers of equity penetration relationships.
[0084] -> or -: The arrow indicates the direction of the relationship. The arrow points from node a to node b, meaning that the relationship between a and b is directional. If it is an equity relationship, it means that a holds shares in b. Without an arrow, it means that the relationship is undirectional, meaning that the relationship between a and b does not consider the direction. If it is a logged-in device relationship, it means that a and b share a device.
[0085] (b): This is another node pattern used to match another node and assign that node to the variable b.
[0086] return p: This statement is used to return the query result. Here, the variable p is returned, which is the matched path. The path p contains all the nodes and relationships that meet the conditions between node a and node b.
[0087] If you want to set some requirements for the shareholding ratio of each layer on the equity penetration path, for example, each layer of the relationship should have at least 20% shareholding to screen out the target company and prevent a large number of irrelevant companies with too low shareholding ratios. Then you can use an advanced query formula.
[0088] Advanced query formula:
[0089] where all(rel in r where rel.weight >= 0.2)
[0090] Among them, where is a filtering condition used to limit the results matched by the match statement. Only the results that meet the where condition will be retained and returned.
[0091] all is a function used to check whether all elements in a set meet a certain condition. It returns true or false.
[0092] rel in r: This part means traversing each element in the set r and assigning each element to the variable rel. In this context, r is the set of relationships defined previously in match (such as r in -[r:LABEL*n..m]->).
[0093] all checks all the relationships rel in the set r and applies the subsequent conditions to them.
[0094] rel.weight: The rel relationship has an attribute called weight. Here, weight is an attribute name. In the equity relationship query, weight is the shareholding ratio between two companies.
[0095] >= 0.2: It means to check whether the value of the weight attribute is greater than or equal to 0.2.
[0096] The condition rel.weight >= 0.2 means that the all function will return true only when the weight attribute value of rel (that is, each relationship in the r set) is greater than or equal to 0.2.
[0097] The basic structures of the five relationships are as shown in Table 1 below:
[0098] Table 1 Basic Structures of the Five Relationships
[0099]
[0100] In the previous group identification technology based on knowledge graphs, the Cypher query rules of the graph database usually rely on manual setting. This method depends on expert experience, is difficult to comprehensively cover diverse risk scenarios, and for different business requirements, it is necessary to manually adjust the query rules, which consumes a lot of time and is inefficient. At the same time, manual rules are difficult to dynamically adapt to changes in business scenarios and may miss hidden or new association relationship patterns. Therefore, in this embodiment, an intelligent rule generation technology based on machine learning is proposed to automatically generate the optimal query rules for specific risk scenarios, thereby improving the accuracy and efficiency of identification. Specifically as follows:
[0101] Using the machine learning technology to automatically generate the optimal graph database query statement and convert the optimal graph database query statement into the query rules of the optimal graph database; identifying the association relationships between enterprises based on the query rules of the optimal graph database includes the following steps:
[0102] S21. Obtain the multi-layer paths between enterprises from the database, count the feature data of each path, and label the association paths in historical risk events as high-risk paths;
[0103] Feature data: Obtain the multi-layer paths between enterprises from the knowledge graph, including the combined paths of equity relationships, entrusted payment relationships, login device relationships, business cooperation relationships, and legal litigation relationships, and count the feature data of each path: enterprise nodes (registered capital, industry), relationship attributes (such as the product of shareholding ratios of each layer, the total amount of entrusted payments of each layer, the total amount of business cooperation of each layer, etc.), path length (including the total length and the length of each relationship).
[0104] Label data: Mark the association paths in historical risk events (such as identified risky enterprises or non-performing loan enterprises) as "high-risk paths", which serve as the training targets for the model.
[0105] S22. Use a machine learning algorithm based on decision trees to model and train the path features. Based on the trained decision tree model, automatically generate the optimal graph database query statement in combination with the input risk scenario;
[0106] Model training: Use a machine learning algorithm based on decision trees to model the path features. The structure of the decision tree is naturally suitable for generating rules. Each path from the root to the leaf of the decision tree can be directly converted into a business rule.
[0107] Training objective: Predict the risk level of the path and mine the pattern features related to the high-risk path.
[0108] Output result: The generated model can dynamically adjust the rule generation according to specific risk scenarios.
[0109] S23. Convert the optimal graph database query statement into the query rules of the optimal graph database, and identify the association relationships between enterprises according to the query rules of the optimal graph database.
[0110] Generation logic: Based on the trained model, extract the pattern features associated with the high-risk path and automatically generate the optimal query rules (Cypher query statements).
[0111] Rule formatting: Convert the pattern features output by the model into graph database query rules. For example:
[0112] Risk path pattern: Penetrate the equity by 3 layers, and the shareholding ratio of each layer ≥ 20%, entrust payment for 1 layer, and the payment frequency ≥ 5 times.
[0113] Convert to query statement:
[0114] MATCH p=(a)-[r:GQ*3]->(b)-[t:PAY]->(c)
[0115] WHERE all(rel in r WHERE rel.weight>=0.2)and t.frequency>=5
[0116] RETURN p
[0117] Among them, the MATCH clause is used to find and match the nodes and relationships that meet the requirements in the graph database;
[0118] p: This is a variable name used to save the matched path, and the path contains nodes and relationships.
[0119] =: means assigning the matched path to the variable p.
[0120] (a): This is a node pattern used to match a node and assign it to the variable a. In this scenario, a refers to a company. Similarly, b and c are similar, each matching a node and referring to a company respectively.
[0121] -[r:GQ*n]-> or -[t:PAY]-
[0122] -[r]: This is a relationship pattern used to match the relationship between nodes and assign the relationship to the variable r.
[0123] *n means there must be at least n layers of such relationships. Here, 3 represents at least 3 layers of relationships. Similarly, t means assigning the relationship to t, representing at least 1 layer of relationship here.
[0124] WHERE all(rel in r WHERE rel.weight >= 0.2)
[0125] Among them, WHERE is a filtering condition used to restrict the results matched by the MATCH statement. Only the results that meet the WHERE condition will be retained and returned.
[0126] all is a function used to check whether all elements in a set meet a certain condition. It returns true or false.
[0127] rel in r: This part means traversing each element in the set r and assigning each element to the variable rel. In this context, r is the set of relationships defined previously in MATCH (such as r in -[r:LABEL*3]->).
[0128] all checks all the relationships rel in the set r and applies the subsequent conditions to them.
[0129] rel.weight: The rel relationship (equity relationship) has an attribute called weight. Here, weight is an attribute name. In the equity relationship query, weight is the shareholding ratio between two companies.
[0130] >= 0.2: means checking whether the value of the weight attribute is greater than or equal to 0.2.
[0131] The condition rel.weight >= 0.2 means that this all function will return true only when the weight attribute value of each rel (that is, each relationship in the r set) is greater than or equal to 0.2.
[0132] t.frequency: The t relationship (entrusted payment relationship) has an attribute named frequency. Here, frequency is an attribute name. In the entrusted payment relationship, frequency is the entrusted payment frequency between two companies.
[0133] t.frequency >= 5: It represents that the entrusted payment frequency is greater than or equal to 5 times.
[0134] return p: This statement is used to return the query result. Here, the variable p is returned, which is the matching path. The path p contains all the nodes and relationships that meet the conditions between node a, passing through node b, to node c.
[0135] S3. Calculate the comprehensive association strength between enterprises according to the comprehensive association strength scoring formula, and determine the association level of the association strength between enterprises based on the comparison result of the comprehensive association strength and the preset association strength threshold; configure corresponding control rules for control based on the association level of the association strength between enterprises.
[0136] Specifically, through calculating the comprehensive association strength, integrating various enterprise relationships, the comprehensiveness and accuracy of enterprise association relationships are further improved. Set the scoring formula for the comprehensive association strength, and calculate through the weighted combination of multi-dimensional features.
[0137] Among them, the comprehensive association strength scoring formula is:
[0138] C 综合 = α1·C 股权 + α2·C 受托支付 + α3·C 设备关系 + α4·C 商业合作 + α5·C 法律诉讼
[0139] In the formula, C 综合 represents the comprehensive association strength;
[0140] C 股权 represents the equity strength, which is calculated according to the product of the shareholding ratios of each layer;
[0141] C 受托支付 represents the entrusted payment strength, which is calculated according to the sum of the payment amounts of each layer × the total payment frequency of each layer;
[0142] C 设备关系 represents the equipment relationship strength, which is calculated according to the minimum value of the login times of each layer;
[0143] C 商业合作 represents the business cooperation strength, which is calculated according to the total cooperation amount of each layer × the total cooperation frequency of each layer;
[0144] C 法律诉讼 Represents the intensity of legal litigation, which is calculated based on the total amount of cases at each level × the total frequency of litigation at each level;
[0145] α1, α2, α3, α4, and α5 respectively represent the weights of equity intensity, entrusted payment intensity, equipment relationship intensity, business cooperation intensity, and legal litigation intensity, which are set according to the business scenario.
[0146] Specifically, through the comprehensive association intensity, the association relationship between enterprises can be stratified. After setting the association relationship threshold, the following stratification can be done:
[0147] Strong association (i.e., the first association level): Comprehensive association intensity > high threshold (i.e., the first association intensity threshold);
[0148] Medium association (i.e., the second association level): Low threshold (i.e., the second association intensity threshold) ≤ comprehensive association intensity ≤ high threshold;
[0149] Weak association (i.e., the third association level): Comprehensive association intensity < low threshold;
[0150] Finally, according to the need, the list of the corresponding association level can be screened out.
[0151] Finally, output the list to the downstream system for risk investigation. Different control rules are configured according to the different association levels of the list.
[0152] 1) The principles for setting control rules according to the comprehensive association intensity are as follows:
[0153] According to the comprehensive association intensity between enterprises, the association relationship is divided into strong association, medium association, and weak association. Differentiated control rules are adopted at different levels. The main principles include:
[0154] · Strong association: Due to the close association relationship, there may be actual control, interest transfer, or concentration risks, which need to be focused on and strict control measures should be taken.
[0155] · Medium association: The association relationship is moderate, and there may be indirect or phased cooperation, which requires appropriate monitoring of potential risks.
[0156] · Weak association: The association is relatively low, the risk is small, but basic rules still need to be set to prevent missed inspections.
[0157] 2) The control rules for different intensities are designed as follows:
[0158] For the comprehensive association intensity at different levels, the set control rules can be designed in the following aspects by combining business requirements and risk levels:
[0159] (1) Strong correlation (Comprehensive correlation strength > high threshold)
[0160] Characteristics: There are significant equity, capital or transaction connections between enterprises, with high frequency and amount, and strong correlation.
[0161] · Key control rules:
[0162] a. Customer merger management: Implement a unified group customer management system for strongly correlated enterprises to avoid potential risk hazards caused by separate credit granting to related enterprises.
[0163] b. Penetrating review: Conduct an equity penetration analysis on all strongly correlated enterprises to identify whether there is an actual controller or concentrated capital flow.
[0164] c. Credit limit merger: Evaluate the credit limits of strongly correlated enterprises together to prevent circumvention of the bank's single customer credit limit through decentralized credit granting.
[0165] d. Transaction behavior monitoring: Focus on monitoring large - amount capital transactions and frequent trading behaviors, and pay attention to abnormal flows.
[0166] e. Risk early warning: Set up a special risk early - warning model and dynamically adjust the risk control strategy according to the comprehensive correlation strength.
[0167] (2) Medium correlation (Low threshold ≤ Comprehensive correlation strength ≤ High threshold)
[0168] Characteristics: There is a certain degree of cooperation or transaction relationship between enterprises, but the intensity is moderate and the correlation is not close enough.
[0169] · Moderate control rules:
[0170] a. Hierarchical management: Classify medium - correlated enterprises into the potential risk customer group, regularly update the comprehensive correlation strength, and prevent risk escalation caused by changes in the correlation relationship.
[0171] b. Strengthened approval process: For business involving medium - correlated enterprises in credit approval, add a "second - level approval" or "risk review" link.
[0172] c. Verification of related behaviors: Conduct appropriate spot checks on the transaction behaviors and capital flows of medium - correlated enterprises, especially for large - amount capital flows or high - frequency trading behaviors.
[0173] d. Credit limit restriction: Set a certain percentage limit (such as 50% of the total limit for a single household) on the total credit limit for medium - correlated enterprises to control the overall risk.
[0174] (3) Weak correlation (Comprehensive correlation strength < low threshold)
[0175] Characteristics: The linkages between enterprises are weak, usually occasional cooperation or indirect relationships, and the risks are relatively low.
[0176] Basic control rules:
[0177] a. Regular assessment: Include weakly associated companies in the scope of regular risk management, regularly assess their comprehensive association strength, and identify whether there is a possibility of upgrading to medium / strong association.
[0178] b. Basic monitoring: Basic monitoring of large transactions is conducted, but there is no need to review daily business activities one by one.
[0179] c. Risk trigger mechanism: A bottom-line risk trigger mechanism is set up to initiate further investigation only when specific conditions are triggered (such as increased correlation intensity or an increase in abnormal transactions).
[0180] According to another embodiment of the present invention, a group customer identification management and control system based on a knowledge graph is provided, and the group customer identification management and control system based on a knowledge graph includes a relationship graph construction module, an association relationship identification module and an enterprise management and control module;
[0181] The relationship map construction module is used to construct a relationship map between enterprises based on the basic data, MAC address data and relationship data of the enterprise, wherein the relationship map includes an equity relationship map, a trustee payment relationship map, a login device relationship map, a business cooperation relationship map and a legal litigation relationship map;
[0182] The association relationship identification module is used to automatically generate an optimal graph database query statement using machine learning technology, and convert the optimal graph database query statement into an optimal graph database query rule; based on the optimal graph database query rule, identify the association relationship between enterprises;
[0183] The enterprise management and control module is used to calculate the comprehensive association strength between enterprises according to the comprehensive association strength scoring formula, and determine the association level of the association strength between enterprises according to the comparison result of the comprehensive association strength and the preset association strength threshold; based on the association level of the association strength between enterprises, configure corresponding management and control rules for management and control.
[0184] To sum up, with the help of the above-mentioned technical scheme of the present invention, the present invention can use machine learning technology to automatically generate optimal graph database query statements, effectively mine and analyze multi-layer equity relationships, entrusted payment relationships, login device relationships, business cooperation relationships, and legal litigation relationships, thereby revealing potential business groups or risk groups, and effectively solving the problem that traditional tabular data analysis methods are often inefficient and not intuitive enough when dealing with complex and multi-level equity relationships.
[0185] Meanwhile, the present invention constructs a graph through five relationships, identifies the association relationships between enterprises from multiple aspects, and proposes a method for determining the strength of the association relationships, stratifies the association relationships between enterprises, and configures different control rules for associated enterprises with different strengths; in addition, the present invention can use graph computing to identify the association relationships between enterprises and can be effectively applied at both the risk identification and compliance management levels.
[0186] The foregoing are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for identifying and controlling group customers based on a knowledge graph, characterized in that It includes the following steps: S1. Based on the basic data, MAC address data and relationship data of enterprises, construct a relationship graph among enterprises. The relationship graph includes an equity relationship graph, a fiduciary payment relationship graph, a login device relationship graph, a business cooperation relationship graph and a legal litigation relationship graph. S2. Use machine learning technology to automatically generate an optimal graph database query statement, and convert the optimal graph database query statement into a query rule of the optimal graph database. Identify the association relationships among enterprises based on the query rules of the optimal graph database. S3. Calculate the comprehensive association strength among enterprises according to the comprehensive association strength scoring formula, and determine the association level of the association strength among enterprises according to the comparison result between the comprehensive association strength and the preset association strength threshold. Configure corresponding control rules for control based on the association level of the association strength among enterprises.
2. The method for identifying and controlling group customers based on a knowledge graph according to claim 1, wherein The construction of the relationship graph among enterprises based on the basic data, MAC address data and relationship data of enterprises includes the following steps: S11. Obtain the basic data and MAC address data of enterprises. The basic data includes enterprise name, unified social credit code, customer identity information, registered capital, industry, customer number and affiliated branch. S12. Obtain relationship data, including the relationships among enterprises and the relationships between enterprises and MAC addresses, and organize the relationship data into the data format input to the graph database. The data format input to the graph database is node-relationship-node. S13. Based on graph database technology, combine the company node data, MAC address node data and relationship data to construct a relationship graph among enterprises. The relationship graph includes an equity relationship graph, a fiduciary payment relationship graph, a login device relationship graph, a business cooperation relationship graph and a legal litigation relationship graph.
3. The group customer identification and control method based on a knowledge graph according to claim 2, wherein In the data format input to the graph database, the nodes represent different companies or MAC addresses, the relationships represent the equity relationship, fiduciary payment relationship, login device relationship, business cooperation relationship and legal litigation relationship among enterprises, and the number of relationship layers represents the penetration layers of the relationship between two nodes.
4. The method for identifying and controlling group customers based on a knowledge graph according to claim 2, wherein When constructing the relationship graph among enterprises, it also includes defining relationship labels for each relationship. Defining relationship labels for each relationship includes: Define the label of the equity relationship as GQ, the label of the fiduciary payment relationship as PAY, the login device relationship as LOGIN, the business cooperation relationship as COOP, and the legal litigation relationship as LIT.
5. The method for identifying and controlling group customers based on a knowledge graph according to claim 2, wherein In the equity relationship graph, the node attribute is the enterprise name, and the relationship attribute is the shareholding ratio. In the fiduciary payment relationship graph, the node attribute is the enterprise name, and the relationship attribute is the fiduciary payment amount. In the login device relationship graph, the node attributes are the enterprise name and MAC address information, and the relationship attribute is the login time. In the business cooperation relationship graph, the node attribute is the enterprise name, and the relationship attributes are the cooperation amount and cooperation time. In the legal litigation relationship graph, the node attribute is the enterprise name, and the relationship attributes are the case amount and case time.
6. The group customer identification and control method based on a knowledge graph according to claim 1, wherein Automatically generating an optimal graph database query statement by using machine learning technology, and converting the optimal graph database query statement into a query rule of the optimal graph database; identifying the association relationship between enterprises based on the query rule of the optimal graph database includes the following steps: S21. Obtain the multi-layer paths between enterprises from the database, count the feature data of each path, and label the associated paths in historical risk events as high-risk paths; S22. Use a machine learning algorithm based on a decision tree to model and train the path features, and automatically generate an optimal graph database query statement based on the trained decision tree model and the input risk scenario; S23. Convert the optimal graph database query statement into a query rule of the optimal graph database, and identify the association relationship between enterprises according to the query rule of the optimal graph database.
7. The group customer identification and control method based on a knowledge graph according to claim 6, characterized in that The multi-layer paths between enterprises include combined paths of equity relationships, payment relationships, equipment relationships, cooperation relationships, and litigation relationships, and the feature data includes enterprise nodes, relationship attributes, and path lengths.
8. The group customer identification and control method based on a knowledge graph according to claim 1, wherein, The comprehensive association strength scoring formula is: C 综合 = α1·C 股权 + α2·C 受托支付 + α3·C 设备关系 + α4·C 商业合作 + α5·C 法律诉讼 Wherein, C 综合 represents the comprehensive correlation strength, C 股权 represents the equity strength, C 受托支付 represents the entrusted payment strength, C 设备关系 represents the equipment relationship strength, C 商业合作 represents the business cooperation strength, C 法律诉讼 represents the legal litigation strength, and α1, α2, α3, α4, and α5 respectively represent the weights of the equity strength, the entrusted payment strength, the equipment relationship strength, the business cooperation strength, and the legal litigation strength.
9. The method for identifying and controlling group customers based on a knowledge graph according to claim 1, characterized in that Determining the association level of the association strength between enterprises according to the comparison result between the comprehensive association strength and the preset association strength threshold includes: When the comprehensive association strength is greater than the first association strength threshold, it is determined as the first association level; When the comprehensive association strength is greater than or equal to the second association strength threshold and less than or equal to the first association strength threshold, it is determined as the second association level; When the comprehensive association strength is less than the second association strength threshold, it is determined as the third association level; Among them, the first association strength threshold is greater than the second association strength threshold, and the association strengths of the first association level, the second association level, and the third association level decrease in turn.
10. A group customer identification and control system based on a knowledge graph, which is used to implement the steps of the group customer identification and control method based on the knowledge graph according to any one of claims 1-9, characterized in that The group customer identification and control system based on the knowledge graph includes a relationship graph construction module, an association relationship identification module, and an enterprise control module; Among them, the relationship graph construction module is used to construct a relationship graph between enterprises based on the basic data, MAC address data, and relationship data of enterprises, where the relationship graph includes an equity relationship graph, a fiduciary payment relationship graph, a login device relationship graph, a business cooperation relationship graph, and a legal litigation relationship graph; The association relationship identification module is used to automatically generate an optimal graph database query statement by using machine learning technology, and convert the optimal graph database query statement into a query rule of the optimal graph database; Identify the association relationship between enterprises based on the query rule of the optimal graph database; The enterprise control module is used to calculate the comprehensive association strength between enterprises according to the comprehensive association strength scoring formula, and determine the association level of the association strength between enterprises according to the comparison result between the comprehensive association strength and the preset association strength threshold; configure corresponding control rules for control based on the association level of the association strength between enterprises.
Citation Information
Cited By
Commercial bank associated party analysis method and device, electronic equipment and storage medium
CN120975213A