Node identification method and device based on knowledge graph, electronic equipment and storage medium
Through the node identification method based on knowledge graph, the edge weights are dynamically adjusted and community division is performed, which solves the accuracy and efficiency problems of suspicious node identification in traditional methods and achieves more efficient suspicious node and group identification.
Patent Information
- Application Number
- CN202510776291.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-09-19
AI Technical Summary
In the field of financial security and anti-money laundering, traditional suspicious node detection methods rely on explicit rules and are easily circumvented by money launderers, resulting in low hit rates and insufficient processing efficiency, making it difficult to effectively identify suspicious nodes.
A knowledge graph-based node identification method is adopted to construct an initial knowledge graph by obtaining suspicious entities captured by expert rules. Combined with transaction association, customer association and account association features, the weight of the edge is dynamically adjusted to perform community division and node identification, thereby improving the accuracy and reliability of node identification.
It significantly improves the accuracy of node identification and the precision of suspicious group identification, can more accurately quantify the transaction correlation strength and potential risks between nodes, and enhances the overall accuracy and reliability of suspicious group identification.
Smart Images

Figure CN120671790A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of computer technology, in particular to the field of financial technology, and specifically to a node identification method, device, electronic device and storage medium based on a knowledge graph. Background Art
[0002] In the fields of financial security and anti-money laundering, detecting suspicious nodes is a core component of maintaining financial order and preventing criminal activity. Existing methods for monitoring suspected money laundering rings rely on expert experience to set association rules and identify connected nodes in transaction chains using traditional relational database techniques such as table joins. However, with the increasing complexity of financial transactions and the explosive growth of data volumes, the limitations of these traditional methods are becoming increasingly apparent. These explicit rule-based monitoring methods are easily circumvented by money launderers through adjustments to transaction patterns, resulting in a low effective hit rate. Furthermore, their inefficiency in processing massive amounts of data forces regulators to devote significant manpower to secondary screening. Therefore, identifying suspicious nodes is crucial. Summary of the Invention
[0003] The present application provides a node identification method, device, electronic device and storage medium based on a knowledge graph to improve the accuracy of node identification.
[0004] In a first aspect, an embodiment of the present application provides a node identification method based on a knowledge graph, comprising:
[0005] Obtaining a suspicious subject captured by expert rules, determining associated nodes of the suspicious subject based on association rules, obtaining an initial knowledge graph including suspicious subject nodes and associated nodes, and using the degree of association with the suspicious subject node as a node label; the association rules include at least one of transaction association, customer association, or account association;
[0006] Determining weights of edges between different nodes based on at least one of transaction association features, customer association features, or account association features between different nodes in the initial knowledge graph, and the node labels;
[0007] Dividing the initial knowledge graph into communities according to the weights of the edges between different nodes to obtain a sub-transaction network;
[0008] The transaction closeness of the node in the sub-transaction network is determined according to the weights of the edges between different nodes, and whether the node is a suspicious node is determined according to the transaction closeness.
[0009] In a second aspect, an embodiment of the present application further provides a node identification device based on a knowledge graph, comprising:
[0010] An initial knowledge graph module is configured to obtain suspicious entities captured by expert rules, determine associated nodes of the suspicious entities based on association rules, obtain an initial knowledge graph including suspicious entity nodes and associated nodes, and use the degree of association with the suspicious entity nodes as node labels; the association rules include at least one of transaction association, customer association, or account association;
[0011] an edge weight module, configured to determine the weights of edges between different nodes based on at least one of transaction association features, customer association features, or account association features between different nodes in the initial knowledge graph, and the node labels;
[0012] A community detection module is used to divide the initial knowledge graph into communities based on the weights of the edges between different nodes to obtain a sub-transaction network;
[0013] The node identification module is used to determine the transaction closeness of the node in the sub-transaction network according to the weights of the edges between different nodes, and to determine whether the node is a suspicious node according to the transaction closeness.
[0014] In a third aspect, an embodiment of the present application further provides an electronic device, the electronic device comprising:
[0015] one or more processors;
[0016] a storage device for storing one or more programs;
[0017] When one or more programs are executed by one or more processors, the one or more processors implement any one of the knowledge graph-based node identification methods provided in the embodiments of the present application.
[0018] In a fourth aspect, an embodiment of the present application further provides a storage medium comprising computer-executable instructions, which, when executed by a computer processor, are used to execute any one of the knowledge graph-based node identification methods provided in the embodiments of the present application.
[0019] Based on the initial knowledge graph, this application not only comprehensively considers the transaction association characteristics, customer association characteristics or account association characteristics between corresponding nodes, but also dynamically and adaptively adjusts the weight of the edge according to the label type of the node, which can more accurately quantify the transaction association strength and potential risk suspicion between different nodes, thereby significantly improving the accuracy of weight calculation; and, based on the optimized edge weights, the initial knowledge graph is divided into communities to obtain multiple sub-transaction networks; in each sub-transaction network, the transaction closeness between nodes is further determined according to the edge weight, and whether the node is a suspicious node is judged according to the transaction closeness of the node, which not only improves the accuracy of single node identification, but also enhances the accuracy and reliability of suspicious group identification as a whole.
[0020] Therefore, the technical solution of this application solves the problem and achieves the desired effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 This is a flowchart of a node identification method based on a knowledge graph provided according to the first embodiment of the present application;
[0022] Figure 2 This is a flowchart of another node identification method based on knowledge graph provided in Example 2 of the present application;
[0023] Figure 3a This is a flowchart of another node identification method based on a knowledge graph according to the third embodiment of the present application;
[0024] Figure 3b This is a schematic diagram of another node identification principle based on the knowledge graph according to the third embodiment of the present application;
[0025] Figure 4 This is a schematic structural diagram of a node identification device based on a knowledge graph according to the fourth embodiment of the present application;
[0026] Figure 5 It is a structural diagram of an electronic device that implements the node identification method based on the knowledge graph of an embodiment of the present application. DETAILED DESCRIPTION
[0027] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0028] It should be noted that the terms "first" and "second" in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0029] Example 1
[0030] Figure 1 This is a flow chart of a node identification method based on a knowledge graph according to the first embodiment of the present application. This embodiment can be applied to the situation where starting from a suspicious subject, the risk node belonging to the same group as the suspicious subject is identified. It can be executed by a node identification device based on a knowledge graph. The node identification device based on the knowledge graph can be implemented in the form of hardware and / or software, and the device can be configured in a computer device. Figure 1 As shown, the method includes:
[0031] S101. Obtain a suspicious subject captured by expert rules, determine associated nodes of the suspicious subject based on association rules, obtain an initial knowledge graph including suspicious subject nodes and associated nodes, and use the degree of association with the suspicious subject node as a node label; the association rules include at least one of transaction association, customer association, or account association;
[0032] S102. Determine the weight of the edge between different nodes based on at least one of the transaction association feature, customer association feature, or account association feature between different nodes in the initial knowledge graph and the node label;
[0033] S103, dividing the initial knowledge graph into communities according to the weights of the edges between different nodes to obtain a sub-transaction network;
[0034] S104: Determine the transaction closeness of the node in the sub-transaction network according to the weights of the edges between different nodes, and determine whether the node is a suspicious node according to the transaction closeness.
[0035] Expert rules are risk indicators and scoring criteria built based on expert experience and knowledge, used to quantify the risk level of an account. For example, by calling the abnormal account interface provided by the regulatory authorities, attribute information such as the number of transactions and amount distribution of abnormal accounts, as well as identity information such as account holder information and associated accounts, can be obtained. Expert rules are then constructed based on this information and used to calculate a risk score for each account. If the risk score of any account exceeds a preset risk threshold, the account is determined to be a suspicious entity, thereby improving the reliability of the suspicious entity and laying the foundation for subsequent identification of risk nodes belonging to the same group as the suspicious entity. Association rules include at least one of transaction association, customer association, or account association. Transaction association includes the same counterparty, the same transaction agent, and the use of the same device for transactions based on Internet Protocol Address (IP address) or Media Access Control Address (MAC address). Customer association includes family relationships, colleagues, neighbors, work relationships, or the same business personnel, business address, and contact information. Account association refers to the same account agent.
[0036] Considering database computing performance issues, association rules are used to determine the preset numerical value of the suspicious subject's associated nodes. Taking the preset value of 3 as an example, the first-degree associated nodes, second-degree associated nodes, and third-degree associated nodes of the suspicious subject can be determined, resulting in an initial knowledge graph including the suspicious subject node, first-degree associated nodes, second-degree associated nodes, and third-degree associated nodes. The corresponding association degree of each node is used as the node label. The node labels of the suspicious subject node, first-degree associated node, second-degree associated node, and third-degree associated node are 0, 1, 2, and 3, respectively.
[0037] The initial knowledge graph can include different suspicious entities. Different suspicious entities can be associated nodes with each other, and the associated nodes of different suspicious entities can be the same. If a node is associated with a single suspicious entity node, it is a single-label node, and its label includes the degree of association with the single suspicious entity. If a node is associated with multiple different suspicious entity nodes simultaneously, it is a multi-label node, and its label includes the corresponding degrees of association with different suspicious entities. For example, if node A is a first-degree associated node with suspicious entity Zhang San and a third-degree associated node with suspicious entity Wang Er, the node labels of node A are Zhang San 1 and Wang Er 3.
[0038] The weights of edges between different nodes are used to characterize the strength of transactional connections and risk suspicion between these nodes. For any edge in the initial knowledge graph, the weight of the edge is determined by combining at least one of the transaction, customer, or account connection characteristics between the two nodes in the edge, as well as the node labels of both nodes. Transactional connection characteristics may include transaction amount and number of transactions; customer connection characteristics may include address similarity and whether the customer is related; and account connection characteristics may include whether the account agents are the same. The process of determining edge weights not only considers the transactional, customer, or account connection characteristics between the corresponding nodes, but also incorporates node label information, enabling adaptive adjustments based on the node's label type. For single-label nodes, the weight is adaptively adjusted based on the node's connection degree with a single suspicious entity; for multi-label entities, the weight is adaptively adjusted based on the node's connection degree with multiple suspicious entities. This allows the edge weight to more accurately reflect the strength of transactional connections and risk suspicion between different nodes, thereby improving the accuracy of weight calculation.
[0039] It should be noted that the relevant information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data comply with relevant laws, regulations and standards of the relevant regions.
[0040] In an embodiment of the present application, the nodes in the initial knowledge graph can be divided into communities based on the weights between different nodes to obtain sub-transaction networks. Exemplarily, the Louvain algorithm can be used to divide the nodes into communities based on modularity to obtain sub-transaction networks. The Louvain algorithm is a heuristic algorithm based on greedy optimization. It does not require pre-labeled data and can perform fast and accurate community detection on large-scale networks. Therefore, it is widely used in fields such as financial network analysis. The goal of the algorithm is to divide the nodes in the initial knowledge graph into several communities to maximize the modularity, and continuously synthesize or split nodes through iterative optimization to identify subgraphs with close internal connections and sparse external connections, i.e., sub-transaction networks. Moreover, based on the weights of the edges between different nodes, the transaction closeness of the node in the sub-transaction network is determined; combined with the transaction closeness, it is determined whether the node is a suspicious node, that is, whether the node in the sub-transaction network belongs to the same group as the suspicious subject.
[0041] The technical solution of this embodiment, based on the initial knowledge graph, not only comprehensively considers the transaction association characteristics, customer association characteristics or account association characteristics between corresponding nodes, but also dynamically and adaptively adjusts the weights of the edges according to the label types of the nodes, which can more accurately quantify the transaction association strength and potential risk suspicion between different nodes, thereby significantly improving the accuracy of weight calculation; and, based on the optimized edge weights, the initial knowledge graph is divided into communities to obtain multiple sub-transaction networks; in each sub-transaction network, the transaction closeness between nodes is further determined according to the edge weights, and whether the node is a suspicious node is judged according to the transaction closeness of the node, which not only improves the accuracy of single node identification, but also enhances the accuracy and reliability of suspicious group identification as a whole.
[0042] Example 2
[0043] Figure 2 This is a flowchart of another node identification method based on knowledge graph provided in Example 2 of this application. The technical solution of this embodiment further refines the weight determination of edges based on the above technical solution. Figure 2 A node identification method based on a knowledge graph is shown, comprising:
[0044] S201. Obtain a suspicious subject captured by expert rules, determine associated nodes of the suspicious subject based on association rules, obtain an initial knowledge graph including suspicious subject nodes and associated nodes, and use the degree of association with the suspicious subject node as a node label; the association rules include at least one of transaction association, customer association, or account association;
[0045] S202. Determine basic weights between different nodes based on at least one of transaction association features, customer association features, or account association features between different nodes in the initial knowledge graph;
[0046] S203, determining a correction coefficient between different nodes based on the node labels; wherein the correction coefficient is negatively correlated with the association degree in the node labels;
[0047] S204, modifying the basic weight according to the modification coefficient to obtain the weights of the edges between different nodes;
[0048] S205: Divide the initial knowledge graph into communities according to the weights of the edges between different nodes to obtain a sub-transaction network;
[0049] S206: Determine the transaction closeness of the node in the sub-transaction network according to the weights of the edges between different nodes, and determine whether the node is a suspicious node according to the transaction closeness.
[0050] Among them, association rules include at least one of transaction association, customer association or account association; transaction association includes being counterparties, transaction agents or using the same transaction equipment; customer association includes relationships such as relatives, colleagues, neighbors, positions, and actual controllers; account association refers to the same account agent.
[0051] For any edge in the initial knowledge graph, the transaction indicators such as the transaction amount and number of transactions between the two nodes corresponding to the edge within a period of time can be normalized and then linearly summed to serve as the transaction association weight; for customer association relationships formed by customer attributes such as kinship and neighborship, a binary classification value can be used as the customer association weight. For example, if there is a kinship or neighbor relationship, the customer association weight is 1; otherwise, it is 0; alternatively, the customer attribute information of the two nodes, such as the address field, is field vectorized to calculate the similarity score between the two. The higher the similarity score, the greater the customer association weight; if a suspicious subject node is included, the risk score calculated in advance for the suspicious subject using expert rules can also be obtained; if different nodes belong to account proxy association, that is, the account agents of the two nodes are the same, a higher account proxy weight can be set, and the preset first value (for example, 0.8) can be used as the corresponding account proxy weight; after normalizing the transaction association weight, customer association weight, risk score, etc., weighted fusion is performed to obtain the basic weight of the edge.
[0052] For example, the basic weight of the edge is calculated using the following formula:
[0053] A ij =α×normalized transaction amount +β×normalized number of transactions +γ×normalized risk score +∑δ k × customer association weight + ε × account proxy weight;
[0054] Among them, A ij is the basic weight of the edge between node i and node j, α, β, γ, δ k and ε are the transaction amount coefficient, transaction number coefficient, risk score coefficient, k-th customer attribute coefficient and account agency coefficient respectively. Each coefficient is a predetermined hyperparameter with a value range of (0,1).
[0055] In the embodiment of the present application, the estimation of the above coefficients can also be integrated into the community detection algorithm, with the goal of maximizing the modularity Q of the objective function, and the optimal hyperparameters α, β, γ, δ are searched by the Bayesian optimization method. kThe final coefficient value is obtained by combining ε and ε. For example, the dataset can be divided into a training set and a test set. On the training set, community detection is performed using different coefficient values and the results are recorded. On the test set, the performance of the community detection results, such as modularity, is evaluated under different coefficient values. The coefficient value with the best performance on the test set is selected as the final coefficient value.
[0056] For any edge in the initial knowledge graph, the correction coefficient between the two nodes is determined according to the node labels of the two nodes in the edge, that is, the correction coefficient between the two nodes is determined in combination with whether the two nodes belong to single-label nodes, multi-label nodes, and the corresponding correlation degree; moreover, the correction coefficient is used to correct the basic weight to obtain the weight of the edge between different nodes; among them, the correction coefficient is negatively correlated with the correlation degree in the node label, the higher the correlation degree, the lower the corresponding correction coefficient, and the smaller the impact of the correction coefficient on the basic weight, so that the adjustment range of the edge weight is smaller, and the actual correlation between the nodes is more accurately reflected.
[0057] In an optional embodiment, the correction coefficient between different nodes is determined based on the node label, including: if the first node is a single-label node and the second node is a multi-label node, then based on the single label of the first node and the multi-label of the second node, determining whether the same common suspicious subject exists between the first node and the second node; if a common suspicious subject exists, the degree of association with the common suspicious subject is used as the degree of association of the second node, and the correction coefficient between the first node and the second node is determined based on a preset first coefficient, the degree of association of the first node and the degree of association of the second node; if no common suspicious subject exists, the minimum degree of association is used as the degree of association of the second node, and the correction coefficient between the first node and the second node is determined based on a preset second coefficient, the degree of association of the first node and the degree of association of the second node; wherein, the first coefficient is less than the second coefficient.
[0058] For an edge between a first node and a second node, if the first node is a single-label node, i.e., the first node is associated with a unique suspicious subject, and the second node is a multi-label node, i.e., the second node is associated with at least two suspicious subjects, then a determination is made as to whether the first node and the second node share a common suspicious subject. If a common suspicious subject exists, then the degree of association with the common suspicious subject is selected from the at least two association degrees of the second node as the association degree of the second node. Furthermore, a correction coefficient between the first node and the second node can be determined based on a preset first coefficient, the association degree of the first node, and the association degree of the second node using the following formula:
[0059]
[0060] Among them, w, β1, d1, and d2 are the correction coefficient, the first coefficient, the correlation degree of the first node, and the correlation degree of the second node respectively.
[0061] If there is no common suspicious subject, the smallest of the at least two association degrees of the second node is selected as the association degree of the second node; and the correction coefficient between the first node and the second node can be determined according to the following formula based on the preset second coefficient, the association degree of the first node and the association degree of the second node:
[0062]
[0063] Among them, w, β2, d1, and d2 are the correction coefficient, the second coefficient, the correlation degree of the first node, and the correlation degree of the second node respectively, and β1 is less than β2.
[0064] By distinguishing between single-label and multi-label nodes and detecting the presence of publicly suspected entities, the method can accurately capture coordinated behavior or group risks within a gang. If a publicly suspected entity exists, the correction coefficient is calculated using the degree of association with the publicly suspected entity and a smaller first coefficient. This not only highlights the directness of the common association, but also weakens the influence of the degree of association to prevent strong associations from overdoing it. If there are no publicly suspected entities, the minimum degree of association among multi-label nodes is selected as the degree of association, combined with a larger second coefficient to amplify the sensitivity of weak associations, thereby effectively identifying dispersed potential risks. This dynamic adjustment mechanism balances the weight differences between strong and weak associations and adapts to the needs of different scenarios through parameter differentiation. In the knowledge graph, it can both focus on core associated groups and explore hidden decentralized risk networks, significantly improving the accuracy and practicality of association analysis.
[0065] In addition, if both the first node and the second node are single-label nodes, it is determined whether there is a common suspicious subject between the first node and the second node; if there is a common suspicious subject, a correction coefficient between the first node and the second node is determined based on a preset first coefficient, the correlation degree of the first node and the correlation degree of the second node; otherwise, a correction coefficient between the first node and the second node is determined based on a preset second coefficient, the correlation degree of the first node and the correlation degree of the second node. If both the first node and the second node are multi-label nodes, it is determined whether there is a common suspicious subject between the first node and the second node. If there is a common suspicious subject, the correlation degree with the common suspicious subject is used as the correlation degree of the first node and the second node, and a correction coefficient between the first node and the second node is determined based on the preset first coefficient, the correlation degree of the first node and the correlation degree of the second node; if there is no common suspicious subject, the minimum correlation degree is used as the correlation degree of the first node and the second node, and a correction coefficient between the first node and the second node is determined based on the preset second coefficient, the correlation degree of the first node and the correlation degree of the second node.
[0066] The technical solution of this embodiment determines the basic weight of an edge based on at least one of the transaction association characteristics, customer association characteristics, or account association characteristics between different nodes based on an initial knowledge graph. It then further distinguishes between single-label and multi-label nodes and detects whether there are any suspicious entities associated with them. It then determines a correction coefficient based on the different detected situations and dynamically adjusts the edge weight using the correction coefficient. This dynamic correction mechanism balances the weight differences between strong and weak associations and adapts to the needs of different scenarios through parameter differentiation, further improving the accuracy of weight calculation and thus node identification.
[0067] Example 3
[0068] Figure 3a This is a flowchart of another node identification method based on knowledge graph according to the third embodiment of the present application. The technical solution of this embodiment further refines the selection of target nodes on the basis of the above technical solution. Figure 3a A node identification method based on a knowledge graph is shown, comprising:
[0069] S301: Obtain a suspicious subject captured by expert rules, determine associated nodes of the suspicious subject based on association rules, obtain an initial knowledge graph including suspicious subject nodes and associated nodes, and use the degree of association with the suspicious subject node as a node label; the association rules include at least one of transaction association, customer association, or account association;
[0070] S302. Determine the weight of the edge between different nodes based on at least one of the transaction association feature, customer association feature, or account association feature between different nodes in the initial knowledge graph and the node label;
[0071] S303: Divide the initial knowledge graph into communities based on the weights of the edges between different nodes to obtain a sub-transaction network;
[0072] S304: Determine the weighted degree centrality, closeness centrality, and eigenvector centrality of each node in the sub-transaction network based on the weights of the edges between different nodes.
[0073] S305: Fusing the weighted degree centrality, closeness centrality, and eigenvector centrality to obtain the transaction closeness of the node in the sub-transaction network;
[0074] S306: Determine whether the node is a suspicious node based on the transaction closeness.
[0075] Among them, weighted degree centrality calculates the sum of the weights of each edge connected to the node in the sub-trading network, and is used to identify core accounts for high-frequency trading or large-scale capital transactions. For a node in a sub-trading network, its weighted degree centrality can be calculated using the following formula:
[0076] score1(v)=∑w(u,v)
[0077] Among them, score1(v) is the weighted degree centrality of node v, and w(u,v) is the weight of the edge between nodes u and v.
[0078] Closeness Centrality calculates the sum of the reciprocals of the shortest distances from a node to other nodes in a sub-transaction network. This is used to locate fast-moving accounts at the center of a sub-transaction network. For a node in a sub-transaction network, its closeness centrality can be calculated using the following formula:
[0079]
[0080] Where score2(v) is the closeness centrality of node v, n is the total number of nodes in the sub-transaction network, and d(u,v) is the shortest path length between nodes u and v, which is obtained by summing the weights of each edge on the shortest path.
[0081] Eigenvector centrality considers not only the direct connections between nodes but also the importance of the neighbors to which they are connected, and is used to identify nodes closely associated with suspicious entities. For nodes in a sub-transaction network, the eigenvector centrality can be determined by the following process: In the sub-transaction network, if there is an edge between nodes i and j, the corresponding element is the weight of the edge between nodes i and j; if there is no edge, the corresponding element is 0, resulting in an adjacency matrix. The maximum eigenvalue of the adjacency matrix and the corresponding eigenvector x are calculated; the elements of the eigenvector x are used as the eigenvector centrality score (score3(v)) for each node.
[0082] Using predetermined weighting coefficients, the weighted degree centrality, closeness centrality and eigenvector centrality of the node are weighted and summed respectively to obtain the transaction closeness of the node in the sub-transaction network. The transaction closeness can more accurately represent the overall influence and status of the node in the sub-transaction network, thereby improving the accuracy of suspicious node detection based on transaction closeness.
[0083] In an optional embodiment, the method further includes: determining, for each node in the sub-transaction network, transaction data between the node and other nodes in the sub-transaction network; wherein the transaction data includes at least one of the number of transactions, transaction amount, loan amount ratio, transaction time, and transaction remarks; determining the similarity between the node and other nodes based on the transaction data, and calculating the sum of the similarities between the node and different other nodes to obtain the cumulative similarity of the node; determining whether a node is a suspicious node based on the transaction closeness includes: determining whether the node is a suspicious node based on the transaction closeness and the cumulative similarity.
[0084] For each node in the sub-transaction network, determine the number of transactions, transaction amounts, loan / debit ratios, transaction times, and transaction comments between that node and all other nodes in the sub-transaction network. Structured data such as transaction numbers, transaction amounts, loan / debit ratios, and transaction times can be concatenated to generate a first transaction vector. Unstructured data such as transaction comments can be converted into vector form using the TF-IDF (Term Frequency-Inverse Document Frequency) method to generate a second transaction vector. The first and second transaction vectors are then concatenated to generate the transaction feature vector for that node. Similarly, the transaction feature vectors for other nodes can be determined.
[0085] The Jaccard similarity between the transaction feature vector of the node and the transaction feature vectors of other nodes is calculated to obtain the similarity between the node and other nodes. The similarities between the node and different other nodes are summed to obtain the node's cumulative similarity. The higher the cumulative similarity, the more likely the node and other nodes in the group have similar transaction patterns. In addition, the transaction closeness and cumulative similarity of the node can be fused, and the fusion result can be used to determine whether the node is a suspicious node. By determining the cumulative similarity of the node based on the transaction data between the node and other nodes, and combining the transaction closeness of the node with its cumulative similarity in the sub-transaction network, the behavioral characteristics of the node in the sub-transaction network can be more comprehensively captured, thereby further improving the accuracy of suspicious node identification.
[0086] In an optional embodiment, the method further includes: for each node in the sub-transaction network, determining the predicted transaction data of the node at each time point in the current time window based on the historical transaction data of the node and other nodes in the sub-transaction network in the previous time window; determining the transaction residual at each time point in the current time window based on the actual transaction data and the predicted transaction data of the node, and selecting the maximum transaction residual; the transaction residual is used to characterize the degree of deviation from the normal transaction pattern; determining whether the node is a suspicious node based on the transaction closeness includes: determining whether the node is a suspicious node based on the transaction closeness and the maximum transaction residual.
[0087] Transaction data is divided into multiple time windows, each containing transaction data for a period of time, such as a day or an hour. For each node in the sub-transaction network, historical transaction data for the previous time window is obtained. This historical transaction data includes at least one of the following: transaction count, transaction amount, loan / debit ratio, transaction time, and transaction comments. This historical transaction data is then fed into a pre-built transaction prediction model to obtain predicted transaction data for the node at each point in the current time window. The prediction model can be constructed using time series analysis, machine learning models, or deep learning models.
[0088] Furthermore, the actual transaction data of the node at each time point within the current time window is obtained. The residual between the actual transaction data and the predicted transaction data is calculated to characterize the degree of deviation between actual and predicted transaction behavior. The largest transaction residual among the residuals at each time point within the current time window is selected as the maximum transaction residual for the node within the current time window. This is used to characterize the degree to which the node's transaction behavior within the current time window deviates from the normal pattern, thereby identifying whether the node has engaged in abnormal transactions within the current time window. Combining the node's transaction closeness and the maximum transaction residual can more comprehensively capture the node's behavioral characteristics within the sub-transaction network, further improving the accuracy of suspicious node identification.
[0089] refer to Figure 3b For nodes in a sub-transaction network, the transaction closeness, cumulative similarity, and maximum transaction residual of the node can be combined to obtain a comprehensive risk score. The weighted weight can be estimated using a linear regression algorithm based on historical data. The comprehensive risk score is used to further screen the nodes in the sub-transaction network. The screening threshold can be adjusted based on actual case development to obtain highly suspicious first-, second-, and third-degree connected nodes, which constitute the final group. The introduction of the comprehensive risk score makes the resulting group more interpretable, assisting anti-money laundering personnel in making judgments and writing reports.
[0090] Alternatively, node transaction data can be combined with various anomaly detection techniques to identify abnormal transaction patterns within the node, and the overall risk score can be adjusted based on these patterns. For example, the first method utilizes time series analysis techniques, employing seasonal decomposition methods to decompose transaction data into three components: trend, seasonality, and residuals. In trend analysis, abnormal trend patterns are identified; seasonal components are tested for cyclical variations that are inconsistent with normal trading cycles; and residuals are tested for unexplained abnormal noise. The second method uses a mean-value algorithm to perform cluster analysis on feature-engineered data to identify outliers, which are then used to characterize abnormal trading patterns.
[0091] For suspicious nodes, firstly, the suspicious nodes and their associations in the gang can be displayed in the abnormal transaction report of the suspicious subject through the knowledge graph provided by the graph database application programming interface (Application Programming Interface, API interface), to assist in the identification and elimination of abnormalities; secondly, an associated suspicious report can be generated for the suspicious nodes to supplement the suspicious subjects that the expert rules failed to capture, thereby reducing the workload of manually adding cases; thirdly, suspicious nodes can be regarded as situations outside the expert rules to supplement the existing expert rules, realize dynamic generation of scores, and solve the problem of explicit and easy evasion of monitoring rules; fourthly, the hit situation of the association rules can be used to supplement the system customer relationship map, assist anti-money laundering business personnel in risk rating judgment, and improve business efficiency.
[0092] The technical solution of this embodiment combines the transaction closeness of the node in the sub-transaction network, the cumulative similarity used to measure similar transaction patterns between the node and other nodes in the sub-transaction network, and the maximum transaction residual used to reflect whether the node itself has an abnormal transaction pattern, to comprehensively determine whether the node is a suspicious node, further improving the accuracy of node identification.
[0093] Example 4
[0094] Figure 4 This is a schematic diagram of the structure of a node identification device based on a knowledge graph according to the fourth embodiment of the present application. This embodiment can be applied to the situation where a suspicious subject is identified as a risk node belonging to the same group as the suspicious subject. The node identification device based on the knowledge graph can be implemented in the form of hardware and / or software, and the device can be configured in a computer device. Figure 4 The specific structure of the node identification device based on the knowledge graph is as follows:
[0095] Initial knowledge graph module 410 is configured to obtain suspicious entities captured by expert rules, determine associated nodes of the suspicious entities based on association rules, obtain an initial knowledge graph including suspicious entity nodes and associated nodes, and use the degree of association with the suspicious entity nodes as node labels; the association rules include at least one of transaction association, customer association, or account association;
[0096] an edge weight module 420 for determining weights of edges between different nodes based on at least one of transaction association features, customer association features, or account association features between different nodes in the initial knowledge graph and the node labels;
[0097] A community detection module 430 is configured to divide the initial knowledge graph into communities based on the weights of the edges between different nodes to obtain sub-transaction networks;
[0098] The node identification module 440 is configured to determine the transaction closeness of the node in the sub-transaction network according to the weights of the edges between different nodes, and determine whether the node is a suspicious node according to the transaction closeness.
[0099] In an optional implementation, the edge weight module 420 includes:
[0100] a basic weight unit, configured to determine basic weights between different nodes in the initial knowledge graph based on at least one of transaction association features, customer association features, or account association features between different nodes;
[0101] A correction coefficient unit, configured to determine a correction coefficient between different nodes based on the node labels; wherein the correction coefficient is negatively correlated with the degree of association in the node labels;
[0102] The edge weight unit is used to correct the basic weight according to the correction coefficient to obtain the weight of the edge between different nodes.
[0103] In an optional implementation, the correction coefficient unit includes:
[0104] a public suspicious subject subunit, configured to determine, if the first node is a single-label node and the second node is a multi-label node, whether the first node and the second node have the same public suspicious subject based on the single label of the first node and the multi-label of the second node;
[0105] a first correction coefficient subunit, configured to, if a public suspicious subject exists, use the correlation degree with the public suspicious subject as the correlation degree of the second node, and determine a correction coefficient between the first node and the second node based on a preset first coefficient, the correlation degree of the first node, and the correlation degree of the second node;
[0106] The second correction coefficient subunit is used to use the minimum correlation degree as the correlation degree of the second node if there is no common suspicious subject, and determine the correction coefficient between the first node and the second node based on the preset second coefficient, the correlation degree of the first node and the correlation degree of the second node; wherein the first coefficient is smaller than the second coefficient.
[0107] In an optional implementation, the node identification module 440 includes:
[0108] a centrality unit, for determining the weighted degree centrality, closeness centrality, and eigenvector centrality of a node in the sub-transaction network, respectively, based on the weights of the edges between different nodes;
[0109] A closeness unit, configured to fuse the weighted degree centrality, closeness centrality, and eigenvector centrality to obtain the transaction closeness of the node in the sub-transaction network;
[0110] The node identification unit is used to determine whether a node is a suspicious node based on the transaction closeness.
[0111] In an optional implementation, the node identification module 440 further includes a similarity accumulation unit, and the similarity accumulation unit includes:
[0112] a transaction data sub-unit, configured to determine, for each node in the sub-transaction network, transaction data between the node and other nodes in the sub-transaction network; wherein the transaction data includes at least one of the number of transactions, transaction amount, loan amount ratio, transaction time, and transaction remarks;
[0113] A cumulative similarity subunit, configured to determine the similarity between the node and other nodes based on the transaction data, and to calculate the sum of the similarities between the node and different other nodes to obtain the cumulative similarity of the node;
[0114] The node identification unit is specifically used for:
[0115] Determine whether the node is a suspicious node based on the transaction closeness and the accumulated similarity.
[0116] In an optional implementation, the node identification module 440 further includes a transaction residual unit, and the transaction residual unit includes:
[0117] The transaction prediction sub-unit is used to determine, for each node in the sub-transaction network, the predicted transaction data of the node at each time point in the current time window based on the historical transaction data of the node and other nodes in the sub-transaction network in the previous time window;
[0118] A transaction residual subunit is used to determine the transaction residual at each time point based on the actual transaction data and predicted transaction data of the node at each time point in the current time window, and select the maximum transaction residual; the transaction residual is used to indicate the degree of deviation from the normal trading pattern;
[0119] The node identification unit is specifically used for:
[0120] Determine whether the node is a suspicious node based on the transaction closeness and the maximum transaction residual.
[0121] The knowledge graph-based node identification device provided in the embodiment of the present application can execute the knowledge graph-based node identification method provided in any embodiment of the present application, and has the corresponding functional modules and beneficial effects for executing the knowledge graph-based node identification method.
[0122] According to an embodiment of the present invention, the present invention further provides an electronic device, a readable storage medium and a computer program product.
[0123] Example 5
[0124] Figure 5 : is a structural diagram of an electronic device 510 that implements the node identification method based on the knowledge graph of an embodiment of the present application. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or required herein.
[0125] like Figure 5 As shown, the electronic device 510 includes at least one processor 511, and a memory connected to the at least one processor 511 in communication, such as a read-only memory (ROM) 512, a random access memory (RAM) 513, etc., wherein the memory stores a computer program that can be executed by at least one processor, and the processor 511 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 512 or the computer program loaded from the storage unit 518 into the random access memory (RAM) 513. In the RAM 513, various programs and data required for the operation of the electronic device 510 can also be stored. The processor 511, ROM 512 and RAM 513 are connected to each other via a bus 514. An input / output (I / O) interface 515 is also connected to the bus 514.
[0126] Multiple components in the electronic device 510 are connected to the I / O interface 515, including an input unit 516, such as a keyboard, a mouse, etc.; an output unit 517, such as various types of displays, speakers, etc.; a storage unit 518, such as a magnetic disk, an optical disk, etc.; and a communication unit 519, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 519 allows the electronic device 510 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0127] The processor 511 can be a variety of general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the processor 511 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 511 executes the various methods and processes described above, such as the node identification method based on the knowledge graph.
[0128] In some embodiments, the node identification method based on the knowledge graph may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 518. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 510 via the ROM 512 and / or the communication unit 519. When the computer program is loaded into the RAM 513 and executed by the processor 511, one or more steps of the node identification method based on the knowledge graph described above may be performed. Alternatively, in other embodiments, the processor 511 may be configured as a node identification method based on the knowledge graph by any other appropriate means (e.g., by means of firmware).
[0129] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0130] Computer programs for implementing the methods of the present application may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0131] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. A computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0132] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0133] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0134] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.
[0135] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this application can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of this application can be achieved. This is not limited herein.
[0136] The above specific embodiments do not constitute a limitation on the scope of protection of this application. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application shall be included within the scope of protection of this application.
Claims
1. A node identification method based on knowledge graph, characterized in that: include: Obtaining a suspicious subject captured by expert rules, determining associated nodes of the suspicious subject based on association rules, obtaining an initial knowledge graph including suspicious subject nodes and associated nodes, and using the degree of association with the suspicious subject node as a node label; the association rules include at least one of transaction association, customer association, or account association; Determining weights of edges between different nodes based on at least one of transaction association features, customer association features, or account association features between different nodes in the initial knowledge graph, and the node labels; Dividing the initial knowledge graph into communities according to the weights of the edges between different nodes to obtain a sub-transaction network; The transaction closeness of the node in the sub-transaction network is determined according to the weights of the edges between different nodes, and whether the node is a suspicious node is determined according to the transaction closeness.
2. The method according to claim 1, wherein Determining the weights of edges between different nodes based on at least one of the transaction association feature, customer association feature, or account association feature between different nodes in the initial knowledge graph, and the node labels, includes: determining basic weights between different nodes based on at least one of transaction association features, customer association features, or account association features between different nodes in the initial knowledge graph; Determining a correction coefficient between different nodes based on the node labels; wherein the correction coefficient is negatively correlated with the degree of association in the node labels; The basic weight is modified according to the correction coefficient to obtain the weight of the edge between different nodes.
3. The method according to claim 2, characterized in that Determining correction coefficients between different nodes according to the node labels includes: If the first node is a single-label node and the second node is a multi-label node, determining whether the first node and the second node have the same common suspicious subject based on the single label of the first node and the multi-label of the second node; If there is a public suspicious subject, the correlation degree with the public suspicious subject is used as the correlation degree of the second node, and a correction coefficient between the first node and the second node is determined according to a preset first coefficient, the correlation degree of the first node and the correlation degree of the second node; If there is no common suspicious subject, the minimum correlation degree is used as the correlation degree of the second node, and the correction coefficient between the first node and the second node is determined based on the preset second coefficient, the correlation degree of the first node and the correlation degree of the second node; wherein the first coefficient is smaller than the second coefficient.
4. The method according to claim 1, wherein Determining the transaction closeness of the node in the sub-transaction network according to the weights of the edges between different nodes, and determining whether the node is a suspicious node according to the transaction closeness, includes: Determine the weighted degree centrality, closeness centrality, and eigenvector centrality of the node in the sub-transaction network according to the weights of the edges between different nodes; The weighted degree centrality, closeness centrality and eigenvector centrality are integrated to obtain the transaction closeness of the node in the sub-transaction network; Determine whether the node is a suspicious node based on the transaction closeness.
5. The method according to claim 1, wherein The method further comprises: For each node in the sub-transaction network, determining transaction data between the node and other nodes in the sub-transaction network; wherein the transaction data includes at least one of the number of transactions, transaction amount, loan amount ratio, transaction time, and transaction remarks; Determine the similarity between the node and other nodes based on the transaction data, and calculate the sum of the similarities between the node and different other nodes to obtain the accumulated similarity of the node; The determining whether a node is a suspicious node according to the transaction closeness includes: Determine whether the node is a suspicious node based on the transaction closeness and the accumulated similarity.
6. The method according to claim 1, characterized in that The method further comprises: For each node in the sub-transaction network, based on the historical transaction data of the node and other nodes in the sub-transaction network in the previous time window, the predicted transaction data of the node at each time point in the current time window is determined; Based on the actual transaction data and predicted transaction data of the node at each time point in the current time window, the transaction residual at each time point is determined, and the maximum transaction residual is selected; the transaction residual is used to indicate the degree of deviation from the normal trading pattern; The determining whether a node is a suspicious node according to the transaction closeness includes: Determine whether the node is a suspicious node based on the transaction closeness and the maximum transaction residual.
7. A node identification device based on knowledge graph, characterized in that: include: An initial knowledge graph module is configured to obtain suspicious entities captured by expert rules, determine associated nodes of the suspicious entities based on association rules, obtain an initial knowledge graph including suspicious entity nodes and associated nodes, and use the degree of association with the suspicious entity nodes as node labels; the association rules include at least one of transaction association, customer association, or account association; an edge weight module, configured to determine the weights of edges between different nodes based on at least one of transaction association features, customer association features, or account association features between different nodes in the initial knowledge graph, and the node labels; A community detection module is used to divide the initial knowledge graph into communities based on the weights of the edges between different nodes to obtain a sub-transaction network; The node identification module is used to determine the transaction closeness of the node in the sub-transaction network according to the weights of the edges between different nodes, and to determine whether the node is a suspicious node according to the transaction closeness.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, it implements the node identification method based on the knowledge graph as described in any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the node identification method based on the knowledge graph as described in any one of claims 1 to 6 is implemented.
10. A computer program product, characterized in that The computer program product comprises a computer program, which, when executed by a processor, implements the knowledge graph-based node identification method according to any one of claims 1 to 6.