Transaction chain abnormal node identification method based on knowledge graph enhancement

By building a transaction chain knowledge graph and assigning dynamic attributes to nodes and edges, combining time series knowledge graphs and topological features, the real-time identification problem of abnormal nodes in the transaction chain is solved, and the accuracy and reliability of abnormal detection are improved.

CN120234582AActive Publication Date: 2025-07-01NANJING BOSHENGYU NETWORK TECH CO LTD

Patent Information

Application Number
CN202510727252.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-07-01
Estimated Expiration
2045-06-03

AI Technical Summary

Technical Problem

The existing technology has static analysis limitations in the transaction chain, and it is impossible to capture dynamic changes in real time and ignore topological relationships, resulting in a lack of contextual support for abnormal detection results and it is difficult to identify cross-node abnormal behavior.

Method used

Build a transaction chain knowledge graph, assign dynamic attributes to nodes and edges, generate dynamic updated knowledge graph representations through the timing knowledge graph embedding algorithm, calculate abnormal scores based on topology and features, and perform secondary verification to output the final list of abnormal nodes.

Benefits of technology

It realizes comprehensive and dynamic identification of abnormal nodes in the transaction chain, improves the accuracy and reliability of abnormal detection, and reduces the risk of misjudgment and misjudgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234582A_ABST
    Figure CN120234582A_ABST
Patent Text Reader

Abstract

The invention discloses a transaction chain abnormal node identification method based on knowledge graph enhancement, and belongs to the technical field of network security, and the method specifically comprises the steps: constructing a transaction chain knowledge graph, endowing nodes and edges with dynamic attributes based on historical data, and calculating a transaction entropy value; collecting transaction data in real time, mapping participants and relationships into low-dimensional vectors by using a time sequence knowledge graph embedding algorithm, generating dynamically updated knowledge graph representation, and updating a transaction entropy value; on the basis of the dynamic knowledge graph, combining topology and features, obtaining an abnormal score through comprehensive calculation, and identifying potential abnormal nodes according to a topological structure; and the transaction chain context information of the potential abnormal nodes and the neighborhood nodes thereof is collected for secondary verification, and a final abnormal node list is output, so that the accuracy and reliability of anomaly detection are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of network security, and specifically relates to a method for identifying abnormal nodes in a transaction chain enhanced by a knowledge graph. Background Art

[0002] In modern various transaction scenarios, a transaction chain often contains a large number of transaction nodes and complex transaction relationships. With the continuous expansion of the transaction scale and the increasing diversification of transaction forms, the emergence of abnormal nodes in the transaction chain may trigger serious problems such as financial fraud and data leakage. Therefore, accurately identifying abnormal nodes in the transaction chain is crucial. Currently, traditional methods for identifying abnormal nodes mainly rely on rule matching and statistical analysis. The rule matching method depends on pre-set rules. When a transaction behavior conforms to a specific rule, it is determined as abnormal. However, this method is difficult to cope with complex and changeable abnormal patterns and has poor flexibility. The statistical analysis method analyzes the statistical characteristics of transaction data and sets thresholds to identify abnormalities. However, for some low-frequency and new abnormal behaviors, it is often unable to effectively identify them. As a semantic network, a knowledge graph can represent entities and their relationships in a structured manner. Applying it to the identification of abnormal nodes in a transaction chain is expected to solve the limitations of traditional methods.

[0003] For example, Chinese Patent with the authorization announcement number CN115345736B discloses a method for detecting abnormal financial transaction behaviors, including: constructing a transaction structure diagram based on historical transaction records, and the nodes in the diagram are accounts; two nodes with transactions are respectively an out-degree node and an in-degree node, and the connection line between the out-degree node and the in-degree node is a transaction route; obtaining the transaction information of each transaction route; inputting the transaction graph structure and the transaction information of each transaction route into a TAD-GCN neural network, and outputting a feature vector of each transaction route through an embedding layer; outputting a description vector of each transaction route through a graph convolutional layer of the TAD-GCN neural network based on the updated weights and the number of convolutional times corresponding to each transaction route; outputting a transaction anomaly recognition result of the transaction route through the description vectors of each transaction route passing through the classification layer of the TAD-GCN neural network. This technical solution can accurately identify the transaction routes of abnormal transaction behaviors.

[0004] The above existing technologies have the following deficiencies: there are limitations in static analysis, that is, relying on static historical data, unable to capture the dynamic changes of the transaction chain in real time and lacking dynamic attributes; ignoring the topological relationship between transaction participants, resulting in difficulty in identifying cross-node abnormal behaviors; the abnormal detection results lack context support and are difficult to assist manual decision-making. Summary of the Invention

[0005] In view of the deficiencies of the prior art, the present invention proposes a method for identifying abnormal nodes in a transaction chain enhanced by a knowledge graph, constructs a transaction chain knowledge graph, and assigns dynamic attributes to nodes and edges based on historical data and calculates the transaction entropy value; real-time collects transaction data, uses a temporal knowledge graph embedding algorithm to map participants and relationships into low-dimensional vectors, generates a dynamically updated knowledge graph representation and updates the transaction entropy value; based on the dynamic knowledge graph, combines topology and features, calculates the abnormal score through comprehensive calculation, and identifies potential abnormal nodes based on the topological structure; collects the transaction chain context information of potential abnormal nodes and their neighborhood nodes for secondary verification, and outputs the final list of abnormal nodes, effectively improving the accuracy and reliability of abnormal detection.

[0006] To achieve the above object, the present invention provides the following technical solutions: A method for identifying abnormal nodes in a transaction chain enhanced by a knowledge graph, comprising: S1: Construct a transaction chain knowledge graph, assign dynamic attributes to the nodes and edges in the transaction chain knowledge graph based on historical transaction chain data, and calculate the transaction entropy value of the dynamic attribute values of the nodes and edges; the nodes represent transaction participants, and the edges represent transaction relationships; S2: Real-time collect the transaction data in the transaction chain, map the transaction participants and transaction relationships into low-dimensional vectors through a temporal knowledge graph embedding algorithm, generate a dynamically updated knowledge graph representation, and update the transaction entropy value of the dynamic attribute values of the nodes and edges; S3: Based on the dynamically updated knowledge graph representation, calculate the abnormal score of each node through an abnormal detection method based on topology and features, and identify potential abnormal nodes in combination with the topological structure of the transaction chain; S4: Collect the transaction chain context information of potential abnormal nodes and their neighborhood nodes, conduct secondary verification on the potential abnormal nodes, and output the final list of abnormal nodes.

[0007] Specifically, the steps of calculating the transaction entropy value of the dynamic attribute values of the nodes and edges in S1 include: S1.1: Create an empty dictionary A, dictionary B, and dictionary C; the dictionaries A, B, and C are in the form of key-value pairs, and the keys in dictionary A are node identifiers, the keys in dictionary B are transaction counterparty node identifiers, the keys in dictionary C are edge identifiers, and the values are initialized to 0; S1.2: For each node i in the transaction chain knowledge graph, count the number of transactions of this node with each transaction counterparty , and obtain the transaction probability of this node with each transaction counterparty by dividing the number of transactions of this node with any transaction counterparty by the total number of transactions of this node , and store it in dictionary B, where j represents the transaction counterparty node identifier; S1.3: Traverse dictionary B. For the trading probability of each counterparty, obtain the trading entropy value of node i according to the information entropy formula , and store it in dictionary A; S1.4: Take the calculated trading entropy value of node i as a new attribute and add it to the nodes of the trading chain knowledge graph; S1.5: For each edge k in the trading chain knowledge graph, by traversing all the trading records corresponding to this edge, count the distribution of different trading attributes to obtain the trading amount interval distribution dictionary and the trading type distribution dictionary; S1.6: By calculating the number of transactions in any trading amount interval divided by the sum of all values in the trading amount interval distribution dictionary of this edge, obtain the probability distribution of this trading amount interval ; S1.7: By calculating the number of transactions of the current type divided by the sum of all values in the trading type distribution dictionary of this edge, obtain the probability distribution of this trading type ; S1.8: According to the information entropy formula in S1.3, calculate the trading amount entropy value and the trading type entropy value respectively. Combine with the preset weights and obtain the edge trading entropy value through weighted summation , and add the edge trading entropy value as a new attribute to the edges of the trading chain knowledge graph. At the same time, store the edge trading entropy value in dictionary C.

[0008] Specifically, the specific steps of S2 include: S2.1: Obtain trading data from the trading system in real time and perform preprocessing; S2.2: Use the dynamic relationship modeling method to model the preprocessed trading data into a dynamic graph with timestamps , where I represents the set of nodes, represents the set of edges, represents the set of dynamic attributes of nodes and edges, and t represents the current moment; S2.3: Use the temporal knowledge graph embedding algorithm to map the nodes and edges in the dynamic graph with timestamps into low-dimensional vectors to obtain the node low-dimensional vectors and the edge low-dimensional vectors ; S2.4: Set the batch threshold for data collection. Each time a new batch of trading data is collected, trigger an update of the embedding vectors, and integrate the newly collected trading data with the existing dynamic graph data with timestamps, and re-execute S2.3 to obtain the updated node low-dimensional vectors and edge low-dimensional vectors; S2.5: Set the time window T. For node i, count the set of counterparties within the time window T , and the number of transactions of this node with each counterparty , and recalculate the transaction probability according to S1.2 and S1.3 and the transaction entropy value ; S2.6: For each edge k, count the distribution dictionary of transaction amount intervals and the distribution dictionary of transaction types within the time window T, and recalculate the edge transaction entropy value according to S1.5 - S1.8 ; S2.7: Add the recalculated node transaction entropy value and the edge transaction entropy value as new attributes to the dynamic graph, and update the dynamic attribute sets of nodes and edges .

[0009] Specifically, the specific steps of S2.3 include: S2.31: Collect dynamic graph data with timestamps; S2.32: Set the embedding dimension d, and create a set of low-dimensional node vectors and a set of low-dimensional edge vectors , and randomly initialize each element in the set of low-dimensional node vectors and the set of low-dimensional edge vectors to a random number in the interval [-0.01, 0.01]. After random initialization, take out the vectors corresponding to the head node h and the tail node e from the set of low-dimensional node vectors as the low-dimensional vector of the head node , the low-dimensional vector of the tail node , where represents the low-dimensional vector of node i, represents the size of the node set, represents the low-dimensional vector of the k-th edge, represents the size of the edge set; S2.33: Set the objective function , and the formula is: ; where represents the time-aware head node embedding vector, represents the time-aware edge embedding vector, represents the time-aware tail node embedding vector, represents the time regularization weight, represents the timestamp, represents the timestamp mapping function, and satisfies , represents the attention mechanism function, represents the attention mechanism weight, represents the structure constraint weight, Represents a structural constraint term and satisfies , represents element-wise multiplication, represents the 1-norm, g represents the embedding dimension index, and .

[0010] Specifically, the specific steps of S2.3 further include: S2.34: Extract positive and negative samples from the dynamic graph with timestamps to form training samples; the positive samples are quadruples of head node, edge, tail node, and timestamp; the negative samples are obtained by randomly replacing the head node, edge, or tail node in the positive samples; S2.35: Input the training samples into the objective function, calculate the loss value corresponding to each training sample, and update the low-dimensional vectors of the head node, tail node, and edge in the real number space according to the loss value , the low-dimensional vector of the tail node and the low-dimensional vector of the edge ; S2.36: After iterative training, when the objective function converges, obtain the optimized low-dimensional vectors of the nodes and the low-dimensional vectors of the edges .

[0011] Specifically, the specific steps of S3 include: S3.1: Load the dynamically updated knowledge graph representation from S2, including the low-dimensional vector representations of nodes and edges, and the dynamic attribute values of nodes and edges; S3.2: Calculate the degree centrality of node i , where u represents the neighbor nodes of node i, represents the neighbor set of node i, containing all nodes directly connected to node i; S3.3: Calculate the closeness centrality of node i , where, represents the shortest path distance from node i to node s, V represents the set of all nodes in the dynamically updated knowledge graph, represents traversing all nodes s except node i in the dynamically updated knowledge graph; S3.4: Calculate the betweenness centrality of node i , where, represents the number of times node i passes through in the shortest path from node s to node l, represents the number of shortest paths from node s to node l; S3.5: Calculate the embedding vector distance between node i and its neighbor node u , where, and represent the embedding vectors of node i and its neighbor node u respectively.

[0012] Specifically, the specific steps of S3 further include: S3.6: Weighted sum the 、 、 、 of each node to obtain the anomaly score of each node ; S3.7: Set an anomaly score threshold according to the historical transaction chain data ; If , then mark the node corresponding to the anomaly score as a potential anomaly node.

[0013] Specifically, the transaction chain context information in S4 includes transaction time series, transaction amount change trend, and relationship change among transaction participants.

[0014] Specifically, the secondary verification in step S4 includes: Conduct logical verification on the transaction behavior of potential anomaly nodes through a rule engine; Perform cross-verification on the associated transactions of potential anomaly nodes by combining external data sources; Dynamically evaluate the transaction patterns of potential anomaly nodes based on time series analysis.

[0015] Specifically, the method for identifying anomaly nodes in a transaction chain enhanced by a knowledge graph further includes: Visually output the final list of anomaly nodes and their associated transaction paths, and generate an anomaly analysis report.

[0016] Compared with the prior art, the beneficial effects of the present invention are: The present invention proposes a method for identifying anomaly nodes in a transaction chain enhanced by a knowledge graph. By constructing a transaction chain knowledge graph containing transaction participants and transaction relationships, and endowing dynamic attributes to nodes and edges, calculating the transaction entropy value, it can comprehensively and dynamically depict transaction characteristics, capture changes and uncertainties in the transaction process, and effectively improve the comprehensiveness and accuracy of anomaly detection.

[0017] The present invention proposes a method for identifying anomaly nodes in a transaction chain enhanced by a knowledge graph. By real-time collecting transaction data and generating a dynamically updated knowledge graph representation, combining topological and feature-based anomaly detection methods to calculate node anomaly scores, and then performing secondary verification on potential anomaly nodes, this solution realizes fine-grained management of the entire process from data update, anomaly detection to verification and confirmation, and can more accurately identify real anomaly nodes, reducing the risks of misjudgment and missed judgment. Description of the Drawings

[0018] Figure 1Schematic diagram of the method for identifying abnormal nodes in a transaction chain enhanced by a knowledge graph according to the present invention; Figure 2 Principle flow chart of the method for identifying abnormal nodes in a transaction chain enhanced by a knowledge graph according to the present invention; Figure 3 Flow chart for implementing the transaction entropy value of the method for identifying abnormal nodes in a transaction chain enhanced by a knowledge graph according to the present invention; Figure 4 Flow chart for identifying potential abnormal nodes of the method for identifying abnormal nodes in a transaction chain enhanced by a knowledge graph according to the present invention. Detailed implementation manner

[0019] Example 1 Please refer to Figure 1 and Figure 2 A method for identifying abnormal nodes in a transaction chain enhanced by a knowledge graph provided by the present invention, the method includes steps S1 to S4, including the following steps: S1: Construct a transaction chain knowledge graph, assign dynamic attributes to the nodes and edges in the transaction chain knowledge graph based on historical transaction chain data, and calculate the transaction entropy value of the dynamic attribute values of the nodes and edges; the nodes represent transaction participants, and the edges represent transaction relationships; Among them, the historical transaction chain data is collected from the enterprise's internal business systems, such as order management systems, financial systems, logistics systems, etc., covering transaction participant information and transaction relationship information; the transaction participant information includes enterprise name, credit rating, and industry category; the transaction relationship information includes transaction time, transaction amount, transaction commodity type, and transaction quantity, where the participants include suppliers, manufacturers, distributors, and retailers.

[0020] Furthermore, the process of constructing the transaction chain knowledge graph includes: (1) Collect multi-source heterogeneous historical supply chain transaction data and perform preprocessing; (2) Define the ontology model of the knowledge graph, including node types and edge types; (3) According to the preprocessed historical supply chain transaction data, add the transaction participants as nodes to the ontology model of the knowledge graph, connect the corresponding nodes with the transaction relationships as edges, and record the transaction information on the edges, and form a transaction chain knowledge graph through the connections between nodes.

[0021] Process of assigning dynamic attributes to nodes: (1) Transaction frequency: Count the number of transactions of each node within a preset time window; (2) Transaction amount volatility: Calculate the ratio of the standard deviation to the average value of the transaction amount of the node within a preset time window; (3)Credit rating change rate: Record the change of the node credit rating and calculate the change range of the credit rating within a preset time.

[0022] Edge dynamic attribute assignment process: (1)Transaction time interval: Calculate the time difference between two adjacent transactions; (2)Transaction quantity change rate: Calculate the change range of the transaction quantity on the edge within a preset time; (3)Transaction price volatility: Calculate the ratio of the standard deviation to the average value of the transaction price on the edge.

[0023] It should be noted that the dynamic attributes are extracted from the knowledge graph. According to the constructed transaction chain knowledge graph, the dynamic attribute data of each node and edge are extracted using the graph query language. For example, use the Cypher query statement to obtain the transaction frequency of the supplier node A in the past month.

[0024] S2: Real-time collect the transaction data in the transaction chain, map the transaction participants and transaction relationships into low-dimensional vectors through the temporal knowledge graph embedding algorithm, generate a dynamically updated knowledge graph representation, and update the transaction entropy values of the node and edge dynamic attributes; S3: Based on the dynamically updated knowledge graph representation, calculate the anomaly score of each node through the anomaly detection method based on topology and features, and identify potential anomaly nodes in combination with the topological structure of the transaction chain; S4: Collect the transaction chain context information of the potential anomaly nodes and their neighborhood nodes, conduct secondary verification on the potential anomaly nodes, and output the final list of anomaly nodes.

[0025] The transaction chain context information includes the transaction time series, the change trend of the transaction amount, and the change of the relationship between the transaction participants.

[0026] Furthermore, the specific steps of S4 include: S4.1: Extract the relevant information of the potential anomaly nodes and their neighborhood nodes from the transaction chain knowledge graph, including transaction records and node attributes; S4.2: According to the collected transaction chain context information, extract the node transaction behavior characteristics and conduct standardization processing. The node transaction behavior characteristics include transaction amount, transaction frequency, and transaction time distribution; S4.3: Divide the data set into a training set and a test set, use the training set to train the secondary verification model, and obtain the trained secondary verification model; S4.4: Input the data of the potential anomaly nodes and their neighborhood nodes to be verified into the trained model to obtain the verification result; S4.5: Generate the final list of anomaly nodes according to the results of the secondary verification.

[0027] The method for identifying abnormal nodes in a transaction chain enhanced by a knowledge graph further includes: Visualize and output the final list of abnormal nodes and their associated transaction paths, and generate an abnormal analysis report.

[0028] Embodiment 2 Please refer to Figure 3 , in this embodiment, the steps of calculating the transaction entropy value of the dynamic attribute values of nodes and edges in S1 include: S1.1: Create an empty dictionary A, dictionary B, and dictionary C; the dictionaries A, B, and C are in the form of key-value pairs, and the keys in dictionary A are node identifiers, the keys in dictionary B are counterparty node identifiers, and the keys in dictionary C are edge identifiers, with the values initialized to 0. Among them, the edge identifier is composed of the starting node and the ending node; S1.2: For each node i in the transaction chain knowledge graph, count the number of transactions between this node and each counterparty , and obtain the transaction probability between this node and each counterparty by dividing the number of transactions between this node and any counterparty by the total number of transactions of this node , and store it in dictionary B, where j represents the counterparty node identifier; Among them, in the transaction chain knowledge graph, the counterparty refers to another node that directly has a transaction relationship with a specific node. It should be noted that the transaction chain knowledge graph is essentially a network composed of nodes and edges, where the nodes represent transaction participants and the edges represent the transaction relationships between these participants. For example, in a simple supply chain transaction chain knowledge graph, there are nodes such as suppliers, manufacturers, distributors, and retailers. When a supplier provides raw materials to a manufacturer, the supplier and the manufacturer are counterparties to each other, and there is a transaction edge between them. And counting the number of transactions between this node and each counterparty is to calculate the frequency of transactions between the specified node and the remaining nodes that have a transaction relationship with it.

[0029] Exemplarily, assume that in the transaction chain knowledge graph, it contains three nodes: retailer , wholesaler , manufacturer , and their transaction relationships are as follows: has three transactions, which are the first transaction amount of 1000 yuan, the second transaction amount of 1500 yuan, and the third transaction amount of 1200 yuan; has two transactions, which are the first transaction amount of 800 yuan and the second transaction amount of 900 yuan, then: For node : The counterparties are and ; By checking the transaction records, it is found that and had 3 transactions between them. So and had 3 transactions; Similarly, and had 2 transactions between them. So and had 2 transactions.

[0030] For node : The counterparty is only ; and had 3 transactions between them. So and had 3 transactions.

[0031] For node : The counterparty is only ; and had 2 transactions between them. So and had 2 transactions.

[0032] S1.3: Traverse dictionary B. For the transaction probability of each counterparty, obtain the transaction entropy value of node i according to the information entropy formula , and store it in dictionary A, where log represents the logarithmic function; S1.4: Take the calculated transaction entropy value of node i as a new attribute and add it to the nodes of the transaction chain knowledge graph; Furthermore, the specific steps of S1.4 include: S1.41: Obtain the transaction entropy value of node i; S1.42: Determine the storage method of the transaction chain knowledge graph; S1.43: In the selected storage method, find the record corresponding to node i, and add the calculated transaction entropy value as a new attribute to the attribute list of this node.

[0033] S1.5: For each edge k in the transaction chain knowledge graph, by traversing all the transaction records corresponding to this edge, count the distribution of different transaction attributes to obtain the transaction amount interval distribution dictionary and the transaction type distribution dictionary; Further, when counting the distribution of different transaction attributes, for the transaction amount, the transaction amount needs to be divided into intervals, and then the number of transactions within each interval of the transaction amount is counted to obtain a transaction amount interval distribution dictionary. Among them, the key in the transaction amount interval distribution dictionary is the amount interval, and the value is the number of transactions.

[0034] For the transaction type, the number of transactions of each transaction type needs to be counted to obtain a transaction type distribution dictionary. Among them, the key in the transaction type distribution dictionary is the transaction type, and the value is the number of transactions.

[0035] S1.6: By calculating the number of transactions in any transaction amount interval divided by the sum of all values in the transaction amount interval distribution dictionary of this side, the probability distribution of this transaction amount interval is obtained ; S1.7: By calculating the number of transactions of the current type divided by the sum of all values in the transaction type distribution dictionary of this side, the probability distribution of this transaction type is obtained ; S1.8: Calculate the transaction amount entropy value and the transaction type entropy value respectively according to the information entropy formula in S1.3, and combine the preset weights to obtain the edge transaction entropy value through weighted summation, and add the edge transaction entropy value as a new attribute to the edge of the transaction chain knowledge graph. At the same time, store the edge transaction entropy value in the dictionary C.

[0036] Embodiment 3 Please refer to Figure 4 , the specific steps of S2 in this embodiment include: S2.1: Obtain transaction data from the transaction system in real time and perform preprocessing; S2.2: Use the dynamic relationship modeling method to model the preprocessed transaction data into a dynamic graph with timestamps , where I represents the set of nodes, represents the set of edges, represents the set of dynamic attributes of nodes and edges, and t represents the current time; Further, the specific steps of S2.2 include: S2.21: Obtain the preprocessed transaction data, and identify entities and relationships from the preprocessed transaction data; the entity is the transaction participant; the relationship is the transaction type; S2.22: Use all different entities involved in the preprocessed transaction data as elements of the node set I, and initialize an empty edge set and a timestamp set ; S2.23: Traverse the preprocessed transaction data in chronological order. For each transaction record, add the corresponding edges to the edge set and record the corresponding timestamps in the timestamp set to obtain a dynamic graph with timestamps.

[0037] S2.3: Use the temporal knowledge graph embedding algorithm to map the nodes and edges in the dynamic graph with timestamps into low-dimensional vectors, obtaining node low-dimensional vectors and edge low-dimensional vectors ; S2.4: Set the batch threshold for data collection. Each time a new batch of transaction data is collected, trigger an update of the embedding vectors, and integrate the newly collected transaction data with the existing dynamic graph data with timestamps, and re-execute S2.3 to obtain updated node low-dimensional vectors and edge low-dimensional vectors; Furthermore, the specific steps of S2.4 include: S2.41: Set the batch threshold M according to the processing capacity of the system and the frequency of data updates. Here, the batch threshold means that when M new transaction data are reached, trigger an update of the embedding vectors; S2.42: Configure a counter in the data collection system; S2.43: Collect new transaction data. Each time a new transaction data is collected, increment the counter. When the counter reaches the batch threshold M, trigger the embedding vector update process and reset the counter; S2.44: Integrate the newly collected transaction data with the existing dynamic graph data with timestamps, including: Add the entities and relationships in the new data to the node set I and the edge set and record the corresponding timestamps in the timestamp set ; S2.45: Re-execute the embedding vector update process of the dynamic graph in S2.3 to update the low-dimensional vector representations of the nodes and edges.

[0038] S2.5: Set the time window T. For node i, count the set of trading counterparts within the time window T, as well as the number of transactions between this node and each trading counterpart , and recalculate the trading probability and the trading entropy value according to S1.2 and S1.3; S2.6: For each edge k, count the distribution dictionary of transaction amount intervals and the distribution dictionary of transaction types within the time window T, and recalculate the edge trading entropy value according to S1.5 - S1.8; S2.7: Add the recalculated node trading entropy value and the edge trading entropy value as new attributes to the dynamic graph, and update the dynamic attribute sets of nodes and edges .

[0039] Furthermore, the specific steps of S2.7 include: S2.71: Obtain the recalculated node trading entropy value and the edge trading entropy value ; S2.72: Determine the storage method of the dynamic graph; S2.73: In the selected storage method, find the record corresponding to node i, and add the calculated node trading entropy value to the dynamic attribute set of node i; S2.74: Find the record corresponding to edge k, and add the calculated edge trading entropy value to the dynamic attribute set of the edge.

[0040] The specific steps of S2.3 include: S2.31: Collect dynamic graph data with timestamps; S2.32: Set the embedding dimension d, and create a set of node low-dimensional vectors and a set of edge low-dimensional vectors , and randomly initialize each element in the set of node low-dimensional vectors and the set of edge low-dimensional vectors to a random number in the interval [-0.01, 0.01]. After random initialization, take out the vectors corresponding to the head node h and the tail node e from the set of node low-dimensional vectors as the head node low-dimensional vector , tail node low-dimensional vector , where represents the low-dimensional vector of node i, represents the size of the node set, represents the low-dimensional vector of the k-th edge, represents the size of the edge set; S2.33: Set the objective function , and the formula is: ; where represents the time-aware head node embedding vector, represents the time-aware edge embedding vector, represents the time-aware tail node embedding vector, represents the time regularization weight, represents the timestamp, represents the timestamp mapping function, and satisfies , represents the attention mechanism function, represents the attention mechanism weight, represents the structure constraint weight, represents the structure constraint term, and satisfies , represents element-wise multiplication, that is, the elements at corresponding positions are multiplied, represents the 1-norm, g represents the embedding dimension index, and ; Furthermore, , , , where, represents the time transformation matrix of the head and tail entities, represents the time transformation matrix of the tail entity, represents the time transformation matrix of the edge, and E represents the identity matrix.

[0041] It should be noted that in the traditional model, the node and relationship vectors are fixed and cannot capture dynamic changes. By introducing the timestamp mapping function the time information is incorporated into the vector to obtain , , , capturing the temporal evolution of nodes and relationships, enabling the same entity to have different representations at different times, enhancing the model's expressive power; at the same time, in the present invention, an attention mechanism is introduced to evaluate the temporal credibility of quadruples, automatically identifying transactions at abnormal time points, such as late-night large transfers with low credibility, alleviating the data sparsity problem, and assigning higher weights to low-frequency but important transactions; finally, adding a DistMult-style structure constraint term to retain the structural information of the static knowledge graph and avoid losing global relationships during temporal modeling.

[0042] S2.34: Extract positive and negative samples from the timestamped dynamic graph to form training samples; The positive sample is a quadruple of head node, edge, tail node, and timestamp; The negative sample is obtained by randomly replacing the head node, edge, or tail node in the positive sample; S2.35: Input the training samples into the objective function, calculate the loss value corresponding to each training sample, and update the low-dimensional vectors of the head node , tail node and edge in the real number space; Furthermore, the specific steps of S2.35 include: (1) Obtain the training samples from S2.34; (2) Randomly initialize a low - dimensional vector representation for each head node, tail node, and edge. These low - dimensional vectors are usually randomly initialized in the real - number space. For example, use a normal distribution or a uniform distribution. (3) Input the training samples into the objective function in S2.33 and calculate the loss value corresponding to each training sample. (4) According to the calculated loss value, use the backpropagation algorithm to calculate the gradient of each embedding vector. According to the gradient calculation result, update the low - dimensional vectors of the head nodes, tail nodes and edge low - dimensional vectors in the real - number space. Among them, the backpropagation algorithm is the prior art content in this field and is not the creative solution of this application, so it will not be elaborated here.

[0043] S2.36: After iterative training, when the objective function converges, obtain the optimized low - dimensional vectors of the nodes and edge low - dimensional vectors .

[0044] The specific steps of S3 include: S3.1: Load the dynamically updated knowledge - graph representation from S2, including the low - dimensional vector representations of nodes and edges, as well as the dynamic attribute values of nodes and edges. S3.2: Calculate the degree centrality of each node , where u represents the neighbor node of node i, represents the neighbor set of node i, which contains all the nodes directly connected to node i; the degree centrality represents the number of connections between a node and its neighbor nodes. It should be noted that is to count each neighbor node u in the neighbor set of node i. If the neighbor node u is in the neighbor set , then add 1.

[0045] S3.3: Calculate the closeness centrality of node i , where represents the shortest - path distance from node i to node s, V represents the set of all nodes in the dynamically updated knowledge graph, represents traversing all nodes s except node i in the dynamically updated knowledge graph; the closeness centrality represents the average distance from a node to all non - self nodes in the dynamically updated knowledge graph. S3.4: Calculate the betweenness centrality of node i , where represents the number of times node i passes through in the shortest path from node s to node l, Denote the number of the shortest paths from node s to node l; the betweenness centrality represents the importance of a node as a mediator of the shortest paths between non-self nodes; S3.5: Calculate the embedding vector distance between node i and its neighbor node u , where and represent the embedding vectors of node i and its neighbor node u respectively; S3.6: Perform weighted summation on , , , of each node to obtain the anomaly score of each node; S3.7: Set an anomaly score threshold according to the historical transaction chain data; If , mark the node corresponding to this anomaly score as a potential anomaly node.

[0046] The secondary verification in step S4 includes: Perform logical verification on the transaction behavior of potential anomaly nodes through a rule engine, where the rule engine is the prior art content in this field and is not the creative solution of this application, so it will not be elaborated here; Perform cross-verification on the associated transactions of potential anomaly nodes in combination with external data sources; Further, performing cross-verification on the associated transactions of potential anomaly nodes in combination with external data sources includes: (1) Determine the potential anomaly nodes to be verified and their associated transactions; (2) Collect and integrate internal data and external data, and perform data cleaning and preprocessing; Among them, collect internal data: Extract the data of potential anomaly nodes and their associated transactions from the internal database; Collect external data: Collect the transaction data related to potential anomaly nodes from external data sources, such as competitor data and industry reports.

[0047] (3) Associate the internal data and external data according to the matching rules, where the matching rules adopt unique identifiers such as customer ID, order number, and product number; (4) Adopt the mean analysis method to perform cross-verification on the associated transactions, where the mean analysis method is the prior art content in this field and is not the creative solution of this application, so it will not be elaborated here; (5) Analyze the results of the cross-verification to identify the true anomaly nodes and anomaly transactions; (6) Optimize product pricing and channel management according to the results of the cross-verification to improve production efficiency.

[0048] Dynamically evaluate the transaction patterns of potential abnormal nodes based on time series analysis. Time series analysis is prior art in this field and not a creative solution of this application, so it will not be elaborated here.

[0049] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make changes, modifications, substitutions, and variations to the above embodiments without departing from the spirit and scope of the present invention, and these all fall within the protection scope of the present invention.

Claims

1. A method for identifying abnormal nodes in a transaction chain enhanced by a knowledge graph, characterized in that, Including: S1: Construct a transaction chain knowledge graph, assign dynamic attributes to the nodes and edges in the transaction chain knowledge graph based on historical transaction chain data, and calculate the transaction entropy values of the dynamic attribute values of the nodes and edges; the nodes represent transaction participants, and the edges represent transaction relationships. S2: Collect transaction data in the transaction chain in real time, map transaction participants and transaction relationships into low-dimensional vectors through a temporal knowledge graph embedding algorithm, generate a dynamically updated knowledge graph representation, and update the transaction entropy values of the dynamic attribute values of the nodes and edges. S3: Based on the dynamically updated knowledge graph representation, calculate the anomaly score of each node through an anomaly detection method based on topology and features, and identify potential anomaly nodes in combination with the topological structure of the transaction chain. S4: Collect the transaction chain context information of potential anomaly nodes and their neighborhood nodes, perform secondary verification on the potential anomaly nodes, and output the final list of anomaly nodes.

2. The method for identifying abnormal nodes in a transaction chain enhanced by a knowledge graph according to claim 1, wherein, The steps of calculating the transaction entropy values of the dynamic attribute values of the nodes and edges in S1 include: S1.1: Create an empty dictionary A, dictionary B, and dictionary C; the dictionaries A, B, and C are in the form of key-value pairs, the key in dictionary A is the node identifier, the key in dictionary B is the counterparty node identifier, the key in dictionary C is the edge identifier, and the value is initialized to 0. S1.2: For each node i in the transaction chain knowledge graph, count the number of transactions between this node and each counterparty , and obtain the transaction probability between this node and each counterparty by dividing the number of transactions between this node and any counterparty by the total number of transactions of this node , and store it in dictionary B, where j represents the counterparty node identifier; S1.3: Traverse dictionary B. For the trading probability of each counterparty, obtain the trading entropy value of node i according to the information entropy formula and store it in dictionary A; S1.4: Add the calculated transaction entropy value of node i as a new attribute to the nodes of the transaction chain knowledge graph; S1.5: For each edge k in the transaction chain knowledge graph, by traversing all the transaction records corresponding to the edge, count the distribution of different transaction attributes to obtain the transaction amount interval distribution dictionary and the transaction type distribution dictionary. S1.6: Obtain the probability distribution of any transaction amount range by dividing the number of transactions in that transaction amount range by the sum of all values in the transaction amount range distribution dictionary of this side. ; S1.7: Obtain the probability distribution of the transaction type by calculating the number of transactions of the current type divided by the sum of all values in the transaction type distribution dictionary of this edge. ; S1.8: Calculate the entropy value of the transaction amount and the entropy value of the transaction type respectively according to the information entropy formula in S1.

3. Combine the preset weights and obtain the edge transaction entropy value through weighted summation. And use the entropy value of the transaction type , combine the preset weights, and obtain the edge transaction entropy value through weighted summation. Add the edge transaction entropy value as a new attribute to the edges of the transaction chain knowledge graph. At the same time, store the edge transaction entropy value in the dictionary C.

3. The method for identifying abnormal nodes in a transaction chain enhanced by a knowledge graph according to claim 2, wherein, The specific steps of S2 include: S2.1: Obtain transaction data from the transaction system in real time and perform preprocessing. S2.2: Model the preprocessed transaction data as a timestamped dynamic graph using a dynamic relationship modeling method , where I represents the set of nodes, represents the set of edges, represents the set of dynamic attributes of nodes and edges, and t represents the current time; S2.3: Map the nodes and edges in the dynamic graph with timestamps into low-dimensional vectors using the temporal knowledge graph embedding algorithm to obtain the low-dimensional vectors of the nodes and the low-dimensional vectors of the edges ; S2.4: Set the batch threshold for data collection. When a new batch of transaction data is collected, trigger an update of the embedding vector, integrate the newly collected transaction data with the existing dynamic graph data with timestamps, and re-execute S2.3 to obtain the updated low-dimensional vectors of the nodes and edges. S2.5: Set the time window T. For node i, count the set of trading counterparts within the time window T , and the number of transactions between this node and each trading counterpart . Then recalculate the trading probability and trading entropy value according to S1.2 and S1.3 ; S2.6: For each edge k, count the distribution dictionary of transaction amount intervals and the distribution dictionary of transaction types within the time window T, and recalculate the edge transaction entropy value according to S1.5 - S1.8 ; S2.7: The recalculated node trading entropy value and the edge trading entropy value are added as new attributes to the dynamic graph, and the dynamic attribute sets of nodes and edges are updated .

4. The method for identifying abnormal nodes in a transaction chain enhanced by a knowledge graph according to claim 3, wherein, The specific steps of S2.3 include: S2.31: Collect dynamic graph data with timestamps. S2.32: Set the embedding dimension d and create a set of low-dimensional node vectors and a set of low-dimensional edge vectors , and randomly initialize each element in the set of low-dimensional node vectors and the set of low-dimensional edge vectors to a random number in the interval [-0.01, 0.01]. After random initialization, take out the vectors corresponding to the head node h and the tail node e from the set of low-dimensional node vectors as the low-dimensional head node vector , the low-dimensional tail node vector , where represents the low-dimensional vector of node i, represents the size of the node set, represents the low-dimensional vector of the k-th edge, represents the size of the edge set; S2.33: Set the objective function , and the formula is: ; Among them, represents the head node embedding vector for time perception, represents the edge embedding vector for time perception, represents the tail node embedding vector for time perception, represents the time regularization weight, represents the timestamp, represents the timestamp mapping function and satisfies , represents the attention mechanism function, represents the attention mechanism weight, represents the structure constraint weight, represents the structure constraint term and satisfies , represents element-wise multiplication, represents the 1-norm, g represents the embedding dimension index, and .

5. The method for identifying abnormal nodes in a transaction chain enhanced by a knowledge graph according to claim 4, wherein, The specific steps of S2.3 also include: S2.34: Extract positive samples and negative samples from the dynamic graph with timestamps to form training samples; the positive samples are quadruples of head node, edge, tail node, and timestamp; the negative samples are obtained by randomly replacing the head node, edge, or tail node in the positive samples. S2.35: Input the training samples into the objective function, calculate the loss value corresponding to each training sample, and update the low-dimensional vectors of the head node, tail node, and edge in the real number space according to the loss value. , the low-dimensional vector of the tail node and the low-dimensional vector of the edge ; S2.36: After iterative training, when the objective function converges, the optimized low-dimensional node vectors and low-dimensional edge vectors are obtained.

6. The method for identifying abnormal nodes in a transaction chain enhanced by a knowledge graph according to claim 5, wherein The specific steps of S3 include: S3.1: Load the dynamically updated knowledge graph representation from S2, including the low-dimensional vector representations of the nodes and edges, and the dynamic attribute values of the nodes and edges. S3.2: Calculate the degree centrality of node i , where u represents the neighbor node of node i, represents the neighbor set of node i, which contains all the nodes directly connected to node i; S3.3: Calculate the closeness centrality of node i , where represents the shortest path distance from node i to node s, V represents the set of all nodes in the dynamically updated knowledge graph, represents traversing all nodes s except node i in the dynamically updated knowledge graph; S3.4: Calculate the betweenness centrality of node i , where represents the number of shortest paths from node s to node l that pass through node i, represents the number of shortest paths from node s to node l; S3.5: Calculate the embedding vector distance between node i and its neighbor node u , where and represent the embedding vectors of node i and its neighbor node u, respectively.

7. The method for identifying abnormal nodes in a transaction chain enhanced by a knowledge graph according to claim 6, wherein The specific steps of S3 also include: S3.6: Sum up the weighted values of , , , for each node to obtain the anomaly score ; S3.7: Set an abnormal score threshold according to the historical transaction chain data ; If , mark the node corresponding to the abnormal score as a potential abnormal node.

8. The method for identifying abnormal nodes in a transaction chain enhanced by a knowledge graph according to claim 7, characterized in that, The transaction chain context information in S4 includes transaction time series, transaction amount change trend, and relationship change between transaction participants.

9. The method for identifying abnormal nodes in a transaction chain enhanced by a knowledge graph according to claim 8, wherein, The secondary verification in step S4 includes: Logically verify the transaction behavior of potential anomaly nodes through a rule engine. Perform cross-verification on the associated transactions of potential anomaly nodes in combination with external data sources. Dynamically evaluate the transaction patterns of potential anomaly nodes based on time series analysis.

10. The method for identifying abnormal nodes in a transaction chain enhanced by a knowledge graph according to claim 9, wherein, Also including: Visually output the final list of anomaly nodes and their associated transaction paths, and generate an anomaly analysis report.

Citation Information

Patent Citations

  • A method for detecting abnormal behavior in financial transactions

    CN115345736B

  • Block chain abnormal node identification method based on dynamic graph convolutional neural network

    CN114240659A

  • Abnormal tissue identification method and device, electronic equipment and medium

    CN115062163A

  • Abnormal transaction detection method and device, computer equipment and storage medium

    CN118861934A

  • Financial transaction anomaly detection and risk assessment method and device based on artificial intelligence

    CN119693111A

Cited By

  • TIP selection method and device based on directed acyclic graph, equipment and medium

    CN120910312A

  • Real-time financial supervision data processing method and system

    CN120912330A

  • Real-time financial regulatory data processing method and system

    CN120912330B