A Context-Based Fraud Node Detection Method for Multi-Financial Relationship Graphs
The method improves fraud detection in financial relationship graphs by categorizing node neighborhoods, calculating global contextual features, and filtering noise edges, enhancing performance and speed in fraud node recognition.
Patent Information
- Application Number
- CN202411401960.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-09
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-10-09
AI Technical Summary
The existing financial fraud model has problems such as unreasonable data analysis, slow analysis speed and insufficient classification performance in e-commerce and digital currency trading scenarios.
The multi-financial relationship graph fraud node detection method is adopted based on context, and the graph structure data set is constructed, node neighborhood feature pooling and global context feature calculation are carried out, global context features are generated using attention mechanism, noise edges are filtered, edge weights are calculated, and the final feature representation of nodes is generated through multi-layer perceptrons, and the model is finally trained using cross-entropy loss function.
It improves the accuracy and computing efficiency of fraud node identification, and can effectively learn the difference in the feature of two types of nodes on the data set, identify noise edges and improve classification performance.
Smart Images

Figure CN119249279B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of graph data mining, and particularly relates to a method for detecting fraud nodes in a multiple financial relationship graph based on context. Background Art
[0002] With the acceleration of the transformation of banking services to the Internet, online financial services have become increasingly popular, but this has also provided more opportunities for fraudsters. Especially in the current new trend of financial fraud showing organization, gang formation, and concealment, fraudsters build complex financial relationship networks and take advantage of information asymmetry and regulatory loopholes to conduct illegal activities. Specifically, fraudsters may control multiple accounts or identities to commit fraud in financial services such as credit and payment. They take advantage of loopholes in the associated relationship graph, establish multiple closely connected nodes to hide their true identities and fraud purposes. For example, in some gang fraud cases, fraudsters share some information or devices to reduce costs, such as using the same mobile phone, IP address, or company phone number, etc., which makes them more difficult to detect and track. To address this issue, financial institutions and regulatory authorities need to rely on advanced technical means, such as associated relationship graphs and intelligent anti-fraud models, to identify and prevent fraud risks. By building a comprehensive risk management system and an efficient anti-fraud mechanism, suspicious fraud behaviors can be detected and processed in a timely manner to protect the legitimate rights and interests of consumers and the stability of the financial market.
[0003] Fraud detection based on graph neural networks is widely used in fields such as fraud loan detection and misleading review detection. For example, using a model based on graph neural networks to identify fraud behaviors affecting product ratings on the Taobao platform, identify users with lower credit scores on e-commerce lending platforms, and identify false reviews on review websites. Another example is to detect companies with illegal phenomena such as financial fraud in a multiple financial relationship graph established through various associated relationships between companies. In the problem of detecting fraud nodes in a multiple financial relationship graph, a financial institution is regarded as a node, the transaction relationship between institutions is regarded as an edge, illegal institutions belong to positive class nodes, and legal institutions belong to negative class nodes. When an edge is established between any two nodes that satisfy a certain relationship, a relationship graph under this relationship is formed. When the number of relationships exceeds one, a multi-relationship graph is formed. In recent years, fraud detection on multi-relationship graphs has attracted wide attention.
[0004] Semi-supervised learning is a machine learning method between supervised learning and unsupervised learning. It uses a small amount of labeled data and a large amount of unlabeled data to jointly train the model, aiming to fully explore the potential information and patterns in the unlabeled data to improve the performance and accuracy of the model. This method is particularly suitable for situations where labeled data is limited. By combining limited labeled resources and the potential of unlabeled data, it can improve model performance and explore the potential value of unlabeled data when the cost of data labeling is high. The research on semi-supervised learning can be traced back to the 1970s and has gradually received attention with the development of fields such as natural language processing, text classification, and computer vision. In the task of fraud detection, because the number of fraud samples themselves is small and difficult to detect and label, it is necessary to combine semi-supervised learning to train a model with better generalization ability.
[0005] Although some semi-supervised financial fraud models based on graph structures have emerged in financial scenarios such as e-commerce and digital currency transactions, these models have many limitations. The assumption that the real data distribution does not conform to the actual situation, as well as the need to improve the running speed and classification performance, are all areas that deserve our improvement. Summary of the invention
[0006] The purpose of the present invention is to provide a context-based method for detecting fraudulent nodes in multiple financial relationship graphs to solve the technical problems of limitations of financial fraud models in the prior art, unreasonable data analysis, and slow analysis speed.
[0007] In order to solve the above technical problems, the specific technical solutions of the present invention are as follows:
[0008] A method for detecting fraudulent nodes in a multi-financial relationship graph based on context, comprising the following steps:
[0009] Step S1: Preprocessing of graph structure data set constructed based on financial field data: In each relationship graph, the neighbors of the node in this relationship graph are divided into 5 sets, namely: the positive node set with first-order known labels, the negative node set with first-order known labels, the positive node set with second-order known labels, the negative node set with second-order known labels, and the node set with second-order unknown labels; after performing feature pooling on each neighborhood node set, the neighborhood pooling feature matrix of the node is obtained by combining;
[0010] Step S2: Calculate global context features for each node: Generate global context features for each node using the neighborhood average feature matrix and attention mechanism;
[0011] Step S3: In each relationship graph, the weight of the edge is calculated using the global context features to filter out the noise edges: the weight of the edge is calculated using the original features of the nodes on both sides of the edge and the global context features;
[0012] Step S4: In each relational graph, after filtering out the noisy edges, calculate the semantic features of each node under this relational graph;
[0013] Step S5: Horizontally splice the node's own features, the node's global context features, and the semantic features of the node under any relational graph, and obtain the final feature representation of the node through a multi-layer perceptron;
[0014] Step S6: Use the cross-entropy loss for backpropagation to update all trainable parameters to train the model: predict the probability that the node is a positive-class node or a negative-class node; calculate the loss function by combining the probability and the true label, and update the parameters to train the model through gradient backpropagation;
[0015] Step S7: Use the trained model to detect fraudulent nodes.
[0016] Furthermore, in Step S1,
[0017] The dataset is graph-structured data where is the node set, is the matrix composed of the original features of all nodes, N represents the number of nodes, and s represents the number of dimensions of the original features. is the edge set, ε r is the edge connection set under the specified relationship type r, is the label set of all nodes. Usually, represents the i-th node in the graph, is the label of this node; in the graph structure G, the types of nodes are consistent, the types of edges are diverse, and the categories of nodes are two types, and there is a serious imbalance in the number of the two types of nodes. For node v i , in the graph of relationship r, the neighbors can be divided into: the set of positive-class nodes with first-order known labels the set of negative-class nodes with first-order known labels the set of positive-class nodes with second-order known labels the set of negative-class nodes with second-order known labels the set of nodes with second-order unknown labels For node v i Perform average pooling on the node features within the set of all neighbor nodes under all relational graphs, and after combination, obtain the neighborhood pooling feature matrix of the i-th node v i R represents the total number of types of relational graphs. R represents the total number of types of relational graphs.
[0018] Furthermore, in Step S2, the global context features are calculated as follows:
[0019]
[0020]
[0021]
[0022]
[0023] where represents the global context feature of the i-th node; att i,j represents the weight between the i-th node and the j-th neighborhood. exp(·) is the exponential function applied to the numerical value e; represents the neighborhood pooling feature matrix of the i-th node, and E represents the standard normal distribution weight matrix, which is used to distinguish different types of neighborhood node sets; added to S i to obtain represents the neighborhood pooling feature matrix of the i-th node after the neighborhood is distinguished; is the neighborhood pooling feature matrix obtained by the i-th node in the j-th neighborhood after being distinguished by E, and is obtained by intercepting the specified dimension features from ; x i is the original feature of the i-th node, W q , W k , W v are all trainable weight matrices, Q i is the feature of x i in the mapping space of W q ; K i,j is in the mapping space of W k ; V i,j is in the mapping space of W v ;
[0024] Furthermore, the weight of the edge connection in step S3 is calculated as follows:
[0025]
[0026] where W w is the trainable weight matrix, x i is the original feature of the i-th node v i ; is the global context feature of the node v i ; x j is the original feature of the j-th node v j ; is the global context feature of the node v j ; Concat(·) is the concatenation function, represents concatenating these four vectors horizontally. Sigmoid(·) is the activation function, and the calculation formula is w i,j is node v i and node v j The weight between them ranges from 0 to 1.
[0027] Furthermore, in step S4, the semantic features are calculated in the following way:
[0028]
[0029]
[0030]
[0031] in is node v i Semantic features extracted from the graph with relation r. r (v i ) represents node v i The set of neighbors on the graph with relation r, N r (v j ) represents node v j The set of neighbors on the graph of relation r, w i,j is node v i and node v j The weight between i,k Represents node v i and node v k The weight between j,k Represents node v j and node v k The weight between j represents the original feature of the jth node, W r is a trainable weight matrix. Used to represent node v i The importance of Used to represent node v j degree of importance.
[0032] Furthermore, the final feature representation in step S5 is calculated in the following way:
[0033]
[0034] Where W c is the weight matrix, z i is the final feature representation of the i-th node.
[0035] Furthermore, in step S6, the probability of predicting whether the node is a positive node or a negative node after feature mapping is calculated by the following method:
[0036]
[0037] wherein is the predicted probability distribution of the i-th node, W1 and W2 are weight matrices, b1 and b2 are biases, and W1, W2, b1, and b2 all belong to trainable parameters. σ(·) is an activation function, and the ReLU activation function is selected. ReLU(x) = max(x, 0);
[0038] The loss function is calculated by combining the probability and the true label, and the parameters are updated through gradient backpropagation to train the model. The loss function for the classification task is calculated by the following formula:
[0039]
[0040] V train represents the set of all nodes in the training set, y i is the true label of the i-th node, is the predicted probability distribution of the i-th node, is the sum of the L2 norms of all trainable parameters involved in the model, which is used to constrain the model to prevent overfitting caused by overtraining. α is a hyperparameter with a positive value range;
[0041] The trainable parameters in the model are derived from the loss function, and an effective model is obtained through multiple gradient updates.
[0042] Compared with the prior art, the present invention has the following beneficial technical effects:
[0043] 1. Compared with the existing methods, in the data preprocessing stage of the present invention, the neighborhood features of all nodes are classified and pooled. This processing method enables the model to effectively learn the differences between the features of the two types of nodes through global context features even in the special case where the feature distribution spaces of positive and negative samples are close. Compared with the existing methods for distinguishing two types of nodes based on the original node features, the present invention effectively learns new discriminative features with a small number of parameters, and can improve the fraud node recognition effect on some datasets by using global context features.
[0044] 2. Compared with the existing methods, the present invention uses the commonalities of different relational graphs to guide feature learning under different relational graphs and calculates the weights of arbitrary edges. For different noise edge ratios within different relational graphs, previous methods attempted to distinguish noise edges by using node feature differences. However, this method is only applicable to the assumption that the feature differences between different types of nodes are large. The present invention can effectively distinguish and identify noise edges by using the learned global context features, achieving a better noise edge filtering effect.
[0045] 3. Compared with existing methods, the present invention has higher computing efficiency and classification performance, and shows better performance and faster running speed than the baseline model in multiple datasets. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0047] Figure 1 Schematic diagram of the steps of the fraud detection model according to the embodiment of the present invention.
[0048] Figure 2 Performance comparison chart of the present invention under different hyperparameters.
[0049] Figure 3 Interpretability analysis chart of the present invention for two datasets in the example
[0050] Figure 4 Running time comparison chart of the present invention and multiple baseline models on the test sets of two datasets in the example. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0051] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0052] This embodiment provides a multi-relational graph fraud detection model based on global context. Refer to Figure 1 , Figure 1Schematic diagram of the steps of the fraud detection model of an embodiment of the present invention. In all relationship graphs of step S1, the first-order neighbors and second-order neighbors of the target node (central node) are classified into: a positive node set with first-order known labels, a negative node set with first-order known labels, a positive node set with second-order known labels, a negative node set with second-order known labels, and a node set with second-order unknown labels. After obtaining the average features of each set, the global context features are learned for the target node through the attention mechanism of step S2. In steps S3-S4, in each relationship graph, the weight of the edge connecting any neighbor to the central node is calculated, and the feature representation in each relationship graph is learned through neighborhood aggregation. In steps S5-S6, the original features of the node, the global context features of the node, and the feature representations of the node in all relationship graphs are spliced and learned to obtain the final feature representation, and it is predicted whether the target node is a fraudulent node.
[0053] A global context-based multi-relationship graph fraud detection method specifically includes the following steps:
[0054] Step S1: Preprocessing of graph structure data set based on financial data. The data set is graph structure data. in is a node set, It is a matrix composed of the original features of all nodes, N represents the number of nodes, and s represents the number of dimensions of the original features; is the edge set, ε r It is the set of edges under the specified relationship type r, where R represents the total number of types of relationship graphs; is the label set of all nodes. Usually used represents the i-th node in the graph, is the label for this node. The dataset used is required to belong to a heterogeneous graph, in which the node types are consistent and the edge relationship types are diverse. The same node can have multiple relationships with any other node. When only focusing on the first relationship, it is considered that the other relationship edges do not exist, thus obtaining the first relationship graph; correspondingly, when only focusing on the Rth relationship, it is considered that the other relationship edges do not exist, thus obtaining the Rth relationship graph. When a node has neighbors based on relationship 1 and relationship R at the same time, it is considered that the node appears in both the first relationship graph and the Rth relationship graph. Depending on whether the node has fraudulent behavior or whether it is fraudulent information, the nodes are divided into two categories, and there is a serious imbalance in the number of nodes in the two categories. In the preprocessing stage, the label information is used to divide the neighborhood of all nodes into 5 node sets, namely: the positive class node set with first-order known labels The set of negative nodes with first-order known labels The set of positive nodes with second-order known labels Second-order known-label negative node set Second-order node set with unknown labels For node v i Average pooling is performed on the node features within the set of all neighborhood nodes under all relational graphs, and after combination, the neighborhood pooling feature matrix of the i-th node v is obtained i The neighborhood pooling feature matrices of all nodes are combined to obtain the global node neighborhood pooling feature matrix s represents the number of dimensions of the original features, R represents the total number of types of relational graphs, and N represents the number of nodes that have appeared in any relational graph
[0055] Step S2: Calculate the global context features for each node: Use the neighborhood pooling feature matrix and the attention mechanism to generate the global context features for each node, and calculate them in the following way
[0056]
[0057]
[0058]
[0059]
[0060] where represents the global context feature of the i-th node; att i,j represents the weight between the i-th node and the j-th neighborhood. exp(·) is the exponential function applied to the numerical e represents the neighborhood pooling feature matrix of the i-th node, E represents the standard normal distribution weight matrix, used to distinguish different types of neighborhood node sets; added to S i to obtain represents the neighborhood pooling feature matrix of the i-th node after the neighborhood is distinguished is the neighborhood pooling feature matrix obtained by the i-th node in the j-th neighborhood after being distinguished by E, obtained by intercepting the specified dimension features from . x i is the original feature of the i-th node, W q , W k , W v are all trainable weight matrices, Q i is the feature of x i in the mapping space of W q , K i,j is in the mapping space of W k , V i,j is in the mapping space of W v . By taking x i , Mapped to W respectively q , W k and W v After being mapped to the Q space, the W space, and the W space respectively, multiply the mapping of the i-th node in the Q space by the mapping of the j-th neighborhood in the K space to obtain the attention score of the i-th node to the j-th neighborhood. After the softmax operation, normalize the score to obtain the weight between the i-th node and the j-th neighborhood. Weightedly sum all the weights with the mapping of the corresponding neighborhood in the V space to obtain the global context feature of the i-th node. This process is the application of the attention mechanism.
[0061] Step S3: In each relational graph, calculate the weights of the edges using the global context features and filter out the noisy edges: Calculate the weights of the edges using the original features of the nodes on both sides of the edge and the global context features.
[0062]
[0063] where W w is a trainable weight matrix, x i is the original feature of the i-th node v i , is the global context feature of node v i , x j is the original feature of the j-th node v j , is the global context feature of node v j . Concat(·) is a concatenation function, indicating that these four vectors are concatenated horizontally. Sigmoid(·) is an activation function, and the calculation formula is w i,j is the weight between node v i and node v j , and its value range is between 0 and 1.
[0064] After assigning weights to each edge, if the weight is closer to 1, the information transmitted by this edge is retained; if the weight is closer to 0, the information transmitted by this edge is filtered.
[0065] Step S4: In each relational graph, after filtering out the noisy edges, calculate the semantic features of each node under this relational graph: The edge weights do not consider the local connection situation of the nodes during the calculation process. For a highly connected node, the importance of any neighbor to this node and the importance of this node to any neighbor are obviously lower than those of a low-connected node. Accordingly, referring to GCN, the edge weights are improved, and the product of the second norms of the importance degrees of two adjacent nodes is used as the denominator of the edge weight. Furthermore, the features of all neighbors are multiplied by the corresponding improved edge weights and then added together to obtain the semantic feature representation of the node. The calculation method of the semantic feature representation is as follows:
[0066]
[0067]
[0068]
[0069] where is the semantic feature extracted from the graph with the relationship r for node v i . N r (v i ) represents the neighbor set of node v i in the graph with the relationship r, N r (v j ) represents the neighbor set of node v j in the graph with the relationship r, w i,j is the weight between node v i and node v j , w i,k represents the weight between node v i and node v k , w j,k represents the weight between node v j and node v k , x j represents the original feature of the j-th node, and W r is a trainable weight matrix. is used to represent the importance degree of node v i , is used to represent the importance degree of node v j .
[0070] Learn new feature representations for all nodes under any relational graph to obtain the semantic features of the nodes under any relational graph.
[0071] Step S5: Horizontally splice the node's own features, the node's global context features, and the semantic features of the node under any relational graph, and obtain the final feature representation of the node through a multi-layer perceptron MLP.
[0072]
[0073] Among which W c is the weight matrix, and z i is the final feature representation of the i-th node.
[0074] Step S6: Training of deep learning model parameters: Training the model constructed according to the above steps. In the present invention, all the above features are concatenated with the initial features of the nodes and then the classifier is used to predict the categories of the nodes to determine whether the nodes belong to fraudulent nodes.
[0075] After the final feature representation of the node is multiplied by the weight matrix multiple times and added with biases, the number of dimensions becomes 2, which is used to predict the probability that the node is a positive-class node or a negative-class node, and is calculated by the following method:
[0076]
[0077] Among which is the predicted probability distribution of the i-th node, W1 and W2 are weight matrices, b1 and b2 are biases,
[0078] W1, W2, b1, and b2 all belong to trainable parameters. σ(·) is an activation function, and the ReLU activation function is selected.
[0079] ReLU(x) = max(x, 0). The softmax(·) is used to control the value of the predicted probability distribution between 0 and 1.
[0080] Combining the probability and the true label to calculate the loss function, and updating the parameters to train the model through gradient backpropagation. The cross-entropy loss function is calculated by the following formula:
[0081]
[0082] V train represents the set of all nodes in the training set, y i is the true label of the i-th node, is the predicted probability distribution of the i-th node, is the sum of the L2 norms of all the trainable parameters involved in the model, which is used to constrain the model to prevent overfitting caused by overtraining. α is a hyperparameter, and its value range is positive. The larger α is, the stronger the resistance of the training process to overfitting; the smaller α is, the weaker the resistance of the training process to overfitting. Taking the derivative of the trainable parameters in the model with respect to the loss function, and obtaining an effective model GIR through multiple gradient updates.
[0083] Step S7: Using the trained model to detect fraudulent nodes
[0084] Calculate the predicted probability distribution for any node using the trained model. For this 2D vector, by default, when the number corresponding to the first dimension is larger, the node is considered a normal node, and when the value corresponding to the second dimension is larger, the node is considered a fraudulent node.
[0085] The overall training process of the present invention is shown in Algorithm 1 below. Given a dataset that meets the requirements as the data, the algorithm first obtains the global context features of the nodes (line 3). Then, within each relational graph, the weights of the edges are calculated using the global context features and the features of the nodes themselves (line 5). A new semantic feature representation is learned for the nodes within each relational graph (line 6). After learning new feature representations in all relational graphs, the original features of the nodes themselves, the global context features, and the new features learned by the nodes in each relational graph are concatenated and processed to obtain the final feature representation (line 7). The new feature representation is used to predict whether the node belongs to a fraudulent node, and the loss function of the model is calculated, and the model is optimized through backpropagation (line 8).
[0086]
[0087] Next, the present invention details the dataset, the baseline model, and the parameter settings.
[0088] Dataset: This example uses three datasets, all constructed based on real-world financial scenario data. The YelpChi dataset is constructed using the hotel reservation data on the Yelp website. There are various deceptive reviews on the hotel reservation recommendation platform, trying to deceive consumers to increase the reservation numbers of certain hotels. The hotel reviews are regarded as nodes, and the connections between the nodes are formed according to three relationships: 1) R-U-R means that two reviews belong to the same user; 2) R-S-R means that two reviews for the same hotel give the same star rating; 3) R-T-R means that two reviews for the same hotel appear in the same time period, such as the same month, etc. The Amazon dataset is constructed based on the product review data on the Amazon shopping platform: 1) U-P-U means that two users give the same rating to the same product; 2) U-S-U means that two users give the same star rating to the same product within a one-week time difference; 3) U-V-U means that the similarity of the product reviews given by two users is extremely high. The Eliptic dataset describes the digital currency transactions between financial institutions, consisting of 2 relationships: 1) E-T-E means that there is a digital currency transfer in or out transaction between two financial institutions; 2) E-F-E means that two financial institutions show extremely high similarity in various financial-related indicators.
[0089] The statistical data of the dataset is shown in Table 1.
[0090] Table 1: Statistical Data Table of the Dataset
[0091]
[0092] The present invention uses several different graph neural networks as baseline models for comparison with the model GIR of the present invention. GCN processes graph data from the spectral domain and is a graph convolutional network that aggregates the first-order neighbor features of nodes. GraphSAGE is a graph neural network model that can be applied to large-scale networks by sampling the neighborhoods of unlabeled nodes and using various methods for neighborhood aggregation. CARE-GNN is a graph fraud detection model that uses reinforcement learning techniques to address the problems of fraud node feature disguise and relationship disguise. FRAUDRE is a fraud detection model composed of four parts, respectively used to address feature inconsistency, topological inconsistency, relationship inconsistency, and class imbalance problems. PC-GNN consists of two steps, namely "pick" and "choose", for node sampling to alleviate the class imbalance problem. RioGNN is an improved model based on CARE-GNN and also uses reinforcement learning for neighborhood sampling within different relational graphs. BHetero-GHRN and Bhomo-GHRN are models that use the relationship between graph frequency and graph heterogeneity to alleviate graph heterogeneity. GDN is a model that uses the maximum gradient to find the important features of nodes and enhances the feature differences of different category nodes for fraud detection. GAGA uses label information in the preprocessing stage, processes the graph structure into a standard structure, and then uses a transformer for fraud detection.
[0093] Training is performed on the dataset. In this case, the embedding dimension is 64, the learning rate of the model parameters is 0.0005, the maximum number of training epochs is 500, and an early stopping mechanism is adopted. When the model obtained in a certain round achieves the optimal performance for 100 consecutive epochs, the training is terminated, and the model of this round is used for testing. The experiments use AUC, F1-Macro, and AP as the accuracy metrics for node classification, and compare the present invention with other centralized methods. The highest score is shown in bold, and the second-highest score is shown underlined. Each experiment is trained for 5 rounds, and the average values and standard deviations are shown in Tables 2 and 3 as follows:
[0094] Table 2: Experimental Data Table for Performing Node Classification on YelpChi and Amazon Datasets
[0095]
[0096] Table 3: Experimental Data Table for Performing Node Classification on Eliptic Dataset
[0097]
[0098] As can be seen from Tables 2 and 3, compared with other baseline models in the three datasets, the model GIR of the present invention has much higher scores in the three evaluation metrics. The scores of AUC, AP, and F1-Macro are 1%-2%, 2%-21%, and 1%-10% higher than the optimal baseline model in the three datasets respectively, which proves the superior performance of the model GIR of the present invention. Each part of the model GIR of the present invention is removed separately. For example, removing the standard normal distribution weight matrix E gives GIR woE and removing the attention mechanism gives GIR woA and removing the edge weight calculation gives GIR woW and removing the global context features when calculating the edge weights gives GIR woWG and finally removing the semantic features obtained from each relational graph in the embedding generation process gives GIR woR and finally removing the global context features in the embedding generation process gives GIR woG The performance of the removed model is compared with GIR. Comparing with multiple baseline models on three metrics in two datasets, it is found that GIR's performance is almost stable in the top two after removing each part, indicating the effectiveness of each component of the GIR model.
[0099] The model is compared and analyzed on the YelpChi dataset for the F1 scores of two categories, and Table 4 is obtained:
[0100] Table 4: F1 score table for different categories on the YelpChi dataset
[0101]
[0102] As can be seen from Table 4, the F1 scores of GIR in the two categories are the largest, and F1-Fraud reaches 0.7189, proving that when the number of fraud nodes is unbalanced with the number of normal nodes, the model GIR can still learn the features to distinguish the two types of nodes. This phenomenon once again verifies the effectiveness of the model.
[0103] Controlling different training ratios and node label rates, training is carried out on two datasets, training 5 times in the same way, taking the average value and standard deviation, the highest score is represented in bold, and the second-best score is underlined, as shown in Table 5;
[0104] Table 5: Experimental data table for node classification at different training ratios and label rates
[0105]
[0106] The model trained by the present invention can still achieve an AUC of 97% on Amazon and 87% on YelpChi under the conditions of low training rate and low label rate, which proves that the model can still achieve excellent classification results under the conditions of low labels and low training rate.
[0107] Adjust the hyperparameters of the model. Set the Embedding Size within the range of {32, 48, 64}, set the LearningRate within the range of {0.0001, 0.0005, 0.001}, set the Early Stop within the range of {10, 20, 30}, and set the Weight Decay within the range of {0.0001, 0.0005, 0.001} to compare the training effects of the models. As Figure 2 shown, Figure 2 in (a) represents the change curves of the evaluation scores F1-score, GMean, and AP of the model during the process of the hidden layer dimension changing among 10, 20, and 30; Figure 2 in (b) represents the change curves of the evaluation scores F1-score, GMean, and AP of the model during the process of the early stop threshold changing among 10, 20, and 30; Figure 2 in (c) represents the change curves of the evaluation scores F1-score, GMean, and AP of the model during the process of the learning rate changing among 0.001, 0.005, and 0.01; Figure 2 in (d) represents the change curves of the evaluation scores F1-score, GMean, and AP of the model during the process of the weight decay coefficient changing among 0.0001, 0.0005, and 0.001. It is found from Figure 2 that the change range of each evaluation index of the model does not exceed 1%, which proves the robustness of the hyperparameters designed for the model.
[0108] Visualize the weights learned by the model in different datasets to obtain Figure 3 . Figure 3 in (a) represents the distribution of the attention coefficients of the model learned in the Amazon dataset for various types of neighbors, Figure 3 in (b) represents the distribution of the attention coefficients of the model learned in the Yelp dataset for various types of neighbors, Figure 3 in (c) represents the attention scores of the model for the positive and negative class nodes in the Amazon dataset for three relationship graphs, Figure 3 in (d) represents the attention scores of the model for the positive and negative class nodes in the Yelp dataset for three relationship graphs, Figure 3 in (e) represents the attention scores of the model for its own features, three relationship graphs, and global context features in the Amazon dataset during the model classification task,Figure 3 The (f) in it represents the attention scores of self - features, three relational graphs, and global context features in the Yelp dataset for the model classification task.
[0109] Figure 3 In figures (a) and (b), it shows that the attention coefficient of positive class nodes to positive class neighbors is higher, and the attention coefficient of negative class nodes to negative class neighbors is higher. This phenomenon is reflected in both Amazon and Yelp datasets. From figures (c) and (d), it is found that the attention scores of the Amazon dataset for the three relational graphs are almost equal, while there are obvious differences in the attention scores of the Yelp dataset for the three relational graphs. By analyzing the datasets, it is found that in Amazon, the number of nodes appearing in each relational graph is not much different, while in the Yelp dataset, the number of nodes appearing in the R - S - R relational graph is significantly more than that in the R - U - R graph. Correspondingly, the attention score of the model in the Yelp dataset for the R - S - R relational graph is also significantly greater than that for the R - U - R relational graph. Therefore, it is analyzed that the number of nodes in the relational graph has a significant impact on the attention score of the model for this relational graph. Figures (d) and (e) show that after using the global context features to guide the calculation of edge weights, the filtering of noisy edges for some relational graphs is indeed effective, making the weight coefficients of the features learned in the corresponding relational graphs larger during the final feature learning.
[0110] Record the average time for the model to be tested once on the test sets of different datasets and draw a bar chart to obtain Figure 4 . From Figure 4 it is found that the running time of the model GIR of the present invention is the shortest in both datasets, which proves that the model has a faster running speed compared to the baseline model.
[0111] It can be understood that the present invention is described through some embodiments. Those skilled in the art know that without departing from the spirit and scope of the present invention, various changes or equivalent replacements can be made to these features and embodiments. Additionally, under the teaching of the present invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application belong to the scope protected by the present invention.
Claims
1. A context-based method for detecting fraud nodes in a multi-financial relationship graph, characterized in that It includes the following steps: Step S1: Preprocessing of the graph structure dataset constructed based on financial domain data: In each relational graph, the neighbors of a node under this relational graph are divided into 5 sets, namely: the positive class node set with first-order known labels, the negative class node set with first-order known labels, the positive class node set with second-order known labels, the negative class node set with second-order known labels, and the node set with second-order unknown labels; after average pooling for each neighborhood node set, the neighborhood pooling feature matrix of the node is obtained by combination; Step S2: Calculate the global context feature for each node: Use the neighborhood pooling feature matrix and the attention mechanism to generate the global context feature for each node; Step S3: In each relational graph, calculate the weight of the edge using the global context feature and filter out the noisy edges: Calculate the weight of the edge using the original features and the global context features of the nodes on both sides of the edge connection; Step S4: In each relational graph, after filtering the noisy edges, calculate the semantic feature of each node under this relational graph; Step S5: Horizontally splice the node's own feature, the node's global context feature, and the semantic feature of the node under any relational graph, and obtain the final feature representation of the node through the multi-layer perceptron MLP; Step S6: After multiplying the final feature representation of the node by the weight matrix multiple times and adding the bias, the number of dimensions becomes 2, which is used to predict the probability that the node is a positive class node or a negative class node; calculate the loss function by combining the probability and the true label, and the loss function is the cross-entropy loss function; update the parameters to train the model through gradient backpropagation; use the cross-entropy loss for backpropagation to update all trainable parameters to train the model; Step S7: Use the trained model to detect fraud nodes; Use the trained model to calculate the predicted probability distribution for any node. For this 2D vector, by default, when the number corresponding to the first dimension is larger, the node is considered to belong to the normal node, and when the value corresponding to the second dimension is larger, the node is considered to belong to the fraud node.
2. The context-based multi-financial relationship graph fraud node detection method according to claim 1, wherein In step S1, The dataset is graph-structured data where is the set of nodes, is the matrix composed of the original features of all nodes, N represents the number of nodes, and s represents the number of dimensions of the original features; is the edge set, ε r is the set of connected edges under the specified relationship type r, R represents the total number of types of relationship graphs, is the set of labels of all nodes; use to represent the i-th node in the graph, is the label of this node; for node v i , in the graph of relationship r, the neighbors are divided into: the set of positive-class nodes with first-order known labels the set of negative-class nodes with first-order known labels the set of positive-class nodes with second-order known labels the set of negative-class nodes with second-order known labels the set of nodes with second-order unknown labels For node v i Average pooling is performed on the node features within the set of all neighbor nodes under all relationship graphs, and after combination, the neighborhood pooling feature matrix i of the i-th node v 3. The context-based multiple financial relationship graph fraud node detection method according to claim 2, wherein, The global context feature in step S2 is calculated in the following way: Among them represents the global context feature of the i-th node; att i,j Indicates the weight between the i-th node and the j-th neighborhood; exp(·) is the exponential function applied to the numerical value e; represents the neighborhood pooling feature matrix of the i-th node, and E represents the standard normal distribution weight matrix, which is used to distinguish different types of neighborhood node sets; added to S i to obtain after addition represents the neighborhood pooling feature matrix of the i-th node after the neighborhood is distinguished; is the neighborhood pooling feature matrix obtained by the i-th node in the j-th neighborhood after being distinguished by E, obtained by intercepting the specified dimension features; x i is the original feature of the i-th node, W q , W k , W v are all trainable weight matrices, Q i is x i in the W q mapping space feature, K i,j is in the W k mapping space feature, V i,j is in the W v mapping space feature.
4. The context-based multiple financial relationship graph fraud node detection method according to claim 3, characterized in that, The weight of the edge connection in step S3 is calculated in the following way: Among which W w is the trainable weight matrix, x i is the original feature of the i-th node v i , is the global context feature of the node v i , x j is the original feature of the j-th node v j , is the global context feature of the node v j ; Concat(·) is a concatenation function, indicating that x i , x j , these four vectors are concatenated horizontally; Sigmoid(·) is an activation function, and its calculation formula is w i,j is the weight between node v i and node v j and its value range is between 0 and 1.
5. The context-based multiple financial relationship graph fraud node detection method according to claim 4, characterized in that The semantic feature in step S4 is calculated in the following way: Among them is the semantic feature extracted from the graph with the relationship r for node v i ; N r (v i ) represents the neighbor set of node v i on the graph with relationship r, N r (v j ) represents the neighbor set of node v j on the graph with relationship r, w i,j is the weight between node v i and node v j , w i,k represents the weight between node v i and node v k , w j,k represents the weight between node v j and node v k , x j represents the original feature of the j-th node, W r is a trainable weight matrix; is used to represent the importance degree of node v i , is used to represent the importance degree of node v j .
6. The context-based multiple financial relationship graph fraud node detection method according to claim 5, characterized in that The final feature representation in step S5 is calculated in the following way: where W c is the weight matrix, and z i is the final feature representation of the i-th node.
7. The context-based multiple financial relationship graph fraud node detection method according to claim 6, characterized in that In step S6, after feature mapping of the final feature representation of the node, the probability that the node is a positive class node or a negative class node is predicted, and it is calculated in the following way: Among them is the predicted probability distribution of the i-th node, W1 and W2 are weight matrices, b1 and b2 are biases, and W1, W2, b1, and b2 all belong to trainable parameters; σ(·) is an activation function, and the ReLU activation function is selected; ReLU(x) = max(x, 0); Calculate the loss function by combining the probability and the true label, and update the parameters to train the model through gradient backpropagation. The loss function for the classification task is calculated by the following formula: V train represents the set of all nodes in the training set, y i is the true label of the i-th node, is the predicted probability distribution of the i-th node, is the sum of the L2 norms of all trainable parameters involved in the model, which is used to constrain the model to prevent overfitting caused by overtraining; α is a hyperparameter with a positive value range; Derive the trainable parameters in the model through the loss function, and obtain an effective model after multiple gradient updates.
Citation Information
Patent Citations
Heterogeneous graph embedding learning method based on attention mechanism
CN113095439A
Medical insurance fraud detection algorithm and system based on multilayer attention mechanism graph neural network
CN114463141A