Semi-supervised graph fraud detection model based on double-space heterogeneity learning

By integrating the attribute space heterogeneity and the structure space heterogeneity into a dual spatial heterogeneity learning model, the node representation is optimized and the category prototype is constructed, which solves the limitations of the existing model in processing heterogeneous graph data and improves the accuracy and efficiency of fraud detection.

CN120654137APending Publication Date: 2025-09-16NANJING AUDIT UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510718394.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing graph fraud detection models based on graph neural networks have limitations in node representation learning and are difficult to handle heterogeneous relationships in fraud graphs, resulting in high missed detection and false detection rates.

Method used

A semi-supervised graph fraud detection model based on dual-space heterogeneity learning is adopted. By fusing attribute space heterogeneity with structural space heterogeneity, the node representation is optimized using the attention mechanism and multi-relation aggregation strategy. The category prototype is constructed through the labels and embeddings of labeled nodes to guide the classification of unlabeled nodes. At the same time, label balanced sampling and semi-supervised learning are used to address the challenges of imbalance between positive and negative samples and scarce labeled data.

Benefits of technology

It significantly improves the fraud detection performance under complex heterogeneous graph data, reduces the missed detection rate and false detection rate, improves the accuracy of fraud detection, and can quickly process heterogeneous relationships in fraud graphs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120654137A_ABST
    Figure CN120654137A_ABST
Patent Text Reader

Abstract

The invention discloses a semi-supervised graph fraud detection model based on double-space heterogeneity learning. According to the method, the attribute space heterogeneity and the structure space heterogeneity are fused, node representation is optimized by using an attention mechanism and a multi-relation aggregation strategy, and classification of unlabeled nodes is guided through labeling and embedding of labeled nodes to construct a category prototype, so that the classification efficiency is improved. The challenges of imbalance of positive and negative samples and scarcity of labeled data are respectively coped with by adopting label balance sampling and semi-supervised learning, the fraud detection performance under complex heterogeneous graph data is remarkably improved, a heterogeneous relationship in a fraud graph can be quickly processed, the detection efficiency is improved, and the detection efficiency is improved based on a double-space heterogeneity learning mechanism. According to the method, node individual feature differences are disclosed, structural space heterogeneity is learned through label directed propagation, node network relation differences are disclosed, the individual feature differences and the network relation differences of the nodes in double spaces are disclosed by means of heterogeneity fusion, the omission ratio and the false detection rate of fraud detection are reduced, and the accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention specifically relates to a semi-supervised graph fraud detection model based on dual-space heterogeneity learning. Background Art

[0002] In today's highly competitive business environment, fraudulent activities such as fake transactions, collusion, ranking manipulation, and malicious ratings are commonplace. These behaviors not only undermine business integrity but also severely harm consumer interests. In recent years, with the widespread adoption of generative artificial intelligence, fraudulent activities have become more realistic and their patterns more subtle. Traditional fraud detection mechanisms based on rules, statistical analysis, and feature engineering have become ineffective in combating these activities.

[0003] Given the advantages of graph data in naturally expressing complex relationships and the powerful graph data processing capabilities of graph neural networks (GNNs), GNN-based graph fraud detection models have attracted considerable attention in recent years. These models model user relationships as graphs and use graph convolutional networks to learn user representations. By deeply analyzing user relationships, they can reveal the complex behavioral characteristics of fraudulent users.

[0004] However, existing GNN-based graph fraud detection models have limitations in node representation learning: models based on homogeneity assumptions have difficulty handling heterogeneous relationships in fraud graphs, while models based on heterogeneity assumptions usually only focus on a single attribute or structural space, resulting in high missed detection and false detection rates for fraud detection.

[0005] Therefore, it is necessary to invent a semi-supervised graph fraud detection model based on dual-space heterogeneity learning to solve the above problems. Summary of the Invention

[0006] The present invention aims to provide a semi-supervised graph fraud detection model based on dual-space heterogeneity learning. By fusing attribute space heterogeneity with structural space heterogeneity, the model optimizes node representation using an attention mechanism and a multi-relation aggregation strategy. Furthermore, the model constructs category prototypes using the labels and embeddings of labeled nodes to guide the classification of unlabeled nodes. Furthermore, the model adopts label-balanced sampling and semi-supervised learning to address the challenges of positive-negative sample imbalance and scarce labeled data, respectively. This significantly improves the fraud detection performance under complex heterogeneous graph data and can quickly process heterogeneous relationships in fraud graphs, thereby improving detection efficiency. Furthermore, a dual-space heterogeneity learning mechanism is proposed. This mechanism learns attribute space heterogeneity based on node attributes to reveal individual feature differences of nodes, and uses directed label propagation to learn structural space heterogeneity to reveal differences in node network relationships. With the help of heterogeneity fusion, the model reveals both individual feature differences and network relationship differences of nodes in the dual spaces, effectively avoiding the single-space limitations faced by traditional GNNs when processing heterogeneous graph data. This reduces the missed detection rate and false detection rate of fraud detection, improves the accuracy of fraud detection, and addresses the above-mentioned shortcomings in the technology.

[0007] To achieve the above objectives, the present invention provides the following technical solution: a semi-supervised graph fraud detection model based on dual-space heterogeneity learning, comprising the following steps:

[0008] Step 1: Given a multi-relation heterogeneous directed graph Let the relation subgraph G r ={V r ,X r ,E r ,Y r ,A r}, perform heterogeneity learning and obtain dual-space heterogeneity

[0009] Step 2: For the relationship subgraph G r Any tail node v∈V r , let N r (v) is the node set formed by all the first nodes corresponding to the tail node. Based on the dual spatial heterogeneity, the lth layer of convolution embeds the tail node and updates it. The overall embedding is obtained by multi-relation aggregation

[0010] Step 3: Overall embedding based on the Lth layer convolution Calculate the probability that the tail node v is a fraudulent node;

[0011] Step 4: Perform end-to-end training on the semi-supervised graph fraud detection model based on dual-space heterogeneity learning.

[0012] In the aforementioned semi-supervised graph fraud detection model based on dual-space heterogeneity learning, in step 1, V is the node set, and n = |V| represents the total number of nodes;

[0013] is the characteristic matrix, the i-th row vector in X is the i-th node v in V i The eigenvector of

[0014] d is the feature dimension;

[0015] E={E1,E2,…,E R} is the edge set;

[0016] R is the number of edge relation types. For each edge relation r∈{1,2,...,R}, and define E r 、 and are the edge set with the rth edge relationship, the homogeneous edge set, and the heterogeneous edge set respectively;

[0017] Y∈{0,1} n is the category vector, the i-th element y in Y i ∈{0,1} is the i-th node v in V i Category label of

[0018] A∈{0,1} n×n is the connection matrix, if there is a node v i from V i Points to the jth node v in V i A ij =1, otherwise A ij =0; i and j are the row and column numbers of matrix A respectively;

[0019] V r is the relation subgraph G r The node set of

[0020] X r is the relation subgraph G r The characteristic matrix of

[0021] E r is the relation subgraph G r The directed edge set of ;

[0022] Y r is the relationship subgraph G r The category label set of

[0023] A r is the relation subgraph G r The adjacency matrix of

[0024] is the relationship subgraph Gr The directed edge is formed by the first node u and the last node v.

[0025] The aforementioned semi-supervised graph fraud detection model based on dual-space heterogeneity learning, in step 1, given a multi-relation heterogeneous directed graph Let the relation subgraph G r ={V r ,X r ,E r ,Y r ,A r}, perform heterogeneity learning and obtain dual-space heterogeneity The specific steps are as follows:

[0026] 1.1. Calculate the relationship subgraph G r The l-th convolution layer has no effect on any directed edge Attribute spatial heterogeneity The specific formula is as follows:

[0027]

[0028] Among them, 1≤l≤L;

[0029] The first node u generated by the convolution layer l is in the relationship subgraph G r Embedding on ;

[0030] The tail node v generated by the convolution layer l is in the relation subgraph G r Embedding on ;

[0031] is the relationship subgraph G r The linear transformation matrix between the l-1th layer of convolution and the lth layer of convolution;

[0032] is the relationship subgraph G r The parameter matrix of the convolution layer l;

[0033] d l is the embedding dimension of the l-th convolution layer;

[0034] || is a vector concatenation operation;

[0035] -It is a vector element-by-element subtraction operation;

[0036] sigmoid() is the S-type activation function;

[0037] tanh() is the hyperbolic tangent activation function;

[0038] is the initial embedding of the first node u;

[0039] is the initial embedding of the tail node v;

[0040] 1.2. Utilize hinge loss Establish constraints so that directed edges Attribute spatial heterogeneity The edge label y that approaches this edge uv , the specific formula is as follows:

[0041]

[0042] in, is the pair-relation subgraph G r The labeled edge set is formed by sampling the positive and negative samples in a balanced manner according to the edge labels;

[0043] 1.3. Calculate the l-th layer convolution in the relation subgraph G r Perform directed propagation and update on the label to obtain the category label Y (l),r , the specific formula is as follows:

[0044]

[0045]

[0046] in, is the relationship subgraph G r The i-th order adjacency matrix of , and Q is the maximum order of the adjacency matrix;

[0047] is the relationship subgraph G r The 1 to Q order adjacency matrix The sum of , and 1≤i≤Q;

[0048] is a matrix The transpose of

[0049] For the l-th layer convolution in the relation subgraph G r The structural spatial heterogeneity matrix of all directed edges learned above;

[0050] For the learned directed edges The structural spatial heterogeneity of

[0051] ⊙ is Hadamard;

[0052] is a matrix The corresponding diagonal in-degree matrix;

[0053] The updated category label vector after the l-th layer label is propagated;

[0054] Y (0),r =[y v ] is the initial category label vector composed of true category labels;

[0055] Scalar y v represents the category of the tail node v, 1 represents fraud and 0 represents benign;

[0056] 1.4. Using Cross Entropy Loss Establish constraints to make the tail node prediction label based on the directed propagation of labels Approaching its true label y v , the specific formula is as follows:

[0057]

[0058] in, is the pair-relation subgraph G r The labeled nodes in the middle are the labeled node sets formed by sampling the positive and negative samples in a balanced manner according to the node labels;

[0059] y v is the label of the tail node v;

[0060] is the relationship subgraph G r The predicted label of the tail node v is obtained based on the directed propagation of the label;

[0061] 1.5. Apply the lth layer of convolution to attribute space heterogeneity and structural spatial heterogeneity Perform weighted fusion to form dual spatial heterogeneity The specific formula is as follows:

[0062]

[0063] in, For directed edges The spatial heterogeneity of attributes;

[0064] For directed edges The structural spatial heterogeneity of

[0065] α is Hyperparameters for weighted fusion;

[0066] β is Hyperparameters for weighted ensemble.

[0067] In the aforementioned semi-supervised graph fraud detection model based on dual-space heterogeneity learning, in step 1.2, the edge label y uvThe allocation rules are:

[0068] If there is an edge The labels of the head node u and the tail node v are the same, that is, y u =y v , then define its edge label y uv =1, called homogeneous edge;

[0069] If the labels are different, y u ≠y v , then define y uv =-1, which is called a heterogeneous edge.

[0070] The aforementioned semi-supervised graph fraud detection model based on dual-space heterogeneity learning, in step 2, for the relationship subgraph G r Any tail node v∈V r , let N r (v) is the node set formed by all the first nodes corresponding to the tail node. Based on the dual spatial heterogeneity, the lth layer of convolution embeds the tail node and updates it. The overall embedding is obtained by multi-relation aggregation The specific steps are as follows:

[0071] 2.1. Calculate in-degree neighbors u∈N based on self-attention mechanism r (v) Importance of the tail node v The specific formula is as follows:

[0072]

[0073] in, and is the trainable weight matrix;

[0074] is the embedding of the first node u obtained by the l-1th layer convolution update;

[0075] is the embedding of the tail node v obtained by the l-1th layer convolution update;

[0076] || is the splicing operation;

[0077] 2.2. Calculate the attention coefficient of the tail node v to each head node u The specific formula is as follows:

[0078]

[0079] Among them, the first node The loop variable of the inner loop;

[0080] First Node The importance of the tail node v;

[0081] LeakyReLU() is a leaky linear rectification function;

[0082] exp{} is the exponential operation;

[0083] 2.3. Update the embedding of the tail node v based on multi-head attention aggregation The specific formula is as follows:

[0084]

[0085] in, is the weight matrix of the kth attention head in the multi-head attention;

[0086] is the attention coefficient of the kth attention head of the tail node v to the first node u;

[0087] K is the total number of heads of multi-head attention;

[0088] It is the concatenation operation of the output of K attention heads;

[0089] 2.4. For the lth layer of convolution, embedding concatenation and mapping are used to form the overall embedding of the tail node v on all relations The specific formula is as follows:

[0090]

[0091] in, is the weight matrix;

[0092] 2.5 Based on overall embedding Use prototype learning to generate the classification probability of the tail node v

[0093] In the aforementioned semi-supervised graph fraud detection model based on dual spatial heterogeneity learning, in step 2.3, for the L-th layer convolution, weighted average aggregation is used for step 2.3 to obtain the embedding of the tail node v The specific formula is as follows:

[0094]

[0095] Among them, L is the last layer of convolution, that is, l=L.

[0096] The aforementioned semi-supervised graph fraud detection model based on dual-space heterogeneity learning, in step 2.5, is based on the overall embedding Use prototype learning to generate the classification probability of the tail node v The specific steps are as follows:

[0097] 2.5.1. The first convolution layer uses the labels and embeddings of the labeled nodes in the graph G to construct the category prototype. The specific formula is as follows:

[0098]

[0099] Among them, V F is the set of labeled fraud nodes;

[0100] V B is the set of labeled benign nodes;

[0101] is the embedding of the tail node v obtained by the l-th layer convolution;

[0102] The embedding of the fraud category prototype obtained for the l-th convolution layer;

[0103] Embedding of the benign category prototype obtained by the l-th convolution layer;

[0104] 2.5.2. Perform balanced sampling of positive and negative samples on the labeled nodes in graph G according to the node labels to form a training set And calculate V tr The Euclidean distance between the embedding of the middle tail node v and the embedding of the prototype of its own category is as follows:

[0105] If v is a fraudulent node, then:

[0106]

[0107] If v is a benign node, then:

[0108]

[0109] Among them, ||·||2 is the L2 norm;

[0110] is the Euclidean distance between the embedding of the fraud node v and the embedding of the fraud category prototype;

[0111] is the Euclidean distance between the embedding of the benign node v and the embedding of the benign category prototype;

[0112] 2.5.3. Convolutional layer l uses the Euclidean distance between the embedding of the tail node v and its own category prototype embedding to generate the classification probability of the tail node v And use cross entropy loss Establish constraints on the classification deviation of nodes. The specific formula is as follows:

[0113]

[0114]

[0115] Among them, y v is the label of the tail node v;

[0116] L(y v )∈{B,F} is the category label indicator of the tail node v.

[0117] In the aforementioned semi-supervised graph fraud detection model based on dual-space heterogeneity learning, in step 2.5.3, the indicator L(y v ) are as follows:

[0118] If y v = 0, then the tail node v belongs to the benign category, that is, the indicator L(y v ) is equal to B;

[0119] If y v =1, then the tail node v belongs to the fraud category, that is, the indicator L(y v ) is equal to F.

[0120] In the aforementioned semi-supervised graph fraud detection model based on dual spatial heterogeneity learning, in step 3, the overall embedding obtained based on the L-th layer convolution Calculate the probability p that the tail node v is ultimately predicted to be a fraudulent node v , the specific formula is as follows:

[0121]

[0122] And use cross entropy loss The node classification effect on the constrained fraud graph is as follows:

[0123]

[0124] in, is the overall embedding of the tail node v obtained by the L-th layer convolution.

[0125] In step 4, the aforementioned semi-supervised graph fraud detection model based on dual-space heterogeneity learning is trained end-to-end. The specific steps are as follows:

[0126] 4.1. Calculate the overall loss of the semi-supervised graph fraud detection model based on dual spatial heterogeneity learning The specific formula is as follows:

[0127]

[0128] Among them, γ1 is the cross entropy loss The weight parameter of

[0129] γ2 is the cross entropy loss The weight parameter of

[0130] γ3 is the cross entropy loss The weight parameter of

[0131] 4.2. Based on the overall loss value The stochastic gradient descent algorithm is used to calculate the gradient through the back propagation algorithm, and the parameters Perform iterative updates.

[0132] Compared with the prior art, the present invention has the following beneficial effects:

[0133] 1. This paper integrates attribute space heterogeneity with structural space heterogeneity, uses an attention mechanism and a multi-relation aggregation strategy to optimize node representation, and constructs category prototypes using the labels and embeddings of labeled nodes to guide the classification of unlabeled nodes. Furthermore, it uses label-balanced sampling and semi-supervised learning to address the challenges of positive and negative sample imbalance and scarce labeled data, respectively. This significantly improves fraud detection performance in complex heterogeneous graph data, can quickly process heterogeneous relationships in fraud graphs, and improves detection efficiency.

[0134] 2. The present invention proposes a dual-space heterogeneity learning mechanism, which learns attribute space heterogeneity based on node attributes to reveal the differences in individual node features, and uses label directed propagation to learn structural space heterogeneity to reveal the differences in node network relationships. With the help of heterogeneity fusion, the individual feature differences and network relationship differences of nodes in the dual space are revealed, effectively avoiding the single space limitation faced by traditional GNN when processing heterogeneous graph data, thereby reducing the missed detection rate and false detection rate of fraud detection and improving the accuracy of fraud detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0135] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments described in the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.

[0136] Figure 1 Flowchart of the present invention

[0137] Figure 2 Schematic diagram of the fraud detection performance of the present invention when using different heterogeneous representation weight combinations on the YelpChi dataset;

[0138] Figure 3 This is a schematic diagram of the fraud detection performance of the present invention when using different heterogeneous representation weight combinations on the Amazon dataset;

[0139] Figure 4 A line graph showing the impact of changes in the attribute space heterogeneity representation loss weight γ1 on the performance of the present invention on the YelpChi dataset;

[0140] Figure 5 A line graph showing the impact of changes in the attribute space heterogeneity representation loss weight γ1 on the performance of the present invention on the Amazon dataset;

[0141] Figure 6 A line graph showing the effect of changes in the loss weight γ2 of the structural spatial heterogeneity representation on the performance of the present invention on the YelpChi dataset;

[0142] Figure 7 A line graph showing the impact of changes in the loss weight γ2 of the structural spatial heterogeneity representation on the performance of the present invention on the Amazon dataset;

[0143] Figure 8 A line graph showing the effect of changes in the category prototype classification loss weight γ3 on the performance of the present invention on the YelpChi dataset;

[0144] Figure 9 A line graph showing the effect of changes in the category prototype classification loss weight γ3 on the performance of the present invention on the Amazon dataset;

[0145] Figure 10 This is a schematic diagram of the visualization of node classification of different models in the verification experiment of the present invention. DETAILED DESCRIPTION

[0146] In order to enable those skilled in the art to better understand the technical solution of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings.

[0147] The present invention provides Figure 1-10 The semi-supervised graph fraud detection model based on dual spatial heterogeneity learning shown in the figure includes the following steps:

[0148] Step 1: Given a multi-relation heterogeneous directed graph Let the relation subgraph G r ={V r , X r , E r , Y r , A r}, perform heterogeneity learning and obtain dual-space heterogeneity

[0149] Where V is the node set, n = |V| represents the total number of nodes;

[0150] is the characteristic matrix, the i-th row vector in X is the i-th node v in V i The eigenvector of

[0151] d is the feature dimension;

[0152] E={E1,E2,…,E R} is the edge set;

[0153] R is the number of edge relation types. For each edge relation r∈{1,2,...,R}, and define E r 、 and are the edge set with the rth edge relationship, the homogeneous edge set, and the heterogeneous edge set respectively;

[0154] Y∈{0,1} n is the category vector, the i-th element y in Y i ∈{0,1} is the i-th node v in V i Category label of

[0155] A∈{0,1} n×n is the connection matrix, if there is a node v i from V i Points to the jth node v in V i A ij =1, otherwise A ij =0; i and j are the row and column numbers of matrix A respectively;

[0156] V r is the relationship subgraph G r The node set of

[0157] X r is the relationship subgraph G r The characteristic matrix of

[0158] E r is the relationship subgraph G r The directed edge set of ;

[0159] Y r is the relationship subgraph G r The category label set of

[0160] A r is the relationship subgraph G r The adjacency matrix of

[0161] is the relationship subgraph G r The directed edge formed by the first node u and the last node v in ;

[0162] The specific steps of heterogeneity learning are as follows:

[0163] 1.1. Calculate the relationship subgraph G r The l-th convolution layer has no effect on any directed edge Attribute spatial heterogeneity The specific formula is as follows:

[0164]

[0165] Among them, 1≤l≤L;

[0166] The first node u generated by the convolution layer l is in the relationship subgraph G r Embedding on ;

[0167] The tail node v generated by the convolution layer l is in the relation subgraph G r Embedding on ;

[0168] is the relationship subgraph G r The linear transformation matrix between the l-1th layer of convolution and the lth layer of convolution;

[0169] is the relationship subgraph G r The parameter matrix of the convolution layer l;

[0170] d l is the embedding dimension of the l-th convolution layer;

[0171] || is a vector concatenation operation;

[0172] -It is a vector element-by-element subtraction operation;

[0173] sigmoid() is the S-type activation function;

[0174] tanh() is the hyperbolic tangent activation function;

[0175] is the initial embedding of the first node u;

[0176] is the initial embedding of the tail node v;

[0177] 1.2. Utilize hinge loss Establish constraints so that directed edges Attribute spatial heterogeneity The edge label y that approaches this edge uv , the specific formula is as follows:

[0178]

[0179] in, is the pair-relation subgraph G rThe labeled edge set is formed by sampling the positive and negative samples in a balanced manner according to the edge labels;

[0180] And the edge label y uv The allocation rules are:

[0181] If there is an edge The labels of the head node u and the tail node v are the same, that is, y u =y v , then define its edge label y uv =1, called homogeneous edge;

[0182] If the labels are different, y u ≠y v , then define y uv = -1, called a heterogeneous edge;

[0183] 1.3. Calculate the l-th layer convolution in the relation subgraph G r Perform directed propagation and update on the label to obtain the category label Y (l),r , the specific formula is as follows:

[0184]

[0185] in, is the relationship subgraph G r The i-th order adjacency matrix of , and Q is the maximum order of the adjacency matrix;

[0186] is the relationship subgraph G r The 1 to Q order adjacency matrix The sum of , and 1≤i≤Q;

[0187] is a matrix The transpose of

[0188] For the l-th layer convolution in the relation subgraph G r The structural spatial heterogeneity matrix of all directed edges learned above;

[0189] For the learned directed edges The structural spatial heterogeneity of

[0190] ⊙ is Hadamard;

[0191] is a matrix The corresponding diagonal in-degree matrix;

[0192] The updated category label vector after the l-th layer label is propagated;

[0193] Y(0),r =[y v ] is the initial category label vector composed of true category labels;

[0194] Scalar y v represents the category of the tail node v, 1 represents fraud and 0 represents benign;

[0195] 1.4. Using Cross Entropy Loss Establish constraints to make the tail node prediction label based on the directed propagation of labels Approaching its true label y v , the specific formula is as follows:

[0196]

[0197] in, is the pair-relation subgraph G r The labeled nodes in the middle are the labeled node sets formed by sampling the positive and negative samples in a balanced manner according to the node labels;

[0198] y v is the label of the tail node v;

[0199] is the relationship subgraph G r The predicted label of the tail node v is obtained based on the directed propagation of the label;

[0200] 1.5. Apply the lth layer of convolution to attribute space heterogeneity and structural spatial heterogeneity Perform weighted fusion to form dual spatial heterogeneity The specific formula is as follows:

[0201]

[0202] in, For directed edges The spatial heterogeneity of attributes;

[0203] For directed edges The structural spatial heterogeneity of

[0204] α is Hyperparameters for weighted fusion;

[0205] β is Hyperparameters for weighted fusion;

[0206] In this step 1, the attribute space heterogeneity is learned through the feature transformation network to reveal the individual feature differences of nodes;

[0207] And use label directed propagation to learn the structural spatial heterogeneity and reveal the network relationship differences of nodes;

[0208] With the help of dual-space heterogeneity fusion, the individual feature differences and network relationship differences of nodes hidden across space are revealed, effectively avoiding the problem of insufficient capture of node differences caused by the existing graph fraud detection model based on heterogeneity assumption being limited to a single space.

[0209] Step 2: For the relationship subgraph G r Any tail node v∈V r , let N r (v) is the node set formed by all the first nodes corresponding to the tail node. Based on the dual spatial heterogeneity, the lth layer of convolution embeds the tail node and updates it. The overall embedding is obtained by multi-relation aggregation The specific steps are as follows;

[0210] 2.1. Calculate in-degree neighbors u∈N based on self-attention mechanism r (v) Importance of the tail node v The specific formula is as follows:

[0211]

[0212] in, and is the trainable weight matrix;

[0213] is the embedding of the first node u obtained by the l-1th layer convolution update;

[0214] is the embedding of the tail node v obtained by the l-1th layer convolution update;

[0215] || is the splicing operation;

[0216] 2.2. Calculate the attention coefficient of the tail node v to each head node u The specific formula is as follows:

[0217]

[0218] Among them, the first node The loop variable of the inner loop;

[0219] First Node The importance of the tail node v;

[0220] LeakyReLU() is a leaky linear rectification function;

[0221] exp{} is the exponential operation;

[0222] 2.3. Update the embedding of the tail node v based on multi-head attention aggregation The specific formula is as follows:

[0223]

[0224] in, is the weight matrix of the kth attention head in the multi-head attention;

[0225] is the attention coefficient of the kth attention head of the tail node v to the first node u;

[0226] K is the total number of heads of multi-head attention;

[0227] It is the concatenation operation of the output of K attention heads;

[0228] In this step 2.3, Is a long vector, cut the vector into K segments, each segment is It can be used by one head, so it is called multi-head. At the same time, the length of each k-th segment is equal, which makes it easy to train in parallel.

[0229] Is a long vector, cut the vector into K segments, each segment is It can be used by one head, so it is called multi-head. At the same time, the length of each k-th segment is equal, which makes it easy to train in parallel.

[0230] In this step 2.3, in order to balance the feature output, for the Lth layer convolution, weighted average aggregation is used for step 2.3 to obtain the embedding of the tail node v The specific formula is as follows:

[0231]

[0232] Among them, L is the last layer of convolution, that is, l = L;

[0233] 2.4. For the lth layer of convolution, embedding concatenation and mapping are used to form the overall embedding of the tail node v on all relations The specific formula is as follows:

[0234]

[0235] in, is the weight matrix;

[0236] 2.5 Based on overall embedding Use prototype learning to generate the classification probability of the tail node v The specific steps are as follows:

[0237] 2.5.1. The first convolution layer uses the labels and embeddings of the labeled nodes in the graph G to construct the category prototype. The specific formula is as follows:

[0238]

[0239] Among them, V F is the set of labeled fraud nodes;

[0240] V B is the set of labeled benign nodes;

[0241] is the embedding of the tail node v obtained by the l-th layer convolution;

[0242] The embedding of the fraud category prototype obtained for the l-th convolution layer;

[0243] Embedding of the benign category prototype obtained by the l-th convolution layer;

[0244] 2.5.2. Perform balanced sampling of positive and negative samples on the labeled nodes in graph G according to the node labels to form a training set And calculate V tr The Euclidean distance between the embedding of the middle tail node v and the embedding of the prototype of its own category is as follows:

[0245] If v is a fraudulent node, then:

[0246]

[0247] If v is a benign node, then:

[0248]

[0249] Among them, ||·||2 is the L2 norm;

[0250] is the Euclidean distance between the embedding of the fraud node v and the embedding of the fraud category prototype;

[0251] is the Euclidean distance between the embedding of the benign node v and the embedding of the benign category prototype;

[0252] In this step 2.5.2, because L(y v )∈{B,F}, so for or

[0253] 2.5.3. Convolutional layer l uses the Euclidean distance between the embedding of the tail node v and its own category prototype embedding to generate the classification probability of the tail node v And use cross entropy loss Establish constraints on the classification deviation of nodes. The specific formula is as follows:

[0254]

[0255] Among them, y v is the label of the tail node v;

[0256] L(y v )∈{B,F} is the category label indicator of the tail node v;

[0257] At the same time, the indicator L(y v ) are as follows:

[0258] If y v = 0, then the tail node v belongs to the benign category, that is, the indicator L(y v ) is equal to B;

[0259] If y v =1, then the tail node v belongs to the fraud category, that is, the indicator L(y v ) is equal to F;

[0260] In step 2.5, we address the classification bias caused by heterogeneous connections in fraud detection by constructing category prototypes and constraining the spatial distribution of node representations. In heterogeneous graphs, prototype learning supplements the shortcomings of local neighborhood aggregation through geometric constraints, significantly improving the ability of this semi-supervised graph fraud detection model based on dual spatial heterogeneity learning to capture fraud patterns.

[0261] In step 2, for a single-relationship graph, the attention coefficient of the tail node to each head node is calculated based on the dual-space heterogeneity. When the tail node embedding is aggregated and updated through the multi-head attention mechanism, the semi-supervised graph fraud detection model based on dual-space heterogeneity learning can differentially process the information transmitted by head nodes with different heterogeneity:

[0262] For the first node with high heterogeneity, reduce its information transmission weight to avoid feature confusion;

[0263] For the first nodes with low heterogeneity, aggregation is enhanced to mine collusion patterns;

[0264] Further combined with the multi-relationship graph aggregation strategy, the tail node embedding can not only retain local feature differences but also take into account the global topological structure, thereby effectively dealing with collaboration and disguise among fraudsters.

[0265] Step 3: Overall embedding based on the Lth layer convolution Calculate the probability that the tail node v is a fraudulent node. The specific formula is as follows:

[0266]

[0267] And use cross entropy loss The node classification effect on the constrained fraud graph is as follows:

[0268]

[0269] in, is the overall embedding of the tail node v obtained by the L-th layer convolution;

[0270] In this step 3, the main classification loss function Directly supervising the fraud detection task can ensure node classification accuracy.

[0271] 4. Perform end-to-end training on the semi-supervised graph fraud detection model based on dual-space heterogeneity learning. The specific steps are as follows:

[0272] 4.1. Calculate the overall loss of the semi-supervised graph fraud detection model based on dual spatial heterogeneity learning The specific formula is as follows:

[0273]

[0274] Among them, γ1 is the cross entropy loss The weight parameter of

[0275] γ2 is the cross entropy loss The weight parameter of

[0276] γ3 is the cross entropy loss The weight parameter of

[0277] 4.2. Based on the overall loss value The stochastic gradient descent algorithm is used to calculate the gradient through the back propagation algorithm, and the parameters Perform iterative updates;

[0278] Among them, the stochastic gradient descent algorithm is a gradient-based optimization algorithm widely used in the fields of machine learning and deep learning. Its core idea is to update model parameters by randomly sampling the gradient of a single or a small number of samples to accelerate the training process and adapt to large-scale data sets. The goal is to gradually converge the loss function to the minimum value by iteratively adjusting the model parameters.

[0279] In this step 4, end-to-end learning is achieved through a multi-task joint optimization framework, and the dual spatial heterogeneity loss function and Modeling heterogeneous relationships in attribute space and structure space separately, capturing complex patterns of fraudulent behavior through edge-level feature difference constraints and label directed propagation learning, and prototype-guided classification loss function The class prototype is used to constrain the node embedding distribution, effectively alleviating the classification bias caused by fraudulent nodes being surrounded by benign nodes;

[0280] The above loss functions are fused through adaptive weights to jointly improve the detection performance of the model on complex heterogeneous graph data, especially showing significant advantages in dealing with category imbalance and fraud pattern diversity.

[0281] Verification experiment

[0282] 5.1 Experimental Setup

[0283] 5.1.1 Dataset

[0284] This experiment uses the YelpChi dataset and the Amazon dataset as datasets. The YelpChi dataset contains reviews of hotels and restaurants. All reviews are labeled as "legitimate" or "spam". The relationships between reviews can be divided into three categories:

[0285] (1) RUR (Comment-User-Comment): comments posted by the same user;

[0286] (2) RSR (Review-Star-Review): reviews that give the same star rating to the same product;

[0287] (3) RTR (Review-Time-Review): reviews of the same product published in the same month;

[0288] The Amazon dataset contains reviews of musical instruments. Some review publishers are labeled as "1" (fraudulent) or "0" (benign). The relationships between review publishers are also divided into three types:

[0289] (1) UPU (User-Product-User): users who have reviewed the same product;

[0290] (2) USU (user-star-user): users who gave the same star rating to the same product within a week;

[0291] (3) UVU (user-text similarity-user): users with the top 5% review text similarity;

[0292] Among them, Table 1 below shows some statistical information of the two data sets.

[0293] Table 1

[0294]

[0295] As shown in Table 1, the positive and negative samples in the two datasets are unbalanced. Therefore, the construction method of the training set, validation set and test set in this experiment is as follows: First, the same number of benign nodes as fraudulent nodes is randomly sampled from the dataset. Then, the fraudulent nodes and benign nodes are divided into training set, validation set and test set in a ratio of 4:2:4 respectively. Finally, the fraudulent node training set and the benign node training set are merged into the training set, the fraudulent node validation set and the benign node validation set are merged into the validation set, and the fraudulent node test set and the benign node test set are merged into the test set.

[0296] 5.1.2 Baseline Model

[0297] This experiment uses the following 9 models as baselines, including 3 classic graph neural networks, namely GCN, SGC, and GAT, 2 heterogeneous graph neural networks, namely GPR-GNN and FAGCN, and 4 graph fraud / anomaly detection models, namely CARE-GNN, PC-GNN, and H 2 -FDetector and GDN;

[0298] Among them, (1) GCN: updates node representation by aggregating node and neighbor information to capture the complex relationships and patterns of graph structures;

[0299] (2) SGC: By removing the nonlinear activation function and weight matrix, node feature aggregation is simplified by using only multiple products of the adjacency matrix;

[0300] (3) GAT: assigns different weights to nodes and their neighbors through the attention mechanism, adaptively aggregating information;

[0301] (4) GPR-GNN: By adaptively learning generalized PageRank weights and jointly optimizing the extraction of node features and graph topology information, it can adapt to homogeneous and heterogeneous relationship patterns and improve the accuracy of node classification;

[0302] (5) FAGCN: assigns weights to different frequency domain components through a frequency adaptive mechanism and dynamically aggregates node information;

[0303] (6) CARE-GNN: It filters neighbor nodes by label-aware similarity, dynamically selects the number of neighbors using reinforcement learning, and aggregates neighbor information of different relationships using a relationship-aware weighted method;

[0304] (7) PC-GNN: It constructs a relational subgraph by using the label-balanced sampler “Pick” to select nodes and edges, uses the neighbor sampler “Choose” to screen candidate neighbors, and aggregates neighbor and relationship information to optimize node representation, thereby effectively distinguishing fraudulent nodes from benign nodes in unbalanced graph data;

[0305] (8)H2 -FDetector: By identifying homogeneous and heterogeneous connections, designing aggregation strategies to propagate similarity and difference information respectively, and introducing prototype priors to guide fraudulent node identification;

[0306] (9) GDN: It extracts and maintains the key features of abnormal nodes to resist heterogeneity, constrains the features of normal nodes to enhance homogeneous connections, and combines the dynamically updated abnormal feature prototype vector to effectively distinguish fraudulent and benign nodes.

[0307] 5.1.3 Parameter settings

[0308] For DualH-FDNet, a semi-supervised graph fraud detection model based on dual spatial heterogeneity learning, this experiment sets the learning rate to 0.1, the weight decay coefficient to 0.00005, the dropout to 0.1, and the number of training iterations to 1000.

[0309] For the YelpChi dataset, this experiment sets γ1, γ2, and γ3 to 1.2, 1.2, and 1.0, respectively, and α and β to 0.5 and 0.5, respectively;

[0310] For the Amazon dataset, this experiment sets γ1, γ2, and γ3 to 0.4, 1.4, and 1.2, respectively, and α and β to 0.6 and 0.4, respectively;

[0311] To ensure fairness in the evaluation, the hyperparameters of the baseline are set as follows:

[0312] (1) CARE-GNN, PC-GNN, H2-FDetector, and GDN: Since the original literature used the same dataset, their optimal parameters were directly used;

[0313] (2) GCN, SGC, GAT, GPR-GNN, and FAGCN: Due to the differences in the original literature datasets, this experiment traverses and tests all parameter combinations and selects the checkpoints with the best performance on Amazon and YelpChi.

[0314] Among them, the parameter optimization of all models is completed based on the Adam optimizer.

[0315] 5.1.4 Model Implementation

[0316] The DualH-FDNet model is implemented based on the PyTorch 2.0.0 framework and uses Python 3.9.15.

[0317] The implementations of the baseline models, including GCN, SGC, GAT, GPR-GNN, FAGCN, CARE-GNN, PC-GNN, H2-FDetector, and GDN, are based on the source code included in the original papers;

[0318] All experiments were conducted on a server equipped with an NVIDIA V100-SXM2-32GB GPU, 43GB memory, and an Intel(R) Xeon(R) Platinum 8255C CPU (clocked at 2.50GHz).

[0319] 5.1.5 Evaluation Metrics

[0320] In fraud detection tasks, four indicators, Recall, F1-macro, AUC, and Gmean, are commonly used to evaluate model performance. Among them, Recall focuses on the recognition ability of positive samples, F1-macro measures the overall performance of the model on positive and negative samples, AUC reflects the classification ability of the model, and Gmean focuses on the balance of the model on positive and negative samples. A practical fraud detection model should achieve a balance among these four indicators.

[0321] 5.2 Performance Comparison

[0322] To evaluate the fraud detection performance of the DualH-FDNet model, we compared its fraud detection performance with nine baselines on the YelpChi and Amazon datasets based on four metrics: Recall, F1-macro, AUC, and Gmean.

[0323] Table 2

[0324]

[0325]

[0326] As shown in Table 2, DualH-FDNet performs well in all indicators on the YelpChi and Amazon datasets, especially in Recall and AUC indicators, significantly outperforming all baseline models. Specifically:

[0327] On the YelpChi dataset, DualH-FDNet achieved a Recall of 0.8495 and an AUC of 0.9154, which were 0.0081 and 0.0335 higher than the second-best H2-FDetector, respectively. On the Amazon dataset, DualH-FDNet achieved a Recall of 0.9515 and an AUC of 0.9743, which were 0.0576 and 0.0142 higher than the second-best H2-FDetector, respectively. This indicates that DualH-FDNet has significant advantages in capturing fraudulent nodes and distinguishing fraudulent from benign nodes.

[0328] Furthermore, the performance of the classic graph neural network models GCN, SGC, and GAT on both datasets was relatively weak. In particular, on the YelpChi dataset, SGC's Recall was only 0.1676, far lower than the other models. This suggests that methods that simplify node feature aggregation by simply multiplying the adjacency matrix multiple times are ineffective when dealing with complex graph structures. In contrast, GAT, by introducing an attention mechanism, achieved Recall and AUC of 0.5566 and 0.6164 on the YelpChi dataset, slightly outperforming GCN and SGC, but still inferior to heterogeneous graph neural networks and specially designed fraud detection models.

[0329] It can also be concluded that the heterogeneous graph neural network models GPR-GNN and FAGCN perform relatively stably on both datasets. GPR-GNN's Recall and AUC on the YelpChi dataset are 0.7756 and 0.8011, respectively, significantly outperforming the classic graph neural network model. FAGCN's Recall and AUC on the Amazon dataset are 0.8268 and 0.9157, respectively, also showing outstanding performance. This shows that through adaptive learning of generalized PageRank weights and frequency adaptation mechanisms, heterogeneous graph neural networks can better capture complex relationships in graph structures and improve the accuracy of fraud detection.

[0330] Meanwhile, the graph fraud / anomaly detection models CARE-GNN, PC-GNN, H2-FDetector, and GDN all performed well on both datasets. In particular, H2-FDetector and GDN achieved Recall and AUC of 0.8414, 0.9076, 0.8199, and 0.8313 on the YelpChi dataset, respectively, approaching DualH-FDNet. This demonstrates that by introducing strategies such as label-aware similarity metrics, label-balanced samplers, prototype priors, and dynamically updating prototype vectors of abnormal features, these models can effectively reduce the interference of feature and relationship camouflage on GNNs, improving the accuracy of fraud detection.

[0331] Finally, DualH-FDNet's excellent performance on both datasets is mainly attributed to its fusion of attribute space and structural space heterogeneity when identifying homogeneous and heterogeneous connections. This improvement enables DualH-FDNet to more accurately capture the key differences between fraudulent nodes and benign nodes, thereby achieving the best performance in all four indicators.

[0332] 5.3 Ablation Study

[0333] 5.3.1 Heterogeneity Ablation Study

[0334] In order to study the impact of different heterogeneity on fraud detection performance, this experiment designed 11 different weighted combinations for attribute space heterogeneity (weight α) and structure space heterogeneity (weight β). The specific results are as follows: Figure 2-3 As shown, Figure 2-3 Demonstrated the fraud detection performance of the DualH-FDNet model on the YelpChi and Amazon datasets using different combinations of heterogeneity weights;

[0335] from Figure 2 It can be seen that on the YelpChi dataset:

[0336] When α = 1.0 and β = 0.0, Recall reaches the highest value, at which point the model has the strongest ability to identify positive samples. As β increases, Recall gradually decreases, especially when β = 0.8, where Recall drops to the lowest value, indicating that the increase in structural spatial heterogeneity significantly reduces the model's ability to identify positive samples.

[0337] When α = 0.5 and β = 0.5, F1-macro reaches its peak, and the model performs best on both positive and negative samples. However, as β increases, F1-macro gradually decreases, indicating that the increase in structural spatial heterogeneity has an adverse effect on the overall classification performance of the model.

[0338] When α = 0.7 and β = 0.3, the AUC reaches the highest value, and the classification ability of the model is the strongest at this time. As β increases, the AUC gradually decreases, especially when β = 0.7, it reaches the lowest value, indicating that the increase in structural spatial heterogeneity significantly reduces the classification ability of the model.

[0339] When α = 0.9 and β = 0.1, Gmean reaches the highest value, and the model performs best in balancing positive and negative samples. As β increases, Gmean gradually decreases, indicating that the increase in structural spatial heterogeneity has a negative impact on the balance performance of the model.

[0340] On the YelpChi dataset, when α = 0.7 and β = 0.3, AUC reaches its highest value. At the same time, Recall, F1-macro, and Gmean are also at high levels. This shows that this combination has a relatively balanced performance on the four indicators. While ensuring high classification ability, it can also take into account the recognition ability of positive samples and the balanced performance of the model on positive and negative samples.

[0341] from Figure 3 It can be seen that on the Amazon dataset:

[0342] When α = 0.8 and β = 0.2, Recall reaches its highest value. At this time, the model has the strongest recognition ability for positive samples. As β increases, Recall gradually decreases, but the decrease is small, indicating that the influence of structural space heterogeneity on the recognition ability of positive samples is relatively small.

[0343] When α = 0.9 and β = 0.1, F1-macro reaches its peak. At this time, the model has the best comprehensive performance on positive and negative samples. As β increases, F1-macro gradually decreases, but the decrease is small, indicating that the impact of structural spatial heterogeneity on the overall classification performance of the model is relatively limited.

[0344] When α = 0.5 and β = 0.5, the AUC reaches the highest value, at which point the model has the strongest classification ability. As β increases, the AUC gradually decreases, but the decrease is small, indicating that the influence of structural spatial heterogeneity on the classification ability of the model is relatively small.

[0345] When α = 0.9 and β = 0.1, Gmean reaches its highest value. At this time, the model has the best balance performance on positive and negative samples. As β increases, GMean gradually decreases, but the decrease is small, indicating that the influence of structural space heterogeneity on the balance performance of the model is relatively limited.

[0346] On the Amazon dataset, when α = 0.9 and β = 0.1, Gmean reaches the highest value, while Recall, F1-macro, and AUC are also at high levels. This shows that the combination has a relatively balanced performance in the four indicators, and can take into account the recognition ability of positive samples, overall classification performance, and classification ability while ensuring a high balanced performance.

[0347] 5.3.2 Loss Function Ablation Study

[0348] This experiment conducts ablation experiments on the three loss function weights in the DualH-FDNet model, namely the weight γ1 of the attribute space heterogeneity loss, the weight γ2 of the structural space heterogeneity loss, and the weight γ3 of the category prototype classification loss, to analyze their impact on the model's fraud detection performance on two datasets: YelpChi and Amazon.

[0349] exist Figure 4-5 In this experiment, the weights of the structural space heterogeneity loss and the category prototype classification loss are fixed, and the weight of the attribute space heterogeneity loss is gradually increased. The performance of the model on the YelpChi and Amazon datasets is observed. The results are as follows:

[0350] from Figure 4It can be seen that on the YelpChi dataset, when γ1 increases from 0.2 to 2.0, the Recall value fluctuates between 0.8138 and 0.8631, showing an overall trend of first increasing and then decreasing. The maximum value occurs when γ1 = 2.0. The changes in F1-macro and Gmean are relatively stable, fluctuating between 0.6864 and 0.7538 and 0.7695 and 0.8401 respectively, indicating that the attribute space heterogeneity loss has little effect on the overall classification ability of the model and the balance of positive and negative samples. The AUC value fluctuates between 0.8664 and 0.9316, with the maximum value occurring when γ1 = 2.0, indicating that a higher attribute space heterogeneity loss weight helps improve the classification ability of the model.

[0351] from Figure 5 As can be seen, on the Amazon dataset, the Recall value fluctuates between 0.9075 and 0.9542, with the maximum value occurring when γ1 = 0.4, indicating that a moderate attribute space heterogeneity loss weight helps improve the model's ability to identify fraudulent nodes. F1-macro and Gmean change relatively steadily, fluctuating between 0.8114 and 0.8605 and 0.8986 and 0.9392, respectively, indicating that attribute space heterogeneity loss has little impact on the model's overall classification ability and the balance between positive and negative samples. The AUC value fluctuates between 0.9334 and 0.9727, with the maximum value occurring when γ1 = 0.4, indicating that a moderate attribute space heterogeneity loss weight helps improve the model's classification ability.

[0352] exist Figure 6-7 In this experiment, we fixed the weights of attribute space heterogeneity loss and category prototype classification loss, gradually increased the weight of structure space heterogeneity loss, and observed the performance of the model on the YelpChi and Amazon datasets. The results are as follows:

[0353] from Figure 6 It can be seen that on the YelpChi dataset, when γ2 increases from 0.2 to 2.0, the Recall value fluctuates between 0.7923 and 0.8558, with the maximum value occurring when γ2 = 1.4, indicating that a moderate structural space heterogeneity loss weight helps improve the model's ability to identify fraudulent nodes. The changes in F1-macro and Gmean are relatively stable, fluctuating between 0.6746 and 0.7513 and 0.7764 and 0.8446, respectively, indicating that the structural space heterogeneity loss has little effect on the overall classification ability of the model and the balance of positive and negative samples. The AUC value fluctuates between 0.8316 and 0.9223, with the maximum value occurring when γ2 = 1.0, indicating that a moderate structural space heterogeneity loss weight helps improve the classification ability of the model.

[0354] from Figure 7As can be seen, on the Amazon dataset, the Recall value fluctuates between 0.8901 and 0.9510, with the maximum value occurring when γ2 = 1.4, indicating that a moderate weighting of the structural spatial heterogeneity loss helps improve the model's ability to identify fraudulent nodes. F1-macro and Gmean change relatively steadily, fluctuating between 0.8016 and 0.8326 and 0.8952 and 0.9347, respectively, indicating that the structural spatial heterogeneity loss has little impact on the model's overall classification ability and the balance between positive and negative samples. The AUC value fluctuates between 0.9084 and 0.9813, with the maximum value occurring when γ2 = 1.0, indicating that a moderate weighting of the structural spatial heterogeneity loss helps improve the model's classification ability.

[0355] exist Figure 8-9 In this experiment, we fixed the weights of attribute space heterogeneity loss and structural space heterogeneity loss, gradually increased the weight of category prototype classification loss, and observed the performance of the model on the YelpChi and Amazon datasets. The results are as follows:

[0356] from Figure 8 It can be seen that on the YelpChi dataset, when γ3 increases from 0.2 to 2.0, the Recall value fluctuates between 0.7095 and 0.8535, with the maximum value occurring when γ3 = 1.2, indicating that a moderate class prototype classification loss weight helps improve the model's ability to identify fraudulent nodes. The changes in F1-macro and Gmean are relatively stable, fluctuating between 0.6925 and 0.7417 and 0.7334 and 0.8537, respectively, indicating that the class prototype classification loss has little effect on the overall classification ability of the model and the balance between positive and negative samples. The AUC value fluctuates between 0.8763 and 0.9234, with the maximum value occurring when γ3 = 1.4, indicating that a moderate class prototype classification loss weight helps improve the classification ability of the model.

[0357] from Figure 9 As can be seen, on the Amazon dataset, the Recall value fluctuates between 0.8368 and 0.9517, with the maximum value occurring when γ3 = 1.2, indicating that a moderate class prototype classification loss weight helps improve the model's ability to identify fraudulent nodes. F1-macro and Gmean change relatively steadily, fluctuating between 0.7273 and 0.8667 and 0.8924 and 0.9535, respectively, indicating that the class prototype classification loss has little impact on the model's overall classification ability and the balance between positive and negative samples. The AUC value fluctuates between 0.9300 and 0.9842, with the maximum value occurring when γ3 = 0.8, indicating that a moderate class prototype classification loss weight helps improve the model's classification ability.

[0358] Based on the above analysis, it can be concluded that the semi-supervised graph fraud detection model based on dual-space heterogeneity learning can show relatively ideal fraud detection performance on the YelpChi dataset and Amazon dataset when the attribute space heterogeneity loss weights are 1.2 and 0.4, respectively. In addition, when the structure space heterogeneity loss weights and the category prototype classification loss weights are 1.0 and 1.2, respectively, the semi-supervised graph fraud detection model based on dual-space heterogeneity learning can show relatively ideal fraud detection performance on both datasets. Finally, the balanced sampling strategy adopted during training effectively improves the model's adaptability to the imbalance of positive and negative samples.

[0359] 5.4 Visualization

[0360] To intuitively demonstrate the classification advantages of DualH-FDNet, this experiment selected five out of nine baselines for comparison: GCN and GAT, representatives of classic graph neural networks; FAGCN, a representative of heterogeneous graph neural networks; and CARE-GNN and H2-FDetector, representatives of graph fraud / anomaly detection models.

[0361] At the same time, the classification effects of the six models (including DualH-FDNet) were visualized on the YelpChi dataset. Due to the large number of nodes in the YelpChi dataset, this experiment selected its test set for display in order to facilitate visualization. That is, the six models were used to obtain the 32-dimensional representation of all nodes in the YelpChi test set. Then, the visualization tool t-SNE was used to map the obtained node representation to a two-dimensional space. The results are shown below. Figure 10 As shown:

[0362] Depend on Figure 10 As can be seen in Figure a, the initial features have limited ability to distinguish fraudulent and benign nodes. The distribution of fraudulent nodes (red) and benign nodes (blue) is relatively mixed, with no obvious separation.

[0363] Depend on Figure 10 As can be seen in Figure b, GCN performs poorly in distinguishing fraudulent nodes from benign nodes. The distribution of the two types of nodes is relatively mixed, and the separation effect is not obvious, especially in the central area where there is a lot of overlap, indicating that GCN is difficult to effectively capture the unique characteristics of fraudulent behavior.

[0364] Depend on Figure 10 As can be seen from Figure c, GAT has a certain effect in distinguishing fraudulent nodes from benign nodes, but the degree of separation is still not obvious enough, and some fraudulent nodes and benign nodes still have overlapping areas;

[0365] Depend on Figure 10As can be seen from Figure d in the figure, the classification effect of FAGCN is improved compared with GAT, and the separation degree of fraudulent nodes and benign nodes is better, but there are still some overlapping areas;

[0366] Depend on Figure 10 As can be seen from Figure e, CARE-GNN's classification effect is better than GCN and GAT. The distribution of fraudulent nodes and benign nodes shows a certain separation trend, but there is still some overlap in some areas, especially in the boundary areas, indicating that its ability to distinguish complex fraud patterns still has room for improvement;

[0367] Depend on Figure 10 As can be seen from the f-graph in the figure, the classification effect of H2-FDetector is further improved, the separation degree of fraudulent nodes and benign nodes is more obvious, and the overlapping area is reduced;

[0368] Depend on Figure 10 As can be seen from the g figure in the figure, the classification effect of DualH-FDNet is significantly better than that of the other five models. Fraudulent nodes and benign nodes form a clear separation boundary in two-dimensional space, with very little overlapping area. This result once again verifies that DualH-FDNet introduces dual spatial heterogeneity, which can more accurately capture the abnormal features of fraudulent nodes, thereby achieving efficient detection of fraudulent behavior.

[0369] In summary

[0370] 1. This paper integrates attribute space heterogeneity with structural space heterogeneity, uses an attention mechanism and a multi-relation aggregation strategy to optimize node representation, and constructs category prototypes using the labels and embeddings of labeled nodes to guide the classification of unlabeled nodes. Furthermore, it uses label-balanced sampling and semi-supervised learning to address the challenges of positive and negative sample imbalance and scarce labeled data, respectively. This significantly improves fraud detection performance in complex heterogeneous graph data, can quickly process heterogeneous relationships in fraud graphs, and improves detection efficiency.

[0371] 2. This paper proposes a dual-space heterogeneity learning mechanism. This mechanism learns attribute space heterogeneity based on node attributes to reveal differences in individual node features. It also uses label directed propagation to learn structural space heterogeneity to reveal differences in node network relationships. With the help of heterogeneity fusion, it reveals differences in individual features and network relationships of nodes in the dual space. This effectively avoids the single-space limitation faced by traditional GNNs when processing heterogeneous graph data, thereby reducing the missed detection rate and false detection rate of fraud detection and improving the accuracy of fraud detection.

[0372] 3. The present invention also conducts generalization tests and ablation studies. The fraud detection performance is generalized and tested on two public datasets based on four indicators. The results show that the performance of the semi-supervised graph fraud detection model based on dual-space heterogeneity learning is significantly better than nine baseline models, and it has a significant performance advantage in revealing complex fraud behaviors. In addition, the ablation study further verifies the contribution of attribute space heterogeneity and structural space heterogeneity to the model performance, proving the effectiveness of dual-space heterogeneity learning.

[0373] Dual-Space Heterophily Learning is a representation learning method for heterogeneous data. Its core is to capture the differences between nodes in the feature space (such as the diversity of node / edge types) and in the structure space (such as the differences in network connectivity between nodes) through collaborative modeling of the feature space and the graph structure space.

[0374] Semi-Supervised Graph Fraud Detection (SSD) is a graph-based fraud detection method that combines semi-supervised learning and graph analysis techniques. It uses structural information in the graph (such as nodes, edges, and relationship subgraphs) and a small amount of annotated data (such as known fraud samples) to identify potential fraudulent behavior.

[0375] The above description is merely illustrative of certain exemplary embodiments of the present invention. It goes without saying that those skilled in the art will be able to modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the above drawings and description are illustrative in nature and should not be construed as limiting the scope of protection of the claims.

Claims

1. A semi-supervised graph fraud detection model based on dual-space heterogeneity learning, characterized by: The following steps are involved: Step 1: Given a multi-relation heterogeneous directed graph Let the relation subgraph G r ={V r ,X r ,E r ,Y r ,A r }, perform heterogeneity learning and obtain dual-space heterogeneity Step 2: For the relationship subgraph G r Any tail node v∈V r , let N r (v) is the node set formed by all the first nodes corresponding to the tail node. Based on the dual spatial heterogeneity, the lth layer of convolution embeds the tail node and updates it. The overall embedding is obtained by multi-relation aggregation Step 3: Overall embedding based on the Lth layer convolution Calculate the probability that the tail node v is a fraudulent node; Step 4: Perform end-to-end training on the semi-supervised graph fraud detection model based on dual-space heterogeneity learning.

2. The semi-supervised graph fraud detection model based on dual-space heterogeneity learning according to claim 1 is characterized by: In step 1, V is the node set, n = |V| represents the total number of nodes; is the characteristic matrix, the i-th row vector in X is the i-th node v in V i The eigenvector of d is the feature dimension; E={E1,E2,…,E R } is the edge set; R is the number of edge relation types. For each edge relation r∈{1,2,...,R}, and define E r 、 and are the edge set with the rth edge relationship, the homogeneous edge set, and the heterogeneous edge set respectively; Y∈{0,1} n is a category vector, the i-th element yi∈{0,1} in Y is the i-th node v in V i Category label of A∈{0,1} n×n is the connection matrix, if there is a node v i from V i Points to the jth node v in V i A ij =1, otherwise A ij =0; i and j are the row and column numbers of matrix A respectively; V r is the relationship subgraph G r The node set of X r is the relationship subgraph G r The characteristic matrix of E r is the relationship subgraph G r The directed edge set of ; Y r is the relationship subgraph G r The category label set of A r is the relationship subgraph G r The adjacency matrix of is the relationship subgraph G r The directed edge is formed by the first node u and the last node v.

3. The semi-supervised graph fraud detection model based on dual-space heterogeneity learning according to claim 2 is characterized by: In step 1, given a multi-relation heterogeneous directed graph Let the relation subgraph G r ={V r ,X r ,E r ,Y r ,A r }, perform heterogeneity learning and obtain dual-space heterogeneity The specific steps are as follows: 1.

1. Calculate the relationship subgraph G r The l-th convolution layer has no effect on any directed edge Attribute spatial heterogeneity The specific formula is as follows: Among them, 1≤l≤L; The first node u generated by the convolution layer l is in the relationship subgraph G r Embedding on ; The tail node v generated by the convolution layer l is in the relation subgraph G r Embedding on ; is the relationship subgraph G r The linear transformation matrix between the l-1th layer of convolution and the lth layer of convolution; is the relationship subgraph G r The parameter matrix of the convolution layer l; d l is the embedding dimension of the l-th convolution layer; || is a vector concatenation operation; -It is a vector element-by-element subtraction operation; sigmoid() is the S-type activation function; tanh() is the hyperbolic tangent activation function; is the initial embedding of the first node u; is the initial embedding of the tail node v; 1.

2. Utilize hinge loss Establish constraints so that directed edges Attribute spatial heterogeneity The edge label y that approaches this edge uv , the specific formula is as follows: in, is the pair-relation subgraph G r The labeled edge set is formed by sampling the positive and negative samples in a balanced manner according to the edge labels; 1.

3. Calculate the l-th layer convolution in the relation subgraph G r Perform directed propagation and update on the label to obtain the category label Y (l),r , the specific formula is as follows: in, is the relationship subgraph G r The i-th order adjacency matrix of , and Q is the maximum order of the adjacency matrix; is the relationship subgraph G r The 1 to Q order adjacency matrix The sum of , and 1≤i≤Q; is a matrix The transpose of For the l-th layer convolution in the relation subgraph G r The structural spatial heterogeneity matrix of all directed edges learned above; For the learned directed edges The structural spatial heterogeneity of ⊙ is Hadamard; is a matrix The corresponding diagonal in-degree matrix; The updated category label vector after the l-th layer label is propagated; Y (0),r =[y v ] is the initial category label vector composed of true category labels; Scalar y v represents the category of the tail node v, 1 represents fraud and 0 represents benign; 1.

4. Using Cross Entropy Loss Establish constraints to make the tail node prediction label based on the directed propagation of labels Approaching its true label y v , the specific formula is as follows: in, is the pair-relation subgraph G r The labeled nodes in the middle are the labeled node sets formed by sampling the positive and negative samples in a balanced manner according to the node labels; y v is the label of the tail node v; is the relationship subgraph G r The predicted label of the tail node v is obtained based on the directed propagation of the label; 1.

5. Apply the lth layer of convolution to attribute space heterogeneity and structural spatial heterogeneity Perform weighted fusion to form dual spatial heterogeneity The specific formula is as follows: in, For directed edges The spatial heterogeneity of attributes; For directed edges The structural spatial heterogeneity of α is Hyperparameters for weighted fusion; β is Hyperparameters for weighted ensemble.

4. The semi-supervised graph fraud detection model based on dual-space heterogeneity learning according to claim 3 is characterized by: In step 1.2, the edge label y uv The allocation rules are: If there is an edge The labels of the head node u and the tail node v are the same, that is, y u =y v , then define its edge label y uv =1, called homogeneous edge; If the labels are different, y u ≠y v , then define y uv =-1, which is called a heterogeneous edge.

5. The semi-supervised graph fraud detection model based on dual-space heterogeneity learning according to claim 1 is characterized by: In step 2, for the relationship subgraph G r Any tail node v∈V r , let N r (v) is the node set formed by all the first nodes corresponding to the tail node. Based on the dual spatial heterogeneity, the lth layer of convolution embeds the tail node and updates it. The overall embedding is obtained by multi-relation aggregation The specific steps are as follows: 2.

1. Calculate in-degree neighbors u∈N based on self-attention mechanism r (v) Importance of the tail node v The specific formula is as follows: in, and is the trainable weight matrix; is the embedding of the first node u obtained by the l-1th layer convolution update; is the embedding of the tail node v obtained by the l-1th layer convolution update; || is the splicing operation; 2.

2. Calculate the attention coefficient of the tail node v to each head node u The specific formula is as follows: Among them, the first node The loop variable of the inner loop; First Node The importance of the tail node v; LeakyReLU() is a leaky linear rectification function; exp{} is the exponential operation; 2.

3. Update the embedding of the tail node v based on multi-head attention aggregation The specific formula is as follows: in, is the weight matrix of the kth attention head in the multi-head attention; is the attention coefficient of the kth attention head of the tail node v to the first node u; K is the total number of heads of multi-head attention; It is the concatenation operation of the output of K attention heads; 2.

4. For the lth layer of convolution, embedding concatenation and mapping are used to form the overall embedding of the tail node v on all relations The specific formula is as follows: in, is the weight matrix; 2.5 Based on overall embedding Use prototype learning to generate the classification probability of the tail node v 6. The semi-supervised graph fraud detection model based on dual-space heterogeneity learning according to claim 5 is characterized by: In step 2.3, for the Lth layer of convolution, weighted average aggregation is used for step 2.3 to obtain the embedding of the tail node v The specific formula is as follows: Among them, L is the last layer of convolution, that is, l=L.

7. The semi-supervised graph fraud detection model based on dual-space heterogeneity learning according to claim 5, characterized in that: In step 2.5, based on the overall embedding Use prototype learning to generate the classification probability of the tail node v The specific steps are as follows: 2.5.

1. The first convolution layer uses the labels and embeddings of the labeled nodes in the graph G to construct the category prototype. The specific formula is as follows: Among them, V F is the set of labeled fraud nodes; V B is the set of labeled benign nodes; is the embedding of the tail node v obtained by the l-th layer convolution; The embedding of the fraud category prototype obtained for the l-th convolution layer; Embedding of the benign category prototype obtained by the l-th convolution layer; 2.5.

2. Perform balanced sampling of positive and negative samples on the labeled nodes in graph G according to the node labels to form a training set And calculate V tr The Euclidean distance between the embedding of the middle tail node v and the embedding of the prototype of its own category is as follows: If v is a fraudulent node, then: If v is a benign node, then: Among them, ||·||2 is the L2 norm; is the Euclidean distance between the embedding of the fraud node v and the embedding of the fraud category prototype; is the Euclidean distance between the embedding of the benign node v and the embedding of the benign category prototype; 2.5.

3. Convolutional layer l uses the Euclidean distance between the embedding of the tail node v and its own category prototype embedding to generate the classification probability of the tail node v And use cross entropy loss Establish constraints on the classification deviation of nodes. The specific formula is as follows: Among them, y v is the label of the tail node v; L(y v )∈{B,F} is the category label indicator of the tail node v.

8. According to the semi-supervised graph fraud detection model based on dual-space heterogeneity learning according to claim 7, in step 2.5.3, the indicator L(y v ) are as follows: If y v = 0, then the tail node v belongs to the benign category, that is, the indicator L(y v ) is equal to B; If y v =1, then the tail node v belongs to the fraud category, that is, the indicator L(y v ) is equal to F.

9. According to the semi-supervised graph fraud detection model based on dual spatial heterogeneity learning in claim 1, in step 3, the overall embedding obtained based on the L-th layer convolution Calculate the probability p that the tail node v is ultimately predicted to be a fraudulent node v , the specific formula is as follows: And use cross entropy loss The node classification effect on the constrained fraud graph is as follows: in, is the overall embedding of the tail node v obtained by the L-th layer convolution.

10. The semi-supervised graph fraud detection model based on dual-space heterogeneity learning according to claim 1, characterized in that: In step 4, the semi-supervised graph fraud detection model based on dual-space heterogeneity learning is trained end-to-end. The specific steps are as follows: 4.

1. Calculate the overall loss of the semi-supervised graph fraud detection model based on dual spatial heterogeneity learning The specific formula is as follows: in, γ1 is the cross entropy loss The weight parameter of γ2 is the cross entropy loss The weight parameter of γ3 is the cross entropy loss The weight parameter of 4.

2. Based on the overall loss value The stochastic gradient descent algorithm is used to calculate the gradient through the back propagation algorithm, and the parameters Perform iterative updates.

Citation Information

Cited By

  • Electronic target behavior intelligent identification method based on multi-domain reconnaissance big data

    CN122112540A

  • An electronic target behavior intelligent identification method based on multi-domain reconnaissance big data

    CN122112540B