A Directed Graph Fraud Detection Method Based on a Jump-Diffusion State-Space Model
Patent Information
- Application Number
- CN202610889517.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-18
- Publication Date
- 2026-09-11
AI Technical Summary
[0006]本发明旨在解决现有方法在有向图欺诈检测中存在的方向信息丢失、序列建模与图结构耦合不足、游走聚合有效性不足以及状态空间模型缺乏结构边界感知能力等技术问题
本发明能够完整保留有向资金流转的因果语义与交易时序,精确感知交易网络中高频拓扑结构突变带来的异常信号,自适应区分正常交易片段与欺诈跳变区域,从而显著提升对复杂钓鱼欺诈和洗钱行为的检测准确性。双向游走建模与缺失流处理有效捕捉交易行为的方向不对称性,增强模型在异构交易网络下的鲁棒性与泛化能力。同时,支持以线性计算复杂度进行长程依赖建模,满足大规模真实交易图谱的高效部署需求。
Smart Images

Figure CN122736746A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of blockchain security and graph neural network technology, specifically to a directed graph fraud detection method based on a jump-diffusion state space model. Background Technology
[0002] Blockchain technology, with its decentralized, immutable, and anonymous characteristics, has been widely applied in finance, the Internet of Things, and other fields. Ethereum, as the largest blockchain platform supporting smart contracts, has seen explosive growth in transaction volume and active users. However, while Ethereum's anonymity and lack of regulation protect user privacy, they also provide a natural shield for cybercrime, leading to frequent phishing scams, money laundering, and Ponzi schemes. The core operation of Ethereum fraud typically manifests as a directed causal transaction chain where funds flow from the victim to the attacker, and then through layers of intermediaries to money laundering addresses. To effectively combat these cybercrimes, researchers often construct Ethereum transaction records as complex transaction network graphs, transforming fraud detection into a task of classifying nodes or detecting anomalies on the graph. An ideal detection system needs the ability to understand the sequence of causal fund flows, the ability to decouple heterogeneous signals at frequencies, the ability to discriminate and aggregate multi-path evidence, and the ability to model the linear complexity of long-range dependencies. Currently, graph representation learning and deep learning-based methods have become the mainstream for Ethereum fraud detection. However, facing the complexity and high heterogeneity of the Ethereum transaction network, existing technologies have revealed the following significant shortcomings in practical applications: (1) Traditional graph neural network methods, such as GCN, GAT, and GraphSAGE, often use GNNs to learn node features by aggregating information in local neighborhoods through message passing mechanisms. However, the 1-hop message passing mechanism of GNNs is limited by the expressive power of the 1-WL test, and essentially performs symmetrical or undirected aggregation of graph topology, which directly loses the directionality and temporality of fund flow. For example, a causal fund chain like "A→B→C" will be compressed into an undirected aggregation representation of "A, B, and C are mutually visible", and the model cannot capture the core fraud temporal causal structure of "where the money comes from, what intermediaries it goes through, and where it goes". At the same time, multi-layer graph neural networks are prone to oversmoothing effects, and such models are essentially spatial low-pass filters. In the real Ethereum transaction graph, fraud nodes generally exhibit heterogeneous characteristics. Normal nodes are usually distributed in dense community structures, and their signal features are mainly low-frequency signals; while abnormal nodes such as phishing and money laundering nodes are manifested as abrupt changes in local topological structures, and their signal features are mainly high-frequency signals. The inherent low-pass filtering characteristic of graph neural networks will eliminate the key discriminative features of heterogeneous nodes as topological abrupt change points, thus leading to a significant decrease in the detection performance of the model in long-range heterogeneous graphs.
[0003] (2) Graph attention network solutions based on the graph Transformer architecture mainly suffer from two technical defects: dilution of semantic features across multiple paths and excessive computational overhead. To address the problems of limited receptive field and oversmoothing effects in traditional graph neural networks, existing research has introduced a global attention mechanism to enable direct information interaction between all nodes. While the global attention mechanism can broaden the model's receptive field, it is highly susceptible to interference from irrelevant contextual information during global information transmission and interaction. When performing global attention calculations, the model cannot distinguish the differences in discriminative abilities among different causal paths, causing the semantic features of key fraud paths, which account for a very small percentage, to be diluted by the contextual information corresponding to a large number of normal communities, making it difficult to retain effective discriminative information. At the same time, the computational complexity of the global attention mechanism is quadratic with the number of nodes. When the model is applied to a real Ethereum transaction graph with tens of thousands of nodes, it will cause a huge consumption of GPU memory resources, making it difficult to deploy and apply in real-world scenarios.
[0004] (3) Detection methods based on state-space models. At present, relevant research applies state-space models, which have linear computational complexity and excellent long sequence modeling capabilities, to financial transactions and blockchain graph data processing in order to solve the aforementioned technical problems. However, such cutting-edge solutions still have shortcomings. First, at the level of local structure perception, although existing graph structure-adapted state-space models have attempted to introduce subgraph encoding strategies to break through the upper bound of subtree expressive power of traditional graph neural networks, their local feature extraction components still rely on the message passing paradigm based on attention mechanisms. Essentially, they are still limited by the homogeneity aggregation assumption and cannot fundamentally avoid the technical limitations brought about by low-pass filtering. Some solutions only achieve adaptive modulation of state parameters through input-dependent selective mechanisms, completely ignoring the impact of abnormal signals formed by topological mutations in the graph frequency domain on the state update process of the state-space model. When money laundering-related abnormal nodes complete behavioral disguise by mixing with a large number of normal transactions, existing models adopt a uniform selective dynamic mechanism for all nodes and cannot rely on the differences in high and low frequency structures of the graph to achieve adaptive identification and response to abnormal states. Secondly, at the global sequence construction level, existing graph structure-adapted state-space models, when flattening graph nodes into sequence inputs, typically only sort them in ascending order based on the sum of the degrees of the subgraph nodes, or directly encode them sequentially along the time steps, without employing long-distance directed random walks to construct effective input sequences that fully preserve the topological semantics of the graph. More critically, such methods fail to strictly distinguish between tracing the source of funds inflows and tracking the outflows of funds, directly causing serious semantic confusion between the funding source patterns of phishing scams and the funding destination patterns of money laundering.
[0005] In summary, existing Ethereum fraud detection technologies have significant shortcomings in preserving the causal semantics of fund flows, perceiving high-frequency topological boundaries, and achieving frequency domain decoupling, making it difficult to meet the detection needs in complex scenarios. Therefore, there is an urgent need for a new graph sequence modeling framework that can comprehensively overcome the above-mentioned deficiencies. Summary of the Invention
[0006] This invention aims to address the technical problems in existing methods for directed graph fraud detection, such as loss of directional information, insufficient coupling between sequence modeling and graph structure, insufficient effectiveness of walk aggregation, and lack of structural boundary awareness in state-space models. By introducing a jump-diffusion stochastic process, the sequence modeling along the walk path is decomposed into continuous diffusion channels and discrete jump channels. Furthermore, a directed gradient-driven topology gating mechanism is used to achieve dynamic perception and adaptive response to changes in the directed graph structure.
[0007] To achieve the above objectives, this invention provides a directed graph fraud detection method based on a jump-diffusion state-space model, comprising the following steps: S1. Decouple and project the original node features of the input directed transaction graph to generate gradient space features and sequence space features respectively. S2. For each target node in the directed transaction graph, sample multiple biased random walk sequences along the inflow and outflow directions to obtain the inflow walk tensor and the outflow walk tensor. S3. Based on the characteristics of the gradient space and the walking tensor, construct a directed gradient tensor, and construct the sequence space input based on the characteristics of the sequence space and the walking tensor. S4. Construct a jump-diffusion state space model, which contains several layers of jump-diffusion modules. Each layer of jump-diffusion module includes a diffusion channel, a jump channel, and a topology gate. The topology gate adaptively adjusts the fusion weights of the diffusion channel output and the jump channel output according to the directed gradient tensor. S5. Calculate the importance weight of each walk using the topology-gated trajectory, and perform weighted aggregation on the multiple walk representations of each target node to obtain the inflow direction aggregation representation and the outflow direction aggregation representation; S6. Fill in the missing flow based on the existence of inflow and outflow, and adaptively fuse the inflow direction aggregation representation and the outflow direction aggregation representation to obtain the final fused representation of the target node; S7. Based on the fraud detection task type, perform node-level or graph-level fraud detection classification based on the final fused representation and the anchor features of gradient space features and sequence space features.
[0008] Preferably, S1 includes: The original feature matrix of the input node is sequentially subjected to linear projection, layer normalization, GELU activation and random deactivation to obtain gradient space features; A single-layer multi-head graph attention aggregation is performed on the undirected adjacency relations of the directed transaction graph, and the sequence space features are obtained by combining residual connections and layer normalization.
[0009] Preferably, S2 includes: For the current node u, the unnormalized transition weight of the candidate next node v is defined as: in, m 1 represents the transaction amount. α The transaction amount bias index controls the degree of wandering's preference for high-value transactions; β It is an anti-centralization penalty index used to suppress wandering from being excessively attracted by central nodes with high height. To prevent extremely small positive numbers from having a value of zero; Represents a node v The reverse degree; After normalizing the transition weights on the candidate set, the transition probability is obtained. Based on the transition probability, K walking sequences of length L are sampled for each target node along the inflow and outflow directions to obtain the inflow walking tensor and the outflow walking tensor and the corresponding time series.
[0010] Preferably, S3 includes: along adjacent walking positions in the walking tensor, extracting corresponding features in the gradient space features, calculating the feature difference between adjacent positions, and obtaining a directed gradient tensor.
[0011] Preferably, in step S4, the processing flow of the m-th layer of the skip-diffusion module is as follows: The working characteristics are obtained by performing layer normalization on the output of the previous layer. The working characteristics are input into the selective state-space model to obtain the diffusion channel output; The working features are passed through a two-layer feedforward network to obtain the skip channel output; The directed gradient tensor is layer normalized and linearly projected, and then the topological gate value is obtained by sigmoid activation. The diffusion channel output and the skip channel output are weighted and fused element-wise based on the topology gating value, and the m-th layer output is obtained through residual connection.
[0012] Preferably, S5 includes: Using the gated trajectory of the last-layer jump-diffusion module, calculate the importance score for each walk, which is the ratio of the sum of the topological gating values on the gated trajectory to the effective position of the walk; Importance scores are converted into normalized weights using a masked Softmax with a learnable temperature parameter, where invalid walks are not included in the normalization. The weighted aggregate representation is calculated based on the normalized weights, and the mean representation of the effective walk is also calculated. The two are then connected by linear projection and residuals and normalized by layer to obtain the aggregate representation in the corresponding direction.
[0013] Preferably, S6 includes: Based on the walk validity mask, determine whether the target node has a valid walk in the inflow and outflow directions, and replace its aggregate representation with a learnable zero-value token for the missing directions; The missing pattern of a node is determined based on the existence combination of inflow and outflow, and the corresponding missing pattern embedding vector is obtained through an embedding lookup table. The inflow and outflow representations after filling are subjected to global average pooling, concatenated and passed through a two-layer bottleneck fully connected network and activated by Sigmoid to obtain channel-level fusion weights; The inflow and outflow representations are weighted and fused according to the channel-level fusion weights. When there is only one direction, the representation of that direction is directly adopted and a learnable missing bias is superimposed. The missing pattern embedding vector is added to the fusion result, and the final fusion representation is obtained after layer normalization.
[0014] Preferably, S7 includes: For node-level classification tasks, the anchor features of the target node in the sequence space features and gradient space features are normalized by layers respectively, and then concatenated with the final fused representation. The node-level prediction probability is obtained by passing through a fully connected network and Softmax normalization. For graph-level classification tasks, all nodes in the graph are used as the target node set. Each node is constructed with a payload vector containing a two-stream difference vector and a missing pattern embedding. The graph-level representation is obtained through a sparse evidence readout mechanism, which includes dense branches and sparse branches. The two branches are adaptively interpolated using a learnable hybrid gate to obtain the graph-level output, and then normalized by Softmax to obtain the graph-level prediction probability.
[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention can fully preserve the causal semantics and transaction timing of directed fund flows, accurately perceive abnormal signals caused by high-frequency topological changes in the transaction network, and adaptively distinguish between normal transaction segments and fraudulent transition regions, thereby significantly improving the detection accuracy of complex phishing fraud and money laundering behaviors. Bidirectional walk modeling and missing flow processing effectively capture the directional asymmetry of transaction behavior, enhancing the model's robustness and generalization ability in heterogeneous transaction networks. Simultaneously, it supports long-range dependency modeling with linear computational complexity, meeting the need for efficient deployment of large-scale real-world transaction graphs. Attached Figure Description
[0016] To more clearly illustrate the technical solution of the present invention, the drawings used in the embodiments are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a schematic diagram of the method flow according to an embodiment of the present invention; Figure 2 This is a diagram of the internal architecture of the jump-diffusion module in the method of this invention. Detailed Implementation
[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0019] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0020] Example This embodiment provides a directed graph fraud detection method based on a jump-diffusion state space model, the method flow of which is as follows: Figure 1 As shown.
[0021] S1. Decouple and project the original node features of the input directed transaction graph to generate gradient space features and sequence space features respectively.
[0022] The original node features of the input directed transaction graph are decoupled and projected to generate gradient space features without graph smoothing for calculating feature differences, and sequence space features after local structure aggregation for providing rich input.
[0023] In this embodiment, the input directed transaction graph is denoted as... ,in For a set of nodes, Let be a set of directed edges, with the total number of nodes denoted as . The original feature matrix of the nodes is denoted as... , No. i The original feature vectors of the nodes are denoted as follows: .side The attributes include transaction amount and timestamp The hidden dimension is denoted as... DSample each target node K 3 walking sequences, each sequence containing L The number of layers in the two-stream jump-diffusion state-space model is denoted as M.
[0024] The core objective of this step is to decouple two complementary representation spaces from the original node features: a gradient space to preserve the unsmoothed differences between nodes, and a sequence space to incorporate the graph structure context. These two spaces serve the subsequent directed gradient calculation and state space sequence modeling, respectively.
[0025] For the original feature matrix of the input node The gradient space feature matrix is obtained by sequentially performing linear projection, layer normalization, GELU activation, and random deactivation operations. : in, , These are learnable parameters. The resulting gradient space features. Its purpose is to preserve the differences between nodes at the original feature level, providing a feature representation that is not affected by graph structure aggregation for subsequent directed gradient calculation.
[0026] Based on this, a single-layer multi-head graph attention aggregation is performed on the undirected adjacency relations of the graph, and combined with residual connectivity and layer normalization, the sequence space feature matrix is obtained. : The gradient space is used for subsequent calculation of directed gradients, and its unsmoothed features can keenly capture feature transitions between adjacent nodes; the sequence space is used for state space modeling of subsequent walk sequences, and its features integrated into the graph structure context help the sequence model understand local topological semantics.
[0027] S2. For each target node in the directed transaction graph, sample multiple biased random walk sequences along the inflow and outflow directions to obtain the inflow walk tensor and the outflow walk tensor.
[0028] This step samples multiple biased random walk sequences for each target node along both the inflow and outflow directions on the directed transaction graph to capture the node's topological neighborhood information in both the inflow and outflow dimensions. The walk process introduces transaction amount bias and decentralization penalty, making the sampling path tend to pass through channels with high transaction amounts, while avoiding excessive attraction to highly centralized nodes.
[0029] Let the set of target nodes to be detected in the current batch be . For each target node K walking sequences of length L were sampled in both the inflow and outflow directions.
[0030] For inbound flows, when the current node is u, the candidate next node is all predecessor nodes that have initiated transactions to u in the graph. For outbound flows, the candidate next node is all successor nodes that have received transactions from u. Regardless of the direction, candidate nodes... v The unnormalized transition weights are all defined as: in, m 1 represents the transaction amount. α The transaction amount bias index controls the degree of wandering's preference for high-value transactions; β It is an anti-centralization penalty index used to suppress wandering from being excessively attracted by central nodes with high height. To prevent extremely small positive numbers from having a value of zero; Represents a node v The reverse degree is calculated as follows: the degree is taken when the flow enters and the in-degree is taken when the flow exits. Normalizing these weights over the candidate set yields the transition probability.
[0031] The inflow wander tensor is obtained through the above sampling process. Outflow wandering tensor and the corresponding time series These walk sequences will serve as input to the subsequent state-space model.
[0032] S3. Based on the gradient space characteristics and the walking tensor, construct a directed gradient tensor, and construct the sequence space input based on the sequence space characteristics and the walking tensor.
[0033] This step constructs the sequence space input and directed gradient input required for the state space model based on the dual-space features obtained in S1 and the walk path obtained in S2.
[0034] For any direction Construct a feature sequence along the walk tensor obtained in step two. For the j-th walk of the b-th target node, the first... k Nodes reached in steps Using the node as an index, its features are extracted from the gradient space matrix E to calculate the directed gradient. The directed gradient tensor is defined as the feature difference between adjacent walk positions in the gradient space: in, This represents the directed gradient tensor. This represents the gradient space characteristics along the walk path.
[0035] The intention behind the directed gradient is that the feature differences between adjacent nodes along the walk path in the gradient space reflect local topological jumps in the direction of the transaction flow. In a normal transaction network, the features of adjacent nodes are usually relatively smooth, resulting in small directed gradient values; however, on fraudulent paths, features often exhibit abrupt changes, leading to larger directed gradient values. This signal will serve as the basis for subsequent topology gating.
[0036] S4. Construct a jump-diffusion state space model, which contains several layers of jump-diffusion modules. Each layer of jump-diffusion module includes a diffusion channel, a jump channel, and a topology gate. The topology gate adaptively adjusts the fusion weights of the diffusion channel output and the jump channel output according to the directed gradient tensor.
[0037] For both the inflow and outflow directions, independent jump-diffusion state-space models are constructed for sequence modeling. The model for each direction consists of... M It is composed of stacked skip-diffusion modules. Each skip-diffusion module contains three core components: a diffusion channel, a skip channel, and a topology gate. Its specific architecture is as follows: Figure 2 As shown.
[0038] Order No. m The input of the layer is The input of layer 0 is the sequence spatial features. The processing flow for each layer is as follows: (1) Prenormalization First, perform layer normalization on the output of the previous layer to obtain the working characteristics of the current layer: (2) Diffusion channels The normalized features are input into the selective state-space model to obtain the diffusion channel output. The selective state-space model models long-range dependencies in a walk sequence through a data-dependent state transition mechanism. When the graph data contains timestamp information, time modulation is further applied to the output of the diffusion channel. The absolute time difference between adjacent walk positions is calculated, and the time difference is mapped into a high-dimensional feature vector through a Fourier time encoder. Then, the channel-wise time modulation factor is obtained through linear projection and Sigmoid activation. And perform element-wise modulation on the diffusion channel output: The introduction of time modulation enables the model to perceive the time interval between transactions, assigning different modeling weights to consecutive transactions occurring frequently within a short period compared to normal transactions with longer intervals. This modulation step is skipped when the data does not contain timestamps.
[0039] (3) Topology gating Normalize and linearly project the directed gradient obtained from S3, and then activate it with a Sigmoid function to obtain the channel-wise topological gating value: in, Indicates the first m The weight matrix for linear projection of topological gating along the layer dir direction. Indicates the first m Layer normalization operation applied to the directed gradient along the layer's dir direction. Indicates the first m The bias term of the topologically gated linear projection along the layer dir direction.
[0040] (4) Jumping passage Based on the normalized input of the current layer, a skip channel output is constructed through a two-layer feedforward network: in, Indicates the first m In the jump channels along the dir direction, the weight matrix of the first layer feedforward network, Indicates the first m The pre-normalized input features along the dir direction of the layer. Indicates the first m In the skip channels along the layer dir direction, the bias term of the first layer feedforward network, Indicates the first m The bias term of the second-layer feedforward network in the jump channel along the layer dir direction.
[0041] Skip channels do not rely on sequence history; they perform nonlinear transformations based solely on the input features at the current position, enabling them to rapidly adjust their representation when topological abrupt changes are detected.
[0042] (5) Skip-diffusion fusion and residual connection. Topology gating is used to perform element-wise weighted fusion of the diffusion channel and the skip channel, and the output of the m-th layer is obtained through residual connection: in, Indicates the first m The topology gating value of the layer.
[0043] When the gate value is close to 0, the output is mainly dominated by the diffusion channel; when the gate value is close to 1, the output is mainly dominated by the jump channel. This adaptive switching mechanism allows the model to flexibly handle normal and abnormal transaction segments on the same walk path. After completing the M-layer stacking, the output of the last position of each walk sequence is taken as the representation of that walk, and then rearranged to obtain... .
[0044] S5. Calculate the importance weight of each walk using the topology-gated trajectory, and perform weighted aggregation on the multiple walk representations of each target node to obtain the inflow direction aggregation representation and the outflow direction aggregation representation.
[0045] This step aggregates the K walk representations for each target node into a single directional representation. This invention utilizes topology-gated trajectories as a measure of the importance of each walk, assigning higher aggregation weights to walks that traverse more topological boundaries. A walk validity mask is defined: a walk is marked as valid if it has made a real movement at at least one location, otherwise it is invalid. The gated trajectory utilizes the last-layer jump-diffusion module. Calculate the importance score for each walk: This importance score measures the average intensity of topology-gated activations along a walk path. More frequent and larger gating activations indicate that the walk traverses more regions of topological abrupt changes and carries richer fraud detection information.
[0046] This importance score is passed through a learnable temperature parameter. The mask Softmax is converted to normalized weights Invalid walks are not included in the normalization.
[0047] Further calculations of the weighted aggregation representation and the mean representation of the effective walks yield the final aggregation representation for this direction, defined as: in, This is the weighted aggregation result based on importance. For the mean pooling result of the effective walk, the two are connected by linear projection and residuals and then output by layer normalization. This aggregation method takes into account both importance weighting and mean pooling, so that the final directional representation can capture key fraud clues without being biased by a single abnormal walk.
[0048] S6. Fill in the missing flow based on the existence of inflow and outflow, and adaptively fuse the inflow direction aggregation representation and the outflow direction aggregation representation to obtain the final fused representation of the target node.
[0049] In real-world transaction networks, some nodes may only have inflow transactions without outflow transactions, or only outflow transactions without inflow transactions, or even isolated nodes with no direct transaction records. These missing patterns are themselves important fraud detection signals. This step processes the missing flows and merges the aggregated representations of the inflow and outflow directions into a unified node representation.
[0050] (1) Missing flow filling Based on the walk validity mask obtained in S5, it is determined whether each target node has a valid walk in both the inflow and outflow directions. For any missing direction, the aggregate representation of that direction is replaced with a learnable zero-value token. Furthermore, based on the existence combination of inflow and outflow, nodes are divided into four missing patterns (both exist, only inflow, only outflow, and both are missing), and the corresponding missing pattern embedding vector is obtained through an embedding lookup table containing four elements. .
[0051] (2) Channel-level adaptive fusion Inflow representation after filling and outflow representation The channel-wise fusion weights are calculated. First, global average pooling is performed on the representations in both directions to obtain channel-level statistics. Then, the two representations are concatenated and input into a two-layer bottleneck fully connected network, and the corresponding fusion weights are obtained by passing the sigmoid activation function. Then, the fusion result is calculated: When only one direction exists, the representation of that direction is directly used and the learnable missing bias is superimposed.
[0052] (3) Final fusion representation The missing pattern is embedded into the fusion result, and then normalized by layers to obtain the final fused representation of the target node: Simultaneously calculate the dual-stream difference vector This vector captures the asymmetry in the behavior of a node in the inflow and outflow directions. The greater the difference between inflow and outflow, the higher the degree of abnormality in the transaction behavior exhibited by the node.
[0053] S7. Based on the fraud detection task type, perform node-level or graph-level fraud detection classification based on the final fused representation and the anchor features of gradient space features and sequence space features.
[0054] Depending on the specific fraud detection task type, this invention supports node-level classification and graph-level classification.
[0055] For node-level fraud detection, the anchor point features of the target node in the gradient space and sequence space are extracted, normalized at each layer, and then concatenated with the fused representation obtained in step six to construct the node-level classification input: in, and These are the normalized anchor features of the target node in the sequence space and gradient space, respectively. This concatenation strategy allows the classifier to simultaneously obtain three complementary information: fused representation. The neighborhood structure information and missing patterns extracted by the target node during bidirectional walks are encoded. Provides local context after graph attention aggregation. The original, unsmoothed features of the nodes are preserved. Subsequently, node-level outputs are obtained through a two-layer fully connected network and normalized using Softmax to obtain the predicted probabilities of the target nodes in each category.
[0056] For graph-level fraud detection tasks, all nodes in the graph are used as the target node set. Each node constructs a payload vector containing five parts: in, For two-stream difference vectors, For missing pattern embedding. Compared to node-level classification, graph-level classification introduces these two additional features to fully utilize global statistical information such as the asymmetry of flow direction and missing patterns among nodes during graph-level aggregation. Graph-level classification employs a sparse evidence readout mechanism, including two paths: dense branch and sparse branch. For the dense branch, the load vector of each node is projected onto the evidence space, and then mean pooling, max pooling, and variance pooling are performed on all nodes in the graph to obtain graph-level statistics.
[0057] The three statistical measures are concatenated and then passed through a fully connected network to obtain a dense branch representation. This branch characterizes the overall distribution features of the transaction graph through global statistical features. For sparse divide and conquer, a graph-level context vector is constructed and broadcast to each node. The evidence score for each node is calculated by combining the node's own evidence. The evidence score is then normalized using a temperature-scaled Softmax method within the graph to obtain attention weights, which are then used to calculate the attention-weighted node aggregation representation. This representation is concatenated with relevant statistics and passed through a fully connected network to obtain the sparse branch representation. The core advantage of this branch lies in its ability to allow a small number of high-evidence nodes to dominate the graph-level representation, even if they represent a small percentage of the entire graph. A learnable hybrid gate is used to adaptively interpolate both branches. The hybrid gate takes dense representations, sparse representations, and node evidence probability statistics as inputs and outputs a scalar gating value. The final graph-level output is: in, This represents the classifier function for dense branches, responsible for mapping the feature representations of the dense branches to the corresponding classification prediction results. Indicates a scalar gate value. The classifier function representing sparse branches.
[0058] Applying Softmax to the results yields the predicted probability of graph-level fraud detection.
[0059] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made to the technical solutions of the present invention by those skilled in the art without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.
Claims
1. A directed graph fraud detection method based on a jump-diffusion state-space model, characterized in that, Includes the following steps: S1. Decouple and project the original node features of the input directed transaction graph to generate gradient space features and sequence space features respectively. S2. For each target node in the directed transaction graph, sample multiple biased random walk sequences along the inflow and outflow directions to obtain the inflow walk tensor and the outflow walk tensor. S3. Based on the characteristics of the gradient space and the walking tensor, construct a directed gradient tensor, and construct the sequence space input based on the characteristics of the sequence space and the walking tensor. S4. Construct a jump-diffusion state space model, which contains several layers of jump-diffusion modules. Each layer of jump-diffusion module includes a diffusion channel, a jump channel, and a topology gate. The topology gate adaptively adjusts the fusion weights of the diffusion channel output and the jump channel output according to the directed gradient tensor. S5. Calculate the importance weight of each walk using the topology-gated trajectory, and perform weighted aggregation on the multiple walk representations of each target node to obtain the inflow direction aggregation representation and the outflow direction aggregation representation; S6. Fill in the missing flow based on the existence of inflow and outflow, and adaptively fuse the inflow direction aggregation representation and the outflow direction aggregation representation to obtain the final fused representation of the target node; S7. Based on the fraud detection task type, perform node-level or graph-level fraud detection classification based on the final fused representation and the anchor features of gradient space features and sequence space features.
2. The directed graph fraud detection method based on the jump-diffusion state-space model according to claim 1, characterized in that, S1 includes: The original feature matrix of the input node is sequentially subjected to linear projection, layer normalization, GELU activation and random deactivation to obtain gradient space features; A single-layer multi-head graph attention aggregation is performed on the undirected adjacency relations of the directed transaction graph, and the sequence space features are obtained by combining residual connections and layer normalization.
3. The directed graph fraud detection method based on the jump-diffusion state-space model according to claim 1, characterized in that, S2 includes: For the current node u, the unnormalized transition weight of the candidate next node v is defined as: in, m 1 represents the transaction amount. α The transaction amount bias index controls the degree of wandering's preference for high-value transactions; β It is an anti-centralization penalty index used to suppress wandering from being excessively attracted by central nodes with high height. To prevent extremely small positive numbers from having a value of zero; Represents a node v The degree of reversal; After normalizing the transition weights on the candidate set, the transition probability is obtained. Based on the transition probability, K walking sequences of length L are sampled for each target node along the inflow and outflow directions to obtain the inflow walking tensor and the outflow walking tensor and the corresponding time series.
4. The directed graph fraud detection method based on the jump-diffusion state-space model according to claim 1, characterized in that, S3 includes: along adjacent walking positions in the walking tensor, extracting corresponding features in the gradient space features, calculating the feature difference between adjacent positions, and obtaining a directed gradient tensor.
5. The directed graph fraud detection method based on the jump-diffusion state-space model according to claim 1, characterized in that, In S4, the processing flow of the m-th layer of the jump-diffusion module is as follows: The working characteristics are obtained by performing layer normalization on the output of the previous layer. The working characteristics are input into the selective state-space model to obtain the diffusion channel output; The working features are passed through a two-layer feedforward network to obtain the skip channel output; The directed gradient tensor is layer normalized and linearly projected, and then the topological gate value is obtained by sigmoid activation. The diffusion channel output and the skip channel output are weighted and fused element-wise based on the topology gating value, and the m-th layer output is obtained through residual connection.
6. The directed graph fraud detection method based on the jump-diffusion state-space model according to claim 5, characterized in that, S5 includes: Using the gated trajectory of the last-layer jump-diffusion module, calculate the importance score for each walk, which is the ratio of the sum of the topological gating values on the gated trajectory to the effective position of the walk; Importance scores are converted into normalized weights using a masked Softmax with a learnable temperature parameter, where invalid walks are not included in the normalization. The weighted aggregate representation is calculated based on the normalized weights, and the mean representation of the effective walk is also calculated. The two are then connected by linear projection and residuals and normalized by layer to obtain the aggregate representation in the corresponding direction.
7. The directed graph fraud detection method based on the jump-diffusion state-space model according to claim 6, characterized in that, S6 includes: Based on the walk validity mask, determine whether the target node has a valid walk in the inflow and outflow directions, and replace its aggregated representation with a learnable zero-value token for the missing directions; The missing pattern of a node is determined based on the existence combination of inflow and outflow, and the corresponding missing pattern embedding vector is obtained through an embedding lookup table. The inflow and outflow representations after filling are subjected to global average pooling, concatenated and passed through a two-layer bottleneck fully connected network and activated by Sigmoid to obtain channel-level fusion weights; The inflow and outflow representations are weighted and fused according to the channel-level fusion weights. When there is only one direction, the representation of that direction is directly adopted and a learnable missing bias is superimposed. The missing pattern embedding vector is added to the fusion result, and the final fusion representation is obtained after layer normalization.
8. The directed graph fraud detection method based on the jump-diffusion state-space model according to claim 7, characterized in that, The S7 includes: For node-level classification tasks, the anchor features of the target node in the sequence space features and gradient space features are normalized by layers respectively, and then concatenated with the final fused representation. The node-level prediction probability is obtained by passing through a fully connected network and Softmax normalization. For graph-level classification tasks, all nodes in the graph are used as the target node set. Each node is constructed with a payload vector containing a two-stream difference vector and a missing pattern embedding. The graph-level representation is obtained through a sparse evidence readout mechanism, which includes dense branches and sparse branches. The two branches are adaptively interpolated using a learnable hybrid gate to obtain the graph-level output, and then normalized by Softmax to obtain the graph-level prediction probability.