A dynamic abnormal subgraph discovery method in a time-series financial network

By modeling subgraph structure and financial distribution, using approximate linear time complexity estimation and reinforcement learning optimization, we solve the problem of misidentification of static abnormal subgraph detection in financial networks and achieve accurate detection of dynamic abnormal subgraphs.

CN119312254BActive Publication Date: 2025-10-24NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411486661.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-23
Publication Date
2025-10-24
Estimated Expiration
2044-10-23

AI Technical Summary

Technical Problem

Existing technologies are unable to effectively detect the changes and trends of user transaction behaviors over time in financial networks, resulting in static anomaly subgraph detection misidentifying legitimate transactions as anomalies.

Method used

By modeling subgraph structure and financial distribution, using approximate linear time complexity to estimate dynamic financial networks, performing chi-square calculation and boundary processing, combined with reinforcement learning to optimize abnormal subgraphs and identify dynamic abnormal subgraphs.

Benefits of technology

It realizes the identification of dynamic abnormal subgraphs in financial networks, detects the changes and trends of user transaction behaviors over time, and improves the accuracy of anomaly detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119312254B_ABST
    Figure CN119312254B_ABST
Patent Text Reader

Abstract

The application discloses a dynamic abnormal subgraph discovery method in a time sequence financial network. The method comprises the following steps: obtaining a dynamic financial network, wherein the dynamic financial network comprises nodes, edges between the nodes, and edges between the nodes are used to represent transaction information between at least one of the nodes; estimating the dynamic financial network through an approximate linear time complexity to obtain dense subgraphs of different time slices, performing chi-square calculation on each edge of the dense subgraph of each time slice to obtain the chi-square of each edge of each time slice, and reserving nodes corresponding to edges with chi-square greater than or equal to the total chi-square of each time slice in the dense subgraph of each time slice. The application solves the technical problem that the prior art can only identify static abnormal subgraphs in a financial network and cannot detect changes and trends of user transaction behaviors over time.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of abnormal subgraph detection, and particularly relates to a dynamic abnormal subgraph discovery method in a time series financial network. BACKGROUND

[0002] Extracting abnormal subgraphs from financial networks is an important problem, and time series abnormal subgraph discovery can capture abnormal situations over time, enabling timely identification of illegal transactions between users. This dynamic monitoring of suspicious activities is crucial for maintaining the security of financial networks. As time anomaly detection is crucial, more and more researchers are focusing on time fraud detection in networks. They propose to detect anomalies based on structural and temporal information that is different from normal patterns in dynamic graphs.

[0003] However, since these methods are designed for general graphs, they fail to take into account the specific transaction details between users in financial networks. Specifically, even if certain nodes frequently transfer money to other nodes at a particular time, resulting in structural and temporal anomalies, these transactions may still be normal. For example, a university account may often disburse stipends or allowances to different student or faculty accounts. If transaction information is ignored, these transactions may be incorrectly identified as abnormal subgraphs, although they are completely legitimate in reality.

[0004] To address the above problems, while taking into account transaction information, the current best method Antibenford suggests using Benford's law to evaluate illegal transactions. Benford's law states that the proportion of numbers starting with a digit follows a monotonic decreasing function in natural data sets (such as tax records, stock quotes, etc.), i.e. . Antibenford [1] also verifies the existence of Benford's law in financial transactions by visualizing the distribution of the first digit in actual transaction data sets.

[0005] However, the latest work, as it is designed based on non-time series graphs, can only identify static abnormal subgraphs in financial networks, which means it cannot detect changes and trends in user transaction behavior over time. In this work, inspired by the static Antibenford method, a time series Antibenford subgraph is defined theoretically. This enables the discovery of time series abnormal subgraphs that deviate from Benford's law in dynamic financial networks with solid theoretical support. SUMMARY

[0006] ​The embodiment of the application provides a dynamic abnormal subgraph discovery method in a time sequence financial network, so as to at least solve the technical problem that the prior art can only identify static abnormal subgraphs in a financial network and cannot detect changes and trends of user transaction behaviors over time.

[0007] According to an aspect of the embodiment of the application, a dynamic abnormal subgraph discovery method in a time sequence financial network is provided, which can include: obtaining a dynamic financial network, wherein the dynamic financial network includes nodes, edges between the nodes, and edges between the nodes are used to represent transaction information of at least one of the nodes; estimating the dynamic financial network by using an approximate linear time complexity to obtain dense subgraphs of different time slices, performing chi-square calculation on each edge of the dense subgraphs of each time slice to obtain chi-square of each edge of each time slice, retaining nodes corresponding to edges of each time slice whose chi-square is greater than or equal to total chi-square of each time slice, and deleting nodes corresponding to edges of each time slice whose chi-square is less than total chi-square of each time slice to obtain top-p abnormal subgraphs of different time slices; encoding the top-p abnormal subgraphs of different time slices to obtain abnormal candidate subgraphs of different time slices, wherein the abnormal candidate subgraphs of different time slices include boundaries; excluding or extending nodes in the abnormal candidate subgraphs of different time slices to obtain abnormal degrees of the abnormal candidate subgraphs of different time slices; maximizing gains of the abnormal degrees of the abnormal candidate subgraphs of different time slices to obtain a maximum reward; and optimizing the abnormal candidate subgraphs of different time slices by using the maximum reward to obtain optimized abnormal candidate subgraphs of different time slices.

[0008] Optionally, the chi-square calculation on each edge of the dense subgraphs of each time slice is performed to obtain an expression of chi-square of the dense subgraphs of each time slice as follows:

[0009]

[0010] wherein, chi-square of the dense subgraphs of each time slice, is the number of time sequence edges of each time slice, is the number of time sequence edges of each time slice, is the number of time sequence edges of each time slice, is the number of time sequence edges of each time slice, .

[0011] Optionally, the process of encoding the top-p abnormal subgraphs of different time slices is: extracting the nodes of the top-p abnormal subgraphs of different time slices by different node numbers to obtain positive top-p abnormal subgraph pairs and negative sample top-p abnormal subgraph pairs in the top-p abnormal subgraphs of different time slices; calculating the similarity between the positive top-p abnormal subgraph pairs and the similarity between the negative top-p abnormal subgraph pairs; and optimizing the target loss function based on the similarity between the positive top-p abnormal subgraph pairs and the similarity between the negative top-p abnormal subgraph pairs.

[0012] Alternatively, the expression for the similarity between positive top-p abnormal subgraph pairs is:

[0013]

[0014] in, Indicates that in the snapshot The top-p abnormal subgraph at any moment Middle The encoding vector of nodes. Similarly, Indicates that in the snapshot The top-p abnormal subgraph at any moment midpoint The encoding vector of the neighborhood of is through the snapshot The top-p abnormal subgraph at any moment midpoint The encoding vectors of all neighbors of Calculated by the function. It is in the snapshot The top-p abnormal subgraph at any moment The number of nodes. And Represents the number of snapshots sampled for each time positive top-p abnormal subgraph, Represents two positive top-p abnormal subgraphs and The financial distribution between Divergence. Positive top-p outlier subgraphs with similar financial distributions lead to lower Score.

[0015] Optionally, the expression for the similarity between negative top-p abnormal subgraph pairs is:

[0016]

[0017] in, Indicates that in the snapshot Time negative top-p abnormal subgraph Nodes randomly sampled fromThe encoding vector of the neighborhood of Since And The positive top-p abnormal subgraph and the negative top-p abnormal subgraph are completely different and have no overlap, so the same node .

[0018] Optionally, based on the similarity between the positive top-p abnormal subgraph pair and the similarity between the negative top-p abnormal subgraph pair, the expression of the target loss function is optimized as:

[0019]

[0020] Wherein, Indicates the temperature parameter.

[0021] Optionally, the expression of the abnormal degree of the abnormal candidate subgraph of different time slices is:

[0022]

[0023] Wherein, The abnormal degree of the abnormal candidate subgraph At step t, The time density of the abnormal candidate subgraph At step t, The abnormal candidate subgraph At step t, The abnormal candidate subgraph At step t, , Indicates the abnormal candidate subgraph in the case of time s, The statistical quantity of the abnormal candidate subgraph At step t,

[0024] Advantages of the present application:

[0025] The present application provides a dynamic abnormal subgraph discovery method in a time series financial network, which encodes the abnormal subgraph by modeling the hidden features in the subgraph structure and financial distribution. Then, through an effective pruning strategy and matching with the trained abnormal, the rough prototype is detected. In addition, a reinforcement learning-based method is used to refine the abnormal subgraph, wherein the search space of the subgraph boundary is limited in range based on the transaction frequency, solving the technical problem that the prior art can only identify static abnormal subgraphs in a financial network and cannot detect changes and trends of user transaction behavior over time. The technical effect of identifying dynamic abnormal subgraphs in a financial network can be achieved, which can be used to detect changes and trends of user transaction behavior over time. BRIEF DESCRIPTION OF DRAWINGS ​

[0026] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0027] Figure 1 is a flowchart of a method for discovering dynamic abnormal subgraphs in a time-series financial network according to an embodiment of the present invention;

[0028] Figure 2 is a schematic diagram of a method for discovering dynamic abnormal subgraphs in a time-series financial network according to an embodiment of the present invention;

[0029] Figure 3 is a schematic diagram of determining abnormal candidate subgraphs of different time slices according to an embodiment of the present invention;

[0030] Figure 4 is a schematic diagram of determining an optimized timing anomaly subgraph according to an embodiment of the present invention;

[0031] Figure 5 is a schematic diagram of collecting positive and negative sample pairs according to an embodiment of the present invention;

[0032] Figure 6 According to an embodiment of the present invention Schematic diagram of the results;

[0033] Figure 7 is a schematic diagram illustrating the relationship between detection time and training time and graph node size according to different methods of an embodiment of the present invention;

[0034] Figure 8 Schematic diagram of time anomaly subgraphs detected by different methods according to an embodiment of the present invention. DETAILED DESCRIPTION

[0035] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0036] It should be noted that the terms "first", "second", and the like in the description and claims of the application and the above drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a particular order or sequence. It should be understood that the terms used in this way can be interchanged under appropriate circumstances, so that the embodiments of the application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0037] Embodiment 1

[0038] According to an embodiment of the application, a dynamic abnormal subgraph discovery method in a time series financial network is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system comprising at least one set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in an order different from that shown herein.

[0039] Figure 1 is a flowchart of a dynamic abnormal subgraph discovery method in a time series financial network according to an embodiment of the application, as shown in Figure 1 The method can include the following steps:

[0040] Step S101, acquiring a dynamic financial network, wherein the dynamic financial network includes nodes, edges between nodes, and edges between nodes, wherein the edges between nodes are used to represent transaction information between at least one of the nodes.

[0041] In the technical solution provided in the above step S101 of the application, Figure 2 is a schematic diagram of a dynamic abnormal subgraph discovery method in a time series financial network according to an embodiment of the application, Figure 2 The leftmost graph in the figure is a dynamic financial network, which includes nodes, edges between nodes, and edges between nodes, wherein the edges between nodes are used to represent transaction information between at least one of the nodes, and wherein the nodes can be users.

[0042] In step S102, the dynamic financial network is estimated by approximate linear time complexity to obtain dense subgraphs of different time slices, chi-square calculation is performed on each edge of the dense subgraph of each time slice to obtain the chi-square of each edge of each time slice, nodes corresponding to edges with chi-square greater than or equal to the total chi-square of each time slice in the dense subgraph of each time slice are retained, and nodes corresponding to edges with chi-square less than the total chi-square of each time slice in the dense subgraph of each time slice are deleted, to obtain top-p abnormal subgraphs of different time slices.

[0043] In the technical solution provided in step S102 of the present application, Figure 3 is a schematic diagram of determining abnormal candidate subgraphs of different time slices according to an embodiment of the present application, and Figure 3 the leftmost dynamic financial network in FIG. Figure 3 the dense subgraph of the second time slice from the left in FIG. 1, that is, the graph behind the temporalDS arrow, wherein the temporal dense subgraph (temporal dense subgraph, referred to as temporalDS) chi-square calculation is performed on each edge of the dense subgraph of each time slice to obtain the chi-square of each edge of each time slice, for example, when the nodes 1, 2 and 3 of the dense subgraph of each time slice are connected by the first edge between node 1 and node 2, and the second edge between node 1 and node 3, when the chi-square of the first edge is greater than or equal to the total chi-square of the dense subgraph of the time slice where it is located, node 1 and node 2 are retained, when the chi-square of the second edge is less than the total chi-square of the dense subgraph of the time slice where it is located, node 3 is deleted, and so on. The dense subgraph corresponding to the nodes that will be retained is determined as the top-p abnormal subgraph of different time slices, and the top-p abnormal subgraph of different time slices (temporalDS) Figure 3 in FIG. 1.

[0044] In step S103, the top-p abnormal subgraphs of different time slices are encoded to obtain abnormal candidate subgraphs of different time slices, wherein the abnormal candidate subgraphs of different time slices contain boundaries.

[0045] In the technical solution provided in step S103 of the present application, the top-p abnormal subgraphs of different time slices are processed by an encoder to obtain abnormal candidate subgraphs of different time slices, such as Figure 3 the rightmost graph in FIG.

[0046] In step S104, the nodes in the abnormal candidate subgraphs of different time slices are excluded or expanded to obtain the abnormal degree of the abnormal candidate subgraphs of different time slices.

[0047] In the technical solution provided in step S104 of the present application, Figure 4 is a schematic diagram of determining the optimized timing anomaly subgraph according to an embodiment of the present application, as Figure 4 shown, for a candidate subgraph in the timing candidate detection module , its connected nodes are obtained: node 1 at snapshot , 5 and 6 at snapshot , and 9 at snapshot . Among them, since compared with node 6 at , node 1 at , 5 at , and 9 at have the top three transaction frequencies, they are selected as the boundary . In general, the boundary should include those nodes connected with the nodes in and having the highest total transaction frequency, which is defined as follows:

[0048]

[0049] Among them, represents the transaction frequency between node and node at snapshot , and represents that the number of nodes in the boundary cannot exceed , is all nodes of the anomaly candidate subgraph at different time slices, represents the anomaly candidate subgraph at snapshot .

[0050] Two actions, i.e. exclusion and expansion, are defined. The action space of exclusion removes the nodes in the anomaly candidate subgraph . While the action space of expansion contains the nodes in the boundary . A multi-layer perception (MLP) is used to learn the features of the nodes in the action space, and the features of the nodes are processed to select the optimal node with the highest probability. In addition, after a certain action is taken, and should become larger, represents the statistics after a certain action is taken, represents the chi-square statistics after a certain action is taken. In order to guide the model to develop in the most abnormal direction, the gain of the abnormality degree caused by a certain action is used as a reward.

[0051] ​Step S105, the gain of the abnormal degree of the abnormal candidate subgraph of different time slices is maximized to obtain the maximum reward.

[0052] In the technical solution provided by the above step S105 of the application, the abnormal candidate subgraph of one time slice should satisfy Therefore, the problem is optimized by policy gradient to maximize the gain of the abnormal degree after taking a certain action, that is, the reward .

[0053] Step S106, the abnormal candidate subgraphs of different time slices are optimized by maximizing the reward, and the optimized abnormal candidate subgraphs of different time slices are obtained.

[0054] In the technical solution provided by the above step S106 of the application, the abnormal candidate subgraphs of different time slices are optimized by maximizing the reward to obtain the optimized abnormal candidate subgraphs of different time slices.

[0055] The above method of the embodiment will be further introduced below.

[0056] As an optional embodiment, step S102, chi-square calculation is performed on each edge of the dense subgraph of each time slice to obtain the expression of the chi-square of the dense subgraph of each time slice as follows:

[0057]

[0058] wherein, is the chi-square of the dense subgraph of each time slice, is the first digit of the transaction amount in the dense subgraph of each time slice is the number of time sequence edges, represents the expected number of edges based on the Benford law, wherein the first digit of the transaction amount is equal to .

[0059] As an optional embodiment, step S103, the process of encoding the top-p abnormal subgraph of different time slices is as follows: the nodes of the top-p abnormal subgraph of different time slices are extracted by different node numbers to obtain positive top-p abnormal subgraph pairs and negative sample top-p abnormal subgraph pairs in the top-p abnormal subgraph of different time slices; the similarity between the positive top-p abnormal subgraph pairs and the similarity between the negative top-p abnormal subgraph pairs are calculated; and the target loss function is optimized based on the similarity between the positive top-p abnormal subgraph pairs and the similarity between the negative top-p abnormal subgraph pairs.

[0060] In this embodiment, Figure 5 ​is a schematic diagram of collecting positive and negative sample pairs according to an embodiment of the present invention, such as Figure 5 As shown, the nodes of the top-p abnormal subgraphs of different time slices are extracted by different node numbers to obtain positive top-p abnormal subgraph pairs and negative sample top-p abnormal subgraph pairs in the top-p abnormal subgraphs of different time slices; the similarity between the positive top-p abnormal subgraph pairs and the similarity between the negative top-p abnormal subgraph pairs are calculated; based on the similarity between the positive top-p abnormal subgraph pairs and the similarity between the negative top-p abnormal subgraph pairs, the target loss function is optimized.

[0061] As an optional embodiment, the similarity between the top-p abnormal subgraph pairs is expressed as follows:

[0062]

[0063] in, Indicates that in the snapshot The top-p abnormal subgraph at any moment Middle The encoding vector of nodes. Similarly, Indicates that in the snapshot The top-p abnormal subgraph at any moment midpoint The encoding vector of the neighborhood of is through the snapshot The top-p abnormal subgraph at any moment midpoint The encoding vectors of all neighbors of Calculated by the function. It is in the snapshot The top-p abnormal subgraph at any moment The number of nodes. And Represents the number of snapshots sampled for each time positive top-p anomaly subgraph, Represents two positive top-p abnormal subgraphs and The financial distribution between Divergence. Positive top-p outlier subgraphs with similar financial distributions lead to lower Score.

[0064] As an optional embodiment, the similarity between the negative top-p abnormal subgraph pairs is expressed as:

[0065]

[0066] in, Indicates that in the snapshot Time negative top-p abnormal subgraph Nodes randomly sampled from The encoding vector of the neighborhood of . and They are completely different and non-overlapping time negative top-p anomaly subgraphs, so the same node cannot be used for the negative top-p anomaly subgraph pair. .

[0067] As an optional embodiment, based on the similarity between the positive top-p abnormal subgraph pairs and the similarity between the negative top-p abnormal subgraph pairs, the expression of the optimization objective loss function is:

[0068]

[0069] in, Represents the temperature parameter.

[0070] As an optional embodiment, in step S104, the expression of the abnormality degree of the abnormal candidate subgraphs of different time slices is:

[0071]

[0072] in, For the Step-time abnormal candidate subgraph The degree of abnormality, Indicates in Step-time abnormal candidate subgraph The time density, ) indicates that Step-time abnormal candidate subgraph , represents the abnormal candidate subgraph at time s, Indicates the Step-time abnormal candidate subgraph of Statistics.

[0073] In this embodiment, if Figure 4 As shown in the leftmost figure, the model selection is such as Expand node 9 at the moment and Always exclude actions like node 4 to get higher rewards. First, define the abnormal candidate subgraph at step i. The degree of abnormality is as follows:

[0074]

[0075] in, Indicates in Step-time abnormal candidate subgraph , ) denotes the number of edges of the abnormal candidate subgraph at step , denotes the abnormal candidate subgraph at time s, denotes the abnormal candidate subgraph at time s, denotes the number of edges of the abnormal candidate subgraph at step , statistic of the abnormal candidate subgraph at step

[0076] Experimental Section:

[0077] Datasets: All methods are evaluated on three real-world financial transaction datasets.

[0078] Data Preprocessing:

[0079] Following the practice of AntiBenford, transactions with values less than 1 unit are excluded in the preprocessing procedure. Since the goal is to find abnormal subgraphs in temporal networks, the timestamps associated with transactions are preserved. The temporal graph is divided into multiple snapshots with equal time intervals, which have sequential relationships between them.

[0080] For each time interval, multiple transactions between two nodes are represented as a single edge. Each edge is associated with a transaction frequency and a financial distribution, which contains the probability of transaction amounts starting with each digit . This forms the input graph for analysis.

[0081] About the abnormal label:

[0082] Since these datasets lack inherent abnormal labels, specifically, the densest temporal subgraph that satisfies ψ(S) ≥ ρT is identified, and 80 of them are randomly selected for the training set.

[0083] Non-temporal anomaly detection methods:

[0084] Non-learning-based abnormal subgraph discovery methods are Holoscope, FlowScope, and AntiBenford.

[0085] Learning-based anomaly detection methods are CLARE, AS-GAE, GCAD, and SIGNET.

[0086] Temporal anomaly detection methods:

[0087] DeepSphere combines deep autoencoders and hyperspherical learning algorithms to detect anomalies in dynamic graphs.

[0088] RustGraph detects anomalies by jointly learning structure-time dependencies in temporal graphs. ​

[0089] AnoGraph proposes a sketching algorithm to detect anomalies in stream graphs.

[0090] Temporal dense subgraph discovery methods:

[0091] FAST-GA proposes a greedy approach to detect time-dense subgraphs.

[0092] OTCD and TopLC propose to use scalable temporal cores for temporal cohesive subgraph mining.

[0093] Hyperparameter settings: Hyperparameter learning is performed by using grid search. Hyperparameter tuning is performed using grid search. The following parameters are adjusted: the time interval length of snapshots is chosen from {5, 10, 20, 30} minutes. The maximum size m of nodes in the boundary is chosen from {100, 150, 200, 250}. The observation window Wr of time-dense subgraphs is chosen from {30, 60, 90, 120} minutes. The size of the top p candidates is chosen from {100, 150, 200, 250}. The number of sampled pairs is chosen from {5, 10, 15, 20}. The number of training subgraphs is chosen from {20, 50, 80, 110}.

[0094] Performance metrics: Performance is evaluated using standard metrics outlined in Antibenford [1], including statistical measures, subgraph density and where is the average value of each node in the subgraph . The metric evaluates the deviation of the transaction distribution of a subgraph from Benford's law, while computes the average deviation, denoted as . This average is important because in larger subgraphs, larger values can lead to seemingly smaller values. Therefore, provides a fair comparison of subgraphs of different sizes.

[0095] Additionally, according to the definition of anomalous subgraphs in Antibenford [1], where a subgraph is considered anomalous if the number of nodes , the anomaly score is used as a metric. A higher indicates a more anomalous subgraph.

[0096] Effectiveness evaluation:

[0097] The top- The effectiveness of the methods is evaluated in terms of their performance on the results, where The subgraph ranks of each method are sorted in descending order of value. To obtain stable results, the average results of five runs are obtained. From Tables 3, 4, and 5, we can observe that:

[0098] (1) The proposed TempASD exhibits higher abnormality scores than other baseline methods on all three datasets. On average, it is 755.9% more effective in abnormality score than the state-of-the-art method. This confirms the ability of TempASD to detect temporal anomalous subgraphs in dynamic financial networks.

[0099] (2) TempASD achieves the highest value on all three datasets compared to other baseline methods. Although it does not have the highest value on the top 5, 15, and 20 values on the Blur dataset, their average value, i.e., the value, is the highest. Since the value measures the average of each node, i.e.,

[0100] (3) The temporal density values of TempASD are comparable to those of other baseline methods, although some are not the highest. However, the abnormality score values of TempASD remain the highest. Some baseline methods, such as FlowScope, can have some of the highest temporal density values, but their abnormality score values are very low, indicating that they are not anomalous subgraphs at all.

[0101] (4) The abnormality score values decrease as the value of a increases. This is because larger a includes more subgraphs with lower abnormality scores. However, even when the value of a is large (e.g., a = 20), TempASD is still able to identify temporal subgraphs with higher abnormality scores.

[0102] The effects of the TempASD model on different hyperparameters are evaluated. Similarly, since the trends of , and density are similar to da(S), the results of Figure 6 are shown in .

[0103] Length of time interval for snapshots: Figure 6 (a) indicates that increasing the time interval performs best at an interval of 10 minutes. However, further increasing the time interval leads to a decrease in performance due to the loss of more detailed transaction information.

[0104] The maximum size of the boundary middle nodes m: Figure 6 (b) Shows that when m exceeds 150, more noise is introduced, leading to worse results.

[0105] The observation window of the time-intensive subgraph The range of: Figure 6 (c) Illustrates that extending beyond 60 leads to larger time subgraphs with lower anomaly degree, which negatively impacts performance.

[0106] The size of the top candidates: Figure 6 (d) Shows that increasing the number of candidates beyond 150 leads to more subgraphs that are not optimal, thus reducing performance.

[0107] The number of sampled pairs: Figure 6 (e) Indicates that increasing the number of sampled pairs beyond 10 leads to decreased performance due to possible increased noise.

[0108] The number of training subgraphs: Figure 6 (f) Shows that increasing the number of training subgraphs improves results up to 80. However, when the number reaches 100, the improvement effect weakens. Due to the scarcity of training subgraphs, this number is not further increased in the experiment.

[0109] Ablation experiments:

[0110] An ablation study was conducted on the TempASD model to evaluate the impact of its key components. The results on the Blur dataset and the top 15 and 20 were excluded due to their similar behavior patterns. The temporal anomaly refinement was omitted. The weights, i.e., transaction frequencies, were not considered in the temporal graph neural network. TempASDnk excluded the KL divergence of the financial distribution between pairs. TempASDnc replaced the time-intensive subgraph with a k-ego network. TempASDnf randomly selected connected nodes as boundary nodes.

[0111] Scalability related to the size of graph nodes:

[0112] Figure 7 (a-b) Show the detection time and training time of all methods in relation to the size of graph nodes. The proposed TempASD exhibits an approximate linear relationship with the size of nodes in financial networks on a logarithmic scale, indicating that its time complexity is polynomial.

[0113] Scalability related to the span of the time dimension:

[0114] Figure 7(c-d) demonstrate the relationship between detection time and training time with respect to days. The node size is fixed at 500,000 to ensure that the variation is solely due to the change in the time dimension. TempASD exhibits an approximate linear trend with respect to the number of days in the dynamic financial network.

[0115] Case Study

[0116] The interpretability of TempASD is investigated by presenting the time anomaly subgraphs detected in the dataset, as shown in Figure 8 (a). In this case study, the default snapshot time interval, i.e., ten minutes, is used. The proposed TempASD successfully identifies time-intensive subgraphs spanning three time intervals, i.e., T1, T2, and T3, whose financial distribution of transaction information deviates significantly from Benford's law. In contrast, AntiBenford, which is designed for static financial networks, can only detect non-time-intensive subgraphs that deviate from Benford's law. This limitation motivates the establishment of a theoretical foundation for detecting anomaly subgraphs in a temporal context. Although AnoGraph is designed for temporal networks, it does not incorporate transaction information, resulting in detected time subgraphs that do not deviate from Benford's law, meaning that the transaction information of the detected time subgraphs is normal.

[0117] In the embodiment of the present application, the dynamic financial network is obtained, wherein the dynamic financial network comprises nodes, edges between the nodes, and the edges between the nodes are used to represent transaction information between at least one of the nodes; the dynamic financial network is estimated by using approximate linear time complexity, dense subgraphs of different time slices are obtained, chi-square calculation is performed on each edge of each time slice of the dense subgraph, the chi-square of each edge of each time slice is obtained, nodes corresponding to edges with chi-square greater than or equal to the total chi-square of each time slice in the dense subgraph of each time slice are retained, and nodes corresponding to edges with chi-square less than the total chi-square of each time slice in the dense subgraph of each time slice are deleted, to obtain top-p abnormal subgraphs of different time slices; the top-p abnormal subgraphs of different time slices are encoded to obtain abnormal candidate subgraphs of different time slices, wherein the abnormal candidate subgraphs of different time slices contain boundaries; nodes in the abnormal candidate subgraphs of different time slices are excluded or expanded to obtain abnormal degrees of the abnormal candidate subgraphs of different time slices; the abnormal degrees of the abnormal candidate subgraphs of different time slices are maximized to obtain maximum rewards; the abnormal candidate subgraphs of different time slices are optimized by using the maximum rewards to obtain optimized abnormal candidate subgraphs of different time slices, thereby solving the technical problem that the prior art can only identify static abnormal subgraphs in the financial network and cannot detect changes and trends of user transaction behaviors over time, and achieving the technical effect that dynamic abnormal subgraphs in the financial network can be identified and can be used to detect changes and trends of user transaction behaviors over time.

[0118] The above-mentioned serial numbers of the embodiments of the present application are only for description, and do not represent the advantages or disadvantages of the embodiments.

[0119] In the above-mentioned embodiments of the present application, the description of each embodiment has its own emphasis, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.

[0120] In several embodiments provided in the present application, it should be understood that the disclosed technical contents can be implemented by other ways. Among them, the above-mentioned device embodiments are only schematic, for example, the division of units can be a logical function division, and actual implementation can have another division way, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed each other can be through some interface, indirect coupling or communication connection between units or modules, which can be electrical or other forms.

[0121] The units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, may be located in one place, or may be distributed to multiple units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.

[0122] In addition, each functional unit in each embodiment of the present application can be integrated in a first processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0123] The above is only the preferred embodiment of the present application, and it should be noted that for those skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, and these improvements and refinements should be considered as the protection scope of the present application.

Claims

1. A method for dynamic anomaly subgraph discovery in temporal financial networks, characterized in that, The method comprises: obtaining a dynamic financial network, wherein the dynamic financial network comprises nodes, edges between the nodes, and the edges between the nodes are used to represent transaction information between at least one of the nodes; estimating the dynamic financial network through an approximate linear time complexity to obtain dense subgraphs of different time slices, performing chi-square calculation on each edge of the dense subgraph of each time slice to obtain chi-square of each edge of each time slice, retaining nodes corresponding to edges of each time slice whose chi-square is greater than or equal to total chi-square of each time slice, and deleting nodes corresponding to edges of each time slice whose chi-square is less than total chi-square of each time slice to obtain top-p abnormal subgraphs of different time slices; encoding the top-p abnormal subgraphs of different time slices to obtain abnormal candidate subgraphs of different time slices, wherein the abnormal candidate subgraphs of different time slices contain boundaries; excluding or expanding nodes in the abnormal candidate subgraphs of different time slices to obtain abnormal degrees of the abnormal candidate subgraphs of different time slices; an expression of the abnormal degrees of the abnormal candidate subgraphs of different time slices is: in, is the abnormal candidate subgraph at step i The degree of abnormality, Indicates the abnormal candidate subgraph at step i The time density of h(S) represents the abnormal candidate subgraph at step i. The number of edges, T s represents the abnormal candidate subgraph at time s, Represents the abnormal candidate subgraph at step i The ψ statistic of maximizing a reward by maximizing a gain of the abnormal degrees of the abnormal candidate subgraphs of different time slices to obtain the reward; and optimizing the abnormal candidate subgraphs of different time slices through the reward to obtain optimized abnormal candidate subgraphs of different time slices.

2. The method of claim 1, wherein, an expression of the chi-square calculation on each edge of the dense subgraph of each time slice to obtain chi-square of the dense subgraph of each time slice is: where the chi-squared of the dense subgraph for each time slice, is the dense subgraph for each time slice the first digit of the transaction amount equals the number of temporal edges for o ∈ {1, ···, 9}, denotes the expected number of edges based on Benford's law, where the first digit of the transaction amount equals o.

3. The method of claim 1, wherein, a process of encoding the top-p abnormal subgraphs of different time slices is: extracting nodes of the top-p abnormal subgraphs of different time slices through different node numbers to obtain positive top-p abnormal subgraph pairs and negative sample top-p abnormal subgraph pairs in the top-p abnormal subgraphs of different time slices; calculating similarity between the positive top-p abnormal subgraph pairs and similarity between the negative top-p abnormal subgraph pairs; and optimizing a target loss function based on the similarity between the positive top-p abnormal subgraph pairs and the similarity between the negative top-p abnormal subgraph pairs.

4. The method of claim 3, wherein, an expression of the similarity between the positive top-p abnormal subgraph pairs is: where, denotes the encoding vector of the i-th node in the top-p anomalous subgraph G[S q ] at snapshot T a ; similarly, denotes the encoding vector of the neighborhood of node i in the top-p anomalous subgraph k at snapshot T , where is computed by applying the readout(·) function on the encoding vectors of all neighbors of node i in the top-p anomalous subgraph k at snapshot T ; is the number of nodes in the top-p anomalous subgraph G[S a ] at snapshot T0; and T V denotes the number of snapshots sampled for each time top-p anomalous subgraph; KL(·) denotes the KL divergence between the financial distributions of two top-p anomalous subgraphs G[S a ] and . Top-p anomalous subgraphs with similar financial distributions will result in lower KL scores.

5. The method of claim 4, wherein, an expression of the similarity between the negative top-p abnormal subgraph pairs is: wherein, represents the encoding vector of the neighborhood of node j randomly sampled in the snapshot T k at time moment negative top-p anomaly subgraph Since G[S a ] and are completely different and have no overlap, the same node i cannot be used for the pair of negative top-p anomaly subgraphs.

6. The method of claim 5, wherein, an expression of the optimization of the target loss function based on the similarity between the positive top-p abnormal subgraph pairs and the similarity between the negative top-p abnormal subgraph pairs is: wherein τ represents a temperature parameter.

7. A computer system, characterized by The method comprises: one or more processors, a computer readable storage medium for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method of claim 1.

8. A computer-readable storage medium, characterized in that a computer executable instruction is stored, and the instruction is used to implement the method of claim 1 when executed.

9. A computer program product, characterised in that a computer executable instruction is stored, and the instruction is used to implement the method of claim 1 when executed.

Citation Information

Patent Citations

  • Illegal financial activity detection method and system, electronic equipment and medium

    CN114782159A

  • Efficient abnormal subgraph discovery method in financial network

    CN117893212A