A financial case fund flow tracing analysis method
By combining social network analysis with graph neural network technology, the efficiency and accuracy issues in tracing the flow of funds in financial cases have been resolved. This enables in-depth analysis and anomaly detection of complex transaction networks, generating detailed reports on fund flow paths and supporting effective asset recovery efforts.
Patent Information
- Application Number
- CN202510237281.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-02
- Publication Date
- 2026-02-17
- Estimated Expiration
- 2045-03-02
AI Technical Summary
Existing technologies are inefficient in tracing and analyzing the flow of funds in financial cases, making it difficult to cope with complex and ever-changing financial crimes. They also have high computational complexity and poor real-time performance.
By combining social network analysis and graph neural network technology, a detailed report on fund flow paths is generated through data collection and preprocessing, transaction network construction, social network analysis, graph neural network analysis, and fund flow tracking.
It improves the accuracy and efficiency of fund flow tracing analysis, can identify complex transaction patterns and abnormal behaviors, provides detailed reports on fund flow paths, and supports the asset recovery efforts of financial institutions and law enforcement agencies.
Smart Images

Figure CN120106956B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer network technology, and specifically relates to a method for tracing and analyzing the flow of funds in financial cases. Background Technology
[0002] Financial crime fund flow tracing analysis refers to the process of tracking and analyzing the flow of funds involved in a case to reveal the entire process from the source to the final destination of the funds. This aims to discover the hiding places and flow patterns of illicit funds, thereby providing strong evidence and support for the recovery of stolen assets and mitigation of losses. Traditional fund flow tracing analysis methods mainly rely on manual investigation and auditing, using manual analysis of bank transaction records, payment platform data, and exchange data to reconstruct the flow of funds. However, this method is inefficient, prone to missing key information, and inadequate when faced with complex, multi-layered, multi-stage financial transaction networks and heterogeneous financial data from multiple sources.
[0003] Currently, the main technical methods for tracing and analyzing fund flows include rule-based analysis, machine learning models, and graph data analysis.
[0004] 1. Rule-based analysis methods rely on pre-set rules and conditions to perform transaction matching and pattern recognition, but this method is inflexible and difficult to deal with complex and ever-changing financial crimes.
[0005] 2. Machine learning models identify suspicious transactions by training classifiers or regression models. While this improves analysis efficiency to some extent, it still faces challenges such as difficulty in feature selection and insufficient model generalization ability.
[0006] 3. Graph data analysis technology treats financial transaction networks as graph structures and analyzes them through the relationships between nodes and edges. However, traditional graph analysis methods have shortcomings such as high computational complexity and poor real-time performance when dealing with large-scale and dynamically changing transaction networks.
[0007] Therefore, how to select advanced and appropriate technologies to further improve the accuracy and efficiency of fund flow tracing analysis has become a hot topic and a difficult point in the current financial case handling field. Summary of the Invention
[0008] (a) Technical problems to be solved
[0009] The technical problem this invention aims to solve is how to provide a method for tracing and analyzing the flow of funds in financial cases, so as to further improve the accuracy and efficiency of fund flow tracing and analysis by selecting advanced and appropriate technologies.
[0010] (II) Technical Solution
[0011] To address the aforementioned technical problems, this invention proposes a method for tracing and analyzing the flow of funds in financial cases. This method includes the following steps: data collection and preprocessing, transaction network construction, social network analysis, graph neural network analysis, and fund flow tracking.
[0012] The data collection and preprocessing steps are used to obtain transaction data related to financial cases from multiple data sources, clean and standardize the data to ensure its integrity and consistency, and provide a high-quality data foundation for subsequent transaction network construction and analysis.
[0013] The transaction network construction steps are used to transform preprocessed financial transaction data into a structured network graph, and to achieve systematic management of the transaction network by constructing a relationship graph of nodes and edges;
[0014] The social network analysis step is used to identify key paths and nodes of fund flows by analyzing nodes and edges in the transaction network, revealing the community structure and potential abnormal transaction behaviors in the transaction network. This step uses social network analysis technology to conduct in-depth analysis of the constructed transaction network graph, providing valuable insights and clues to support subsequent fund flow tracking.
[0015] The graph neural network analysis step is used to perform deep learning analysis on the transaction network using graph neural network technology, extract transaction features, and identify complex transaction patterns and abnormal transaction behaviors. By efficiently extracting and learning features from the network structure and node attributes, it is possible to deeply explore the implicit relationships and patterns in the transaction network.
[0016] The fund flow tracking step combines the results of social network analysis and graph neural network analysis to trace the flow path of suspicious funds. This step utilizes graph search algorithms and comprehensive analysis of transaction data and network structure to generate detailed fund flow path reports, assisting financial institutions and law enforcement agencies in taking effective actions to recover stolen assets and mitigate losses.
[0017] (III) Beneficial Effects
[0018] This invention proposes a method for tracing and analyzing the flow of funds in financial cases. It comprehensively utilizes social network analysis and graph neural network (GNN) technologies to achieve efficient feature extraction and model training. The invention combines social network analysis with GNN technology, leveraging the advantages of both to analyze and track fund flows within financial transaction networks. Social network analysis identifies key nodes and group structures in the transaction network through computational node centrality and community discovery. GNN analysis further extracts complex transaction features and patterns through deep learning models to identify abnormal transaction behaviors. Furthermore, by utilizing advanced GNN models such as graph convolutional networks and graph attention networks, combined with parallel computing and distributed training techniques, the model training process is efficient and possesses good scalability, adapting to the processing needs of large-scale financial transaction data.
[0019] This invention provides precise tracking and visualization of fund flows, generating detailed fund flow reports. In the fund flow tracking phase, a graph search algorithm reveals the destination and flow patterns of funds in detail. The system-generated fund flow map visually displays the flow of funds from source to destination, identifies key nodes and high-risk paths, and generates a detailed fund flow report. This report allows users to analyze and explore fund flow paths in depth as needed, providing flexible filtering and query functions. It offers specific action plans and evidence for financial institutions and law enforcement agencies, greatly enhancing the ability to understand and act on suspicious fund flows. Attached Figure Description
[0020] Figure 1 This is a general flowchart of the financial case fund flow tracing analysis method of the present invention;
[0021] Figure 2 A diagram illustrating the components of the data collection and preprocessing workflow;
[0022] Figure 3 A diagram illustrating the components of the process for building a transaction network;
[0023] Figure 4 A diagram illustrating the components of a social network analysis process;
[0024] Figure 5 This is a flowchart illustrating the components of a graph neural network analysis process.
[0025] Figure 6 This is a diagram showing the components of the fund flow tracking process. Detailed Implementation
[0026] To make the objectives, contents, and advantages of the present invention clearer, the specific embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and examples.
[0027] To address the aforementioned issues, this invention provides a method for tracing and analyzing the flow of funds in financial cases. This method integrates social network analysis and graph neural network technology, combined with fund flow path tracking and visualization, enabling in-depth analysis and anomaly detection of electronic financial data, thereby improving the accuracy and efficiency of fund flow tracing.
[0028] The method of this invention mainly includes five processes: data collection and preprocessing, transaction network construction, social network analysis, graph neural network analysis, and fund flow tracking, such as... Figure 1 As shown below, each process is described in detail.
[0029] 1. Data Collection and Preprocessing Steps
[0030] The main function of the data collection and preprocessing steps is to obtain transaction data related to financial cases from multiple data sources, clean and standardize the data to ensure its integrity and consistency, and provide a high-quality data foundation for subsequent transaction network construction and analysis.
[0031] The data collection and preprocessing process includes five parts: data collection, data cleaning, data standardization, data validation, and data storage. For example... Figure 2 As shown.
[0032] (1) Data collection
[0033] The data collection sub-step retrieves transaction data related to financial cases from multiple financial data sources, ensuring the comprehensiveness and accuracy of data collection.
[0034] In terms of technical implementation, transaction data from different financial institutions is acquired in real time or periodically through API interfaces, database connections, and batch data import. By establishing connections with various data sources, the required data can be extracted automatically, and preliminary format conversion and storage can be performed.
[0035] (2) Data cleaning
[0036] The data cleaning sub-step cleans and filters the raw transaction data collected from various data sources to provide high-quality analytical data.
[0037] In terms of technical implementation, data preprocessing techniques can be used to identify and correct errors and redundancies in the data through methods such as deduplication, noise reduction, and handling missing values.
[0038] (3) Data standardization
[0039] The data standardization sub-step transforms transaction data from different data sources into a unified format and structure. By standardizing data fields, formats, and types, differences between data sources are eliminated, enabling seamless integration and comparison of data in subsequent analysis.
[0040] In terms of technical implementation, common data processing tools can be used to achieve data field unification, data format unification, and data type conversion. These standardized processing steps can be automated. Through a series of predefined rules and algorithms, the system can efficiently convert multi-source data into a unified format.
[0041] (4) Data validation
[0042] The data validation sub-step ensures the integrity, consistency, and accuracy of data during data collection and preprocessing. Integrity validation ensures that each transaction record contains all necessary fields and information; consistency validation ensures that information for the same account is consistent across different data sources; and accuracy validation verifies the accuracy of the data through a series of verification rules and algorithms. By comprehensively validating and verifying the processed data, potential errors and problems can be identified and corrected promptly.
[0043] (5) Data storage
[0044] The data storage sub-step persistently stores high-quality data that has undergone cleaning, standardization, and verification. It manages the data's storage location and method, providing a stable and reliable data foundation for subsequent transaction network construction and analysis.
[0045] In terms of technical implementation, distributed databases or distributed file systems, such as Hadoop HDFS, Apache Cassandra, or Amazon S3, are typically used to ensure the security and scalability of data storage.
[0046] 2. Steps for building a trading network
[0047] The main function of the transaction network construction step is to transform preprocessed financial transaction data into a structured network graph. This step achieves systematic management of the transaction network by constructing a graph showing the relationships between nodes and edges.
[0048] The transaction network construction process includes four parts: node construction, edge construction, weighted processing, and storage and management. For example... Figure 3 As shown.
[0049] (1) Node Construction
[0050] The node construction sub-step maps the preprocessed financial account data to nodes in the transaction network graph. Each node represents a financial account and contains its detailed information and attributes. This process ensures that all relevant accounts in the transaction network accurately reflect their relationships and attributes.
[0051] (2) Edge construction
[0052] The edge construction sub-step maps the preprocessed transaction data into edges in the transaction network graph. Each edge represents a transaction, connects two financial accounts (nodes), and contains detailed information and attributes of the transaction. This process ensures that all transaction relationships in the network graph accurately reflect the path and characteristics of fund flows, providing a complete transaction network structure for subsequent analysis.
[0053] (3) Weighting
[0054] The weighted processing sub-step assigns weights to the edges in the transaction network to reflect the importance and frequency of transactions. Setting these weights helps identify and highlight key transaction paths and nodes in network analysis, enabling a more realistic representation of the characteristics and patterns of fund flows.
[0055] (4) Storage and Management
[0056] The storage and management sub-step persistently stores the constructed transaction network graph and provides efficient management and query functions, enabling the system to quickly access and process large-scale financial data and support complex network analysis tasks.
[0057] 3. Social Network Analysis Steps
[0058] The social network analysis step is one of the core processes of this invention. Its main function is to identify key paths and nodes of fund flows by analyzing nodes and edges in the transaction network, revealing the community structure and potential abnormal transaction behaviors within the network. This step utilizes social network analysis technology to conduct in-depth analysis of the constructed transaction network graph, providing valuable insights and clues to support subsequent fund flow tracking.
[0059] Social network analysis involves three parts: node centrality analysis, community detection, and pattern recognition and anomaly detection. For example... Figure 4 As shown.
[0060] (1) Node centrality analysis
[0061] The main function of the node centrality analysis sub-step is to calculate and analyze the centrality index of each node in the transaction network, identifying nodes that play a key role in the flow of funds. By analyzing node centrality, including degree centrality, intermediateness centrality, and proximity centrality, it is possible to determine which accounts hold significant positions in the transaction network and may be hubs or major participants in the flow of funds. This information is crucial for identifying critical paths and detecting potential illicit fund transfers.
[0062] (2) Community Discovery
[0063] The primary function of the community detection sub-step is to analyze nodes and edges in a transaction network using community detection algorithms to identify community structures within the network. Community detection algorithms include the Louvain algorithm, the Girvan-Newman algorithm, and spectral clustering algorithms. Community structure refers to closely related subgroups of nodes within a network, where transactions occur frequently between nodes within these groups but less frequently with external nodes. Identifying these communities helps uncover potential money laundering groups, criminal organizations, or other high-risk account groups, providing crucial clues for subsequent detailed investigations and analysis.
[0064] (3) Pattern recognition and anomaly detection
[0065] The main function of the pattern recognition and anomaly detection sub-step is to combine pattern recognition algorithms and anomaly detection techniques to analyze common transaction patterns in the transaction network and identify abnormal transaction behaviors, revealing potential illegal activities and suspicious fund flow paths. Pattern recognition algorithms include frequent pattern mining algorithms and clustering analysis algorithms; anomaly detection techniques include statistical anomaly detection and machine learning-based anomaly detection.
[0066] 4. Steps for analyzing graph neural networks
[0067] The graph neural network analysis step is also one of the core processes of this invention. Its main function is to use graph neural network (GNN) technology to perform deep learning analysis on the transaction network, extract transaction features, and identify complex transaction patterns and abnormal transaction behaviors. By efficiently extracting and learning features from the network structure and node attributes, it is possible to deeply explore the implicit relationships and patterns in the transaction network.
[0068] Graph neural network analysis and social network analysis work closely together to improve the accuracy and depth of transaction network analysis. Social network analysis provides preliminary structural information about the transaction network and identifies key nodes and communities by calculating node centrality and community discovery. Graph neural network analysis then builds upon this, using this structural information for deep learning to extract more complex transaction patterns and implicit relationships.
[0069] The centrality indicators and community structure features output from the social network analysis steps will serve as input to the graph neural network model, further enriching the model's feature space. By learning and training on these features, graph neural network analysis can more accurately identify abnormal transaction behavior. Furthermore, the key nodes and community structures identified in the social network analysis steps also provide graph neural network analysis with focused objects and regions, improving the model's targeting and efficiency.
[0070] Through this collaboration and integration, the social network analysis step and the graph neural network analysis step form a complete analysis chain, making full use of the structured information of the social network analysis step and the deep learning capabilities of the graph neural network analysis step, thereby comprehensively improving the effectiveness of tracking the flow of funds in the transaction network.
[0071] Graph neural network analysis steps include three parts: feature extraction, model training, and model prediction and application. For example... Figure 5 As shown.
[0072] (1) Feature extraction
[0073] The main function of the feature extraction sub-step is to extract features from nodes and edges in the transaction network. These features will serve as input data for graph neural network analysis. By extracting features from nodes and edges, important attributes and relationships in the transaction network can be captured, providing high-quality input for subsequent model training and prediction.
[0074] (2) Model Training
[0075] The main function of the model training sub-step is to train the graph neural network model using the node and edge feature data obtained from the feature extraction stage. By repeatedly adjusting the model parameters, the model can learn the complex relationships and patterns in the transaction network, improving its ability to identify normal and abnormal transaction behaviors.
[0076] (3) Model Prediction and Application
[0077] The main function of the model prediction and application sub-step is to use the pre-trained graph neural network model to predict and analyze new transaction data.
[0078] 5. Steps for tracking fund flows
[0079] The main function of the fund flow tracking step is to combine the results of the social network analysis and graph neural network analysis steps to trace the flow path of suspicious funds. This step utilizes graph search algorithms and comprehensive analysis of transaction data and network structure to generate detailed fund flow path reports, assisting financial institutions and law enforcement agencies in taking effective actions to recover stolen assets and mitigate losses.
[0080] The fund flow tracking process includes two parts: fund flow path tracking and fund flow path report generation. For example... Figure 6 As shown.
[0081] (1) Tracking the flow of funds
[0082] The main function of the fund flow tracking sub-step is to trace the flow path of suspicious funds within the transaction network, recording in detail the entire process from the source account to the final destination. By revealing the flow patterns and paths of funds, key nodes and potential transit accounts can be identified, providing specific action guidelines for recovering stolen assets and mitigating losses.
[0083] (2) Generation of Fund Flow Path Report
[0084] The main function of the fund flow path report generation sub-step is to organize and visualize the analysis results of the fund flow path tracking process, generating a detailed fund flow path report. This report intuitively displays the flow process of suspicious funds, the key nodes involved, and the transaction paths, providing clear guidance for financial institutions and law enforcement agencies to recover stolen assets and mitigate losses.
[0085] Example 1:
[0086] A method for tracing and analyzing the flow of funds in financial cases, with the following typical usage process.
[0087] Step 1: Data Collection and Preprocessing
[0088] Data collection and preprocessing is a crucial first step, ensuring the acquisition of high-quality financial transaction data to lay the foundation for subsequent transaction network construction and analysis. This process involves the following key steps.
[0089] (1) Data collection
[0090] 1) First, it is necessary to identify and connect to various financial data sources, including banks, payment platforms, and exchanges. Identifying data sources requires listing the data types and formats provided by each data source in detail to ensure comprehensive coverage of the required information. Establish stable and secure connections with each data source through API interfaces, JDBC / ODBC, and other database connection technologies, and use encryption protocols to protect the security of data during transmission.
[0091] 2) During data acquisition: For data sources that provide API interfaces, the system will write programs to periodically call the API to obtain data, and obtain the latest transaction data in real time through periodic polling or event-triggered methods; for data sources that provide database access permissions, SQL queries will be written to extract data from the database, and batch data export function will be used to obtain data in batches during offline periods; for data sources that provide log files or other forms of batch data, scripts will be written to parse the contents of the log files, extract transaction records, and process common file formats such as CSV, JSON, XML, etc.
[0092] 3) After data collection, the system stores the raw data in a temporary database or distributed file system, such as HDFS, to ensure data storage reliability and scalability. For ease of subsequent processing, the system performs preliminary format conversion on the data, such as standardizing data fields from different sources to standard field names and different time formats to the ISO standard time format. Furthermore, the system records a timestamp for each data collection operation to ensure data timeliness and traceability. These steps guarantee data consistency and integrity throughout the entire processing flow, laying the foundation for subsequent cleaning and standardization.
[0093] (2) Data cleaning
[0094] 1) First, the data is deduplicated by using a hash algorithm or a matching algorithm based on a unique identifier to identify and delete duplicate transaction records.
[0095] 2) Next, the noise filtering step filters out outliers and erroneous data by setting reasonable thresholds and rules, such as records with transaction amounts exceeding a reasonable range or incorrectly formatted timestamps.
[0096] 3) To handle missing values, the system will use a variety of methods, such as mean imputation, interpolation, or direct deletion of records containing a large number of missing values. The specific method depends on the characteristics of the data and the analysis requirements.
[0097] (3) Data standardization
[0098] 1) First, the data fields across different data sources are standardized. The same concepts in different data sources may use different field names and representations. For example, the "transaction time" field in one data source may be called "timestamp" in another. To resolve this inconsistency, the system will map and convert all fields to a standardized field name. For example, the "transaction time" field will be uniformly named "transaction_time".
[0099] 2) Next is the standardization of formatting. Date and time formats in transaction data often differ depending on the data source, such as ISO 8601 format, UNIX timestamps, or other localized date and time representations. The system will convert all date and time fields to the standard ISO 8601 format to ensure data consistency. In addition, currency amounts also need to be standardized. Different data sources may use different currency symbols and decimal places; the system will unify them to the specified decimal places and currency symbols.
[0100] 3) Data type conversion is another crucial step in the standardization process. Different data sources may use different data types for the same field; for example, representing transaction amounts as strings or numeric types. The system will convert these fields to appropriate data types based on standardization requirements, ensuring data consistency and operability. For example, all transaction amount fields will be converted to floating-point numbers, and all date and time fields will be converted to date and time types.
[0101] All these standardized processing steps are automated. Through a series of predefined rules and algorithms, the system can efficiently convert multi-source data into a unified format, providing a high-quality data foundation for subsequent analysis.
[0102] (4) Data validation
[0103] 1) First, an integrity verification is performed to ensure that each transaction record contains all necessary fields and information. For example, in financial transaction data, each record should at least include key fields such as transaction time, transaction amount, and account information. The system uses predefined verification rules to scan all records and check for any missing required fields. If any are found to be missing, the system will mark these records and take appropriate action, such as supplementing the missing information or isolating them for further inspection.
[0104] 2) Next, consistency verification is performed to ensure that information for the same account is consistent across different data sources. For example, the account balance and transaction records for the same account should be consistent across different banking systems. The system cross-validates this information across data sources, comparing records of the same account in different data sources to check for inconsistencies. If inconsistencies are found, the system records these differences and provides a detailed report for further investigation and correction.
[0105] 3) Finally, accuracy verification is performed using a series of validation rules and algorithms. For example, the system checks whether the transaction amount is within a reasonable range and whether the date and time are valid and within the expected time period. The system also uses historical data for trend analysis and anomaly detection to identify potential errors and abnormal behavior. For instance, by comparing historical transaction patterns, the system can detect unusually large transaction amounts or frequent transaction behavior, which may indicate data errors or potential fraud.
[0106] (5) Data storage
[0107] 1) Choose a suitable storage system to store the preprocessed data. For large-scale financial transaction data, distributed databases or distributed file systems, such as Hadoop HDFS, Apache Cassandra, or Amazon S3, are typically used to ensure the reliability and scalability of data storage. Distributed systems can handle large amounts of data and provide high availability and fault recovery capabilities, ensuring data distribution and redundant backup across different nodes.
[0108] 2) The cleaned, standardized, and validated data is written in batches to the distributed storage system. Each write operation is accompanied by a timestamp and a unique identifier to ensure data timeliness and traceability. This metadata helps the system accurately locate and access the required data in future data queries and analyses. To optimize data access performance, the system partitions and indexes data based on its characteristics. For example, data can be partitioned based on transaction date or account ID, allowing related data to be stored in adjacent physical locations, thereby speeding up data retrieval.
[0109] 3) Data access control and permission management ensure that only authorized users and system components can access and modify data. By setting different access permissions and encryption technologies, the system protects data security and privacy, preventing unauthorized access and data leakage.
[0110] 4) In addition to handling persistent data storage, it also provides fast data access and query capabilities. Leveraging the parallel processing capabilities of the distributed storage system, the system can efficiently execute complex query operations and support real-time analysis and processing of large-scale data.
[0111] Step 2: Building the Transaction Network
[0112] Next, the preprocessed financial transaction data is transformed into a structured transaction network diagram. This process involves the following key steps.
[0113] (1) Node Construction
[0114] 1) First, extract all financial account information from the preprocessed dataset. This information typically includes account ID, account holder name, account type (e.g., savings account, credit card account), bank name, and account status. The system assigns a unique node identifier to each account to ensure that each node in the network graph is unique and identifiable.
[0115] 2) Next, the system creates a node object containing all relevant account information and attributes. To efficiently process and store this node data, the system typically uses an object-oriented programming language to define the structure of the node object. Each node object includes multiple fields that store various attribute information of the account, such as account ID, holder name, account type, etc.
[0116] 3) After the nodes are created, the system will insert all node objects into the graph database. During the insertion process, the system will ensure the integrity and consistency of the node data, and at the same time, create necessary indexes to speed up subsequent queries and access.
[0117] Through these steps, the node construction process successfully transforms financial account data into structured network nodes, forming a complete set of nodes. These node sets serve as the foundation of the transaction network graph, providing necessary support for subsequent edge construction and weighted processing.
[0118] (2) Edge construction
[0119] 1) First, extract all transaction records from the preprocessed dataset. Each transaction record typically includes detailed information such as transaction ID, source account ID, target account ID, transaction amount, transaction time, and transaction type. The system assigns a unique edge identifier to each transaction to ensure that each edge in the network graph is unique and identifiable.
[0120] 2) Next, the system creates edge objects, which contain all relevant attributes from the transaction records. To efficiently process and store this edge data, the system typically uses an object-oriented programming language to define the structure of the edge objects. Each edge object includes multiple fields that store various attribute information about the transaction, such as transaction ID, source account ID, target account ID, transaction amount, and transaction time.
[0121] 3) After the edge objects are created, the system will insert them into the graph database. During the insertion process, the system ensures the integrity and consistency of the edge data, and establishes necessary indexes to speed up subsequent queries and access. The source account ID and target account ID in the edge object will correspond to two nodes in the network graph, representing the transaction relationship between the two accounts.
[0122] Through these steps, the edge construction sub-stage successfully transforms transaction data into structured network edges, forming a complete set of edges. These edge sets, as the core part of the transaction network graph, together with the node set, constitute the complete transaction network graph.
[0123] (3) Weighting
[0124] 1) Collect relevant data for each transaction edge, including attributes such as transaction amount, transaction frequency, and transaction time. The system calculates the weight of each edge based on the comprehensive information from these attributes. The weight calculation is typically based on a predefined weight model that considers factors such as transaction amount and frequency to comprehensively assess the importance of each transaction.
[0125] The design of a weighted model needs to be customized according to actual business requirements. A common method for calculating weights is to combine transaction amount and transaction frequency. The system first normalizes the transaction amount, distributing it between 0 and 1, and then performs a similar process on the transaction frequency. Normalization can employ methods such as max-min normalization or standard deviation normalization to eliminate magnitude differences between different transaction data. The normalized transaction amount and frequency are then weighted and summed according to the set weight ratios to obtain the final edge weights. For example, the edge weight can be expressed as: Edge weight = α * Normalized transaction amount + β * Normalized transaction frequency, where α and β are weight ratio coefficients that are adjusted according to actual needs.
[0126] In practice, the system calculates the weight for each edge and stores the result in an edge object. This weight represents the relative importance of the transaction; edges with high weights often represent paths for large or high-frequency transactions and deserve close attention. After the calculation is complete, the system re-inserts the weighted edge objects into the graph database, updating the existing transaction network graph.
[0127] 2) Weighted processing also requires handling weight adjustments for the time dimension. For example, the system can set a time window, considering only transaction data within a specific time period to address changes in trading patterns over time. This allows for the generation of different weighted network diagrams for different periods, reflecting the temporal characteristics of capital flows.
[0128] 3) To ensure the efficiency of weighted processing, the system will use parallel computing technology, especially when processing large-scale transaction data, and use multi-threaded or distributed computing frameworks (such as Apache Spark) to accelerate the weight calculation process.
[0129] Through these technical means, the weighted processing stage can quickly and accurately assign weights to each edge in the transaction network, thereby enhancing the analytical value of the entire network graph.
[0130] (4) Storage and Management
[0131] 1) First, choose a suitable graph database to store the constructed transaction network graph. Graph databases such as Neo4j and Amazon Neptune are suitable choices, as they are specifically designed for handling complex network relationship data and can efficiently store and manage large-scale nodes and edges. The system will insert the constructed node and edge objects into the graph database to ensure persistent data storage.
[0132] 2) During the insertion process, as mentioned earlier, the system will create an index for each node and edge to accelerate data retrieval. Indexes are typically based on frequently used query fields, such as account ID and transaction time, which can significantly improve data access efficiency. The system will also design and create appropriate graph structures according to business needs, making data storage and management more logical and hierarchical.
[0133] 3) To ensure data consistency, the system employs transaction management and data backup mechanisms. Transaction management guarantees that during data insertion, updating, and deletion, operations either all succeed or all fail, preventing data inconsistencies. Data backup ensures rapid data recovery in the event of loss or corruption through regular full and incremental backups. Backup data is typically stored on off-site servers or in cloud storage to mitigate the risk of data loss due to physical disasters.
[0134] 4) In addition, the storage and management system provides rich data query and retrieval functions. Users can perform complex query operations using the query language provided by the graph database (such as Cypher forNeo4j) to quickly retrieve and analyze specific nodes and edges in the transaction network graph. The system supports multiple query modes, such as path query, subgraph query, and pattern matching, to meet the needs of different analysis tasks.
[0135] 5) To improve query performance, the system leverages the parallel processing capabilities and distributed storage architecture of graph databases. By executing query tasks in parallel, the system can handle large-scale query requests in a short time. The distributed storage architecture ensures that data is stored and computed across multiple nodes, improving the system's scalability and fault tolerance.
[0136] Through these technical means, the storage and management process successfully persists the transaction network graph and provides efficient management and query functions.
[0137] Step 3: Social Network Analysis
[0138] Once the transaction network is constructed, the system loads and parses it to perform social network analysis. This process involves the following key steps.
[0139] (1) Node centrality analysis
[0140] Node centrality analysis is primarily based on structural analysis of the transaction network graph. The transaction network graph consists of nodes (accounts) and edges (transactions). The system first loads and parses the transaction network graph stored in the graph database. Next, the system calculates the centrality metric for each node; commonly used centrality metrics include degree centrality, betweenness centrality, and proximity centrality.
[0141] 1) Degree centrality is the most intuitive indicator of centrality, representing the number of connections a node has, i.e., the number of edges directly connected to that node. Nodes with high degree centrality are typically accounts with frequent transactions. The system iterates through each node in the network graph, counts the degree of each node, and records it in the node's attributes. This process can be efficiently implemented using a graph database query language (such as Cypher forNeo4j).
[0142] 2) Betweenness centrality measures the degree to which a node acts as a mediator in a network, i.e., the frequency with which the node appears in the shortest paths between other node pairs. Nodes with high betweenness centrality are often hubs for fund flows and may be key points for fund transfers. Calculating betweenness centrality requires considering the shortest paths between all node pairs in the network graph, which is usually implemented using the Floyd-Warshall algorithm or Dijkstra's algorithm. The system calculates the frequency of each node's appearance in all shortest paths to obtain the betweenness centrality of each node.
[0143] 3) Proximity centrality measures the average shortest path length from a node to all other nodes in a network. Nodes with high proximity centrality can typically reach other nodes in the network quickly, indicating efficient transmission capabilities in fund flows. Calculating proximity centrality also requires the use of a shortest path algorithm. The system traverses the network graph, calculates the shortest path length from each node to all other nodes, and takes the average to obtain the proximity centrality of each node.
[0144] To improve computational efficiency, the system employs parallel computing techniques, especially when processing large-scale transaction networks. It leverages multi-threaded or distributed computing frameworks (such as Apache Spark GraphX) to accelerate the calculation of centrality metrics. The calculated centrality metrics are stored in a graph database as node attributes for easy querying and analysis.
[0145] (2) Community Discovery
[0146] 1) Analyze transaction network graphs using various community detection algorithms to identify and classify community structures within the network. Commonly used community detection algorithms include the Louvain algorithm, the Girvan-Newman algorithm, and spectral clustering algorithms.
[0147] a) The Louvain algorithm is a greedy optimization algorithm that identifies community structure by maximizing modularity. Modularity is a metric for the quality of network partitioning, representing the density of node connections in the network compared to a random network. The Louvain algorithm first treats each node as an independent community, then repeatedly merges adjacent nodes, gradually building larger communities until the modularity no longer increases. This algorithm performs exceptionally well in handling large-scale networks, efficiently identifying community structure. In its implementation, the system traverses each node in the network, continuously adjusting the communities to which each node belongs to maximizes modularity, ultimately determining the community to which each node belongs.
[0148] (b) The Girvan-Newman algorithm is a community detection algorithm based on edge betweenness centrality. It splits communities by progressively removing edges with high betweenness centrality from the network. Edges with high betweenness centrality typically connect different communities; by removing these edges, the network gradually splits into multiple subgroups. In its implementation, the system first calculates the betweenness centrality of all edges in the network, and then sequentially removes these high betweenness centrality edges until the network splits into multiple connected components. Each connected component represents a community. Although the Girvan-Newman algorithm has high computational complexity, it performs well in revealing the hierarchical structure of networks.
[0149] c) Spectral clustering is a community detection method based on graph spectral decomposition. It divides nodes into different communities using the eigenvalues and eigenvectors of the graph's Laplacian matrix. In its implementation, the system first calculates the Laplacian matrix of the network graph, then performs eigenvalue decomposition, selecting the top k eigenvectors to embed the nodes into a k-dimensional space. By applying the k-means clustering algorithm in this k-dimensional space, the system can divide the nodes into different communities. Spectral clustering has advantages in handling complex network structures and discovering hidden communities.
[0150] 2) In practical applications, the system may combine multiple algorithms for community detection to improve the accuracy and robustness of identification. After calculation, the system stores the community partitioning results in a graph database as node attributes for easy subsequent querying and analysis. Each node is labeled with its associated community, and the system also generates a relationship graph between communities, displaying transaction relationships and interaction strength between communities.
[0151] 3) To improve computational efficiency, especially when dealing with large-scale transaction networks, the system employs parallel and distributed computing techniques. Utilizing multi-threaded or distributed computing frameworks (such as Apache Spark GraphX), the system can quickly complete community discovery computation tasks, ensuring the timeliness and accuracy of analysis results.
[0152] (3) Pattern recognition and anomaly detection
[0153] By combining pattern recognition algorithms and anomaly detection technology, in-depth analysis of transaction networks is conducted to provide early warnings and clues about potential financial crimes. Specifically, this includes the following components.
[0154] 1) First, a comprehensive data analysis of the transaction network is conducted to extract common transaction patterns and behavioral characteristics. The pattern recognition part mainly employs techniques such as frequent pattern mining and cluster analysis to identify common patterns in the transaction network. For example, through frequent pattern mining, the system can identify frequently occurring transaction combinations between certain accounts within a specific time period. These combinations may represent legitimate business transaction patterns or potential abnormal patterns. Cluster analysis identifies common patterns in the transaction network by grouping similar transaction behaviors together through clustering transaction data.
[0155] 2) In order to identify abnormal transaction behavior, the system will use statistical and machine learning-based anomaly detection techniques.
[0156] a) Statistical anomaly detection techniques identify transactions that significantly deviate from normal behavior by establishing statistical models of trading behavior. For example, the system can build a transaction amount distribution model for each account and use statistical analysis to identify transactions with amounts significantly higher or lower than the normal range. Commonly used statistical anomaly detection methods include z-score and IQR (interquartile range), which can effectively identify abnormal behavior based on a single indicator.
[0157] b) Machine learning-based anomaly detection techniques identify complex, multi-dimensional anomaly patterns by training models. The system first trains anomaly detection models using historical trading data, such as Isolation Forest, Local Outlier Factor (LOF), or deep learning-based autoencoders. These models learn the characteristics of normal trading behavior and identify anomalous trades that deviate from these characteristics. Isolation Forest constructs multiple isolated trees by randomly selecting features and split points to isolate data points and calculate anomaly scores. The Local Outlier Factor identifies local anomalies by comparing the density differences between data points and their neighbors. The autoencoder constructs an encoder-decoder network to learn a low-dimensional representation of the data, and data points with large reconstruction errors are considered anomalies.
[0158] 3) The system also incorporates time series analysis techniques to identify temporal anomalies in trading behavior. For example, using time series anomaly detection algorithms (such as ARIMA and LSTM), the system can analyze the trends of trading amounts and frequencies over time, identifying sudden surges or drops in trading activity. By analyzing the trends and changes in trading time series, the system can provide more accurate anomaly detection results.
[0159] Step 4: Graph Neural Network Analysis
[0160] Building upon social network analysis, graph neural network analysis is then performed to extract more complex transaction patterns and implicit relationships. This process involves the following key steps.
[0161] (1) Feature extraction
[0162] 1) First, load the transaction network data from the graph database, including information on all nodes (accounts) and edges (transactions).
[0163] a) For nodes, the features the system needs to extract include account ID, account type, transaction frequency, transaction amount, and centrality indicators. Transaction frequency refers to the number of transactions an account makes within a certain period, which can be obtained by calculating the node's degree centrality. Transaction amount is the sum or average of all transactions for that account. Centrality indicators such as degree centrality, betweenness centrality, and proximity centrality can be calculated using social network analysis processes, reflecting the node's importance and connectivity in the network.
[0164] b) For edges, the features the system needs to extract include transaction amount, transaction time, transaction type, and transaction frequency. Transaction amount refers to the specific monetary value of each transaction, transaction time refers to the timestamp of the transaction, and transaction type refers to the specific nature of the transaction (such as transfer, payment, etc.). Transaction frequency refers to the number of transactions between two nodes (accounts) within a certain period of time, which can be calculated by counting the number of transactions between pairs of the same nodes.
[0165] 2) During feature extraction, the system standardizes these features to ensure consistency in their numerical range and distribution. Standardization includes normalizing numerical features to the range of 0 to 1, or converting them to a standard normal distribution (mean 0, standard deviation 1). This method helps eliminate differences in magnitude between different features, improving the effectiveness of model training. For time-related features, the system converts timestamps into periodic features, such as daily, weekly, and monthly, to capture the temporal patterns of transaction behavior.
[0166] 3) The feature extraction stage also requires processing the high-dimensional features of nodes and edges, converting them into low-dimensional embedding representations. The system can use pre-trained embedding models, such as Word2Vec or Node2Vec, to map the high-dimensional features of nodes and edges into a low-dimensional space, generating dense feature vectors. These low-dimensional feature vectors can be used as input to the graph neural network, helping the model better capture the complex relationships and patterns in the transaction network.
[0167] 4) In actual feature extraction operations, the system processes node and edge feature extraction tasks in parallel. Especially when dealing with large-scale transaction networks, it utilizes multi-threading or distributed computing frameworks (such as Apache Spark) to accelerate the feature extraction process. The system stores the extracted features in memory or a distributed storage system for subsequent model training and prediction.
[0168] These technical means enable the feature extraction process to accurately extract the features of nodes and edges from the transaction network.
[0169] (2) Model Training
[0170] 1) First, the system receives the feature data generated in the previous step, including the feature vectors of nodes and edges. Then, the system selects an appropriate graph neural network model for training. Commonly used graph neural network models include Graph Convolutional Network (GCN) and Graph Attention Network (GAT).
[0171] a) For Graph Convolutional Networks (GCNs), the system learns the global features of nodes by aggregating neighbor information layer by layer through convolution operations. At the start of training, the system initializes model parameters, including the weight matrix and bias terms. Then, the system inputs feature data into the model, calculates the convolution operation at each layer, and obtains a new feature representation for each node by weighted summation of the features of the node and its neighboring nodes. After each convolution operation, the system applies a non-linear activation function (such as ReLU) to further enhance the model's expressive power.
[0172] b) Graph Attention Networks (GAT) use an attention mechanism to assign different weights to the neighbors of each node, allowing for more flexible learning of relationships between nodes. The system first initializes the attention weights and inputs node features into the model. For each node, the system calculates its attention score with its neighboring nodes, obtaining a new feature representation of the node by multiplying the node features by the attention weights. Similar to GCNs, GAT also applies a non-linear activation function after each layer operation.
[0173] 2) During training, the system uses the training data for forward propagation to calculate the model's output. For supervised learning, the system compares the model output with the true labels and calculates the loss function (such as cross-entropy loss or mean squared error). Then, the system uses the backpropagation algorithm to calculate the gradient of the loss function with respect to the model parameters, updates the model parameters, and minimizes the loss function. This process is repeated until the loss function converges or the preset number of training epochs is reached.
[0174] 3) Model training also requires handling batch training and validation of data. The system divides the training data into multiple mini-batches, inputting each batch into the model for training to improve computational efficiency and the model's generalization ability. After each training cycle, the system uses a validation dataset to evaluate model performance and prevent overfitting. The validation process includes calculating validation loss and accuracy, and monitoring the model's performance on the validation set.
[0175] 4) To improve training efficiency, the system employs parallel computing and distributed training techniques. Through multi-threaded or distributed computing frameworks (such as TensorFlow Distributed or PyTorch Distributed), the system can accelerate the model training process and handle large-scale transaction network data. Furthermore, the system uses optimization algorithms (such as Adam or RMSprop) to adjust the learning rate and model parameters, thereby improving training performance.
[0176] 5) After the model training is completed, the system will save the trained model parameters and structure as the basis for subsequent predictions and applications. The trained model can accurately capture the complex relationships and patterns in the trading network and identify normal and abnormal trading behaviors.
[0177] (3) Model Prediction and Application
[0178] 1) First, new transaction data is received, and features of nodes and edges are extracted from it. This process is similar to feature extraction, including collecting node attributes (such as account ID, account type, transaction frequency, transaction amount, etc.) and edge attributes (such as transaction amount, transaction time, transaction type, etc.), and performing standardization processing. The extracted feature data is then used as input to the pre-trained graph neural network model.
[0179] 2) Before inputting data into the model, the system loads the pre-trained model parameters and structure. Model parameters include the weight matrix and bias terms, while the model structure includes information such as the number of network layers, the number of nodes per layer, and the activation function. After loading, the system inputs the features of the new data into the model for forward propagation computation. During forward propagation, the system computes the feature representations of nodes layer by layer, aggregating information from neighboring nodes through convolution operations or attention mechanisms to generate new node feature representations.
[0180] 3) Model Inference Calculation. For Graph Convolutional Networks (GCNs), the system aggregates the neighbor information of nodes layer by layer through convolution operations to obtain a new feature representation for each node. For Graph Attention Networks (GATs), the system calculates attention scores and assigns different weights to the neighbors of each node, flexibly learning the relationships between nodes. Regardless of the model used, the final output is the predicted score or classification label for each node.
[0181] 4) Post-process the model output. For classification tasks, the model output classification label indicates the category of the node or edge, such as normal transaction or abnormal transaction. For regression tasks, the model output prediction score indicates the degree of suspicion of the transaction; the higher the score, the more suspicious the transaction. The system will mark transactions with scores higher than a preset threshold as abnormal transactions and record detailed information about these abnormal transactions.
[0182] Step 5: Tracking Fund Flows
[0183] Fund flow tracking is the final step of this invention. It uses the results of social network analysis and graph neural network analysis to trace the flow path of suspicious funds. This process involves the following key steps.
[0184] (1) Tracking the flow of funds
[0185] 1) The system first extracts transaction records marked as abnormal from the transaction network as a starting point. These abnormal transaction records originate from previous social network analysis and graph neural network analysis processes. The system then traces the flow of funds step by step based on the source account of each abnormal transaction.
[0186] 2) In practice, the system first loads the graph structure of the transaction network, including detailed information on all nodes (accounts) and edges (transactions). Then, the system selects either Depth-First Search (DFS) or Breadth-First Search (BFS) algorithms for path tracing. The DFS algorithm recursively explores from a node (account) along the transaction path (edge) to the furthest possible node, then backtracks to continue exploring other paths. This approach is effective for discovering long chains of transaction paths. Conversely, the BFS algorithm starts from the initial node and expands outwards layer by layer, exploring all adjacent nodes before further expansion. This approach is suitable for discovering extensive, multi-level transaction networks.
[0187] 3) During path tracing, the system records detailed transaction information for each step, including transaction amount, transaction time, source account, and target account. The system also calculates the cumulative transaction amount and number of transactions for each path to help identify high-frequency and large-value transaction paths. To improve efficiency, the system uses data structures (such as queues or stacks) to manage the search process and avoids duplicate searches by marking visited nodes.
[0188] 4) During the path tracing process, the system not only records the transaction path, but also combines the results from the social network analysis process to identify key nodes on the path and highlight them in the path map.
[0189] 5) In addition, the system will further evaluate transactions along the path based on predefined rules and anomaly patterns. For example, the system will check whether the transaction amount is abnormally large, the transaction frequency is abnormally high, and the transaction time is concentrated in specific time periods. These evaluation results help to further confirm and rule out suspicious fund flow paths.
[0190] Through these technological means, fund flow tracking can comprehensively and accurately track the flow path of suspicious funds, record every step of the fund flow process in detail, and reveal the destination and flow pattern of funds.
[0191] (2) Generation of Fund Flow Path Report
[0192] 1) First, the system receives tracking results from fund flow path tracing, including detailed path information, transaction records, and key nodes. This information is then processed and analyzed to ensure the accuracy and comprehensiveness of the report.
[0193] 2) During the information processing, the system will summarize each fund flow path, including all related transaction records and account information. The summarized information for each path includes detailed information such as transaction ID, source account, target account, transaction amount, and transaction time. In addition, the system will calculate the cumulative transaction amount and number of transactions along the path, identifying high-frequency and high-value transaction paths. The system will then categorize and sort these paths, highlighting the most important and suspicious paths based on their degree of anomaly and risk level.
[0194] 3) To visually represent the flow of funds, the system generates a visualized fund flow path diagram. The visualization process utilizes graph drawing tools (such as D3.js or Graphviz) and the visualization capabilities of a graph database to present the fund flow path graphically. The fund flow path diagram shows the complete flow of funds from the source account to the final destination, including each transaction node and transaction path. The system uses different colors and line styles to distinguish between normal and abnormal transactions, and highlights key nodes and high-risk paths. The size and color of nodes indicate their importance, such as transaction frequency and amount, while the thickness and color of edges indicate the weight and risk level of the transaction.
[0195] 4) In addition to the path diagram, the report includes detailed text descriptions explaining the key details and anomalies of each fund flow path. The text descriptions provide analysis of each key node, including its role, trading behavior, and risk assessment. For example, the system will explain in detail why a node is identified as a key node, and how its trading frequency and amount are abnormal. The text descriptions also provide a comprehensive assessment of each path, indicating its importance and risk level within the overall fund flow.
[0196] 5) The system also generates interactive reports, allowing users to conduct in-depth analysis and exploration as needed. These interactive reports are displayed through a graphical user interface, allowing users to click on nodes and edges in the path graph to view detailed transaction information and analysis results. The interactive reports also provide filtering and search functions, allowing users to filter and query fund flow paths based on different criteria (such as transaction amount, time range, account type, etc.). The generation of interactive reports relies on modern web technologies and data visualization tools, such as HTML5, JavaScript, and D3.js, to achieve real-time interaction and data display.
[0197] 6) Finally, the system will generate an official document version of the fund flow path report, output in PDF or other document formats. These document versions of the report include all path diagrams, text descriptions, and analysis results, and can be easily distributed and archived as official report files.
[0198] The beneficial effects of this invention are:
[0199] This invention integrates social network analysis and graph neural network (GNN) technologies to achieve efficient feature extraction and model training. It combines the advantages of both to analyze and track fund flows within financial transaction networks. Social network analysis identifies key nodes and group structures in the transaction network through node centrality and community discovery. GNN analysis, using deep learning models, further extracts complex transaction features and patterns to identify abnormal trading behavior. Furthermore, by utilizing advanced GNN models such as graph convolutional networks and graph attention networks, combined with parallel computing and distributed training techniques, the model training process is highly efficient and scalable, adapting to the processing needs of large-scale financial transaction data.
[0200] This invention provides precise tracking and visualization of fund flows, generating detailed fund flow reports. In the fund flow tracking phase, a graph search algorithm reveals the destination and flow patterns of funds in detail. The system-generated fund flow map visually displays the flow of funds from source to destination, identifies key nodes and high-risk paths, and generates a detailed fund flow report. This report allows users to analyze and explore fund flow paths in depth as needed, providing flexible filtering and query functions. It offers specific action plans and evidence for financial institutions and law enforcement agencies, greatly enhancing the ability to understand and act on suspicious fund flows.
[0201] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the technical principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for tracing financial case money flow, characterized in that, The method comprises the following steps: data collection and preprocessing, transaction network construction, social network analysis, graph neural network analysis, and fund flow tracking; The data collection and preprocessing step is used to obtain transaction data related to financial cases from multiple data sources, clean and standardize the data, and ensure the integrity and consistency of the data, providing a high-quality data foundation for subsequent transaction network construction and analysis; The transaction network construction step is used to convert the preprocessed financial transaction data into a structured network graph, and to realize systematic management of the transaction network by constructing a relationship graph of nodes and edges; The social network analysis step is used to identify key paths and nodes of fund flow by analyzing nodes and edges in the transaction network, and to reveal community structures and potential abnormal transaction behaviors in the transaction network; this step uses social network analysis techniques to conduct in-depth analysis of the constructed transaction network graph, providing valuable insights and clues to support subsequent fund flow tracking; The graph neural network analysis step is used to use graph neural network GNN technology to conduct deep learning analysis on the transaction network, extract transaction features, and identify complex transaction patterns and abnormal transaction behaviors; Through efficient feature extraction and learning of network structure and node attributes, the hidden relationships and patterns in the transaction network can be deeply mined; The fund flow tracking step combines the results of the social network analysis step and the graph neural network analysis step to track the flow path of suspicious funds; this step uses graph search algorithms and analyzes transaction data and network structure to generate a detailed fund flow path report, assisting financial institutions and law enforcement departments in taking effective actions to recover losses; Among them, The graph neural network analysis and the social network analysis work closely together to improve the accuracy and depth of transaction network analysis; the social network analysis provides preliminary structural information of the transaction network and identification of key nodes and communities by calculating node centrality and community discovery; the graph neural network analysis uses this structural information for deep learning to extract more complex transaction patterns and hidden relationships; the centrality indicators and community structure features output by the social network analysis step will be used as input for the graph neural network model to further enrich the feature space of the model; the graph neural network analysis identifies abnormal transaction behaviors by learning and training these features; In addition, the key nodes and community structures identified by the social network analysis step also provide objects and areas of focus for the graph neural network analysis, improving the relevance and efficiency of the model.
2. The financial case fund flow tracing analysis method of claim 1, wherein, The data collection and preprocessing step comprises five parts: data collection, data cleaning, data standardization, data verification, and data storage: The data collection sub-step obtains transaction data related to financial cases from multiple financial data sources to ensure the comprehensiveness and accuracy of data collection; The data cleaning sub-step cleans and filters the original transaction data collected from various data sources to provide high-quality analysis data; The data standardization sub-step uniformly converts transaction data from different data sources into consistent formats and structures by standardizing data fields, formats, and types, eliminating differences between data sources, so that data can be seamlessly integrated and compared in subsequent analysis processes; The data verification sub-step ensures the integrity, consistency, and accuracy of data during data collection and preprocessing; The data storage sub-step persistently stores high-quality data that has been cleaned, standardized, and verified, manages the storage location and method of data, and provides a stable and reliable data foundation for subsequent transaction network construction and analysis.
3. The financial case fund flow tracing analysis method of claim 1, wherein, The transaction network construction step includes node construction, edge construction, weighting processing, storage and management; The node construction sub-step maps preprocessed financial account data into nodes in the transaction network graph, with each node representing a financial account and containing detailed information and attributes of the account. This process ensures that all related accounts in the transaction network accurately reflect their relationships and attributes; The edge construction sub-step maps preprocessed transaction data into edges in the transaction network graph, with each edge representing a transaction, connecting two financial accounts, and containing detailed information and attributes of the transaction. This process ensures that all transaction relationships in the network graph accurately reflect the path and characteristics of fund flow, providing a complete transaction network structure for subsequent analysis; The weighting processing sub-step assigns weights to edges in the transaction network to reflect the importance and frequency of transactions. The setting of weights helps to identify and highlight key transaction paths and nodes in network analysis, and can more realistically show the characteristics and patterns of fund flow; The storage and management sub-step persistently stores the constructed transaction network graph and provides efficient management and query functions, enabling the system to quickly access and process large-scale financial data and support complex network analysis tasks.
4. The financial case fund flow tracing analysis method of claim 3, wherein, The weight calculation method combines transaction amount and transaction frequency; the system first normalizes the transaction amount to distribute between 0 and 1, then processes the transaction frequency similarly; the normalization process uses the maximum and minimum normalization or standard deviation normalization method to eliminate the magnitude difference between different transaction data; The normalized transaction amount and frequency are weighted and summed according to the set weight ratio to obtain the final edge weight; The edge weight is represented as: edge weight = α * normalized transaction amount + β * normalized transaction frequency, where α and β are weight ratio coefficients adjusted according to actual needs.
5. The financial case fund flow tracing analysis method of claim 1, wherein, The social network analysis step includes node centrality analysis, community discovery, pattern recognition and anomaly detection; The node centrality analysis sub-step calculates and analyzes the centrality indicators of each node in the transaction network, identifies nodes that play a key role in fund flow, and determines which accounts have an important position in the transaction network by analyzing the centrality of nodes, including degree centrality, betweenness centrality, and closeness centrality, which may be the hub or main participant of fund flow. The community discovery sub-step analyzes the nodes and edges in the transaction network through a community discovery algorithm to identify the community structure in the network; the community structure refers to subgroups of nodes in the network that have close relationships, and the nodes within these groups frequently transact with each other while transacting less with external nodes; identifying these communities helps to reveal potential money laundering gangs, criminal organizations, or other high-risk account groups; The pattern recognition and anomaly detection sub-step combines pattern recognition algorithms and anomaly detection techniques to analyze common transaction patterns and identify abnormal transaction behaviors in the transaction network, revealing potential illegal activities and suspicious fund flow paths.
6. The financial case fund flow tracing analysis method of claim 5, wherein, The community discovery algorithm includes Louvain algorithm, Girvan-Newman algorithm, and spectral clustering algorithm; the pattern recognition algorithm includes frequent pattern mining algorithm and clustering analysis algorithm; the anomaly detection technique includes statistical-based anomaly detection and machine learning-based anomaly detection.
7. The financial case fund flow tracing analysis method of claim 6, wherein, Degree centrality represents the number of connections of a node, i.e., the number of edges directly connected to the node; a node with high degree centrality is a frequently transacting account; the system will traverse each node in the network graph, count the degree of each node, and record it in the node's attributes; Betweenness centrality measures the degree to which a node acts as an intermediary in the network, i.e., the frequency of its appearance in the shortest paths between other node pairs; a node with high betweenness centrality is a hub of fund flow and may be a key fund transfer point; calculating betweenness centrality requires considering the shortest paths between all node pairs in the network graph, which can be achieved through the Floyd-Warshall algorithm or Dijkstra algorithm; the system will calculate the frequency of each node in all shortest paths to obtain the betweenness centrality of each node; Closeness centrality measures the average shortest path length from a node to all other nodes in the network; A node with high closeness centrality can quickly reach other nodes in the network, indicating its high efficiency in fund flow transmission; calculating closeness centrality also requires using the shortest path algorithm, and the system will traverse the network graph to calculate the shortest path length from each node to all other nodes and take the average to obtain the closeness centrality of each node; Use multiple community discovery algorithms to analyze the transaction network graph, identify and divide the community structure in the network; After the calculation is completed, the system will store the community division results in the graph database as the attributes of the nodes, facilitating subsequent queries and analysis; each node will be labeled as belonging to a community, and the system will also generate a relationship graph between communities to show the transaction relationships and interaction strength between communities; Combine pattern recognition algorithms and anomaly detection techniques to conduct in-depth analysis of the transaction network and provide early warnings and clues for potential financial crimes; Specifically includes the following parts; First, conduct a comprehensive data analysis of the transaction network to extract common transaction patterns and behavioral characteristics; The pattern recognition part uses frequent pattern mining and cluster analysis techniques to identify common patterns in the transaction network; cluster analysis identifies common patterns in the transaction network by grouping similar transaction behaviors into one category through clustering transaction data. Statistical anomaly detection techniques identify transactions that significantly deviate from normal behavior by establishing statistical models of trading behavior; machine learning-based anomaly detection techniques identify complex, multi-dimensional anomaly patterns by training models. The system first uses historical transaction data to train anomaly detection models. These models can learn the characteristics of normal trading behavior and identify abnormal transactions that deviate from these characteristics. The system will also incorporate time series analysis techniques to identify timing anomalies in trading behavior.
8. The financial case fund flow tracing analysis method of claim 1, wherein, The graph neural network analysis steps include: feature extraction, model training, model prediction, and application. The feature extraction sub-step extracts features of nodes and edges from the transaction network. These features will serve as input data for graph neural network analysis. By extracting features from nodes and edges, the important attributes and relationships in the transaction network are captured, providing high-quality input for subsequent model training and prediction. The model training sub-step uses node and edge feature data obtained from the feature extraction stage to train the graph neural network model. By repeatedly adjusting the model parameters, it learns the complex relationships and patterns in the transaction network, thereby improving the ability to identify normal and abnormal transaction behaviors. The model prediction and application sub-step utilizes the pre-trained graph neural network model to predict and analyze new transaction data.
9. The method for tracing and analyzing the flow of funds in financial cases as described in claim 8, characterized in that, In the feature extraction sub-step, transaction network data, including information on all nodes and edges, is loaded from the graph database. For nodes, the features that the system needs to extract include account ID, account type, transaction frequency, transaction amount, and centrality index. Transaction frequency refers to the number of transactions of an account within a certain period of time, which is obtained by calculating the degree centrality of the node. Transaction amount is the sum or average of all transaction amounts of the account. Centrality index is calculated through the social network analysis process and reflects the importance and connectivity of the node in the network. For edges, the features that the system needs to extract include transaction amount, transaction time, transaction type, and transaction frequency. Transaction amount refers to the specific amount of each transaction, transaction time refers to the timestamp of the transaction, transaction type refers to the specific nature of the transaction, and transaction frequency refers to the number of transactions between two nodes within a certain period of time, which is calculated by counting the number of transactions between pairs of the same nodes. The feature extraction stage also needs to process the high-dimensional features of nodes and edges, converting them into low-dimensional embedding representations; the system will process the feature extraction tasks of nodes and edges in parallel; the system will store the extracted features in memory or a distributed storage system for subsequent model training and prediction; In the model training sub-step, the feature data generated in the previous step is first received, including the feature vectors of nodes and the feature vectors of edges; Then, the system will select an appropriate graph neural network model for training, including: Graph Convolutional Network (GCN) and Graph Attention Network (GAT); For Graph Convolutional Networks (GCNs), the system learns the global features of nodes by aggregating the neighbor information of nodes layer by layer through convolution operations. At the beginning of the training process, the system initializes the model parameters, including the weight matrix and bias terms. Then, the system inputs the feature data into the model, calculates the convolution operation of each layer, and obtains a new feature representation of each node by weighted summation of the features of the node and its neighboring nodes. After the convolution operation of each layer, the system applies a non-linear activation function to further enhance the expressive power of the model. The Graph Attention Network (GAT) uses an attention mechanism to assign different weights to the neighbors of each node, enabling more flexible learning of relationships between nodes. The system first initializes the attention weights and inputs the node features into the model. For each node, the system calculates its attention score with its neighboring nodes, and obtains a new feature representation of the node by multiplying the node features with the attention weights. GAT applies a non-linear activation function after each layer operation. During training, the system uses training data for forward propagation to calculate the model's output. For supervised learning, the system compares the model output with the true labels and calculates the loss function. Then, the system uses the backpropagation algorithm to calculate the gradient of the loss function with respect to the model parameters and updates the model parameters to minimize the loss function. This process is repeated until the loss function converges or the preset number of training rounds is reached. The system divides the training data into multiple mini-batches and inputs them into the model one batch at a time for training to improve computational efficiency and the model's generalization ability. After each training cycle, the system uses a validation dataset to evaluate the model's performance and prevent overfitting. The validation process includes calculating the validation loss and accuracy and monitoring the model's performance on the validation set. The system will employ parallel computing and distributed training techniques; through a multi-threaded or distributed computing framework, the system can accelerate the model training process and process large-scale transaction network data; in addition, the system will use optimization algorithms to adjust the learning rate and model parameters. After the model is trained, the system will save the trained model parameters and structure as the basis for subsequent prediction and application; the trained model can accurately capture the complex relationships and patterns in the trading network and identify normal and abnormal trading behaviors. In the model prediction and application sub-step, new transaction data is first received, and the features of nodes and edges are extracted and standardized. The extracted feature data is then used as input to the pre-trained graph neural network model. Before inputting data into the model, the system first loads the trained model parameters and structure. The model parameters include the weight matrix and bias terms, while the model structure includes the number of network layers, the number of nodes in each layer, and activation function information. After loading, the system inputs the features of the new data into the model and performs forward propagation calculations. During forward propagation, the system calculates the feature representations of nodes layer by layer, and aggregates the information of neighboring nodes through convolution operations or attention mechanisms to generate new node feature representations. Model inference computation; for Graph Convolutional Network (GCN), the system aggregates the neighbor information of nodes layer by layer through convolution operations to obtain a new feature representation of each node; for Graph Attention Network (GAT), the system calculates attention scores and assigns different weights to the neighbors of each node to flexibly learn the relationships between nodes; regardless of which model is used, the final output is the predicted score or classification label of each node. The model output is post-processed. For classification tasks, the model output classification label indicates the category of the node or edge, including: normal transaction or abnormal transaction. For regression tasks, the model output prediction score indicates the degree of suspicion of the transaction. The higher the score, the more suspicious the transaction. The system will mark transactions with scores higher than the preset threshold as abnormal transactions according to the preset threshold and record the detailed information of these abnormal transactions.
10. The financial case fund flow tracing analysis method of claim 1, wherein, The fund flow tracking steps include: fund flow path tracking and fund flow path report generation; The fund flow path tracking sub-step tracks the flow path of suspicious funds in the transaction network, records the entire process of funds from the source account to the final destination in detail, and identifies key nodes and potential transit accounts by revealing the flow patterns and paths of funds, providing specific action guidelines for recovering stolen assets and mitigating losses. The fund flow path report generation sub-step organizes and visualizes the analysis results of the fund flow path tracking process, generating a detailed fund flow path report. This report intuitively displays the flow process of suspicious funds, the key nodes involved, and the transaction path, providing clear guidance for financial institutions and law enforcement agencies to recover stolen assets and mitigate losses.
Citation Information
Patent Citations
Wallet address mining method based on graph neural network
CN117764581A