Suspicious money laundering gang identification method and device, electronic equipment and storage medium

CN122736741APending Publication Date: 2026-09-11MINSHENG BANKING CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610513616.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-17
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

近年来,电信诈骗、地下钱庄、跨境赌博、虚假贸易及虚拟资产交易等犯罪活动不断演化,犯罪团伙通常通过多账户协同操作、分层转移资金以及复杂交易网络进行资金清洗,使得资金流动路径更加隐蔽、交易链条更加复杂,传统以单账户或单笔交易规则为核心的反洗钱监测方式难以有效识别团伙化洗钱行为

Benefits of technology

本申请实施例中,通过构建客户交易关系网络,完整捕捉客户间的交易关联及资金流转特征,为洗钱团伙识别提供全面的数据支撑。借助洗钱风险预测模型同步提取局部时序特征与全局关联特征,实现对每个客户洗钱风险概率的精准量化,有效筛选出高风险客户形成风险客户列表,避免低风险客户干扰与高风险客户漏判。再采用自底向上的识别方法,以高风险客户为起点,通过挖掘节点间直接或间接交易关联形成候选团伙,能够完整保留隐蔽性强、层级复杂、分散化洗钱团伙的拓扑结构,规避传统识别方法拆分团伙导致的识别偏差,进而综合提高此类洗钱团伙的识别准确率,降低识别过程中的误判率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122736741A_ABST
    Figure CN122736741A_ABST
Patent Text Reader

Abstract

This application provides a method, apparatus, electronic device, and storage medium for identifying suspicious money laundering groups. The method includes: constructing a customer transaction relationship network based on customer identifiers and transaction information of all customers within an enterprise; extracting local temporal features and global correlation features from the customer transaction relationship network using a money laundering risk prediction model, and processing the local temporal features and global correlation features to obtain the money laundering risk probability for each customer; filtering out risky customers whose money laundering risk probability is greater than a probability threshold to obtain a risky customer list; and using a bottom-up identification method, identifying suspicious money laundering groups based on the risky customer list and the customer transaction relationship network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a method, apparatus, electronic device, and storage medium for identifying suspicious money laundering groups. Background Technology

[0002] With the accelerating digitalization of the global financial system and the continuous strengthening of my country's financial regulatory system, money laundering crimes are becoming increasingly organized, professional, and covert. In recent years, criminal activities such as telecommunications fraud, underground banks, cross-border gambling, fraudulent trade, and virtual asset transactions have evolved. Criminal gangs typically launder funds through multi-account coordinated operations, tiered fund transfers, and complex transaction networks, making fund flows more concealed and transaction chains more complex. Traditional anti-money laundering monitoring methods, which focus on single accounts or single transactions, are insufficient to effectively identify organized money laundering activities.

[0003] Therefore, how to improve the accuracy of identifying money laundering gangs that are highly concealed, have complex hierarchies, and are decentralized, and how to reduce the rate of missed identification, are urgent problems to be solved. Summary of the Invention

[0004] The technical problem to be solved by the embodiments of this application is to provide a method, device, electronic device and storage medium for identifying suspicious money laundering gangs, so as to improve the identification accuracy of money laundering gangs that are highly concealed, hierarchical and decentralized, and reduce the misjudgment rate of money laundering gang identification.

[0005] In a first aspect, embodiments of this application provide a method for identifying suspicious money laundering groups, the method comprising: Based on the customer identifiers and transaction information of all customers within the enterprise, a customer transaction relationship network is constructed. The money laundering risk prediction model extracts local temporal features and global correlation features from the customer transaction relationship network, and processes the local temporal features and global correlation features to obtain the money laundering risk probability of each customer. Filter out all customers whose money laundering risk probability is greater than the probability threshold to obtain a list of risky customers; A bottom-up identification method is used to identify suspicious money laundering groups based on the risk customer list and the customer transaction relationship network.

[0006] Secondly, embodiments of this application provide a device for identifying suspicious money laundering groups, the device comprising: The network construction module is used to build a customer transaction relationship network based on the customer identifiers and customer transaction information of all customers within the enterprise; The probability acquisition module is used to extract local temporal features and global correlation features from the customer transaction relationship network through the money laundering risk prediction model, and to process the local temporal features and global correlation features to obtain the money laundering risk probability of each customer. The list retrieval module is used to filter out all customers whose money laundering risk probability is greater than the probability threshold, and obtain a list of risky customers. The gang identification module is used to identify suspicious money laundering gangs using a bottom-up identification method based on the risk customer list and the customer transaction relationship network.

[0007] Thirdly, embodiments of this application provide an electronic device, including: A processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method for identifying suspicious money laundering groups as described above.

[0008] Fourthly, embodiments of this application provide a computer-readable storage medium that, when the instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to perform the identification method for suspicious money laundering groups described in any of the preceding claims.

[0009] Compared with the prior art, the embodiments of this application have the following advantages: In this embodiment, a customer transaction relationship network is constructed to fully capture the transaction connections and fund flow characteristics among customers, providing comprehensive data support for the identification of money laundering gangs. By simultaneously extracting local temporal features and global correlation features using a money laundering risk prediction model, the probability of money laundering risk for each customer is accurately quantified, effectively screening high-risk customers to form a risk customer list, avoiding interference from low-risk customers and missing high-risk customers. Furthermore, a bottom-up identification method is adopted, starting with high-risk customers, and forming candidate gangs by mining direct or indirect transaction connections between nodes. This method can fully preserve the topological structure of highly concealed, hierarchical, and decentralized money laundering gangs, avoiding the identification bias caused by splitting gangs in traditional identification methods, thereby comprehensively improving the identification accuracy of such money laundering gangs and reducing the false judgment rate during the identification process.

[0010] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description

[0011] Figure 1 A flowchart illustrating the steps of a method for identifying a suspicious money laundering group, as provided in this application embodiment; Figure 2 A schematic diagram of a money laundering risk prediction model provided in an embodiment of this application; Figure 3 A schematic diagram of a bottom-up money laundering gang detection technology process provided for an embodiment of this application; Figure 4 A schematic diagram illustrating the implementation process of a transaction network construction module provided in an embodiment of this application; Figure 5 A schematic diagram of a dual-view customer money laundering risk assessment module provided for an embodiment of this application; Figure 6 A schematic diagram of the structure of a device for identifying a suspected money laundering gang, provided in an embodiment of this application; Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0012] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0013] The terminology used in the embodiments of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. The singular forms “a,” “the,” and “the” used in the embodiments of this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.

[0014] Reference Figure 1 The diagram illustrates a flowchart of the steps involved in identifying a suspicious money laundering group, as provided in an embodiment of this application. Figure 1 As shown, the identification method for this suspicious money laundering group may include steps 101 to 104.

[0015] Step 101: Construct a customer transaction relationship network based on the customer identifiers and transaction information of all customers within the enterprise.

[0016] The embodiments of this application can be applied to the scenario of money laundering gang detection. Money laundering gang detection refers to the technical process of using multi-dimensional data such as financial transaction flows, account relationships, and fund transfer paths, employing techniques such as graph mining, anomaly detection, and machine learning to extract features, perform correlation analysis, and structural clustering on massive amounts of financial data. This process automatically identifies suspicious money laundering organizations that operate collaboratively by multiple accounts or entities, exhibiting hierarchical and covert characteristics, reconstructs the fund flow network, and achieves gang location, risk assessment, and behavior tracing.

[0017] A customer identifier is a unique identity identifier assigned by a financial institution to each customer (such as a customer number, ID card number associated code, or terminal identifier of the customer's terminal). It is used to uniquely distinguish different customers and serves as the core identifier of nodes in the customer transaction relationship network.

[0018] Customer transaction information refers to the complete data collected by financial institutions related to customer transactions, including basic customer information, account statistics, transaction statistics, transaction sequence characteristics, transaction amount, number of transactions, transaction cycle, etc. It is the basic data for building customer transaction relationship networks and extracting risk characteristics.

[0019] A customer transaction relationship network is a homogeneous graph network (denoted as G=(V,E), where V is the set of nodes and E is the set of edges) constructed with customer identifiers as network nodes and the fund transaction relationships between customers as directed edges. Nodes carry customer base, account, and transaction-related attributes, while edges carry attributes such as transaction amount, number of transactions, and cycle, which are used to represent the transaction relationships and fund flow relationships between customers.

[0020] When investigating money laundering groups, data collection can begin: collect customer identifiers for all customers within financial institutions, as well as multi-source heterogeneous initial transaction-related data. The collection dimensions include: basic customer information (customer identity information, customer type, customer level, etc.), account statistics (account type, account status, account opening / closing time, number of account uses, etc.), transaction statistics (transaction frequency, transaction time period, abnormal transaction markers, etc.), and transaction flow information (transaction time, transaction amount, counterparty, transaction summary, etc.). This ensures that the data comprehensively covers the fund flow, information flow, and account association flow data required for investigating money laundering groups.

[0021] Next, data preprocessing is performed: the initial collected data is standardized, outliers are removed, and data quality is ensured. Specific operations include: ① Missing value handling: a small number of missing fields are filled in, and a large number of missing fields are removed; ② Redundant data removal: duplicate transaction records, duplicate account information, and duplicate counterparty information are merged and deleted; ③ Standardized coding: transaction currencies, timestamp formats, etc. are standardized to transform unstructured data into standardized computable data, providing standardized input for network construction.

[0022] Finally, the customer transaction relationship network is constructed, that is, the customer identifier is used as the node and the transaction information is used as the attribute to build the customer transaction relationship network. The specific implementation process of network construction will be described in detail in the following embodiments, and will not be repeated here.

[0023] Step 102: Extract local temporal features and global correlation features from the customer transaction relationship network using the money laundering risk prediction model, and process the local temporal features and global correlation features to obtain the money laundering risk probability for each customer.

[0024] Money laundering risk prediction models are models used to quantitatively assess a customer's money laundering risk. They include local feature extraction networks (deep temporal networks with integrated attention mechanisms), graph neural networks (one of GCN, GraphSAGE, or GAT), feature fusion networks, and fully connected networks. Their core function is to extract the customer's local temporal features and global correlation features, and output the customer's money laundering risk probability.

[0025] Local temporal features refer to the dynamic evolution features of a single customer's transaction behavior extracted from the data. These features can characterize the temporal dependencies, abnormal transaction periods, and amount characteristics of customer transactions, and reflect the risk attributes of individual customer transaction behavior.

[0026] Global correlation features refer to the correlation features of customers extracted in the global transaction network. They can characterize the strength of the transaction correlation between customers and other customers, the risk transmission effect, and the network topology features, reflecting the global risk attributes of customers in the transaction network.

[0027] Money laundering risk probability refers to a value (range 0-1) output by a money laundering risk prediction model, used to quantify the likelihood of a customer engaging in money laundering activities. The higher the value, the greater the likelihood of the customer engaging in money laundering activities.

[0028] After obtaining the customer transaction relationship network, it can be input into the money laundering risk prediction model. Specifically, local temporal features of individual customers can be extracted using a deep temporal network in the model, and global correlation features of individual customers can be extracted using a graph neural network in the model. Then, the local temporal features and global correlation features are fused to obtain fused features, which are then processed by a fully connected network to output the money laundering risk probability of each customer. This implementation process will be described in detail in the following embodiments, and will not be repeated here.

[0029] Step 103: Filter out all customers whose money laundering risk probability is greater than the probability threshold to obtain a list of risky customers.

[0030] The probability threshold refers to a risk assessment threshold (such as 0.7) preset by financial institutions based on anti-money laundering business needs and historical data verification results. It is used to screen high-risk customers, and the threshold can be dynamically adjusted according to business scenarios.

[0031] The risk customer list refers to the set of customers whose money laundering risk probability is greater than the probability threshold. It includes customer identification, money laundering risk probability, risk score details (local time-series feature contribution, global correlation feature contribution), etc., and serves as the core seed node source for bottom-up mining of money laundering gangs.

[0032] After obtaining the money laundering risk probability for each customer, a list of high-risk customers can be generated by filtering out those with a probability greater than a threshold. Specifically, a money laundering risk probability threshold (e.g., 0.7) can be preset based on the anti-money laundering business needs of financial institutions and the verification results of historical money laundering case data. This threshold can be dynamically adjusted according to business scenarios and regulatory requirements. The money laundering risk probabilities of all customers are iterated through, and customers with probability values ​​greater than the preset threshold are selected as high-risk customers. Finally, the relevant information of all high-risk customers is integrated to generate a list of high-risk customers. This list includes customer identification, money laundering risk probability, and risk score details (local temporal feature contribution, global correlation feature contribution), which are used for gang identification in step 104.

[0033] Step 104: Using a bottom-up identification method, identify suspicious money laundering groups based on the risk customer list and the customer transaction relationship network.

[0034] The bottom-up identification method uses high-risk customers as core seed nodes, relies on the customer transaction relationship network, gradually explores the direct / indirect transaction connections between seed nodes, aggregates them to form candidate money laundering groups, and then screens suspicious money laundering groups through risk ranking. This method is different from the traditional top-down model of "first splitting the network and then judging the group".

[0035] Suspicious money laundering groups consist of multiple nodes with direct / indirect transactional connections, all of which are high-risk clients. They are characterized by concealment and coordination. High-risk candidate groups selected after risk ranking have a high probability that their members are engaged in coordinated money laundering activities.

[0036] After obtaining the list of high-risk customers, a bottom-up identification method can be used to identify suspicious money laundering groups based on the list and the customer transaction relationship network. Specifically, the customers with the highest risk probability can be selected from the high-risk customer list as seed nodes. Then, starting from the seed nodes, customers in the customer transaction relationship network that have a direct or indirect relationship with the seed nodes and belong to the high-risk customer list can be identified, thereby obtaining suspicious money laundering groups. The identification process of suspicious money laundering groups will be described in detail in the following embodiments, and will not be repeated here.

[0037] In this embodiment, the data generated in the above steps undergoes a triple check of validity, completeness, and timeliness to eliminate abnormal data, ensuring data consistency and real-time availability. Then, a multi-dimensional interactive interface tailored to analytical habits is constructed to achieve layered and interconnected data presentation. The overview layer presents core monitoring indicators of the group in dashboard format, while the details layer displays auxiliary information for identification, such as group members, fund flows, and risk transmission chains. It supports exporting results in multiple formats, analysis annotation, and risk warnings, ensuring traceability and timely response in analysis; it also supports integration with existing monitoring platforms, establishing a user feedback mechanism to continuously optimize the interface and functionality.

[0038] This application's embodiments construct a customer transaction relationship network to fully capture the transaction connections and fund flow characteristics among customers, providing comprehensive data support for the identification of money laundering gangs. By leveraging a money laundering risk prediction model to simultaneously extract local temporal features and global correlation features, it achieves accurate quantification of the money laundering risk probability for each customer, effectively screening high-risk customers to form a risk customer list, avoiding interference from low-risk customers and the omission of high-risk customers. Furthermore, a bottom-up identification method is employed, starting with high-risk customers, and forming candidate gangs by mining direct or indirect transaction connections between nodes. This method can fully preserve the topological structure of highly concealed, hierarchical, and decentralized money laundering gangs, avoiding the identification bias caused by the fragmentation of gangs in traditional identification methods. This comprehensively improves the identification accuracy of such money laundering gangs and reduces the false positive rate during the identification process.

[0039] In one implementation of this application, step 101 may include sub-steps A1 to A3.

[0040] Sub-step A1: Obtain customer identifiers and initial transaction data for all customers, and perform preprocessing on the initial transaction data, including missing value handling, redundancy removal, and standardized coding, to obtain the preprocessed customer transaction information.

[0041] In this embodiment, initial transaction data refers to raw transaction-related data collected directly from various business systems of financial institutions (such as payment systems, online banking, interbank clearing systems, etc.) without any processing. It includes multi-source heterogeneous data such as customer transaction records and account association records, and serves as the basis for subsequent preprocessing.

[0042] When investigating money laundering groups, unique customer identifiers for all clients within financial institutions can be collected in batches, along with initial transaction data for each client simultaneously. The collection scope covers all raw data related to client transactions, including basic information, account information, and transaction history, ensuring comprehensive coverage of all information required for subsequent node and edge attribute construction, with no missing core data. Afterward, the following preprocessing operations can be performed: 1. Missing value handling: Classify and process fields with missing values ​​in the initial transaction data. For fields with a small amount of missing data (e.g., missing percentage less than 5%), fill them with the mean, median, or interpolation of similar customer data. For fields with a large amount of missing data (e.g., missing percentage greater than 50%), and for which accuracy cannot be guaranteed by filling, remove them directly to avoid missing data affecting the accuracy of subsequent processing results.

[0043] 2. Redundancy Removal: Redundant records in the initial transaction data are screened and deleted, with a focus on removing duplicate transaction records (records that are entered repeatedly for the same transaction), duplicate account information (duplicate registration information for the same type of account for the same customer), and duplicate counterparty information. At the same time, related transaction data of the same customer are merged to ensure that the data is unique and concise, and to reduce the amount of redundant calculations in subsequent processing.

[0044] 3. Standardized Coding: The dataset after missing value processing and redundancy removal is uniformly standardized and encoded. Specifically, this includes: converting transaction amounts in different currencies to a unified currency (such as RMB), converting timestamps in different formats (such as year-month-day-hour-minute-second format and timestamp format) to a standard time format, encoding and converting unstructured transaction summaries, customer types, and other information, and converting all data into a standardized format that can be directly used for subsequent node and edge attribute construction, ultimately obtaining preprocessed customer transaction information.

[0045] Sub-step A2: Using each customer's customer identifier as a network node, and determining the node attributes of each network node based on the first and second features in the customer transaction information; wherein, the first feature includes at least: customer basic information, account statistics information and transaction statistics information, and the second feature includes: transaction sequence features.

[0046] The first feature refers to the core feature set used to determine the attributes of network nodes. It includes at least customer basic information, account statistics, and transaction statistics, focusing on customer static and basic transaction-related attributes to provide basic support for node risk feature analysis.

[0047] The second feature refers to the feature used to supplement the attributes of network nodes. It only includes transaction sequence features, focuses on the temporal change pattern of customer transaction behavior, and provides a data foundation for subsequent local temporal feature extraction.

[0048] Node attributes refer to the characteristic information of each node (customer identifier) ​​in the customer transaction relationship network. They are composed of the first feature and the second feature and are used to describe the customer's basic information, account status, transaction patterns and other core attributes.

[0049] Each customer's unique identifier serves as an independent node in the customer transaction relationship network, with each customer identifier corresponding to a unique node, ensuring that each node can uniquely distinguish different customers.

[0050] From the preprocessed customer transaction information, the first feature corresponding to each node is extracted, which specifically includes: basic customer information (customer identity information, customer type, customer level, etc.), account statistics (number of accounts, account type, account usage frequency, account opening and closing time, etc.), and transaction statistics (total transaction amount, transaction frequency, transaction time distribution, number of abnormal transactions, etc.). The extracted first features are organized and verified to ensure that the feature data is accurate and complete.

[0051] Subsequently, from the preprocessed customer transaction information, the second feature (transaction sequence feature) corresponding to each node is extracted. Specifically, it is the time-series transaction data of each customer within the preset business monitoring window, including the transaction amount, counterparty, transaction frequency and other information of the customer at different time points, forming a complete transaction time-series sequence. This ensures that the time-series sequence can fully cover the customer's transaction behavior cycle and capture the dynamic changes in transaction behavior.

[0052] Then, the first and second features corresponding to each node can be merged and integrated into the complete node attributes of that node. The attribute information corresponds one-to-one with the node, realizing a comprehensive characterization of each customer's basic information, account status, transaction patterns and time series changes, providing reliable node basic data for subsequent extraction of local time series features and global correlation features.

[0053] Sub-step A3: Establish connection edges between nodes in the network that have transaction-related behaviors, and determine the edge attributes of the connection edges according to the third feature in the customer transaction information to obtain the customer transaction relationship network; wherein, the third feature includes at least: transaction amount information, number of transactions information, and transaction cycle information.

[0054] The third feature refers to the feature set used to determine the attributes of network connection edges, which includes at least transaction amount information, number of transactions information, and transaction cycle information, representing the strength and pattern of transaction associations between customers.

[0055] Transaction-related behavior refers to the actual flow of funds between two customers, that is, one customer conducting a transaction with another customer, which is the core basis for establishing network connection edges.

[0056] A connecting edge is a link used to connect two nodes with related transaction behaviors in a customer transaction relationship network. It represents the transaction relationship between nodes and provides a structural basis for subsequent global association feature extraction.

[0057] Edge attributes refer to the characteristic information attached to the network connection edge. They are determined by the third feature and are used to characterize details such as the amount, frequency, and cycle of transactions between two nodes, reflecting the closeness of the transaction association.

[0058] Traverse all network nodes and combine them with preprocessed customer transaction information to determine whether there is a transaction association between any two nodes. The determination criteria are: there is a real flow of funds between the customers corresponding to the two nodes, that is, one customer has had at least one valid transaction with the other customer (after removing invalid and duplicate transactions).

[0059] For two nodes that are determined to have a transaction relationship, a connection edge is established between them. The transaction flow is represented by a directed edge (i.e., from the node corresponding to the paying customer to the node corresponding to the receiving customer) to ensure that the connection edge can accurately reflect the direction and relationship of fund transfer between customers.

[0060] Subsequently, from the preprocessed customer transaction information, the third feature corresponding to each connection edge is extracted, which specifically includes: transaction amount information (cumulative transaction amount between customers corresponding to two nodes, highest / lowest single transaction amount, etc.), number of transactions information (cumulative number of transactions between customers corresponding to two nodes, average daily / monthly number of transactions, etc.), and transaction cycle information (transaction interval between customers corresponding to two nodes, transaction frequency pattern, etc.). The extracted third features are organized and quantified to form standardized edge feature data.

[0061] Finally, the third feature corresponding to each connecting edge can be used as the edge attribute of that connecting edge, so as to accurately characterize the transaction correlation strength and transaction pattern between nodes; by integrating all network nodes, connecting edges and their corresponding node attributes and edge attributes, a complete customer transaction relationship network can be formed, providing structured network data support for subsequent money laundering risk prediction and money laundering gang identification.

[0062] The data acquisition and network construction process can be as follows: Figure 4 As shown, the process may include: S1.1 Data Acquisition, which involves collecting basic customer information, customer account information, customer transaction history information, and customer transaction behavior information. S1.2 Data Preprocessing: Handling missing values, deleting redundant data, and standardizing the encoding of the collected data. S1.3 Network Construction: Constructing graph nodes and edges based on the preprocessed data to build a customer transaction relationship network.

[0063] The embodiments of this application can achieve standardized processing and structured transformation of customer transaction data, effectively eliminating abnormal and redundant data and ensuring data quality. The constructed customer transaction relationship network can completely and accurately represent the transaction relationships between customers and the core attributes of each customer and each transaction.

[0064] In one implementation of this application, step 102 may include sub-steps B1 to B4.

[0065] Sub-step B1: Input the transaction sequence features of a single customer into the local feature extraction network for temporal feature encoding to obtain the local temporal features of the corresponding customer.

[0066] In this embodiment, the structure of the money laundering risk prediction model can be as follows: Figure 2 As shown, it can include: local feature extraction networks, graph neural networks, feature fusion networks, and fully connected networks.

[0067] Among them, the local feature extraction network can be a deep temporal network with integrated attention mechanism (such as RNN, GRU, LSTM, etc.). Its core function is to perform temporal encoding on the transaction sequence features of a single customer and extract local temporal features that can characterize the individual transaction risk of the customer.

[0068] Graph neural networks are a core component of money laundering risk prediction models. Any one of GCN, GraphSAGE, or GAT can be selected. The core function is to extract global association features of customers based on the topology of the customer transaction relationship network and by aggregating the features of the node neighborhood.

[0069] Feature fusion network is a component of money laundering risk prediction model. Its core function is to use an attention mechanism to weightedly fuse local temporal features and global correlation features, eliminating the limitations of single features and forming a fused feature vector that can comprehensively characterize the money laundering risk of customers.

[0070] The fully connected network is the output link of the money laundering risk prediction model. Its core function is to perform money laundering risk quantification calculation on the fused feature vector, output the money laundering risk probability of each customer, and achieve accurate determination of the customer's money laundering risk.

[0071] Temporal feature encoding refers to the process by which a local feature extraction network processes the features of a customer's transaction sequence. By capturing the temporal dependencies and abnormal transaction features of the transaction sequence, it transforms unstructured temporal transaction data into standardized feature vectors, thereby achieving a quantitative representation of the temporal features of the transaction.

[0072] After obtaining the customer transaction relationship network, the transaction sequence features of individual customers can be extracted from the network. These features are then input into a local feature extraction network for temporal feature encoding to obtain the corresponding customer's local temporal features. The processing of transaction sequence features by the local feature extraction network will be described in detail in the following embodiments, and will not be repeated here.

[0073] Sub-step B2: Input the customer transaction relationship network into the graph neural network, and obtain the global association features of each customer through node neighborhood feature aggregation.

[0074] Node neighborhood feature aggregation is a core operation of graph neural networks. It refers to taking a single client node as the core, collecting feature information of its surrounding related nodes, and integrating the association features of the surrounding nodes and the core node through a specific algorithm to achieve an accurate characterization of the global association risk of the core node.

[0075] After obtaining the customer transaction relationship network, it can be input into a graph neural network. The graph neural network aggregates the node neighborhood features to obtain the global association features of each customer. That is, taking each customer node as the core, it samples the local transaction network corresponding to that customer from the customer transaction relationship network. Through the neighborhood aggregation algorithm of the graph neural network, it integrates the attribute information of the surrounding related nodes of the core node and the edge attribute information of the connecting edges, and mines the transaction association strength and risk transmission rules between the core node and the surrounding nodes, so as to achieve a comprehensive characterization of the global association risk of the core node. The implementation process will be described in detail in the following embodiments, and will not be repeated here.

[0076] Sub-step B3: The local temporal features and global correlation features are fused using the attention mechanism through the feature fusion network to obtain a fused feature vector.

[0077] Attention mechanism refers to the core algorithm used for feature fusion. By assigning different weights to local temporal features and globally related features, it highlights features that have a greater impact on the risk of money laundering by customers, thereby improving the targeting and effectiveness of the fused features.

[0078] Collect the local temporal feature vector of each customer output by sub-step B1, and the global associated feature vector of the corresponding customer output by sub-step B2, to ensure that the dimensions of the two feature vectors match and the data is synchronized, corresponding to the dual-view features of the same customer.

[0079] The local temporal feature vector and the global associated feature vector are simultaneously input into the feature fusion network, and the attention mechanism fusion algorithm is activated. This algorithm can automatically assign reasonable weights to the two feature vectors according to the degree of influence of the features on the money laundering risk.

[0080] By employing an attention mechanism, weighted calculations are performed on local temporal feature vectors and global correlation feature vectors. This approach emphasizes features that significantly impact customer money laundering risks (such as abnormal transaction features in local temporal features and high-risk correlation node features in global correlation features), while mitigating the influence of irrelevant features. This achieves deep fusion of dual-view features and eliminates the blind spots in evaluation based on a single feature dimension.

[0081] The final output is a fusion feature vector that integrates local temporal features and global correlation features. This vector combines the dynamic risk characteristics of individual customer transaction behavior with the correlation risk characteristics of the global transaction network, which can comprehensively and objectively characterize the customer's money laundering risk level and provide core input for subsequent money laundering risk prediction calculations.

[0082] Sub-step B4: Input the fused feature vector into the fully connected network to perform money laundering risk prediction calculation, and output the money laundering risk probability for each customer.

[0083] Money laundering risk prediction calculation refers to the process of processing fused feature vectors by a fully connected network. Through a preset risk assessment algorithm, the fused feature vectors are transformed into probability values ​​in the range of 0-1, which quantifies the possibility that a customer may be engaging in money laundering.

[0084] For each customer's fused feature vector, the vector is standardized and validated to ensure that the vector format is correct, the data is valid, there are no outliers, and the input requirements of a fully connected network are met.

[0085] The fused feature vector is input into the trained fully connected network, which contains two fully connected layers. It has been trained and optimized using fused feature data from historical risk customers and normal customers. The model parameters have reached the preset loss threshold, enabling accurate quantitative assessment of money laundering risks.

[0086] Through forward propagation calculations in a fully connected network, the money laundering risk of the fused feature vector is quantitatively analyzed. Combined with a pre-set risk assessment logic, the fused feature vector is transformed into a probability value in the range of 0-1. This probability value represents the likelihood that a customer is engaging in money laundering activities. The higher the probability value, the higher the risk of money laundering for the customer.

[0087] The final output is the money laundering risk probability for each customer, along with a detailed risk score (including the contribution of local time-series features and the contribution of global correlation features), providing accurate quantitative basis for subsequent screening of high-risk customers and generation of a risk customer list.

[0088] The process for data input and core seed customer identification can be as follows: Figure 5As shown, the overall process can include: S2.1 Data Input: Customer transaction relationship network G. S2.2 Local Temporal Feature Extraction: Acquiring single-customer time-series data, and simultaneously acquiring a trained deep temporal model. The acquired single-customer time-series data is processed by the deep temporal model to output local temporal features. S2.3 Global Association Feature Extraction: First, customer transaction network sampling is performed to obtain the local transaction network of a single customer. Then, a trained graph neural network model is acquired, and the local transaction network is processed by this model to output the global association features of each customer. S2.4 Core Seed Customer Mining: Based on the output money laundering risk probability of each customer, customers with a probability greater than a probability threshold are selected as core seed customers.

[0089] This application embodiment captures the time-series features of individual customer transactions and the global correlation features through a local feature extraction network and a graph neural network, respectively. It achieves deep fusion of the two features with the help of an attention mechanism, and then completes the risk prediction calculation through a fully connected network. This effectively solves the problems of insufficient dimensions and low accuracy of single feature evaluation. The output money laundering risk probability can objectively and comprehensively represent the level of money laundering risk of customers.

[0090] In one implementation of this application, the above sub-step B1 may include: sub-steps C1 to C3.

[0091] Sub-step C1: Obtain the characteristics of a single customer transaction sequence that match the business monitoring window; wherein, the business monitoring window is a pre-set time window.

[0092] In this embodiment, the business monitoring window refers to a fixed time window that is pre-set to carry out anti-money laundering risk monitoring and capture the complete transaction behavior cycle of customers. Its duration and start time can be flexibly adjusted according to the anti-money laundering business needs of financial institutions, regulatory requirements and historical transaction data patterns to ensure that it can fully cover the typical transaction cycle of customers.

[0093] The business monitoring window parameters have been set according to the anti-money laundering business needs of financial institutions (such as monthly monitoring and quarterly monitoring), regulatory requirements, and customer transaction behavior patterns. The start time, end time, and duration of the time window are clearly defined to ensure that the window duration can fully cover the typical transaction cycle of customers and avoid data loss due to a window that is too short or redundancy due to a window that is too long.

[0094] From the preprocessed customer transaction information, locate all transaction data corresponding to a single customer, filter out transaction data whose transaction time falls within the business monitoring window according to the time range of the business monitoring window, and remove historical transaction data outside the window to ensure that the filtered transaction data completely matches the business monitoring window.

[0095] The selected individual customer transaction data are sorted according to the order of transaction time and integrated to form a transaction sequence feature. This feature includes information such as transaction amount, transaction frequency, counterparty, and transaction time period at each point in time within the window, ensuring that the time sequence is accurate, the data is complete, and can fully reflect the dynamic transaction behavior of the customer within the business monitoring window.

[0096] Finally, the output is a single customer transaction sequence feature that matches the business monitoring window. This serves as the core input data for subsequent deep time series networks to extract time-dependent features, ensuring the temporal completeness and relevance of the input data.

[0097] Sub-step C2: Input the transaction sequence features into the deep temporal network to obtain the temporal hidden states at each time point in the transaction sequence features.

[0098] Temporal hidden states refer to the hidden feature vectors output by a deep temporal network after parsing the transaction sequence time by time. They can represent the transaction behavior features and temporal correlations of a single transaction time and are the basic processing objects for the weighted aggregation of attention mechanisms.

[0099] The transaction sequence features of a single customer are validated to confirm that the time sequence is correct, the data format is standardized, there are no missing values ​​or outliers, and the input requirements of deep time series networks are met, thus avoiding the impact of abnormal data on feature extraction results.

[0100] The verified transaction sequence features are input into the trained deep temporal network. Through the recurrent structure of the deep temporal network, the transaction sequence features are analyzed one by one in time sequence. For each transaction time point in the business monitoring window, the corresponding temporal hidden state is output, resulting in a set of temporal hidden states that is consistent with the number of transaction time points.

[0101] Finally, the set of hidden states at each time point corresponding to a single customer's transaction sequence is output. This set can fully reflect the customer's transaction behavior characteristics and time sequence evolution at different transaction time points, providing a basic processing object for subsequent attention mechanism weighted aggregation.

[0102] Sub-step C3: Calculate the weight of each of the temporal hidden states through the attention mechanism, and perform weighted aggregation on each of the temporal hidden states to obtain the local temporal features.

[0103] The attention mechanism is used to adaptively weight and aggregate the latent states of each transaction point in the output of the deep temporal network. By automatically learning and assigning weights to the latent states at different points in time, it strengthens the feature representation of key transaction points for money laundering risk identification and weakens the feature influence of ordinary transaction points, thereby achieving accurate aggregation of temporal features.

[0104] Obtain the set of time-series hidden states for each transaction point output by sub-step C2, and standardize each time-series hidden state to ensure that the feature dimensions of each hidden state are consistent and the numerical range is standardized, so as to meet the weighted aggregation requirements of the attention mechanism.

[0105] The attention mechanism integrated into the deep temporal network is invoked. This mechanism optimizes parameters through model training and can automatically calculate the weight corresponding to each temporal hidden state based on the impact of transaction behavior on money laundering risk at each point in time.

[0106] Through an attention mechanism, adaptive weights are assigned to the temporal hidden states corresponding to each transaction point. The weights of the hidden states at key risk points, such as abnormal transaction periods, large-amount split transactions, and frequent transactions with multiple unfamiliar accounts in a short period of time, are increased, while the weights of the hidden states at regular transaction points are decreased. Then, all the weighted temporal hidden states are aggregated and calculated to obtain a single standardized feature vector.

[0107] The final output is a local temporal feature vector after attention-weighted aggregation. This vector integrates the temporal features of the entire trading period and highlights the features of key risk points, which can accurately characterize the temporal anomalies and dynamic risks of a single client's trading behavior.

[0108] This application embodiment ensures the integrity and relevance of time-series data by matching business monitoring windows, captures the temporal patterns of transaction behavior by leveraging deep time-series networks, and strengthens the impact of high-risk features through an attention mechanism, effectively solving the problems of insufficient relevance and masking of key risk features in traditional time-series feature extraction.

[0109] In one implementation of this application, the above sub-step B2 may include: sub-steps D1 to D3.

[0110] Sub-step D1: Using a single customer as the target node, sample the local transaction network of the corresponding customer from the customer transaction relationship network.

[0111] In this embodiment, the target node refers to the network node corresponding to a single customer currently undergoing global correlation feature extraction. It is the core benchmark for local transaction network sampling. Each target node corresponds to a unique customer and is used to focus on the global correlation risk analysis of a single customer.

[0112] A local transaction network refers to a sub-network sampled from the complete customer transaction relationship network with the target node as the core. It includes the target node, surrounding nodes that have direct or indirect transactional relationships with the target node, as well as the connecting edges and corresponding attributes between these nodes, and is used to accurately mine the neighborhood association information of the target node.

[0113] Select a single customer from whom global correlation features need to be extracted, and use the network node corresponding to that customer as the target node. Define the unique identifier of the target node to ensure that the target node corresponds one-to-one with the single customer, and focus on the global correlation risk analysis of the single customer.

[0114] The neighborhood sampling strategy is adopted, with a preset sampling radius (i.e., the range of related nodes around the target node, which can be adjusted according to business needs) to ensure that the sampling range can cover the core related nodes of the target node, while avoiding the problem that the node representation is too smooth due to the sampling range being too large, making it impossible to effectively distinguish customer risks.

[0115] Based on a preset sampling radius, from the complete customer transaction relationship network, we select the surrounding nodes that have direct transaction associations (direct connection edges) and indirect transaction associations (connected through one or more intermediate nodes) with the target node. At the same time, we extract all connection edges between the target node and the surrounding nodes, as well as the node attributes of each node and the edge attributes of each connection edge.

[0116] Finally, the target node, the selected surrounding nodes, connecting edges, and corresponding attributes are integrated to form a local transaction network with the target node as the core. This ensures that the sub-network structure is complete and the data is accurate, fully reflecting the neighborhood association pattern of the target node, and serving as input data for subsequent graph neural network feature aggregation.

[0117] Sub-step D2: Input the local transaction network into the trained graph neural network to extract the neighborhood association features and network topology features of the target node.

[0118] Domain association features refer to the attribute features of surrounding nodes of the target node (such as customer basic information and transaction statistics of surrounding nodes), as well as the attribute features of the connection edges between the target node and surrounding nodes (such as transaction amount and number of transactions), reflecting the transaction association strength and risk transmission relationship between the target node and surrounding nodes.

[0119] Network topology features refer to the structural characteristics of a local transaction network, including the number of nodes, the number of connecting edges, the association paths between nodes, and the coreness of the target node in the local network. These features are used to characterize the position of the target node in the local network and the association pattern in the global transaction network.

[0120] The local transaction network output by sub-step D1 is verified to ensure that the nodes, connecting edges and corresponding attributes are complete and without missing parts, the network structure is standardized and meets the input requirements of graph neural networks, and to avoid abnormal data affecting the feature aggregation effect.

[0121] The verified local transaction network is input into the trained graph neural network. This network has been trained and optimized using network feature data from historical money laundering gang-related accounts and normal related accounts. The model parameters have reached the preset loss threshold and have the ability to accurately aggregate node neighborhood information and network topology features.

[0122] By using a neighborhood aggregation algorithm based on graph neural networks, with the target node as the core, the algorithm integrates the node attributes (customer basic information, transaction statistics, etc.) of surrounding nodes and the edge attributes (transaction amount, number of transactions, etc.) between the target node and surrounding nodes, and mines the transaction correlation strength and risk transmission pattern between the target node and surrounding nodes, thereby achieving accurate integration of neighborhood correlation features.

[0123] Simultaneously, a graph neural network is used to extract the topological features of the local transaction network, including the coreness of the target node in the local network, the length of the association path between nodes, and the density of the connecting edges. The topological features are then fused and aggregated with the neighborhood association information to comprehensively depict the global association pattern of the target node.

[0124] Sub-step D3: Perform an aggregation operation on the neighborhood association features and the network topology features to obtain the global association features of the target node.

[0125] After extracting the neighborhood association features and network topology features of the target node, an aggregation operation can be performed on these features. The aggregation results can then be collected, from which key features characterizing the global association risk of the target node can be extracted. These features include the association strength between the target node and high-risk nodes, the overall risk level of surrounding nodes, and the association influence of the target node in the transaction network. These key features are then quantified and encoded into standardized feature data. The quantified key features are integrated to generate an M-dimensional global association feature vector. This vector can completely and accurately characterize the association risk of the target node (a single customer) in the global transaction network, reflecting the transaction associations and risk transmission effects between the customer and other customers.

[0126] The system outputs a global correlation feature vector corresponding to a single customer, which serves as the core input for subsequent feature fusion networks and local temporal feature fusion, providing global correlation dimension support for a comprehensive quantitative assessment of customer money laundering risk.

[0127] This application embodiment leverages graph neural networks to aggregate neighborhood association features and network topology features, effectively mining the association risks and transmission patterns of customers in the global transaction network. It solves the problems of incomplete and insufficient targeting of traditional association feature extraction. The output global association features can accurately characterize the global association risks of customers, providing reliable support for dual-view feature fusion and accurate money laundering risk assessment, and improving the comprehensiveness of customer risk assessment.

[0128] In one implementation of this application, step 104 may include sub-step E1 and sub-step E2.

[0129] Sub-step E1: Remove customers from the whitelist from the risk customer list, and add customers who are not included in the risk customer list and belong to the gray list or blacklist as target customers to the risk customer list to obtain an updated risk customer list.

[0130] In this embodiment, the whitelist refers to a set of customers pre-defined by financial institutions that are confirmed to have no money laundering risk or extremely low risk. These customers are included in the whitelist after strict review and do not need to be the focus of money laundering gangs. It is used to remove low-risk interference items from the risk customer list.

[0131] The gray list refers to a set of customers set up by financial institutions that have potential money laundering risks but have not been clearly identified as high-risk. These customers have certain abnormal transaction characteristics and need to be included in the risk customer list for further analysis.

[0132] A blacklist refers to a set of clients established by financial institutions that are clearly at high risk of money laundering or are associated with money laundering-related illegal activities. These clients must be directly included in the risk client list and are the core targets for money laundering gangs to investigate.

[0133] Target customers refer to customers who are not included in the initial risk customer list but belong to the gray list or black list. These customers have a high money laundering risk and need to be added to the risk customer list to improve the coverage of high-risk customers.

[0134] The updated risk customer list refers to the risk customer list after processing in sub-step E1, which removes whitelist customers and adds target customers. Compared with the initial list, it has higher accuracy and completeness, and serves as the core seed node source for bottom-up identification of money laundering gangs.

[0135] The system retrieves a pre-set whitelist of customers from financial institutions, compares it with an initial list of high-risk customers, and filters out customers from the initial list who are on the whitelist. These customers are then removed from the high-risk list to prevent low-risk customers from interfering with subsequent money laundering investigations and to improve the accuracy of seed nodes.

[0136] The system retrieves the pre-set gray list and black list of customers from financial institutions, compares these two lists with the initial risk customer list, and filters out customers who are not included in the initial risk customer list but belong to the gray list or black list. These customers are then identified as target customers to ensure that no high-risk customers are missed.

[0137] The selected target customers are added in batches to the risk customer list after removing whitelisted customers. The list is then sorted and updated with customer identifiers and money laundering risk information for the target customers to form an updated risk customer list. This ensures that the list only contains high-risk customers, providing accurate seed nodes for subsequent bottom-up mining.

[0138] The updated list of high-risk customers is verified to ensure that there are no whitelisted customers remaining, no target customers are omitted, and that customer information is complete and accurate, thus ensuring the accuracy and completeness of the list and meeting the input requirements for subsequent gang identification.

[0139] Sub-step E2: Using a bottom-up identification method, based on the updated list of risky customers and the customer transaction relationship network, identify suspicious money laundering groups.

[0140] After obtaining the updated list of high-risk customers, a bottom-up identification method can be used to identify suspicious money laundering groups based on the updated list and customer transaction relationship network. Specifically, the customer with the highest risk probability in the list can be used as a seed node, and then customers associated with the seed node and belonging to the updated list of high-risk customers can be identified from the customer transaction relationship network to obtain suspicious money laundering groups. This implementation process will be described in detail in the following embodiments, and will not be repeated here.

[0141] This application embodiment avoids interference from low-risk customers in gang detection by removing low-risk customers from the whitelist, while supplementing target customers in the gray list and blacklist to fill the gap of high-risk customers missing from the initial risk customer list. This ensures that the updated risk customer list focuses only on high-risk customers, providing accurate core seed nodes for subsequent gang detection and laying the foundation for accurate identification.

[0142] In one implementation of this application, the above sub-step E2 may include: sub-steps F1 to F5.

[0143] Sub-step F1: Select the customer node with the highest money laundering risk from the updated list of risky customers as the starting point for mining.

[0144] In this embodiment, the starting point for data mining refers to the customer node with the highest money laundering risk selected from the updated list of high-risk customers. This node serves as the core starting point for identifying money laundering groups from the bottom up and forms the basis for expanding the group's associations.

[0145] For the updated list of high-risk customers, the money laundering risk probability and risk score details for each customer node in the list can be extracted. All customer nodes can be sorted in descending order of money laundering risk probability from high to low, thus clarifying the risk priority of each node.

[0146] From the sorted list of customer nodes, select the top-ranked customer node (with the highest probability of money laundering risk) as the starting point for mining. Record the unique customer identifier, money laundering risk probability, and node attributes of this starting point node to ensure that the starting point node is a high-risk core node in the list, providing an accurate starting point for subsequent gang mining.

[0147] The selected starting point node is marked to avoid repeated selection in subsequent iterations, ensuring the orderly progress of bottom-up gang identification.

[0148] Sub-step F2: In the customer transaction relationship network, a priority search strategy algorithm is used to mine the remaining core seed customer nodes that have transaction associations with the mining starting point and belong to the updated risk customer list through connected subgraph extraction.

[0149] Priority search strategy algorithms refer to core algorithms used to discover and identify transaction-related nodes at the starting point in a customer transaction relationship network. These include, but are not limited to, two strategies: breadth-first search (BFS) and depth-first search (DFS). These strategies can be flexibly selected according to the anti-money laundering business needs of financial institutions to ensure the efficiency and comprehensiveness of node discovery.

[0150] A connected subgraph is a basic structure in graph theory. It refers to a subgraph consisting of several nodes and connecting edges within a directed or undirected graph, where there is a reachable path between any two nodes and no isolated nodes or branches. In financial risk control and gang detection, it often corresponds to a set of accounts with close relationships, interconnected funds, and coordinated behavior, and is an important structural unit for identifying money laundering gangs.

[0151] Based on the anti-money laundering business needs of financial institutions, breadth-first search (BFS) or depth-first search (DFS) is selected as the preferred search strategy, and the search scope (the neighborhood association range with the starting node as the core) and the search termination condition (traversing to the association boundary of all core seed nodes) are clearly defined.

[0152] Taking the starting node as the core, a connected subgraph containing the starting node is extracted from the customer transaction relationship network. All nodes in the subgraph that have direct or indirect transactional relationships with the starting node are filtered out. At the same time, the connecting edges, node attributes, and edge attributes of these nodes are also filtered out to ensure that the subgraph covers all the core relationships of the starting node.

[0153] From all associated nodes in the connected subgraph, select nodes that belong to the updated risk customer list and identify them as the remaining core seed customer nodes. Remove irrelevant nodes from the non-risk customer list to avoid redundant node interference, and retain only high-risk associated nodes related to gang mining.

[0154] The final output is a list of all core seed customer nodes that have transactional relationships with the starting node, providing node data support for the construction of candidate groups.

[0155] Sub-step F3: Combine the mining starting point with the core seed customer nodes obtained from the mining to form a candidate money laundering gang.

[0156] Candidate money laundering groups refer to a set of customer nodes with potential money laundering risks, formed by the starting point of the data mining and the core seed customer nodes obtained through mining. They need to be confirmed as suspicious money laundering groups after subsequent sorting and screening.

[0157] The starting node is integrated with the core seed customer nodes obtained through mining to form a complete set of customer nodes. A unique gang identifier is assigned to this set, and the customer identifiers, money laundering risk probabilities, risk score details, and transaction relationships between nodes are recorded for all nodes in the set.

[0158] Initialize relevant statistical data for candidate money laundering groups, including group size (number of nodes), average warning probability (average risk probability of nodes), and total transaction amount (cumulative transaction amount between nodes), to provide basic data for subsequent group ranking. Sub-step F4: Remove customer nodes that have been included in the candidate group from the updated risk customer list, and repeat the step of selecting the customer node with the highest money laundering risk from the updated risk customer list as the mining starting point, until the mining starting point is combined with the core seed customer node obtained by mining to form a candidate money laundering group, until no new candidate money laundering group can be generated.

[0159] Obtain the core seed customer node list of the currently generated candidate money laundering gangs, remove these nodes that have been included in the gangs from the updated risk customer list, update the risk customer list, and ensure that subsequent iterations only target high-risk nodes that have not been discovered.

[0160] Check if the updated list of risky customers is empty, or if there are still high-risk customer nodes that can be used as a starting point for mining: If the list is not empty and there are high-risk nodes that have not been mined, repeat the process of sub-steps F1-F3, select the node with the highest risk in the current list as the new starting point for mining, mine related core seed nodes and build new candidate money laundering gangs; if the list is empty or there are no new high-risk nodes to mine, stop the iteration process and complete the construction of all candidate money laundering gangs.

[0161] All candidate money laundering groups generated during the iteration process are aggregated to form a complete set of candidate money laundering groups, ensuring that no groups are omitted or duplicated.

[0162] Sub-step F5: Sort the candidate money laundering groups according to their size, average warning probability, and total transaction amount, and determine the suspicious money laundering groups based on the sorting results.

[0163] The average warning probability refers to the average money laundering risk probability of all core seed customer nodes within a candidate money laundering gang. It is used to quantitatively characterize the overall money laundering risk level of the gang and is one of the core dimensions for gang ranking.

[0164] Total transaction amount refers to the cumulative transaction amount between all core seed customer nodes within the candidate money laundering group and with external nodes. It is used to characterize the group's transaction scale and capital size, and to help assess the severity of the group's money laundering risk.

[0165] A multi-dimensional priority ranking strategy is adopted to rank all candidate money laundering groups. The ranking priorities are as follows: average warning probability of candidate groups (highest priority): the higher the probability, the higher the overall money laundering risk of the group; total transaction amount (secondary priority): the higher the amount, the larger the group's funds, and the stronger the harm of money laundering; group size (supplementary priority): the larger the size, the more group members, and the more obvious the coordinated money laundering behavior.

[0166] Based on the ranking results, the top-ranked candidate money laundering groups are selected as suspected money laundering groups. The screening criteria can be dynamically adjusted by financial institutions according to anti-money laundering regulatory requirements and actual business operations (such as selecting the top N% of groups, or selecting groups with an average warning probability ≥ the threshold and whose size / amount meets the requirements).

[0167] Finally, the system can output the relevant information of the identified suspicious money laundering groups, including the group's unique identifier, a list of core seed customer nodes (including identity identifiers, account information, and risk scores), group transaction details, risk score details, and a network topology diagram of the group's transaction relationships, thus completing the identification and output of information about suspicious money laundering groups.

[0168] Understandably, in the above scheme, since the updated list of risky customers has added customers from the blacklist and / or graylist, when identifying money laundering groups, it is necessary to check whether the newly added customers in the list have a transaction relationship network with other customers. If so, they are marked (there will inevitably be isolated newly added customers in the list), and then money laundering group identification is performed. Isolated newly added customers in the list are not processed.

[0169] This application uses the highest-risk node as the starting point for data mining, identifying the core high-risk nodes of the group from the source. This avoids the drawbacks of traditional top-down methods that break down the group's topology, laying the foundation for the complete construction of the group. By iteratively eliminating mined nodes and repeatedly expanding the group, it achieves full coverage of all high-risk nodes in the updated risk customer list, avoiding the omission of covert and decentralized money laundering groups.

[0170] Next, combined Figure 3 The four modules shown describe the process of money laundering gang detection as follows. The bottom-up money laundering gang detection technology proposed in this application mainly includes the following four modules: M1 Transaction Network Construction Module, M2 Dual-View Customer Money Laundering Risk Assessment Module, M3 Bottom-Up Money Laundering Gang Identification Module, and M4 Visualization Display Module. Each module works collaboratively and performs its specific function. Through standardized implementation steps, it achieves accurate identification of money laundering gangs, quantitative risk assessment, and visual presentation. The overall process is as follows: Figure 3 As shown.

[0171] The M1 transaction network construction module serves as the basic data support unit of this application. Its core technical purpose is to build a standardized customer transaction relationship network, providing a high-quality customer transaction network for subsequent functional modules, used for money laundering gang detection and visualization. The specific implementation steps are as follows: (1) Data collection: Collect multi-source heterogeneous data such as customer basic information and transaction flow through various internal systems of the bank to ensure that the collected data fully covers all the related data required for money laundering gang detection, providing a data foundation for subsequent data processing and analysis. (2) Data preprocessing: Perform preprocessing operations on the collected raw data to filter out invalid, duplicate and abnormal data, ensuring the accuracy of subsequent module calculations. (3) Network construction: With customers as network nodes and transaction relationships between customers as edges, construct a homogeneous transaction network to realize the structured representation of customer transaction relationships, providing reliable data support for customer money laundering risk assessment in the M2 module and money laundering gang identification in the M3 module.

[0172] The M2 dual-view customer money laundering risk assessment module is the core innovative module of this application. Its core technical purpose is to achieve accurate quantitative rating of the money laundering risk of a single customer in the transaction network and to screen high-risk customers as seed nodes for money laundering gangs. The specific implementation steps are as follows: (1) Data input: Based on the isomorphic transaction network data output by the M1 module, extract the local transaction data and global correlation data of a single customer, respectively, as the basic data for dual-view risk assessment, to ensure the comprehensiveness and accuracy of the assessment data. (2) Local time series feature extraction: Use a deep time series model to perform in-depth analysis on the transaction flow data of a single customer, capture the dynamic evolution law of customer transactions, mine the hidden abnormal transaction behavior of customers, and output the local time series features of a single customer. (3) Global correlation feature extraction: Use graph neural network technology to extract the correlation features of customers in the global transaction network, analyze the risk transmission effect between customers and other accounts, and output the global correlation features of a single customer. (4) Core seed customer mining: The local time-series features and global correlation features obtained in steps (2) and (3) are fused together to quantitatively predict the final money laundering risk score of each customer. High-scoring customers will be used as core seed nodes for money laundering gang mining from the bottom up in the M3 module.

[0173] The M3 bottom-up money laundering gang identification module is the core functional module of this application. Its core technical purpose is to break through the technical limitations of the existing top-down money laundering gang mining mode and realize the accurate identification and positioning of money laundering gangs. The specific implementation steps are as follows: (1) Seed node determination: Receive the high-risk customer list output by the M2 module and use it as the core seed node for bottom-up money laundering gang mining, providing the original node for subsequent association expansion. (2) Association expansion: Based on the isomorphic transaction network constructed by the M1 module, mine the transaction association between various sub-nodes, and gradually merge and expand the associated nodes through the connected subgraph extraction technology to form candidate money laundering gangs. (3) Candidate money laundering gang ranking and output: According to the candidate gang size, total transaction amount and other dimensions, the candidate money laundering gangs are risk-scored and ranked, and the information of high-risk money laundering gangs is output.

[0174] The M4 visualization module, as the output presentation unit of this application, aims to present the results of money laundering gang investigations in a visual form, providing convenient and efficient analysis support for anti-money laundering monitoring personnel. The specific implementation steps are as follows: (1) Data reception: Receive the output data from the three modules M1, M2, and M3 to ensure that the data presented in the visualization is comprehensive and accurate. (2) Visualization result presentation: Based on the received data, construct a multi-dimensional visualization interactive interface to present the results of money laundering gang investigations in an intuitive and clear visual form, providing accurate and efficient analysis support for anti-money laundering monitoring personnel, and significantly improving the intelligence level and analysis efficiency of anti-money laundering monitoring work.

[0175] Next, the training process of the money laundering risk prediction model will be described in detail. The model training process may include: Step 1: Construction and preprocessing of training base data.

[0176] 1. Collect historical transaction data from financial institutions (including full transaction information of confirmed money laundering customers, normal customers, and basic customer information); 2. Perform data preprocessing: process the raw transaction data for missing values, remove redundancy, and standardize the coding to obtain standardized customer transaction information; 3. Extract transaction sequence features of all customers according to the business monitoring window to ensure that the training data and the actual inference data are in the same format.

[0177] Step 2: Training label construction and dataset partitioning.

[0178] 1. Labeling: Using financial institutions' regulatory labels and verified money laundering clients as positive samples (label=1), and normal compliant clients as negative samples (label=0), a true risk label for the samples is constructed; 2. Dataset partitioning: All preprocessed samples are divided into training set (70%), validation set (20%), and test set (10%) according to the proportions to ensure a balanced distribution of positive and negative samples; 3. Remove risk-free customers from the whitelist and retain customers from the graylist and blacklist as core training samples.

[0179] Step 3: Training customer transaction relationship network construction.

[0180] 1. Based on the customer data in the training set, configure node attributes using customer identifiers as network nodes; 2. Establish connection edges based on customer transaction relationships, configure edge attributes, and construct a dedicated customer transaction relationship network for training; 3. This network provides the topology, node attributes, and edge attribute data required for training graph neural networks.

[0181] Step 4: Model initialization.

[0182] 1. Initialize all network modules of the money laundering risk prediction model: 2. Initialize the local feature extraction network: Set the parameters of the deep temporal network (LSTM / GRU) and the parameters of the temporal attention layer; 3. Initialize the graph neural network: Set the parameters for the graph convolutional layer and the neighborhood feature aggregation algorithm; 4. Initialize the feature fusion network: Set the parameters for the cross-feature attention mechanism; 5. Initialize the fully connected prediction network: Set the parameters for the fully connected layers and the activation function (Sigmoid); 6. Configure training hyperparameters such as optimizer (Adam / AdamW), learning rate, and batch size.

[0183] Step 5: Model forward propagation computation (core training inference).

[0184] 1. Input the training set data into the model, and perform feature extraction and risk prediction in separate branches: 2. Local Temporal Feature Extraction: Input customer transaction sequence features into a deep temporal network to obtain the temporal hidden state at each transaction time point; adaptively calculate the weight of the hidden state at each time point through a temporal attention mechanism, and output the local temporal features after weighted aggregation; 3. Global Relationship Feature Extraction: Taking customer nodes as targets, sample the local transaction network; input the graph neural network to complete the aggregation of neighborhood relationship information and network topology features, and output global relationship features; 4. Feature Fusion: Local temporal features and global correlation features are input into the feature fusion network. Weights are assigned through an attention mechanism to complete the weighted fusion of the two features and obtain the fused feature vector. 5. Risk Prediction: Input the fused feature vector into the fully connected network to calculate and output the predicted money laundering risk probability of the customer.

[0185] Step 6: Calculate the loss function.

[0186] 1. The binary cross-entropy loss function is used as the target loss function for model training; 2. Compare the predicted money laundering risk probability output by the model with the actual risk labels constructed in step 2, and calculate the overall predicted loss; 3. The loss value represents the error between the model's prediction and the true label. The larger the error, the greater the space for optimizing the model parameters.

[0187] Step 7: Model validation and performance evaluation.

[0188] 1. After each round of training, the validation set data is input into the model to calculate core anti-money laundering evaluation metrics such as precision, recall, and AUC; 2. Monitor the changes in validation set loss to determine if the model is overfitting / underfitting; 3. Retain the temporary model parameters that yield the best performance on the validation set.

[0189] Step 8: Iterative training and hyperparameter tuning.

[0190] Repeat steps 4-7 to complete multiple rounds of iterative training; If the model performance does not meet the requirements, adjust hyperparameters such as learning rate, number of network layers, and sampling radius, and retrain. Continue optimizing until the model's performance on the validation set stabilizes.

[0191] Step 9: Model convergence judgment and final solidification.

[0192] The model is considered converged when the validation set loss does not decrease for multiple consecutive rounds and the risk prediction accuracy reaches the preset business threshold. Save the optimal model parameters after convergence to complete the training of the money laundering risk prediction model; The solidified model can be directly used in actual business operations for predicting customer money laundering risks and identifying suspicious money laundering groups.

[0193] Reference Figure 6 The diagram illustrates the structure of a device for identifying a suspected money laundering gang, as provided in an embodiment of this application. Figure 6 As shown, the identification device 600 for the suspected money laundering group may include the following modules: Network building module 610 is used to build a customer transaction relationship network based on the customer identifiers and customer transaction information of all customers within the enterprise; The probability acquisition module 620 is used to extract local temporal features and global correlation features from the customer transaction relationship network through the money laundering risk prediction model, and process the local temporal features and global correlation features to obtain the money laundering risk probability of each customer. The list acquisition module 630 is used to filter out all customers whose money laundering risk probability is greater than the probability threshold, and obtain a list of risky customers. The gang identification module 640 is used to identify suspicious money laundering gangs using a bottom-up identification method based on the risk customer list and the customer transaction relationship network.

[0194] Optionally, the network building module includes: The transaction information acquisition unit is used to acquire the customer identifiers and initial transaction data of all customers, and to perform preprocessing on the initial transaction data, including missing value processing, redundancy removal, and standardized coding, to obtain the preprocessed customer transaction information. A node attribute determination unit is used to determine the node attributes of each network node based on the customer identifier of each customer as a network node and the first and second features in the customer transaction information; wherein the first feature includes at least: customer basic information, account statistics information and transaction statistics information, and the second feature includes: transaction sequence features; A network construction unit is used to establish connection edges between nodes in the network nodes that have transaction-related behaviors, and to determine the edge attributes of the connection edges based on the third feature in the customer transaction information to obtain the customer transaction relationship network; wherein, the third feature includes at least: transaction amount information, number of transactions information, and transaction cycle information.

[0195] Optionally, the money laundering risk prediction model includes: a local feature extraction network, a graph neural network, a feature fusion network, and a fully connected network. The probability acquisition module includes: The local feature acquisition unit is used to input the transaction sequence features of a single customer into the local feature extraction network for temporal feature encoding to obtain the local temporal features of the corresponding customer. The global feature acquisition unit is used to input the customer transaction relationship network into the graph neural network and obtain the global association features of each customer through node neighborhood feature aggregation. The feature fusion acquisition unit is used to fuse the local temporal features and the global correlation features through the feature fusion network using an attention mechanism to obtain a fused feature vector. The probability output unit is used to input the fused feature vector into the fully connected network to perform money laundering risk prediction calculation and output the money laundering risk probability for each customer.

[0196] Optionally, the local feature extraction network is a deep temporal network integrating an attention mechanism. The local feature acquisition unit includes: The sequence feature acquisition subunit is used to acquire the sequence features of a single customer transaction that match the business monitoring window; wherein, the business monitoring window is a pre-set time window; A feature extraction subunit is used to input the transaction sequence features into the deep temporal network to obtain the temporal hidden states corresponding to each time point in the transaction sequence features. The local feature output subunit is used to calculate the weight of each of the temporal hidden states through the attention mechanism, and to perform weighted aggregation on each of the temporal hidden states to obtain the local temporal features.

[0197] Optionally, the global feature acquisition unit includes: The local network acquisition subunit is used to sample the local transaction network of a corresponding customer from the customer transaction relationship network, with a single customer as the target node; The feature aggregation subunit is used to input the local transaction network into the trained graph neural network and extract the neighborhood association features and network topology features of the target node; A global feature generation subunit is used to perform an aggregation operation on the neighborhood association features and the network topology features to obtain the global association features of the target node.

[0198] Optionally, the gang identification module includes: An Yuan, which updates the list, is used to remove customers from the whitelist from the risk customer list and add customers who are not included in the risk customer list and belong to the gray list or blacklist as target customers to the risk customer list, thus obtaining an updated risk customer list. The gang identification unit is used to identify suspicious money laundering gangs using a bottom-up identification method based on the updated list of risky customers and the customer transaction relationship network.

[0199] Optionally, the gang identification unit includes: The starting point acquisition subunit is used to select the customer node with the highest money laundering risk from the updated list of risky customers as the mining starting point. Seed node acquisition subunit is used to mine, in the customer transaction relationship network, other core seed customer nodes that have transaction association with the mining starting point and belong to the updated risk customer list by using a priority search strategy algorithm and a connected subgraph extraction method; The candidate group constitutes a sub-unit, which is used to combine the mining starting point with the core seed customer nodes obtained from the mining to form a candidate money laundering group. Remove customer nodes that have been included in the candidate group from the updated list of risky customers, and repeat the process of obtaining sub-units from the starting point, obtaining sub-units from the seed node, and forming sub-units from the candidate group until no new candidate money laundering groups can be generated. The gang identification subunit is used to sort the candidate money laundering gangs according to their size, average warning probability, and total transaction amount, and to identify the suspicious money laundering gangs based on the sorting results.

[0200] This application's embodiments construct a customer transaction relationship network to fully capture the transaction connections and fund flow characteristics among customers, providing comprehensive data support for the identification of money laundering gangs. By leveraging a money laundering risk prediction model to simultaneously extract local temporal features and global correlation features, it achieves accurate quantification of the money laundering risk probability for each customer, effectively screening high-risk customers to form a risk customer list, avoiding interference from low-risk customers and the omission of high-risk customers. Furthermore, a bottom-up identification method is employed, starting with high-risk customers, and forming candidate gangs by mining direct or indirect transaction connections between nodes. This method can fully preserve the topological structure of highly concealed, hierarchical, and decentralized money laundering gangs, avoiding the identification bias caused by the fragmentation of gangs in traditional identification methods. This comprehensively improves the identification accuracy of such money laundering gangs and reduces the false positive rate during the identification process.

[0201] This application also provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the above-mentioned method for identifying suspicious money laundering groups.

[0202] Figure 7 A schematic diagram of the structure of an electronic device 700 according to an embodiment of the present invention is shown. Figure 7 As shown, the electronic device 700 includes a central processing unit (CPU) 701, which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (ROM) 702 or loaded from storage unit 708 into random access memory (RAM) 703. The RAM 703 can also store various programs and data required for the operation of the electronic device 700. The CPU 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0203] Multiple components in electronic device 700 are connected to I / O interface 705, including: input unit 706, such as keyboard, mouse, microphone, etc.; output unit 707, such as various types of displays, speakers, etc.; storage unit 708, such as disk, optical disk, etc.; and communication unit 709, such as network card, modem, wireless transceiver, etc. Communication unit 709 allows electronic device 700 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0204] The various processes and handling described above can be executed by processing unit 701. For example, the methods of any of the above embodiments can be implemented as computer software programs, which are tangibly contained in a computer-readable medium, such as storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 700 via ROM 702 and / or communication unit 709. When the computer program is loaded into RAM 703 and executed by CPU 701, one or more actions of the methods described above can be performed.

[0205] Additionally, embodiments of this application also provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned method for identifying suspicious money laundering groups.

[0206] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal that includes said element.

Claims

1. A method for identifying suspicious money laundering groups, characterized in that, include: A customer transaction relationship network is constructed based on the customer identifiers and customer transaction information of all customers within the enterprise. The money laundering risk prediction model extracts local temporal features and global correlation features from the customer transaction relationship network, and processes the local temporal features and global correlation features to obtain the money laundering risk probability of each customer. Filter out all customers whose money laundering risk probability is greater than the probability threshold to obtain a list of risky customers; A bottom-up identification method is used to identify suspicious money laundering groups based on the risk customer list and the customer transaction relationship network.

2. The method according to claim 1, characterized in that, The process of constructing a customer transaction relationship network based on the customer identifiers and transaction information of all customers within the enterprise includes: Obtain customer identifiers and initial transaction data for all customers, and perform preprocessing on the initial transaction data, including missing value handling, redundancy removal, and standardized coding, to obtain the preprocessed customer transaction information; Each customer's customer identifier is used as a network node, and the node attributes of each network node are determined according to the first feature and the second feature in the customer transaction information; wherein, the first feature includes at least: customer basic information, account statistics information and transaction statistics information, and the second feature includes: transaction sequence features; Establish connection edges between nodes in the network that have transaction-related behaviors, and determine the edge attributes of the connection edges based on the third feature in the customer transaction information to obtain the customer transaction relationship network; wherein, the third feature includes at least: transaction amount information, number of transactions information, and transaction cycle information.

3. The method according to claim 1, characterized in that, The money laundering risk prediction model includes: a local feature extraction network, a graph neural network, a feature fusion network, and a fully connected network. The method involves extracting local temporal features and global correlation features from the customer transaction relationship network using a money laundering risk prediction model, and processing these features to obtain the money laundering risk probability for each customer, including: The transaction sequence features of a single customer are input into the local feature extraction network for temporal feature encoding to obtain the local temporal features of the corresponding customer. The customer transaction relationship network is input into the graph neural network, and the global association features of each customer are obtained by aggregating the node neighborhood features. The feature fusion network uses an attention mechanism to fuse the local temporal features and the global related features to obtain a fused feature vector. The fused feature vector is input into the fully connected network to calculate money laundering risk and output the money laundering risk probability for each customer.

4. The method according to claim 3, characterized in that, The local feature extraction network is a deep temporal network integrating an attention mechanism. The step of inputting the transaction sequence features of a single customer into the local feature extraction network for temporal feature encoding to obtain the local temporal features of the corresponding customer includes: Acquire the characteristics of a single customer transaction sequence that match a business monitoring window; wherein, the business monitoring window is a pre-set time window; The transaction sequence features are input into the deep temporal network to obtain the temporal hidden states corresponding to each time point in the transaction sequence features. The weights of each temporal hidden state are calculated using the attention mechanism, and the weighted aggregation of each temporal hidden state is performed to obtain the local temporal features.

5. The method according to claim 3, characterized in that, The step of inputting the customer transaction relationship network into the graph neural network and obtaining the global association features of each customer through node neighborhood feature aggregation includes: Using a single customer as the target node, a local transaction network for the corresponding customer is obtained by sampling from the customer transaction relationship network; The local transaction network is input into the trained graph neural network to extract the neighborhood association features and network topology features of the target node; An aggregation operation is performed on the neighborhood association features and the network topology features to obtain the global association features of the target node.

6. The method according to claim 1, characterized in that, The bottom-up identification method, based on the risk customer list and the customer transaction relationship network, identifies suspicious money laundering groups, including: Remove customers from the whitelist from the risk customer list, and add customers who are not included in the risk customer list and belong to the gray list or blacklist as target customers to the risk customer list to obtain an updated risk customer list; A bottom-up identification method is used to identify suspicious money laundering groups based on the updated list of risky customers and the customer transaction relationship network.

7. The method according to claim 6, characterized in that, The bottom-up identification method, based on the updated list of high-risk customers and the customer transaction relationship network, identifies suspicious money laundering groups, including: Select the customer node with the highest money laundering risk from the updated list of high-risk customers as the starting point for the analysis; In the customer transaction relationship network, a priority search strategy algorithm is used to mine the remaining core seed customer nodes that have transaction associations with the mining starting point and belong to the updated risk customer list through connected subgraph extraction. The starting point of the mining is combined with the core seed customer nodes obtained from the mining to form a candidate money laundering gang; Remove customer nodes that have been included in the candidate group from the updated risk customer list, and repeat the step of selecting the customer node with the highest money laundering risk from the updated risk customer list as the starting point for mining, until the starting point for mining is combined with the core seed customer nodes mined to form a candidate money laundering group, until no new candidate money laundering group can be generated. The candidate money laundering groups are ranked according to their size, average warning probability, and total transaction amount, and the suspicious money laundering groups are determined based on the ranking results.

8. A device for identifying suspicious money laundering groups, characterized in that, include: The network construction module is used to build a customer transaction relationship network based on the customer identifiers and customer transaction information of all customers within the enterprise; The probability acquisition module is used to extract local temporal features and global correlation features from the customer transaction relationship network through the money laundering risk prediction model, and to process the local temporal features and global correlation features to obtain the money laundering risk probability of each customer. The list retrieval module is used to filter out all customers whose money laundering risk probability is greater than the probability threshold, and obtain a list of risky customers. The gang identification module is used to identify suspicious money laundering gangs using a bottom-up identification method based on the risk customer list and the customer transaction relationship network.

9. An electronic device, characterized in that, include: A processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method for identifying a suspicious money laundering group as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device is able to perform the identification method for a suspected money laundering group as described in any one of claims 1 to 7.